Tributary ↗ · Validation Study · Trust Report v6 · 25 Jul 2026
A validation study of the Tributary forensic engine — known error rates, measured against labeled ground truth, a real 324-page bank statement, and the DOJ's 1MDB complaint.
Why a validation study
Expert analysis offered in federal court is measured against Daubert and Rule 702 today — which ask, among other things, about a technique's known or potential error rate and whether it rests on a documented, reliably applied methodology. That standard is current law, not a forecast.
Where machine-generated evidence fits is still being settled. Proposed Federal Rule of Evidence 707 would hold AI output offered without a sponsoring expert to those same Rule 702 requirements. It was published for public comment in August 2025; comment closed in February 2026; and at its May 2026 meeting the Advisory Committee on Evidence Rules declined to recommend adoption, instead revising the proposal and carrying it forward for further study alongside the deepfakes question. Even on the original schedule it could not have taken effect before December 2027. It is not law, and its timeline has slipped.
Most tools in this market claim "99%+ accuracy" with no methodology, no ground truth, and no error definition. This report is the opposite: every number below comes from a repeatable evaluation with an answer key written before the engine ran — the validation-study record that admissibility will ask for, kept current as the engine changes.
Context: ACFE 2026 benchmarking — 75% of fraud examiners say AI accuracy matters · only 18% test their models
Methodology · stated up front
Famous-case validation · new in v4
The engine analyzed a 45,000-character section of the Department of Justice's 2016 forfeiture complaint from the 1MDB case — the "Good Star phase," in which more than $1 billion left a Malaysian sovereign fund. Because the case is public and extensively adjudicated, every claim the engine makes can be checked against established record.
Every headline fact of the phase was recovered: the $1 billion transfer split into $700M to Good Star and $300M to the joint venture; the $330M follow-on transfers; and the $1.03B total — which the source never states as one number, and which the engine correctly reported as its own computation, labeled derived, with both components shown. The central deception of the case — a 1MDB officer telling a bank that "Good Star is owned 100% by PetroSaudi" when its sole signatory was Jho Low — was caught as a contradiction.
The finding · a hole in the DOJ's own complaint
The complaint states that "more than $85 million" flowed from the Good Star account into personal spending, then itemizes only part of it — $12M to Caesars Palace, $13.4M to Las Vegas Sands, $11M to an associate, and so on: $53.75M in named payments. The engine summed the itemized components, noticed they fall short of the stated total, and emitted:
Nothing in the prompt asked for this. The unaccounted remainder wasn't smoothed over or silently absorbed into a total — it became a line item that says go find out where this went. That is the product thesis, demonstrated on the most scrutinized financial document of the decade.
A new measured property · containment
The DOJ complaint anonymizes its subjects. The engine — like any modern AI model — knows from years of public reporting who they are. That creates a measurable risk: does outside knowledge leak into the analysis as if it were in the document? Scored across the full output: zero real-world identifications appeared in the entities, money flows, timeline, or contradictions. Names from public reporting surfaced only in the investigative-leads section, each explicitly attributed — "widely reported in open sources as…" — where a suggestion is supposed to live.
One honest miss inside the clearly-fenced leads section: the engine described an executive as serving during "the relevant period" who in fact joined the company years after it. Contained, attributed, and speculative by design — but factually wrong, and reported here because reporting it is what this document is for.
New benchmarks since v3 · coverage & tables
Recall tells you what was found of what you checked; completeness asks the harder question — of every dollar amount in the document, how many did the engine surface at all? And financial tables add a distinct failure mode: pulling a figure from the wrong column (requested vs. approved vs. disbursed) produces a number that is in the document but means the wrong thing.
In the column-trap case the engine also reported the sums of the other columns — but labeled them as its own derived totals with full breakdowns, rather than presenting them as stated figures. The first automated check scored those as errors; they weren't. Third time this evaluation has had to learn the same lesson: the metric must credit labeled derivation, or the score punishes honesty.
Real-case validation
The engine was run against a real Electronic Funds Transfer Analysis — 324 pages of OCR'd business bank statements. On the tested statements, every dollar figure it reported was scored against the source text:
The finding · unprompted, on real data
From the raw statements, with no prompting, the engine surfaced:
Checked against the source, every part holds:
It also caught a reconciliation gap: "$75.15 of the stated $488.75 in April electronic payments is unaccounted for." That is the entire product in one example — it doesn't just read the numbers, it reasons about them and tells you where to look.
Resolved in v5 · from caveat to result
Earlier versions of this report carried a caveat: naive exact-match scoring read entity grounding at 58%, because the engine fixes what it reads — expanding acronyms (IRS → Internal Revenue Service), repairing OCR damage ("MERCIAL" → "340 Commercial Street") — and a literal string comparison scores those corrections as inventions.
v5 replaces that metric with a three-tier score in which every credit is auditable: verbatim (name appears in the source as-is), normalized (a credited transformation — token overlap, acronym ↔ expansion, OCR-fragment match — with the matching evidence recorded per entity), and ungrounded. Crediting never hides: the tiers are always reported separately.
The single ungrounded entity is itself worth reading: "LLC (unnamed)" — the document's LLC name is redacted, and the engine described it as unnamed rather than guessing. The one measured entity error is the engine declining to invent a name. It is counted as an error anyway, because a metric that starts making exceptions stops being a metric.
Controlled benchmark · synthetic first pass
Three synthetic-but-realistic forensic documents — a wire-transfer record, a deposition with one deliberately planted contradiction, and a corporate-ownership filing describing an interconnected shell-company structure. Every correct entity, figure, date, and contradiction was labeled by hand before the engine saw the text, so each output can be scored against ground truth rather than opinion.
Headline results · aggregate
✓ Planted contradiction caught
Per-document
| Document | Entity recall | Fig. grounding | Key figs | Contra. |
|---|---|---|---|---|
| D1 · wire records | 100% | 75% | 100% | — |
| D2 · deposition | 75% | 100% | 100% | caught |
| D3 · ownership | 100% | n/a | n/a | — |
The finding · why measurement matters
The one drag on figure-grounding wasn't a hallucination. It was the engine being helpful in a way a court can't accept. In the wire-transfer document it reported:
The source never states $430,000 — it states $250,000 and $180,000. The engine added them. The sum is correct, but it was presented as if extracted from the document when it was actually derived. For a lawyer, that difference — a stated figure vs. the engine's arithmetic — is the difference between evidence and analysis. Under a Daubert lens, an unlabeled derived figure is precisely the kind of methodology gap opposing counsel goes hunting for.
The evaluation caught this automatically, on the first run. That is the entire thesis of the product, demonstrated: the measurement surfaces what the engine hides.
What it changes
Limitations · stated plainly