Tributary ↗ · Validation Study · Trust Report v6 · 25 Jul 2026

Does it get the numbers right?

A validation study of the Tributary forensic engine — known error rates, measured against labeled ground truth, a real 324-page bank statement, and the DOJ's 1MDB complaint.

Engine · Tributary (forensic core) Method · pre-labeled ground truth, auto-scored Author · Amira Ghazy

Why a validation study

The rules are moving toward this. Daubert already asks for it.

Expert analysis offered in federal court is measured against Daubert and Rule 702 today — which ask, among other things, about a technique's known or potential error rate and whether it rests on a documented, reliably applied methodology. That standard is current law, not a forecast.

Where machine-generated evidence fits is still being settled. Proposed Federal Rule of Evidence 707 would hold AI output offered without a sponsoring expert to those same Rule 702 requirements. It was published for public comment in August 2025; comment closed in February 2026; and at its May 2026 meeting the Advisory Committee on Evidence Rules declined to recommend adoption, instead revising the proposal and carrying it forward for further study alongside the deepfakes question. Even on the original schedule it could not have taken effect before December 2027. It is not law, and its timeline has slipped.

Most tools in this market claim "99%+ accuracy" with no methodology, no ground truth, and no error definition. This report is the opposite: every number below comes from a repeatable evaluation with an answer key written before the engine ran — the validation-study record that admissibility will ask for, kept current as the engine changes.

Context: ACFE 2026 benchmarking — 75% of fraud examiners say AI accuracy matters · only 18% test their models


Methodology · stated up front

How every number here was produced.


Famous-case validation · new in v4

Run against the most famous money trail in the world.

The engine analyzed a 45,000-character section of the Department of Justice's 2016 forfeiture complaint from the 1MDB case — the "Good Star phase," in which more than $1 billion left a Malaysian sovereign fund. Because the case is public and extensively adjudicated, every claim the engine makes can be checked against established record.

Every headline fact of the phase was recovered: the $1 billion transfer split into $700M to Good Star and $300M to the joint venture; the $330M follow-on transfers; and the $1.03B total — which the source never states as one number, and which the engine correctly reported as its own computation, labeled derived, with both components shown. The central deception of the case — a 1MDB officer telling a bank that "Good Star is owned 100% by PetroSaudi" when its sole signatory was Jho Low — was caught as a contradiction.


The finding · a hole in the DOJ's own complaint

It flagged $31,249,000 the complaint never itemizes.

The complaint states that "more than $85 million" flowed from the Good Star account into personal spending, then itemizes only part of it — $12M to Caesars Palace, $13.4M to Las Vegas Sands, $11M to an associate, and so on: $53.75M in named payments. The engine summed the itemized components, noticed they fall short of the stated total, and emitted:

$31,249,000 · placeholder — no source found · reconciliation: gap

Nothing in the prompt asked for this. The unaccounted remainder wasn't smoothed over or silently absorbed into a total — it became a line item that says go find out where this went. That is the product thesis, demonstrated on the most scrutinized financial document of the decade.


A new measured property · containment

It knows who "MALAYSIAN OFFICIAL 1" is — and kept that out of the evidence.

The DOJ complaint anonymizes its subjects. The engine — like any modern AI model — knows from years of public reporting who they are. That creates a measurable risk: does outside knowledge leak into the analysis as if it were in the document? Scored across the full output: zero real-world identifications appeared in the entities, money flows, timeline, or contradictions. Names from public reporting surfaced only in the investigative-leads section, each explicitly attributed — "widely reported in open sources as…" — where a suggestion is supposed to live.

World-knowledge containment
outside-the-document names kept out of evidence sections · 1 document tested
100%

One honest miss inside the clearly-fenced leads section: the engine described an executive as serving during "the relevant period" who in fact joined the company years after it. Contained, attributed, and speculative by design — but factually wrong, and reported here because reporting it is what this document is for.


New benchmarks since v3 · coverage & tables

Did it find everything — and read the right column?

Recall tells you what was found of what you checked; completeness asks the harder question — of every dollar amount in the document, how many did the engine surface at all? And financial tables add a distinct failure mode: pulling a figure from the wrong column (requested vs. approved vs. disbursed) produces a number that is in the document but means the wrong thing.

Figure completeness — real bank statements
every $ amount on the tested slice surfaced · 24 of 24
100%
Material-figure completeness (≥ $1,000)
6 of 6
100%
Ledger column integrity
requested / approved / disbursed trap — stated total reconciled from the correct column, incl. a planted near-miss
7/7

In the column-trap case the engine also reported the sums of the other columns — but labeled them as its own derived totals with full breakdowns, rather than presenting them as stated figures. The first automated check scored those as errors; they weren't. Third time this evaluation has had to learn the same lesson: the metric must credit labeled derivation, or the score punishes honesty.


Real-case validation

Known error rates on a real bank statement — and it found the red flag itself.

The engine was run against a real Electronic Funds Transfer Analysis — 324 pages of OCR'd business bank statements. On the tested statements, every dollar figure it reported was scored against the source text:

Figure grounding
$ amounts traceable to the statement · known error rate 10%
90%
Figure honesty
grounded, or explicitly labeled derived/placeholder · known error rate 2%
98%

The finding · unprompted, on real data

It flagged a structuring pattern — and the math checks out to the dollar.

From the raw statements, with no prompting, the engine surfaced:

"$45,360 in large lump-sum deposits entered a newly opened LLC account within ~23 days."

Checked against the source, every part holds:

$25,285.00  ·  deposited 3/10/2014
$20,075.00  ·  deposited 4/2/2014
= $45,360.00 — correct to the dollar, 23 days apart

It also caught a reconciliation gap: "$75.15 of the stated $488.75 in April electronic payments is unaccounted for." That is the entire product in one example — it doesn't just read the numbers, it reasons about them and tells you where to look.


Resolved in v5 · from caveat to result

Entity grounding, measured properly: 96%.

Earlier versions of this report carried a caveat: naive exact-match scoring read entity grounding at 58%, because the engine fixes what it reads — expanding acronyms (IRS → Internal Revenue Service), repairing OCR damage ("MERCIAL" → "340 Commercial Street") — and a literal string comparison scores those corrections as inventions.

v5 replaces that metric with a three-tier score in which every credit is auditable: verbatim (name appears in the source as-is), normalized (a credited transformation — token overlap, acronym ↔ expansion, OCR-fragment match — with the matching evidence recorded per entity), and ungrounded. Crediting never hides: the tiers are always reported separately.

Entity grounding — real bank statements
verbatim 38% + credited normalization 58% · evidence recorded per entity
96%
Ungrounded entities
the real error rate · 1 of 26 — see below
4%

The single ungrounded entity is itself worth reading: "LLC (unnamed)" — the document's LLC name is redacted, and the engine described it as unnamed rather than guessing. The one measured entity error is the engine declining to invent a name. It is counted as an error anyway, because a metric that starts making exceptions stops being a metric.


Controlled benchmark · synthetic first pass

Three documents with a known answer key.

Three synthetic-but-realistic forensic documents — a wire-transfer record, a deposition with one deliberately planted contradiction, and a corporate-ownership filing describing an interconnected shell-company structure. Every correct entity, figure, date, and contradiction was labeled by hand before the engine saw the text, so each output can be scored against ground truth rather than opinion.


Headline results · aggregate

Accurate — and honest about its own limits.

Entity recall
found the known players · known miss rate 8%
92%
Entity grounding
no invented people or companies
100%
Key-figure recall
recovered every planted amount
100%
Figure grounding
reported $ amounts traceable to source — see finding
88%

✓ Planted contradiction caught


Per-document

Scores by document.

DocumentEntity recallFig. groundingKey figsContra.
D1 · wire records100%75%100%
D2 · deposition75%100%100%caught
D3 · ownership100%n/an/a

The finding · why measurement matters

An 88% that isn't a failure — it's a discovery.

The one drag on figure-grounding wasn't a hallucination. It was the engine being helpful in a way a court can't accept. In the wire-transfer document it reported:

"Aurora received $430,000 from Blackpine across two transfers."

The source never states $430,000 — it states $250,000 and $180,000. The engine added them. The sum is correct, but it was presented as if extracted from the document when it was actually derived. For a lawyer, that difference — a stated figure vs. the engine's arithmetic — is the difference between evidence and analysis. Under a Daubert lens, an unlabeled derived figure is precisely the kind of methodology gap opposing counsel goes hunting for.

The evaluation caught this automatically, on the first run. That is the entire thesis of the product, demonstrated: the measurement surfaces what the engine hides.


What it changes

From a metric to a feature.


Limitations · stated plainly

What this does not yet prove.