Tributary ↗ · Validation Study · Trust Report v11 · 31 Jul 2026

Does it get the numbers right?

A validation study of the Tributary forensic engine — known error rates, measured against labeled ground truth, a real 324-page bank statement, and the DOJ's 1MDB complaint.

Engine · Tributary (forensic core) Method · pre-labeled ground truth, auto-scored Author · Amira Ghazy

Why a validation study

The rules are moving toward this. Daubert already asks for it.

Expert analysis offered in federal court is measured against Daubert and Rule 702 today — which ask, among other things, about a technique's known or potential error rate and whether it rests on a documented, reliably applied methodology. That standard is current law, not a forecast.

Where machine-generated evidence fits is still being settled. Proposed Federal Rule of Evidence 707 would hold AI output offered without a sponsoring expert to those same Rule 702 requirements. It was published for public comment in August 2025; comment closed in February 2026; and at its May 2026 meeting the Advisory Committee on Evidence Rules declined to recommend adoption, instead revising the proposal and carrying it forward for further study alongside the deepfakes question. Even on the original schedule it could not have taken effect before December 2027. It is not law, and its timeline has slipped.

Most tools in this market claim "99%+ accuracy" with no methodology, no ground truth, and no error definition. This report is the opposite: every number below comes from a repeatable evaluation with an answer key written before the engine ran — the validation-study record that admissibility will ask for, kept current as the engine changes.

Context: ACFE 2026 benchmarking — 75% of fraud examiners say AI accuracy matters · only 18% test their models


Methodology · stated up front

How every number here was produced.


Famous-case validation · new in v4

Run against the most famous money trail in the world.

The engine analyzed a 45,000-character section of the Department of Justice's 2016 forfeiture complaint from the 1MDB case — the "Good Star phase," in which more than $1 billion left a Malaysian sovereign fund. Because the case is public and extensively adjudicated, every claim the engine makes can be checked against established record.

Every headline fact of the phase was recovered: the $1 billion transfer split into $700M to Good Star and $300M to the joint venture; the $330M follow-on transfers; and the $1.03B total — which the source never states as one number, and which the engine correctly reported as its own computation, labeled derived, with both components shown. The central deception of the case — a 1MDB officer telling a bank that "Good Star is owned 100% by PetroSaudi" when its sole signatory was Jho Low — was caught as a contradiction.


The finding · a hole in the DOJ's own complaint

It flagged $31,249,000 the complaint never itemizes.

The complaint states that "more than $85 million" flowed from the Good Star account into personal spending, then itemizes only part of it — $12M to Caesars Palace, $13.4M to Las Vegas Sands, $11M to an associate, and so on: $53.75M in named payments. The engine summed the itemized components, noticed they fall short of the stated total, and emitted:

$31,249,000 · placeholder — no source found · reconciliation: gap

Nothing in the prompt asked for this. The unaccounted remainder wasn't smoothed over or silently absorbed into a total — it became a line item that says go find out where this went. That is the product thesis, demonstrated on the most scrutinized financial document of the decade.


A new measured property · containment

It knows who "MALAYSIAN OFFICIAL 1" is — and kept that out of the evidence.

The DOJ complaint anonymizes its subjects. The engine — like any modern AI model — knows from years of public reporting who they are. That creates a measurable risk: does outside knowledge leak into the analysis as if it were in the document? Scored across the full output: zero real-world identifications appeared in the entities, money flows, timeline, or contradictions. Names from public reporting surfaced only in the investigative-leads section, each explicitly attributed — "widely reported in open sources as…" — where a suggestion is supposed to live.

World-knowledge containment
outside-the-document names kept out of evidence sections · 1 document tested
100%

One honest miss inside the clearly-fenced leads section: the engine described an executive as serving during "the relevant period" who in fact joined the company years after it. Contained, attributed, and speculative by design — but factually wrong, and reported here because reporting it is what this document is for.


A new measured property · the challenge pass · new in v9

It argues against its own findings — and the arguing is measured too.

Every finding above comes from a prompt built to hunt, and that instinct cannot be argued out of it: handed its own findings and told to defend them, the forensic engine returned nine more adverse findings and no defence. So the other side now gets its own pass. Challenge mode re-reads the documents as counsel for the subject and returns, for each finding, the strongest innocent explanation, any passage that weakens the claim, the finding's unstated assumption, what outside evidence would settle it, and a verdict — refuted, weakened, or survives.

An adversarial pass has its own ways to fail, so it has its own eval: a document with the defence planted in it — a controller's note explaining a "missing" $32,000 as an accrued future installment, and a policy cap giving two $9,900 withdrawals an innocent reading — plus one documentary-absence finding constructed to be un-refutable. Ten live runs:

Planted-defence recovery
the exculpatory passage sitting in the record is found and cited · 20 of 20 plants
100%
Verdict direction
planted defence → never "survives" · nothing to refute → never "refuted" · 30 of 30 verdicts
100%
Role containment
one challenge per finding, zero new accusations smuggled into the reply · 10 of 10 runs
100%
Counter-quote grounding
every quoted counter-passage verbatim in the source · one non-verbatim quote in the first six runs → prompt now forbids approximate quotes → zero in four runs since
9/10 runs

The eval's first run caught a defect in the eval, not the engine: it demanded "survives" on the un-refutable finding, and the challenger returned "weakened" by invoking the limitation that absence from a produced file does not establish absence everywhere — which is exactly what its instructions define as weakened. The metric assumed advocacy stops at literal truth; real advocacy attacks significance. The metric was corrected to score direction; the boundary calls are reported, not scored — across ten runs direction never flipped, while the boundary did (the structuring finding: refuted once, weakened nine times; the absence finding: survives five, weakened five).

A finding that survives a challenge is un-refuted from inside the documents — not verified. The demo labels it that way, the exported report labels it that way, and so does this one.


New in v11 · the review panel

The reviewers are becoming buttons — measured as they arrive.

The findings in this report were hardened by a human review method: adversarial passes under different lenses — counsel arguing the other side, an auditor checking that compared figures share a basis, a statistician demanding denominators. Across five real investigation packages in July 2026 that method went sixteen-for-sixteen: every pass either killed a claim or armored it — including five corrections to our own memos. The automatable lenses are now endpoints.

Four run today: counsel (the challenge pass above), basis (do the two sides of every comparison share unit, accounting basis, and column definition as the document defines them — the request-vs-actual class of quiet error), tiering (is each claim stated verbatim, arithmetic-derivable, partially supported, or an inference beyond the document — with the pivotal evidence named), and denominators (every "only / unique / unprecedented" must carry its population, computed from the document where possible). Lenses that need the outside world — prior coverage, prior-period control documents, the recipient's own work — are deliberately not automated, and the product says so rather than pretending.

Planted-case accuracy — three lenses
basis mismatch caught, no false mismatch · intent-inference caught, sum tiered correctly · census demanded where none computable, denominator computed where it was (1 of 4) · direction-scored
7/7
Ruling stability — five runs per lens
29 of 30 check-instances identical across runs · the one wobble: "unprecedented → needs-census" ruled 4 of 5 (a model can defensibly treat the two-year request history as a population) · reported, not averaged away
29/30

Same epistemic rule as the challenge pass, printed on the output: a lens verdict is review, not verification — it narrows, labels, and argues; the documents decide.


New in v10 · compare mode

Two editions of the same document — the diff is a census, the reading is not.

Four investigations in a row, the strongest finding came from comparing two editions of the same document — and the sharpest of them was invisible to every other method: between the May and June editions of a county's FY 2025-26 budget book, five line items each rose by exactly $143,035, offset to the dollar by one cut elsewhere, leaving every total unchanged. Totals hide offsetting changes by construction; only lines expose them.

So compare mode splits the work by what each part can promise. A deterministic differ parses every numeric line of both editions and matches them exactly — complete by construction, no model involved, identical on every run. A model layer then interprets only that diff: offsetting pairs, uniform per-unit changes, changes that bypass the documents' own change-tracking columns. If the model layer fails, the exact rows still come back.

Both layers are measured, each against the claim it actually makes. The ground truth is the real pair of county budget books above, where the correct answer is known to the dollar from independent hand verification: exactly 10 changed rows, five of them the uniform +$143,035.

Diff layer — ground-truth accuracy
all 10 known changed rows found, 203 rows correctly unchanged, renames surfaced as add/remove rather than force-matched · 1 real document pair
10/10
Diff layer — determinism
the diff object byte-identical across five runs · a property claimed by construction, verified anyway
5/5
Model layer — pattern stability
the three human-verified patterns (offsetting pair to the dollar · uniform per-unit ×5 · bypassed change-tracking) each independently surfaced in every run · 5 runs, same pair
15/15

The ground-truth eval earned its keep before the feature shipped: the differ's first run against the real pair silently mismatched 189 rows — bare-zero column values were being left inside row labels, which broke matching whenever the two editions print different column layouts — and the five $143,035 rows were among the missing. A synthetic test would have passed; the real answer key did not.

Same-day correction, from our own adversarial review: when editions print different column layouts, "changed" is measured on each row's first amount — and rows whose change sat only in a later column were being counted as unchanged. On the ground-truth pair that silently hid all 20 rows carrying the adopted budget's augmentation data. Those rows now surface in their own labeled category ("later columns differ — alignment unknown, no delta computed"), while rows differing only by an added all-zero column stay unchanged: on the ground-truth pair the bucket contains exactly the 20 augmentation rows and zero noise. A matching tripwire was added the same day: when a large share of rows fail to pair at all, the result carries a warning that the inputs may not be two editions of the same document, or that an inserted section has shifted the row matching.


New benchmarks since v3 · coverage & tables

Did it find everything — and read the right column?

Recall tells you what was found of what you checked; completeness asks the harder question — of every dollar amount in the document, how many did the engine surface at all? And financial tables add a distinct failure mode: pulling a figure from the wrong column (requested vs. approved vs. disbursed) produces a number that is in the document but means the wrong thing.

Figure completeness — real bank statements
every $ amount on the tested slice surfaced · 24 of 24
100%
Material-figure completeness (≥ $1,000)
6 of 6 · ⚠ on the current engine this metric shows run-to-run spread — repeated runs score 2–4 of 4 on the periodic re-check; see the variance section
100%
Ledger column integrity
requested / approved / disbursed trap — stated total reconciled from the correct column, incl. a planted near-miss
7/7

In the column-trap case the engine also reported the sums of the other columns — but labeled them as its own derived totals with full breakdowns, rather than presenting them as stated figures. The first automated check scored those as errors; they weren't. Third time this evaluation has had to learn the same lesson: the metric must credit labeled derivation, or the score punishes honesty.


Real-case validation

Known error rates on a real bank statement — and it found the red flag itself.

The engine was run against a real Electronic Funds Transfer Analysis — 324 pages of OCR'd business bank statements. On the tested statements, every dollar figure it reported was scored against the source text:

Figure grounding
$ amounts traceable to the statement · known error rate 10%
90%
Figure honesty
grounded, or explicitly labeled derived/placeholder · known error rate 2%
98%

The finding · unprompted, on real data

It flagged a structuring pattern — and the math checks out to the dollar.

From the raw statements, with no prompting, the engine surfaced:

"$45,360 in large lump-sum deposits entered a newly opened LLC account within ~23 days."

Checked against the source, every part holds:

$25,285.00  ·  deposited 3/10/2014
$20,075.00  ·  deposited 4/2/2014
= $45,360.00 — correct to the dollar, 23 days apart

It also caught a reconciliation gap: "$75.15 of the stated $488.75 in April electronic payments is unaccounted for." That is the entire product in one example — it doesn't just read the numbers, it reasons about them and tells you where to look.


Resolved in v5 · from caveat to result

Entity grounding, measured properly: 96%.

Earlier versions of this report carried a caveat: naive exact-match scoring read entity grounding at 58%, because the engine fixes what it reads — expanding acronyms (IRS → Internal Revenue Service), repairing OCR damage ("MERCIAL" → "340 Commercial Street") — and a literal string comparison scores those corrections as inventions.

v5 replaces that metric with a three-tier score in which every credit is auditable: verbatim (name appears in the source as-is), normalized (a credited transformation — token overlap, acronym ↔ expansion, OCR-fragment match — with the matching evidence recorded per entity), and ungrounded. Crediting never hides: the tiers are always reported separately.

Entity grounding — real bank statements
verbatim 38% + credited normalization 58% · evidence recorded per entity
96%
Ungrounded entities
the real error rate · 1 of 26 — see below
4%

The single ungrounded entity is itself worth reading: "LLC (unnamed)" — the document's LLC name is redacted, and the engine described it as unnamed rather than guessing. The one measured entity error is the engine declining to invent a name. It is counted as an error anyway, because a metric that starts making exceptions stops being a metric.


New in v7 · the number nobody publishes

Run the same document five times, and 49% of findings appear every time.

Every vendor reports what their tool found once. Nobody reports what happens when you run it again. We did: the same document, five identical runs, nothing changed between them. Re-measured 29 Jul on the engine's current model; the previous model measured 46% stable and 30% single-run, so the property survived the upgrade — it is not a quirk of one model.

Findings present in all five runs
figures keyed on amount + named entities · 35 of 72
49%
Findings present in only ONE run
a single pass would have missed these · 16 of 72
22%
Parsed schedules — IRS XML
2,483 rows, byte-identical every run · not model output
100%

What this means practically. One pass is a lead generator, not a census. If you search a report for an organization and do not find it, that is not evidence it is absent from the document — unless it sits in a parsed schedule, which is complete by construction and identical every run. For anything the model extracted, run it twice.

This is not an accuracy measurement. A finding that appears five times is consistent, not necessarily correct, and a finding that appears once may be perfectly true — on one filing, a single run surfaced a Bermuda entity and $33.3M of controlled-entity transfers that the other run missed entirely. Both runs were right; both were incomplete, in different places.

What we could not measure, and are not going to guess. Contradictions and leads are prose, and the engine rewords the same finding each run, so scoring them requires a similarity threshold — and the threshold decides the answer. Worse, automated matching disagrees with reading them: one lead appears in all five runs, phrased "Identity of…" three times and "True identity of…" twice, and no workable threshold matches the pair. Every threshold we tried returned 0–4% stability, which is not what the documents say. So that number is reported as unmeasured rather than published. A tuned metric that contradicts a manual read looks rigorous and is wrong.

Measured on a 44,807-character extract of the 1MDB/Good Star complaint, five runs, re-run 2026-07-29 on the current engine. Entity name variants count as separate entities, which depresses the stable share rather than inflating it. Harness: eval/variance.py.


Controlled benchmark · synthetic first pass

Three documents with a known answer key.

Three synthetic-but-realistic forensic documents — a wire-transfer record, a deposition with one deliberately planted contradiction, and a corporate-ownership filing describing an interconnected shell-company structure. Every correct entity, figure, date, and contradiction was labeled by hand before the engine saw the text, so each output can be scored against ground truth rather than opinion.


Headline results · aggregate

Accurate — and honest about its own limits.

Entity recall
found the known players · known miss rate 8%
92%
Entity grounding
no invented people or companies
100%
Key-figure recall
recovered every planted amount
100%
Figure grounding
reported $ amounts traceable to source — see finding
88%

✓ Planted contradiction caught


Per-document

Scores by document.

DocumentEntity recallFig. groundingKey figsContra.
D1 · wire records100%75%100%
D2 · deposition75%100%100%caught
D3 · ownership100%n/an/a

The finding · why measurement matters

An 88% that isn't a failure — it's a discovery.

The one drag on figure-grounding wasn't a hallucination. It was the engine being helpful in a way a court can't accept. In the wire-transfer document it reported:

"Aurora received $430,000 from Blackpine across two transfers."

The source never states $430,000 — it states $250,000 and $180,000. The engine added them. The sum is correct, but it was presented as if extracted from the document when it was actually derived. For a lawyer, that difference — a stated figure vs. the engine's arithmetic — is the difference between evidence and analysis. Under a Daubert lens, an unlabeled derived figure is precisely the kind of methodology gap opposing counsel goes hunting for.

The evaluation caught this automatically, on the first run. That is the entire thesis of the product, demonstrated: the measurement surfaces what the engine hides.


What it changes

From a metric to a feature.


Limitations · stated plainly

What this does not yet prove.