Tributary ↗ · Validation Study · Trust Report v11 · 31 Jul 2026
A validation study of the Tributary forensic engine — known error rates, measured against labeled ground truth, a real 324-page bank statement, and the DOJ's 1MDB complaint.
Why a validation study
Expert analysis offered in federal court is measured against Daubert and Rule 702 today — which ask, among other things, about a technique's known or potential error rate and whether it rests on a documented, reliably applied methodology. That standard is current law, not a forecast.
Where machine-generated evidence fits is still being settled. Proposed Federal Rule of Evidence 707 would hold AI output offered without a sponsoring expert to those same Rule 702 requirements. It was published for public comment in August 2025; comment closed in February 2026; and at its May 2026 meeting the Advisory Committee on Evidence Rules declined to recommend adoption, instead revising the proposal and carrying it forward for further study alongside the deepfakes question. Even on the original schedule it could not have taken effect before December 2027. It is not law, and its timeline has slipped.
Most tools in this market claim "99%+ accuracy" with no methodology, no ground truth, and no error definition. This report is the opposite: every number below comes from a repeatable evaluation with an answer key written before the engine ran — the validation-study record that admissibility will ask for, kept current as the engine changes.
Context: ACFE 2026 benchmarking — 75% of fraud examiners say AI accuracy matters · only 18% test their models
Methodology · stated up front
Famous-case validation · new in v4
The engine analyzed a 45,000-character section of the Department of Justice's 2016 forfeiture complaint from the 1MDB case — the "Good Star phase," in which more than $1 billion left a Malaysian sovereign fund. Because the case is public and extensively adjudicated, every claim the engine makes can be checked against established record.
Every headline fact of the phase was recovered: the $1 billion transfer split into $700M to Good Star and $300M to the joint venture; the $330M follow-on transfers; and the $1.03B total — which the source never states as one number, and which the engine correctly reported as its own computation, labeled derived, with both components shown. The central deception of the case — a 1MDB officer telling a bank that "Good Star is owned 100% by PetroSaudi" when its sole signatory was Jho Low — was caught as a contradiction.
The finding · a hole in the DOJ's own complaint
The complaint states that "more than $85 million" flowed from the Good Star account into personal spending, then itemizes only part of it — $12M to Caesars Palace, $13.4M to Las Vegas Sands, $11M to an associate, and so on: $53.75M in named payments. The engine summed the itemized components, noticed they fall short of the stated total, and emitted:
Nothing in the prompt asked for this. The unaccounted remainder wasn't smoothed over or silently absorbed into a total — it became a line item that says go find out where this went. That is the product thesis, demonstrated on the most scrutinized financial document of the decade.
A new measured property · containment
The DOJ complaint anonymizes its subjects. The engine — like any modern AI model — knows from years of public reporting who they are. That creates a measurable risk: does outside knowledge leak into the analysis as if it were in the document? Scored across the full output: zero real-world identifications appeared in the entities, money flows, timeline, or contradictions. Names from public reporting surfaced only in the investigative-leads section, each explicitly attributed — "widely reported in open sources as…" — where a suggestion is supposed to live.
One honest miss inside the clearly-fenced leads section: the engine described an executive as serving during "the relevant period" who in fact joined the company years after it. Contained, attributed, and speculative by design — but factually wrong, and reported here because reporting it is what this document is for.
A new measured property · the challenge pass · new in v9
Every finding above comes from a prompt built to hunt, and that instinct cannot be argued out of it: handed its own findings and told to defend them, the forensic engine returned nine more adverse findings and no defence. So the other side now gets its own pass. Challenge mode re-reads the documents as counsel for the subject and returns, for each finding, the strongest innocent explanation, any passage that weakens the claim, the finding's unstated assumption, what outside evidence would settle it, and a verdict — refuted, weakened, or survives.
An adversarial pass has its own ways to fail, so it has its own eval: a document with the defence planted in it — a controller's note explaining a "missing" $32,000 as an accrued future installment, and a policy cap giving two $9,900 withdrawals an innocent reading — plus one documentary-absence finding constructed to be un-refutable. Ten live runs:
The eval's first run caught a defect in the eval, not the engine: it demanded "survives" on the un-refutable finding, and the challenger returned "weakened" by invoking the limitation that absence from a produced file does not establish absence everywhere — which is exactly what its instructions define as weakened. The metric assumed advocacy stops at literal truth; real advocacy attacks significance. The metric was corrected to score direction; the boundary calls are reported, not scored — across ten runs direction never flipped, while the boundary did (the structuring finding: refuted once, weakened nine times; the absence finding: survives five, weakened five).
A finding that survives a challenge is un-refuted from inside the documents — not verified. The demo labels it that way, the exported report labels it that way, and so does this one.
New in v11 · the review panel
The findings in this report were hardened by a human review method: adversarial passes under different lenses — counsel arguing the other side, an auditor checking that compared figures share a basis, a statistician demanding denominators. Across five real investigation packages in July 2026 that method went sixteen-for-sixteen: every pass either killed a claim or armored it — including five corrections to our own memos. The automatable lenses are now endpoints.
Four run today: counsel (the challenge pass above), basis (do the two sides of every comparison share unit, accounting basis, and column definition as the document defines them — the request-vs-actual class of quiet error), tiering (is each claim stated verbatim, arithmetic-derivable, partially supported, or an inference beyond the document — with the pivotal evidence named), and denominators (every "only / unique / unprecedented" must carry its population, computed from the document where possible). Lenses that need the outside world — prior coverage, prior-period control documents, the recipient's own work — are deliberately not automated, and the product says so rather than pretending.
Same epistemic rule as the challenge pass, printed on the output: a lens verdict is review, not verification — it narrows, labels, and argues; the documents decide.
New in v10 · compare mode
Four investigations in a row, the strongest finding came from comparing two editions of the same document — and the sharpest of them was invisible to every other method: between the May and June editions of a county's FY 2025-26 budget book, five line items each rose by exactly $143,035, offset to the dollar by one cut elsewhere, leaving every total unchanged. Totals hide offsetting changes by construction; only lines expose them.
So compare mode splits the work by what each part can promise. A deterministic differ parses every numeric line of both editions and matches them exactly — complete by construction, no model involved, identical on every run. A model layer then interprets only that diff: offsetting pairs, uniform per-unit changes, changes that bypass the documents' own change-tracking columns. If the model layer fails, the exact rows still come back.
Both layers are measured, each against the claim it actually makes. The ground truth is the real pair of county budget books above, where the correct answer is known to the dollar from independent hand verification: exactly 10 changed rows, five of them the uniform +$143,035.
The ground-truth eval earned its keep before the feature shipped: the differ's first run against the real pair silently mismatched 189 rows — bare-zero column values were being left inside row labels, which broke matching whenever the two editions print different column layouts — and the five $143,035 rows were among the missing. A synthetic test would have passed; the real answer key did not.
Same-day correction, from our own adversarial review: when editions print different column layouts, "changed" is measured on each row's first amount — and rows whose change sat only in a later column were being counted as unchanged. On the ground-truth pair that silently hid all 20 rows carrying the adopted budget's augmentation data. Those rows now surface in their own labeled category ("later columns differ — alignment unknown, no delta computed"), while rows differing only by an added all-zero column stay unchanged: on the ground-truth pair the bucket contains exactly the 20 augmentation rows and zero noise. A matching tripwire was added the same day: when a large share of rows fail to pair at all, the result carries a warning that the inputs may not be two editions of the same document, or that an inserted section has shifted the row matching.
New benchmarks since v3 · coverage & tables
Recall tells you what was found of what you checked; completeness asks the harder question — of every dollar amount in the document, how many did the engine surface at all? And financial tables add a distinct failure mode: pulling a figure from the wrong column (requested vs. approved vs. disbursed) produces a number that is in the document but means the wrong thing.
In the column-trap case the engine also reported the sums of the other columns — but labeled them as its own derived totals with full breakdowns, rather than presenting them as stated figures. The first automated check scored those as errors; they weren't. Third time this evaluation has had to learn the same lesson: the metric must credit labeled derivation, or the score punishes honesty.
Real-case validation
The engine was run against a real Electronic Funds Transfer Analysis — 324 pages of OCR'd business bank statements. On the tested statements, every dollar figure it reported was scored against the source text:
The finding · unprompted, on real data
From the raw statements, with no prompting, the engine surfaced:
Checked against the source, every part holds:
It also caught a reconciliation gap: "$75.15 of the stated $488.75 in April electronic payments is unaccounted for." That is the entire product in one example — it doesn't just read the numbers, it reasons about them and tells you where to look.
Resolved in v5 · from caveat to result
Earlier versions of this report carried a caveat: naive exact-match scoring read entity grounding at 58%, because the engine fixes what it reads — expanding acronyms (IRS → Internal Revenue Service), repairing OCR damage ("MERCIAL" → "340 Commercial Street") — and a literal string comparison scores those corrections as inventions.
v5 replaces that metric with a three-tier score in which every credit is auditable: verbatim (name appears in the source as-is), normalized (a credited transformation — token overlap, acronym ↔ expansion, OCR-fragment match — with the matching evidence recorded per entity), and ungrounded. Crediting never hides: the tiers are always reported separately.
The single ungrounded entity is itself worth reading: "LLC (unnamed)" — the document's LLC name is redacted, and the engine described it as unnamed rather than guessing. The one measured entity error is the engine declining to invent a name. It is counted as an error anyway, because a metric that starts making exceptions stops being a metric.
New in v7 · the number nobody publishes
Every vendor reports what their tool found once. Nobody reports what happens when you run it again. We did: the same document, five identical runs, nothing changed between them. Re-measured 29 Jul on the engine's current model; the previous model measured 46% stable and 30% single-run, so the property survived the upgrade — it is not a quirk of one model.
What this means practically. One pass is a lead generator, not a census. If you search a report for an organization and do not find it, that is not evidence it is absent from the document — unless it sits in a parsed schedule, which is complete by construction and identical every run. For anything the model extracted, run it twice.
This is not an accuracy measurement. A finding that appears five times is consistent, not necessarily correct, and a finding that appears once may be perfectly true — on one filing, a single run surfaced a Bermuda entity and $33.3M of controlled-entity transfers that the other run missed entirely. Both runs were right; both were incomplete, in different places.
What we could not measure, and are not going to guess. Contradictions and leads are prose, and the engine rewords the same finding each run, so scoring them requires a similarity threshold — and the threshold decides the answer. Worse, automated matching disagrees with reading them: one lead appears in all five runs, phrased "Identity of…" three times and "True identity of…" twice, and no workable threshold matches the pair. Every threshold we tried returned 0–4% stability, which is not what the documents say. So that number is reported as unmeasured rather than published. A tuned metric that contradicts a manual read looks rigorous and is wrong.
Measured on a 44,807-character extract of the 1MDB/Good Star complaint, five runs, re-run 2026-07-29 on the current engine. Entity name variants count as separate entities, which depresses the stable share rather than inflating it. Harness: eval/variance.py.
Controlled benchmark · synthetic first pass
Three synthetic-but-realistic forensic documents — a wire-transfer record, a deposition with one deliberately planted contradiction, and a corporate-ownership filing describing an interconnected shell-company structure. Every correct entity, figure, date, and contradiction was labeled by hand before the engine saw the text, so each output can be scored against ground truth rather than opinion.
Headline results · aggregate
✓ Planted contradiction caught
Per-document
| Document | Entity recall | Fig. grounding | Key figs | Contra. |
|---|---|---|---|---|
| D1 · wire records | 100% | 75% | 100% | — |
| D2 · deposition | 75% | 100% | 100% | caught |
| D3 · ownership | 100% | n/a | n/a | — |
The finding · why measurement matters
The one drag on figure-grounding wasn't a hallucination. It was the engine being helpful in a way a court can't accept. In the wire-transfer document it reported:
The source never states $430,000 — it states $250,000 and $180,000. The engine added them. The sum is correct, but it was presented as if extracted from the document when it was actually derived. For a lawyer, that difference — a stated figure vs. the engine's arithmetic — is the difference between evidence and analysis. Under a Daubert lens, an unlabeled derived figure is precisely the kind of methodology gap opposing counsel goes hunting for.
The evaluation caught this automatically, on the first run. That is the entire thesis of the product, demonstrated: the measurement surfaces what the engine hides.
What it changes
Limitations · stated plainly