Kaia Benchmarks
The serving system, measured, on the record.
We publish the current accuracy of the serving model on frozen, versioned evaluation sets, release over release — the same number a GC's diligence team, a payer's compliance office, or a regulator should be able to demand. Vol. 1 is live below; the cadence is release over release. No cherry-picked demos, no one-time lab results.
August 2026
Vol. 1Legal E-Discovery
95.7%
Privilege-detection accuracy — measured internally on 810 documents, published as-is.
Status: MEASURED (internal) — internal measurement on a frozen evaluation set, published as-is
Model: Kaia’s native legal model — the model serving in production since July 31, 2026
Evaluation set: 810 held-out gold-standard records across 94 distinct document bodies — frozen, never used in training
Method: The exact production prompt and production response parser, temperature 0, run against the full set — not a sample
Serving provenance: The prompt measured here is the prompt serving in production — pinned by cryptographic hash in the codebase and enforced in CI. A single changed byte fails the build until the change is re-accepted on a fresh benchmark run, and every live classification stamps the hash it served under.
Privilege detection is the axis that matters most in regulated e-discovery: a missed privileged document is the failure mode that ends careers and waives rights. That is why it is the axis we publish first.
The full results — including the axes we're not proud of yet
Transparency means the whole table, not the best row. Per-record results on the same 810-document frozen set:
| Axis (per record, n=810) | Correct | Rate |
|---|---|---|
| Privilege(the regulated claim) | 775 | 95.7% |
| Sensitivity | 742 | 91.6% |
| Relevance | 546 | 67.4% |
| All three axes exact | 516 | 63.7% |
| Unparseable responses | 19 | 2.3% |
What we say about the weaker rows, on the record
Relevance is flat at 67.4%. This is a known, tracked limitation. Relevance scoring is a separate training-corpus workstream now underway; it is not hidden, and we will publish its movement in this series — up or down.
2.3% of responses fail parsing. A parse failure is fail-closed by design: the document routes to human review. The system never silently guesses. We count these against ourselves here.
How to read this number honestly
This is an internal measurement. The evaluation set is frozen and held out from training, but the measurement is run by us. We publish it because buyers deserve the number and the method — and we invite design partners to run their own documents through the system and hold us to it.
The number is re-measured. Corrections route through the Intelligence Engine, inside each client's own tenant. As engagements go live, this series will publish any movement — up or down — measured on the same frozen evals, in public, release over release.
A note on metric history: our July acceptance bar was privilege recall on the privileged class (94.74% against an 84.7% bar, up from 74.9% for the v1 model). This publication reports per-axis accuracy. Both metrics stand on the record; going forward this series standardizes on per-axis accuracy over the frozen set.
Methodology appendix
Set construction
810 records / 94 distinct bodies, gold-standard labeled, held out from all training runs. Gate-verified (#425).
Serving path
Bedrock Custom Model Import, production registry entry — the measurement exercised the same prompt, parser, and temperature that every production classification runs under.
Cold start
The run exercised real cold-start behavior (retry-until-warm, then ~1.4–1.9s per call) — production conditions, not lab conditions.
Per-body consistency
41 of 94 distinct document bodies scored perfectly across all their records; macro-average exact-three-axes rate 62.5%.
Next in this series
- Next release close:movement on all axes; first read on the relevance training-corpus workstream.
- Now published below:Accounts Payable and Healthcare Claims on the same frozen-set standard — and, once client engagements are running, measured per-engagement improvement from the correction loop.
Questions from diligence teams: benchmarks@kaiaai.ai. We answer with the raw data.
Accounts Payable
92.75%
Exact disposition accuracy — the routing decision for every invoice (pay straight through, exception, held or blocked), matched exactly, measured internally on 1,200 held-out records across 35 distinct documents.
Status: MEASURED (internal) — internal measurement on a frozen evaluation set, published as-is
Model: Kaia’s native accounts-payable model — the model serving the accounts-payable lane in production since August 5, 2026
Evaluation set: 1,200 held-out gold-standard records across 35 distinct document bodies — frozen, never used in training
Method: The exact production adapter path and production response parser, run against the full set — not a sample. An exact score means everything matched: a partially right invoice counts as wrong.
Per-body consistency: 30 of the 35 distinct document bodies scored exactly on every one of their records; the macro-average exact rate across bodies was 93.13%.
The residual errors run conservative by design. On this set, sanctions-screening recall was 100% and straight-through-payment precision was 100% — no invoice that should have been held was ever routed to pay straight through. The failure mode we refuse is the expensive one.
Healthcare Claims
99.66%
Fraud-signal recall — 292 of 293 flagged-class claims caught on the held-out set, at 100% precision, measured internally.
Status: MEASURED (internal) — internal measurement on a frozen evaluation set, published as-is
Model: Kaia’s native claims model — the model serving the claims lane in production
Evaluation set: 860 held-out gold-standard records — frozen, never used in training
Method: The exact production adapter path and production response parser, run against the full set. The acceptance bar was pre-registered before the measurement ran — the number had to clear a written-down falsifier, not a retrospective judgment.
The whole-set number, same run: overall exact accuracy across all decision classes on the same 860 records was 96.05%. Both numbers stand on the record together.
Oil & Gas — published to a different standard
There is no model-accuracy score in this section — deliberately. The reserves lane is a deterministic engineering engine, not a trained classifier. Every figure it produces carries the calculation that produced it, and every result is reproducible from stored evidence — the inputs, the method, and the number, together.
For reserves estimation and disclosure work, the standard is the benchmark: a figure either carries its audit trail or it does not ship. When this series publishes for the Oil & Gas lane, it will publish engineering-verification results — never a model score dressed as one.
in build Electrical design (construction): the same engineering-grade standard, extended to construction electrical design. Produced to date: a design-basis report, load calculations, and a bill of quantities. No benchmark until the lane earns one.
Run your own documents. Hold us to it.
Design partners get access to the serving model and frozen evaluation methodology. We publish the results — yours and ours.