Kaia Benchmarks

The serving system, measured, on the record.

We publish the current accuracy of the serving model on frozen, versioned evaluation sets, release over release — the same number a GC's diligence team, a payer's compliance office, or a regulator should be able to demand. Vol. 1 is live below; the cadence is release over release. No cherry-picked demos, no one-time lab results.

Accounts Payable

92.75%

Exact disposition accuracy — the routing decision for every invoice (pay straight through, exception, held or blocked), matched exactly, measured internally on 1,200 held-out records across 35 distinct documents.

Status: MEASURED (internal) — internal measurement on a frozen evaluation set, published as-is

Model: Kaia’s native accounts-payable model — the model serving the accounts-payable lane in production since August 5, 2026

Evaluation set: 1,200 held-out gold-standard records across 35 distinct document bodies — frozen, never used in training

Method: The exact production adapter path and production response parser, run against the full set — not a sample. An exact score means everything matched: a partially right invoice counts as wrong.

Per-body consistency: 30 of the 35 distinct document bodies scored exactly on every one of their records; the macro-average exact rate across bodies was 93.13%.

The residual errors run conservative by design. On this set, sanctions-screening recall was 100% and straight-through-payment precision was 100% — no invoice that should have been held was ever routed to pay straight through. The failure mode we refuse is the expensive one.

Healthcare Claims

99.66%

Fraud-signal recall — 292 of 293 flagged-class claims caught on the held-out set, at 100% precision, measured internally.

Status: MEASURED (internal) — internal measurement on a frozen evaluation set, published as-is

Model: Kaia’s native claims model — the model serving the claims lane in production

Evaluation set: 860 held-out gold-standard records — frozen, never used in training

Method: The exact production adapter path and production response parser, run against the full set. The acceptance bar was pre-registered before the measurement ran — the number had to clear a written-down falsifier, not a retrospective judgment.

The whole-set number, same run: overall exact accuracy across all decision classes on the same 860 records was 96.05%. Both numbers stand on the record together.

Oil & Gas — published to a different standard

There is no model-accuracy score in this section — deliberately. The reserves lane is a deterministic engineering engine, not a trained classifier. Every figure it produces carries the calculation that produced it, and every result is reproducible from stored evidence — the inputs, the method, and the number, together.

For reserves estimation and disclosure work, the standard is the benchmark: a figure either carries its audit trail or it does not ship. When this series publishes for the Oil & Gas lane, it will publish engineering-verification results — never a model score dressed as one.

in build Electrical design (construction): the same engineering-grade standard, extended to construction electrical design. Produced to date: a design-basis report, load calculations, and a bill of quantities. No benchmark until the lane earns one.

Run your own documents. Hold us to it.

Design partners get access to the serving model and frozen evaluation methodology. We publish the results — yours and ours.