Kaia Benchmarks

The serving system, measured, on the record.

We publish the current accuracy of the serving model on frozen, versioned evaluation sets, release over release — the same number a GC's diligence team, a payer's compliance office, or a regulator should be able to demand. Vol. 1 is live below; monthly is the cadence we intend. No cherry-picked demos, no one-time lab results.

Accounts Payable

92.75%

Exact extraction accuracy — the complete extracted record and routing decision, matched exactly, measured internally on 1,200 held-out invoices.

Status: MEASURED (internal) — internal measurement on a frozen evaluation set, published as-is

Model: kaia-native-ap-8b — the model serving the accounts-payable lane in production since August 5, 2026

Evaluation set: 1,200 held-out gold-standard records — frozen, never used in training

Method: The exact production adapter path and production response parser, run against the full set — not a sample. An exact score means everything matched: a partially right invoice counts as wrong.

The residual errors run conservative by design. On this set, sanctions-screening recall was 100% and straight-through-payment precision was 100% — no invoice that should have been held was ever paid. The failure mode we refuse is the expensive one.

Healthcare Claims

99.66%

Fraud-signal recall — 292 of 293 flagged-class claims caught on the held-out set, at 100% precision, measured internally.

Status: MEASURED (internal) — internal measurement on a frozen evaluation set, published as-is

Model: kaia-native-claims-8b — the model serving the claims lane in production

Evaluation set: 860 held-out gold-standard records — frozen, never used in training

Method: The exact production adapter path and production response parser, run against the full set. The acceptance bar was pre-registered before the measurement ran — the number had to clear a written-down falsifier, not a retrospective judgment.

The whole-set number, same run: overall exact accuracy across all decision classes on the same 860 records was 96.05%. Both numbers stand on the record together.

Oil & Gas — published to a different standard

There is no model-accuracy score in this section — deliberately. The reserves lane is a deterministic engineering engine, not a trained classifier. Every figure it produces carries the calculation that produced it, and every result is reproducible from stored evidence — the inputs, the method, and the number, together.

For reserves estimation and disclosure work, the standard is the benchmark: a figure either carries its audit trail or it does not ship. When this series publishes for the Oil & Gas lane, it will publish engineering-verification results — never a model score dressed as one.

in build Electrical design (construction): the same engineering-grade standard, extended to construction electrical design. Produced to date: a design-basis report, load calculations, and a bill of quantities. No benchmark until the lane earns one.

Run your own documents. Hold us to it.

Design partners get access to the serving model and frozen evaluation methodology. We publish the results — yours and ours.