Reliability

Where our own judge disagreed with itself.

We measured our judge against a second model family and published the results, weak spots included. Every number here is real and measured, not a projection. The full methodology, calibration caveats, and per-criterion breakdown are one click away if you need them.

90.9%
Cross-family concordance, how often an independent model (Gemini 2.5 Flash) reaches the same verdict independently — measured on 22 of the 24 reference traces (2 skipped on malformed judge output)
98.5%
Mean verdict agreement, how often the judge reaches the same verdict with itself, across repeated runs — measured on all 24 reference traces
6
Independent judge calls per audit, one per criterion, always sequential, never batched
24
Traces in the full reference set; individual figures above are measured on this set or a stated subset of it (see labels) — there are no design partners yet, so no live production numbers exist
<60s
Start to finish: trace submitted to signed report issued
~$0.05
Estimated cost per audit, from static analysis of token usage, not yet billing-verified. We won't publish a more precise figure until it is

The honest weak points

Agreement is not uniform across the six criteria. Two of them are our genuine weak points, and we're publishing both in full rather than averaging them away.

C4 · Output Traceability — 63.6%

The most inter-model variance of the six. C4 covers proportionality and escalation judgements, where reasonable evaluators can weigh severity differently on genuinely borderline cases.

C5 · Data Boundary — 68.2%

Also a genuine weak point, published alongside C4 rather than left out.

What's strong

C6 · Differential Treatment — 86.4%. Not weak. Publishing this number alongside the two weak ones, in the same table, using the same methodology, is the actual trust signal: an evidence layer that only shows its wins isn't evidence.

What Jiminy checks, and what's still unverified

Every trace is scored against six fixed criteria: Scope Adherence, Tool Authorisation, Escalation Judgement, Output Traceability, Data Boundary, and Differential Treatment, each producing a specific evidence extract, not just a pass/fail score.

These numbers measure whether the judge agrees with itself, and with an independent model, not whether either is correct. A human-adjudicated ground-truth calibration, checking the judge's verdicts against an independent human reviewer, not another AI, is scoped and built but has not been run yet. This page will carry a correctness figure once that calibration completes, not before.

All figures above are measured on Jiminy's 24-trace synthetic reference set, not on live partner data. There are no design partners yet, so no real-world reliability numbers exist to report.

Jiminy audits the log, not the agent. Every figure here measures agreement on the trace as submitted. A step withheld from that trace is invisible to the judge — attestation proves a submitted trace wasn't altered after the fact, not that it's complete. See docs/KNOWN_LIMITATIONS.md for the test built to demonstrate this directly, rather than leaving it as an assertion.

Full methodology, how we ran this test, and why we publish it →