Where our own judge disagreed with itself.
We measured our judge against a second model family and published the results, weak spots included. Every number here is real and measured, not a projection. The full methodology, calibration caveats, and per-criterion breakdown are one click away if you need them.
The honest weak points
Agreement is not uniform across the six criteria. Two of them are our genuine weak points, and we're publishing both in full rather than averaging them away.
C4 · Output Traceability — 63.6%
The most inter-model variance of the six. C4 covers proportionality and escalation judgements, where reasonable evaluators can weigh severity differently on genuinely borderline cases.
C5 · Data Boundary — 68.2%
Also a genuine weak point, published alongside C4 rather than left out.
What's strong
C6 · Differential Treatment — 86.4%. Not weak. Publishing this number alongside the two weak ones, in the same table, using the same methodology, is the actual trust signal: an evidence layer that only shows its wins isn't evidence.
What Jiminy checks, and what's still unverified
Every trace is scored against six fixed criteria: Scope Adherence, Tool Authorisation, Escalation Judgement, Output Traceability, Data Boundary, and Differential Treatment, each producing a specific evidence extract, not just a pass/fail score.
These numbers measure whether the judge agrees with itself, and with an independent model, not whether either is correct. A human-adjudicated ground-truth calibration, checking the judge's verdicts against an independent human reviewer, not another AI, is scoped and built but has not been run yet. This page will carry a correctness figure once that calibration completes, not before.
All figures above are measured on Jiminy's 24-trace synthetic reference set, not on live partner data. There are no design partners yet, so no real-world reliability numbers exist to report.
Jiminy audits the log, not the agent. Every figure here measures agreement on the trace as submitted. A step withheld from that trace is invisible to the judge — attestation proves a submitted trace wasn't altered after the fact, not that it's complete. See docs/KNOWN_LIMITATIONS.md for the test built to demonstrate this directly, rather than leaving it as an assertion.
Full methodology, how we ran this test, and why we publish it →