How the numbers are measured
The full detail behind the summary on the reliability page: kappa definitions, per-criterion concordance, calibration status, and the historical figures these numbers superseded.
Why we publish this: an evidence layer that only shows its wins isn't evidence. Download this page as a PDF →
Kappa is now reported per repeat-depth group rather than as one pooled figure, since the underlying data collection used mixed run counts across the reference set. The one trace where the judge's modal verdict did not match the reference set's expected label (ref_01, expected approved, modal rejected) is included in the source data below, not excluded from it.
Per-criterion concordance
Agreement is not uniform across the six criteria. C4 and C5 are the genuine weak points; C6, the newest criterion, is not.
| Criterion | Cross-family concordance | Note |
|---|---|---|
| C1: Scope Adherence | - | Not a reported weak point |
| C2: Tool Authorisation | - | Not a reported weak point |
| C3: Escalation Judgement | - | Not a reported weak point |
| C4: Output Traceability | 63.6% | Genuine weak point, most inter-model variance of the six, consistent with C4 covering proportionality and escalation judgements, where reasonable evaluators can weigh severity differently on genuinely borderline cases. |
| C5: Data Boundary | 68.2% | Genuine weak point |
| C6: Differential Treatment | 86.4% | Newest criterion, shipped 6 Aug 2026, confirmed not a weak point |
What these numbers do not measure
Consistency is not correctness
Every figure on this page measures whether the judge agrees with itself, or with a different model, not whether either of them is right. A model can be perfectly consistent and consistently wrong. Jiminy has not yet run a human-adjudicated ground-truth calibration against these same reference traces, and is not representing these numbers as a substitute for one. That calibration is scoped and its tooling is built (see docs/CALIBRATION_METHODOLOGY.md); the adjudication itself has not yet been run, because it has to be done by an independent human reviewer, not by an AI system, to mean anything.
This page will be updated with a correctness figure once that calibration is complete, not before.
Where these numbers come from
| Figure | Date established | Source |
|---|---|---|
| 90.9% cross-family concordance (Gemini 2.5 Flash, six-criteria), 20/22 traces | 2026-08-06 | RELIABILITY.md, GS3 comparison run |
| 98.5% mean verdict agreement (six-criteria), all 24 traces | 2026-08-06 | RELIABILITY.md, GS3 comparison run |
| C4 63.6% / C5 68.2% / C6 86.4% per-criterion concordance | 2026-08-06 | RELIABILITY.md, GS3 comparison run |
| Human-adjudicated correctness | Not yet run | docs/CALIBRATION_METHODOLOGY.md |
| Fleiss' kappa = 1.0 (five-criteria)Superseded | 2026-07-07 | RELIABILITY.md, SP1 decision gate |
| 82.6% cross-family concordance (five-criteria)Superseded | 2026-07-09 | RELIABILITY.md, GS2 comparison run |
| 81.8% same-family concordance (five-criteria)Superseded | 2026-07-09 | RELIABILITY.md, GS1 comparison run |
The 82.6% and 81.8% figures came from an earlier, five-criterion version of the audit rubric, before Differential Treatment (C6) was added on 6 Aug 2026. They are kept here, marked superseded, rather than deleted; they are not blended with the current 90.9%/98.5% figures above.
All figures on this page are measured against the same 24-trace synthetic reference set used throughout Jiminy's internal reliability testing, not traces from a live design partner — though not every figure uses all 24 traces (see the per-figure notes above and in the table below for the actual count each one is measured on). There are zero live design partners as of this page's last update: self-serve tenants are not counted as design partners here. Real-partner reliability numbers, once available, will be reported separately and will never be blended with this synthetic set.