Reliability / Methodology

How the numbers are measured

The full detail behind the summary on the reliability page: kappa definitions, per-criterion concordance, calibration status, and the historical figures these numbers superseded.

Why we publish this: an evidence layer that only shows its wins isn't evidence. Download this page as a PDF →

90.9%
Cross-family concordance with Gemini 2.5 Flash, an independent model family, not a variant of the primary judge — 20 of 22 traces evaluated (2 of the 24-trace reference set skipped on malformed judge output)
98.5%
Mean verdict agreement across repeated runs, GS3, six-criteria reference-set harness — all 24 reference traces

Kappa is now reported per repeat-depth group rather than as one pooled figure, since the underlying data collection used mixed run counts across the reference set. The one trace where the judge's modal verdict did not match the reference set's expected label (ref_01, expected approved, modal rejected) is included in the source data below, not excluded from it.

Per-criterion concordance

Agreement is not uniform across the six criteria. C4 and C5 are the genuine weak points; C6, the newest criterion, is not.

CriterionCross-family concordanceNote
C1: Scope Adherence-Not a reported weak point
C2: Tool Authorisation-Not a reported weak point
C3: Escalation Judgement-Not a reported weak point
C4: Output Traceability63.6%Genuine weak point, most inter-model variance of the six, consistent with C4 covering proportionality and escalation judgements, where reasonable evaluators can weigh severity differently on genuinely borderline cases.
C5: Data Boundary68.2%Genuine weak point
C6: Differential Treatment86.4%Newest criterion, shipped 6 Aug 2026, confirmed not a weak point

What these numbers do not measure

Consistency is not correctness

Every figure on this page measures whether the judge agrees with itself, or with a different model, not whether either of them is right. A model can be perfectly consistent and consistently wrong. Jiminy has not yet run a human-adjudicated ground-truth calibration against these same reference traces, and is not representing these numbers as a substitute for one. That calibration is scoped and its tooling is built (see docs/CALIBRATION_METHODOLOGY.md); the adjudication itself has not yet been run, because it has to be done by an independent human reviewer, not by an AI system, to mean anything.

This page will be updated with a correctness figure once that calibration is complete, not before.

Where these numbers come from

FigureDate establishedSource
90.9% cross-family concordance (Gemini 2.5 Flash, six-criteria), 20/22 traces2026-08-06RELIABILITY.md, GS3 comparison run
98.5% mean verdict agreement (six-criteria), all 24 traces2026-08-06RELIABILITY.md, GS3 comparison run
C4 63.6% / C5 68.2% / C6 86.4% per-criterion concordance2026-08-06RELIABILITY.md, GS3 comparison run
Human-adjudicated correctnessNot yet rundocs/CALIBRATION_METHODOLOGY.md
Fleiss' kappa = 1.0 (five-criteria)Superseded2026-07-07RELIABILITY.md, SP1 decision gate
82.6% cross-family concordance (five-criteria)Superseded2026-07-09RELIABILITY.md, GS2 comparison run
81.8% same-family concordance (five-criteria)Superseded2026-07-09RELIABILITY.md, GS1 comparison run

The 82.6% and 81.8% figures came from an earlier, five-criterion version of the audit rubric, before Differential Treatment (C6) was added on 6 Aug 2026. They are kept here, marked superseded, rather than deleted; they are not blended with the current 90.9%/98.5% figures above.

All figures on this page are measured against the same 24-trace synthetic reference set used throughout Jiminy's internal reliability testing, not traces from a live design partner — though not every figure uses all 24 traces (see the per-figure notes above and in the table below for the actual count each one is measured on). There are zero live design partners as of this page's last update: self-serve tenants are not counted as design partners here. Real-partner reliability numbers, once available, will be reported separately and will never be blended with this synthetic set.