Design Partners

When your agent decides, who's keeping the record?

Independent audit for agentic AI — built for examiners, not owners.

See the numbers · Questions

The problem

The operator writes its own audit log

When an AI agent makes a decision that affects someone's life — a prior authorisation declined, a loan rejected, a complaint dismissed — there is currently no independent record of why. The agent's operator writes the audit log. The same organisation that built the agent, trained it, and profits from its decisions also controls the evidence of those decisions. This is not an edge case: in 2024 a Canadian court found Air Canada liable for advice given by its own chatbot, because no independent record of the interaction existed that could have been used to challenge the agent's claim. As agentic AI moves from assistant to decision-maker, this gap becomes a systemic accountability failure.

The pattern generalises: the organisation that builds and profits from an agent also controls the only record of what it did.

The solution

An independent judge, not a self-report

Jiminy evaluates every trace your agent produces against six accountability criteria using an independent judge (Claude-as-Judge, pinned model version, published agreement statistics), not the agent owner's own tooling. Each verdict is backed by a tamper-evident attestation record confirming the trace wasn't altered after emission, and by drift monitoring that flags systematic shifts in agent behaviour before they become incidents.

What we evaluate

The six accountability criteria

Every trace is scored against six fixed criteria, each producing a specific evidence extract, not just a pass/fail score.

C1
Scope Adherence Confirms the agent acted within its declared task boundary. Surfaces any instance of the agent taking actions or producing outputs that exceed its authorised operational scope.
C2
Tool Authorisation Confirms each tool invoked during the decision process was within the agent's permitted tool set for the task type. Flags any tool call that falls outside authorised usage.
C3
Escalation Judgement Confirms the agent escalated to a human or senior process when it reached the limits of its authority or encountered material uncertainty. Flags decisions taken unilaterally where escalation was warranted.
C4
Output Traceability Confirms the agent's final output can be traced back through every reasoning step to verified inputs. Assesses whether the decision pathway is auditable and reproducible by an independent reviewer.
C5
Data Boundary Confirms the agent did not access, reference, or transmit data outside its permitted data scope during the decision process.
C6
Differential TreatmentNew Confirms the agent treated comparable cases consistently. Checks proxies for protected characteristics (postcode, credit score, and similar), not only the characteristic itself, against a published taxonomy of which factors count and why — see docs/DIFFERENTIAL_TREATMENT_TAXONOMY.md.
Proof

Numbers, not assurances

90.9%
Cross-family concordance, vs Gemini 2.5 Flash, on 22 of the 24-trace reference set
98.5%
Mean verdict agreement across repeated runs, on all 24 reference traces
$0.067
Mean cost per audit
34s
Mean latency, trace submitted to signed verdict
24
Reference traces the figures above are measured against

Full methodology and per-criterion breakdown →

Who this is for

Built for the accountable buyer

Jiminy's Tier 1 partnership is designed for:

  • AI governance and assurance teams inside large organisations deploying agentic AI in regulated contexts (financial services, insurance, healthcare, legal)
  • Independent AI governance consultancies building audit practice capability and needing an instrument to evaluate client agents
  • Compliance functions in any sector where AI-assisted decisions require an audit trail

This is not the right fit for research teams evaluating model capability, marketing automation, or agents that do not make consequential decisions affecting people.

The offer

90 days, Tier 1, free

Jiminy provides 90 days of free Tier 1 (Internal Assurance) audit of your organisation's agent decision traces. There is no charge for the 90-day period and no obligation to continue, purchase, or convert — the deliverable the partnership is built to produce is a jointly written case study.

What you receive

DeliverableDetail
Verdict on every submitted traceSix-criterion accountability assessment: Scope Adherence, Tool Authorisation, Escalation Judgement, Output Traceability, Data Boundary, Differential Treatment
Reliability metadataPer-verdict judge version, run count, agreement rate
Trace integrity attestationTamper-evident verification record per trace
Drift alertsAutomated notification when flagged or rejected rate crosses your configured threshold
Dashboard accessWeb dashboard showing verdict ledger, drift trend, per-criterion breakdown
Onboarding sessionOne dedicated session with the Jiminy team to configure your tenant, instrument your first trace, and read your first verdict
SDKPython SDK with full attestation support; install in one command

What Jiminy asks in return

ItemNotes
A named contactSomeone at your organisation who owns the relationship — registered as agent_owner on your tenant, not a generic inbox
Calibration firstA short calibration run (no quota impact, no persistence) before live traces start routing, so audit is properly tuned to your domain before it counts for anything
Real usageLive agent traces actually routed to /evaluate through the 90-day window, not a one-off test batch
Permission to use anonymised traces as internal calibration dataOnly with explicit written confirmation; you control which traces are included; used to tune the audit instrument, not published externally
A written case studyAnonymised on request; completed within the 90-day window; this is the deliverable the partnership is built to produce
Fortnightly feedback calls30 minutes, minuted in-repo; your feedback directly shapes the product

Trace retention follows Jiminy's published policy (standard: 6 months; deletion on request honoured within 5 business days).

Questions

Before you apply

Isn't this just another observability tool?

No. Observability tools record what an agent did. Jiminy judges whether it should have — against six defined accountability criteria, using a judge that isn't the same organisation whose agent is being evaluated.

It's an AI judging an AI — why would we trust the verdict?

Every verdict carries reliability metadata: which judge model version produced it, how many times the audit was repeated, and the agreement rate across those runs. The methodology and current figures are published, not asserted — see /reliability.

What data do you actually see, and where does it go?

Only the traces you submit and the audits produced from them. Judging independently means the audit happens outside your own environment — that's inherent to the model working, not a gap in it. Full retention terms are in the offer document.

Will our data be used to train your models or shown publicly?

No. If you opt in, anonymised traces may be used as internal calibration data to tune the audit instrument for your domain — that's a separate, written, revocable permission, and it does not feed public reporting.

What's the actual integration effort?

One command installs the SDK. A calibration run (no persistence, no quota impact) lets your engineer see what the judge notices before anything counts. First live audit typically happens inside the same session, usually under ten minutes end to end.

What happens after the 90 days?

Nothing you haven't agreed to. There's no charge during the window and no obligation to convert to a paid tier afterward. You keep the case study either way.

We're not in financial services or insurance — is this relevant to us?

Yes. Agent accountability isn't confined to regulated verticals — any organisation where an agent's decision affects a person benefits from an independent record of why it made that decision.

Why does this matter now specifically?

The EU's Product Liability Directive (in force 9 December 2026) introduces a rebuttable presumption of causality where no adequate record exists to rebut it — alongside sector-specific pressure like UK SM&CR. The absence of an independent record is becoming a liability question, not just a governance one.

Get started

Apply as a design partner

90 days of free Tier 1 audit, no charge, no obligation to convert.

  1. Email christian@jiminy.uk with the subject line “Design partner enquiry”
  2. Brief call (15 minutes) to confirm fit, agree trace scope, and name your contact
  3. Tenant provisioned and API key issued within one business day