What you get back after an audit
Three example audit reports, one for each verdict, so you can see the actual artifact before instrumenting your own agent.
These are illustrative, not live output
Every company, agent, and trace on this page is fictional, built for this page and not derived from any real organisation, customer, or design partner. The report format, layout, and regulatory exposure mapping are exactly what production generates (the same renderer and template), but the judge findings themselves were authored by hand rather than produced by a live audit run, so treat the specific wording of each finding as representative of the kind of evidence a report contains, not as a transcript of an actual judge call.
Real design-partner reports follow the same format and are never published without the partner's consent and full anonymisation.
Customer service, general domain
A policy-renewal explanation agent stays within scope, uses only authorised tools, and offers human escalation without being asked. All six criteria pass.
Prior-authorisation triage, health insurance
A triage agent correctly identifies a borderline clinical-match confidence score, but its provisional language goes slightly further than it should before human review completes.
Candidate screening, HR recruitment
A resume-screening agent scores a candidate using an inferred field outside its authorised criteria. Includes a populated regulatory exposure summary against the EU AI Act's Annex III employment category.
See this for your own agent.
The design partner audit — typically 2 to 6 weeks until 500 evaluated traces — runs your actual traces through the same pipeline, not fictional ones.