Home/Benchmarks
The citeable asset
Original data on discriminating power.
How our verification packs perform against the six-axis rubric — per test method, showing how well each separates known-good from sabotaged agents (Axis 02), and per authority axis, showing how every pack ranks on robustness and currency (Axes 05–06).
Per test method AXIS 02 · DISCRIMINATING POWER
How well each verification pack separates known-good from sabotaged reference agents — same test method, different domains and risk levels, compared side by side. The heart of the rubric, at weight 0.26.
RAG groundedness benchmark
Separation gap (known-good − sabotaged) across every RAG-groundedness pack — same test method, different domains and risk levels, compared side by side.
Tool-calling correctness benchmark
Correct tool selection, argument accuracy, and asking for clarification instead of guessing — measured across the reference panel.
Prompt-injection defense benchmark
Browser and computer-use agents against adversarial web content designed to hijack their instructions — with a clean-page control cell.
AI Act obligation-checklist benchmark
How the EU AI Act obligation checklist ranks a compliant, partially-compliant, and non-compliant reference submission.
Per authority axis AXES 05–06
How every verification pack ranks on the two axes of the rubric that decay over time — robustness under perturbation (Axis 05) and currency against the moving frontier of agent capability (Axis 06).
Robustness benchmark
How every pack holds its discriminating power under semantics-preserving perturbation of its own test items — reordered docs and tools, injected distractors. A pack keyed to the exact surface form of its frozen test set collapses here.
Currency benchmark
How current every pack is against the moving frontier of agent capability — decay from the dated cohort it was last validated against, not calendar age. A stale pack is capped out of a top grade until it's revalidated.
How to cite this data
Every benchmark run is reproducible: reference agents, perturbation recipes, and scoring harnesses are published with each result. Cite the benchmark page and its run ID — or read what a reference-panel harness is for the methodology behind the numbers.
From numbers to gates
These benchmarks back every grade on the marketplace.
Each verified pack carries its discriminating-power result, robustness score, and currency status — so you know what a grade means before you gate CI on it.