Home/Benchmarks

The citeable asset

Original data on discriminating power.

How our verification packs perform against the six-axis rubric — per test method, showing how well each separates known-good from sabotaged agents (Axis 02), and per authority axis, showing how every pack ranks on robustness and currency (Axes 05–06).

Pack leaderboardre-ranked every run
1AE-commerce policy RAG0.97
2AEU AI Act0.97
3AAPI-docs RAG0.96
4ACustomer-support RAG0.95
verification bar
FGeneral RAG groundedness — draft0.50held
Weighted six-axis scorerun #128
6
Published benchmarks
6
Grading axes, every pack
3
Axes benchmarked in public
3×
Reference agents per run
The six-axis rubric — every pack, one weighted score3 of 6 axes benchmarked on this page
AXIS 01
Structural validity
0.24
AXIS 02ON THIS PAGE
Discriminating power
0.26
AXIS 03
Standard coverage
0.12
AXIS 04
Thoroughness
0.22
AXIS 05ON THIS PAGE
Robustness
0.10
AXIS 06ON THIS PAGE
Currency
0.06

From numbers to gates

These benchmarks back every grade on the marketplace.

Each verified pack carries its discriminating-power result, robustness score, and currency status — so you know what a grade means before you gate CI on it.