Original data · discriminating power
RAG groundedness benchmark
Discriminating-power results across every RAG-groundedness pack — same test method, different domains and risk levels, compared side by side.
Reference panel · known quality vs. pack score
General RAG groundedness — draft submission did not meet the verification bar and was held back rather than approved for sale — shown here for transparency alongside the packs that passed.
A good pack scores the known-good agent high and the sabotaged one near zero. That gap is the evidence the meter works — this is mutation testing applied to evals: does the pack catch the planted bug?
Methodology — How to measure RAG groundedness explains how this metric is scored and what these numbers mean.