The verified eval marketplace

Buy evals you can
actually trust.

A marketplace of AI-agent eval packs — where nothing gets listed until it passes the authority. Every pack carries a grade, a discriminating-power benchmark, and a public audit trail.

How a pack gets here
1
Submit
Anyone uploads a pack — cases, harness, and claimed standard.
2
Validate
We test it on six axes: structural validity, discriminating power, standard coverage, thoroughness, robustness, and currency.
3
List & earn
Pass, and it's listed with its grade. You set the price; we handle the sale.
14
Verified packs for sale
6
Verification axes, every pack
1
In review — held from sale
100%
Listed packs pass the authority

All-access pass

Every verified pack, one subscription.

Unlimited access to all verified packs and unlimited runs — plus every new pack the day it clears the authority. Best for teams gating more than two evals in CI.

  • All 14 packs included
  • Unlimited runs
  • New packs auto-added
150/mo
billed monthly · cancel anytime
15 packs
Capability
B

Legal contract RAG — groundedness & abstention

Verifies a contract-QA agent cites the actual clause it answers from, and refuses to answer questions the contract doesn't cover.

LegalMedium riskragas
Discriminating power100%
AGAgentGrading LabsFirst-party · 15 cases
79one-time · or all-access
Included in All-access
Capability
B

Insurance policy RAG — groundedness & abstention

Checks that a policy/claims assistant answers from the exact clause it cites — deductibles, limits, exclusions — and abstains on claim-specific questions it has no evidence for, instead of inventing coverage.

InsuranceHigh riskEU AI Actragas
Discriminating power100%
AGAgentGrading LabsFirst-party · 15 cases
89one-time · or all-access
Included in All-access
Capability
B

Public-sector benefits RAG — groundedness & abstention

Verifies a benefits-eligibility assistant grounds answers in the actual program rule cited — thresholds, deadlines, required documents — and declines case-specific questions it can't evidence.

Public sectorHigh riskEU AI Actragas
Discriminating power100%
AGAgentGrading LabsFirst-party · 15 cases
89one-time · or all-access
Included in All-access
Capability
A

HR policy RAG — groundedness & abstention

Verifies an employee-handbook assistant cites the actual policy section it answers from — leave, notice, benefits — and refuses personal HR questions it has no evidence for.

HRMedium riskragas
Discriminating power100%
AGAgentGrading LabsFirst-party · 15 cases
79one-time · or all-access
Included in All-access
Capability
A

E-commerce policy RAG — groundedness & abstention

Verifies a shopping-support bot answers returns, shipping, and warranty questions from the actual policy cited, and declines order-specific questions (where's my package, why was I declined) it has no evidence for.

E-commerceLow riskragas
Discriminating power100%
AGAgentGrading LabsFirst-party · 15 cases
49one-time · or all-access
Included in All-access
Capability
B

Tool-calling correctness

Verifies function/tool selection and argument correctness, and checks the agent asks for clarification instead of guessing when a request is ambiguous.

Cross-domainMedium riskpromptfoo
Discriminating power100%
AGAgentGrading LabsFirst-party · 15 cases
49one-time · or all-access
Included in All-access
Capability
B

Financial reporting RAG — groundedness & abstention

Checks that an accounting-policy assistant grounds answers in the actual policy note cited, and doesn't fabricate figures for uncovered questions.

FinanceHigh riskEU AI Actragas
Discriminating power99%
AGAgentGrading LabsFirst-party · 15 cases
89one-time · or all-access
Included in All-access
Capability
B

General-knowledge RAG — groundedness & abstention

A control cell: general-reference Q&A used to baseline groundedness and abstention scoring before applying the harness to specialized domains.

GeneralLow riskragas
Discriminating power99%
AGAgentGrading LabsFirst-party · 15 cases
29one-time · or all-access
Included in All-access
Capability
A

Security-policy RAG — groundedness & abstention

Checks that an infosec-policy assistant grounds answers in the actual control cited — rotation windows, encryption standards, incident SLAs — and abstains on live operational questions it can't answer from policy.

CybersecurityMedium riskragas
Discriminating power99%
AGAgentGrading LabsFirst-party · 15 cases
79one-time · or all-access
Included in All-access
Conformance
A

EU AI Act — high-risk conformance

Evidence-based checklist for high-risk obligations: logging, human oversight, transparency, data governance, and conformity assessment. Anchored to the Act, not to opinion.

Cross-domainHigh riskEU AI ActISO 42001deepeval
Discriminating power96%
AGAgentGrading LabsFirst-party · 8 cases
99one-time · or all-access
Included in All-access
Capability
A

Customer-support RAG — groundedness & abstention

Verifies a support bot answers policy questions from the actual help-doc cited, and declines account-specific questions it has no evidence for.

SupportLow riskragas
Discriminating power94%
AGAgentGrading LabsFirst-party · 15 cases
49one-time · or all-access
Included in All-access
Capability
A

API-docs RAG — groundedness & abstention

Checks that a developer-docs assistant answers from the actual API reference cited — rate limits, status codes, auth — and abstains on environment-specific questions it can't answer from the docs.

Developer toolsLow riskragas
Discriminating power94%
AGAgentGrading LabsFirst-party · 15 cases
39one-time · or all-access
Included in All-access
Capability
B

Medical RAG — groundedness & abstention

Catches confident fabrication with fake citations. Scores groundedness, citation accuracy, and whether the agent abstains when evidence is missing.

MedicalHigh riskEU AI Actragas
Discriminating power88%
AGAgentGrading LabsFirst-party · 15 cases
79one-time · or all-access
Included in All-access
Safety
B

Browser agent — prompt-injection red-team

Adversarial web content that tries to make a computer-use agent exfiltrate data or take destructive actions. The test set is the attack, not a Q&A.

Cross-domainHigh riskOWASPpromptfoo
Discriminating power83%
AGAgentGrading LabsFirst-party · 15 cases
89one-time · or all-access
Included in All-access
Capability Failed verification
F

General RAG groundedness — draft submission

An early-draft submission that hasn't cleared verification yet — held back rather than approved for sale. The report below shows exactly which checks it didn't pass.

GeneralLow riskragas
Discriminating power50%
?Community submissionRejected · 2 cases
Not for salesubmission rejected
Withheld from the marketplaceView failure report

For pack authors

Built an eval that works? List it and earn.

Turn a battle-tested eval pack into revenue. Submit once — if it clears the six-axis verification, it's listed on the marketplace with a grade buyers trust. You keep the majority of every sale.

Set your own price

You price it; we take a flat marketplace fee. Payouts on every sale.

The grade is the trust

Buyers see your grade, benchmark and audit trail — not just a description.

Fail fast, iterate

Don't pass? You get the exact failing checks back — fix and resubmit free.