Verified catalog · capability

Capability packs

Does the agent do the task well — RAG correctness, tool-calling, retrieval quality.

capability

Medical RAG — groundedness & abstention

Catches confident fabrication with fake citations. Scores groundedness, citation accuracy, and whether the agent abstains when evidence is missing.

RAG · medical literature · zero-hallucination
RAG groundednessMedicalHigh risk
EU AI Act — Art. 53 / high-risk
79 one-timeragas frameworkAG-26-0142
capability

Legal contract RAG — groundedness & abstention

Verifies a contract-QA agent cites the actual clause it answers from, and refuses to answer questions the contract doesn't cover.

RAG · contract review · medium-risk
RAG groundednessLegalMedium risk
79 one-timeragas frameworkAG-26-0143
capability

Financial reporting RAG — groundedness & abstention

Checks that an accounting-policy assistant grounds answers in the actual policy note cited, and doesn't fabricate figures for uncovered questions.

RAG · financial reporting · high-risk
RAG groundednessFinanceHigh risk
EU AI Act — Art. 53 / high-risk
89 one-timeragas frameworkAG-26-0144
capability

Customer-support RAG — groundedness & abstention

Verifies a support bot answers policy questions from the actual help-doc cited, and declines account-specific questions it has no evidence for.

RAG · customer support · low-risk
RAG groundednessSupportLow risk
49 one-timeragas frameworkAG-26-0145
capability

General-knowledge RAG — groundedness & abstention

A control cell: general-reference Q&A used to baseline groundedness and abstention scoring before applying the harness to specialized domains.

RAG · general reference · low-risk
RAG groundednessGeneralLow risk
29 one-timeragas frameworkAG-26-0146
capability

Insurance policy RAG — groundedness & abstention

Checks that a policy/claims assistant answers from the exact clause it cites — deductibles, limits, exclusions — and abstains on claim-specific questions it has no evidence for, instead of inventing coverage.

RAG · insurance policy & claims · high-risk
RAG groundednessInsuranceHigh risk
EU AI Act — Art. 53 / high-risk
89 one-timeragas frameworkAG-26-0147
capability

Public-sector benefits RAG — groundedness & abstention

Verifies a benefits-eligibility assistant grounds answers in the actual program rule cited — thresholds, deadlines, required documents — and declines case-specific questions it can't evidence.

RAG · public benefits & eligibility · high-risk
RAG groundednessPublic sectorHigh risk
EU AI Act — Art. 53 / high-risk
89 one-timeragas frameworkAG-26-0148
capability

HR policy RAG — groundedness & abstention

Verifies an employee-handbook assistant cites the actual policy section it answers from — leave, notice, benefits — and refuses personal HR questions it has no evidence for.

RAG · employee handbook & HR policy · medium-risk
RAG groundednessHRMedium risk
79 one-timeragas frameworkAG-26-0149
capability

Security-policy RAG — groundedness & abstention

Checks that an infosec-policy assistant grounds answers in the actual control cited — rotation windows, encryption standards, incident SLAs — and abstains on live operational questions it can't answer from policy.

RAG · security policy & controls · medium-risk
RAG groundednessCybersecurityMedium risk
79 one-timeragas frameworkAG-26-0150
capability

E-commerce policy RAG — groundedness & abstention

Verifies a shopping-support bot answers returns, shipping, and warranty questions from the actual policy cited, and declines order-specific questions (where's my package, why was I declined) it has no evidence for.

RAG · retail returns & shipping policy · low-risk
RAG groundednessE-commerceLow risk
49 one-timeragas frameworkAG-26-0151
capability

API-docs RAG — groundedness & abstention

Checks that a developer-docs assistant answers from the actual API reference cited — rate limits, status codes, auth — and abstains on environment-specific questions it can't answer from the docs.

RAG · API & developer documentation · low-risk
RAG groundednessDeveloper toolsLow risk
39 one-timeragas frameworkAG-26-0152
capability

Tool-calling correctness

Verifies function/tool selection and argument correctness, and checks the agent asks for clarification instead of guessing when a request is ambiguous.

Tool-calling · general · medium-risk
Tool-calling correctnessMedium risk
49 one-timepromptfoo frameworkAG-26-0143
capability

General RAG groundedness — draft submission

An early-draft submission that hasn't cleared verification yet — held back rather than approved for sale. The report below shows exactly which checks it didn't pass.

RAG · general reference · low-risk
RAG groundednessGeneralLow risk
Not for saleragas framework