How to Test Whether a RAG Answer Is Grounded
Evaluate the retrieved evidence and the final answer as separate steps.
On this page
A fluent answer can still be unsupported. Build a test set with questions, expected source passages, acceptable answers, and explicit 'do not answer' cases. Review whether the system retrieved the right material, used it accurately, and respected permissions.
Label retrieval
For each test question, identify the document section that should be found. Record missed passages and irrelevant results before blaming the generation step.
Check answer support
Compare every material claim with the retrieved passages. A citation that points to a document without supporting the sentence is a failure, not a pass.
Include negative cases
Test outdated documents, conflicting versions, missing permissions, ambiguous questions, and information absent from the corpus. The safe outcome may be an explicit inability to answer.
Build a small evidence ledger
For every test question, name the passage that should support the answer, its version, and the claim the answer may safely make. Include questions with several relevant passages, ambiguous wording, and no supporting material. Run retrieval alone first to see whether the right evidence appears at all. Then inspect each generated claim against exact sentences rather than accepting a citation that merely points to the right file. Keep access rules in the test: a passage that exists but is not permitted for the requester should not count as a successful retrieval.
Use an evidence record for every evaluation case. Store the question, requester permission, expected passage, acceptable claim, required citation, and whether abstention is correct. After a run, record the retrieved passages and the exact answer sentence that each supports. Give a separate mark for retrieval success and answer support. Add a short failure label such as missed source, stale version, unsupported inference, or permission leak. A stable record lets the team compare a document refresh with a model change without confusing their effects. It also makes a reviewer able to explain why an answer was accepted or rejected.
Use failures to choose the repair
A missing passage suggests extraction, indexing, metadata, or query work. A good passage paired with an unsupported answer suggests prompt, model, or answer validation work. A correct answer from an obsolete policy is a version-control failure. Record these separately and route each to the responsible owner. Re-run a stable set after document updates and model changes, and sample production questions under privacy controls to find new cases. The release criterion should include safe abstention and accurate citations, not only the percentage of questions that received a fluent response.
Decision checklist
- Keep a stable evaluation set for releases.
- Score retrieval and answer support separately.
- Inspect citations against exact passages.
- Review failure patterns with the content owner.
A small test before committing
Create a release set with answerable questions, absent answers, conflicting documents, and a user lacking permission. For each answer, mark the exact sentence or table cell that supports every material claim. Score retrieval separately from final response, and treat a citation to an irrelevant document as unsupported. Keep the same set for each release so improvements are comparable. When a source changes, update expected evidence before trusting the new score.
Worked scenario
Suppose a hypothetical assistant is asked for a warranty period but the retrieved document contains only an installation date. A grounded system should say the warranty term was not found, even if a typical industry period seems plausible.
For a scoped application of this decision, see RAG & Knowledge Systems.