# Does better retrieval produce a better answer?

Canonical: https://shivanshsen.com/blogs/evaluating-hybrid-rag

Compare retrieved evidence, supported answers, citations, and the cost of each improvement.

![A reviewer checks an answer against source scrolls beside a sand timer.](https://shivanshsen.com/illustrations/articles/evaluating-hybrid-rag-960.webp)

A reviewer checks an answer against source scrolls beside a sand timer.

AI-generated contemporary Phad-inspired illustration; not traditional artisan authorship.

A search result can contain the right paragraph while the answer still gives the wrong instruction. If I change a retriever and see more relevant documents, I have evidence about retrieval. I still need to check what the generator tells the reader.

Use ParcelSync, a synthetic support corpus shared across this series. Its version 2.4 guide says ERR-AUTH-217 can follow credential rotation because a consumer keeps cached credentials; restarting that consumer reloads them. Version 2.5 reloads credentials automatically. ERR-AUTH-218 concerns permissions. These are invented documents for explaining a test, not reports from a deployed product.

## Fix the comparison before choosing a winner

Compare keyword search, vector search, hybrid search, and hybrid search with reranking against the same corpus snapshot and questions. Keep the generator model, prompt, generation settings, and context token budget fixed. A retriever that gets twice the context allowance has a different advantage from one that finds better evidence.

Record the final context after filtering, deduplication, and truncation. A document in the candidate list cannot support an answer if the prompt never contains it. Keep the chunking policy fixed for this comparison; test chunking changes in a separate experiment.

Write the expected evidence and acceptable answer before running the models. Include an exact error-code question, a paraphrase about failed authentication after rotation, a version-specific question, and a permissions question. Add a question whose answer the corpus does not contain, such as the service’s credential-rotation SLA.

## Score retrieval and answers separately

For retrieval, label which passages support each question, then measure how many appear within the context budget. Recall at a fixed cutoff shows whether search found the labelled evidence. A ranking metric can show whether that evidence appears early. Inspect false matches too: ERR-AUTH-218 shares vocabulary with ERR-AUTH-217 but supports a different diagnosis.

For the answer, check each material claim against the supplied context. Does it preserve the requested version? Does it distinguish cached credentials from permissions? Does it recommend a restart only where the guide supports one? An answer can faithfully repeat a version 2.4 passage and still fail a question about version 2.5. Context faithfulness and task correctness need separate judgments.

The Ragas paper separates relevant context, faithful use of that context, and generation quality. That separation is useful here: it lets a retrieval improvement and a generation failure appear in the same report.[ Ragas paper.](https://arxiv.org/abs/2309.15217)

Check citations at the claim level. Confirm that each cited document exists, that its passage supports the attached claim, and that material claims have support. A valid URL alone earns no credit for a recommendation. The FActScore paper evaluates factual precision through individual facts; I would borrow that claim-by-claim discipline without treating its biography benchmark as a support-system benchmark.[ FActScore paper.](https://arxiv.org/abs/2305.14251)

For the missing SLA, the acceptable answer states that the supplied documents do not establish one. Count invented values and unsupported assurances as failures. Also count unnecessary refusals on answerable questions. A generator that refuses every question can avoid unsupported claims while doing no useful support work.

## Use controls to locate the failure

Run an oracle-context control: give the same generator the passages a reviewer selected as sufficient evidence. If it still recommends restarting a version 2.5 consumer, retrieval cannot explain that failure. The prompt, generation behavior, or answer rubric needs attention. Oracle context is a diagnostic control, not a realistic search result.

Run a closed-book control with no retrieved context. It shows how the generator answers without your documents. In a synthetic corpus, a confident explanation of ERR-AUTH-217 deserves scrutiny because the model has no supplied basis for that product-specific claim. Compare support and correctness under the same rubric.

## Keep the cost of improvement visible

Store each question, ranked candidates, final context, answer, citations, grader decisions, model versions, and run settings. Record retrieval and reranking time separately from generation time. Report end-to-end latency, token use, and cost per request, with the pricing date and the accounting method.

Repeat generation runs and retain failures. Review close calls and a sample of automated judgments against human labels. Keep some questions outside the tuning loop; once a question guides a revision, it no longer serves as an unseen check.

This is an evaluation design. I have not run the benchmark or measured a hybrid-search improvement. A useful result would show which questions gained supported answers, which still failed, and the extra time and cost. Until those runs exist, I would keep every performance claim out of the post.

Trace the failure from search to prompt to answer in the first post.[ Which part of your RAG pipeline is failing?](https://shivanshsen.com/blogs/rag-pipeline-failures)

Compare retrieval methods on exact identifiers and paraphrases.[ Why an error code and a paraphrase need different searches.](https://shivanshsen.com/blogs/hybrid-search-error-codes)

Apply the same controls when testing relationship-based retrieval.[ When does GraphRAG help?](https://shivanshsen.com/blogs/graphrag-relationships)

## Sources

- https://arxiv.org/abs/2309.15217

- https://arxiv.org/abs/2305.14251
