When a retrieval-augmented answer fails, I want to see the evidence that reached the model. Did the search miss the right document? Did the pipeline discard it? Did the model read it and still give the wrong advice? Each failure calls for a different change.

A useful starting point is a plain baseline: split documents into passages, index them, retrieve a small set for a question, then put those passages into the generator’s prompt. Save the retrieved IDs and final context alongside the answer. That trace makes the next decision inspectable.

The original RAG paper describes a trained architecture combining a sequence-to-sequence model with a dense retrieval index. A search-and-prompt application borrows the wider idea of generation with retrieved evidence; it does not reproduce that training recipe. I would keep that distinction explicit before comparing approaches.

Lewis and colleagues: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Keep one question in view

Consider a synthetic service called ParcelSync. Its version 2.4 runbook says ERR-AUTH-217 can occur when a consumer retains cached credentials after secret rotation. For that version, the documented remedy is to restart the consumer. Version 2.5 reloads credentials automatically. These are invented documents for reasoning through retrieval, not production guidance or measured results.

The question is: “ParcelSync 2.4 started returning ERR-AUTH-217 after we rotated the secret. What should I check?” A useful answer needs the matching error, the matching version, the cause, and the supported action. A passage about version 2.5 may share most of those words while supporting different advice.

Keep those documents, their version metadata, and the expected supporting passages fixed across comparisons. Add distractors about queue backlogs and unrelated authentication errors. Then give the same underlying problem several phrasings. This makes vocabulary changes visible without quietly changing the task.

Rewrite when the question is incomplete

A follow-up such as “Does that apply after rotation?” lacks the service, version, and earlier error. A rewriter can use conversation context to make a standalone search question. Ma and colleagues study this adaptation of the query before retrieval, including a learned rewriter; their results support that approach on the tasks they tested.

Ma and colleagues: Query Rewriting for Retrieval-Augmented Large Language Models

For ParcelSync, check that rewriting preserves 2.4, ERR-AUTH-217, and secret rotation. A fluent rewrite that drops the version can retrieve the wrong runbook. Log the rewritten query so a reviewer can locate that mistake. Rewriting is useful only if it preserves what the user asked.

Use several queries for several phrasings

Multi-query retrieval searches several variants and combines their results. One variant might retain the exact error; another might ask about stale consumer credentials. This can expose passages that one wording missed. It also adds searches and can introduce irrelevant candidates.

Generating variants, merging duplicate documents, and combining ranked lists are distinct steps. LangChain’s MultiQueryRetriever documents the multiple-search pattern. A fusion method needs its own rule for ordering the combined candidates. Keep the original query available, and inspect whether generated variants invent a cause before evidence supports it.

LangChain: MultiQueryRetriever reference

Rerank evidence already found

Suppose the 2.4 runbook reaches the first twenty candidates but misses the five passages sent to the generator. A reranker can score each query–passage pair and reorder that pool. Sentence Transformers documents this two-stage arrangement: efficient retrieval followed by a cross-encoder that reads the query and candidate together.

Sentence Transformers: Retrieve and Re-Rank

The reranker cannot recover a document absent from its input. Check candidate recall first: did the needed evidence enter the pool? Then check its final position. More candidates create more scoring work, so record latency and context size alongside any quality change.

Check the search and the answer separately

Retrieval quality asks whether the selected passages support the task. Answer quality asks whether the generator uses that support correctly and completely. RAGAS treats context relevance, faithful use of context, and generation quality as separate dimensions. A correct-looking restart recommendation with only a 2.5 citation should fail our evidence check.

RAGAS: Automated Evaluation of Retrieval Augmented Generation

Hold the generator, prompt, document snapshot, and final context budget constant. Compare the baseline with one change at a time. Add a run with no retrieved evidence and another supplied with the known supporting passages. If the second still fails, improving search alone will not settle the problem.

This example proposes checks; it reports no experiment. I would start with the missing passage or unsupported instruction visible in the trace, then test the smallest change that could repair it.

Next: compare hybrid search on error codes and paraphrases

Measure the trade-offs: evaluating hybrid RAG

When the question depends on relationships: GraphRAG