Why an error code and a paraphrase need different searches
Combine lexical and dense retrieval, then inspect what reaches the reranker.

An error code gives a search engine something precise to match. A description such as “the connector stopped signing in after we changed its credentials” gives it meaning to interpret. I want both routes available when an engineer looks for a runbook. Combining them is useful only if we can inspect what each route contributes.
Consider ParcelSync, a fictional service used throughout this series. In its synthetic troubleshooting corpus, ERR-AUTH-217 means version 2.4 cached credentials after a rotation. Version 2.5 reloads them automatically. ERR-AUTH-218 concerns permission denied. A queue-backlog document provides another plausible distraction. These facts are authored examples, not reports from a deployed service.
Keep the identifier and the symptom
For “ParcelSync v2.4 ERR-AUTH-217 after rotation,” lexical retrieval has several useful terms. BM25 scores matching words using their frequency, rarity across the corpus, and document length. Check the analyzer: splitting an error code into common fragments can weaken the distinction between 217 and 218. Exact identifier fields can preserve that distinction.
Dense retrieval represents the query and passages as embedding vectors, then ranks their similarity. It can connect “changed its credentials” with “secret rotation” even when those phrases share few words. That is a reason to test it, not a promise that the correct passage will win. Similar troubleshooting language can also pull in the permission-denied document.
Elastic’s hybrid-search documentation
Hybrid search combines retrieval signals. Here I mean BM25 plus dense retrieval; other combinations exist. It returns candidates. A RAG pipeline can pass those candidates to a generator, as described in the companion post on retrieval and answer failures. GraphRAG adds graph-derived context and query strategies. Those choices can coexist with hybrid search.
When relationships help retrieval
Merge ranks without adding incompatible scores
A BM25 score and a vector similarity score have different meanings. Adding them directly makes their relative influence depend on their scales. Reciprocal rank fusion, or RRF, instead uses each document’s position in each retrieved list.
RRF(d) = sum over lists containing d of 1 / (c + rank(d))
rank starts at 1; a missing document contributes 0.
c is the rank constant, not the number of retrieved candidates.Suppose I supply two illustrative rankings: the credential-cache passage appears first in the lexical list and third in the dense list. With c = 60, its fused score is 1/61 + 1/63, about 0.032266. This is arithmetic on supplied ranks. It is not an embedding measurement or evidence that hybrid search outperformed either retriever.
RRF rewards a document that ranks well in several lists. Agreement can also promote a distractor. The rank constant changes how much rank differences matter, and the candidate window determines which documents can contribute. Avoid calling this configuration-free just because it does not require normalizing the original scores.
Elastic’s RRF formula and candidate-window reference
A reranker can only inspect what arrives
If each retriever returns five passages, fusion works on their union. A useful passage ranked sixth in both lists never reaches that union. A reranker cannot rescue it. Increasing candidate depth may expose it, but also adds work; measure that trade-off on actual queries.
A cross-encoder reranker reads each query and candidate passage together and assigns a relevance score. Apply it after retrieval to reorder a manageable candidate set. It can distinguish wording more closely than an independent embedding comparison, but its scores still need checking against labelled cases.
Sentence Transformers’ retrieve-and-rerank guide
Similarity does not settle the version
For a v2.4 incident, the v2.5 automatic-reload guidance can be a bad answer despite matching the topic. Preserve service and version metadata. Apply access restrictions consistently to both retrieval branches before exposing candidates. Decide explicitly whether the query needs current guidance or historical guidance; a date filter alone cannot make that decision.
Use the same passages and queries when comparing BM25, dense retrieval, fusion, and reranking. Label the required version, inspect the wrong error code, and include a question the corpus cannot answer. Keep retrieval quality separate from the generator’s ability to cite and use a passage.
How to evaluate the retrieval and answer together
This example reports no benchmark. Before choosing a configuration for ParcelSync, I would need the ranked passages for both query forms, the relevance labels, and the retrieval cost. The first check is concrete: did the v2.4 credential-cache guidance reach the candidate set?
Try the rank-fusion arithmetic
These two ranked lists are supplied examples for the fictional ParcelSync corpus. They are not search results from a model or a benchmark. Change the candidate window to see which contributions reach fusion.
Lexical example
- ERR-AUTH-217: credential cache, version 2.4
- ERR-AUTH-218: permission denied
- Credential reload, version 2.5
- Queue backlog
Dense example
- Credential reload, version 2.5
- Queue backlog
- ERR-AUTH-217: credential cache, version 2.4
- ERR-AUTH-218: permission denied
Constant 60, four candidates per list. Scores add 1 / (constant + rank) for each included list.
| Passage | Lexical contribution | Dense contribution | Sum |
|---|---|---|---|
| Credential reload, version 2.5 | 0.015873 | 0.016393 | 0.032266 |
| ERR-AUTH-217: credential cache, version 2.4 | 0.016393 | 0.015873 | 0.032266 |
| ERR-AUTH-218: permission denied | 0.016129 | 0.015625 | 0.031754 |
| Queue backlog | 0.015625 | 0.016129 | 0.031754 |
Read the supplied passages
ERR-AUTH-217: credential cache, version 2.4
ParcelSync 2.4 consumers cache credentials. After key rotation, ERR-AUTH-217 requires the consumer restart described in the 2.4 runbook.
Credential reload, version 2.5
ParcelSync 2.5 reloads rotated credentials automatically. Its advice does not establish the behavior of a 2.4 consumer.
ERR-AUTH-218: permission denied
Check the caller’s permission to access the resource. This is a different error from ERR-AUTH-217.
Queue backlog
A queue can accumulate work when consumers are slow. This passage provides no evidence about rotated credentials.
The highest fused score does not prove a passage answers the question. Check the error code, version, source, and permission to read it before assembling context.
