# Why an error code and a paraphrase need different searches

Canonical: https://shivanshsen.com/blogs/hybrid-search-error-codes

Combine lexical and dense retrieval, then inspect what reaches the reranker.

![Two archive keepers search by seal and subject, then compare their selected scrolls.](https://shivanshsen.com/illustrations/articles/hybrid-search-error-codes-960.webp)

Two archive keepers search by seal and subject, then compare their selected scrolls.

AI-generated contemporary Phad-inspired illustration; not traditional artisan authorship.

An error code gives a search engine something precise to match. A description such as “the connector stopped signing in after we changed its credentials” gives it meaning to interpret. I want both routes available when an engineer looks for a runbook. Combining them is useful only if we can inspect what each route contributes.

Consider ParcelSync, a fictional service used throughout this series. In its synthetic troubleshooting corpus, ERR-AUTH-217 means version 2.4 cached credentials after a rotation. Version 2.5 reloads them automatically. ERR-AUTH-218 concerns permission denied. A queue-backlog document provides another plausible distraction. These facts are authored examples, not reports from a deployed service.

## Keep the identifier and the symptom

For “ParcelSync v2.4 ERR-AUTH-217 after rotation,” lexical retrieval has several useful terms. BM25 scores matching words using their frequency, rarity across the corpus, and document length. Check the analyzer: splitting an error code into common fragments can weaken the distinction between 217 and 218. Exact identifier fields can preserve that distinction.

Dense retrieval represents the query and passages as embedding vectors, then ranks their similarity. It can connect “changed its credentials” with “secret rotation” even when those phrases share few words. That is a reason to test it, not a promise that the correct passage will win. Similar troubleshooting language can also pull in the permission-denied document.

[Elastic’s hybrid-search documentation](https://www.elastic.co/docs/solutions/search/hybrid-search)

Hybrid search combines retrieval signals. Here I mean BM25 plus dense retrieval; other combinations exist. It returns candidates. A RAG pipeline can pass those candidates to a generator, as described in the companion post on retrieval and answer failures. GraphRAG adds graph-derived context and query strategies. Those choices can coexist with hybrid search.

[Where a RAG pipeline fails](https://shivanshsen.com/blogs/rag-pipeline-failures)

[When relationships help retrieval](https://shivanshsen.com/blogs/graphrag-relationships)

## Merge ranks without adding incompatible scores

A BM25 score and a vector similarity score have different meanings. Adding them directly makes their relative influence depend on their scales. Reciprocal rank fusion, or RRF, instead uses each document’s position in each retrieved list.

```text
RRF(d) = sum over lists containing d of 1 / (c + rank(d))
rank starts at 1; a missing document contributes 0.
c is the rank constant, not the number of retrieved candidates.
```

Suppose I supply two illustrative rankings: the credential-cache passage appears first in the lexical list and third in the dense list. With c = 60, its fused score is 1/61 + 1/63, about 0.032266. This is arithmetic on supplied ranks. It is not an embedding measurement or evidence that hybrid search outperformed either retriever.

RRF rewards a document that ranks well in several lists. Agreement can also promote a distractor. The rank constant changes how much rank differences matter, and the candidate window determines which documents can contribute. Avoid calling this configuration-free just because it does not require normalizing the original scores.

[Elastic’s RRF formula and candidate-window reference](https://www.elastic.co/docs/reference/elasticsearch/rest-apis/reciprocal-rank-fusion)

## A reranker can only inspect what arrives

If each retriever returns five passages, fusion works on their union. A useful passage ranked sixth in both lists never reaches that union. A reranker cannot rescue it. Increasing candidate depth may expose it, but also adds work; measure that trade-off on actual queries.

A cross-encoder reranker reads each query and candidate passage together and assigns a relevance score. Apply it after retrieval to reorder a manageable candidate set. It can distinguish wording more closely than an independent embedding comparison, but its scores still need checking against labelled cases.

[Sentence Transformers’ retrieve-and-rerank guide](https://www.sbert.net/examples/sentence_transformer/applications/retrieve_rerank/README.html)

## Similarity does not settle the version

For a v2.4 incident, the v2.5 automatic-reload guidance can be a bad answer despite matching the topic. Preserve service and version metadata. Apply access restrictions consistently to both retrieval branches before exposing candidates. Decide explicitly whether the query needs current guidance or historical guidance; a date filter alone cannot make that decision.

Use the same passages and queries when comparing BM25, dense retrieval, fusion, and reranking. Label the required version, inspect the wrong error code, and include a question the corpus cannot answer. Keep retrieval quality separate from the generator’s ability to cite and use a passage.

[How to evaluate the retrieval and answer together](https://shivanshsen.com/blogs/evaluating-hybrid-rag)

This example reports no benchmark. Before choosing a configuration for ParcelSync, I would need the ranked passages for both query forms, the relevance labels, and the retrieval cost. The first check is concrete: did the v2.4 credential-cache guidance reach the candidate set?

## Sources

- https://www.elastic.co/docs/solutions/search/hybrid-search

- https://www.elastic.co/docs/reference/elasticsearch/rest-apis/reciprocal-rank-fusion

- https://www.sbert.net/examples/sentence_transformer/applications/retrieve_rerank/README.html
