# When does a question need GraphRAG?

Canonical: https://shivanshsen.com/blogs/graphrag-relationships

Test the relationship question before paying to extract and maintain a graph.

![An archive keeper follows connections between artisans, workshops, and their source scrolls.](https://shivanshsen.com/illustrations/articles/graphrag-relationships-960.webp)

An archive keeper follows connections between artisans, workshops, and their source scrolls.

AI-generated contemporary Phad-inspired illustration; not traditional artisan authorship.

An error code gives search something exact to match. A dependency question asks for connections that may span several documents. Before adding GraphRAG, I want to know which of those jobs the reader needs done. A graph can help connect evidence. It can also preserve a mistaken connection and make it easy to repeat.

Use the synthetic ParcelSync example from this series. Its v2.4 consumer caches credentials; rotating an identity key can leave it reporting ERR-AUTH-217 until credentials reload. The v2.5 release notes say the consumer reloads them automatically. These are invented facts for explaining retrieval, not a production incident or a measured result.

## Change the question first

“What does ERR-AUTH-217 mean in v2.4?” asks for a specific explanation. Keyword and vector retrieval, combined through hybrid search, may find the runbook and release note. Hybrid search combines retrieval signals; GraphRAG adds derived relationship data and summaries. The two can coexist. A graph adds little if those passages already supply the answer.

Now ask: “Which consumers depend on the rotated identity key?” A connector configuration, a consumer inventory, and an identity mapping might each hold part of the answer. The useful structure is a chain of documented dependencies: connector to consumer to identity. The order and meaning of each relation matter more than whether the documents use similar words.

A graph could expose those connections for retrieval. But the two version notes alone cannot name every affected consumer. If the corpus lacks the consumer inventory, the answer must say so. A retrieved edge is a claim to inspect against its source; drawing it does not establish that it is true or current.

## Local questions and corpus-wide questions

Microsoft’s GraphRAG local search starts from relevant entities and gathers connected graph information alongside source text. For ParcelSync, a named identity could provide a starting point for finding consumers and checking the passages behind each dependency. The retrieved context still needs to fit the question and retain the relevant version.

[Microsoft GraphRAG local search](https://microsoft.github.io/graphrag/query/local_search/)

Global search answers a different kind of request. It works over generated community reports, which summarize groups in the graph, and combines intermediate answers. “What recurring causes appear across these incident reports?” needs coverage across the corpus. Retrieving a few passages close to the question may miss themes described elsewhere.

[Microsoft GraphRAG global search](https://microsoft.github.io/graphrag/query/global_search/)

A community summary can help surface such patterns, but it compresses the underlying records. I would check a reported pattern against the incidents that support it before presenting it as a finding. A summary of several incidents cannot establish the complete list of consumers affected by one key rotation.

DRIFT adds community context to local search and uses it to develop follow-up questions. That offers another way to explore a broad question before pursuing details. It also adds work at query time. Choosing it should follow from a question that benefits from that exploration.

[Microsoft GraphRAG DRIFT search](https://microsoft.github.io/graphrag/query/drift_search/)

## The index becomes another thing to maintain

Microsoft’s standard indexing pipeline extracts entities and relationships, finds communities, generates reports, and creates embeddings. Those steps add model calls and stored artifacts before a reader asks anything. The documents, extracted graph, and reports need a clear refresh policy when a source changes.

[Microsoft GraphRAG indexing overview](https://microsoft.github.io/graphrag/index/overview/)

In our example, an extractor might merge two identities with similar names, omit a consumer, or lose the version attached to a dependency. A community report might repeat that mistake. Keep source references for relations, inspect ambiguous entity matches, and test a v2.5 update against questions previously answered from v2.4. Updating the text alone is insufficient if retrieval still uses stale derived records.

## Give the graph a test it can fail

I would compare ordinary hybrid retrieval with graph-assisted retrieval on the same fixed corpus. Include exact-code questions, documented dependency chains, corpus-wide questions, and questions whose required relation is absent. Review the answer and supporting sources, missing consumers, unsupported edges, indexing cost, refresh effort, and query latency.

This series reports no GraphRAG benchmark. For ParcelSync, the first decision is whether we have reliable dependency evidence and enough relationship questions to justify maintaining it. If we do, the consumer-to-identity question gives us a concrete test. If we do not, extracting more edges will not supply the missing inventory.

[Start with the RAG pipeline failure](https://shivanshsen.com/blogs/rag-pipeline-failures)

[Match error codes and paraphrases with hybrid search](https://shivanshsen.com/blogs/hybrid-search-error-codes)

[Evaluate retrieval and the resulting answer](https://shivanshsen.com/blogs/evaluating-hybrid-rag)

## Sources

- https://microsoft.github.io/graphrag/query/local_search/

- https://microsoft.github.io/graphrag/query/global_search/

- https://microsoft.github.io/graphrag/query/drift_search/

- https://microsoft.github.io/graphrag/index/overview/
