GraphRAG in Practice: What It's Actually Good At

RAG Architecture Series · Part 4

GraphRAG in Practice

Part 1 said most GraphRAG builds are expensive ways to store the same chunks in a worse format. This is the argument in full, plus the cases where a graph is the only thing that works. The hard parts are all at index time.

Reviewed August 2026 / Developer track / 25 min read

I want to open with the failure mode nobody writes postmortems about, because nothing exploded.

A team ships a graph. The demo is excellent, and I mean genuinely excellent, the kind where someone in the room says "wait, do that again." Four months later the graph looks different. There are four nodes for the same supplier, spelled four ways. There's one node with an edge count that makes no physical sense, sitting in the middle of everything like a black hole. A community summary confidently describes a document that was deleted in March. Nobody filed an incident, because nothing failed. Someone quietly put a router in front of it, sent the boring queries to hybrid search, watched the numbers not get worse, and after that the graph was a line item that got renewed twice out of politeness.

I've watched versions of this more than once and I'm deliberately describing it as a pattern rather than pinning it on one company, because the details change and the shape doesn't. The graph was fine on day one. Keeping it true was the job nobody scoped.

That's the whole piece, really, and you can stop here if you want. Everything below is the mechanics of why keeping it true is hard, which parts of the pipeline eat the budget, and what I'd ask a team to show me before I believed their graph was earning its keep.

What this isn't: a tutorial. I'm not going to hand you a config file. The frameworks change every quarter, the parameter names will have moved by the time you read this, and the copy-paste version of this article would be worse than useless because it would look authoritative while being stale. Part 1 was the architecture decision, aimed at the people who sign the invoice. This one is the anatomy, aimed at whoever has to keep the thing alive afterwards.

02 · Three jobs

One word, three jobs

The single most expensive misunderstanding in this space is that "GraphRAG" names one thing. It names at least three, and they have different cost curves, different failure modes, and different reasons to exist.

Local entity lookup. You have a question about a specific thing. The graph gives you that entity plus its neighbourhood, which is often better context than whatever the vector index would have handed you, because it's assembled by relationship rather than by surface similarity.

Path reasoning. The answer requires connecting A to B through C, and no single chunk contains the connection. This is what most people mean when they say "multi-hop," and it's what most people are actually shopping for.

Corpus sensemaking. What are the recurring themes here? What changed between these two years of filings? Who keeps appearing next to whom? No chunk contains the answer because the answer isn't in any chunk, it's a property of the collection.

FIG.01 — ONE LABEL, THREE JOBS LOCAL LOOKUP QUESTION "what do we know about X" NEEDS entity linking, 1 to 2 hops INDEX COST graph, no summaries PATH REASONING QUESTION "how does X connect to Y" NEEDS correct edges, correct merges INDEX COST graph, resolution critical SENSEMAKING QUESTION "what are the themes here" NEEDS communities and summaries INDEX COST the full pipeline MOST TEAMS BUY THE THIRD AND POINT IT AT THE SECOND
FIG.01 Three different jobs share one word. Pick which one you are buying, out loud, before anyone opens a terminal.

Microsoft's original paper is explicitly aimed at the third one. It frames global questions over an entire corpus as a query-focused summarization problem rather than a retrieval problem, and reports its gains on datasets in the million-token range against comprehensiveness and diversity of the answers1. Read that sentence twice, because comprehensiveness and diversity are not accuracy, and the paper never pretended otherwise. The industry did that part on its own.

So when a team tells me GraphRAG underperformed, my first question is which job they bought it for. Usually they bought sensemaking machinery and pointed it at path reasoning, or worse, at lookups.

The independent evaluations back this up more cleanly than the marketing does. A systematic comparison from Michigan State and Meta ran plain RAG against four GraphRAG families under one protocol, same chunking, same embeddings, same generation, and found no universal winner: plain RAG came out slightly ahead on single-hop factual lookup, 64.8 F1 against 63.0 for the best graph method, while graph-guided retrieval led on multi-hop at 70.3 against 67.02. Those are not the numbers of a revolution. They're the numbers of a specialized tool.

And the benchmarks themselves are shaky in a specific way. GraphRAG-Bench argues that the standard evaluation sets, HotpotQA and MultiHopRAG and UltraDomain among them, don't actually isolate what the graph structure contributes, because of how the questions and corpora are built3. Which means the honest position for most of us is: the published numbers tell you less about your corpus than you'd like, and you're going to have to measure this yourself.

03 · The pipeline

Drawn with the costs on it

Seven stages, and it's worth internalizing where they sit relative to your invoice.

  1. Chunk for extraction
  2. Extract entities and relations
  3. Resolve entities against each other
  4. Construct the graph
  5. Cluster it into communities
  6. Summarize those communities
  7. Embed everything for retrieval

Now the part that reorganizes how you think about this: in the standard pipeline, every one of those seven happens before anyone asks a question. Not most of them. All of them, embedding included14. GraphRAG is an index-time architecture wearing a retrieval-time costume, and query-time traversal, the part every demo shows you, the part with the pretty node animation, isn't on that list at all.

You don't have to take my word for how heavy the index side is, because the people who built it went back and attacked it twice. Microsoft's own follow-up, LazyGraphRAG, defers summarization to query time and reports data indexing costs identical to vector RAG, which is to say 0.1% of full GraphRAG, while matching global answer quality at more than 700 times lower query cost in their configuration4. When the original authors ship something that undercuts their own indexing pipeline by three orders of magnitude, that is not a scandal. It's a confession about where the weight was, and it's the most useful single data point in this entire article.

FIG.02 — SEVEN STAGES, ALL BEFORE THE QUESTION STAGE MODEL COST FAILURE RISK 01 chunk for extraction 02 extract entities and relations 03 resolve entities ← 04 construct the graph 05 cluster into communities 06 summarize communities 07 embed for retrieval
FIG.02 Model cost and failure risk do not sit on the same stages. The arrow marks stage 03, which is cheap in tokens and expensive in everything else.

So the mental model I'd hold: extraction and summarization are the two model-heavy stages and they scale with corpus size, not query volume. Resolution is the stage that's cheap in tokens and expensive in engineering time. Clustering is cheap in both and dangerous anyway, for reasons in section 06. And every one of them runs again, in full, the day you decide to re-index.

Aside

I keep meeting teams who have never actually costed a full re-index, and treat it as a maintenance task rather than a purchase decision. It's a purchase decision. Put the number on a slide before you need it, because the meeting where you find out is not a good meeting.

04 · Extraction

What the extractor gets wrong

Extraction is where every tutorial spends its time, which is roughly inverse to where the difficulty lives. But it does have real failure modes, and they're recognizable on sight once you know the shapes.

Open extraction gives you a vocabulary problem

Let the model name its own relation types and you'll finish with two hundred predicates where you wanted twelve: supplies, supplies_to, is_supplier_of, provides_components_for, all pointing at the same fact. Nothing downstream can traverse that. Your queries either miss two thirds of the relevant edges or you write a mapping table by hand afterwards, which is the same work you avoided, done later, with less context.

Close the list. Six lines of schema, and the constraint is the entire point:

class Relation(BaseModel):
    source: str
    target: str
    predicate: Literal["supplies", "acquired", "employs",
                       "located_in", "supersedes", "references"]
    valid_from: date | None
    valid_to: date | None

Twelve predicates you chose beats two hundred the model invented

If you can't enumerate your predicates, that's worth knowing before you index a corpus, not after.

Granularity is a real tuning axis, not a default

The original paper's own experiments found that a 600-token extraction chunk pulled out roughly twice as many entities as a 2400-token chunk, with additional gleaning passes as a second axis on top1. Which is the sort of finding that should make you nervous about anyone quoting graph statistics without saying what chunk size produced them. Note also that your extraction chunks and your retrieval chunks want different things: extraction wants enough surrounding context to get relation direction right, retrieval wants to fit an embedding window. Reusing one for the other is a decision, and mostly an unexamined one.

The four errors worth looking for by hand

Direction inversion, where the model records that the supplier was acquired by the customer. Dropped negation, where "the board did not approve the merger" arrives in your graph as an approved edge, cheerfully, with no hedge. Confidence flattening, where "sources suggest" and "the filing states" become the same edge with the same weight. And bridging, where the model supplies a relation that makes a paragraph tidier and was never asserted by anyone.

Only one of those is visible in aggregate statistics. The other three you find by reading a hundred extracted triples next to their source sentences, which takes an afternoon and which almost nobody does before shipping.

Temporal validity is not a later feature

If your edges don't carry a validity window, your graph answers "who is the CEO" with a superposition of everyone who has ever held the role. Retrofitting this is a re-extraction of the entire corpus, so decide at the start. The design worth stealing here is bi-temporal: Graphiti tracks when a fact was true in the world separately from when the system learned it, and when new information contradicts an old edge it closes the old edge's validity window instead of deleting it5. Two timelines, and superseded facts are invalidated rather than discarded. That distinction sounds academic until an auditor asks what your system believed last March, at which point it's the whole ballgame.

05 · Resolution

Entity resolution is the project

Here's the section I'd keep if I could only keep one.

Everything above assumes that when the extractor emits "ACME Corp" and "ACME Corporation" and "Acme" and "the supplier," something downstream figures out those are one node. That something is entity resolution, and it is the difference between a knowledge graph and an expensive pile of unconnected assertions.

It's also not new, and I want to be pointed about that, because the LLM era has a habit of rediscovering solved problems with worse tools. There are survey papers on end-to-end entity resolution covering blocking, matching, and batch versus incremental workflows6. There are dedicated surveys just on the blocking stage7. There's a standard textbook that predates all of this by more than a decade8. If you're building a graph from prose, you have inherited a mature subfield, and the correct move is to read a survey rather than to invent a similarity threshold on a Tuesday.

The shape of the problem, briefly, because the details are in those references and I'd rather spend the space on the parts that bite.

Blocking is candidate generation: you cannot compare every pair, so you produce plausible pairs cheaply, by shared trigrams, minhash, or nearest neighbours in embedding space. Cheap and high recall, because anything not proposed here is never merged.

Pairwise scoring decides whether a proposed pair is the same thing, and the useful signal is never just string distance. Type compatibility, shared neighbours in the graph, overlap in the source documents they came from, all of it matters, and shared neighbours in particular does work that string similarity cannot.

Clustering turns pairwise decisions into groups, and this is where it gets genuinely nasty, because pairwise similarity is not transitive. A matches B, B matches C, and A absolutely does not match C. Naive connected components will happily merge all three, and then keep going. That's how you get one node called something like "Systems" that has absorbed nine unrelated organizations and now sits at the centre of your graph radiating false connections.

Which brings me to the opinion I'll defend hardest in this whole article: over-merging is worse than under-merging, and it isn't close.

A missed merge costs you recall. The answer is incomplete, a path doesn't connect, someone gets a partial response. Annoying, bounded, and visible in evaluation.

An incorrect merge corrupts three things at once. Path reasoning now finds routes that don't exist in the world, and states them with the same confidence as real ones. Community detection reshapes itself around a node that isn't real, so your clusters are wrong. And the community summaries built on top of those clusters are wrong in a way that reads perfectly fluent, because the summarizer had no way to know the node was fictional. One bad merge, three corrupted layers, and none of it announces itself.

So tune asymmetrically. Set your merge threshold where you'd rather miss than merge. Then route the high-consequence cases, the ones where the merge would create a high-degree node, into a human review queue, because those are exactly the merges that do the most damage when wrong and there are far fewer of them than you fear.

And run the cheapest diagnostic in the entire pipeline, which is three lines and a plot:

degrees = [d for _, d in graph.degree()]   # or a GROUP BY over your edge table
top = sorted(degrees, reverse=True)[:20]
print(top, statistics.median(degrees))

Top node three orders of magnitude above median? That is not a hub

Real graphs have hubs. Real hubs are explicable: you can name the entity and say why everything touches it. If your top node is orders of magnitude above the median and nobody in the room can explain what it is, you're looking at a merge bug, and you were about to ship it.

Aside

The merge review queue I've had the most success with was, and I'm sorry about this, a spreadsheet. Candidate pairs, sorted by resulting degree, one human column. Same shape as the retrieval debugging sheet from Part 1. I've stopped being embarrassed about this and started just budgeting for it.

06 · Communities

Communities move under you

Once you have a graph, you partition it into communities, summarize each one, and those summaries become the material that answers corpus-level questions. Three things about this stage are worth knowing before you rely on it.

Use Leiden, and know why. The Louvain algorithm was the default for years and it has a defect that went largely unnoticed: it can produce communities that are badly connected, and in the worst case internally disconnected. The paper introducing Leiden measured up to 25% of communities badly connected and up to 16% disconnected on real networks, and proves that Leiden's output is guaranteed connected9. A disconnected community is not a subtle statistical wobble. It's a cluster containing two groups of entities that have nothing to do with each other, which you are about to hand to a language model and ask for a coherent summary. It will produce one. That's the problem.

The resolution parameter is a product decision wearing a maths costume. Higher resolution gives you more, smaller communities; lower gives you fewer, broader ones9. Because communities are hierarchical, you also pick which level to answer from, and that choice moves both your answer quality and your bill. Microsoft's own guidance points at C2, the third level of the hierarchy, as the level suited to most applications4. Fine as a starting point. Not a fact about your corpus.

And now the part that surprises people. Re-run community detection on a graph that changed slightly and two things happen at once, which are worth keeping separate in your head. The partition itself moves: entities land in different groups than before. And the identifiers get reassigned, so community 41 may now name a different set, or nothing at all.

The renumbering on its own is harmless if every join is recomputed together. It stops being harmless the moment something outside the index holds one of those IDs: a cache, a stored citation, a "this answer came from community 41" audit record, a diff between index versions. Then the label points somewhere it didn't before and nothing complains, because the system has no way to know the ID used to mean something else. Quietly wrong rather than loudly broken.

Two habits fix this and neither is expensive. Seed the algorithm. Then, after each re-cluster, map old communities to new ones by membership overlap and keep a stable public identifier that survives the remap. You want the ID in your logs to mean something six months later, which is the same argument Part 1 made about chunk provenance, arriving one layer up.

FIG.03 — THE DERIVATION CHAIN TEXT UNIT source ENTITY summarized COMMUNITY clustered REPORT summarized ANSWER shipped text_unit_ids carried at every step: the only path back WITHOUT THE DASHED PATH YOU CANNOT ANSWER "WHICH CHUNK" AND YOU CANNOT DELETE ANYTHING WITHOUT A FULL REBUILD
FIG.03 Two summarization steps sit between a source document and a shipped answer. The dashed path is what makes the system defensible, and what makes erasure possible at all.

The same discipline applies to summaries. A community report is a summary of entities that were themselves summarized out of text. That's two derivation steps away from any source document, and if the report doesn't carry the identifiers of the text units it came from, nobody can ever answer "which chunk?" The reference implementation does carry these links, and its own docs are explicit that the relationship weights matter because Leiden uses them10. If you build your own version of this stage, keep that plumbing. It is the difference between a graph you can defend and one you can only apologize for.

Aside

I still don't have a satisfying answer for how much community structure should be allowed to change between index versions before you treat it as a new system requiring re-evaluation. Ten percent of nodes moving? Half the top-level communities reshaping? I've been eyeballing it, which is not a methodology, and I haven't found anyone who has published a threshold. Might not be a real problem. Feels like one.

You are probably buying the wrong hard thing

Here's my fight, and I'd like to have it in public.

The graph database is not the difficult part of GraphRAG. It is the purchasable part, which is not the same thing, and the gap between those two facts is where a lot of money goes.

Look at what the reference implementation actually does. Microsoft's default pipeline writes its output tables, entities, relationships, communities, community reports, text units, as parquet files on disk10. No graph database. The system that defined this whole category ships a directory of files. You can hold that fact next to any vendor deck you're currently reading and see what it does to the argument.

At the corpus sizes most teams actually have, low millions of chunks, a relational database with a node table, an edge table, and a provenance table will carry you further than you expect. Your traversals are typically one or two hops with a degree cap, which is a recursive query, not a research problem. You get transactions, backups, a migration story, and an ops team that already knows how to run it.

A graph database earns its place when you genuinely need deep variable-length traversal, when the query language is doing real expressive work you'd otherwise write badly by hand, or when you already run one and the marginal cost is near zero. Those are real cases. They are not most cases, and adopting one at the start means you've bought an operational commitment to solve a problem you have not yet demonstrated you have.

Meanwhile the actual difficulty, entity resolution, is sitting right there in section 05, unglamorous, undemoable, and largely absent from both the vendor decks and the quickstarts. Nobody sells you a resolution layer because you cannot put it on a slide. It just quietly determines whether your graph is a knowledge structure or a heap.

Related pet peeve

"Knowledge graph" now means any structure with an adjacency list in it. I've seen a lookup table with two foreign keys described as a knowledge graph in a slide deck, by a person who I am fairly sure knew better. If the term covers everything, it distinguishes nothing, and we all lose a useful word.

08 · Query paths

Where the bill actually lands

Three paths, and they have almost nothing in common except the index they share.

Local search starts by linking your query to entities in the graph, then expands to their neighbourhood, then assembles a context window out of entities, relationships, source text units, and relevant community reports. The interesting design decision here is not the traversal, it's the budget: what proportion of your context goes to each of those categories. Skew it toward source text and you've built a slightly fancier vector search. Skew it toward relationships and community reports and you get more synthesis, and more distance from the underlying documents. There's no correct answer, but there is an unexamined default, and unexamined defaults are how systems end up behaving in ways nobody chose.

Entity linking is the silent failure, and it deserves its own paragraph. If the linking step returns nothing useful, if the query names something your extractor never captured, or names it differently, the whole graph apparatus degrades to whatever fallback you configured. Usually plain vector search. The system keeps answering. Latency looks normal. Quality drops. And nothing in your logs says "the graph contributed nothing to this response," because nobody instrumented that. Log the seed set size on every query. Alert on zero. This is a two-hour change that will tell you more about whether your graph is earning its cost than a quarter of eval work.

Global search is a different animal. It takes your question to the community reports at a chosen level, generates partial answers from them in parallel, then reduces those into a final response. The cost of that scales with the number of communities you're mapping over, which means it scales with your corpus and your resolution choice, and has almost nothing to do with the length of the question. One question can be dozens of model calls.

FIG.04 — THREE PATHS, ONE INDEX LATENCY TOKEN COST PER QUESTION LOCAL seed, expand, assemble DRIFT community seed, then local GLOBAL map over communities, then reduce RELATIVE AND SCHEMATIC — NOT MEASURED ON YOUR CORPUS
FIG.04 Global search costs scale with the number of communities, not the length of the question. Positions are relative, not measurements.

Do that arithmetic once, honestly, then multiply by realistic traffic. The number that comes out is the reason Microsoft published a dynamic community selection method that uses cheaper models to rate relevance before spending on the expensive stage12, and it's the reason LazyGraphRAG exists at all, reporting more than 700 times lower query cost than global search while matching its answer quality on global questions4. When the same team ships two separate optimizations aimed at one code path, that path was expensive.

DRIFT sits between the two, using community information to broaden a local query's starting point and generate follow-up questions, then refining downward11. Useful, and more expensive than local search. Worth knowing it exists so you don't build a worse version of it by accident.

The conclusion I'd draw, and I'd draw it firmly: global sensemaking is not a chat feature. The latency and cost profile belong in asynchronous surfaces. Scheduled briefings. Weekly digests. An analyst clicking "generate report" and going to get coffee. Every team I've seen put global search behind an interactive chat box has ended up either capping it, caching it into uselessness, or hiding it behind a button that most users never press. Design for the shape of the cost instead of fighting it.

09 · Update

And the one nobody writes about, deletion

Everything so far assumed a static corpus. Yours isn't.

Delta ingest is the easy half. New documents arrive, you extract from them, and, this is the important bit, you resolve the new entities against your existing canonical entities rather than against the batch. Graphiti does exactly this: it ingests continuously, resolves against what's already in the graph on arrival, and uses temporal metadata to invalidate superseded facts rather than discard them, specifically to avoid recomputing the whole graph5. That's the right shape. Append edges, close validity windows, move on.

Community drift is the hard half. Your communities were computed over a graph that no longer exists. Append enough new nodes and the partition that made sense in March is describing a corpus from March. You have three options and you should pick one deliberately rather than by neglect: append until quality decays and accept the drift, re-cluster only the affected subgraphs, or schedule full rebuilds. Whichever you pick, define the trigger in advance. Share of nodes added since last clustering, change in modularity, or simply a calendar. A trigger you wrote down is a policy. A trigger you didn't is a surprise.

Summary invalidation follows from the provenance links, and it's subtler than it first looks. The obvious trigger is a community whose membership changed. But a report is built from the entity descriptions, relationships and claims inside that community, not from the membership list, so a report can go stale while its membership is identical: an entity description gets richer because new documents mentioned it, a relationship weight moves, a claim is added14. And because the community hierarchy is built upward, a stale report at one level makes every ancestor report above it stale too.

So invalidate by dependency rather than by membership. Something changed underneath this report, directly or through a descendant, therefore regenerate this report and everything above it. That's still far cheaper than a full rebuild, which is the point, but the naive version misses exactly the cases where a summary drifts quietly instead of obviously.

And then deletion. A document is retracted. A customer exercises a right to erasure. A contract expires and its terms must no longer inform answers. In a chunk-based system this is tractable: find the chunks, remove them, re-embed. In a graph, that document's claims have already been abstracted into entity descriptions, then into a community report, then possibly into a parent report summarizing that one. The text is gone. Its influence isn't.

The only way this is solvable is if every derived object carries the identifiers of the text units beneath it, so you can compute the affected set and regenerate. Which means the provenance links from section 06 aren't an audit nicety, they're the deletion mechanism. Build without them and erasure becomes a full re-index, every time, at full cost.

I'll note honestly that I have not found a good published treatment of erasure in graph-based RAG. Not in the framework docs, not in the papers. If someone has solved this properly I'd like to read it, and if nobody has, that's a gap worth someone's attention, because the obligation exists whether or not the tooling acknowledges it.

Two operational habits close this out. Run re-indexes blue/green, because a rebuild that swaps in place will serve half-built state to somebody. And put the index version in every log line, so that when a question comes back six weeks later, the sort of question that opened this series, you can reconstruct which graph the system was actually looking at.

10 · Cost

The cost shape

I'm not going to give you a dollar figure, because any number I invent would be wrong for your corpus and would get quoted back at me for two years. What's useful is the shape.

Index-time cost has two dominant terms. Extraction is the sum of input tokens times the input price plus output tokens times the output price, across the initial pass over every chunk and every gleaning call after it:

C_extract = Σcalls ( T_in × P_in + T_out × P_out )
calls = chunks × (1 + gleaning passes)
note : each gleaning pass re-sends the accumulated context,
            so T_in grows pass over pass

Summarization scales with the number of communities you generate, which is a function of graph size and your resolution setting, and it compounds because higher levels summarize the summaries below them.

Query-time cost depends entirely on which path you're on. Local search is close to ordinary RAG plus a traversal. Global search is many model calls per question. DRIFT sits between them.

Two levers move the index number more than anything else: gleaning passes, and how many community levels you generate and summarize. If your bill is uncomfortable, those are the first two dials, and you can reason about both without a benchmark.

FIG.05 — WHERE THE MONEY SITS INDEX TIME QUERY TIME SMALL CORPUS MEDIUM LARGE RELATIVE COST SCHEMATIC a re-index spends the whole violet block again, in one afternoon, on purpose
FIG.05 Proportions are schematic, not measured. The point is which block grows, and that it is spent again in full on every rebuild.

Here's the part worth saying out loud. Part 1 told a story about runaway cost from an agentic retry loop, which was a query-time failure: it accumulated invisibly, one request at a time, until a finance meeting. Graph systems fail the opposite way. Index-time cost arrives as a single decision, made in an afternoon, that spends the whole indexing budget again. Nobody watches it accumulate because it doesn't accumulate. Someone types a command, and either you knew what it cost or you find out.

Which is why I'd rather a team know their re-index cost to within a factor of two before launch than have a beautiful eval suite. The eval tells you whether the system is good. The re-index number tells you whether you can afford to keep it good.

11 · Evaluation

Three scorecards, and one of them is genuinely unsolved

Part 1 made the case for splitting retrieval and generation into separate numbers. A graph needs three, because you've added a layer that can be wrong on its own.

Graph quality. This is measured on the graph, without asking any questions at all. Take a sample of chunks, extract from them, and have a human check the triples against the source: precision and recall on entities and relations. Then measure resolution separately, as pairwise precision and recall on a hand-labeled set of candidate pairs. Then look at the structural diagnostics from section 05, degree distribution and relation type distribution, which cost nothing and catch the loudest failures. A hundred labeled chunks and a few hundred labeled pairs is not a research project. It's a week, once, and it tells you whether anything downstream is worth measuring.

Local retrieval quality. Ordinary recall against labeled relevant evidence. Nothing exotic here, and that's the point: if your graph can't beat hybrid search on your own question set, you've learned something cheaply.

Global answer quality, where I have to be honest with you. There isn't a good answer yet. The original paper evaluated global sensemaking on comprehensiveness and diversity using a language model as judge1, and the benchmark literature has since argued that the standard evaluation sets don't cleanly isolate what the graph contributes in the first place3. So the field's own instruments are unsettled.

FIG.06 — THREE SCORECARDS GRAPH QUALITY MEASURES extraction, resolution, shape METHOD labeled sample, one week settled LOCAL RETRIEVAL MEASURES did it fetch the evidence METHOD recall against labels settled GLOBAL ANSWERS MEASURES is the synthesis any good METHOD claim checks, theme coverage no settled method VERSION THE GOLDEN SET AND THE GRAPH SNAPSHOT TOGETHER, OR YOU MEASURED THE INDEX
FIG.06 Two of the three are ordinary work. The third is where the field's own instruments are still unsettled, and pretending otherwise helps nobody.

What I'd actually do, absent a solved method: check claims rather than answers. Decompose the response into individual assertions and verify each one traces to source text units through the derivation chain. That's tractable, it's mechanical, and it tests the property you care about, which is groundedness rather than eloquence. Then, separately, build a small hand-written list of themes you know are in the corpus and check coverage. Crude. Better than nothing, and considerably better than a model rating its own summary out of ten.

If you do use a model as judge, at least don't use the same model that wrote the summary, and say so in your write-up. The circularity is real and mostly goes unmentioned.

One discipline that matters more here than in chunk-based systems: version your golden set and your graph snapshot together. If the graph changed between eval runs, you measured the index, not your change. In a chunk system that's a minor confound. In a graph system, where a re-cluster reshapes everything downstream, it invalidates the comparison entirely.

Aside

I mentioned in Part 1 that I still don't have a golden dataset for my own side project, and that I'd been meaning to build one for about eight months. It has now been longer than eight months. I'm not going to keep updating this number.

12 · Diagnosis

The failure table

The part I'd actually bookmark. Symptom, likely cause, first thing to check.

SymptomLikely causeFirst check
Answers no better than hybrid search Entity linking returning nothing useful Log seed set size per query, look for zeros
One entity connected to everything Over-merge in resolution Degree distribution against median
Traversals miss obvious connections Predicate sprawl from open extraction Count distinct relation types
Summary describes deleted content No invalidation on dependency change Check whether reports carry text unit IDs
Relationship direction reversed in answers Extraction direction errors Read fifty triples against source sentences
"Who is the X" returns several people No temporal validity on edges Look for validity fields, expect to find none
Same question, different answer after re-index Community partition changed, or a stale cache keyed on unstable community IDs Compare community membership across versions, then check what your caches are keyed on
Index cost far above estimate Gleaning passes, or too many community levels Count model calls per document

Most of these take under an hour to check and most of them are, in my experience, present in some form on first inspection.

13 · Alternatives

The cheaper things, and the graph you already have

Before you build any of the above, three alternatives deserve a real attempt.

Query decomposition with iterative retrieval handles a surprising share of what people call multi-hop. Break the question into sub-questions, retrieve for each, feed the results forward. No index, no maintenance, and you can build it in a week on top of the hybrid stack you already have.

Hierarchical summarization gets you much of the corpus-level benefit without entities or relations at all. RAPTOR recursively clusters and summarizes chunks into a tree, then retrieves at whichever level of abstraction fits the question13. Compare that with the graph pipeline in section 03 and notice how much you're not building: no extraction schema, no resolution, no community stability problem. If your global questions are really "summarize across this collection," that may be the whole answer.

Use the graph you already have. This is the one I'd shout if the format allowed. Relationships are already sitting in structured form all over most organizations: citations, ticket parent and child links, foreign keys, bill of materials structures, org charts, contract references, hyperlinks between wiki pages. Those edges are free, they're already correct, and using them skips stages two and three of the pipeline, which are the two most expensive stages in both money and engineering time. Extraction from prose is what you do when the relationships exist only in prose. It's the last resort. Every quickstart presents it as the starting line.

And I say that as someone who got it wrong in exactly that direction. On one build I went straight to model-based extraction over documents, wrote a schema, tuned prompts, ran gleaning passes, argued about chunk sizes for extraction, and produced a graph I was reasonably proud of. Most of those relationships already existed, cleanly, in systems the client owned. I'd been so interested in the extraction problem that I never asked whether the extraction problem needed solving. It cost weeks, and the embarrassing part is that the boring version would have been better as well as cheaper: real foreign keys don't invert direction or hallucinate bridges. Same mistake as the semantic search one from Part 1, wearing a different hat. I prefer the interesting path and I have to be talked out of it, so now I ask the structured-sources question first, out loud, in the first meeting.

14 · Close

Where this leaves you

Run the baseline and name the failure class before you build anything. Prefer the edges you already have over the ones a model can infer. Budget resolution and re-indexing rather than extraction, because that's where the work actually is. Keep chunk-level provenance on every derived object, not for the auditor but because it's the only thing that makes deletion possible. And put the index version in every log line.

What I still can't tell you is where the crossover is. There must be a point, some combination of corpus size, update rate and question mix, where the graph starts paying for its own maintenance, and I don't know how to compute it. Nobody I've read has published it against a genuinely messy internal knowledge base, the kind with duplicated wiki pages and half-abandoned ticket threads and three competing product naming conventions. Maybe the answer is that it's too corpus-specific to generalize. Maybe someone has done it and I haven't found it.

If you've built one that's still running after a year, and you can say what it costs to keep true, I'd like to hear it. Especially the maintenance number. That's the one nobody publishes, and it's the one that decides this.

Notes and sources
  1. Edge, D., Trinh, H., Cheng, N., Bradley, J., Chao, A., Mody, A., Truitt, S., Metropolitansky, D., Ness, R. O., Larson, J. (2024). From Local to Global: A Graph RAG Approach to Query-Focused Summarization. arXiv:2404.16130. arxiv.org/abs/2404.16130↩↩↩
  2. RAG vs. GraphRAG: A Systematic Evaluation and Key Insights. arXiv:2502.11371. arxiv.org/abs/2502.11371↩
  3. When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation (GraphRAG-Bench). arXiv:2506.05690. arxiv.org/abs/2506.05690↩↩
  4. Microsoft Research (2024). LazyGraphRAG: Setting a New Standard for Quality and Cost. microsoft.com/en-us/research/blog/lazygraphrag-setting-a-new-standard-for-quality-and-cost↩↩↩
  5. Rasmussen, P., et al. (2025). Zep: A Temporal Knowledge Graph Architecture for Agent Memory. arXiv:2501.13956. arxiv.org/abs/2501.13956↩↩
  6. Christophides, V., Efthymiou, V., Palpanas, T., Papadakis, G., Stefanidis, K. (2020). An Overview of End-to-End Entity Resolution for Big Data. ACM Computing Surveys 53(6), Article 127. dl.acm.org/doi/10.1145/3418896↩
  7. Papadakis, G., Skoutas, D., Thanos, E., Palpanas, T. (2020). Blocking and Filtering Techniques for Entity Resolution: A Survey. ACM Computing Surveys 53(2). arxiv.org/abs/1905.06167↩
  8. Christen, P. (2012). Data Matching: Concepts and Techniques for Record Linkage, Entity Resolution, and Duplicate Detection. Springer. link.springer.com/book/10.1007/978-3-642-31164-2↩
  9. Traag, V. A., Waltman, L., van Eck, N. J. (2019). From Louvain to Leiden: Guaranteeing Well-Connected Communities. Scientific Reports 9, 5233. nature.com/articles/s41598-019-41695-z↩↩
  10. Microsoft. GraphRAG documentation: Outputs. microsoft.github.io/graphrag/index/outputs↩↩
  11. Microsoft Research (2024). Introducing DRIFT Search: Combining Global and Local Search Methods to Improve Quality and Efficiency. microsoft.com/en-us/research/blog/introducing-drift-search↩
  12. Microsoft Research (2024). GraphRAG: Improving Global Search via Dynamic Community Selection. microsoft.com/en-us/research/blog/graphrag-improving-global-search-via-dynamic-community-selection↩
  13. Sarthi, P., Abdullah, S., Tuli, A., Khanna, S., Goldie, A., Manning, C. D. (2024). RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval. ICLR 2024. arXiv:2401.18059. arxiv.org/abs/2401.18059↩
  14. Microsoft. GraphRAG documentation: Default Dataflow. microsoft.github.io/graphrag/index/default_dataflow↩↩

Framework behaviour described here reflects the reference implementation as documented at the time of review. Parameter names and defaults move; the failure modes don't.

  • 01 · Choosing an architecture
  • 02 · Hybrid & reranking
  • 02a · Addendum: the middle ground
  • 03 · Corrective & Self-RAG
  • 04 · GraphRAG in practice
  • 05 · Adaptive & agentic
  • 06 · Multimodal
  • 07 · Evaluation & golden datasets

I publish a new article in this series every week: mostly what breaks in production, not what looked good in the demo. Part 5 is Adaptive and Agentic RAG, which has its own version of the cost problem, except that one accumulates in the dark instead of arriving in a single invoice. Follow along if you want it when it lands; I post here first, before anywhere else.

And if you've built a graph that's still running after two years: tell me what it costs to keep true. Comment, DM, whatever. That maintenance number is the one nobody publishes.