Vector search answers one question well: which passages sound like this question? Many real questions are not like that. "Which of our products contain a part from a supplier that had a recall in 2025?" has no single passage that sounds like the answer, because the answer is a chain of facts spread across a bill of materials, a supplier register and a stream of recall notices. Graph RAG stores those facts as explicit nodes and edges, so retrieval can follow the chain instead of hoping one chunk contains it.

This article explains how a graph RAG system is put together next to a vector index, which question shapes a graph actually wins on, how to design a schema where every edge carries its evidence, why entity resolution decides whether the whole thing works, and how to route queries between the two paths. It ends with an evaluation harness so the decision to build or keep a graph is made from your own data. The maths of community detection and summary costs is covered separately in Graph RAG Math; this page is about architecture and judgement.

Advertisement

What vector search retrieves, and what it cannot

A dense retriever embeds the question and every chunk into the same vector space and returns the chunks nearest to the question. That is powerful for paraphrase and topical match, and combined with keyword search and a reranker, as described in hybrid search with BM25, dense retrieval and reranking, it handles most lookup questions.

It has three structural blind spots. First, multi-hop joins: when the answer needs fact A from one document and fact B from another, and B is only relevant because of A, nothing about B's text resembles the question. Second, aggregation: "how many suppliers in Vietnam supply more than three of our products" needs counting over many facts; top-k retrieval returns a sample, and the model then counts the sample. Third, corpus-wide themes: "what are the recurring causes of our recalls" needs a view across hundreds of documents, more than any context window holds.

A graph addresses the first two by making relationships data that can be traversed and counted exactly, and the third by precomputing summaries over clusters. It pays with an extraction pipeline, a second store to keep fresh, and a new class of errors: wrong or missing edges.

The architecture

The diagram shows the usual shape. Documents are chunked with stable ids, exactly as for vector RAG, and the chunks are embedded into a vector index. In parallel, an extractor, usually an LLM with a fixed schema prompt, reads each chunk and emits typed entities and relations. Entity resolution merges mentions that refer to the same real thing. The result is loaded into a property graph in which every edge records the chunk it came from. Optionally, clusters of the graph are summarised for corpus-wide questions.

Graph RAG next to a vector index: one corpus, two retrieval pathsINGESTDocumentschunked, ids stableExtractor (LLM)entities + relationsEntity resolutionmerge aliases, log mergesSTORESChunk storetext + metadataVector indexchunk embeddingsProperty graphedges cite chunk idsCommunity summariesoptional, global questionsQUERYQuestionRouterlinked entities + shapeContext assemblerpaths + chunks, budgetedLLM answercites chunk idsvectorgraphglobalevidenceEvaluation harness: golden questions labelled by shape, scored per strategyevidence recall, answer correctness, tokens, latency -> the routing table
Two retrieval paths over one corpus. The graph never replaces the chunk store: graph paths are evidence pointers, and the chunks they cite are what the model reads. The evaluation harness decides the router's rules.

At query time a router links entities in the question and classifies its shape, then sends it to vector retrieval, graph traversal, community summaries or a combination. The context assembler turns graph paths back into the source chunks they cite, so the model reads real text rather than bare triples, and every claim in the answer can point at a chunk id.

Advertisement

Which question shapes a graph wins

Classify questions by shape before arguing about technology. This table is the core decision aid; the evaluation section shows how to fill in the last column with your own measurements rather than intuition.

ShapeExampleBetter pathWhy
LookupWhat is our refund window for EU customers?Vector / hybridOne passage holds the answer
Fuzzy or paraphrasedHow do I get my money back?VectorSemantic match is the whole problem
Multi-hopWhich products use parts from recalled suppliers?GraphAnswer is a path across documents
AggregationHow many suppliers serve more than three products?GraphCounting must be exact, not sampled
RelationshipHow is team A connected to vendor X?GraphShortest path is a graph primitive
Global themeWhat recurring causes appear in recalls?Community summariesNeeds a view across the corpus
Recent or volatileWhat changed in yesterday's release notes?VectorExtraction lag makes the graph stale

If most of your traffic is lookup and paraphrase, as in most support assistants, a graph adds cost and little accuracy. If a meaningful share is multi-hop or aggregation over entities your business already models, such as parts, suppliers, contracts, customers, services and incidents, a graph usually pays for itself.

A worked example: supplier recall exposure

Consider a manufacturer with three document sets: bill-of-materials pages (product contains component), a supplier register (component made by supplier, supplier located in country) and recall notices (recall concerns supplier, with a date and cause). The question is: which products are exposed to a supplier recall in 2025?

Vector retrieval returns the recall notices, because they share words with the question. It rarely returns the bill-of-materials pages, because "Product P-220 contains assembly A-17" says nothing about recalls. The model then either says it cannot tell or, worse, guesses. A graph answers with a traversal:

MATCH path = (p:Product)-[:CONTAINS*1..3]->(c:Component)-[m:MADE_BY]->(s:Supplier)
             <-[:CONCERNS]-(r:Recall)
WHERE r.date >= date('2025-01-01') AND m.confidence >= 0.7
RETURN p.name AS product, c.name AS component, s.name AS supplier, r.id AS recall,
       [e IN relationships(path) | e.source_chunk] AS evidence
LIMIT 50

The variable-length CONTAINS*1..3 follows sub-assemblies down three levels, which no top-k retriever can do. Binding the whole path lets the evidence column list the source chunk of every edge on it; the assembler fetches those chunks and the prompt asks the model to explain exposure product by product, citing them. The same graph answers "how many products per supplier" with a count() instead of a guess.

Design the schema so every edge carries evidence

The most important schema decision is that edges are claims, not facts. Each one was produced by an extractor from a specific chunk, with some confidence, at some time. Store that on the edge:

MERGE (c:Component {id: $component_id})
  ON CREATE SET c.name = $component_name
MERGE (s:Supplier {id: $supplier_id})
  ON CREATE SET s.name = $supplier_name
MERGE (c)-[m:MADE_BY {source_chunk: $chunk_id}]->(s)
SET m.confidence   = $confidence,
    m.extractor    = $extractor_version,
    m.observed_at  = datetime($observed_at),
    m.doc_version  = $doc_version

Keying the relationship on source_chunk means two documents asserting the same fact create two edges, which is what you want: support can be counted, contradictions stay visible, and deleting or re-extracting one document removes exactly its edges. Keep the schema small, with a closed list of node and edge types written into the extraction prompt. Open-ended extraction produces hundreds of near-duplicate relation names such as supplies, is_supplier_of and provides_parts_to, which cannot be queried reliably. Start from the entities your business already has identifiers for; those give you free ground truth for resolution.

Entity resolution decides whether it works

Extraction produces mentions: "Acme GmbH", "ACME", "Acme Components", "the supplier". If they are not merged, the graph fragments and traversals silently miss paths; if unrelated entities are merged, traversals invent paths. Both errors are invisible in the final answer.

def resolve(mention, index, embed, auto=0.92, review=0.80):
    key = normalise(mention.name)                 # case, punctuation, legal suffixes
    if mention.external_id:                        # a registry id beats any similarity
        return index.by_external_id(mention.external_id), 1.0
    candidates = index.block(mention.type, key[:4])  # cheap blocking, same type only
    scored = []
    for c in candidates:
        name_sim = string_similarity(key, c.key)
        ctx_sim = cosine(embed(mention.context), c.context_centroid)
        scored.append((0.6 * name_sim + 0.4 * ctx_sim, c))
    score, best = max(scored, default=(0.0, None), key=lambda t: t[0])
    if score >= auto:
        index.log_merge(mention, best, score, decided_by="auto")
        return best, score
    if score >= review:
        index.queue_for_review(mention, best, score)   # human or stronger model decides
    return index.create(mention), score

Three habits matter more than the similarity function. Use external identifiers such as part numbers and tax ids whenever they exist. Log every merge with its score so a bad merge can be reversed. And sample the review band weekly: the threshold that looks right on day one drifts as the corpus grows.

Retrieval patterns on the graph

Three patterns cover most needs. Entity-anchored expansion links entities in the question, then expands one or two hops with a type filter and a fan-out cap, and returns the cited chunks. Text-to-query asks an LLM to write the graph query from the question and the schema; it handles aggregations well but needs guardrails: read-only credentials, a whitelist of labels, a mandatory LIMIT, a timeout, and a fallback when the query returns nothing. Community summaries precompute summaries of densely connected clusters and answer global questions by map-reduce over them; the cost model is in the maths article linked above.

def expand(graph, entity_ids, hops=2, max_edges=200, min_conf=0.7):
    frontier, seen_edges = set(entity_ids), []
    for _ in range(hops):
        edges = graph.edges_from(frontier, min_confidence=min_conf,
                                 limit=max_edges - len(seen_edges))
        seen_edges += edges
        frontier = {e.target for e in edges} - frontier
        if len(seen_edges) >= max_edges or not frontier:
            break
    chunk_ids = {e.source_chunk for e in seen_edges}
    return seen_edges, chunk_ids        # paths for reasoning, chunks for citation

The fan-out cap is not optional. Hub entities, such as a large supplier or a common component, connect to thousands of nodes, and an uncapped two-hop expansion from them floods the context with irrelevant edges.

Routing between the paths

A router keeps graph cost off questions that do not need it. A small classifier, or a cheap LLM call with a fixed label set, predicts the shape; an entity linker checks whether the question mentions anything in the graph.

SHAPES = ["lookup", "fuzzy", "multi_hop", "aggregate", "relationship", "global_theme"]

def route(question, classify, link_entities):
    shape = classify(question, labels=SHAPES)
    entities = link_entities(question)
    if shape in ("multi_hop", "relationship") and entities:
        return ["graph_expand", "vector"]           # graph first, vector as backstop
    if shape == "aggregate" and entities:
        return ["graph_query"]
    if shape == "global_theme":
        return ["community_summaries"]
    return ["vector"]

def retrieve(question, plan, budget_tokens=6000):
    for strategy in plan:
        ctx = STRATEGIES[strategy](question)
        if ctx.chunks:
            return ctx.trim(budget_tokens)
    return STRATEGIES["vector"](question).trim(budget_tokens)

Always fall back to vector retrieval when the graph returns nothing, and record which path answered. Silent graph misses, where the linker found no entity or a traversal hit a missing edge, are the most common production failure, and the path log is how you find them.

Decide with an evaluation harness

Whether a graph beats vector search is an empirical question about your corpus and your traffic. Build a golden set of 100 to 300 real questions, label each with its shape and the chunks that contain the evidence, and score every strategy on every question.

rows = []
for q in golden:                                   # q.text, q.shape, q.gold_chunks, q.reference
    for strategy in ("vector", "hybrid", "graph_expand", "router"):
        ctx = run_strategy(strategy, q.text)
        answer = generate(q.text, ctx)
        rows.append({
            "shape": q.shape, "strategy": strategy,
            "evidence_recall": len(set(ctx.chunk_ids) & q.gold_chunks) / len(q.gold_chunks),
            "correct": judge(answer, q.reference),    # rubric-based, spot-checked by people
            "tokens": ctx.tokens, "latency_ms": ctx.latency_ms,
        })
report = pivot(rows, index="shape", columns="strategy",
               values=["evidence_recall", "correct", "tokens"])

Read the report by shape, not overall: a graph that gains on multi-hop questions but loses a little on lookups is a routing opportunity. Evidence recall tells you whether retrieval found the facts; correctness tells you whether the model used them. Weight each shape by its share of real traffic before deciding, because a large gain on a rare shape may not justify a second store.

Failure modes

  • Extraction misses. An edge that was never extracted makes a traversal return a confident, incomplete answer. Measure extraction recall on a labelled sample, and keep vector retrieval as a backstop.
  • Over-merged entities. Two different suppliers named "Delta" become one node and the graph invents exposure. Merge logs and external ids are the defence.
  • Stale graph. Extraction runs in batch, so the graph lags the chunk store. Record extraction time per document and route recency questions to vector search.
  • Hub explosion. Traversals through popular nodes return thousands of edges. Cap fan-out and filter by edge type.
  • Schema drift. A prompt change renames a relation and old and new edges coexist. Version the extractor and re-extract, or migrate edges, when the schema changes.
  • Unsafe generated queries. Text-to-query can write expensive or destructive queries. Use a read-only role, timeouts and a label whitelist.

Costs and operations

The dominant cost is extraction: an LLM call per chunk at ingest, repeated whenever the extractor or schema changes, so budget for full re-extraction as routine. Community summaries add a second batch job. The graph store needs its own backups and access control, and that control must match the chunk store, or a traversal can reveal a relationship from a document the user may not read. Filter edges by the permissions on their source chunk.

Monitor extraction throughput and cost, the resolution review queue, the share of queries per route, graph-empty fallbacks, and evidence recall on a weekly golden-set run. For the surrounding pipeline and index operations, see RAG pipeline design at scale, vector search infrastructure and context and RAG libraries in the agentic stack. A companion article in this category, on data lineage and contracts for AI knowledge systems, covers how to trace every edge and chunk back to its source.

What to do next

  1. Pull 200 real questions from logs and label each with one of the shapes in the table above.
  2. Estimate the share of multi-hop, aggregation and relationship questions. If it is small, improve hybrid retrieval first.
  3. Pick five to ten entity types you already have identifiers for and write a closed extraction schema.
  4. Extract a pilot slice of the corpus with source_chunk, confidence and extractor version on every edge.
  5. Build resolution with external ids first, logged merges and a review band.
  6. Implement entity-anchored expansion with a fan-out cap and a vector fallback, then a simple router.
  7. Run the harness per shape, weight by traffic, and decide whether to expand, route narrowly or drop the graph.
Key takeaway: Graph RAG beats vector search on multi-hop, aggregation and relationship questions over entities your business already models, and loses on lookups, paraphrase and fast-changing content. Build it as an evidence index next to the chunk store: every edge cites a chunk, resolution is logged and reversible, traversals are capped, and vector retrieval remains the fallback. Let a per-shape evaluation on your own questions decide the routing, and budget for re-extraction as routine.