The Chunk Was Never the Document: Multimodal RAG in Production

RAG Architecture Series · Part 6

The Chunk Was Never the Document

Five articles of retrieval tuning assumed a chunk is a span of text you can embed, cite, and delete. Point that assumption at a scanned invoice or a 400 page equipment manual and watch it come apart. This is what replaces it, and what the replacement costs.

Reviewed September 2026 / Developer track / 21 min read

Before any of this: go open twenty documents your users actually query, and count how many answers live in something that isn't extractable text.

For most corpora the number is small. Which means the request behind "we need multimodal" is usually "our tables come out as word soup," and that's a parsing problem with a much cheaper fix than a vision pipeline. Do the count first. It takes an afternoon and it decides the rest of this article.

But when the number isn't small, you have a real problem, and it's bigger than it looks, because everything in Parts 1 through 5 quietly assumed the same thing. A chunk is a contiguous span of text. It has an ID. You can embed it, rank it, show it to a grader, cite it in an answer, and delete it when someone asks you to. Every technique we've built rests on that object existing.

Put a page in front of it where the answer is a value in the third column of a table that broke across a page boundary, and the object stops existing. Not degrades. Stops. There's no span to point at, because the thing you'd want to point at was never text in the first place.

Everything below assumes hybrid retrieval and a reranker from Part 2, grading from Part 3, and a router from Part 5. If you don't have those, this is premature, and pointing a vision model at retrieval that was never good enough gets you the same wrong page, in higher resolution, for more money.

Terminology

Image retrieval and document understanding are not the same job

The demo is a photo library and a CLIP model. Somebody types "red bicycle" and gets a red bicycle, and everyone agrees this is impressive, and it is.

Your corpus is not that. Your corpus is equipment manuals with exploded parts diagrams, invoices that were faxed and rescanned, slide exports where half the content is in text boxes and the other half is baked into a PNG somebody pasted in 2019, and screenshots in wiki pages documenting a UI that no longer exists.

Those are different problems. The first is matching a description to a picture. The second is reading a document that happens to convey meaning through layout, and where the visual element usually isn't the answer so much as the container the answer is sitting in. They got the same word because both involve pixels, which is roughly as useful a category as "things that are blue."

Part 4 made the same complaint about GraphRAG naming three different jobs. I'm not going to pretend that's a new insight arriving fresh. It's the same failure in a new domain, and the fix is the same: say which one you're buying, out loud, in a sentence, before anyone opens a terminal.

Aside

Every multimodal retrieval demo runs on curated images with clean captions. Product photos. Stock scenes. Charts rendered at 300 DPI by matplotlib last week. Nobody demos a fax. Nobody demos the third-generation photocopy with the coffee ring and the handwritten "SEE ATTACHED" across the corner, which in a lot of industries is the actual document of record. I'd take one honest demo on a bad scan over ten on an infographic.

Architecture

Four patterns, and where each one puts the cost

Caption at ingest

A vision model looks at each figure or page and writes a description. That description becomes text, and from there it's an ordinary text pipeline: same embeddings, same reranker, same everything. Query time is cheap because there's no vision model in the request path. Ingest is where it hurts. The ColPali authors note that captioning-based approaches can run to dozens of seconds per page, and their own measured pipeline breaks down as 0.81 seconds for layout detection, 2.67 for OCR and 3.71 for captioning, 7.22 seconds per page all in, against 0.39 for embedding the page image directly1. Then look at what the 3.71 bought: on their benchmark, adding a full captioning stage to a parsed pipeline landed within about a point of just running OCR over the visual elements1. Half your ingestion budget for a point. That's not nothing, but it isn't what the architecture diagram implied either.

Unified multimodal embeddings

One vector space, text and images together, cosine similarity across the whole thing. Conceptually lovely. On documents it falls over, and not by a little. On ViDoRe, the general-purpose contrastive vision-language models scored 17.7 and 12.9 nDCG@5, against 81.3 for a model trained for document retrieval1. Those aren't tuning differences, they're a different sport. The paper's own explanation is the boring one and probably the right one: these models were trained to match captions to pictures, their vision encoders were never optimized for reading dense text, and document retrieval asks them to do the thing they weren't built for1. There's also a geometric story in the literature, which is that image and text representations in contrastively trained models occupy separate regions of the space from initialization onward, and that the training objective preserves the separation rather than closing it2. I'd be careful about how much weight to put on that second one, for reasons I get into in section 06. It's a plausible contributing factor, not a proof.

Late interaction over page images

Skip parsing entirely. Render the page, embed it as patches, match query tokens against patches at retrieval time. This works startlingly well on exactly the documents that break parsers, and it deletes the whole brittle ingestion stage. The bill is storage: ColPali reports 257.5 KB per page for its multi-vector index1. Run the arithmetic on your own corpus before you get excited, because that's per page, and it's a different order of magnitude from one dense vector per chunk. Pooling recovers a lot of it, 66.7 percent fewer vectors while keeping 97.8 percent of performance at a pool factor of 3, with the important caveat that the most text-dense dataset in their evaluation degraded worse than the rest1. Dense pages have fewer redundant patches to pool away. If your corpus is dense text, you are the outlier case.

Keep the text pipeline, add an image path, merge at rerank

Parse what parses. Route the visual elements down a separate path. Bring both back together at rerank time, with a caveat that matters: the text cross-encoder from Part 2 cannot do this. Hand it an image candidate and a text candidate and there is nothing for it to score. You need one of two things. Either a multimodal reranker, which now exists as an off-the-shelf component category, or you rerank a shared textual representation of every candidate, the caption or table serialization, while carrying the visual source along as provenance and showing that to the user. The second option is cheaper, works with the reranker you already run, and fits the result object in section 05. Boring, incremental, and consistent with the position this series has held since Part 1.

FIG.01 — FOUR PATTERNS, FOUR INVOICES INGEST STORAGE QUERY CHARACTERISTIC FAILURE CAPTION AT INGEST HIGH LOW LOW prompt change = full re-index UNIFIED EMBEDDINGS LOW LOW LOW dense document text PAGE-IMAGE LATE INTERACTION MED HIGH MED storage, and nothing left to grep TEXT PIPELINE + IMAGE PATH MED MED MED needs a multimodal rerank step SHAPE, NOT BENCHMARK. MEASURE YOURS.
FIG.01 Same inputs, same outputs, four completely different invoices.
PatternProvenance you end up withWins when
Caption at ingestWeak. The answer traces to a description, and the description is the only artifact anyone kept.Few visual elements, stable corpus, and you are never changing the prompt.
Unified embeddingsWeak, and the ranking is hard to reason about besides.Genuine image search. Not documents.
Page-image late interactionStrong. The page is the citation, and you can show it.Scans, dense layout, and parsing that keeps breaking.
Text pipeline + image pathStrong, if you keep element IDs and the crop.Most corpora, most of the time.

Here's the part worth carrying out of this section. Patterns one and three are index-time architectures. The work happens before anyone asks a question, it scales with corpus size rather than traffic, and it runs again in full the day you change your captioning prompt or swap your vision model.

That should feel familiar, because it's the exact shape of GraphRAG from Part 4, and it fails the same way. Query-time cost accumulates quietly until a finance meeting. Index-time cost is one decision, made in an afternoon, that spends the whole indexing budget again. Nobody watches it accumulate because it doesn't accumulate. Somebody types a command, and either you knew the number or you're about to.

So the same discipline applies. Know your re-caption cost to within a factor of two before you launch. And note that this is now the second architecture in this series with that property, which makes me think the useful question to ask about any new RAG technique isn't "how good is it," it's "when does the expensive part run."

Tables are the actual problem

If I had to bet on the single most common root cause behind "the model made up a number," I'd bet on a table that got serialized badly, and it wouldn't be close.

Three things go wrong and they compound. Reading order, where a two-column layout gets flattened left to right across the fold, so half of row four is now sitting inside row nine. Merged cells, where one header spans three columns and the serializer either drops the span or repeats it, and either way the column-to-header mapping is now fiction. And page breaks, where the header row is on page eleven and the row you need is on page twelve, which arrives in your index as a naked sequence of numbers with nothing to say what they measure.

That last one is the killer, because the chunk is not obviously broken. It's fluent. It has numbers in it. It'll rank fine for a query about those numbers. The generator will read it and answer confidently, and the answer will be a value from the wrong column.

FIG.02 — ONE TABLE, THREE SERIALIZATIONS SOURCE REGION | Q3 UNITS | Q3 REVENUE EMEA | 4,120 | 1,884,000 APAC | 2,905 | 1,102,900 PAGE 11 PAGE BREAK LATAM | 1,340 | 602,300 NORAM | 8,715 | 4,006,900 PAGE 12, NO HEADER QUERY "What was NORAM revenue in Q3?" TRUE ANSWER: 4,006,900 SERIALIZED THREE WAYS NAIVE EXTRACTION NORAM 8,715 4,006,900 read left to right across the fold WRONG COLUMN: 8,715 MARKDOWN, SPLIT | NORAM | 8,715 | | 4,006,900 | header on page 11 NO SCHEMA: UNGROUNDED HEADER CARRIED REGION: NORAM Q3 UNITS: 8,715 Q3 REVENUE: 4,006,900 CORRECT: 4,006,900
FIG.02 Same table, three serializations. Two of them answer the question with a wrong number, and both look fine in the index.

For born-digital documents, this is a parsing problem rather than a reason to build multimodal retrieval, and the distinction is worth holding onto because the two get conflated constantly. What fixes it is a parser that preserves structure, a serialization that keeps every cell attached to its header, and the invariant that matters more than any chunking rule: a table is never split without carrying its header and structural context into every fragment. Note that's not "never split a table." Big tables have to be split, they exceed limits. The thing you must never do is sever rows from the schema that makes them mean anything.

Scans are different and I don't want to overclaim. Recovering rows, cells, spans and labels from a scanned table is genuinely a document-vision problem, and there's no parsing-only path to it. But note what kind of vision problem it is: table structure recognition, which outputs a grid with headers attached. That is not the same activity as handing the page to a generative model and asking it to describe what it sees, which is what "we need multimodal" usually turns into in practice. One produces structure you can serialize and cite. The other produces prose.

I want to be pointed about the history here, because it's the same complaint I made about entity resolution in Part 4. I was doing this in 2014. Machine learning to detect table regions in PDFs, then to recover the cell structure inside them, then to get the values out with the headers still attached. No language models, no RAG, nobody had a word for any of this. It was tedious and it mostly worked, and the failure modes were the three I just listed, in the same order of severity.

Which means table extraction is not a frontier problem that arrived with multimodal RAG. It's a mature subfield that predates the entire vocabulary we're using, and the current generation of tooling is genuinely better at it than what I had, but it is better at a known problem, not solving a new one. If you're building this and you feel like a pioneer, you aren't, and that's good news: there's a decade of prior work to read before you invent a heuristic on a Tuesday.

Opinion

I was going to tell you the model can't read your chart

I had a version of this section written that argued vision models don't actually read charts, they recognize the genre of chart and generate a plausible description of that genre, confidently inventing the numbers. I had the citation lined up. CharXiv, 2,323 real charts pulled from arXiv papers rather than the tidy templated ones the earlier benchmarks used, and on the reasoning questions the strongest model at the time managed 47.1 percent against a human baseline of 80.53.

Then I checked whether that still holds, and it doesn't. Anthropic's Opus 4.7 system card reports 82.1 percent on CharXiv reasoning without tools and 91.0 with a simple image-cropping tool, against 69.1 and 84.7 for the previous generation measured in the same document under the same settings5. I wouldn't line any of those up against the 80.5 percent human figure and declare a winner, because the harness has changed and human baselines don't transfer across harnesses. But the direction isn't ambiguous. The 47.1 percent result I was going to build this argument on is not representative of anything current, and if I'd shipped this piece on the schedule I originally wanted, I'd have shipped it wrong.

Small thing worth noticing on the way past: the earlier system card scored that same previous-generation model at 68.5 and 77.4, and the newer one scores it at 69.1 and 84.7, because the evaluation settings moved between documents5. Same model, same benchmark name, seven points apart in the tool-assisted column. Keep that in mind the next time someone puts two numbers from two sources on one slide, including when the someone is me.

One detail in those numbers is worth pausing on, because it points at where this section ends up. The tool that adds roughly nine points is an image cropper. The model gets to zoom in. That is the model looking harder at the same pixels, not the model going and fetching the data the chart was drawn from, and nothing about it produces anything a reader could check.

So let me make the argument that survives, which I think is the better one anyway.

The problem was never just that the model might read the chart badly. It's that a number read off pixels has a weaker audit trail than a value retrieved from structured data, and no accuracy improvement changes that.

Three tiers, and it's worth being precise about what separates them.

  1. Structured data. Provenance and a source value. You can point at the row it came from, and compare the number in the answer against the number in the table, mechanically, in a test, at three in the morning, with no human in the loop.
  2. Text. Provenance and an inspectable source. A skeptical reader opens the chunk and reads the sentence. Not machine-checkable in the same way, and it doesn't need to be, because the check is fast and unambiguous and anybody can do it.
  3. Chart pixels. Provenance is achievable, and section 05 is entirely about making sure you have it: document, page, bounding box, rendered crop. Build that. What you don't have, unless the underlying data still exists somewhere, is a source value to check the reading against. The final verification is another interpretation of the same pixels.

So the failure mode isn't "the model gets it wrong." At current accuracy it usually doesn't. The failure mode is that when it does, nothing in your system is structurally capable of noticing, and no dashboard will tell you which answers landed on the wrong side. That's the same complaint I've made about routers and graders in every article since Part 3, arriving somewhere it's harder to fix.

Which is why I'd retrieve charts and show them rather than summarize them into the context window. Put the figure in front of the user with its caption and its page and let them read the axis themselves. That moves the interpretation step to the person who has the context to judge it, and it's the strongest version of provenance available when the pixels are all you have. The model's job is finding the right figure, which it is good at now and was already good at when it was bad at reading them.

The other half of the CharXiv paper survives untouched, by the way, and it's the half that matters for production: small perturbations to the chart or the question dropped one model's accuracy by 34.5 points3. Benchmark charts are clean, rendered, consistent. Your charts are a screenshot of a chart, in a slide, exported to PDF, at 40 percent scale. Robustness to exactly that kind of degradation is the thing the headline number doesn't measure and your corpus consists entirely of.

And then the part that annoys me most

The numbers behind that chart almost always still exist somewhere structured. The CSV it was built from. The warehouse table behind the dashboard. The API the report generator called. Somebody rendered data into pixels, and now we're paying a model to render the pixels back into data, and losing precision in both directions.

Extracting values from an image is what you do when the values exist only as an image. That's the last resort. Every multimodal quickstart presents it as the starting line. This is the identical mistake I made with GraphRAG in Part 4, where I built an extraction pipeline over prose for relationships that already existed as foreign keys in a database the client owned. Same shape, different pixels. So now I ask the boring question first, in the first meeting: does the underlying data still exist anywhere, and can we just join to it?

Aside, unfinished

Something I've never properly measured and probably should. VLM captions converge on a register. Same sentence shapes, same vocabulary, same "this figure illustrates the relationship between" opener. If every figure in your corpus is described in the same voice, they all embed into roughly the same neighborhood, and figure-to-figure discrimination collapses even though each individual caption is accurate.

I've eyeballed the embeddings on one project and thought "that looks tighter than it should," and then the project ended. My guess is it's worse on technical corpora, where the figures are genuinely similar to begin with, than on the varied stuff benchmarks are built from. Might be a real effect worth a paper. Might be that reranking cleans it up and nobody cares. I don't know.

Provenance

What "which chunk?" means now

This series opened, in Part 1, with a claims bot citing a policy exclusion that didn't exist, and nobody being able to say which chunk it came from. Every article since has kept pulling on that thread. This is where the thread changes shape.

When the answer lives in a figure on page 47, a chunk ID isn't a citation. What you need is the page, the region of the page, the rendered crop, and the version of the parser that produced all three. Minimum viable result object:

@dataclass
class VisualResult:
    doc_id: str
    page: int
    bbox: tuple[float, float, float, float]   # normalized, origin top-left
    element_type: Literal["table", "figure", "chart", "text", "page"]
    render_uri: str          # the crop you show the user
    text_repr: str | None    # caption or serialization, if any
    parser_version: str      # which pipeline produced this
    caption_model: str | None

Everything you need to render the citation, and to delete it later

Two fields there are the ones people skip. render_uri, because showing the user the actual crop is the cheapest trust mechanism you will ever build, and it costs one image tag. And parser_version, for the same reason Part 4 wanted the index version in every log line: when a bad answer surfaces six weeks later, you need to know which pipeline was looking at that page, and by then you've changed the pipeline twice.

Then deletion, which is worse here than anywhere else in this series

Part 4 made the case that in a graph, a retracted document's text is deletable but its influence isn't, because it's already been abstracted into entity descriptions and community reports. Multimodal has the same problem with more debris. That page has been rendered to PNG, possibly at two resolutions, captioned into text, embedded as patches, thumbnailed for the UI, and cached by a CDN. The text is one delete statement. The derived artifacts are scattered across object storage that nobody wrote a manifest for, because they were generated by a batch job in week three and nobody thought of them as data.

FIG.03 — THE CHAIN RUNS BOTH WAYS FORWARD / INGEST SOURCE PDF doc_id PAGE RENDER page ELEMENT + BBOX bbox CROP render_uri EMBEDDING parser_version ANSWER EVERY ARROW CARRIES AN IDENTIFIER, OR THE CHAIN IS ONE-WAY REVERSE / DELETION SOURCE PDF PAGE RENDER ELEMENT + BBOX CROP EMBEDDING ANSWER Dashed nodes are the derived artifacts that usually have no back-pointer: the crop in object storage, the thumbnail in the CDN, the vectors in the index. The text is one DELETE. Its influence is not.
FIG.03 The chain has to work in both directions. Most pipelines only build it forwards.

Same conclusion as Part 4, arriving from a different direction: provenance links aren't an audit nicety, they're the deletion mechanism. Every derived artifact carries the ID of the source it came from, or erasure becomes a full re-index at full cost, every time. I'll repeat the honest note too: I have not found a good published treatment of erasure in multimodal RAG. If someone has written one I'd like to read it.

Field note

One index, two modalities, and a ranked list that was lying

This one isn't mine. I got called in on it, which is a more comfortable seat and I want to be clear that's the seat I was in.

A team had built visual retrieval into an existing text pipeline. Page images embedded alongside the text chunks, everything in one index, one query embedding, one ranked list. Elegant, one retrieval call, and exactly what I would have built if I'd got there first, which is worth admitting before I describe what was wrong with it.

The complaint was vague in the way the interesting ones always are: the visual search "doesn't really work." No error, no incident, no reproducible case anyone had written down. So the first hour went on the boring instrumentation, which was logging the modality of every result in the top 20 across a few thousand real queries.

The image side essentially never won. Not "won less often." Effectively never surfaced above the text results, including on queries where somebody had already confirmed by hand that the answer was in a figure.

What the instrumentation actually established was narrower than a diagnosis, and I want to keep those separate, because I've watched people skip this step and then fix the wrong thing. What we had was a modality-dependent score distribution: image candidates and text candidates were scoring on different scales, and a single threshold cutting across both was never going to select fairly between them. Their one ranked list was a text ranking with image entries appended below the fold, and the ranking had been saying so the whole time in a language nobody had asked it to speak.

Why the distributions diverged is a separate question and I'd hold it more loosely. The modality-gap literature offers a plausible geometric account: image and text representations occupy separate regions of a contrastive embedding space, present at initialization and preserved rather than closed by the training objective2. That's suggestive. It is not proof that this happens to your index, and it can't be, because these models are specifically trained to make cross-modal similarity meaningful in spite of that geometry, and plainly they often succeed. Training data, the retrieval task, and the score normalization all plausibly contributed here. I didn't run the ablations that would separate them, so I'm not going to tell you I know which one it was.

The useful part is that the fix didn't depend on knowing.

The reason it had gone months without being caught is the part I'd put on a wall. Text recall on the existing eval set was unchanged, because the text path was untouched. Latency was fine. Cost was fine. Every dashboard they had said the change was neutral, and the change was neutral, for everything anyone was measuring. What it wasn't neutral for was the queries the feature had been built for, which were a small fraction of traffic and weren't in the eval set, because the eval set predated the feature.

I recognized the shape of that within about ten minutes, and only because I'd shipped the same disease myself in Part 5: a router that failed by routing down, which made latency and cost look better while quality quietly dropped. Different mechanism, identical signature. So the rule from Part 5 transfers intact and I'll restate it, since I now have evidence it generalizes beyond my own mistakes: if this decision is wrong, what does it look like on the dashboard? If the answer is "fine," or "better," you need labeled examples before you ship, and the labels have to cover the case the feature exists for.

Retrieve each modality separately, with its own budget, so nothing depends on two score distributions being comparable, and merge afterwards, either at a multimodal reranker or over a shared textual representation with the image carried as provenance. That's the fourth pattern from section 02, and it sidesteps the calibration question instead of trying to win it. They'd been one architectural decision away from it the entire time.

Production

Most queries never need the vision path

A cheap modality router in front of the vision path is worth building, and Part 5's three options carry over unchanged: an embedding classifier over labeled queries, a small dedicated model, or an LLM router that you'll reach for first and profile last.

But I have to be more careful here than I was in Part 5, because a modality router is being asked a question it structurally cannot answer. Part 5's router predicted how much work a query needed, which is at least a property of the query. A modality router is being asked where the answer is stored, which is a property of your corpus. "What was revenue in Q3" could be answered by a sentence, a table or a chart, and nothing in those six words tells you which. The router can estimate that a question is likely to want visual evidence. It cannot know.

So treat the output as a prediction, not a fact. Use it to suppress the obviously unnecessary vision traffic, which is real volume and worth having. For anything the classifier isn't confident about, either retrieve both paths cheaply and let the merge stage sort it out, or decide after first-stage retrieval, when you can see what actually came back. That's the Part 5 rule arriving in a new domain, and I should have applied it here from the start: route on what's obvious from the query, grade on what actually came back. A modality router that silently downgrades ambiguous queries to text-only is precisely the failure the previous section was about.

The cost argument is the one from Part 5 and it hasn't changed. Whatever the vision path costs, running it on all traffic to serve the fraction that needs it adds that cost to your floor, and adds it to the queries that were supposed to be fast. The visual floor is higher than text people expect. Images are billed by patch area rather than by anything you'd intuit from file size, and current Claude models cap a single image at 1,568 visual tokens on the standard tier or 4,784 on the high-resolution tier, after any downscaling4. So ten candidate pages is on the order of 15,600 visual tokens at the standard cap and up to roughly 47,800 at the high-resolution one, before the question, the retrieved text and the answer are counted.

That is not the whole context window. Current models have room for it. It is a real number for cost, for time to first token, and for how much irrelevant material you're asking the generator to hold in mind, which is the one that doesn't show up on an invoice. Every provider tokenizes images differently, so run yours, but the shape holds: a page image is worth a great deal of text, and precision in what you retrieve pays for itself faster here than anywhere else in this series.

The field nobody logs

Per query, at minimum:

  1. Which modality path fired, and the router's confidence when it decided.
  2. Parser version and caption model version, on every visual result.
  3. Whether the image path contributed anything to the final answer.

That third one is the whole game and almost nobody has it. Not "was an image retrieved." Whether it survived reranking and made it into the generated context. That distribution is your health metric for this entire architecture, the same way termination reason was for agentic loops in Part 5. If images are retrieved on 30 percent of queries and contribute to 2 percent of answers, you are running a vision pipeline as an expensive no-op, and you'd never know from any other number you're collecting.

Aside

On parsers: I've now watched several teams buy the expensive document extraction product without testing it against the free one on their own documents, on the grounds that the expensive one must be better. Sometimes it is. On one corpus of scanned forms I looked at, it wasn't, and the gap ran the other way on tables specifically. The test costs you fifty representative pages and a spreadsheet with a human column. It's the same spreadsheet from Part 2 and Part 5, which is at this point less a recommendation than a personality trait, but I'm still right about it.

Evaluation

Three scorecards, and the labeling is harder

Short section. Part 7 is the whole story.

Three things can be wrong independently and need separate numbers. Parse quality, measured on the parser output against the source page, with no queries involved. Retrieval quality, ordinary recall against labeled relevant elements. And whether the retrieved visual element actually supported the answer, which is the hard one, because "was this the right image" is a genuinely harder judgment for a human labeler than "was this the right paragraph." Paragraphs argue for themselves. A figure requires the labeler to decide what the figure was even claiming.

Version your parser with your golden set, for the same reason Part 4 said to version the graph snapshot. If the parser changed between eval runs, you measured the parser, not your change.

Aside, running

I still do not have a golden dataset for my own side project. It was eight months in Part 1, nine in Part 5, and it is now ten, which means the rate of not doing it has been perfectly steady for the entire run of this series. In Part 4 I said I'd stop updating the number. Evidently not. Part 7 is about evaluation and I am going to write it anyway.

The thing I can't resolve, and I'd like someone to resolve it for me.

Page-image retrieval is the most convincing multimodal approach I've read about, and I have not met anyone running it at real scale on a real corpus. Every production system I've looked at closely is captioning at ingest and hoping, or parsing carefully and treating the figures as decoration. When I ask why, the answer is always some version of storage, or "we'd have to re-index everything," which are both real answers and neither of them is a measurement.

So: has anyone run it in production, on more than a pilot corpus, long enough to know what the index actually costs to keep current? Not a benchmark. Not a demo on 1,000 pages. The number I want is the one Part 4 asked for about graphs, which is what it costs to keep the thing true after the first year, and it's the number nobody publishes about anything.

If you've got it, I'd like to see it. And if you've measured caption clustering in your own index, the thing in the aside up in section 04, send me the numbers even if they show nothing, because a clean negative would let me stop wondering. Comment, DM, whatever's easiest.

Notes
  1. Faysse, M., Sibille, H., Wu, T., Omrani, B., Viaud, G., Hudelot, C., Colombo, P. (2025). ColPali: Efficient Document Retrieval with Vision Language Models. ICLR 2025. arXiv:2407.01449. arxiv.org/abs/2407.01449. Per-page indexing latencies are Table 5, Appendix B.4; the remark that captioning-based approaches can run to dozens of seconds per page is from the benchmark design section and describes the field rather than their own measurement. ↩↩↩
  2. Liang, V. W., Zhang, Y., Kwon, Y., Yeung, S., Zou, J. (2022). Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning. NeurIPS 2022. arXiv:2203.02053. arxiv.org/abs/2203.02053. The paper establishes the geometric separation and its persistence under contrastive training. It does not establish that cross-modal scores are therefore uncalibrated in any particular index, which is why both places it appears above treat it as a contributing account rather than a mechanism. ↩↩
  3. Wang, Z., Xia, M., He, L., Chen, H., Liu, Y., Zhu, R., Liang, K., Wu, X., et al. (2024). CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs. NeurIPS 2024 Datasets and Benchmarks Track. arXiv:2406.18521. arxiv.org/abs/2406.18521. The 34.5 point perturbation result is SPHINX V2 dropping from 63.2 to 28.6. ↩↩
  4. Anthropic, Vision, Claude Platform Docs. platform.claude.com/docs/en/build-with-claude/vision. Images are billed as visual tokens over 28 by 28 pixel patches. Verified September 2026; provider tokenization changes, so check before relying on the arithmetic. ↩
  5. Anthropic, Claude Opus 4.7 System Card (April 2026) and Claude Opus 4.6 System Card (February 2026). Both generations are measured under one harness in the 4.7 card, which is the comparison used above; the 4.6 card's own figures of 68.5 and 77.4 appear only in the paragraph about reporting drift. Scores use adaptive thinking at max effort, averaged over five runs. ↩↩

Benchmark figures move. Everything above was checked in September 2026, and the two system card numbers had already shifted once between documents, which is the point of that paragraph.

  • 01 · Choosing an architecture
  • 02 · Hybrid & reranking
  • 02a · Addendum: the middle ground
  • 03 · Corrective & Self-RAG
  • 04 · GraphRAG in practice
  • 05 · Adaptive & agentic
  • 06 · Multimodal
  • 07 · Evaluation & golden datasets

I publish a new article in this series every week: mostly what breaks in production, not what looked good in the demo. Part 7 is Evaluation and Golden Datasets, which is the article I've been deferring since Part 1 and referencing in every piece since, so it had better be good. Follow along if you want it when it lands; I post here first, before anywhere else.