How to Build Reliable Multimodal RAG for Visual Documents
Build multimodal RAG with two views of every document: the original visual page and a structured, searchable representation of its text, layout, tables, figures, and metadata. Retrieve from both, expand around the winning evidence, and ask a vision-language model to reason only over the smallest useful set of source pages. The answer should cite the exact page or region, not merely the document that contains it.
By Sahil Maheshwari
The short answer: preserve two views of every page
A visual document communicates through more than its words. Position can distinguish a table header from a footnote. Colour and line style can encode a chart series. An arrow can make a diagram directional. Extracting plain text can make those relationships disappear while still producing output that looks complete.
Reliable multimodal RAG therefore keeps two linked views. The source view is the original page image or rendered PDF page, with stable page and region coordinates. The retrieval view contains OCR text, native text when available, layout blocks, table cells, figure captions, visual embeddings, and document metadata. Neither view replaces the other. One makes search efficient; the other preserves what a reviewer actually saw.
The ColPali paper frames document retrieval as a visual problem and represents page images with multiple embeddings rather than reducing them to a single text passage. Its ViDoRe benchmark covers documents whose meaning depends on layout, figures, tables, and typography. That does not mean every system needs ColPali, but it does establish a useful design rule: page appearance can be retrieval evidence, not just presentation.
Start by separating three questions. Can the system find the relevant page? Can it identify the relevant region or relationship? Can it answer without inventing missing detail? A strong answer model cannot recover evidence that ingestion discarded or retrieval never supplied.
Classify what the question asks the document to do
Not every question about a PDF is multimodal. ‘What is the policy number?’ may be exact text lookup. ‘Which building is shaded as high risk?’ depends on a legend and a map. ‘Did incidents fall after the intervention?’ may require reading axes, comparing chart values, and linking the result to a paragraph that defines the intervention. Route these cases differently instead of sending every page through the most expensive model.
A practical router can distinguish exact text, semantic text, table lookup, chart reasoning, diagram or map interpretation, image inspection, and cross-page synthesis. The categories may overlap. Use query terms, document type, retrieved block types, and a lightweight classifier to select candidate indexes. When confidence is low, search more than one representation and fuse the results.
This distinction matters because visual questions are not solved by OCR alone. ChartQA includes questions that require both visual and logical reasoning over charts, and its models combine visual features with the underlying data table. A chart pipeline should preserve titles, axes, units, legends, labels, plotted series, and any recoverable data—not just the text printed around the figure.
Likewise, document question answering depends on structure. The original DocVQA benchmark assembled more than 50,000 questions over 12,000 document images and reported a notable gap on questions that need document structure. When a user asks ‘Who approved the amount in the right-hand column?’, the spatial relationship between the name, signature, and column is part of the query.
Build linked representations, not one universal chunk
Render each page at a legible resolution and retain the original file. Extract native text where available, run OCR where needed, and record bounding boxes. Detect headings, paragraphs, lists, tables, figures, captions, headers, footers, and page numbers. Store every derived object with document ID, page, region coordinates, extraction method, and confidence.
Keep tables as cell structures as well as page regions. Keep a chart image beside any extracted data series and description. Link a figure to its caption and to nearby paragraphs that introduce or interpret it. If a diagram has nodes and edges that can be extracted reliably, store that graph as a derived view while preserving the image. These links allow retrieval to move from a matching caption to the figure, or from a chart to the paragraph that defines its units.
OCR remains useful, but it is an error-prone dependency. The Donut paper developed an OCR-free document understanding model partly because OCR errors can propagate into later stages. You do not need to choose one camp. Index native or OCR text for precise search, and keep a visual route for cases where recognition is uncertain, handwriting matters, or layout carries the answer.
Do not turn generated image descriptions into source truth. A caption produced by a model can improve recall, but it is a derived claim that may omit a label or misread a relationship. Mark it accordingly and require the answering step to inspect the original region before using it as evidence.
- Source layer: original file, page image, page number, region coordinates, and checksum.
- Text layer: native text, OCR spans, reading order, headings, captions, and confidence.
- Structure layer: tables, chart data, figure links, layout blocks, and document hierarchy.
- Search layer: lexical, semantic, visual, and metadata indexes that point back to the same source regions.
Retrieve narrowly, then expand around the evidence
Use parallel retrieval paths. Lexical search is strong for identifiers, labels, quoted phrases, and rare terms. Text embeddings find paraphrases in OCR or native text. Visual page embeddings can find layouts and figures that text extraction represents poorly. Structured queries work best for a known table column, date, or document type. Merge candidates by document and page, then rerank with the query and the available representations.
The first match is often a doorway rather than the complete answer. A chart may rely on a legend above it, a caption below it, and a definition on the previous page. Expand to linked regions and a small page neighbourhood. Follow explicit references such as ‘see Figure 4’ or ‘continued on next page’. Keep expansion bounded so the model sees coherent evidence rather than dozens of loosely related pages.
Long documents make this retrieval step unavoidable. M3DocRAG addresses questions over multiple visually rich pages and documents by combining a multimodal retriever with a multimodal language model. Its motivation is practical: single-page visual models cannot simply ingest a large collection, while OCR-only RAG can lose visual evidence. A useful production pattern is therefore retrieval first, visual reasoning second.
Imagine an engineering inspection report. A text section names a cracked support; a plan marks its location with a symbol; a photograph shows the crack; and a table assigns the repair priority. Search may begin from the support name, but the answer needs four linked regions. Preserve those relations in the evidence bundle and show them to the user rather than collapsing them into one opaque summary.
Retrieve the page that looks relevant, then gather the caption, legend, neighbouring region, and referenced page that make it interpretable.
Generate from a compact, visible evidence bundle
Pass the answering model page crops or full pages, plus extracted text and structured data. Label every item with document, page, and region identifiers, and distinguish original from derived content. For numeric chart questions, ask it to state the units and calculation. Where exact data is available, calculate outside the language model and use the visual model to verify the selected series and labels.
Citations should land at the page and, when the interface supports it, highlight the relevant region. A document-level citation is too broad when a 90-page report contains several similar charts. If a claim combines modalities, cite each supporting region: the paragraph that defines the measure, the chart that shows the trend, and the table that supplies the exact value.
Separate observation from inference. ‘The legend labels the red zone high risk’ is a direct visual observation. ‘This site should close’ is a recommendation that may require policy or expertise not present in the document. If a label is unreadable, the answer should say so and request a clearer scan. The goal is not to force every visual question into a fluent response; it is to keep the response within the evidence that can actually be inspected.
Multimodal evidence can also conflict. OCR may read a chart label differently from the page image, or an extracted table may omit a merged cell. Preserve both results, surface the disagreement, and prefer the original page for adjudication. Do not silently let a derived representation overrule its source.
Evaluate retrieval and reasoning by modality
Build an evaluation set from the documents people use. Label required pages and regions, modalities, the expected answer, and whether enough evidence exists. Include digital PDFs, scans, rotated pages, small fonts, dense tables, similar charts, handwritten marks, multi-page figures, and absent answers.
Measure retrieval before answer quality. Track whether the required page and region appear in the candidate set, whether the correct legend or caption was included, and how much irrelevant material was added. Then score OCR or extraction accuracy, table and chart reasoning, cross-page synthesis, citation completeness, faithfulness, and correct abstention. Break results out by document type and modality; one average score can hide a system that handles prose well but fails on diagrams.
MMLongBench-Doc evaluates long-document questions whose evidence can span text, images, charts, tables, layout, and pages, including unanswerable cases. MMDocRAG goes further into multimodal evidence selection and integration. Public benchmarks are useful stress tests, but a rollout decision should still depend on a smaller, reviewed set that reflects your own files and consequences.
Record cost and latency as part of quality. Cache immutable page representations, process documents incrementally, and reserve expensive visual reasoning for candidates that survive retrieval. Test whether a cheaper text route answers straightforward questions just as well.
Know when multimodal RAG is worth the complexity
A text-first system is usually enough for born-digital documents whose questions are answered by clean paragraphs. Multimodal RAG earns its cost when layout changes meaning, figures carry evidence, scans resist extraction, or an answer must connect text with a chart, image, table, or diagram. Use observed failure cases to justify each new modality rather than adding visual models by default.
A sensible first release handles one document family well. Choose, for example, inspection reports, research papers, claims files, or technical manuals. Map the recurring question types, build the linked representations they require, and make citations visually inspectable. Only then widen the corpus. A small system that reliably shows why an answer is grounded is more useful than a broad system that merely accepts every PDF.
If a real visual-document question keeps losing the chart, diagram, or page relationship that makes the answer clear, share the anonymised document shape and question at sahil@granveo.com. The failure usually reveals whether the missing layer is ingestion, retrieval, evidence expansion, or visual reasoning.
Use multimodal RAG when the page itself is evidence. Keep text-only retrieval when text already preserves the answer.
Sources and further reading
- 1ColPali: Efficient Document Retrieval with Vision Language Models
Faysse et al., 2024 — Introduces multi-vector visual page retrieval and the ViDoRe benchmark for visually rich documents.
- 2M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
Cho et al., 2024 — Combines multimodal retrieval and generation for questions over collections of visually rich documents.
- 3DocVQA: A Dataset for VQA on Document Images
Mathew et al., 2020 — Introduces a large document-image question answering benchmark and analyses structural reasoning.
- 4ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Masry et al., Findings of ACL 2022 — Benchmarks chart questions that require visual interpretation, logical reasoning, and chart data.
- 5OCR-free Document Understanding Transformer
Kim et al., ECCV 2022 — Presents an OCR-free approach motivated by the cost and error propagation of OCR pipelines.
- 6MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
Ma et al., NeurIPS 2024 — Evaluates long-document questions across text, images, charts, tables, layout, pages, and unanswerable cases.
- 7Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
Dong et al., NeurIPS 2025 — Benchmarks multimodal evidence selection, integration, and grounded answer generation.
Continue the conversation
Where does context get lost in your work?
I am speaking with researchers, founders, and operators about the handoffs, evidence, and decisions that are hardest to keep connected.