All field notes
Retrieval and context10 min read

How to Build RAG for Spreadsheets and Tabular Data

Build spreadsheet RAG as a retrieve–execute–explain system. Preserve workbook, sheet, table, row, column, formula, unit, and cell identities during ingestion. Retrieve the relevant table and schema first, use a deterministic engine to filter and calculate over the selected data, then ask the model to explain the result with citations to the exact cells and operations used.

By Sahil Maheshwari

A spreadsheet question routed through schema retrieval, row filtering, calculation, and cell-level citation before producing an answerA DURABLE RESEARCH LOOP1Collectfiles + notes2Connectclaims + links3Questiongaps + tension4Createbrief + draftnew questions return to the map

The short answer: a spreadsheet is not a bag of text

Ordinary RAG works with passages because their meaning is carried largely by words and nearby sentences. A spreadsheet distributes meaning across two dimensions. The value `12.4` may mean nothing without its column header, row label, unit, sheet, reporting period, formula, and workbook version. Flatten those elements into unrelated text chunks and retrieval may return the right number with the wrong meaning.

The reliable pattern has three stages. Retrieve the table, schema, and notes that define the question. Execute the selection, join, aggregation, or arithmetic over typed values. Then generate an answer with the cell ranges and operations that produced it.

Consider a hypothetical broker workbook with one sheet for submissions, another for quotes, and a third for endorsements. “Which open property risks have only one quote, and what is their total insured value?” requires a status filter, a product filter, a group count by risk, a second filter, and a sum over a different column. Similarity search alone does not perform those operations reliably.

Use the language model to interpret the question, locate candidate data, and explain the result. Use spreadsheet formulas, SQL, dataframe operations, or another constrained execution engine for the calculation. The model should not silently reconstruct a table in its prompt and do the arithmetic from memory.

Retrieve the evidence, execute the operation, and cite the cells. Do not ask one prompt to impersonate all three stages.

Preserve the tabular model during ingestion

Keep the original hierarchy: workbook, sheet, table, header rows, data rows, columns, and cells. Give each level a stable ID. Record coordinates, raw and display values, types, number formats, formulas, merged ranges, hidden state, comments, and errors. Preserve the formula and last calculated result, plus when and by which engine it was produced.

Column metadata should include the original label, normalised name, type, unit, allowed values, null semantics, and business meaning. A blank premium may mean unknown, not applicable, not yet quoted, or zero. Those states must not collapse into one empty string. Likewise, `5` could mean 5%, ₹5 crore, or five claims.

The W3C Model for Tabular Data and Metadata on the Web provides a useful standard vocabulary for tables, rows, columns, cells, annotations, datatypes, and metadata. A production implementation can use SQL, Parquet, JSON, or a property graph instead of RDF. The portable lesson is to represent the table explicitly rather than infer its structure again for every question.

Keep source location beside every normalised record: workbook ID and version, sheet name, table or range, row key, column key, and cell address. If a spreadsheet is re-uploaded, do not overwrite its identity. Link versions and decide whether each query needs the current view or the historically effective one.

  • Workbook: owner, version, checksum, permissions, calculation mode, and retrieved time.
  • Sheet or table: title, range, header depth, row keys, surrounding notes, and relationships.
  • Column: stable ID, source label, type, unit, vocabulary, and null rules.
  • Cell: coordinate, raw and displayed values, formula, error state, and provenance.

Retrieve schema and values at different levels

Indexing every cell as an independent embedding creates noise and destroys context. Start with a catalogue of workbooks, tables, column descriptions, representative values, date coverage, owners, and access rules. Retrieve likely tables and columns before searching rows or cells within them.

Use several retrieval signals. Lexical search is strong for policy numbers, codes, exact headers, and quoted values. Embeddings help match “amount covered” to a column labelled “TIV” when the glossary links total insured value to that abbreviation. Metadata filters enforce workbook, date, client, product, and permission boundaries. Value indexes find entities stored in the table, while schema retrieval finds where the answer can be computed.

For very large tables, retrieve a compact subtable rather than serialising the whole sheet. The NeurIPS paper TableRAG: Million-Token Table Understanding with Language Models combines query expansion with separate schema and cell retrieval. Its results support this direction on large benchmark tables, but they do not determine the best index or thresholds for a messy operational workbook.

Return enough context to interpret the selected cells: hierarchical headers, units, row identifiers, footnotes, and the formulas behind derived values. Retrieval that finds a cell but omits its denominator, period, or qualifier has not found usable evidence.

Schema retrieval answers where and how to query. Cell retrieval supplies values. Treating them as the same search problem weakens both.

Route each question to the right execution path

Classify the requested operation before generating a query. A lookup asks for one recorded value. A filter asks which rows satisfy conditions. An aggregation needs count, sum, average, minimum, maximum, grouping, or sorting. A derived calculation combines values with an explicit formula. A hybrid question may require a paragraph or footnote to define which rows or operation apply.

Generate a constrained plan against the retrieved schema, not arbitrary code against the environment. Resolve columns to stable IDs, validate types and units, enforce permissions, and allow only approved operations. Run it in a sandboxed SQL, spreadsheet, or dataframe engine and capture the query, rows, warnings, and result.

Table question-answering research shows why structure and execution matter. TaPas encodes table structure, selects cells, and can apply an aggregation operator. FinQA pairs financial questions with annotated reasoning programs. These are research tasks rather than turnkey architectures, but both separate evidence selection from the operation needed to reach the answer.

Do not recalculate a workbook casually. Some formulas depend on macros, external links, volatile functions, hidden assumptions, or a calculation engine with different semantics. Decide whether the source’s stored result is authoritative or whether your system may recompute it. If recomputation is allowed, record the formula, inputs, engine, version, locale, and time.

  • Lookup: retrieve an existing value with its row, column, and version.
  • Filter or join: execute conditions over typed fields and explicit keys.
  • Aggregation: compute over the eligible row set, preserving included rows.
  • Derived answer: store the formula or program and every input cell.
  • Hybrid answer: combine executed table results with cited explanatory text.

Keep tables connected to surrounding text

Spreadsheet questions often depend on definitions outside the cells. A footnote may exclude cancelled records. An email may define “open” for the current review. A policy document may explain whether a deductible is per occurrence or aggregate. Store these materials as separate sources and connect them to the table, column, range, or formula they qualify.

TAT-QA was built from financial reports containing both tables and text, with questions that often require arithmetic, comparison, counting, or sorting. Its baseline extracts relevant cells and text spans before symbolic reasoning. The result is a useful reminder: retrieval must find both the numbers and the language that gives them meaning.

An SQL-based approach can also support multi-step questions. The 2025 EMNLP paper TableRAG: A Retrieval Augmented Generation Framework for Heterogeneous Document Reasoning alternates query decomposition, text retrieval, SQL generation and execution, and intermediate answers. Treat that as one demonstrated design, not evidence that every workbook needs agentic iteration.

Build explicit links such as `column defined by glossary term`, `range extracted from document table`, `formula depends on cells`, and `sheet attached to email thread`. Those connections let the system retrieve an explanation with the calculation instead of presenting a precise number stripped of its business context.

Cite the cells and the operation

A useful answer should show more than a workbook filename. Cite the workbook version, sheet, table or range, relevant row keys, column labels, and exact cells. For an aggregation, expose the filtered row set or a reviewable summary of it. For a derived answer, show the operation and input ranges in a readable form.

Separate retrieved facts from computed facts. “Quote count is 1” may be computed by grouping quote rows. “Total insured value is ₹40 crore” may sum four cells. The explanation should say that those values were calculated, identify the operations, and link back to each input rather than implying the final number appeared verbatim in the sheet.

Keep display precision distinct from calculation precision. Preserve currency, percentages, negative-number conventions, and locale. Never convert `1.2` into `1.2 million` unless the unit comes from the schema or source. If two candidate tables use different currencies or periods, do not combine them until an explicit conversion or alignment rule is applied.

For sensitive workbooks, citations must obey the same access controls as execution. A user who can see an aggregate may not be allowed to inspect every contributing row. Define whether the system can return a cell, a redacted range, an aggregate-only explanation, or no answer at all. Hidden sheets and hidden columns are presentation features, not authorization rules.

The evidence trail for a computed answer is the selected cells plus the operation—not a citation to the spreadsheet as a whole.

Evaluate spreadsheet RAG from retrieval to calculation

Test each stage. Did routing identify the question type? Did retrieval find the right workbook, table, columns, and rows? Was the generated operation valid? Did execution produce the expected result? Did the answer preserve units, scope, and citations?

Use real failure cases: repeated headers, merged cells, multiple tables on one sheet, blank-versus-zero values, hidden rows, stale formula results, dates stored as text, duplicated IDs, totals mixed with detail rows, changed column names, external links, and conflicting workbook versions. Add adversarial questions whose requested metric is not defined so the system must ask or abstain.

The 2026 T²-RAGBench evaluates retrieval before numerical reasoning over text-and-table data, rather than assuming the correct context is already supplied. Public benchmarks remain useful probes, but your evaluation set must reproduce the spreadsheets, permissions, definitions, and question language of the real workflow.

Track table and column recall, row-set precision and recall, execution success, numerical accuracy, unit accuracy, citation completeness, correct abstention, latency, and cost. Compare against a direct SQL or spreadsheet baseline. If users ask a fixed set of deterministic metrics over a clean database, a governed report or BI query may be safer and cheaper than RAG.

Spreadsheet RAG earns its complexity when people ask changing natural-language questions across many tables and need the result connected to notes, documents, and evidence. It is not a substitute for repairing an undefined metric, a broken workbook, or missing source data.

If a real spreadsheet question keeps returning the right-looking number from the wrong rows, share the workbook shape and question at sahil@granveo.com. That failure usually reveals whether the missing layer is schema, retrieval, execution, or provenance.

  • Retrieval: correct workbook, table, schema, rows, and explanatory text.
  • Execution: valid plan, permitted operations, correct types, and reproducible result.
  • Grounding: every stated and computed value maps to cells and transformations.
  • Robustness: versions, nulls, units, formulas, ambiguity, and missing data.
  • Safety: row and column permissions remain enforced in results and citations.

Sources and further reading

  1. 1
    Model for Tabular Data and Metadata on the Web

    W3C — Defines an abstract model for annotated tables, columns, rows, cells, datatypes, and metadata.

  2. 2
    TaPas: Weakly Supervised Table Parsing via Pre-training

    Herzig et al., ACL 2020 — Demonstrates table-aware encoding, cell selection, and optional aggregation for table question answering.

  3. 3
    TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance

    Zhu et al., ACL 2021 — Provides questions that combine financial tables, text spans, and symbolic numerical reasoning.

  4. 4
    FinQA: A Dataset of Numerical Reasoning over Financial Data

    Chen et al., EMNLP 2021 — Pairs expert-written financial questions with annotated reasoning programs for explainable calculation.

  5. 5
    TableRAG: Million-Token Table Understanding with Language Models

    Lin et al., NeurIPS 2024 — Separates schema and cell retrieval to reduce context size for very large tables.

  6. 6
    TableRAG: A Retrieval Augmented Generation Framework for Heterogeneous Document Reasoning

    Yu et al., EMNLP 2025 — Presents an iterative text-retrieval and SQL-execution framework for questions spanning tables and text.

  7. 7
    T²-RAGBench: Text-and-Table Benchmark for Evaluating Retrieval-Augmented Generation

    Strich et al., EACL 2026 — Evaluates retrieval and numerical reasoning together over real-world text-and-table question-answer pairs.

Continue the conversation

Where does context get lost in your work?

I am speaking with researchers, founders, and operators about the handoffs, evidence, and decisions that are hardest to keep connected.