All field notes
Retrieval and context10 min read

How to Build RAG for Email and Conversation Threads

Build email RAG around conversations, not isolated messages. Preserve message headers, reply relationships, participants, timestamps, quoted spans, and attachments; retrieve at message and thread level; reconstruct the sequence that changed the state of the work; then answer with citations to the exact messages and attachment passages that support each claim.

By Sahil Maheshwari

An email question traced through messages, quoted replies, attachments, and decisions to a chronological source-backed answerTRACE A CLAIMSOURCE 03WORKING CLAIMContext is more usefulwhen its origin survives.3 sources · 2 relationships · 1 open question

The short answer: reconstruct before retrieval

Email looks like a text corpus, but its meaning is carried by sequence and reply structure. ‘Approved’ may reverse an earlier objection. ‘Please use the attached version’ can make every previous attachment obsolete. A forwarded message may supply evidence from another conversation while the current sender adds a new interpretation. Index those messages independently and similarity search can retrieve the right words from the wrong state of the work.

A reliable email RAG system has four layers. First, ingest each message and attachment as an immutable source. Second, reconstruct threads and links between quoted and replied-to passages. Third, derive reviewable objects such as decisions, requests, commitments, entities, and document versions. Finally, retrieve both the relevant evidence and enough conversational context to interpret it.

Consider a hypothetical commercial insurance placement. A broker sends risk details, an underwriter asks for a survey, the broker supplies it in a later attachment, and the underwriter offers terms subject to one condition. The question ‘Did the insurer quote, and what remains open?’ cannot be answered from the latest email alone. It needs the offer, the condition, the attachment history, and the absence or presence of a later response.

Do not begin by embedding every inbox message. Begin with the questions people need answered and identify which message relationships make those answers trustworthy.

The retrieval unit is not necessarily the email. It may be a message, a quoted passage, an attachment section, a decision, or a complete thread state.

Store each message as a structured event

Keep the raw message and a parsed representation. Record the provider message ID, Internet Message-ID, thread ID, In-Reply-To and References headers, sender, recipients, sent and received times, subject, labels or folders, and source mailbox. Normalise addresses for lookup, but retain the original display names and header values.

RFC 5322 defines the Internet message format and shows how Message-ID, In-Reply-To, and References connect reply messages. Use those fields before falling back to subject lines, participant overlap, or time proximity. Provider thread IDs are useful, but they express a provider’s grouping decision rather than a complete model of the underlying conversation.

Parse the body into authored text, quoted text, signature, disclaimer, and forwarded-message blocks without destroying the raw form. Preserve HTML and plain-text alternatives, character set, content type, and byte-level attachment identity. RFC 2045 specifies the MIME foundations for non-ASCII text, non-text content, and multipart message bodies. Treat MIME structure as evidence about what the sender actually included.

An attachment is not a bag of extracted text. Store its filename, media type, checksum, parent message, sender, time, and extraction status. Parse the document using the method appropriate to its format, then cite back to its page, cell, slide, image region, or other stable locator. If the same file appears under different names, link the copies; do not silently assume they are the same version without comparing their content.

  • Message identity: provider ID, Message-ID, thread ID, mailbox, and immutable raw source.
  • Conversation edges: replies to, references, forwards, quotes, and same-subject fallback links.
  • People and time: sender, recipients, roles, sent time, received time, and timezone.
  • Content parts: authored text, quoted spans, HTML, plain text, signatures, inline media, and attachments.

Retrieve messages, threads, and attachments together

Create indexes at several levels. Message-level search finds a precise statement. Thread-level representations capture the subject, participants, entities, and evolving topic. Attachment indexes find evidence inside files. Structured indexes cover exact addresses, policy numbers, project names, dates, and status fields. Search them together and merge candidates by thread and source identity.

Use lexical retrieval for names, identifiers, exact phrases, and subject lines. Use embeddings for paraphrases and intent. Apply mailbox, participant, time, project, label, and permission filters before reranking. Recency can help, but it cannot replace authority: the newest email may be a reminder rather than the message that approved the decision.

When a message is retrieved, expand around it deliberately. Include its parent, the quoted passage it answers, later replies that may change its status, and referenced attachment sections. The Gmail API, for example, exposes a thread resource that groups replies and allows all messages in a conversation to be retrieved in order. Other providers differ, so keep the internal conversation model independent of one API.

Avoid sending an entire long thread to the model by default. Retrieve candidate evidence first, then build a compact timeline containing the relevant turns, their relationship, and unresolved gaps. Preserve links to excluded messages so the system can expand when the question changes.

A message match finds a statement. Thread expansion tells you whether that statement was answered, superseded, or left open.

Model decisions and changing work state

Most useful inbox questions ask about state: what was requested, who owns it, what was decided, which version is current, and what remains unresolved. Extract these objects as candidates with source links, not as free-floating facts. A commitment should include actor, action, due date if stated, status, and the message span that created or changed it.

Email dialogue is not only informational. Messages propose, challenge, accept, reject, defer, clarify, and withdraw. The LEDA dataset defines dialogue acts for studying decision-making in public IETF email archives. Its labels are not a universal business ontology, but the work supports a useful design choice: identify the communicative role of a message before converting it into a durable decision.

Represent changes rather than overwriting history. ‘Quote valid until Friday’ creates a time-bounded state. A later extension changes its validity; it does not edit the original message. ‘Use v3’ supersedes a document version, while ‘I reviewed v3’ does not necessarily approve it. Keep those distinctions visible.

Summaries are derived views, not sources. EmailSum reports challenges in identifying sender intent and participant roles, and found weak correlation between common automatic summarization metrics and human judgments on its dataset. Use a thread summary to orient retrieval, but return to the underlying messages before asserting a decision or obligation.

  • Request: who asked for what, from whom, when, and with which stated deadline.
  • Decision: proposal, authority, outcome, scope, conditions, and supporting messages.
  • Commitment: owner, action, due date, completion evidence, and current status.
  • Document state: version, sender, attachment identity, review status, and supersession link.

Synchronize changes and deletion correctly

Inbox knowledge changes continuously. New replies alter thread state; labels and folder placement change; drafts become sent messages; attachments are replaced; users delete content or lose access. The index must update derived objects and summaries when their supporting evidence changes.

Use incremental provider history where available and keep a recovery path for full reconciliation. Google’s Gmail synchronization guidance distinguishes full and partial sync, with history records for changes after a stored history ID and a required full sync when that ID falls outside the available range. The provider-specific mechanism will vary, but the invariant is the same: an index is not current merely because ingestion once succeeded.

Propagate deletion and permission changes through message chunks, attachment extracts, caches, summaries, and decision objects. Mark derived state invalid when its only evidence disappears. Do not let a cached answer expose a thread the current user cannot retrieve.

Keep personal mailboxes separate unless policy explicitly authorises a shared view. Recipient lists are not a complete access-control system: forwarded content, delegated mailboxes, confidential attachments, and legal holds may follow different rules. Reuse the organisation’s identity and retention controls rather than inventing access from the text of the email.

Cite and evaluate RAG for email threads

An answer should distinguish recorded facts, inferred state, and missing evidence. Cite the exact message with sender, date, subject, and the relevant authored span. For attachment claims, cite both the parent message and the attachment location. For an open item, show the request and state that no completion evidence was found within the searched scope and cutoff time.

Build evaluation questions from real inbox work: find the latest approved attachment, identify unanswered requests, reconstruct why a decision changed, compare two offers, list commitments due this week, and explain what evidence is still missing. Include broken threads, altered subjects, forwards, duplicate attachments, timezone boundaries, ambiguous ‘yes’ replies, and a decision later reversed.

Score message and attachment recall, thread reconstruction, chronology, participant attribution, version selection, decision-state accuracy, citation completeness, permission enforcement, deletion propagation, and correct abstention. Review summaries with humans who understand the workflow; surface similarity is a weak proxy for whether a thread’s state was captured correctly.

Email RAG is unnecessary when the task is a deterministic filter that the provider already supports, such as messages from one sender during a fixed date range. It earns its complexity when the answer depends on connecting messages, files, people, and changes over time—and when that connection remains inspectable.

If a real inbox question keeps returning the right message but the wrong thread state, share the question and anonymised conversation shape at sahil@granveo.com. The reply pattern usually reveals whether the missing layer is threading, versioning, retrieval, or decision tracking.

A trustworthy email answer is a small, cited timeline—not a confident summary detached from the conversation that produced it.

Sources and further reading

  1. 1
    RFC 5322: Internet Message Format

    RFC Editor — Defines Internet message fields, including Message-ID, In-Reply-To, and References used to connect replies.

  2. 2
    RFC 2045: Multipurpose Internet Mail Extensions Part One

    RFC Editor — Defines MIME foundations for character sets, non-text content, and multipart message bodies.

  3. 3
    Manage threads

    Google for Developers — Documents Gmail’s thread resource and ordered retrieval of messages within a conversation.

  4. 4
    Synchronize clients with Gmail

    Google for Developers — Explains full and partial synchronization, history IDs, change records, and recovery through full sync.

  5. 5
    EmailSum: Abstractive Email Thread Summarization

    Zhang et al., ACL-IJCNLP 2021 — Provides a human-annotated thread summarization dataset and analyses intent, participant-role, and evaluation challenges.

  6. 6
    LEDA: a Large-Organization Email-Based Decision-Dialogue-Act Analysis Dataset

    Karan et al., Findings of ACL 2023 — Introduces dialogue acts for analysing decision-making in real public organisational email archives.

  7. 7
    Identifying Implicit Quotes for Unsupervised Extractive Summarization of Conversations

    Kano et al., AACL-IJCNLP 2020 — Studies how reply relationships and implicit quotations can reveal salient passages in email and conversation data.

Continue the conversation

Where does context get lost in your work?

I am speaking with researchers, founders, and operators about the handoffs, evidence, and decisions that are hardest to keep connected.