All field notes
Retrieval and context10 min read

How to Keep a RAG Knowledge Base Fresh as Sources Change

Keep a RAG knowledge base fresh by treating every source as a versioned stream, not a folder that is periodically copied. Give each source a stable identity, record its revisions and effective dates, consume provider change feeds where possible, and propagate every update or deletion through chunks, embeddings, summaries, permissions, caches, and citations. Freshness becomes dependable only when the system can show which revision answered a question and when that revision was valid.

By Sahil Maheshwari

A source change moving through detection, versioning, parsing, indexing, retrieval, and freshness validationA DURABLE RESEARCH LOOP1Collectfiles + notes2Connectclaims + links3Questiongaps + tension4Createbrief + draftnew questions return to the map

The short answer: freshness is a change-processing problem

Re-embedding the corpus every night sounds safe, but it leaves important questions unanswered. What changed since the last successful run? Did a deleted page disappear from the vector index? Did a revised table invalidate a summary generated from the old version? Can the system answer both ‘What is the current limit?’ and ‘What limit applied on the claim date?’ A rebuild schedule alone cannot express those guarantees.

Model the pipeline as a sequence of state transitions. A source revision arrives; the system records it, parses it, derives retrievable units, publishes a new searchable state, and retires artifacts derived from the previous revision. The update is complete only when all downstream representations agree. A current document paired with an old embedding is still a stale system.

The need is measurable. StreamingQA evaluates question answering against a timestamped stream of news, while FreshLLMs introduces FreshQA for questions whose answers may change over time. Both make the same operational point: retrieving external evidence helps only when the evidence collection itself reflects the relevant time.

Start with a freshness contract for each collection: how quickly changes should become searchable, which source is authoritative, how deletions behave, how much history must remain available, and what the system should say when synchronization is late. Different sources deserve different targets. A policy manual may tolerate minutes; a slowly revised research archive may tolerate a day.

A successful connector run is not the freshness goal. The goal is a searchable state whose content, permissions, history, and citations all refer to the right source revisions.

Separate source time from system time

Every revision needs at least two clocks. Valid time says when the information applies in the source domain. System time says when your knowledge base observed and published it. Suppose an insurance endorsement uploaded on 12 June states that it became effective on 1 June. The two dates answer different questions. Current guidance should use the effective revision; an audit of what staff could have known on 5 June must respect the later observation date.

Do not overwrite the previous document in place. Store an immutable revision with a stable source ID, a provider revision or ETag, a content hash, observed-at time, effective interval when known, permission snapshot, and status such as current, superseded, or deleted. OWL-Time provides a useful vocabulary for reasoning about instants and intervals, even if your implementation uses ordinary database columns rather than an ontology.

Keep provenance between revisions and derived objects. The W3C PROV-O standard describes derivation, revision, generation, and invalidation relationships. In practical terms, each chunk, embedding, extracted claim, generated summary, and cache entry should point back to the exact source revision that produced it. That dependency graph tells the update worker what must be regenerated or invalidated.

Some sources do not expose an effective date. Record that uncertainty instead of copying the file's modified timestamp into a business-time field. Modified time may reflect an upload, format conversion, or metadata edit. It is evidence about the repository, not necessarily about when a policy or fact became true.

Detect changes with cursors, validators, and reconciliation

Prefer a provider's change protocol over listing every object repeatedly. The Google Drive changes API uses page tokens to retrieve changes after a known state; its watch notifications signal that changes are available but do not contain the changes themselves. Microsoft Graph delta query similarly returns created, updated, and removed entities plus an opaque link for the next synchronization. Persist the cursor only after every page before it has been durably processed.

For web resources without a change feed, use conditional requests. RFC 9110 defines ETag and Last-Modified validators as well as If-None-Match and If-Modified-Since requests. ETags are preferable when available because modification dates can be coarse or inconsistently maintained. Hash the normalized content after download so metadata-only updates do not trigger expensive parsing and embedding.

Webhooks reduce detection delay; they do not replace a change ledger. Notifications can be duplicated, delivered out of order, or missed during an outage. Make events idempotent by source ID and revision, then run bounded reconciliation on a slower schedule. The reconciliation compares the provider's current inventory with local state, repairs missing updates, detects vanished objects, and resets an expired cursor through an explicit full scan.

Track connector health separately from document age. Useful fields include last attempted sync, last successful cursor advance, oldest unprocessed event, and last full reconciliation. A green process heartbeat can hide a cursor that has not advanced for hours.

  • Stable source ID: survives renames and moves.
  • Provider revision or validator: distinguishes source states.
  • Content hash: avoids rebuilding unchanged content.
  • Durable cursor: resumes without silently skipping a page.
  • Reconciliation checkpoint: proves the incremental view matches the source.

Propagate each change through every derived artifact

Once a revision is accepted, build its replacement artifacts before making them visible. Parse the new content, assign deterministic block and chunk IDs where structure is unchanged, compute embeddings, extract metadata, and run validation. Then publish a manifest that switches retrieval from the old revision to the new one. Readers should see either complete state, not a mixture produced halfway through the job.

A dependency record makes selective work safe. If one section changed, its chunks and summaries need regeneration; an unchanged appendix can retain derived artifacts only if their inputs, parser version, embedding model, and access rules are identical. Store those transformation versions alongside the content hash. Otherwise an apparent incremental optimization can preserve artifacts created by incompatible code.

Permissions travel with content. Re-evaluate access controls on every change event, including events where the bytes did not change. Remove newly forbidden chunks from eligible search immediately, invalidate answer caches that used them, and test retrieval as the affected user. Content freshness without permission freshness is a data leak.

Treat deletion as a first-class event rather than absence from one poll. Mark the revision invalid, remove or exclude its chunks and embeddings, expire summaries and cached answers derived from it, and retain only the audit record allowed by the retention policy. Soft deletion, legal hold, source unlinking, and hard erasure are different states; one generic `deleted` flag usually cannot enforce all four.

Publish a revision atomically: source record, derived artifacts, permissions, and retrieval eligibility must change as one observable state.

Retrieve current or historical evidence deliberately

Current-state retrieval should exclude superseded, deleted, quarantined, and not-yet-effective revisions before semantic ranking. This is an eligibility rule, not a recency boost. A newer blog post does not outrank an older governing contract merely because its timestamp is later. Apply authority, jurisdiction, approval status, and valid-time filters before asking the ranker which eligible passage best answers the question.

Historical questions require a different filter. Parse phrases such as ‘as of 31 March’ or accept an explicit date from the workflow, then retrieve revisions whose effective intervals cover that date. If the source provides only publication dates, label the answer accordingly. Do not imply business-time precision the evidence does not contain.

Return revision metadata with every citation: source title, stable ID, revision or validator, effective date when available, and retrieval timestamp. If the connector is outside its freshness target, disclose that before answering or abstain for workflows where stale guidance is unsafe. The generation model should not be asked to infer index health from document prose.

Keep current and historical indexes logically separate even if they share storage. A default search should not surface obsolete wording because it happens to be semantically closer. Conversely, deleting history to simplify current retrieval makes comparisons, investigations, and ‘what changed?’ questions impossible.

Design for partial failure instead of hiding it

Freshness pipelines fail between stages. A source fetch succeeds but parsing fails; embeddings time out; a permission lookup is throttled; the new manifest is written but the answer cache is not cleared. Give every revision a state machine such as detected, fetched, transformed, validated, published, or failed, with retry counts and an error reason. Never advance the durable cursor past work that can still be lost.

Quarantine malformed revisions and keep the last fully valid version available only when the freshness contract permits it. Show its age and the failed update status. In a low-risk research workspace, a labelled older copy may be more useful than no result. For current operating procedures, serving the previous version after a known change may be unacceptable.

Use idempotent workers so the same event can safely run twice, and compare expected inputs before publishing. If revision C finishes before revision B, C should remain current; a slow retry must not roll the source backward. This is why revision ordering and atomic manifests matter more than job completion time.

Test the freshness contract with controlled mutations

A freshness test should mutate a fixture source and observe the whole path. Edit a paragraph, change only permissions, rename and move the file, delete it, restore it, deliver events out of order, expire a cursor, and make parsing fail after download. Ask both a current question and an ‘as of’ question after each mutation. Confirm the cited revision as well as the answer text.

Measure detection latency, publish latency, stale retrieval rate, deletion leakage, permission leakage, historical-answer correctness, and reconciliation parity. Track the age of the oldest unprocessed event and the share of sources beyond their freshness target. These measures reveal different failures; a good average update time can coexist with one permanently stuck source.

Periodically rebuild a small reference collection from scratch and compare its searchable manifest with the incrementally maintained one. The two should agree on current source IDs, revisions, permissions, and content hashes. A mismatch is evidence that an event was missed or a transformation was not invalidated.

For a small, slow-moving corpus, a verified full rebuild may be the simplest reliable design. Use incremental synchronization when volume, update frequency, deletion speed, or history requirements justify the added state. The right architecture is the least complex one that can demonstrate its freshness contract under failure—not the one with the most streaming components.

If a real knowledge workflow keeps surfacing a superseded source or cannot answer an ‘as of’ question cleanly, share the anonymised change path at sahil@granveo.com. A concrete source, revision, and failure sequence is enough to diagnose where freshness breaks.

Sources and further reading

  1. 1
    StreamingQA: A Benchmark for Adaptation to New Knowledge over Time in Question Answering Models

    Liska et al., ICML 2022 — Evaluates question-answering systems against a timestamped stream of news and changing knowledge.

  2. 2
    FreshLLMs: Refreshing Large Language Models with Search Engine Augmentation

    Vu et al., Findings of ACL 2024 — Introduces FreshQA, a dynamic benchmark for questions that require current knowledge or rejection of a false premise.

  3. 3
    Time Ontology in OWL

    W3C and OGC — Defines temporal concepts for representing instants, intervals, durations, and relationships in time.

  4. 4
    PROV-O: The PROV Ontology

    W3C — Defines provenance relationships including derivation, revision, generation, and invalidation.

  5. 5
    Retrieve changes

    Google Drive API documentation — Documents page-token-based change retrieval and the relationship between watch notifications and the change feed.

  6. 6
    Use delta query to track changes in Microsoft Graph data

    Microsoft Graph documentation — Documents incremental synchronization of created, updated, and deleted resources with opaque state tokens.

  7. 7
    RFC 9110: HTTP Semantics

    IETF — Defines ETag and Last-Modified validators and the conditional request semantics used to detect changed web representations.

Continue the conversation

Where does context get lost in your work?

I am speaking with researchers, founders, and operators about the handoffs, evidence, and decisions that are hardest to keep connected.