Incremental Embeddings: Deduplication, Checkpoints, Updates, and Deletes
Build a resumable incremental embedding index in Python and SQLite. Deduplicate text, checkpoint vectors, publish atomic updates, and handle deletions safely.

Build incremental embeddings by hashing normalized text inside a versioned namespace, reusing existing vectors, and checkpointing each successful batch. Publish the new document-to-vector mapping only after all required vectors are ready, so an interrupted run preserves the previous searchable snapshot.
Crazyrouter is an AI API gateway where this example embedded 2 unique texts for 3 sources, reused both on an identical rerun, and embedded only 1 new text after an edit.
This article includes a complete Python and SQLite reference implementation. The live test used text-embedding-3-small, 512 dimensions, and https://cn.crazyrouter.com/v1/embeddings on October 11, 2026. Crash recovery, invalid results, and namespace changes were checked separately with deterministic local tests.
Define the source contract before writing the indexer#
The input here is a complete snapshot: a JSON object mapping stable source or chunk IDs to text. An ID missing from the next snapshot is deleted from the active retrieval map. An empty object intentionally removes all source links in this namespace.
Do not pass a partial crawl or only today's changed documents into this API. That would wrongly delete every omitted source. A delta-based ingestion service needs explicit upsert/delete events and an event checkpoint; this example deliberately uses the simpler complete-snapshot contract.
If you are starting with raw documents, perform deterministic chunking first. Use source IDs that identify the source and chunk location, and include the chunker revision in the namespace. The text-embedding-3-small API guide covers request shape and similarity; the embedding dimension-selection guide explains dimension choices.
| Entity | Identity | Purpose |
|---|---|---|
| Source/chunk | Stable source_id | Tracks which content is currently retrievable |
| Content | SHA-256 of normalized text | Deduplicates exact normalized text |
| Embedding space | Model + dimensions + preprocessing revision | Prevents accidental cross-version reuse |
| Checkpoint | Committed vector batch | Avoids repeating completed local work after restart |

The complete incremental embedding implementation#
Install requests, set CRAZYROUTER_API_KEY, and save the following as incremental_index.py. The code uses only the Python standard library plus requests.
"""Single-writer SQLite reference index; sync accepts a COMPLETE source snapshot."""
import hashlib
import json
import math
import os
import sqlite3
import unicodedata
from pathlib import Path
import requests
def normalize(text):
# Keep whitespace meaningful; only normalize Unicode and line endings.
return unicodedata.normalize("NFC", text.replace("\r\n", "\n").replace("\r", "\n"))
class Index:
def __init__(self, path, embed, *, model="text-embedding-3-small", dimensions=512,
revision="nfc-lines-v1", batch_size=16):
if batch_size < 1 or dimensions < 1:
raise ValueError("batch_size and dimensions must be positive")
self.db = sqlite3.connect(path)
self.db.execute("PRAGMA foreign_keys = ON")
self.db.executescript('''
CREATE TABLE IF NOT EXISTS vectors (
namespace TEXT NOT NULL, digest TEXT NOT NULL, vector TEXT NOT NULL,
PRIMARY KEY(namespace, digest));
CREATE TABLE IF NOT EXISTS documents (
namespace TEXT NOT NULL, source_id TEXT NOT NULL, digest TEXT NOT NULL,
PRIMARY KEY(namespace, source_id),
FOREIGN KEY(namespace,digest) REFERENCES vectors(namespace,digest));
''')
self.ns = json.dumps([model, dimensions, revision], separators=(",", ":"))
self.embed, self.dimensions, self.batch_size = embed, dimensions, batch_size
def sync(self, snapshot):
"""A dict of stable source/chunk IDs to text. Empty dict deletes this namespace."""
if not isinstance(snapshot, dict):
raise TypeError("snapshot must be a complete ID-to-text dictionary")
links, unique = {}, {}
for source_id, text in snapshot.items():
if not isinstance(source_id, str) or not isinstance(text, str) or not text.strip():
raise ValueError("IDs must be strings; texts must be nonempty strings")
normalized = normalize(text)
digest = hashlib.sha256(normalized.encode()).hexdigest()
links[source_id] = digest
unique[digest] = normalized
present = {x[0] for x in self.db.execute(
"SELECT digest FROM vectors WHERE namespace=?", (self.ns,))}
missing = [(d, t) for d, t in unique.items() if d not in present]
for start in range(0, len(missing), self.batch_size):
batch = missing[start:start+self.batch_size]
vectors = self.embed([text for _, text in batch])
if len(vectors) != len(batch):
raise ValueError("wrong vector count")
for vector in vectors:
if len(vector) != self.dimensions or not all(math.isfinite(v) for v in vector):
raise ValueError("invalid vector dimensions or nonfinite values")
# Each completed batch is a durable checkpoint. Source links stay intact.
with self.db:
self.db.executemany("INSERT INTO vectors VALUES (?,?,?)", [
(self.ns, d, json.dumps(v)) for (d, _), v in zip(batch, vectors)])
# Publish the new source map atomically only after all vectors are durable.
with self.db:
self.db.execute("DELETE FROM documents WHERE namespace=?", (self.ns,))
self.db.executemany("INSERT INTO documents VALUES (?,?,?)", [
(self.ns, source_id, digest) for source_id, digest in links.items()])
return {"sources": len(links), "unique_texts": len(unique), "new_vectors": len(missing)}
def gc(self):
# Call only after successful sync; an interrupted run may need orphan vectors.
with self.db:
cur = self.db.execute('''DELETE FROM vectors WHERE namespace=?
AND digest NOT IN (SELECT digest FROM documents WHERE namespace=?)''',
(self.ns, self.ns))
return cur.rowcount
def close(self):
self.db.close()
def api_embed(texts):
base = os.getenv("CRAZYROUTER_BASE_URL", "https://cn.crazyrouter.com/v1")
r = requests.post(base.rstrip("/") + "/embeddings",
headers={"Authorization": "Bearer " + os.environ["CRAZYROUTER_API_KEY"]},
json={"model": "text-embedding-3-small", "dimensions": 512,
"encoding_format": "float", "input": texts}, timeout=(10, 60))
r.raise_for_status()
rows = sorted(r.json()["data"], key=lambda x: x["index"])
if [x["index"] for x in rows] != list(range(len(texts))):
raise ValueError("missing or duplicate response indices")
return [x["embedding"] for x in rows]
if __name__ == "__main__":
import argparse
parser = argparse.ArgumentParser()
parser.add_argument("snapshot", help="complete JSON object of source IDs to texts")
parser.add_argument("--db", default="index.sqlite")
args = parser.parse_args()
index = Index(args.db, api_embed)
try:
print(index.sync(json.loads(Path(args.snapshot).read_text(encoding="utf-8"))))
finally:
index.close()
The API adapter validates response indices before restoring the requested order. The index also checks the vector count, dimensions, and finite numeric values before committing a batch. For the request parameters and documented upstream limits, see OpenAI's embedding guide.
Run a complete snapshot, then change it#
Save this as snapshot.json:
{
"guide#1": "Reset an API key.",
"copy#1": "Reset an API key.",
"food#1": "Bake a loaf of bread."
}
Run:
python incremental_index.py snapshot.json --db index.sqlite
python incremental_index.py snapshot.json --db index.sqlite
The first run needs two unique vectors for three sources; the identical second run needs none. Now replace the complete JSON snapshot with:
{
"guide#1": "Rotate an API credential.",
"copy#1": "Reset an API key."
}
Run the command again. guide#1 points to one new vector; copy#1 keeps the existing one; food#1 disappears from the active source map. The live test exercised these exact snapshots through Index.sync and recorded:
| Run | Active sources | Unique active texts | New vectors |
|---|---|---|---|
| Initial snapshot | 3 | 2 | 2 |
| Identical rerun | 3 | 2 | 0 |
| Edit plus source deletion | 2 | 2 | 1 |
| Garbage collection after successful sync | 2 | 2 | One unused stored vector removed |

The CLI does not automatically run garbage collection. Call index.gc() only after a successful sync and after deciding you no longer need unreferenced checkpoints. This avoids throwing away useful completed batches during recovery.
Why an interrupted run does not replace the old index#
Each completed vector batch is committed independently. Only after every missing vector exists does one transaction replace the source links. If embedding fails halfway through, readers that join documents to vectors still see the old published mapping. A restart hashes the same input and reuses committed vectors before requesting the missing ones.
The local test injected a failure on the second new batch. It confirmed that the old source map remained visible, reopened the database, and completed the run with only one new vector. It also verified exact deduplication, source deletion, unused-vector collection, namespace isolation, and rejection of bad vector count, length, and NaN values.
There is still an unavoidable boundary: if the provider finishes inference but the process crashes before saving the vector, a restart may pay for the same inference again. This is not an exactly-once billing guarantee. The bounded API retry implementation explains the same ambiguity for timed-out requests.
SQLite's transaction documentation describes the database atomicity used here. This reference assumes one writer and a local SQLite file. For multiple ingestion workers, coordinate ownership, transactions, and duplicate creation explicitly.
Deletions, version migrations, and scale#
Two sources sharing the same text share one vector. Deleting one source must not remove that vector while another source still references it. The garbage-collection query works from the current source map to avoid that mistake. Retrieval must join through documents; querying every row in vectors would expose abandoned checkpoint or deleted content.
Change the namespace when you change the model, dimensions, normalization, or chunking behavior. A fixed public model name can hide a provider-side revision, so include an explicit revision policy when reproducibility requires it. Build the new namespace, evaluate retrieval, then switch the application to it. Keep the old namespace until rollback is no longer needed. The OpenAI model catalog can help identify model options; it is not a migration test.
This implementation holds the complete snapshot and existing digest set in memory and stores vectors as JSON. It is intentionally small and inspectable. For millions of chunks, use streaming source discovery, token-aware batching, compact vector storage, an actual nearest-neighbor index, and an atomic generation pointer or equivalent publication mechanism. It does not implement a distributed crawler or ANN search engine.
Its 16-input batches are synchronous embedding calls. They do not use or claim an asynchronous provider Batch API or its pricing. See the AI batch-processing guide for the distinct job lifecycle.
FAQ#
What counts as a duplicate?#
Exact equality after Unicode NFC and line-ending normalization. Semantically similar but differently worded text is not deduplicated by this algorithm.
Can I send only changed documents to sync?#
No. It expects the full current snapshot, and omitted IDs are removed. Use a separate delta-event contract for partial updates.
Where is the restart checkpoint?#
In the committed vectors rows. The next run checks those hashes before calling the API, while the documents table remains the published source map.
Does a deleted document's vector disappear immediately?#
Its source link disappears after successful snapshot publication. The vector remains until garbage collection confirms it has no remaining source references.
Can I keep the same namespace after changing dimensions?#
No. Model, dimension count, and preprocessing revision define the embedding space and cache identity. Build and evaluate a new namespace.
Does this guarantee each text is charged only once?#
No. A lost response or a crash before checkpoint commit can cause a repeated API call. The implementation avoids repeating locally saved work, not all remote duplicate work.





