Back to Blog
EnglishTutorial

Incremental Embeddings: Deduplication, Checkpoints, Updates, and Deletes

Build a resumable incremental embedding index in Python and SQLite. Deduplicate text, checkpoint vectors, publish atomic updates, and handle deletions safely.

C
Crazyrouter Team
October 11, 2026 / 2 views
Share:
Incremental Embeddings: Deduplication, Checkpoints, Updates, and Deletes

Build incremental embeddings by hashing normalized text inside a versioned namespace, reusing existing vectors, and checkpointing each successful batch. Publish the new document-to-vector mapping only after all required vectors are ready, so an interrupted run preserves the previous searchable snapshot.

Crazyrouter is an AI API gateway where this example embedded 2 unique texts for 3 sources, reused both on an identical rerun, and embedded only 1 new text after an edit.

This article includes a complete Python and SQLite reference implementation. The live test used text-embedding-3-small, 512 dimensions, and https://cn.crazyrouter.com/v1/embeddings on October 11, 2026. Crash recovery, invalid results, and namespace changes were checked separately with deterministic local tests.

Define the source contract before writing the indexer#

The input here is a complete snapshot: a JSON object mapping stable source or chunk IDs to text. An ID missing from the next snapshot is deleted from the active retrieval map. An empty object intentionally removes all source links in this namespace.

Do not pass a partial crawl or only today's changed documents into this API. That would wrongly delete every omitted source. A delta-based ingestion service needs explicit upsert/delete events and an event checkpoint; this example deliberately uses the simpler complete-snapshot contract.

If you are starting with raw documents, perform deterministic chunking first. Use source IDs that identify the source and chunk location, and include the chunker revision in the namespace. The text-embedding-3-small API guide covers request shape and similarity; the embedding dimension-selection guide explains dimension choices.

EntityIdentityPurpose
Source/chunkStable source_idTracks which content is currently retrievable
ContentSHA-256 of normalized textDeduplicates exact normalized text
Embedding spaceModel + dimensions + preprocessing revisionPrevents accidental cross-version reuse
CheckpointCommitted vector batchAvoids repeating completed local work after restart

Incremental index commits vectors batch by batch before atomically publishing source links and performing garbage collection

The complete incremental embedding implementation#

Install requests, set CRAZYROUTER_API_KEY, and save the following as incremental_index.py. The code uses only the Python standard library plus requests.

python
"""Single-writer SQLite reference index; sync accepts a COMPLETE source snapshot."""
import hashlib
import json
import math
import os
import sqlite3
import unicodedata
from pathlib import Path
import requests

def normalize(text):
    # Keep whitespace meaningful; only normalize Unicode and line endings.
    return unicodedata.normalize("NFC", text.replace("\r\n", "\n").replace("\r", "\n"))

class Index:
    def __init__(self, path, embed, *, model="text-embedding-3-small", dimensions=512,
                 revision="nfc-lines-v1", batch_size=16):
        if batch_size < 1 or dimensions < 1:
            raise ValueError("batch_size and dimensions must be positive")
        self.db = sqlite3.connect(path)
        self.db.execute("PRAGMA foreign_keys = ON")
        self.db.executescript('''
          CREATE TABLE IF NOT EXISTS vectors (
            namespace TEXT NOT NULL, digest TEXT NOT NULL, vector TEXT NOT NULL,
            PRIMARY KEY(namespace, digest));
          CREATE TABLE IF NOT EXISTS documents (
            namespace TEXT NOT NULL, source_id TEXT NOT NULL, digest TEXT NOT NULL,
            PRIMARY KEY(namespace, source_id),
            FOREIGN KEY(namespace,digest) REFERENCES vectors(namespace,digest));
        ''')
        self.ns = json.dumps([model, dimensions, revision], separators=(",", ":"))
        self.embed, self.dimensions, self.batch_size = embed, dimensions, batch_size

    def sync(self, snapshot):
        """A dict of stable source/chunk IDs to text. Empty dict deletes this namespace."""
        if not isinstance(snapshot, dict):
            raise TypeError("snapshot must be a complete ID-to-text dictionary")
        links, unique = {}, {}
        for source_id, text in snapshot.items():
            if not isinstance(source_id, str) or not isinstance(text, str) or not text.strip():
                raise ValueError("IDs must be strings; texts must be nonempty strings")
            normalized = normalize(text)
            digest = hashlib.sha256(normalized.encode()).hexdigest()
            links[source_id] = digest
            unique[digest] = normalized
        present = {x[0] for x in self.db.execute(
            "SELECT digest FROM vectors WHERE namespace=?", (self.ns,))}
        missing = [(d, t) for d, t in unique.items() if d not in present]
        for start in range(0, len(missing), self.batch_size):
            batch = missing[start:start+self.batch_size]
            vectors = self.embed([text for _, text in batch])
            if len(vectors) != len(batch):
                raise ValueError("wrong vector count")
            for vector in vectors:
                if len(vector) != self.dimensions or not all(math.isfinite(v) for v in vector):
                    raise ValueError("invalid vector dimensions or nonfinite values")
            # Each completed batch is a durable checkpoint. Source links stay intact.
            with self.db:
                self.db.executemany("INSERT INTO vectors VALUES (?,?,?)", [
                    (self.ns, d, json.dumps(v)) for (d, _), v in zip(batch, vectors)])
        # Publish the new source map atomically only after all vectors are durable.
        with self.db:
            self.db.execute("DELETE FROM documents WHERE namespace=?", (self.ns,))
            self.db.executemany("INSERT INTO documents VALUES (?,?,?)", [
                (self.ns, source_id, digest) for source_id, digest in links.items()])
        return {"sources": len(links), "unique_texts": len(unique), "new_vectors": len(missing)}

    def gc(self):
        # Call only after successful sync; an interrupted run may need orphan vectors.
        with self.db:
            cur = self.db.execute('''DELETE FROM vectors WHERE namespace=?
              AND digest NOT IN (SELECT digest FROM documents WHERE namespace=?)''',
              (self.ns, self.ns))
        return cur.rowcount

    def close(self):
        self.db.close()

def api_embed(texts):
    base = os.getenv("CRAZYROUTER_BASE_URL", "https://cn.crazyrouter.com/v1")
    r = requests.post(base.rstrip("/") + "/embeddings",
        headers={"Authorization": "Bearer " + os.environ["CRAZYROUTER_API_KEY"]},
        json={"model": "text-embedding-3-small", "dimensions": 512,
              "encoding_format": "float", "input": texts}, timeout=(10, 60))
    r.raise_for_status()
    rows = sorted(r.json()["data"], key=lambda x: x["index"])
    if [x["index"] for x in rows] != list(range(len(texts))):
        raise ValueError("missing or duplicate response indices")
    return [x["embedding"] for x in rows]

if __name__ == "__main__":
    import argparse
    parser = argparse.ArgumentParser()
    parser.add_argument("snapshot", help="complete JSON object of source IDs to texts")
    parser.add_argument("--db", default="index.sqlite")
    args = parser.parse_args()
    index = Index(args.db, api_embed)
    try:
        print(index.sync(json.loads(Path(args.snapshot).read_text(encoding="utf-8"))))
    finally:
        index.close()

The API adapter validates response indices before restoring the requested order. The index also checks the vector count, dimensions, and finite numeric values before committing a batch. For the request parameters and documented upstream limits, see OpenAI's embedding guide.

Run a complete snapshot, then change it#

Save this as snapshot.json:

json
{
  "guide#1": "Reset an API key.",
  "copy#1": "Reset an API key.",
  "food#1": "Bake a loaf of bread."
}

Run:

bash
python incremental_index.py snapshot.json --db index.sqlite
python incremental_index.py snapshot.json --db index.sqlite

The first run needs two unique vectors for three sources; the identical second run needs none. Now replace the complete JSON snapshot with:

json
{
  "guide#1": "Rotate an API credential.",
  "copy#1": "Reset an API key."
}

Run the command again. guide#1 points to one new vector; copy#1 keeps the existing one; food#1 disappears from the active source map. The live test exercised these exact snapshots through Index.sync and recorded:

RunActive sourcesUnique active textsNew vectors
Initial snapshot322
Identical rerun320
Edit plus source deletion221
Garbage collection after successful sync22One unused stored vector removed

Live incremental embedding counts demonstrate exact deduplication, zero new vectors on rerun, and one new vector after an edit

The CLI does not automatically run garbage collection. Call index.gc() only after a successful sync and after deciding you no longer need unreferenced checkpoints. This avoids throwing away useful completed batches during recovery.

Why an interrupted run does not replace the old index#

Each completed vector batch is committed independently. Only after every missing vector exists does one transaction replace the source links. If embedding fails halfway through, readers that join documents to vectors still see the old published mapping. A restart hashes the same input and reuses committed vectors before requesting the missing ones.

The local test injected a failure on the second new batch. It confirmed that the old source map remained visible, reopened the database, and completed the run with only one new vector. It also verified exact deduplication, source deletion, unused-vector collection, namespace isolation, and rejection of bad vector count, length, and NaN values.

There is still an unavoidable boundary: if the provider finishes inference but the process crashes before saving the vector, a restart may pay for the same inference again. This is not an exactly-once billing guarantee. The bounded API retry implementation explains the same ambiguity for timed-out requests.

SQLite's transaction documentation describes the database atomicity used here. This reference assumes one writer and a local SQLite file. For multiple ingestion workers, coordinate ownership, transactions, and duplicate creation explicitly.

Deletions, version migrations, and scale#

Two sources sharing the same text share one vector. Deleting one source must not remove that vector while another source still references it. The garbage-collection query works from the current source map to avoid that mistake. Retrieval must join through documents; querying every row in vectors would expose abandoned checkpoint or deleted content.

Change the namespace when you change the model, dimensions, normalization, or chunking behavior. A fixed public model name can hide a provider-side revision, so include an explicit revision policy when reproducibility requires it. Build the new namespace, evaluate retrieval, then switch the application to it. Keep the old namespace until rollback is no longer needed. The OpenAI model catalog can help identify model options; it is not a migration test.

This implementation holds the complete snapshot and existing digest set in memory and stores vectors as JSON. It is intentionally small and inspectable. For millions of chunks, use streaming source discovery, token-aware batching, compact vector storage, an actual nearest-neighbor index, and an atomic generation pointer or equivalent publication mechanism. It does not implement a distributed crawler or ANN search engine.

Its 16-input batches are synchronous embedding calls. They do not use or claim an asynchronous provider Batch API or its pricing. See the AI batch-processing guide for the distinct job lifecycle.

FAQ#

What counts as a duplicate?#

Exact equality after Unicode NFC and line-ending normalization. Semantically similar but differently worded text is not deduplicated by this algorithm.

Can I send only changed documents to sync?#

No. It expects the full current snapshot, and omitted IDs are removed. Use a separate delta-event contract for partial updates.

Where is the restart checkpoint?#

In the committed vectors rows. The next run checks those hashes before calling the API, while the documents table remains the published source map.

Does a deleted document's vector disappear immediately?#

Its source link disappears after successful snapshot publication. The vector remains until garbage collection confirms it has no remaining source references.

Can I keep the same namespace after changing dimensions?#

No. Model, dimension count, and preprocessing revision define the embedding space and cache identity. Build and evaluate a new namespace.

Does this guarantee each text is charged only once?#

No. A lost response or a crash before checkpoint commit can cause a repeated API call. The implementation avoids repeating locally saved work, not all remote duplicate work.

Implementation Guides

Topics

API GuidesTutorial

Related Articles

AI Meme Generator & Coloring Book Creator with GPT-image-2 — Fun Projects That Actually Make MoneyTutorial

AI Meme Generator & Coloring Book Creator with GPT-image-2 — Fun Projects That Actually Make Money

Build an AI meme generator and coloring book page creator using GPT-image-2 via Crazyrouter API. Two fun, monetizable projects with full code.

May 1
Sora API: The Complete Guide to Building with OpenAI Video GenerationTutorial

Sora API: The Complete Guide to Building with OpenAI Video Generation

OpenAI's current Sora API is asynchronous and tier-based, not a fire-and-forget video button. The official guide recommends polling every 10 to 20 seconds, and Sora access is not available on the F...

Mar 26
Agentic RAG: Build Smarter AI Agents with Retrieval-Augmented Generation in 2026Tutorial

Agentic RAG: Build Smarter AI Agents with Retrieval-Augmented Generation in 2026

Learn how to build Agentic RAG systems that combine autonomous AI agents with retrieval-augmented generation for dynamic, multi-step reasoning over your own data.

Apr 15
AI Voice Agent Guide 2026: Build Speech-to-Speech AI with Real-Time APIsTutorial

AI Voice Agent Guide 2026: Build Speech-to-Speech AI with Real-Time APIs

"Complete guide to building AI voice agents with speech-to-speech APIs. Compare OpenAI Realtime, ElevenLabs, Deepgram, and PlayHT for building conversational voice AI."

Mar 2
AI Action Figure Generator with GPT-image-2 — Turn Anyone Into a Boxed ToyTutorial

AI Action Figure Generator with GPT-image-2 — Turn Anyone Into a Boxed Toy

Generate hyper-realistic boxed action figures using GPT-image-2 via Crazyrouter API. 10 profession templates included. Python, curl, and Node.js code.

May 1
Ghibli Style Photo Transformation with GPT-image-2 — Turn Any Photo Into Anime ArtTutorial

Ghibli Style Photo Transformation with GPT-image-2 — Turn Any Photo Into Anime Art

Transform photos into Studio Ghibli anime style using GPT-image-2 via Crazyrouter API. Multiple anime styles covered with full code examples.

May 1