Skip to main content

How It Works

DriftGuard has two main pipelines: recording and retrieval.

Recording pipeline​

When you call guard.record(action, feedback, outcome):

action text
↓
normalize (spaCy lemmatize + stopword removal)
↓
embed (sentence-transformers)
↓
find similar existing node (cosine similarity > threshold)
↓
merge into existing node OR create new node
↓
add edges: action → feedback → outcome
↓
persist to JSON or SQLite

Semantic merge​

If "add more salt" arrives after "increase salt" is already in the graph, the merge engine computes cosine similarity between the two normalized embeddings. If similarity exceeds the action threshold (default 0.72), the new text merges into the existing node — incrementing its frequency — rather than creating a duplicate.

This keeps the graph compact and ensures repeated paraphrased mistakes reinforce one signal rather than fragmenting it.

Retrieval pipeline​

When you call guard.before_step(context):

context text
↓
embed
↓
top-k cosine similarity search over action nodes
↓
walk causal chains from each matched node (DFS)
↓
build warnings from (trigger, risk) pairs
↓
score confidence = f(node_freq, edge_freq, similarity, recency)
↓
deduplicate + sort by confidence descending
↓
return RetrievalResponse

The same pipeline runs a second time against the success graph, building reinforcements from (trigger, recommendation) pairs using the same confidence scoring. Both results — warnings and reinforcements — are attached to the same RetrievalResponse. See Success Memory.

Confidence scoring​

confidence = 0.65 × reinforcement_score
+ 0.35 × similarity_score
+ recency_weight × recency_score

Reinforcement score tiers:

Combined frequencyScore
≥ 50.95
≥ 30.85
≥ 20.75
10.60

Graph structure​

The in-memory graph is a directed networkx.DiGraph. Each node stores:

AttributeTypePurpose
typestraction, feedback, outcome
embeddingnp.ndarrayNormalized sentence embedding
frequencyintHow many times merged into
first_seendatetimeWhen first recorded
last_seendatetimeWhen last reinforced

Each edge stores frequency, weight, and created_at.

Pruning​

Deep pruning runs on a schedule and removes:

  1. Weak edges — frequency below prune_edge_min_frequency
  2. Stale nodes — not seen within prune_node_stale_days
  3. Isolated nodes — no incoming or outgoing edges after edge removal