In daily use · architecture case study
Hippocampus
A hybrid AI memory engine combining FTS5, exact vector search, rank fusion, diversity selection, typed freshness, and source recovery inside SQLite.
My role Memory architecture, retrieval implementation, consolidation, and operating workflow.
The design question
An AI assistant with no memory starts every conversation from zero. Loading every old conversation solves that problem badly: context becomes expensive, relevant details compete with noise, and the model still has to discover which statements are current.
The real question is narrower and harder: out of everything the system has seen, which few memories belong in front of the agent for this task, and how can the original source remain available when the summary is not enough?
Hippocampus is my answer. It is the memory subsystem inside the Agentic System, built in Python on one SQLite database with two distinct stores and several retrieval stages.
Database counts inspected September 14, 2026. They describe corpus size, not recall accuracy.
Finds exact names, commands, phrases, and file references.
Top 30 candidatesFinds related decisions even when the question uses different words.
sqlite vec · top 30 candidatesWhy naive memory fails
Embedding every message and returning the nearest results looks attractive because it is easy to explain. It fails in three predictable ways.
Exact identifiers disappear
A filename, command flag, error code, or product name may need a literal match. An embedding converts text into meaning, which can blur the exact string the task depends on. Hippocampus therefore searches both meaning and words.
Repetition crowds out variety
The same decision can be restated across many sessions. A plain nearest neighbor query may return five versions of one fact and exclude four other facts the task needs. Hippocampus explicitly penalizes redundant results before selection.
Old truth competes with new truth
Yesterday’s correction should usually outrank an old project note. A durable safety rule should not vanish simply because it is old. Hippocampus assigns different freshness behavior to different memory types instead of using one global aging policy.
Core principle. Rules, decisions, project facts, temporary tasks, and monitoring pings are different kinds of information. Their storage, ranking, consolidation, and retirement behavior should reflect that difference.
Four layers, loaded at different times
The system does not treat memory as one database query. It separates information according to when the agent should receive it.
| Layer | What it contains | When it loads |
|---|---|---|
| Rule floor | Small nonnegotiable operating and communication rules | At session start |
| Domain guidance | Project procedures, skills, and specialist instructions | When the task route requires them |
| Echo archive | Original conversation chunks | Only during explicit transcript search |
| Curated retrieval | Facts, rules, decisions, tasks, and learned procedures | At task start or when the agent requests memory |
This split controls cost and noise. The small rule floor is paid often, so it stays small. Project guidance loads selectively. The large transcript archive stays silent until exact history matters. Curated retrieval contributes only a focused result.
Two stores in one SQLite file
The curated Brain and the Echo archive share one physical database for operational simplicity, but they use separate tables and retrieval paths.
Brain stores concise memory records with type, project, tags, importance, creation time, and consolidation state. It is designed for decisions and reusable context.
Echo stores chunks from original sessions. It is deliberately larger and more literal. A distilled memory may say that a deployment decision changed. Echo can recover the conversation that explains who changed it, what alternatives were considered, and what wording was used.
Keeping the stores separate avoids two opposite failures. Raw transcripts do not flood normal retrieval, and concise memory does not pretend to preserve every detail.
The retrieval pipeline
Every Brain query follows an explicit sequence.
1. Embed the question
The query is converted into a 1,536 dimension embedding with OpenAI text embedding 3 small. Query embeddings are cached inside the running process.
2. Run two retrievers
sqlite vec performs exact cosine nearest neighbor search. FTS5 with BM25 performs lexical search. Each returns up to 30 unconsolidated candidates, optionally filtered by memory tier.
The vector search is exact rather than approximate. That keeps the implementation simple and predictable at the current scale, while creating a clear future boundary if the corpus outgrows brute force lookup.
3. Fuse the rankings
Reciprocal rank fusion combines the two ordered lists with a denominator constant of 60. The fused scores are normalized before later stages.
Rank fusion is useful because BM25 and cosine distance do not share a meaningful score scale. The system combines their positions rather than pretending that the raw numbers are directly comparable.
4. Apply a small query hint
Queries containing words such as task, reminder, deadline, or next give task and todo memories a small relevance nudge. This is a bounded heuristic, not a general intent classifier.
5. Diversify with MMR
Maximal marginal relevance selects up to 10 memories. Its default relevance weight is 0.65, leaving 0.35 to penalize similarity with results already selected.
This stage is why a repeated fact does not consume the entire answer. The system can keep the strongest version while making room for other useful context.
6. Score relevance, importance, and freshness
The current weights are 0.75 for fused relevance, 0.10 for stored importance, and 0.15 for freshness. Freshness uses exponential decay with a type specific time constant.
Rules, architecture, and decisions use 3,650 day constants. Tools, configuration, and people use 730. General facts use 365. Insights use 180. Tasks use 90. Observability records use 7.
The implementation names these values as half lives, but the formula does not include the factor required for a mathematical half life. I describe them as decay constants on this page because that matches the actual calculation.
7. Optionally rerank
A Voyage cross encoder path can rerank the final diverse candidates. It has an 800 millisecond timeout, a circuit breaker after three consecutive failures, a five minute backoff, a daily call cap, and graceful fallback to the local ordering.
That path is currently dormant because its key is not provisioned. Normal search ends after MMR and typed freshness.
8. Return the focused result
The default interface emits five memories with type, project, date, importance, and score. The agent can then inspect the stored source or explicitly query Echo when those five are insufficient.
Memory types and authorship
Hippocampus separates procedural, semantic, and episodic tiers while also preserving more specific record types such as rule, decision, fact, task, tool, insight, and observability.
Authorship matters as much as type. Human curated rules and decisions should not be silently rewritten by an automated cleanup job. Machine extracted facts and procedures can merge when repeated evidence supports consolidation.
That boundary lets automated extraction operate without contaminating deliberate policy. The adaptive store can improve while the human authored operating floor remains controlled.
How memory enters the system
An hourly watcher scans completed sessions after a 30 minute idle gate. It extracts durable facts, decisions, tasks, and procedures instead of storing every sentence as a memory. A daily backstop processes sessions missed by the normal watcher.
Watermarks record how far each session has been read. Tail resume behavior continues from the last processed position. SQLite write ahead logging supports concurrent readers and safer recovery.
The raw transcript path is different. Meaningful messages are chunked, embedded, and added to Echo so that the original record remains searchable even when extraction chose not to create a curated memory.
Consolidation, retirement, and self maintenance
A nightly consolidation task compares compatible memory types. Similar records at or above a cosine threshold of 0.93 can merge into a stronger survivor. Rules, architecture, and decisions remain frozen from this automated merge path.
The survivor keeps the higher importance record, receives the merged statement and embedding, and can gain confidence when a learned procedure repeats. The retired row is marked consolidated and removed from vector matching. FTS5 is rebuilt as a consistency backstop.
Old task and todo memories can retire after their useful window unless marked canonical or pinned. The system therefore manages accumulation instead of assuming the database can grow forever without affecting retrieval quality.
Failure behavior and honest limits
Lexical failure
If FTS5 is unavailable, search can continue with vector candidates. Exact identifier recall becomes weaker.
Reranker failure
The optional remote reranker degrades to the MMR and freshness ordering after timeout or circuit opening.
Embedding failure
Query embedding is the hard dependency. An OpenAI embedding outage currently stops a new semantic query.
Retrieval uncertainty
A relevant memory can still be missed or outdated. Source metadata and Echo provide an inspection path, not perfect recall.
Hippocampus is a working architecture for a single user workstation, not a hosted memory product or a claim of unlimited scale. Its value is the explicit engineering: multiple stores, multiple retrievers, rank fusion, diversity, typed lifecycle policy, consolidation, provenance, and visible failure boundaries.