All projects

In daily use · architecture case study

Hippocampus

A hybrid AI memory engine combining FTS5, exact vector search, rank fusion, diversity selection, typed freshness, and source recovery inside SQLite.

Published 5 minute technical case study

My role Memory architecture, retrieval implementation, consolidation, and operating workflow.

Hippocampus retrieval architecture showing lexical and semantic search merged into a focused memory result.

The design question

An AI assistant with no memory starts every conversation from zero. Loading every old conversation solves that problem badly: context becomes expensive, relevant details compete with noise, and the model still has to discover which statements are current.

The real question is narrower and harder: out of everything the system has seen, which few memories belong in front of the agent for this task, and how can the original source remain available when the summary is not enough?

Hippocampus is my answer. It is the memory subsystem inside the Agentic System, built in Python on one SQLite database with two distinct stores and several retrieval stages.

25,863curated and extracted memory rows
200,796raw transcript chunks in Echo
60rank fusion denominator constant
5focused memories returned by default

Database counts inspected September 14, 2026. They describe corpus size, not recall accuracy.

Retrieval architecture Exact words and related meaning meet in one ranking The pipeline keeps useful context while removing repeated or stale results.
QuestionWhat did we decide, and why?
LEXICAL FTS5 with BM25

Finds exact names, commands, phrases, and file references.

Top 30 candidates
SEMANTIC Exact cosine search

Finds related decisions even when the question uses different words.

sqlite vec · top 30 candidates
RRF merges both rankings MMR reduces repetition Decay weighs type and freshness Result returns five memories
Curated BrainRules · facts · decisions
Echo archiveOriginal transcript chunks

Why naive memory fails

Embedding every message and returning the nearest results looks attractive because it is easy to explain. It fails in three predictable ways.

Exact identifiers disappear

A filename, command flag, error code, or product name may need a literal match. An embedding converts text into meaning, which can blur the exact string the task depends on. Hippocampus therefore searches both meaning and words.

Repetition crowds out variety

The same decision can be restated across many sessions. A plain nearest neighbor query may return five versions of one fact and exclude four other facts the task needs. Hippocampus explicitly penalizes redundant results before selection.

Old truth competes with new truth

Yesterday’s correction should usually outrank an old project note. A durable safety rule should not vanish simply because it is old. Hippocampus assigns different freshness behavior to different memory types instead of using one global aging policy.

Core principle. Rules, decisions, project facts, temporary tasks, and monitoring pings are different kinds of information. Their storage, ranking, consolidation, and retirement behavior should reflect that difference.

Four layers, loaded at different times

The system does not treat memory as one database query. It separates information according to when the agent should receive it.

LayerWhat it containsWhen it loads
Rule floorSmall nonnegotiable operating and communication rulesAt session start
Domain guidanceProject procedures, skills, and specialist instructionsWhen the task route requires them
Echo archiveOriginal conversation chunksOnly during explicit transcript search
Curated retrievalFacts, rules, decisions, tasks, and learned proceduresAt task start or when the agent requests memory

This split controls cost and noise. The small rule floor is paid often, so it stays small. Project guidance loads selectively. The large transcript archive stays silent until exact history matters. Curated retrieval contributes only a focused result.

Two stores in one SQLite file

The curated Brain and the Echo archive share one physical database for operational simplicity, but they use separate tables and retrieval paths.

Brain stores concise memory records with type, project, tags, importance, creation time, and consolidation state. It is designed for decisions and reusable context.

Echo stores chunks from original sessions. It is deliberately larger and more literal. A distilled memory may say that a deployment decision changed. Echo can recover the conversation that explains who changed it, what alternatives were considered, and what wording was used.

Keeping the stores separate avoids two opposite failures. Raw transcripts do not flood normal retrieval, and concise memory does not pretend to preserve every detail.

The retrieval pipeline

Every Brain query follows an explicit sequence.

1. Embed the question

The query is converted into a 1,536 dimension embedding with OpenAI text embedding 3 small. Query embeddings are cached inside the running process.

2. Run two retrievers

sqlite vec performs exact cosine nearest neighbor search. FTS5 with BM25 performs lexical search. Each returns up to 30 unconsolidated candidates, optionally filtered by memory tier.

The vector search is exact rather than approximate. That keeps the implementation simple and predictable at the current scale, while creating a clear future boundary if the corpus outgrows brute force lookup.

3. Fuse the rankings

Reciprocal rank fusion combines the two ordered lists with a denominator constant of 60. The fused scores are normalized before later stages.

Rank fusion is useful because BM25 and cosine distance do not share a meaningful score scale. The system combines their positions rather than pretending that the raw numbers are directly comparable.

4. Apply a small query hint

Queries containing words such as task, reminder, deadline, or next give task and todo memories a small relevance nudge. This is a bounded heuristic, not a general intent classifier.

5. Diversify with MMR

Maximal marginal relevance selects up to 10 memories. Its default relevance weight is 0.65, leaving 0.35 to penalize similarity with results already selected.

This stage is why a repeated fact does not consume the entire answer. The system can keep the strongest version while making room for other useful context.

6. Score relevance, importance, and freshness

The current weights are 0.75 for fused relevance, 0.10 for stored importance, and 0.15 for freshness. Freshness uses exponential decay with a type specific time constant.

Rules, architecture, and decisions use 3,650 day constants. Tools, configuration, and people use 730. General facts use 365. Insights use 180. Tasks use 90. Observability records use 7.

The implementation names these values as half lives, but the formula does not include the factor required for a mathematical half life. I describe them as decay constants on this page because that matches the actual calculation.

7. Optionally rerank

A Voyage cross encoder path can rerank the final diverse candidates. It has an 800 millisecond timeout, a circuit breaker after three consecutive failures, a five minute backoff, a daily call cap, and graceful fallback to the local ordering.

That path is currently dormant because its key is not provisioned. Normal search ends after MMR and typed freshness.

8. Return the focused result

The default interface emits five memories with type, project, date, importance, and score. The agent can then inspect the stored source or explicitly query Echo when those five are insufficient.

Memory types and authorship

Hippocampus separates procedural, semantic, and episodic tiers while also preserving more specific record types such as rule, decision, fact, task, tool, insight, and observability.

Authorship matters as much as type. Human curated rules and decisions should not be silently rewritten by an automated cleanup job. Machine extracted facts and procedures can merge when repeated evidence supports consolidation.

That boundary lets automated extraction operate without contaminating deliberate policy. The adaptive store can improve while the human authored operating floor remains controlled.

How memory enters the system

An hourly watcher scans completed sessions after a 30 minute idle gate. It extracts durable facts, decisions, tasks, and procedures instead of storing every sentence as a memory. A daily backstop processes sessions missed by the normal watcher.

Watermarks record how far each session has been read. Tail resume behavior continues from the last processed position. SQLite write ahead logging supports concurrent readers and safer recovery.

The raw transcript path is different. Meaningful messages are chunked, embedded, and added to Echo so that the original record remains searchable even when extraction chose not to create a curated memory.

Consolidation, retirement, and self maintenance

A nightly consolidation task compares compatible memory types. Similar records at or above a cosine threshold of 0.93 can merge into a stronger survivor. Rules, architecture, and decisions remain frozen from this automated merge path.

The survivor keeps the higher importance record, receives the merged statement and embedding, and can gain confidence when a learned procedure repeats. The retired row is marked consolidated and removed from vector matching. FTS5 is rebuilt as a consistency backstop.

Old task and todo memories can retire after their useful window unless marked canonical or pinned. The system therefore manages accumulation instead of assuming the database can grow forever without affecting retrieval quality.

Failure behavior and honest limits

Lexical failure

If FTS5 is unavailable, search can continue with vector candidates. Exact identifier recall becomes weaker.

Reranker failure

The optional remote reranker degrades to the MMR and freshness ordering after timeout or circuit opening.

Embedding failure

Query embedding is the hard dependency. An OpenAI embedding outage currently stops a new semantic query.

Retrieval uncertainty

A relevant memory can still be missed or outdated. Source metadata and Echo provide an inspection path, not perfect recall.

Hippocampus is a working architecture for a single user workstation, not a hosted memory product or a claim of unlimited scale. Its value is the explicit engineering: multiple stores, multiple retrievers, rank fusion, diversity, typed lifecycle policy, consolidation, provenance, and visible failure boundaries.