All projects

In daily use

Agentic System

A Windows native orchestration platform where Claude and Codex share rules, hybrid memory, MCP tools, persistent run state, and adversarial review.

Published 5 minute technical case study

My role System architecture, implementation, integration, and daily operation.

Agentic System architecture showing shared context, two model providers, tools, review, and operational recovery.

The system behind the AI work

A model can answer a prompt. A dependable engineering environment must also recover project context, select tools, coordinate specialists, survive interruptions, control cost, and preserve a record that a person can inspect.

I built this Windows native system to support product development and the daily operation of Ktisis Arc. Claude and Codex run on top of the same provider neutral protocol. The active model can change while the rules, memory, tools, project maps, and review process remain shared.

25,863memory rows in the September 14 snapshot
200,796searchable transcript chunks
12named tools exposed through MCP
78top level Python utilities in the tool layer
System architecture One operating system for two AI providers The model can change while rules, memory, tools, review, and recovery remain shared.
Task Boot loads the smallest useful context

A root map routes the agent to project rules, memory, and specialist instructions only when needed.

CONTEXT Shared protocol

Rules, project maps, procedural memory, decisions, and searchable transcripts.

Markdown · SQLite · FTS5 · sqlite vec
EXECUTION Claude or Codex

Provider adapters expose the same tools and allow specialist agents to work in bounded scopes.

Python · MCP · local scripts
ASSURANCE Adversarial review

Providers critique the same artifact, cluster findings, vote, and record accepted decisions.

Schemas · bounded rounds · decision log
Persistent run ledger Duplicate work guard Cost circuit breaker Heartbeats and recovery

How the system evolved

The architecture did not appear in one pass. Each stage solved a problem exposed by actual use.

1. Persistent gateway

The first version was a long lived agent gateway running in Docker inside WSL Ubuntu. It established messaging, tools, persistent state, and work that could continue across conversations.

2. Windows native core

I rebuilt the active path around the machine I operate every day. Windows services, Task Scheduler, local storage, GPU access, Telegram, and first party command line agents replaced the old gateway.

3. Compounding memory

Procedural, semantic, and episodic memory gained a separate raw transcript archive. Scheduled extraction, consolidation, and retrieval let the system recover decisions without loading every conversation.

4. One protocol, two providers

I separated the operating brain from the model vendor. Claude and Codex now source the same rule floor and can be selected for a session or a bounded specialist task.

5. Provider diverse review

Structured review lets the providers critique the same artifact, challenge findings, combine related issues, and stop after a bounded round adds nothing material.

The original Docker gateway is retired. A separate Linux desktop still runs in Docker as an operator environment. A portable core has a reviewed blueprint, but it is not presented as a running deployment.

Follow a real task through the system

A request can begin in a terminal or arrive through Telegram. Voice input is converted and transcribed locally with faster Whisper large v3 on the GPU. The transcript becomes the instruction, so speech does not require a separate manual handoff.

Boot loads the smallest rule floor that can operate safely. A root map routes the agent to the project index, relevant procedures, skills, and memory only when the task needs them. The agent can then call MCP tools or direct scripts for search, files, browser capture, transcription, monitoring, and maintenance.

Before parallel work expands, a duplicate guard checks whether equivalent work is already running. A cost circuit breaker applies the configured limit. Specialist agents receive bounded scopes and append structured findings to shared state. The primary agent reviews their output and relevant diffs before an accepted decision enters the durable log.

Heartbeats make liveness observable. Scheduled monitoring detects stale state and sends escalating Telegram alerts when work appears stuck. The run ledger supports recovery after an interrupted process. These controls contain and expose failure; they do not remove operator responsibility.

Why this architecture matters. The model is one replaceable execution component. Context, evidence, tools, and decisions remain part of the system instead of disappearing inside a chat window.

Context routing and measured boot reduction

The earlier workspace loaded a large general map before the task was known. That spent context on unrelated projects and made every session pay for information it might never use.

The replacement uses a small root map with selective topic routes. At the September 7, 2026 implementation snapshot:

MeasurementEarlier pathSelective pathResult
Root map text8,316 tokens773 tokens90.7 percent less
Default boot floor13,821 tokens2,614 tokens81.1 percent less
Cold lookup comparison16,633 tokens5,470 tokensSmaller start with routed lookup

Sixteen bounded lookup cases tested whether the smaller map could still find required material. The first candidate resolved fifteen. The missed route was corrected, then all sixteen targeted cases resolved.

This is evidence for the tested routes and context size. It is not a universal recall benchmark, a latency benchmark, or proof that the provider bill falls by the same percentage.

Memory that stays useful

The Hippocampus memory layer keeps curated rules, facts, decisions, and procedures separate from the Echo transcript archive. A normal memory query combines FTS5 keyword retrieval with exact vector similarity, then fuses and diversifies the rankings.

An hourly watcher processes sessions after an idle window. A daily backstop catches work the normal schedule missed. Watermarks, tail resume behavior, and SQLite write ahead logging let extraction continue after interruption without starting the corpus again.

The September 14 database inspection found 25,863 memory rows and 200,796 transcript chunks. Those counts show operating scale, not retrieval accuracy. Source inspection remains available when the focused result is incomplete.

Tools, skills, and provider adapters

MCP exposes twelve named capabilities through a stable interface. The local tool directory contains 78 top level Python utilities, and the skill registry contains 60 directories in the same snapshot. The surface covers memory, context, transcription, browser work, monitoring, document delivery, and maintenance.

Provider adapters keep Claude specific and Codex specific behavior outside the shared protocol. Both use their first party command line tools and plan based authentication. This makes provider switching practical and keeps the rule system independent from one API contract.

The wider environment is still hybrid. Models, embeddings, search, and selected services can use the cloud. Local storage, orchestration, GPU transcription, state, and many tools run on the workstation.

Adversarial review as an engineering subsystem

The review layer is more than asking a second model for an opinion. Each reviewer produces findings in a common schema. The system validates the records, clusters overlapping concerns, records agreement, and allows another bounded round when material disagreement remains.

Two historical operating runs recorded costs of $0.62 for three rounds and $0.17 for two rounds. They demonstrate that the mechanism ran with a dry stop, minimal seats, and a round cap. They do not prove that two providers always beat one strong model at the same budget.

This system has reviewed its own memory changes and agent team designs. That is useful because the architecture can apply its assurance process to the machinery that defines the process.

Reliability and operating boundaries

Persistent state

SQLite and appendable run artifacts keep memory, task progress, findings, and decisions available after the model session ends.

Bounded execution

Duplicate checks, cost limits, scoped workers, and round caps constrain expansion before it becomes invisible or expensive.

Observable failure

Heartbeats, scheduled checks, and Telegram alerts expose stalled work and preserve a route for recovery.

Human authority

Agents contribute evidence and proposed changes. I remain responsible for consequential decisions, release approval, and failure resolution.

The system is in daily use for StreakUp, Theophonia, this portfolio, research, document production, and operational checks. It demonstrates engineering across orchestration, retrieval, provider abstraction, evaluation, observability, and recovery.

For Ktisis Arc, the same design discipline applies to client work: automate the repeatable parts, preserve the source of a decision, expose failure, and keep control with the business owner.