In daily use
Agentic System
A Windows native orchestration platform where Claude and Codex share rules, hybrid memory, MCP tools, persistent run state, and adversarial review.
My role System architecture, implementation, integration, and daily operation.
The system behind the AI work
A model can answer a prompt. A dependable engineering environment must also recover project context, select tools, coordinate specialists, survive interruptions, control cost, and preserve a record that a person can inspect.
I built this Windows native system to support product development and the daily operation of Ktisis Arc. Claude and Codex run on top of the same provider neutral protocol. The active model can change while the rules, memory, tools, project maps, and review process remain shared.
A root map routes the agent to project rules, memory, and specialist instructions only when needed.
Rules, project maps, procedural memory, decisions, and searchable transcripts.
Markdown · SQLite · FTS5 · sqlite vecProvider adapters expose the same tools and allow specialist agents to work in bounded scopes.
Python · MCP · local scriptsProviders critique the same artifact, cluster findings, vote, and record accepted decisions.
Schemas · bounded rounds · decision logHow the system evolved
The architecture did not appear in one pass. Each stage solved a problem exposed by actual use.
1. Persistent gateway
The first version was a long lived agent gateway running in Docker inside WSL Ubuntu. It established messaging, tools, persistent state, and work that could continue across conversations.
2. Windows native core
I rebuilt the active path around the machine I operate every day. Windows services, Task Scheduler, local storage, GPU access, Telegram, and first party command line agents replaced the old gateway.
3. Compounding memory
Procedural, semantic, and episodic memory gained a separate raw transcript archive. Scheduled extraction, consolidation, and retrieval let the system recover decisions without loading every conversation.
4. One protocol, two providers
I separated the operating brain from the model vendor. Claude and Codex now source the same rule floor and can be selected for a session or a bounded specialist task.
5. Provider diverse review
Structured review lets the providers critique the same artifact, challenge findings, combine related issues, and stop after a bounded round adds nothing material.
The original Docker gateway is retired. A separate Linux desktop still runs in Docker as an operator environment. A portable core has a reviewed blueprint, but it is not presented as a running deployment.
Follow a real task through the system
A request can begin in a terminal or arrive through Telegram. Voice input is converted and transcribed locally with faster Whisper large v3 on the GPU. The transcript becomes the instruction, so speech does not require a separate manual handoff.
Boot loads the smallest rule floor that can operate safely. A root map routes the agent to the project index, relevant procedures, skills, and memory only when the task needs them. The agent can then call MCP tools or direct scripts for search, files, browser capture, transcription, monitoring, and maintenance.
Before parallel work expands, a duplicate guard checks whether equivalent work is already running. A cost circuit breaker applies the configured limit. Specialist agents receive bounded scopes and append structured findings to shared state. The primary agent reviews their output and relevant diffs before an accepted decision enters the durable log.
Heartbeats make liveness observable. Scheduled monitoring detects stale state and sends escalating Telegram alerts when work appears stuck. The run ledger supports recovery after an interrupted process. These controls contain and expose failure; they do not remove operator responsibility.
Why this architecture matters. The model is one replaceable execution component. Context, evidence, tools, and decisions remain part of the system instead of disappearing inside a chat window.
Context routing and measured boot reduction
The earlier workspace loaded a large general map before the task was known. That spent context on unrelated projects and made every session pay for information it might never use.
The replacement uses a small root map with selective topic routes. At the September 7, 2026 implementation snapshot:
| Measurement | Earlier path | Selective path | Result |
|---|---|---|---|
| Root map text | 8,316 tokens | 773 tokens | 90.7 percent less |
| Default boot floor | 13,821 tokens | 2,614 tokens | 81.1 percent less |
| Cold lookup comparison | 16,633 tokens | 5,470 tokens | Smaller start with routed lookup |
Sixteen bounded lookup cases tested whether the smaller map could still find required material. The first candidate resolved fifteen. The missed route was corrected, then all sixteen targeted cases resolved.
This is evidence for the tested routes and context size. It is not a universal recall benchmark, a latency benchmark, or proof that the provider bill falls by the same percentage.
Memory that stays useful
The Hippocampus memory layer keeps curated rules, facts, decisions, and procedures separate from the Echo transcript archive. A normal memory query combines FTS5 keyword retrieval with exact vector similarity, then fuses and diversifies the rankings.
An hourly watcher processes sessions after an idle window. A daily backstop catches work the normal schedule missed. Watermarks, tail resume behavior, and SQLite write ahead logging let extraction continue after interruption without starting the corpus again.
The September 14 database inspection found 25,863 memory rows and 200,796 transcript chunks. Those counts show operating scale, not retrieval accuracy. Source inspection remains available when the focused result is incomplete.
Tools, skills, and provider adapters
MCP exposes twelve named capabilities through a stable interface. The local tool directory contains 78 top level Python utilities, and the skill registry contains 60 directories in the same snapshot. The surface covers memory, context, transcription, browser work, monitoring, document delivery, and maintenance.
Provider adapters keep Claude specific and Codex specific behavior outside the shared protocol. Both use their first party command line tools and plan based authentication. This makes provider switching practical and keeps the rule system independent from one API contract.
The wider environment is still hybrid. Models, embeddings, search, and selected services can use the cloud. Local storage, orchestration, GPU transcription, state, and many tools run on the workstation.
Adversarial review as an engineering subsystem
The review layer is more than asking a second model for an opinion. Each reviewer produces findings in a common schema. The system validates the records, clusters overlapping concerns, records agreement, and allows another bounded round when material disagreement remains.
Two historical operating runs recorded costs of $0.62 for three rounds and $0.17 for two rounds. They demonstrate that the mechanism ran with a dry stop, minimal seats, and a round cap. They do not prove that two providers always beat one strong model at the same budget.
This system has reviewed its own memory changes and agent team designs. That is useful because the architecture can apply its assurance process to the machinery that defines the process.
Reliability and operating boundaries
Persistent state
SQLite and appendable run artifacts keep memory, task progress, findings, and decisions available after the model session ends.
Bounded execution
Duplicate checks, cost limits, scoped workers, and round caps constrain expansion before it becomes invisible or expensive.
Observable failure
Heartbeats, scheduled checks, and Telegram alerts expose stalled work and preserve a route for recovery.
Human authority
Agents contribute evidence and proposed changes. I remain responsible for consequential decisions, release approval, and failure resolution.
The system is in daily use for StreakUp, Theophonia, this portfolio, research, document production, and operational checks. It demonstrates engineering across orchestration, retrieval, provider abstraction, evaluation, observability, and recovery.
For Ktisis Arc, the same design discipline applies to client work: automate the repeatable parts, preserve the source of a decision, expose failure, and keep control with the business owner.