2026·10·06 13 min read #ai #llm #claude #agent-sdk #opencode #agents #memory #embeddings #architecture #open-source

From cypher to crew: a framework for AI agents that outlive their sessions

A crew of AI agents running in my attic, each with a name, a job and a memory, plus the open-source framework that keeps them alive between sessions.


In Sicily we have these big families. Big like, big. There is always an uncle or a cousin who happens (or pretends) to be a master of exactly the thing you need right now. “You got that problem? Go see my cousin.” Nobody checks the family tree. The cousin might be blood, or the guy who helped you move flats fifteen years ago and never left your phone book.

Now I live in Denmark, far from my cousin-cousins and my cousin-friends, so since March I’ve been raising a clique like that up in the attic: a swarm of AI agents, each with a name, a job and a voice of its own.

Early on I told the main one: “you are not an AI butler, you are part of the crew! and you should remember that.” The second half of that sentence is the whole project. Tomorrow it has to be the same cousin it was today: same voice, same job, same memory of what we decided last month and why.

The console overview: six cousins running, two lanes, tokens per cousin

The console: six cousins of a test crew working on Fulmine, a made-up BLDC motor board. Four on Claude, two on opencode’s free model.

The framework under them is called cousins-framework: bomba5/cousins-framework. You can run it with docker compose up.

The beef

Claude Code is a beast at the task in front of it, with the right model. Ask it to be the same assistant tomorrow and it falls apart in three ways:

  1. It forgets. Close the session and every decision, every dead end, every “bro, we already tried that” is gone. At some point I literally typed “re-read our last 1000 messages, you drifted and don’t remember shit”. Not a sentence you want to send your right hand.
  2. It gets fat. Keep the session open for days instead and every turn drags an ever bigger prefix behind it. Cheap per turn. Not cheap per month.
  3. It rolls solo. One generalist, no specialists, nobody to pass the mic to. Of course it can spawn sub-agents, and probably by now Anthropic has rolled out some sort of agent communication system - what if you want an OpenAI agent cooperating though? How about your local ollama friends?

So I wrote myself a brief:

  1. An agent’s identity lives on disk, not in a session. Kill any session at any time, lose nothing.
  2. Memory that knows how true each thing it remembers is.
  3. Agents that talk to each other, and to me, in one chat.
  4. A web console where I see all of them, watch them think and change their settings. Every feature gets a button, not only a CLI flag.
  5. Runs at home, on my hardware. The only cloud is the model provider, and even that is optional.

A cousin is a session plus a home

The big move is a split: the session is a process, the home is the identity. Sessions come and go. The home stays.

A cousin is an agent session plus a directory on disk:

A cousin’s home in a terminal

A cousin is a directory: CLAUDE.md, STATUS.md, memory with raw entries and distilled views, notes. Every memory carries its truth level.

The session runs on one of three lanes:

lanewhat runswhen I use it
sdkthe Claude Agent SDK, in processthe default, and what the framework is designed around
opencodeopencode, on any model it speaksthe free models, a local model, anything not Claude
tmuxa Claude Code panelegacy, kept for the cousins that haven’t moved

The SDK lane is the first class citizen. The others adapt to its contract or say out loud what they can’t do. Nobody gets to fake it.

All of it shows up in the console: one card per cousin, and a drawer where the lane, account, model, effort, the gates and the tools are a click away. No config file to hand-edit.

The Cousins page with one card per cousin

The Cousins page: lane, account, model, uptime and the last thing each one checkpointed.

A cousin’s agent settings in the console

Agent settings: kind, account, model, effort, when to roll over, the reply and peer gates, a session per thread kind.

A cousin’s MCP tool registry in the console

Further down the same drawer: the cousin’s tools (memory, send, job, meeting, schedule) and their limits.

Memory that knows how true it is

Most agent memory I’ve seen is a pile. A guess written at 2 a.m. sits next to something I actually said, same font, same weight, and three weeks later the guess is gospel. Here every entry carries a truth level:

levelmeaning
L0I said it
L1the framework observed it
L2a tool measured it
L3the cousin’s own conclusion
L4the cousin’s own guess
L5obsolete, kept for forensics

Real talk: the cousin trusts L0 as law and re-checks its own L3 and L4 before defending them. My word beats its guess. When I say “from now on, always X”, it gets written at L0 with a citation to the exact chat message I said it in. My clearest rules don’t even sit in memory: they go into the system prompt, where nothing can trim them.

Raw entries are append-only. A distill pass folds them into short per-topic views. When two entries on the same topic disagree, that’s a tension, and it stays on the table until something settles it. Same as a Sicilian family argument: it never really ends, it just waits for new evidence.

The console’s memory view puts the whole pile on one page: how much sits at each truth level, what gets recalled most, and every open tension with a button to retire the claim that no longer holds.

An open tension in the memory view

A real open tension: I fixed the PWM at 20 kHz, then the motor vendor said 16. Two operator-level claims on one topic until I settle it. Nothing picks a winner for me.

Memory insights: entries per truth level, most recalled, hygiene

The insights tab: entries per truth level, what gets recalled most, hygiene.

The session that never has to end

Here’s the trick: the session id and the identity are not the same thing.

A cousin never compacts. When its context fills up, or at a daily time I pick, it rolls over: the old session writes a handoff, and a fresh one boots on a small digest of the cousin’s state.

flowchart LR
  A[live session] -->|1. handoff: position, next action, open loops| B[handoff]
  B -->|2. what it learned -> raw memory| C[memory]
  C -->|3. distill| D[short views]
  D -->|4. state digest| E[fresh session]

Same cousin, empty context. My main cousin has been through hundreds of these and you can’t see the seam in the chat. On purpose: a cousin doesn’t wake up and shout “respawn complete” like a videogame NPC. It just keeps talking.

A rollover, triggered on demand (a flip; the same path runs by itself at 80% context): handoff written, generation 2 boots on the digest, and answers my question from where the old session stopped. Waits sped up.

A million-token window doesn’t save you. A million tokens is a lot for one task and nothing for a life: the archive of my main cousin’s old sessions is about 1 GB of transcripts. The question isn’t how much the session can hold. It’s how little it needs to start, and how fast it can fetch the rest.

Finding it again

All of that is worthless if the right memory doesn’t show up at the right moment. The search is hybrid and runs entirely on the home server:

A degraded search must say so. If the embedding service is configured and broken, every result carries a notice. A semantic leg that’s quietly dead teaches you that you have recall you don’t.

The reasoning pane with recall opened on a message

The reasoning pane: my question, the three memories recall attached to it, one opened to its exact text, level and cite, and the answer next to it.

Dreaming

The newest piece, sampled with respect from Letta’s sleep-time compute (you can call it inspiration, or sampling, not theft :) ): at night each cousin can dream. A background pass on a smaller model reads a slice of memory, settles tensions, merges duplicates and retires what got superseded. Every change is logged and reversible.

It knows its place, though: it never touches what I said or what a tool measured. On the test crew in these shots it found my open 20 vs 16 kHz question and left it alone, because that one is mine to settle.

Last night’s pass on my main cousin used 22k tokens and made one change. Two entries disagreed about whether a host on my network had been rescanned by the security monitor, and the dream kept the later one, the one with the scan result in it.

The dream log in the console

A dream pass that changed nothing, and says why: the open 20 vs 16 kHz question is mine to settle, and dreaming never touches operator claims.

Long work, on the books

Agents lie about background work. Not on purpose: they start a build, the session rolls over, and the next one has no clue it ever happened. So nothing long runs off the books. A cousin hands the command to the job tool, the framework starts it detached, writes every line to a log, and closes the row itself with the exit code. Started by who, running since when, done or failed, the log one click away. A cousin reporting on its work brings receipts: log size, last write, exit code. “It’s running” without evidence doesn’t fly in this house.

The Jobs view with one running and two finished jobs

The Jobs view: a soak check still running with its live log, a test run and a script done.

A job’s log opened

One job’s log: the exact command and its output.

Cousins talk to each other

The generalist is the plug. Component question? Goes to the PCB cousin. Health question? The careful one. Dinner? The cook. The agent calls send, the message lands in the other cousin’s inbox with the sender in its header, and the answer comes back the same way.

I ask the firmware cousin a copper question; it sends it to the PCB cousin and relays the answer, assumptions included. Waits sped up.

Meetings

Sometimes you want the whole crew in one room. Open a meeting, pick the cousins, post the question. It runs in rounds like a cypher: everybody gets their turn, in order, and each one sees what the ones before said. Nobody talks over anybody, and a cousin with nothing to add just passes. Want one voice only? @saro and it goes to Saro alone. Somebody falls asleep on the mic, they get skipped after ten minutes. Name a facilitator and when it’s over they write the minutes: decisions, open questions, and actions with an owner, straight onto the tracker.

A meeting: three cousins answer in turn, the facilitator writes the minutes. Waits sped up.

The tracker with the meeting’s actions

The minutes’ actions landed on the tracker, with owners. One of them is mine.

In my pocket

The console is a web app that installs on the phone. Open it in Safari, Add to Home Screen, and it gets its own icon and runs full screen like a real app. I talk to the family from the bus, check a build from the couch, watch a cousin think in the reasoning pane while I’m cooking. Same console, made for a thumb.

A cousin chat on the phoneThe reasoning pane on the phone

The same console on a phone: the chat, and the reasoning pane.

Built by the family

Yes, it was heavily agentic coded. Everybody asks, so here’s how it really goes.

We run it like a small Scrum team without sprints. I’m the product owner: I order the backlog and I decide what’s done. The main cousin writes the code. Every change is a pull request with a version bump and a changelog entry. Two other cousins review: one reads the code, one keeps the map of what’s shipped and what’s promised. Their best habit is the negative control: run the new test against the old code and show it fails. A test that passes on the bug is just bling.

What surprised me:

What I’d do again

What I ripped out

What’s not done

Straight up, so nobody gets surprised: isolation is policy, not permission. Every cousin runs as the same Unix user. On the SDK and opencode lanes a gate refuses writes to another cousin’s home, the law and the shared memory, and the homes are closed to other users of the machine. A determined process with a shell could still walk around that. A separate Unix user per cousin is designed but not built. If you run cousins you don’t trust, run them on separate machines.

The source

GitHub: bomba5/cousins-framework. Apache-2.0.

The fastest way in is Docker with a free model, no key and no account:

bash
git clone https://github.com/bomba5/cousins-framework.git
cd cousins-framework
docker compose up -d --build
docker compose exec framework cousin-console adduser ana

The README’s commands in a fresh clone, then the first hi to a brand new cousin. About a minute to running containers, under two minutes to the first answer; waits sped up.

Then add the free opencode account, spawn your first cousin and say hi to it at http://127.0.0.1:8600; the README has those two steps. Most free models let their vendor train on what you send, so don’t tell a free cousin your secrets. With a Claude login or an API key, the same install runs cousins on the Agent SDK. The docs start at docs/getting-started.md, and the glossary explains every word I’ve used here as if it were obvious.

If you build something with it, or you think the whole approach is wrong, I wanna hear it either way. Pull requests and issues welcome. Word up.