Introducing Ouroboros: a runtime for AI coding agents
Why I built Ouroboros with Elixir, Rust and WebAssembly, and how durable sessions, bounded self-repair and auditable tools fit together.
TLDR: Ouroboros is my open-source runtime for long-running AI coding work. It can recover sessions, run subagents across your machines, and build new capabilities under human control. The runtime is Elixir, the terminal client is Rust, and extensions run as WebAssembly components. Here is why.
In Demystifying Agents, we built a small agent in TypeScript. A model asks for a tool, our program executes it, and we send the result back. Repeat until we have an answer, or until we decide that we have spent enough money on this conversation.
That loop is still there in Ouroboros.
But give it a repository, a shell and a task that takes a while, and the interesting questions move outside the model. The terminal closes. A process crashes. A command needs approval. Two agents want to edit the same files. The model says it fixed something, and you would quite like to see the test that proves it.
I wanted a coding agent where I could inspect and control those things without treating the conversation as the only record of what happened.
Ouroboros is now available under the MIT license, with a terminal client and a web interface. It is a working research project, still evolving. I am announcing something you can run and inspect, not a claim that unattended software development is solved. That would be a rather ambitious README :)
Start with a coding task
The launch page has recordings from the development build. One shows a small JavaScript project with a broken task summary: completed tasks are always counted as zero.
The agent inspects the code, runs the tests after approval, edits the implementation and adds an all-completed regression case. The final run passes, and an independent rerun passes too. You can follow the commands and the edit in the web interface.
It is deliberately a small example. We can understand the bug and check the result without trusting a paragraph written by the agent.
Another recording kills the agent process while it is idle. The runtime restores the session, and the agent answers a follow-up using the conversation it retained. That demonstrates a different property: the work has an identity beyond the process currently doing it.
The capture notes describe exactly what was tested. The recovery clip shows one process crash, not a machine losing power halfway through a shell command. The fleet clip uses two nodes on one Mac, not a network-partition experiment. The recordings also show a development build, which can be ahead of the announced stable release.
Those distinctions matter when your main feature is being able to trust what the runtime tells you.
Why Elixir for the runtime?
A coding agent spends a lot of time waiting. It waits for the model, a test process, a human approval or another agent. Meanwhile, the interface should remain usable, cancellation should work, and an unrelated session should keep going.
Elixir runs on the BEAM, the Erlang virtual machine. Its lightweight processes communicate through messages. OTP, the libraries and conventions around that runtime, gives us supervisors that monitor processes and decide what should restart when something dies.
If you are coming from NodeJS, do not think of every BEAM process as a separate operating-system process. It is a lightweight unit of execution managed by the VM, with its own state and mailbox. We can use those processes to give different parts of the agent separate owners.
In Ouroboros, the native session owns the live conversation and turn scheduling. A coordinator owns the public session state and workspace admission. The execution of a turn runs in a task, so the session process can still answer while the model is working.
That separation is useful during a crash. A coordinator can disappear while the native session remains alive, with output waiting to be collected. Recovery can acquire fresh admission and reconnect to it instead of blindly launching another model call.
A supervisor cannot remember your conversation
You will often see Erlang described with the phrase “let it crash”. It does not mean “throw away the user’s work and hope they ask again”.
Restarting a process gives you another process. You still have to decide where its state comes from.
Ouroboros keeps durable session records and checkpoints. The coordinator checkpoints output before acknowledging it and broadcasting it to clients. The file adapter writes a temporary file, syncs it, renames it over the checkpoint, then syncs the parent directory.
Even that has an awkward case: the rename can succeed and the directory sync can fail. The new file is visible, but the commit has not been proven durable. The adapter returns commit_outcome_unknown, rather than pretending the write definitely failed and continuing with stale state.
This is the kind of detail I wanted the runtime to own. A better prompt cannot fix a dishonest storage contract.
Restart order is part of the design
The application’s top-level supervisor uses OTP’s rest_for_one strategy. When a child dies, the supervisor restarts it and the children that were started after it. Independent subtrees use one_for_one where a failure should stay local.
The ordering is intentional. The effect ledger and permission authorities sit above the sessions that depend on them. If an authority restarts, those consumers must not carry on using assumptions from its previous incarnation. A web-interface failure, on the other hand, should not take the durable stores down with it.
You can read that ordering in application.ex. It is an executable description of which failures are allowed to affect which work.
This is why I chose Elixir. I wanted process ownership, failure handling and supervision to be ordinary parts of the program, rather than a collection of recovery callbacks added after the agent loop.
It does not make the model smarter. It makes the program around the model easier to reason about when something goes wrong.
Why keep the agent loop in-process?
The current runtime has one agent provider, Ouroboros.Provider.Native. It calls model APIs through ReqLLM. Here, “one provider” means one execution implementation, not one model vendor: OpenAI, Anthropic and xAI are exposed directly, with other ReqLLM model lanes configurable too.
The model adapter translates messages, tool schemas and streamed responses. It does not decide whether a shell command is permitted. That decision belongs to Ouroboros.
Owning the loop gives us a place to stop a tool before execution, wait for approval, accept steering between tool calls, and record the result. We do not have to infer what another coding CLI is doing by watching its terminal output.
The project previously carried several wrapped vendor CLIs. That compatibility layer was removed. It had its own lifecycle and approval differences, while the features I cared about needed control of the actual execution path.
There is a trade-off here. Supporting a model API does not mean inheriting every feature of that vendor’s coding application. I prefer that limitation to presenting several very different execution engines as though they had identical guarantees.
Rust where the user meets the runtime
I still like Rust. If you have read my web backend articles, this will probably not surprise you.
The ouro terminal client uses Ratatui for rendering, Crossterm for terminal interaction, and Tokio for asynchronous I/O. It handles the terminal, commands and runtime connection; the durable coding session belongs to the Elixir side.
A terminal client has a different set of problems from a session coordinator. It needs to render a changing transcript, handle keyboard input, restore the terminal correctly and deal with Unicode widths without making the cursor wander off. Rust is a good fit for that explicit state handling and native distribution.
The release packages the BEAM runtime inside the client binary. Users do not need an Elixir or Rust development environment just to run it. That convenience costs build and packaging work: the artifacts are platform-specific, and the project has to test the combinations it distributes.
I did not choose two languages to make the installation instructions more interesting. The split lets each side own a different job without moving the session’s lifetime into the UI.
Why Phoenix LiveView for the browser?
The browser interface uses Phoenix LiveView, served by the same Elixir runtime, with Bandit as the HTTP server.
A graphical interface is useful for reading a long transcript or reviewing an edit. I did not want it to become a second implementation of the runtime’s permissions. The web layer goes through the gateway’s shared method and authorization checks; it does not call privileged internals directly because a button happens to be visible.
LiveView also gives the project headless interface tests. The earlier native desktop surface had a more awkward testing story, and maintaining that rendering layer was not the problem I wanted to spend time on.
There is no Node build step in the production web asset path. The release copies Phoenix’s prebuilt browser assets and ships the application’s JavaScript and handwritten CSS. Node and Playwright are development tools for browser testing.
Some of the dependency choices are quite boring on purpose. Bandit is pure Elixir. Markdown uses Earmark rather than adding another Rust native dependency to the BEAM release for these payload sizes. Agent output is untrusted, so the renderer escapes raw HTML.
Using Rust in the terminal does not mean every other part of the program needs a Rust dependency too.
Security has to happen before the tool runs
Our tiny TypeScript agent could call any function we put in its tool array. In a coding agent, one of those functions can launch a shell. We need more than a system prompt saying “please be careful”.
Ouroboros separates permission decisions from execution containment.
Permission rules decide whether an action is allowed, denied or needs a human answer. Cluster grants cover what an agent may ask the cluster to do. A live tool call must have its attempt recorded in the effect ledger before execution. If that write fails, the tool does not run. Recording a refusal or settling a finished tool call is best-effort on this operational path, so it is not a promise that every outcome survives a storage failure.
Containment limits what the executing code can reach. The model’s shell uses sandbox-exec on macOS or bubblewrap on Linux. A session asking for workspace-write containment must not silently become an unrestricted shell because the machine lacks a working backend. The runtime refuses it instead. Unrestricted execution is an explicit operator choice.
These layers solve different problems. An approval is permission for an action; it is not an operating-system sandbox. And a sandbox does not tell you whether the action was a sensible response to the user’s request.
WebAssembly for extensions
Extensions make this more interesting. If an agent can write code that extends its own runtime, loading that code directly into the trusted process would give it far too much authority.
Ouroboros runs extension code as WebAssembly components through a separate Rust helper using Wasmtime.
A component declares the host functions it imports. The current host interface exposes a log-line function, not a filesystem, network connection or clock. An import the host does not provide is refused at load time. Execution also has resource limits.
That makes a repository-supplied hook much less alarming. We can give it an event and receive a verdict without giving it arbitrary shell access. A repository hook can tighten a decision, but cannot loosen the existing policy.
There are signed, deployed capabilities as well: components with messages, replies and retained state. The signature authorizes a particular artifact; the import boundary limits what that artifact can do. Signing arbitrary code would not, by itself, make it safe.
There is still a compiler involved in producing the component. Cargo build scripts execute on the build host, before Wasmtime ever sees the resulting artifact. The forge therefore builds under an OS sandbox with networking disabled. The resulting component’s containment does not magically protect the build that produced it.
What I mean by self-repair
There are two different mechanisms here, and I do not want to hide them behind the same word.
The first is runtime recovery. A process dies, the runtime examines retained state, and it restores what it can safely restore. Automatic native-session resume is bounded to one attempt per coordinator incarnation. If an in-flight turn was interrupted, the runtime does not assume that every effect is safe to repeat.
Imagine a command completed just before its process died, but the result never reached the coordinator. Retrying because we did not receive “success” could execute it twice. An unfinished acknowledged effect is therefore recovered as ambiguous. Sometimes the correct recovery is to stop and ask for reconciliation.
The second is agent-written improvement. An agent can work on Ouroboros itself and forge a component that the runtime can use. The self-development workflow keeps evidence, and the outer change script can prepare a commit or pull request but never merges it.
Component signing is a policy service using operator-controlled keys, not necessarily a human clicking approve for every signature. It validates the submitted artifact and must persist its signing decision before returning a signature. Human review and control of promotion remain separate responsibilities; a signature is not a code review.
A passing test is evidence for that test. A component answering a health probe is evidence that it answered the probe. Neither means the agent has proven its own change universally correct.
The useful loop is to make a bounded change, validate it against declared checks, review the evidence and decide whether to promote it. Failed rollout checks can trigger rollback; ambiguous evidence can require quarantine. The agent does not get to approve itself merely because it wrote an enthusiastic summary.
The outer workflow can also run a benchmark, but that step is optional and does not enforce a positive before-and-after result. I would not call a change an improvement just because that script completed. We still need to read the measurements and decide whether the change made the agent better.
The self-development documentation describes the machinery and its limits. This is supervised self-development, not an agent with unlimited permission to rewrite the program that constrains it.
Auditability is more than saving the chat
A transcript is useful for reading the conversation. It is not enough to establish whether a command was authorized, whether it ran, or whether its result was lost.
Ouroboros keeps records with different jobs. The effect ledger is the authority record. The conversation checkpoint is the model’s retained context. A turn journal holds the execution record used for verified replay.
The additional audit evidence system is opt-in. Standard mode keeps the operational ledger and journal. Local audit mode adds evidence recording but can continue when recording fails, with that failure visible. Required mode refuses dispatch or stops dependent execution when required evidence cannot be committed.
That stronger mode comes with constraints: it requires named identities and encryption configuration, and currently excludes MCP, WASM capabilities and component hooks from execution. Those paths do not yet meet its full recording and containment contract. You cannot just enable every feature and assume it carries the same audit guarantee.
Audit facilities add inspection and export, with an optional SQLite index for local search.
SQLite is an index here, not a distributed session database. The runtime does not require a separate database server, and the index is not a substitute for the underlying records.
Verified replay is particularly useful. It runs recorded turns through the real agent loop, substituting the recorded model output and tool results. It checks whether that loop reconstructs the expected requests and conversation. It does not call the model again or re-execute the tools.
If we call the model again to try a different answer, that is a fork, not a replay.
Replay also has limits. A missing result, a journal gap, or a changed workspace can prevent verification. The result should name that boundary instead of showing a reassuring green check for a run it could not reconstruct.
These records can contain sensitive project content. Auditability does not make them suitable for public sharing, and a local hash chain is not proof against an attacker who can rewrite the entire local history. Stronger custody requires a separately protected record outside that attacker’s control.
More than one machine, without pretending it is a cloud
BEAM distribution lets Ouroboros place subagents on other machines in the same cluster. A child can be placed according to machine or toolchain tags, and it holds its own Git worktree lease.
A worktree gives the child a separate checkout. The lease is how the runtime represents ownership of that workspace. Keeping ownership separate from a process ID matters because process IDs change after a restart. Uncommitted work also must not disappear just because a process reached cleanup.
This is useful if you want to use machines you already own. It is not a multi-tenant security model.
Every node that completes the Erlang distribution handshake is inside the same trust domain. Builder and signer roles reduce what normally runs on a node; they do not make a malicious cluster member harmless. TLS protects the connection, not the other nodes from a member you have already trusted.
The current system also does not promise highly available replicated state or partition-safe consensus. Durable local state and distributed execution are useful without claiming either of those things.
If you are building something similar
Start with one coding task you can verify. A failing test and a small repository are more useful than an impressive conversation you cannot check.
Give the session a stable identity, then decide which process owns each piece of state. Persist the transitions that matter before telling a client they succeeded. Record an effect before executing it, and represent an unknown outcome separately from a failure.
Keep authorization outside the model. If you add extensions, define their host imports before you define an extension marketplace. Keep the UI on the same authorization path as your other clients.
Then deliberately break the runtime. Kill a coordinator while execution is still alive. Fail a checkpoint write before and after its rename. Try to resume with a pending approval. Check that cleanup preserves an uncommitted edit. Ouroboros has tests around these cases because a successful chat does not exercise them.
You do not need to begin with a cluster or a self-improvement system. You need to know what one interrupted task is allowed to do next.
Try Ouroboros
The installation guide covers macOS and GNU/Linux releases for ARM64 and x86_64. The launch page provides this installer command; it downloads and executes a script, so inspect it first if that is not how you normally install software:
curl -fsSL https://github.com/monocursive/ouroboros/releases/latest/download/install.sh | bash
Then, from the project you want to work on:
ouro
Or open the browser interface:
ouro web
You can use ChatGPT sign-in or configure model API access. API usage is billed by the model provider; a Claude or Grok application subscription is not an API key.
I would start with a disposable checkout and a small bug. Read the proposed commands, inspect the diff, and run the tests yourself. Before giving it sensitive or unattended work, read the architecture and safety boundaries.
The code is on GitHub. If you try it, I am especially interested in the cases where recovery stops, an approval is confusing, or the records do not explain what happened. Those are useful bugs to bring back to this project.
Michaël Mazurczak
Fullstack developer, Lyon, France