Durable agent execution is a state and authority problem
A prompt loop can choose the next action. A durable agent runtime must also know what happened, who may act, and what to do when neither answer is certain.
TLDR: An agent loop can decide what to do next. It cannot, by itself, prove which work already happened, decide whether an interrupted effect is safe to repeat, retain a workspace claim across a crash, or establish who is allowed to run a command. Those are state and authority problems. While building Ouroboros, I found that durable execution comes from making both explicit: stable identities, committed checkpoints, an effect ledger, permission boundaries, workspace leases and a supervision tree that restarts consumers when their authority disappears.
In Demystifying Agents, I built a small agent in TypeScript. Its central loop is deliberately uncomplicated: send messages and tools to a model, execute the tool call, append the result, then ask the model what to do next.
That is enough to make an agent useful.
It is not enough to make one durable.
Imagine that a coding agent runs a command which opens a pull request. The command succeeds on GitHub, then the runtime crashes before it records the result. When the process starts again, the conversation still ends before the tool result.
Should it run the command again?
The prompt loop has no correct answer. Retrying may create a duplicate pull request. Continuing as if it succeeded may invent state. Reporting failure may also be false: the pull request exists, even if the runtime did not hear about it.
The model did not reason badly. The runtime lost the boundary between intention and effect.
This article is about that boundary. I will use parts of Ouroboros, my Elixir runtime for coding agents, but the design problem is not specific to Elixir or to coding. Any agent which survives restarts and can change the world eventually needs answers to two different questions:
- State: what has been committed, what is still in flight, and what can be recovered?
- Authority: which component, process or person may perform, approve, publish or revoke an action?
A longer prompt does not answer either one.
The loop owns a turn, not the session
The simplest agent architecture often gives one process too many jobs:
receive a prompt
-> call the model
-> execute tools
-> append messages
-> stream events
-> save the conversation
-> repeat
This is attractive because all the relevant values are already in memory. There is one message array, one current tool call and one stream going back to the client.
It also quietly makes the lifetime of the work equal to the lifetime of that process.
A process identifier is not a durable identity. After a restart, a new process can be given the same input, but it is not automatically the same execution. It does not know whether the previous process reached the model, whether the model response was charged, whether a tool started, or whether an event reached the user.
In Ouroboros I separate three identities:
logical session
owns the durable history and public progress
coordinator process
owns recovery, publication and workspace admission
execution runtime
owns the live conversation and turn scheduling
The logical session must outlive either process. The coordinator and execution runtime can disappear independently, and that distinction changes recovery.
If the coordinator crashes while the execution runtime remains alive, recovery should reattach to the same runtime. Starting another model turn would duplicate work. If the runtime disappears, recovery may need a replacement, but that replacement must start from a committed checkpoint rather than from whatever the old process last held in memory.
The native session therefore has separate logical and runtime identifiers, plus an attachment cursor. Its output buffer is useful for delivery, but it is not the public source of truth. The coordinator’s durable checkpoint is.
You can see this split in Ouroboros.Provider.Native.Session and the durable projection in Ouroboros.Interactive.Store.
This gives recovery something better than “run the prompt again”. It can ask:
- Which logical session is this?
- Which runtime generation produced the pending output?
- Up to which cursor did the coordinator commit?
- Can the same runtime be adopted?
- If not, which checkpoint is authoritative for its replacement?
Durability starts with identity because every other record needs to say whose work it describes.
Commit progress before publishing it
Streaming makes a runtime feel immediate. It also creates two histories:
- what the runtime has produced;
- what the durable session says it has produced.
Those histories must not drift silently.
Suppose a runtime emits three events, the coordinator sends them to a client, then its checkpoint fails. The user has seen progress which recovery cannot prove. On restart, the coordinator may emit those events again, omit them, or continue from an older conversation.
The transport got ahead of the state.
Ouroboros uses a stricter publication order:
runtime produces a bounded batch
-> coordinator drains the batch
-> coordinator updates its durable checkpoint
-> runtime acknowledges the committed cursor
-> coordinator publishes the batch
If the checkpoint fails, the batch remains unacknowledged and unpublished. The coordinator can drain it again after recovery. This does not make every client delivery exactly once; a subscriber can still disconnect at an awkward moment. It does make the runtime’s public history follow one durable cursor.
The distinction matters. “I received these bytes” is a transport fact. “These events are part of the session” is a state transition.
The file adapter has to be honest about that transition too. Ouroboros.Storage.DurableFile writes a checkpoint through a temporary file:
open temporary file exclusively
-> write the checkpoint
-> sync the file
-> close it
-> rename it over the previous checkpoint
-> sync the parent directory
A failure before the rename is relatively simple. The previous checkpoint is still authoritative.
A failure after the rename is not. The new file may already be visible, but if the parent-directory sync fails, the adapter cannot prove whether the new directory entry will survive a machine failure. Ouroboros reports commit_outcome_unknown instead of turning uncertainty into an ordinary rejection.
That state is inconvenient, which is why it is useful.
If the caller treated it as “the write failed”, it could continue from the old value while the new value is actually present. If it treated it as success, it could publish progress which was never made durable. The only safe next action is reconciliation: read the authority again and determine which value won before dependent execution continues.
Atomic rename is a mechanism. A truthful commit contract is the design.
Record an effect before allowing it to happen
Conversation checkpoints tell us what the agent believed. They do not tell us enough about external effects.
A tool call can change a file, send a request, start another agent or execute a shell command. By the time its result becomes another assistant message, the important part may already have happened.
For effects, Ouroboros uses a write-ahead ledger. The reduced lifecycle looks like this:
permission decision
-> record attempt as started
-> commit that record
-> execute the effect
-> record ok, failed, denied or ambiguous
The initial started record is an admission boundary. If it cannot be committed, the effect does not run.
The ledger entry contains more than a tool name. It binds the attempt to stable identities, an effect fingerprint and the authority snapshot under which it was admitted. A later process can distinguish “this exact request was denied” from “a similar request is being attempted under a different rule”.
The public interface in Ouroboros.Agent.EffectLedger makes the ordering visible:
record_started(attrs)
# only after this durable write succeeds:
run_effect()
record_settled(id, outcome)
This is simplified, but the invariant is not: no durable start, no admissible effect.
Now return to the pull request which may have been created before the crash. Recovery finds a started entry without a settlement. It does not convert that entry to failed, because process death is not proof that the external action failed. It marks the attempt ambiguous.
Ambiguity is a terminal runtime fact, not necessarily a terminal business fact. A GitHub-specific reconciler could search for a pull request carrying the idempotency marker and settle the attempt later. A shell command may require an operator to inspect the workspace. Some effects cannot be resolved at all.
The ledger does not promise exactly-once side effects. It promises not to hide when exactly-once cannot be proven.
That is a more useful guarantee.
Recovery should adopt before it repeats
“Retry on crash” sounds sensible until the crashed operation has effects.
For pure computation, retrying is usually fine. For a paid model call, a file mutation or an API request, it can be expensive or destructive. A durable runtime needs a recovery policy based on the state it can prove, not a universal retry policy.
Ouroboros recovery first looks for work which can be adopted:
- a surviving execution runtime with the expected logical identity;
- a pending output batch beyond the committed cursor;
- a workspace reservation belonging to the recovering session;
- a checkpoint from which a replacement can safely resume;
- effect attempts which need reconciliation rather than repetition.
This also means recovery may refuse to proceed. A runtime with the wrong generation, a workspace claim which cannot be authenticated, or a checkpoint with an unknown commit outcome is not “close enough”. Durable systems become dangerous when availability pressure encourages them to guess.
The recovery path in Ouroboros.Session.Recovery reconstructs ownership from persisted records. Tests cover coordinator crashes which reattach the same runtime and checkpoint failures which keep an output batch unpublished until a later commit.
The prompt is the last thing recovery should repeat, not the first.
A journal is evidence, not another conversation
A saved message array is helpful for resuming a model. It is a poor audit record.
Messages do not naturally tell us which system prompt was used, which model request was assembled, when control input arrived, which tool result came from the world, or whether part of the history was dropped after a write failure. Editing the array in place also makes it difficult to detect tampering or truncation.
Ouroboros keeps an append-only turn journal. Each record has a sequence number, its predecessor’s hash and its own SHA-256 hash. The chain lets verification name the first record which no longer follows from the previous one.
There is one detail I particularly like: if an append fails, the next successful write can contain a gap record describing which kinds of records were lost.
It cannot reconstruct their content. That would be fiction. It can preserve the fact that the record is incomplete.
The journal then supports a constrained replay. Ouroboros.Provider.Native.Replay runs recorded turns through the shipped loop while replacing nondeterministic inputs:
- recorded model chunks replace live inference;
- recorded tool results replace live tool dispatch;
- recorded steering replaces mailbox timing;
- recorded timestamps replace the current clock;
- prompt and conversation digests are derived again and compared.
The important part is what replay does not do. It does not execute tools, write checkpoints or add ledger entries. Replaying an audit trail should not create another pull request to verify whether the first pull request happened. That would be admirably recursive and operationally unhelpful.
Replay also stops at honest boundaries: a journal gap, a truncated prefix, an unsettled turn, unavailable attachment bytes, a compaction it cannot reproduce, or a tool seam which is no longer wired. The turns before the boundary can still be verified. The turns after it are not guessed.
A durable runtime needs both recovery and replay, but they answer different questions:
- recovery asks, “What work may safely continue?”
- replay asks, “Can this implementation derive the recorded result without repeating effects?”
Neither answer belongs in the prompt.
Permissions must sit in front of execution
A system prompt can tell a model not to edit files. This may reduce how often it asks to edit files. It does not remove the write tool’s authority.
The model is one participant in the system, not its security boundary.
A runtime permission check must happen after the model requests an action and before the tool performs it. The check needs the structured action, its effect class, the current session posture, the workspace and the applicable rules.
Ouroboros routes native tool requests through Ouroboros.Provider.Native.Permissions. A planning session, for example, refuses write and execute actions at this layer even if an ordinary repository rule would allow them. Read and network actions remain available because a plan which cannot inspect its subject is mostly a confident guess.
This ordering makes posture an authority:
model requests a tool
-> session posture applies stricter bounds
-> permission engine evaluates rules
-> runtime allows, denies or asks
-> admitted effect enters the ledger
-> tool may run
Notice that each arrow narrows what may happen. A lower layer cannot quietly restore authority removed by a higher one.
Human approval is another authority transition. An approval request needs an identity, a deadline, a bounded queue and a durable answer. Ouroboros records the answer before telling the waiting caller that it was approved. If the coordinator restarts while questions are pending, it denies the requests it can no longer honestly associate with a live interaction.
That is less convenient than resurrecting a button from an old screen. It is also less likely to turn yesterday’s click into today’s shell command.
An approval is not merely another message in the conversation. It changes who has authorized an effect.
The workspace is an owned resource
Coding agents make authority physical. Two sessions can have perfectly coherent conversations while overwriting the same file.
The current working directory is not enough as a boundary. Paths can contain .., symbolic links can escape an expected root, and two different strings can resolve to overlapping locations. A durable runtime needs to canonicalize the root before admitting it, then grant an owned claim.
Ouroboros.Workspace.Manager issues leases for canonical workspace roots. Shared reads can coexist, while an exclusive lease rejects overlapping claims. The manager monitors the owning process so ordinary process death releases its lease.
Recovery complicates that pleasant rule.
If every lease vanished immediately with its process, another session could acquire the workspace while the durable owner was still recoverable. Ouroboros therefore reconstructs reservations from durable session records. A replacement coordinator must prove the expected task identity and ownership before it can convert that reservation back into a live lease.
The difference is small in code and large in consequence:
process died, therefore the workspace is free
is not equivalent to:
process died, therefore consult durable ownership
For parallel coding work, Ouroboros can also provision separate Git worktrees. This reduces physical overlap, but it does not replace leases. Each worktree still has an owner, and cleanup retains dirty trees rather than deleting changes because a marker says they are old.
The workspace manager is deliberately node-local. It is not a distributed lock. If several machines can mutate the same physical filesystem, they must route claims through one authority or use a consensus-backed mechanism.
Calling a local lease “distributed” does not improve it. It only makes the incident report more surprising.
Policy and containment are different boundaries
Permission answers whether an action is admitted. A sandbox limits what an admitted process can actually reach.
These are related, but they are not substitutes.
A policy can approve a command because its declared action looks reasonable while the command touches more of the machine than expected. A sandbox can block that access. Conversely, a sandbox error cannot explain whether the command was forbidden by a project rule, refused because the session was planning, or merely unsupported by the current platform.
Ouroboros therefore keeps a permission decision in front of execution and applies operating-system containment separately. The current sandbox support uses Seatbelt on macOS, bubblewrap on Linux, or an explicit unsandboxed mode. Those backends do not provide identical guarantees. Network fencing is coarse, bubblewrap does not add a seccomp filter, and Apple’s sandbox-exec is deprecated.
“Sandboxed” is not a portable boolean.
The runtime has to describe the actual boundary it established, fail closed for unknown modes, and avoid presenting unrestricted execution as containment. Security comes from composing named authorities and limits, not from attaching one reassuring adjective to the tool runner.
Supervision order is part of the security model
OTP supervision is usually introduced as a reliability feature: when a process crashes, restart it.
For an agent runtime, the more interesting question is which other processes must stop when it crashes.
Ouroboros’s top-level supervisor uses rest_for_one. Children start in an intentional order. Durable effect, maintenance, permission and workspace authorities sit before the execution and coordinator subtrees which consume them.
A reduced view looks like this:
durable storage boundaries
-> effect ledger
-> maintenance authority
-> grants and permission authority
-> workspace authority
-> execution subtree
-> coordinator subtree
If the permission authority restarts, live sessions below it must not continue using decisions from the previous generation. rest_for_one stops and restarts those consumers. If a coordinator crashes, however, the independent execution subtree can remain alive so recovery can reattach instead of repeating the turn.
You can read the complete ordering in application.ex. It is not just a startup list. It encodes which state owners dominate which consumers, and which failures should remain local.
This is where state and authority meet.
A checkpoint without an owner can be written concurrently. A permission rule without a generation can outlive the authority which interpreted it. A lease without recovery semantics can be released too early or held forever. A supervised process without a durable identity can restart cleanly into the wrong work.
Restarting is easy. Restoring the right relationships is the runtime.
Design the state machine before the prompt loop
If I were building another durable agent runtime, I would still begin with a small model-and-tool loop. It is the fastest way to discover whether the product is useful.
Before allowing long-running or consequential work, I would then make these contracts explicit:
- Give work a stable identity. Separate the logical session from every process, socket and worker which may temporarily execute it.
- Name the publication point. Decide exactly which durable transition makes output part of the public session.
- Write ahead of effects. Record admission before execution and settlement afterwards.
- Represent uncertainty. Keep
ambiguousandcommit_outcome_unknowndistinct from success and failure. - Recover by adoption. Reattach to surviving work and reconcile uncertain work before retrying anything.
- Put permissions outside the model. Treat prompts as guidance and runtime checks as authority.
- Own shared resources. Canonicalize workspaces, lease them, and rebuild reservations from durable state.
- Keep evidence append-only. Record gaps when evidence is missing and stop replay where verification becomes impossible.
- Order failure domains. Restart consumers when the authority they relied on has been replaced.
- State the limits. Local locks, platform sandboxes and effect-specific reconciliation should not be advertised as stronger guarantees than they provide.
None of this makes the model more intelligent. It makes the surrounding program less willing to lie.
The prompt loop remains important. It decides the next action, incorporates tool results and eventually produces an answer. But once an agent can run for hours, survive a terminal, spend money, edit a repository or ask a human for approval, the loop is no longer the architecture.
The architecture is a state-and-authority machine with an LLM inside it.
Michaël Mazurczak
Fullstack developer, Lyon, France