> work with me

Production Agent Reliability Sprint

Your agent completes the workflow. Then a tool times out, a process restarts, or someone presses cancel. We need to know what happens next.

I spend two weeks working with you on one critical workflow: understand where it can fail, test those cases, and implement the change that matters most before you put it in front of more users.

Duration
Two weeks
Scope
One workflow
Tell me about your agent

Booking for November and December.

> the problem

A timeout leaves an awkward question.

Imagine your agent updates a customer record. The API accepts the change, but the connection drops before the response arrives. The agent sees a timeout. The customer record has already changed.

Retrying might apply the action twice. Calling it a failure would also be wrong. Somewhere, your program needs enough information to work out what happened, or to stop and let a person check.

These are the details I like working on. What survives a restart. What a tool is allowed to touch. What stops a loop from spending money all night. What you can inspect when the agent says it is finished.

The sprint gives us time to answer those questions in your system and fix the most consequential gap we agree to tackle.

> the work

What we will work through.

We start with an agent that already does something useful. You show me the workflow, the code, and the part you are least comfortable relying on.

  1. A review of the system around the model

    I trace the workflow through your code, tools, accounts and stored state. We map the failure modes and threats, decide what the agent may change, and write down which actions need a human's approval and who owns those decisions.

  2. Tests for the workflow you care about

    We define what a correct result looks like, then test the real execution path. That includes interruptions, failed tools and repeated requests. You get the cases, the results, and a clear account of what they do and don't prove.

  3. A plan for retries, recovery and running costs

    We work out what can be retried, how cancellation should behave, and how to avoid repeating an action that already happened. I review token budgets, spend limits and runaway loops, along with the records an operator needs to reconstruct what the agent actually did.

  4. The most important fix, implemented

    I implement the agreed change that addresses the highest production risk for this workflow, and check it against the relevant tests. Depending on your system, that might be a permission check before tool execution, recovery after a crash, or a limit that stops an expensive loop.

  5. A handoff you can keep working from

    You get the code, the findings, and a production roadmap ordered by risk. We go through what changed, what remains uncertain, and which conditions should stop a rollout. The support window and channel are agreed in writing before we start.

Two weeks will not settle every production question. You will know what we tested, have one important improvement in place, and know what still needs work.

> two weeks

First understand it. Then change it.

Before we start, we agree on the workflow, the improvement to deliver, access to the system, and who can make decisions. The scope and exclusions go in writing.

Week 1

Read, trace, test.

I review the architecture and permissions, follow the workflow through the system, and establish the evaluation cases and current behaviour. We check the proposed fix against what we find.

Week 2

Implement, verify, hand over.

I implement the agreed improvement, run the checks, and document the remaining risks. We finish by going through the code, the results and the next steps together.

> working together

Small enough to finish properly.

This works best when you have a working prototype, one workflow that matters, and an engineer or founder available to work through decisions with me. The aim is to give your team a system they can understand, operate and keep improving.

The sprint covers the review, tests, one implemented improvement and the roadmap. Additional workflows, new product surfaces, a platform rebuild or a product built from scratch need a separate scope. Marketing, fundraising materials and promises about adoption are outside this work.

After the handoff

There is a support window for questions about the delivered work. We agree on its duration and channel in the engagement letter. Ongoing operations, on-call coverage, an SLA, staff augmentation and unlimited revisions are outside the sprint. If more work makes sense, we agree on it separately.

> who you will work with

I'm Michaël.

I'm a software engineer in Lyon, shipping software since 2012. I build Ouroboros, Hubfluencer Studio and Benzaiflow. Monocursive is where I bring that engineering work to other teams.

While building Ouroboros, I have spent a lot of time on the awkward parts of agent execution: recovering a session, checking permission before a tool runs, and keeping a useful record when something fails. I write about those decisions, including their limits.

If you want to see how I approach the work, start with Durable agent execution is a state and authority problem or the introduction to Ouroboros.

> price and timing

€12,000–€20,000

A fixed fee for the two-week sprint. The amount depends on the depth of implementation in your stack; we agree on the scope and exact price before starting.

I take one engagement at a time, with one slot in November and one in December. If those are taken, we can discuss a later date.

Tell me what your agent does.

A short email is enough: the workflow, your stack, and what worries you about running it in production. We will work out whether two weeks together would be useful.

Tell me about your agent

contact@monocursive.com