How to build a model testbed inside an AI agent
A model benchmark is only useful if it runs your real agent, exercises believable user situations, and measures what your product actually needs. Let's build one from a production testbed.
TLDR: Do not benchmark an LLM in isolation. Run several models through the real prompts, tools and application services of your agent. Give them believable scenarios, assert on observable behaviour, repeat every case, and keep the transcripts for everything a regex cannot judge.
I recently needed to choose which models should power Benzaiflow, an ADHD-friendly planner with an AI companion called Noji.
Opening three playgrounds and asking each model the same question would have been easy.
It would also have told me almost nothing.
An agent is not just a model. It is a model surrounded by a system prompt, user context, tools, schemas, retries, timeouts, a database and a user interface. A model can write a very nice answer while calling the wrong tool. It can select the right tool with arguments that fail validation. It can return valid JSON, but take fifteen seconds to do it.
So I built a testbed inside the application.
Let’s see how it works, what was useful, and how you can build one without creating a second fake version of your agent.
What are we actually benchmarking?
The first question is not “Which model is the smartest?”
The useful question is:
Which model best fulfils this application’s contracts, with its real prompts and tools, at an acceptable speed and cost?
Benzaiflow has several AI lanes with different jobs:
- Chat handles conversations and can call tools that create, complete, schedule or decompose tasks.
- Onboarding guides a new user through a sensitive ADHD intake conversation.
- Decompose turns an overwhelming task into manageable steps.
- Enrich extracts structured information such as duration, energy and cognitive load.
Those lanes should not necessarily use the same model. Chat needs reliable tool use and good conversation. Enrichment needs valid structured output. Onboarding needs to respect a conversational flow without sounding like a form wearing a hat.
A single global score would hide all of that.
The first useful design decision was therefore to benchmark product lanes, not abstract capabilities.
Keep the production path, add one injection seam
The most dangerous eval harness is one that reimplements the agent.
You copy the prompt into a script, recreate a few tool schemas and call a provider directly. It works. Six weeks later the production prompt changes, a tool gains a required field, and your benchmark keeps testing the old system with great confidence.
Nice dashboard, wrong program.
Instead, the Benzaiflow services accept an optional model configuration:
export type InjectedModelConfig = {
model: LanguageModel;
providerOptions?: ProviderOptions;
temperature?: number;
};
export async function runAgent(
context: AgentContext,
modelOverride?: InjectedModelConfig,
) {
const config = modelOverride ?? getProductionModelConfig();
return streamText({
model: config.model,
providerOptions: config.providerOptions,
temperature: config.temperature,
// The real system prompt, messages and tools are assembled here.
});
}
Production does not pass an override. The testbed does.
This small seam lets the harness run Gemini, Claude or GPT through the exact same prompt assembly, tools and application logic used by real users. It also keeps provider-specific options attached to the model configuration. That matters more than it looks: caching, reasoning modes, JSON settings and temperatures can change latency, cost and behaviour.
The model name is not the complete experimental configuration.
Keep the seam narrow. We are replacing the engine, not rebuilding the car.
Build situations, not prompt collections
A list of isolated prompts is a start, but an agent acts inside a world.
When a user says “Do the low-energy one”, Noji needs the user’s current tasks, energy, calendar and previous messages. When a user asks it to schedule something it just created, the second tool call must reference the identifier returned by the first one.
The testbed uses scenarios containing:
- a stable identifier;
- a language;
- conversation history;
- realistic application context;
- expected tool calls and arguments;
- forbidden actions;
- text requirements;
- optional case-specific validation.
A simplified scenario looks like this:
const scenario = {
id: "complete-task-fr",
language: "fr",
messages: [
{ role: "user", content: "J'ai terminé la déclaration d'impôts" },
],
expect: {
calls: [
{
anyOf: ["completeTask"],
validate: ({ taskId }) =>
taskId === "t-impots" ? null : "wrong task selected",
},
],
allowedMutatingTools: ["completeTask"],
maxMutatingCalls: 1,
text: { language: "fr", minChars: 10 },
},
};
This is not testing whether the model knows what taxes are. It is testing whether our whole agent can resolve a user’s intent against the data the product actually provides.
Good scenarios come from real usage:
- users changing their mind halfway through a conversation;
- ambiguous references which should trigger clarification;
- French and English conversations;
- requests that must not mutate anything;
- a task created in one turn and referenced in the next;
- a low-energy user asking for a realistic next action;
- an overwhelmed user who needs decomposition, not a motivational paragraph.
The boring edge cases are usually the product.
Use a real database where identity matters
My first instinct was to mock every tool callback and return convenient objects.
That works until the agent chains tools.
Imagine the model correctly creates a task and then asks the decomposition tool to split it. If the fake createTask callback returns a made-up identifier, the production decomposition service looks it up and returns TASK_NOT_FOUND. The model behaved correctly, but the harness marks it as a failure.
Benzaiflow’s fixture creates a fresh in-memory PGlite database for every scenario. It seeds a believable day with tasks, a project, energy levels and calendar context. Calls which create tasks go through the production task service and write real rows. Later tools can resolve the returned identifiers exactly as they do in the application.
const testDb = await createTestDb();
const user = await createTestUser(testDb.db);
const project = await createTestProject(testDb.db, user.id);
await createTestTask(testDb.db, user.id, {
id: "t-report",
originalInput: "Write the quarterly report draft",
duration: 90,
energyLevel: "high",
cognitiveLoad: "high",
projectId: project.id,
});
const executionContext = {
db: testDb.db,
userId: user.id,
createTask: (input) => createTask(testDb.db, user.id, input),
};
Not every dependency needs to be real. Sending calendar invitations during a benchmark would be a creative way to lose friends.
The practical boundary is:
- use real local services and persistence when their behaviour affects the agent loop;
- replace external side effects with recorders;
- log every attempted mutation;
- create a fresh fixture for every run;
- dispose of it afterwards.
This gives us realism without letting a benchmark escape into the world.
Score behaviour before prose
Agent outputs have two layers:
- what the agent did;
- what the agent said.
The first layer is usually easier and more important to score.
For tool-using chat scenarios, the Benzaiflow scorer checks things such as:
- was the expected tool called?
- were its arguments valid?
- did calls happen in the right order?
- did any tool return an error?
- did the model call a forbidden mutating tool?
- did it mutate more than once?
- did it repeat the same call and enter a loop?
- did it return visible text in the user’s language?
The negative assertions matter. If the user says “Complete it” without a clear antecedent, an agent that confidently completes a random task is worse than one that asks a question. Testing only for expected calls would miss this class of failure.
if (!allowedMutatingTools.includes(call.name)) {
failures.push(`unexpected mutating tool called: ${call.name}`);
}
if (mutationCount > maxMutatingCalls) {
failures.push(`too many mutating tool calls`);
}
Structured-output lanes use different contracts. A decomposition can be valid JSON and still be useless. We check its schema, but also product constraints: sensible step counts, useful labels, realistic durations and whether a vague task exploded into a huge plan.
The scorer should explain failures, not just output zero.
“Expected createTask, got no call” is actionable. “Score: 0.73” is a small mystery you now own.
Do not automate taste by accident
Onboarding is mostly prose. It would be tempting to ask another LLM whether the reply is warm, supportive and on-brand.
That can be useful, but it also replaces one opaque model with two opaque models and a rubric.
For the Benzaiflow onboarding lane, I kept automated checks mechanical and tied them to real product behaviour:
- the response is not empty;
- it was not truncated by the token limit;
- it uses the user’s language;
- it lands on the one question selected by the server;
- it does not echo the option labels already rendered as buttons;
- it does not repeat the previous assistant message;
- it does not use explicit shaming, diagnostic or medical language;
- it does not ask another question after the user ends onboarding.
Warmth, humour, rhythm and whether it actually sounds like Noji are deliberately not reduced to regexes.
The harness writes transcripts for humans to read.
This distinction is important:
Automate contractual properties. Review subjective qualities as subjective qualities.
Otherwise your eval slowly becomes a machine for selecting the model that writes most like the person who authored the regexes.
Freeze everything you can
Models are already non-deterministic. There is no need to add accidental randomness.
A model comparison can take long enough to cross an hour or midnight. If your prompt says “today” or includes the next calendar event, later models may receive a different situation.
The Benzaiflow testbed freezes one reference instant and one timezone:
export const EVAL_REFERENCE_NOW_ISO = "2026-07-30T13:00:00.000Z";
export const EVAL_TIMEZONE = "Europe/Paris";
It also uses stable task identifiers and runs each scenario several times.
One run tells you that a model can pass. Repetitions tell you whether you can rely on it.
A 2/3 result is more informative than averaging three runs into 67%. It immediately tells us that the case is flaky. In an agent that mutates user data, occasional creativity is not always charming.
Freeze or record:
- time and timezone;
- fixture identifiers;
- exact model snapshots, not floating aliases;
- provider routing;
- temperature and reasoning settings;
- prompt and scoring revisions;
- number of repetitions;
- timeout and concurrency.
If any of those change, you may be running a different experiment.
Capture the whole execution
A final answer is not enough to debug an agent.
For every run, the testbed captures:
- visible text;
- tool calls and arguments;
- tool results and errors;
- finish reason;
- time to first token;
- total duration;
- input, output and cached tokens;
- structured-output fallback usage;
- loops, stalls and scripted fallbacks;
- estimated cost.
The aggregate report stays readable, while JSON captures and transcripts preserve the details.
This is the same principle as application observability: summaries tell you where to look, traces tell you what happened.
Cost also needs context. Provider list prices are not enough when one model caches most of a large system prompt and another does not. The testbed prices input, output and cached tokens separately when providers expose them. The table is dated because prices change, and unknown models still run with cost displayed as n/a.
An honest unknown is better than invented precision.
What the testbed changed in production
This was not an academic benchmark. It changed Benzaiflow’s routing.
In one revision-6 run, with 20 chat cases repeated three times:
| Model | Chat pass rate | Median first token | Median total | Estimated cost per run |
|---|---|---|---|---|
| Claude Haiku 4.5 | 57/60 (95%) | 1.3s | 2.7s | $0.0078 |
| Gemini 3 Flash Preview | 50/60 (83%) | 1.5s | 2.6s | $0.0085 |
| GPT 5.6 Luna | 49/60 (82%) | 2.5s | 4.1s | n/a |
The interesting part was not that one model “won”.
The failure digest showed how they differed. One model guessed which task to complete when the request was ambiguous. Another skipped the tool needed to decompose an overwhelming task. Some cases failed only once in three runs. A French low-energy planning case failed across all three models, which was a strong hint that the problem might be our tool contract or scenario rather than model selection.
The result supported moving the chat lane to a pinned Claude Haiku snapshot, while keeping structured extraction on Gemini. Prompt caching made the chat option economically viable, so the production configuration includes the same caching breakpoint measured by the testbed.
Different jobs, different models.
More importantly, the decision is documented as a reproducible product trade-off, not “Claude felt nicer when I tried it on Tuesday.”
A practical testbed structure
You do not need a platform to start. A folder and a command are enough.
scripts/evals/
├── run.ts
├── models.ts
├── score.ts
├── report.ts
├── clock.ts
├── chat/
│ ├── fixtures.ts
│ ├── scenarios.ts
│ └── run-chat.ts
├── onboarding/
│ ├── cases.ts
│ └── run-onboarding.ts
├── object/
│ ├── cases.ts
│ └── run-object.ts
└── results/
The runner accepts a model list, suite, repetition count, concurrency and timeout. A typical command looks like this:
bun run eval -- \
--models google:gemini-3-flash-preview,anthropic:claude-haiku-4-5-20251001 \
--suite chat \
--reps 3 \
--concurrency 2
Start with this loop:
- List your agent’s lanes. Separate conversation, extraction, planning and any other job with a different contract.
- Add a model injection seam. Keep all production prompt and tool assembly intact.
- Create five real scenarios per lane. Use support issues, manual testing and failures you have already seen.
- Seed believable state. Include the database records and context needed for tools to behave normally.
- Record side effects. External mutations become callbacks which append to a log.
- Write mechanical expectations. Expected calls, forbidden calls, argument checks, schema rules and language.
- Run each case at least three times. Mark mixed outcomes as flaky instead of hiding them in an average.
- Save reports and transcripts. You will need both when a number changes.
- Version the scorer. Never compare totals across incompatible scenario or rubric revisions.
- Turn failures into new cases. The testbed should grow from real product surprises.
Do not begin with a hundred scenarios and an LLM judge.
Begin with the five ways your agent can hurt, confuse or disappoint a user.
A few traps I fell into
Mocking too much
A mock returning synthetic identifiers made correct tool chains fail. Use real local persistence when later steps depend on earlier results.
Testing the model instead of the product
Generic reasoning questions do not tell us whether the model can operate our tools. Every scored rule should map to a user-visible product contract.
Scoring only the happy path
Checking that the expected tool appeared is not enough. Also check forbidden mutations, duplicate calls, tool errors and loops.
Comparing different revisions
Adding six harder scenarios changes the denominator and the difficulty. Put a scoring revision in every report and compare like with like.
Trusting one run
LLMs are stochastic systems. Run repetitions and keep per-case results visible.
Hiding all nuance in one score
Chat pass rate, structured-output validity, latency, fallbacks and cost describe different risks. Keep lane-specific tables and a failure digest.
Automating what should be read
Tone is part of the product, but a brittle style regex is not taste. Save transcripts and read them.
What now?
A useful model testbed is not a leaderboard.
It is an executable description of what your agent owes its users.
Once it exists, you can use it to compare models, upgrade a pinned snapshot, change a prompt, modify a tool schema or test a caching strategy. More importantly, failures stop being vague impressions. They become scenarios you can reproduce.
Start small. Use the real agent. Freeze the world. Record what it does.
Then read the transcripts.
The numbers will help you choose a model. The failures will help you build a better agent.
Michaël Mazurczak
Fullstack developer, Lyon, France