How to Write Evals for an AI Agent Before Shipping It
A step-by-step walk through building a real eval set for an agent, starting from actual failure examples rather than hypothetical test cases nobody will ever see in production.
AI & Tech Insights Team
September 30, 2026 · 3 min read
If you already know what an eval is conceptually, this is the practical part: the actual steps to build a working eval set for an agent before it ships, starting from real examples rather than cases you imagine might happen.
Step 1: Collect real failure examples first
Before writing any new test cases, pull together every example you already have of the agent behaving badly, flagged support tickets, manually noticed bad responses during development, edge cases a teammate hit while testing. This is more valuable starting material than hypothetical cases, because it reflects failures that actually occurred with real inputs, not ones a developer imagined in isolation.
Step 2: Turn each failure into a reusable test case
For each real failure, capture the exact input that triggered it and define what a correct or acceptable response would have looked like, either a specific reference answer, a rubric describing acceptable characteristics, or a rule-based check (did it call the right tool, did it refuse appropriately). Write this so it can run automatically and repeatedly, not just document the one-off incident.
Step 3: Add coverage for known-risky categories, even without a prior failure
Beyond documented failures, deliberately write test cases for categories known to be risky for agents generally: ambiguous requests, requests that should trigger a refusal, multi-step tasks where an early wrong turn compounds, and inputs designed to test whether the agent stays within its intended scope.
Step 4: Decide how each test case gets scored
Some cases can be checked with a simple rule (did the correct tool get called, is a required disclaimer present). Others need a more nuanced judgment, an automated grader model scoring against a rubric, or human review for the highest-stakes cases where automated grading isn't reliable enough yet. Mixing scoring methods across a test suite is normal, forcing everything into one scoring approach usually produces a weaker eval set than matching the method to what each case actually needs.
Step 5: Run the full set before every meaningful change
Not just before a major release, before any change to the prompt, the model, or the available tools, since small changes can shift behavior in ways that are hard to predict without actually testing. Track the score over time as a trend, not just a pass or fail snapshot, so a gradual decline in quality is visible before it becomes a real incident.
Step 6: Add a new test case every time something breaks in production
This is the step that determines whether an eval set stays useful over time or quietly goes stale. Every real production failure, once diagnosed and fixed, should become a permanent addition to the eval set, so the exact same failure literally cannot ship again undetected.
The realistic time investment
Building a genuinely useful eval set from scratch takes real, deliberate effort, and it's tempting to skip in favor of shipping faster. The teams that skip it consistently end up rediscovering the same categories of failure repeatedly in production, at a real cost in user trust, that the upfront eval investment would have caught before it ever reached a real user.
© 2026 AI & Tech Insights. All rights reserved. This article may not be reproduced without permission. See our disclaimer.
← Previous
How to Structure a Page So AI Search Actually Cites It: A Practical AEO Checklist
Next →
Langfuse vs Braintrust vs Arize Phoenix: Agent Observability Tools Compared
Related articles
What Breaks When You Scale an AI Agent from Demo to Production
A working demo tested a handful of times by the team that built it survives contact with real users surprisingly poorly. Here's specifically what tends to break, and why.
Sep 30 · 3 min read
How to Version-Control and Test Prompts Like Real Code
A prompt that gets edited directly in a dashboard with no history, no review, and no tests is exactly the kind of untracked change that causes production incidents nobody can trace.
Sep 30 · 3 min read
Structured Output and Function Calling: Common Failure Modes
Function calling is reliable enough that it's easy to stop checking it carefully. The failures that do happen tend to be subtle, wrong-but-valid outputs rather than obvious crashes.
Sep 30 · 3 min read
Get new guides by email
Useful AI and tech guides, occasionally. No unnecessary emails.