Developers

How to Write Evals for an AI Agent Before Shipping It

A step-by-step walk through building a real eval set for an agent, starting from actual failure examples rather than hypothetical test cases nobody will ever see in production.

A&

AI & Tech Insights Team

September 30, 2026 · 3 min read

If you already know what an eval is conceptually, this is the practical part: the actual steps to build a working eval set for an agent before it ships, starting from real examples rather than cases you imagine might happen.

Step 1: Collect real failure examples first

Before writing any new test cases, pull together every example you already have of the agent behaving badly, flagged support tickets, manually noticed bad responses during development, edge cases a teammate hit while testing. This is more valuable starting material than hypothetical cases, because it reflects failures that actually occurred with real inputs, not ones a developer imagined in isolation.

Step 2: Turn each failure into a reusable test case

For each real failure, capture the exact input that triggered it and define what a correct or acceptable response would have looked like, either a specific reference answer, a rubric describing acceptable characteristics, or a rule-based check (did it call the right tool, did it refuse appropriately). Write this so it can run automatically and repeatedly, not just document the one-off incident.

Step 3: Add coverage for known-risky categories, even without a prior failure

Beyond documented failures, deliberately write test cases for categories known to be risky for agents generally: ambiguous requests, requests that should trigger a refusal, multi-step tasks where an early wrong turn compounds, and inputs designed to test whether the agent stays within its intended scope.

Step 4: Decide how each test case gets scored

Some cases can be checked with a simple rule (did the correct tool get called, is a required disclaimer present). Others need a more nuanced judgment, an automated grader model scoring against a rubric, or human review for the highest-stakes cases where automated grading isn't reliable enough yet. Mixing scoring methods across a test suite is normal, forcing everything into one scoring approach usually produces a weaker eval set than matching the method to what each case actually needs.

Step 5: Run the full set before every meaningful change

Not just before a major release, before any change to the prompt, the model, or the available tools, since small changes can shift behavior in ways that are hard to predict without actually testing. Track the score over time as a trend, not just a pass or fail snapshot, so a gradual decline in quality is visible before it becomes a real incident.

Step 6: Add a new test case every time something breaks in production

This is the step that determines whether an eval set stays useful over time or quietly goes stale. Every real production failure, once diagnosed and fixed, should become a permanent addition to the eval set, so the exact same failure literally cannot ship again undetected.

The realistic time investment

Building a genuinely useful eval set from scratch takes real, deliberate effort, and it's tempting to skip in favor of shipping faster. The teams that skip it consistently end up rediscovering the same categories of failure repeatedly in production, at a real cost in user trust, that the upfront eval investment would have caught before it ever reached a real user.

Share:XLinkedInWhatsApp

© 2026 AI & Tech Insights. All rights reserved. This article may not be reproduced without permission. See our disclaimer.

Related articles

Get new guides by email

Useful AI and tech guides, occasionally. No unnecessary emails.