How to Actually Evaluate an Agent
LLM evals grade answers; agents have permissions. A working framework for testing outcomes, process and safety before things break.
contents
Almost every project I’ve touched lately is building an agent. But the hard part isn’t getting one running — it’s whether it can finish tasks stably, reliably, and safely once it does.
Traditional LLM evals mostly grade the answer. Agents are different: they hold far more power — calling tools, hitting servers, connecting to third-party services — so the risk expands well beyond simple answer accuracy.
So here’s a quick walk through the dimensions an agent eval should actually cover.
Three Things to Test
Test only the outcome and you never see the agent’s bad habits. Test only the process and you may block it from finding a better solution. Skip safety entirely, and the more permissions it holds, the worse things get. All three legs matter.
“The bigger its permissions, the bigger the risk.”
Anatomy of a Good Case
A good test case has at least six parts:
- User input — how a user would actually phrase the request
- Initial environment — database, files, tool return values, page state
- Success criteria — what counts as the task being done
- Failure criteria — what must be ruled a failure
- Scoring method — deterministic checks, an LLM judge, or human review
- Risk flags — does it touch privacy, permissions, funds, deletion, or outbound sends
Designing the Scorer
| Scorer | Best for | Cost | Stability |
|---|---|---|---|
| Deterministic | Anything a program can judge: code tests, field extraction, database state, tool-parameter matching | Low | High |
| LLM judge | Open-ended content: is the answer professional, is the tone right, is the summary complete, does the report hold together — needs periodic calibration | Medium | Medium |
| Human review | High-value, high-risk tasks: medical, legal, financial, enterprise approvals — irreplaceable when the boundaries are fuzzy | High | High |
“Use code where code works, hand what code can’t test to an LLM, and keep humans for the critical and contested samples.”
This layered setup is the mainstream scoring design in AI right now — OpenAI’s Evals and Graders docs follow the same logic, supporting string checks, similarity scoring, model graders, Python code graders, and so on.
Btw, my personal take: in most cases an LLM judge is necessary. Only rarely can you get away with pure deterministic code scoring and keep the LLM out of the loop entirely.
Capability vs. Regression
Two eval types really matter for agents: capability evals and regression evals.
- Capability evals — what’s the hardest thing this Agent can pull off? Can a coding Agent fix a gnarly bug? Can a research Agent complete a multi-source investigation?
- Regression evals — the things it used to do well, does it still do them? Every prompt tweak, model swap, or new tool can quietly degrade old functionality.
Testing the Trajectory
Yes, Agents call tools. No, you shouldn’t dictate a fixed calling order.
Anthropic’s advice: don’t over-constrain the Agent’s creativity. Often it won’t take the path you imagined — and still finishes the task a better way. So process evals should lean on key constraints instead. For example:
- It must ground its answer in the policy-lookup result
- It must not call the tool that deletes user data
- The amount in tool parameters can’t exceed the order amount
- Total turns can’t exceed 10
- Token cost can’t blow past the budget
Unless you’re in a hard-process business — approvals, payments, compliance review — don’t casually write rules like “step one must call A, step two must call B.”
Safety Evals
Agents connect to tools, databases, browsers, file systems, and third-party services. One malicious input is all it takes to trick them into doing the wrong thing. So a safety eval needs at least five categories:
- Prompt injection — malicious instructions smuggled inside web pages, documents, or tool return values
- Privilege escalation — a regular user requesting admin operations
- Privacy leaks — demands to dump raw customer records, keys, or internal data
- High-risk operations — does it ask for confirmation before deleting, sending, paying, or approving?
- Tool abuse — does it call unnecessary or dangerous tools?
Closing Notes
This isn’t a long piece — it just lays out the main dimensions of Agent evals. Any real evaluation will need to go a lot deeper on each one.
One more thing: don’t build an Agent just to have built an Agent. An Agent only means something if it has a real use case. Otherwise you’re just burning time and tokens.
First posted as a thread on X, June 2026. Corrections and replies: hi@0xtz.us