AI Agent Evals in Production: Tool Use, RAG, MCP, and Regression Testing
A practical framework for evaluating AI agents before model upgrades, prompt changes, tool additions, or MCP server rollouts.
Built for: Teams shipping coding agents, research agents, internal copilots, support automations, and tool-using assistants.
key takeaways
- +Agent evals need stateful tasks, not only single-turn Q&A.
- +Measure tool choice, arguments, side effects, recovery, refusal, and final answer quality separately.
- +MCP-style tool surfaces make eval fixtures more portable across clients.
- +Run regression evals before model upgrades because better models can still break workflows.
architecture - agent stack layers
User workflow
Task, risk, approval, expected artifact
Agent runtime
Graph, handoffs, memory, retries, traces
Protocols
MCP for tools and context; A2A for remote agents
Tools
Search, files, code, databases, SaaS APIs
Governance
Identity, permissions, audit, evals, rollback
Why agent evals are different
A normal LLM eval asks whether an answer is good. An agent eval asks whether a sequence of decisions was good: which tool was chosen, what arguments were passed, whether the agent noticed failure, whether it retried safely, and whether the final user-visible answer reflects the tool result.
Single-turn benchmark scores are not enough for production agents. Your eval set should look like the work the agent actually performs.
Eval case schema
Represent each eval as a task with allowed tools, fixture data, expected tool calls, expected final state, and a grader. This makes tests runnable in CI and comparable across model versions.
Agent eval case
json
{
"id": "billing-refund-policy-001",
"input": "Can we refund this annual account? Customer bought it 9 days ago.",
"fixtures": {
"customer": {"plan": "annual", "purchase_days_ago": 9},
"policy": "Annual plans are refundable within 14 days."
},
"expected_tools": [
{"name": "get_customer_account"},
{"name": "search_policy"}
],
"expected_final": "eligible_with_policy_citation",
"must_not": ["issue_refund_without_confirmation"]
}Metrics that matter
Score the chain, not just the final prose. A correct answer with an unauthorized tool call is a failure. A safe refusal after a tool outage might be a pass. This is where agent evals become product-specific.
- -Tool selection: did it call the right tool or avoid tools when unnecessary?
- -Argument quality: were IDs, dates, filters, and permissions correct?
- -Recovery: did it handle 404s, timeouts, empty results, and conflicting evidence?
- -Grounding: are claims supported by retrieved documents or tool output?
- -Side effects: did it ask before mutating external systems?
MCP and eval portability
The Model Context Protocol matters because tools are becoming portable surfaces rather than one-off SDK calls inside a single app. As the protocol adds stateless operation, authorization hardening, and cacheable list results, evals should assert both model behavior and tool contract behavior.
Treat every new MCP server as a dependency that needs fixtures, permission tests, and failure-mode evals.
CI loop
Run a small blocking eval suite on every prompt or tool change. Run a wider suite before model upgrades. Log failures by category so the fix is obvious: retrieval, tool schema, prompt, model, or product policy.
Regression gate shape
bash
npm run agent:evaluate -- --suite smoke --model candidate
npm run agent:evaluate -- --suite tool-regression --model candidate
npm run agent:evaluate -- --suite rag-grounding --model candidateSources and further reading
OpenAI Evals guide
Programmatic evaluation guidance and model-output testing.
Hugging Face Lighteval
Toolkit for evaluating LLMs across inference backends.
Ragas eval guide
LLM application evaluation workflow.
MCP 2026-07-28 specification announcement
Current MCP changes including stateless protocol core and authorization hardening.