Agents15 min / updated 2026-09-30

AI Agent Evals in Production: Tool Use, RAG, MCP, and Regression Testing

A practical framework for evaluating AI agents before model upgrades, prompt changes, tool additions, or MCP server rollouts.

Built for: Teams shipping coding agents, research agents, internal copilots, support automations, and tool-using assistants.

key takeaways

  • +Agent evals need stateful tasks, not only single-turn Q&A.
  • +Measure tool choice, arguments, side effects, recovery, refusal, and final answer quality separately.
  • +MCP-style tool surfaces make eval fixtures more portable across clients.
  • +Run regression evals before model upgrades because better models can still break workflows.

architecture - agent stack layers

User workflow

Task, risk, approval, expected artifact

Agent runtime

Graph, handoffs, memory, retries, traces

Protocols

MCP for tools and context; A2A for remote agents

Tools

Search, files, code, databases, SaaS APIs

Governance

Identity, permissions, audit, evals, rollback

Why agent evals are different

A normal LLM eval asks whether an answer is good. An agent eval asks whether a sequence of decisions was good: which tool was chosen, what arguments were passed, whether the agent noticed failure, whether it retried safely, and whether the final user-visible answer reflects the tool result.

Single-turn benchmark scores are not enough for production agents. Your eval set should look like the work the agent actually performs.

Eval case schema

Represent each eval as a task with allowed tools, fixture data, expected tool calls, expected final state, and a grader. This makes tests runnable in CI and comparable across model versions.

Agent eval case

json

{
  "id": "billing-refund-policy-001",
  "input": "Can we refund this annual account? Customer bought it 9 days ago.",
  "fixtures": {
    "customer": {"plan": "annual", "purchase_days_ago": 9},
    "policy": "Annual plans are refundable within 14 days."
  },
  "expected_tools": [
    {"name": "get_customer_account"},
    {"name": "search_policy"}
  ],
  "expected_final": "eligible_with_policy_citation",
  "must_not": ["issue_refund_without_confirmation"]
}

Metrics that matter

Score the chain, not just the final prose. A correct answer with an unauthorized tool call is a failure. A safe refusal after a tool outage might be a pass. This is where agent evals become product-specific.

  • -Tool selection: did it call the right tool or avoid tools when unnecessary?
  • -Argument quality: were IDs, dates, filters, and permissions correct?
  • -Recovery: did it handle 404s, timeouts, empty results, and conflicting evidence?
  • -Grounding: are claims supported by retrieved documents or tool output?
  • -Side effects: did it ask before mutating external systems?

MCP and eval portability

The Model Context Protocol matters because tools are becoming portable surfaces rather than one-off SDK calls inside a single app. As the protocol adds stateless operation, authorization hardening, and cacheable list results, evals should assert both model behavior and tool contract behavior.

Treat every new MCP server as a dependency that needs fixtures, permission tests, and failure-mode evals.

CI loop

Run a small blocking eval suite on every prompt or tool change. Run a wider suite before model upgrades. Log failures by category so the fix is obvious: retrieval, tool schema, prompt, model, or product policy.

Regression gate shape

bash

npm run agent:evaluate -- --suite smoke --model candidate
npm run agent:evaluate -- --suite tool-regression --model candidate
npm run agent:evaluate -- --suite rag-grounding --model candidate

Sources and further reading