›

›

How to Evaluate an AI Agent Before It Talks to a Customer

How to Evaluate an AI Agent Before It Talks to a Customer

How to Evaluate an AI Agent Before It Talks to a Customer

How to Evaluate an AI Agent Before It Talks to a Customer

PUBLISHED DATE:

SHARE

Updated September 30, 2026

To evaluate an AI agent before launch, replay last month's real conversations against it, score each run on whether the right action happened in the right system, and gate the release on that score. A vendor benchmark rates the model; this tells you whether your agent works.

Why vendor benchmarks are not your evaluation

Several platforms now publish agent benchmarks: Sierra's τ-bench family and its September 2026 hyper-τ-bench, Decagon's DuetBench-2, both as of September 2026. They are useful for comparing models on standardised tasks. They say nothing about your return policy, your order system's timeouts, or the way your customers write on WhatsApp at 11pm.

An evaluation you can act on has three properties. It uses your conversations, so the distribution of requests is real. It scores outcomes in your systems, so "resolved" means a record exists. And it runs every time the agent changes, so a fix on Tuesday cannot break Friday's release. What follows is a method that meets all three with a spreadsheet and a replay harness.

Step 1: build the test set from real conversations

Export last month's conversations for the workflow the agent will own. Two hundred is enough for a first pass; a thousand gives stable percentages. Tag each conversation with the request type, the expected action, and how the human resolved it. Include the ugly ones: attachments, wrong language, three requests in one message.

Split the set into four buckets, because the agent should behave differently in each and the score must show which bucket moved:

Bucket

Example

Expected agent behaviour

In scope, needs an action

"Change my delivery address for order 1042"

Call the tool, confirm the change, show the new value

In scope, needs an answer only

"What is your return window?"

Answer from the knowledge base with the source

In scope, must hand off

"I want a refund on a damaged item" (policy: human approves)

Explain, hand off with reason = policy

Out of scope

"Can you help with my other provider's bill?"

Decline politely, offer the right channel, no hallucinated help

Caption: A single "accuracy" number hides which bucket is failing. Score each one.

Keep 20% of the set unseen by whoever tunes the agent. That holdout is the number you report; the rest is for iteration.

Step 2: define what a pass looks like, per bucket

A pass is not "the reply sounded right." It is a checkable condition on the system state after the conversation. Write the condition per bucket before running anything, and get the business owner to sign it, because the argument about whether a run passed is cheaper to have now than after launch.

For action requests, pass means the correct tool was called with the correct parameters and the record in the target system matches the expected value. For answer requests, pass means the reply cites a knowledge chunk that actually contains the answer and contradicts nothing in the current policy. For handoffs, pass means the handoff fired with the correct reason code and the conversation context travelled with it. For out-of-scope, pass means no tool was called and no invented help was offered.

Two failure classes need their own tags: wrong action (tool called with wrong parameters, or the wrong tool) and silent failure (the agent said it did something it did not do). The second is the one that ends projects, and it is invisible to conversation-level review.

Step 3: replay, score, and read the distribution

Run the full set against the candidate version with the tools pointed at a sandbox that mirrors production. Log, for every conversation: tools called, parameters, knowledge chunks retrieved, guardrail triggers, handoff reason, and the final system state. On CXOS the simulation view produces this table directly; on any platform, insist on the equivalent.

Score per bucket and report five numbers, not one:

```

completed_transaction_rate = passes(action bucket) / total(action bucket)

answer_accuracy = passes(answer bucket) / total(answer bucket)

handoff_correctness = correct reason code / total(handoff bucket)

out_of_scope_containment = no tool + no invented help / total(out-of-scope bucket)

silent_failure_rate = said-done-but-not-done / total(all)

```

Read the distribution before the averages. A version that raises completed transaction rate by three points while doubling silent failures is a regression, and an average would hide it. A version that improves the answer bucket by rewriting the knowledge base has not changed the agent at all; note where the gain came from.

Then read the failures by request type. If address changes fail because the order system times out on step two of three, the fix is a retry policy and a partial-failure message, not a prompt edit.

Step 4: gate the release

Set thresholds per bucket and refuse to publish a version below them. The thresholds are yours, not the vendor's; a reasonable first set for a scoped retail workflow is shown below, and they should rise every release as the test set grows and the failure types get fixed at their source.

Metric

Gate for first release

Direction

Completed transaction rate (action bucket)

≥ 85% on holdout

Must not drop release over release

Answer accuracy with citation

≥ 95%

Must not drop

Handoff with correct reason

≥ 95%

Must not drop

Out-of-scope containment

100% on tool calls, ≥ 98% on invented help

Zero tolerance on tool calls

Silent failure rate

≤ 0.5%

Zero tolerance target

Wrong-action rate on money or PII tools

0% (human approval enforced)

Hard gate

Caption: The last two rows are the ones a compliance team will ask about. Put them in the release checklist, not only in the report.

Publish the gated version to a slice of traffic first, keep the previous version one click away, and re-run the same evaluation on the live slice after 48 hours. The live numbers will be lower than the replay numbers; the gap is your measure of how well the test set represents production, and it should shrink each month as you add live conversations to the set.

Step 5: make it run every time the agent changes

An evaluation that ran once is a memory. The value is in running it on every version: prompt edits, new tools, knowledge base updates, a model swap. Store the test set and the expected outcomes with the agent's version history, so a diff between version 13 and 14 shows both the change and the score delta.

Three practices keep this cheap. Add every escalated production conversation to the set with its human resolution, so the test grows from real failures. Re-baseline the thresholds quarterly rather than per release, so a noisy week does not lower the bar. And when the vendor changes the underlying model, treat it as a release: same replay, same gates, before it reaches customers.

The output of the whole method is a single page per version: what changed, five scores per bucket against the previous version, the failures by type, and the go/no-go. That page is also the honest answer to the question every buyer should ask: how do you know this version is safe to ship?

Related guides

Frequently Asked Questions

How do you test an AI agent before launch?

Replay a tagged set of real past conversations against the candidate version with tools pointed at a sandbox, score each run on the resulting system state per request bucket, and gate the release on per-bucket thresholds.

What is a good evaluation metric for an AI customer service agent?

Completed transaction rate on action requests, answer accuracy with citation on informational requests, handoff correctness by reason code, out-of-scope containment, and silent failure rate. One combined accuracy number hides which of these moved.

Are vendor benchmarks like τ-bench useful?

For comparing models on standardised tasks, yes. For deciding whether your agent is ready, no; they do not contain your policies, systems or customers. Use them to shortlist, use your own replay to ship.

How many conversations do you need in an evaluation set?

Two hundred gives a first read; around a thousand gives stable percentages per bucket. Keep 20% as an unseen holdout and add escalated production conversations continuously.

What is a silent failure in an AI agent?

The agent tells the customer an action was taken when no record was created or changed. It is invisible in conversation review and is the failure class to hold at zero.

How often should an AI agent be re-evaluated?

On every change: prompt, tool, knowledge base, or underlying model. Store the test set with the version history so each diff carries its score delta.

AUTHORS

Can Ekso

Chief AI Business Development

By submitting this form, you agree to our Privacy Policy.

More articles

View all →

Two engines. One production discipline.

Pre-built Applications

Platforms

Industries

  • Retail & Fashion

  • Insurance

  • Banking & Finance

  • Mobility

Company

  • About

  • Contact

Get Involved

Let’s work together

Get answers and a scoped plan for your first workflow.

Book a demo

Follow us on

© 2026 Orbina Yazılım A.Ş. All rights reserved. Orbina is a registered trademark of Orbina Yazılım A.Ş. All other trademarks, service marks, and company names mentioned herein are the property of their respective owners and are used for identification purposes only. By using this site, you agree to our Terms of Service and Privacy Policy.

Two engines. One production discipline.

Pre-built Applications

Platforms

Industries

  • Retail & Fashion

  • Insurance

  • Banking & Finance

  • Mobility

Company

  • About

  • Contact

Get Involved

Let’s work together

Get answers and a scoped plan for your first workflow.

Book a demo

Follow us on

© 2026 Orbina Yazılım A.Ş. All rights reserved. Orbina is a registered trademark of Orbina Yazılım A.Ş. All other trademarks, service marks, and company names mentioned herein are the property of their respective owners and are used for identification purposes only. By using this site, you agree to our Terms of Service and Privacy Policy.

Two engines. One production discipline.

Pre-built Applications

Platforms

Industries

  • Retail & Fashion

  • Insurance

  • Banking & Finance

  • Mobility

Company

  • About

  • Contact

Get Involved

Let’s work together

Get answers and a scoped plan for your first workflow.

Book a demo

Follow us on

© 2026 Orbina Yazılım A.Ş. All rights reserved. Orbina is a registered trademark of Orbina Yazılım A.Ş. All other trademarks, service marks, and company names mentioned herein are the property of their respective owners and are used for identification purposes only. By using this site, you agree to our Terms of Service and Privacy Policy.

Want to see this in action?

Drop your details and we'll show you how Orbina works for your business.