AI, Plain English · Post 009

Test the judge before you trust the score.

An AI support agent can pass every test when the evaluator misunderstands the real outcome. Calibrate the judge first, then test the agent.

Direct answer

A support agent score becomes useful only after domain experts agree on representative cases, an explicit rubric, and the failures the automated judge must catch.

A calibration bench connecting expert labels, a rubric, an automated judge, and a support agent release gate.
Ahmad Bukhari · Post 009
A calibration bench connecting expert labels, a rubric, an automated judge, and a support agent release gate.

By Ahmad Bukhari · Founder, Aixcel Solutions · Published 30 July 2026

Key takeaways

  • Test the evaluator against expert labels before using it to compare support agents.
  • Measure the final customer outcome and the quality of the interaction, not fluency alone.
  • Use separate gates for evaluator quality, agent quality, and live operational performance.
  • Keep policy exceptions, consequential account changes, and release ownership with named people.

What the research changes

A June 2026 paper accepted at KDD 2026 describes evaluation practices used across five production customer support agent deployments. Three operations analysts independently labelled cases, majority vote established the reference answer, and written rationales helped refine the evaluator rubric.

In one reported evaluator task, the majority baseline scored 77.78, a short manually written judge prompt scored 68.88, and an optimized judge prompt scored 88.89 on a held out set. The lesson is not that 88.89 is a universal target. The lesson is that the judge itself can be wrong enough to reverse a release decision.

Why one overall score can hide failure

A fluent answer can still violate policy, miss the requested outcome, use the wrong account state, or create a second contact. One average score compresses those failures into a number that looks precise while hiding what matters.

  • Outcome correctness: did the customer receive the right final result?
  • Policy compliance: did the agent stay inside the approved rules?
  • Tool correctness: did reads and writes match the case and account state?
  • Escalation quality: did the agent stop and hand over at the right moment?
  • Interaction quality: was the conversation clear, efficient, and respectful?

Calibrate the judge with expert labels

Start with a small set of real and costly cases. Ask several domain experts to label each case independently. Where they disagree, resolve the rule before automating the judgment. Their agreed labels and rationales become the reference set for testing the automated judge.

Track false passes as carefully as false failures. A false pass is dangerous because it tells the release owner that an unsafe or ineffective answer is acceptable. Refine the rubric until the judge catches the failures that experts consider material, then freeze the version used for a release decision.

Worked example: a subscription downgrade

A customer asks to downgrade immediately and avoid the next charge. A weak judge may reward a polite explanation even if the agent changes the wrong plan, misses the billing cutoff, or promises a refund outside policy.

The evaluator should inspect the requested outcome, the actual account change, policy compliance, the explanation given to the customer, and whether the case required a human decision. The customer outcome is the unit of evaluation. The wording is supporting evidence.

Connect offline tests to a small live release

The paper authors report that offline improvements correlated with online metrics, then describe small initial launches before broader rollout. Use that as a release pattern, not as a promise. Your traffic, policies, languages, systems, and failure costs will differ.

  • Judge gate: the evaluator agrees with experts on representative and costly cases.
  • Agent gate: the candidate meets thresholds for outcome, policy, tools, escalation, and interaction.
  • Live gate: a small release confirms that offline gains survive real customer behavior and system conditions.
  • Stop condition: named owners can pause the release when a material failure appears.

Opportunities, risks, and limitations

A calibrated evaluator makes faster iteration possible. Teams can compare prompts, tools, policies, and agent versions with evidence that is closer to expert judgment. It also creates a repeatable release record for quality review.

The method still depends on the quality of cases, expert agreement, and access to the true customer outcome. Policy drift can make an old rubric stale. A judge can also overfit to the reference set. Recheck it when workflows, policies, models, tools, or customer segments change.

Support data may contain sensitive customer information. Minimize collection, pseudonymize where practical, restrict access by role, document retention, and verify contractual and legal duties before using transcripts for evaluation.

Who should act now and who should wait

Act now if you have a stable support workflow, named policy owners, representative cases, access to final outcomes, and enough expert time to resolve disagreements. These conditions make evaluator calibration practical and useful.

Wait if the policy changes weekly, experts cannot agree on a correct result, the agent cannot observe whether its action succeeded, or no one owns live exceptions. Fix those operating conditions before treating an automated score as release evidence.

A practical 30, 60, and 90 day framework

Days 1 to 30: choose one support intent. Gather representative, costly, ambiguous, and policy sensitive cases. Have domain experts label outcomes independently, resolve disagreements, and write the first rubric.

Days 31 to 60: test the automated judge against the expert reference set. Review false passes, false failures, and disagreement by case type. Version the rubric and set separate thresholds for the judge and the agent.

Days 61 to 90: release the strongest agent to a small share of eligible traffic. Compare live outcomes with the offline prediction, inspect every material exception, and expand only when the named release owner accepts the evidence.

Questions decision makers ask.

Clear answers before a platform choice becomes an operational commitment.

01Why test the evaluator before the agent?

Because a weak evaluator can reward the wrong behavior or reject a good result. Its agreement with domain experts is part of the measurement system.

02How many experts are needed?

There is no universal number. The lead paper used three operations analysts and majority vote. Use enough independent expertise to expose disagreement and document how it is resolved.

03Should one score decide a launch?

No. Keep separate evidence for outcome correctness, policy, tool use, escalation, interaction quality, and live operational performance.

04When should the rubric be reviewed?

Review it when policies, workflows, customer segments, models, tools, or failure patterns change, and on a fixed schedule even when they appear stable.

Continue your evaluation.

Compare adjacent systems, inspect evidence, or see how Aixcel delivers the work.

Bring us the constraint. Leave with a clearer next move.

In 25 focused minutes, we will map where work or revenue is getting stuck, test whether AI is the right intervention, and identify the highest leverage first step.

Book a free systems audit