Back to Blog

Operations

Your AI Quality Score Needs a Quality Check

Automated AI evaluation needs calibration, visible coverage, and an owned response. A passing score should earn trust before it guides operations.

Revival Group•10/7/2026•5 min read

An AI assistant answers a customer question. The application stays online, the response arrives quickly, and an automated evaluator gives it a passing score. A service manager later discovers that the answer promised something the business cannot deliver.

In this hypothetical case, the team needs to investigate two decisions: why the assistant made the promise and why the evaluator accepted it. Adding another model to assess the first can help find problems, but that assessment needs its own operating standard.

On October 6, New Relic announced AI Evaluation, a planned capability that attaches automated quality assessments to application traces. The company describes an asynchronous service that uses a language model as a judge, alongside checks on sampled inputs and outputs. Public preview is scheduled for November. These are announced capabilities, not a generally available product we have independently evaluated.

The practical opportunity is connecting a questionable answer to the surrounding workflow so a team can investigate it. Our recommendation is to treat automated scores as signals that must be calibrated against business judgments before they drive operating decisions.

Define the failure before choosing the score

A single label such as response quality can conceal several different requirements. An answer might accurately summarize a document while applying it to the wrong customer. It might be relevant and polite while making an unauthorized commitment.

For a support workflow, start with the decisions that matter to the business. Can the assistant state the customer's eligibility correctly? Does it distinguish an approved refund from a request awaiting approval? Does it recognize when the available evidence cannot support an answer?

Write those requirements in terms a qualified reviewer can apply to a specific case. Include examples of acceptable answers and failures that would require intervention. The goal is agreement about what constitutes a consequential error before anyone selects a scoring threshold.

The business owner should approve that standard. Engineering can implement the evaluation, but it should not have to infer the company's service commitments from a dashboard metric.

Check the judge against reviewed examples

Build a small, representative set of cases whose outcomes have been reviewed by people who understand the workflow. Include routine requests, difficult exceptions, and answers that sound plausible but violate an important rule. Use authorized material and remove unnecessary sensitive information.

Run the automated evaluator against those cases and inspect its disagreements with the reviewers. A false alarm consumes review time. A missed failure can leave a customer exposed to an incorrect answer. Those consequences are different, so report them separately rather than combining everything into one agreement rate.

Also examine whether the evaluator had enough evidence to make its judgment. If it sees only the final response, it may be unable to determine whether a promise was authorized. If it sees a source document but not the relevant customer status, its assessment may be incomplete.

As we discussed in cross-system source coverage, a supported statement can still be insufficient for the decision at hand. The same problem applies to the system judging that statement.

Monitoring after the fact cannot substitute for approval

Timing determines what a control can accomplish. An asynchronous review can flag a response after the application has delivered it. That can support investigation and correction, but it cannot retroactively prevent the customer from relying on the message.

For consequential actions, keep the required approval or policy check in the execution path. Sending a refund, changing a booking, or making a binding service commitment should follow the workflow's authorization rules before the action occurs.

Use retrospective scoring to identify patterns, select cases for review, and improve future behavior. Define what happens when a concerning score arrives after completion. Someone may need to inspect the transaction, correct a customer record, or contact the responsible service team.

These are separate responsibilities. The person maintaining the evaluator does not automatically own every customer consequence it identifies.

Make sampling visible

When only some interactions are evaluated, an absence of alerts says little about the interactions that were never checked. A dashboard should distinguish eligible traffic, evaluated traffic, and cases awaiting evaluation.

Choose sampling with the workflow's risks in mind. Routine answers can be sampled differently from unusual exceptions or steps near an approval boundary. Document the approach so a falling alert count is not mistaken for an improvement when evaluation coverage has simply declined.

There is a cost tradeoff. More coverage can increase model usage, storage, and review work. Lower coverage can leave important failures unseen. Set the balance explicitly and inspect the queue of flagged cases as well as the number of scores produced.

The exception path should carry the relevant evidence and assign a next action. A growing collection of alerts without a responsible reviewer is unfinished work, even if the monitoring system is functioning.

Version the evaluator as part of the workflow

Changing the judge model, its instructions, or the evidence supplied to it can change the meaning of a score. A quality trend becomes hard to interpret if the measuring system changed halfway through the period.

Record the evaluator version alongside the workflow version and the result. Before adopting a revised evaluator, run it against the same reviewed cases and compare disagreements. Investigate whether apparent improvement reflects better business outcomes or merely a more permissive judge.

Keep enough information to reconstruct significant decisions, subject to appropriate access and retention controls. Evaluation records can contain the same sensitive material as the original workflow, so collection should be deliberate.

For a first implementation, choose one workflow, name its business reviewer, and agree on the errors that matter. Validate the evaluator, define coverage, and connect alerts to an owned response before relying on the score in production reporting.

If you're adding automated evaluation to an AI workflow, talk with Revival Group about making the evidence, review process, and operating response work together.

Enjoyed this article?

Explore more Revival Group perspectives on AI operating systems and operational transformation.