Who Checks the Agent's Work?

||By Jen Spencer

When you hire a person to do a job, quality control is built into how you think about the role from day one. There is a ramp, there is a manager reviewing early work, there are spot checks, and there is a feedback loop when something goes wrong. Nobody questions whether this is necessary. It is just how you bring a new person into real work.

Then a company drops an AI agent into the same workflow and skips all of it. The agent goes live, it produces output at a speed no human could match, and the speed itself becomes the reassurance. The work is getting done, and the queue is moving, so it must be fine. The question of who is actually checking the agent's output, and how, tends to get answered much later, usually by a customer.

This is the gap in most AI-supported operations right now. Companies put a lot of thought into deploying the agent and almost none into governing it. They treat quality as a property the model either has or does not, rather than something you have to measure continuously, the same way you would with any team doing work that matters.

The trouble is that agents fail differently than people do. A new human employee who is unsure will usually hesitate, ask a question, or flag that they are out of their depth. An agent generates a fluent, confident answer whether or not it is right. The errors do not announce themselves. They look exactly like the correct outputs until someone with context reads them closely, which is why a drifting agent can run for weeks before anyone notices the quality has slipped. Research on large language models keeps confirming that fluency and accuracy are not the same thing, and confident wrong answers are the hardest kind to catch.

 

A Quiet Failure, In Slow Motion

Here is a composite, blended from conversations with operations leaders over the past year so that no single company is identifiable. The pattern shows up again and again.

A team automates a workflow that produces customer-facing output: drafting responses, resolving routine cases, and processing requests against a set of rules. The agent is good. For the clean, common cases, it is clearly better than the queue it replaced, faster and tireless. Leadership is pleased. The team that used to do the work is redeployed or reduced, and the agent runs.

What no one set up was a way to know whether the output was still good three months in. There was no sampling, no rubric for what a correct response looked like, and no one whose job was to read a slice of the agent's work each week and grade it. So when a subset of cases started going subtly wrong, because an upstream system changed, or because a new type of request started arriving that the agent handled badly, nothing caught it. The errors accumulated quietly. The first real signal was a frustrated customer escalating something that should never have left the building, and by then the same mistake had gone out hundreds of times.

The failure here was not the agent. The agent did exactly what it was built to do. The failure was that the company removed the quality control that a human team had been providing without realizing that is what those people had partly been doing. Experienced staff is not only producing the work. They are noticing when something is off, and that noticing is a function nobody had explicitly accounted for.

 

Quality Control is a System, Not a Vibe

The fix is to treat the agent's output the way a serious operation treats any output: with quality control that is designed in, not bolted on after a complaint. In practice that means a few concrete things.

You define what good looks like, specifically enough to grade against. A rubric, not a feeling. You sample the agent's work continuously and score it, the same way a contact center has always audited a percentage of calls, so you have a live read on quality rather than a quarterly surprise. You route the cases the agent is least sure about, or that are most expensive to get wrong, to a human before they go out, not after. And you feed what you learn back into the system, tightening the rules and the rubric as new failure modes show up.

None of this is exotic. It is the same statistical-quality discipline that manufacturing and customer service have used for decades, applied to a new kind of worker. The mistake companies make is assuming that because the agent is software, it inherits software's reliability. A deterministic system you can test once and trust. A generative one drifts with its inputs and needs to be watched the way you watch a process, not the way you trust a calculator.

 

Where Embedded QA Changes the Math

This is one of the less obvious advantages of running AI-supported work through a managed operation rather than standing it up from scratch in-house. A team whose entire business is delivering this work has to solve quality control to survive, so the QA discipline is built into the delivery model by design: defined rubrics, continuous sampling, a human reviewer in the loop, and the failure modes for a class of workflow understood from running it at scale. The right version puts the agent inside that quality system rather than letting it run unsupervised. You get more than the automation. You get the apparatus that tells you, every week, whether the automation is still doing its job, and a person is accountable for the answer. 

That accountability is the part that is hard to improvise. Someone has to own the quality number, be able to explain a miss, and have the authority to pull a workflow back to human handling when the agent is not ready for it. Speed is easy to buy. The discipline to verify the speed is producing good work is the harder thing, and it is what separates an AI deployment that holds up from one that quietly erodes trust until a customer forces the issue.

 

The Operator's Read

If you have an agent in production, or you are about to, the most useful question is about quality. How would you know if it started getting things wrong? If the honest answer is that you would find out when a customer complained, you do not yet have an AI-supported operation. You have an unsupervised one.

Define good before you go live. Sample and score the output on a schedule, not on a hunch. Keep a human reviewing the hard and high-stakes cases, and give that person the authority to act on what they find. The agent can carry an enormous amount of the work. Someone still has to be answerable for whether the work is any good.

 

Jen Spencer

About the Author:

Jen Spencer is the Chief Growth Officer at Booth, where she helps growth-stage companies build hybrid human-AI teams without compromising on quality, culture, or control. She's the author of Lead Anyway: How to Build the Career You Want When the Rules Weren't Written for You, a six-week workbook on reclaiming leadership in rooms that weren't built for you. Before Booth, Jen was the CEO of a 300-person digital agency. She advises and invests in early-stage SaaS, is a Pavilion member, and speaks on operations strategy in the AI era. Author. Operator. Recovering impostor.

Back to blog
08-FeaturedBlogPosts