In the Lab/Research

The Speed Trap: Why AI's Biggest Strength Is Creating Its Biggest Quality Problem

The LitmusLab Team

The LitmusLab Team

July 28, 2026·5 min read

The Speed Trap: Why AI's Biggest Strength Is Creating Its Biggest Quality Problem

Google's 2026 AI Agent Trends report dropped a number worth celebrating: 88% of agentic AI early adopters are seeing positive ROI. 52% of executives in gen AI-using organizations already have AI agents in production.

That's remarkable adoption. That's real business value.

But buried in the same report is a quieter observation that nobody is putting on a slide deck:

"Verify quality — Act as the final checkpoint for quality, accuracy, and tone."

That's listed as a core human responsibility in the new agentic workflow. It's framed as straightforward. Expected. Obvious, even.

And it raises a question the report doesn't answer:

How?


Before you balk at the stat

Before you balk at that 52% figure, it's worth understanding what's likely driving it. Most of that adoption isn't autonomous multi-step AI reasoning across enterprise systems. It's Copilot Studio agents embedded in Microsoft 365, Agentforce workflows in Salesforce, Gemini agents in Google Workspace. Real deployments, real value, but a long way from the fully autonomous agentic systems the report describes.

The bar for "agent in production" is lower than the headline implies.

And yet, even at that lower bar, the verification problem is real. Which means at true agentic scale, it's only harder.


The tension nobody is talking about

The ROI case for AI is built almost entirely on speed. Agents complete in minutes what used to take days. Teams ship faster. Costs drop. Output scales. The business case is real.

But speed and meaningful oversight are in direct tension.

When a consultant delivered a report, you had time to review it. The pace of the work created natural checkpoints. When AI delivers the same output in seconds (and does it dozens or hundreds of times a day across your organization), the pressure to accept it and move on is enormous. The very thing that makes AI valuable is what makes verification feel like an obstacle.

To be clear, this problem has existed in other forms before. Humans have always been responsible for the work that leaves their hands, regardless of where it came from. What's new is the volume, which outpaces any team's capacity to review; the speed, which eliminates the natural checkpoints that used to exist; and the AI response confidence that makes it harder for teams to detect what's wrong using traditional instincts.

Traditional tools fail visibly: a broken formula, a missing data source, a consultant who says "I'm not sure." AI fails invisibly, and it does so continuously, at the same pace that makes it valuable in the first place. The instincts people have built over careers were calibrated for visible failures they could see and slow down for. They don't hold when the volume never stops and the output always sounds certain.


The structural problem

Here's what makes this more than a training issue or a culture issue: verification doesn't scale with production. People do.

At some point (and many organizations are already past this point), the volume of AI output outpaces any human's ability to review it meaningfully. You can tell your team to verify what AI generates. You can build checklists and policies and guidelines. But if the system produces output faster than people can evaluate it, the gap between what ships and what was actually reviewed will widen regardless of intent.

That 88% ROI figure measures efficiency, speed, and cost reduction. It does not measure accuracy. It does not measure quality. It does not measure what was missed.

Teams are faster. But are they better? Are they catching the errors they're generating at 10x speed?

Most organizations don't know. Because they aren't measuring it.


Flipping the incentives

The pace of AI adoption isn't reversing, and the productivity gains are too real to walk away from. The work is in building systems and incentives that make verification as natural as production, so quality is built into how work moves rather than added at the end.

A few places worth considering:

Make the cost of AI errors visible. Most organizations don't track what bad AI output actually costs: rework, customer trust, support tickets, and regulatory exposure. When that number is invisible, verification feels like friction. When it's on a dashboard, it feels like ROI.

Reward catch rate, not just ship rate. Speed is measured. Being the person who slows down to verify AI output looks like a bottleneck, not a quality contribution.

Worth sitting with the premise here: measuring catch rate means accepting that errors will happen. That's a different starting point from treating every error as something to be eliminated through better prompting or tighter guidelines. Organizations chasing zero errors tend to underinvest in verification because they're always one improvement away from solving the problem. Organizations that track catch rate have already moved past that assumption. They're building systems that reflect how AI actually behaves in production: generating output at high volume and high confidence, some fraction of which will be wrong.

If teams are measured on errors caught before production, rather than velocity alone, verification becomes a performance metric rather than a tax on productivity.

Tie AI output quality to customer metrics. NPS, churn, support volume. When AI quality is explicitly connected to the numbers leaders already care about, the business case for investing in verification infrastructure writes itself.

Make verification the path of least resistance. Right now, verification is extra work that competes with everything else on the list. If a quality gate is built into the workflow (something that runs automatically before output ships), it stops being optional and starts being just how things work.


The gap worth closing

Google is right that humans need to be the final checkpoint for quality, accuracy, and tone. The question was never whether that's true. It's whether the systems, incentives, and tools exist to make it actually happen at the speed AI now demands.

For most organizations right now, they don't.

The ROI of AI is real. The quality infrastructure to protect it is still catching up. Teams building with AI are doing genuinely hard work, and the gap they're running into isn't a failure of effort. It's a gap in the tooling.