Human review doesn’t make AI agents better or safer
Enterprise conversations about AI agents arrive, usually within minutes, at the same reassurance: don’t worry, a human stays in the loop. It calms security teams and procurement alike, and it likely appears in many vendor decks. Gartner expects more than 40 percent of agentic AI projects to be canceled by the end of 2027, with inadequate risk controls among the leading causes, so the anxiety underneath is well founded, even if the reassurance is not.
The boring work is the hard work
Enterprise content production runs on unglamorous problems. Product data arrives in hundreds of unpredictable variants. Artwork files carry claims, compliance text, and product specs that all need to be extracted and normalized into a reliable source of truth before a single image can be generated. Anyone who has shipped a packshot pipeline across forty markets knows the shape of this problem. Grip runs agents that reason through the material: reading the artwork, then reconciling naming conventions that different regional teams invented independently, and checking whether the copy matches what the brand’s rules actually allow.
None of this is a lookup table problem. The moment you map every known variant, a new one appears. Data inconsistency is open-ended by nature, and rules-based automation stalls on exactly this long tail. The only approach that holds up is an agent that understands the brand and the data well enough to reason through inputs it has never seen before.
So the agents make decisions. Real ones, thousands per day, about how a product name should be normalized or whether a claim belongs in a market. Which raises the question every enterprise asks next: shouldn’t a human be checking all of that?
A reviewer who approves ten thousand agent decisions a day supervises nothing
The standard industry answer is yes. Put a human in the loop, and the risk is handled. That answer deserves more scrutiny than it usually gets.
The writer Cory Doctorow draws a distinction between centaurs and reverse centaurs. A centaur is a person assisted by a machine: a human head on a tireless body, getting more done than they ever could alone. Flip it and you get the reverse centaur, a machine using a person as its assistant. The human works as a peripheral at the machine’s pace and, in Doctorow’s sharpest formulation, is paid mainly to take the blame for the machine’s mistakes.
Doctorow is a critic of the AI industry, and here the critique is on point. Many “human in the loop” deployments run the clear risk of producing reverse centaurs. A reviewer facing that volume of approvals supervises nothing. The stamp moves at machine speed, and the company knows it, because genuine review at human pace would erase the savings the automation was bought for.
The research here is older than the AI boom. Lisanne Bainbridge described the pattern back in 1983 in “Ironies of Automation”: when you automate a process, the human left over becomes a monitor, and monitoring is a task people are physiologically bad at. Vigilance collapses within about half an hour. Skills atrophy exactly when a rare intervention needs them. Modern AI makes it worse, because a language model’s errors are statistically the most probable ways of being wrong. So the reviewer is stuck solving a what’s-wrong-with-this-picture puzzle, thousands of times a day, against mistakes that probability has dressed up to look correct.
Deployed this way, the human in the loop is a compliance prop: everyone feels better while a bored person absorbs the liability. The very thing that appears to be reassuring turns out to be quite the dangerous thing.
Safety has to live in the infrastructure
If token review doesn’t make agents safe, what does? Constraining what the agent can do in the first place. Across serious agent deployments, that constraint is settling into multiple mechanisms:
- Isolated sandboxes. An agent can only touch what it is explicitly given. A misjudgment stays contained instead of cascading through production systems.
- Policy-based permissions. What an agent may do is defined in advance, not decided on the fly. The policy is a versioned artifact with a named owner, which means a bad outcome traces to a specific line someone can change, rather than to whichever operator was nearest when it happened.
- Privacy boundaries. Lightweight tasks run close to the data. Heavier reasoning goes to larger models through a controlled boundary, and sensitive data does not move unless policy allows it.
- A system of rules that can be accessed at run-time. Agents can (if allowed) access repositories of rules they should follow. If needed, other agents can be orchestrated to verify. The rules can be monitored and updated whenever, again by a named owner.
Take away the rubber-stamp human and someone still answers for the system, just in a place where answering can actually work. In policies people author and audit trails people read, rather than on the desk of a reviewer who never had a real chance of catching the error.
Workslop puts every human in the loop
The other path, handing everyone AI and hoping for the best, has data now as well. Researchers at BetterUp and Stanford coined the term “workslop” for AI output that looks like finished work but shifts the real effort downstream. In their September 2025 survey, 40 percent of workers had received it in the past month, and each instance took nearly two hours to untangle. Ungoverned AI doesn’t remove the human from the loop. It quietly makes everyone the loop.
A guardrailed agent inverts that. By the time a person sees its output, it has already cleared the sandbox, passed policy, and satisfied the brand rules. The agent finishes work instead of generating it.
Agents in the loop, humans in charge: what people should do instead
When agents absorb the normalization drudgery, the people in content production stay on the work that was always theirs. Art direction. Brand judgment, and deciding what good looks like for a product hero shot in a market they know. The taste decisions, in other words, which no amount of infrastructure makes automatic.
Which of the two the people around an AI system become, centaur or reverse centaur, is a design decision rather than a property of the technology. The infrastructure now exists to make that choice deliberately. What’s left is making it in favor of the humans.
References
- Kate Niederhoffer, Gabriella Rosen Kellerman, Angela Lee, Alex Liebscher, Kristina Rapuano, and Jeffrey T. Hancock (BetterUp Labs and Stanford Social Media Lab), “AI-Generated ‘Workslop’ Is Destroying Productivity,” Harvard Business Review, 22 September 2025.
- Lisanne Bainbridge, “Ironies of Automation,” Automatica, Vol. 19, No. 6, 1983.
About Grip
As INDG’s software branch, Grip is a visual content configuration engine powered by NVIDIA Omniverse that makes it possible for large enterprises to use AI at scale. It breaks existing content down into configurable modules, allowing brands to swap out any element, including products, talent, accessories, and branding assets, with complete control and accuracy. Grip integrates with existing workflows to automate product swaps and generate endless, hero-quality content variation, without disrupting established content production processes.



