AI
Manual review of AI agents is not a cost you can remove
Collibra and The Harris Poll report that 51% of organisations spend significant staff hours reviewing what agents produce. The figure gets read as waste. It is the last control that exists before risk moves from a wrong answer to a wrong action.

What happened
Collibra published a survey on 23 September 2026, run by The Harris Poll among data and AI decision-makers. Three figures were pulled out: 76% had hit critical roadblocks moving agents from pilot to production over the past year, 87% said their teams regularly re-verify whether the context an agent works from is correct, and 51% spend significant staff hours reviewing what the agent produced.
Why the figure gets read wrong
The figures were presented as a cost to bring down, a “hallucination tax” on enterprise AI. That reading assumes the review is a temporary weakness in the model, and that better models will remove it.
The reading misses what the review actually is. A result can only be reviewed for as long as the system still produces a result before it acts. The control sits in the gap between a proposal and an action. It has little to do with the quality of the model.
Governance teams carry growing manual workload, which is exactly why Maestro automates governance operations directly.
What it means if the risk moves
If risk moves from a wrong answer to a wrong action once software acts on its own, then two of the figures in the survey are the same figure seen twice. 76% stall on the way from pilot to production. 51% spend staff hours reviewing the result. What stops a pilot is usually that nobody will sign for an action nobody has seen.
Removing the review does not remove the cost. The cost moves from an hour of someone's time to an action already carried out in a system outside the agent. That bill arrives somewhere else and later, and it has no line in any budget.
What we do differently
The publishing agent on apps.no can carry out zero irreversible actions. The agent reads the whole article archive, picks the subject, writes in two languages and makes the cover image, and every single write lands as a draft. On 21 September 2026 the agent delivered two complete drafts that passed every machine check in the CMS. A person rejected both: one because the evidence was second-hand, the other because the figures belonged to a different company than the blog belongs to.
Neither of those two reasons has a field in any form. So Apps AS does not budget for removing the review. Apps AS budgets for making it cheap, and for keeping the list of irreversible actions short. What that requires of permissions, logging and escalation is set out in What changes when software acts on its own. The figures on where agentic AI actually pays off are in How Will Agentic AI Improve My Business?, and why the log is also the documentation regulators will ask for is in Agentic AI in a country with no AI law.
Count how many of your reviews change something
Do you have an agent running in production? Count how many of your reviews actually changed the result over the past month, and send the number to hello@apps.no. We will reply with which of the steps we would let the agent do on its own, and which we would leave alone.
Sources
- Collibra, Collibra Launches New Capabilities to Reduce the Hallucination Tax on Enterprise AI, press release 23 September 2026. prnewswire.com
- Apps AS, What changes when software acts on its own, read 25 September 2026. apps.no
About this article. Written by Espen Hareide, co-founder and partner at Apps. A take using the pillar A new reality with agentic AI as its lens. The figures from Collibra and The Harris Poll come from Collibra's own press release of 23 September 2026, read 25 September 2026. The survey's sample size and the countries respondents work in are not stated in the press release, and are not estimated here. The publishing agent example draws on the runs of 21 September 2026. No client figures are used in this article.
FAQ
- How many reviews need to change something before they are worth the time?
- There is no universal threshold, but the share is worth counting. If the share of reviews that actually change the result sits near zero over several weeks, and the action can be rolled back, that is the first place to let the agent act on its own. On actions that cannot be rolled back the share does not matter, because the review there is an approval rather than a quality check.
- Does this mean an agent can never act on its own?
- No. The point is to separate actions that can be rolled back from actions that cannot. Reversible actions can be left to the agent, because a mistake can be corrected. Irreversible actions need an approval however good the model is, because there is nothing to correct afterwards.
- What do we do when the volume is too large to review everything?
- Move to sampling on the reversible steps and keep full approval on the irreversible ones. A person assesses a random selection of what the agent did, and you measure outcomes rather than completion: the share of cases reopened, or the share of results corrected later. Operational monitoring will show green either way.
- Who should do the review?
- Someone who has done the work themselves long enough to recognise a mistake that looks plausible. Those skills are harder to hire than the ones they replace, and they do not belong to the cheapest person available. A review done by someone without that grounding only confirms that the fields have been filled in.
- Do the Collibra figures apply to companies outside the United States?
- We do not know. The press release states neither the sample size nor which countries the respondents work in, so the figures are better read as a direction than as a benchmark to compare yourself against. The number that matters for you is how many of your own reviews changed something.


