All insights
Fundamentals6 min read

The Reviewer Fails on Schedule

Vishal Sachar

Co-Founder & CEO of CLRT

Somewhere in your organisation one sentence does the work of an entire AI governance programme: a human reviews every output. It appears in the board paper, the vendor contract and the policy the compliance team circulated, and it is what permits an agent to draft the contract, triage the claim or write the code. The sentence carries a hidden assumption, which is that the reviewer is reliable in proportion to how much rides on the review. The evidence says the opposite. The human checker fails in a pattern that is measurable, repeatable and, for anyone who has read the literature, predictable to within a few percentage points. It catches the errors a machine would catch and passes the errors that matter. It approves more with each month of exposure. And it does this whether the reviewer is a junior annotator or a consultant radiologist.

31%1
Share of conceptual errors in AI pre-annotations that 2,784 human reviewers corrected, against 82 percent of simple digit swaps, in the Beck et al. randomised experiment
Beck et al., arXiv, 2025
36.8%2
Approval rate for AI-agent code in the same 400 reviewers' late reviews, up from 30.1 percent seven months earlier, per Yu et al.
Yu et al., arXiv, 2026
45.5%3
Accuracy of very experienced radiologists when the AI suggestion was wrong, against 82.3 percent when it was right, in the Radiology mammography study by Dratsch et al.
Radiology, 2023

Start with the cleanest measurement of what a reviewer actually catches. In September 2025, researchers at LMU Munich and the University of Maryland posted a preprint of a randomised experiment in which 2,784 people checked readings of corporate emissions tables presented as AI-generated and seeded with planted mistakes. The reviewers corrected 82 percent of the swapped digits. They corrected roughly 77 percent of cases where the right number sat in the wrong cell. When spotting the error required understanding the reporting rules, the catch rate fell to 31 percent. On one further table the AI was right, but confirming it needed the distinction between market-based and location-based Scope 2 emissions, and only 21 percent of reviewers handled it correctly; most rejected a correct answer. The reviewer is reliable exactly where the machine is already good, on the surface, and unreliable where the machine is weakest, on meaning, in both directions. That is the inverse of what the sign-off is supposed to provide. The task was annotation, not a boardroom decision, but the shape of the curve recurs everywhere anyone has measured it.

FIG. 01The share of planted errors that 2,784 human reviewers corrected, by error type, and one further table on which the AI was right. The 77 percent is approximate in the source. Source: Beck, Eckman, Kern and Kreuter, Bias in the Loop, arXiv, September 2025, a preprint.
01An old finding

The shape was known long before language models. Parasuraman and Manzey's 2010 review in Human Factors compressed two decades of automation research into one finding: automation bias occurs in both naive and expert participants, cannot be prevented by training or instructions, and affects teams as well as individuals. A 2012 systematic review in the Journal of the American Medical Informatics Association put a base rate on it. Across four clinical studies, between 6 and 11 percent of decisions that were correct before the decision-support system spoke became wrong after it did, and pooled, erroneous advice was 26 percent more likely to be followed when a system delivered it. A later review in the same journal found the bias even in single tasks, whenever verifying the aid is cognitively expensive. The harder an output is to check, the more the checker defers. Agent output, long, fluent and finished-looking, is the hardest class of output a reviewer has ever been handed.

02The stamp forms

The checker's second property is that it degrades with exposure, and this has now been measured in production. A June 2026 preprint tracked 400 repeat reviewers across 11,429 reviews of AI-agent code in open-source repositories over seven months. Between their early and late reviews, approval rose from 30.1 percent to 36.8 percent. The gap from the first to the tenth decile of experience was 14.5 points. Inline comments fell by 22 percent, and the time a request waited for review rose 3.5 times. The effect held after controlling for calendar time and did not appear for human-written code, where approval fell. The authors read this as most consistent with reflexive habituation under a growing workload rather than trust calibration on its own. Nobody decided to stop reading. The queue grew, the output looked fine, and the stamp came down a little faster each month.

FIG. 02Early versus late reviews by the same 400 reviewers of AI-agent code over seven months, each measure indexed to its own early value. Source: Yu et al., Habituation at the Gate, arXiv, June 2026, a preprint on open-source repositories.

The instinctive response is to put a better reviewer on it, and here the evidence is least comfortable. In a 2023 study in Radiology, 27 radiologists read 40 test mammograms with a purported AI suggestion attached, twelve of which were deliberately wrong. With correct suggestions, accuracy was around 80 percent at every experience level. With wrong ones, inexperienced readers fell to 19.8 percent, moderately experienced readers to 24.8 percent, and the very experienced to 45.5 percent. Seniority bought back some accuracy and still lost more than a third. Training does no better. A 2025 randomised trial in Pakistan, a preprint, recruited 44 physicians who had completed a 20-hour AI-literacy course and let all of them consult a model; for half, its advice carried planted errors in three of six cases. The physicians whose model was seeded with errors scored 14 points lower after adjustment, 73.3 percent against 84.9 for those given error-free advice. The standard corporate remedy, an experienced person who has been on the course, is the configuration these studies tested.

FIG. 03Radiologists' accuracy with a correct versus a wrong AI suggestion by experience level (Dratsch et al., Radiology, 2023), with AI-trained physicians from a separate randomised trial shown on their own scale (Qazi et al., medRxiv, 2025, preprint).
03The specification

Put the findings together and the checker has a specification. It catches surface errors and passes errors of meaning. Its approval rate drifts upward with exposure while its inspection effort drifts down. Its accuracy collapses in proportion to how convincing the wrong answer is; neither seniority nor training restores it. It also responds to friction in the wrong direction: in the emissions experiment, requiring reviewers to type a correction for every flagged error produced fewer corrections and more acceptance of wrong suggestions, and attitude towards AI predicted accuracy better than any demographic; the sceptics caught more. A policy that says a human reviews every output has therefore made a claim about a component whose failure curve is published. The question that decides whether an agent is safe to run is not whether a person is in the loop but which errors that person can physically detect, at what volume, for how long, and what catches the rest.

The reviewer is reliable exactly where the machine is already good, and unreliable exactly where it is not.

A deeper dive

The mechanism is attentional, not moral, which is why exhortation does not work. Reviewing is a cost-benefit decision made continuously and mostly below awareness. A 2023 series of five studies with 731 participants showed that people engage with verification only when engaging is cheaper than deferring, and that anything which raises the cost of checking, a harder task, a longer explanation, tips them towards the AI's answer. A bigger queue does the same thing by the same arithmetic. Agent output raises that cost by construction: it is long, it is fluent, it arrives with its reasoning attached, and it is usually right. Explanations make this worse. The CHI study that tested them found that showing the AI's reasoning increased the chance a reviewer accepted its recommendation regardless of whether it was correct. A reviewer facing a hundred outputs that were fine learns, correctly for ninety-nine of them, that reading closely is wasted effort. The hundred-and-first is the one that reaches the customer, and by then the reviewer's behaviour has been trained by the ninety-nine. A recent test of a hiring pipeline made the same point in miniature: the human checkpoint removed every flagrant fabrication and let the subtle ones through roughly half the time.

The second-order trap is that the fixes which work are the ones the organisation will remove. The one intervention with consistent evidence is cognitive forcing: making the reviewer commit to an answer before the AI's is shown, or delaying the suggestion. In a 199-person experiment it significantly reduced over-reliance, and participants gave those designs the least favourable ratings of any tested. A control that works feels like obstruction, so it is the first thing a team under delivery pressure asks to switch off. The fashionable alternative, letting the model abstain when unsure, moves the error rather than removing it. Among 259 clinicians, wrong AI cut accuracy from 66 to 56 percent, abstention recovered most of it, and the abstaining cases produced 18 percent more missed diagnoses and 35 percent more missed treatments than unaided decisions. What survives is a different design principle: the human check is reserved for the errors a human can actually detect, and everything else is caught by verification that runs on measurements the model does not control and that does not habituate. Deciding where that line falls, workflow by workflow, is the judgement most governance programmes have never made.

Work with CLRT

CLRT builds verification on the assumption the evidence supports: that the reviewer will fail on schedule, in known ways, and that the system has to catch what the person cannot. That begins with a judgement about where a human check is real and where it is theatre, workflow by workflow, and it ends with the measurements, sampling and stop conditions that do not habituate. If your AI governance currently rests on someone signing it off, the CLRT Ascent diagnostic at ascent.clrtstudio.com is where we find out what that signature is actually worth. Bring us the workflow, and we will show you which errors your reviewer can catch and which ones are already getting through.

Vishal Sachar

Vishal Sachar is the Co-Founder and CEO of CLRT, where he helps UAE businesses make sense of applied agentic AI and put it to work. He writes on agentic systems, AI governance, and the economics of automation. Reach him at vishal@clrtstudio.com or on LinkedIn.

Start here

Skip the reading. See where your leverage leaks.

Ascent is our free diagnostic. Ten minutes, and you have the one workflow worth building first.