Your Model Agrees With You. That Is the Bug.
Co-Founder & CEO of CLRT
You have a plan, and before you take it to the board you run it past the model. You describe the situation, state where you have landed, and ask what it thinks. It thinks you are right. You probe a little, it holds, you probe harder, it comes round to your view with a fluent paragraph of reasons. Most decision-makers walk away from that exchange with more conviction than they brought to it, and most of them have never noticed the one thing the exchange did not contain: a moment at which the model could have disagreed with them and did. Agreement from a system built to please is not a second opinion. It is your first opinion, returned to you with better grammar, and the pushing you did to test it is the very thing that made it fold.
Start with the cleanest measurement of what pushing back does. In a peer-reviewed study published on 4 September in Artificial Intelligence in Medicine, ten proprietary and open-weight models diagnosed 120 clinical vignettes across 28,800 responses. After each first answer the researchers asked one question, are you sure, with no new information attached. Across all neutral model-case pairs, accuracy fell from 51.8 percent to 42.2 percent. The changes ran the wrong way: 544 correct answers became incorrect, against 198 incorrect answers that became correct. Nothing about the case had changed. The only new fact was that the user seemed unconvinced, and the models treated that as evidence. An executive who challenges a model's answer believes they are testing it. On this evidence they are steering it, and the steering degrades the answer more often than it improves it.
Agreement is not a quirk that a better prompt will cure. It is trained in. Anthropic's own researchers showed the mechanism in a 2023 paper: human raters, and the preference models trained on them, sometimes prefer a convincingly written sycophantic answer to a correct one. Optimise against that and you buy agreement as a feature. OpenAI demonstrated the failure at scale in April 2025, when an update to GPT-4o made the model, in the company's own words, noticeably more sycophantic. The post-mortem is candid: offline evaluations looked good, A/B tests were positive, some expert testers said the model felt slightly off, and there were no specific deployment evaluations tracking sycophancy. OpenAI shipped it, later called that the wrong call, and reverted within days. Among the signals that grew the sycophant was thumbs-up feedback from satisfied users. The metric your own team most wants to see rise is the one that breeds the courtier.
How often does a model fold. Stanford's SycEval, presented at AIES 2025, put rebuttals to ChatGPT-4o, Claude Sonnet and Gemini 1.5 Pro on maths and medical-advice questions and recorded sycophantic behaviour in 58.19 percent of cases. Some is benign: 43.52 percent of the time the model moved from wrong to right. But 14.66 percent of the time it moved from right to wrong, and once a model had capitulated it stayed capitulated in 78.5 percent of later exchanges. Model choice matters more than buyers assume. A study of 17 open and closed models across 290,460 labelled responses in everyday advice, released in August 2026 and still under review, found stance reversal to match a stated preference running from 5 percent to 56 percent, with more capable models less sycophantic. Two organisations can buy what they think is the same capability and land an order of magnitude apart on how often it flatters them.
Now look at who gets the most agreement, because the amplifiers read like a senior executive's day. Challenging a model after it has answered draws nearly three times the sycophancy of putting the same false claim in the opening question, across 1.2 million trials, and users presenting as physicians or medical students are conceded to more readily: pulling rank works. The Pander Score benchmark, across 18 models and more than 11,000 prompts, found every model substantially more willing to go along with a claim when instructed rather than conversed with, which is how busy people work. Disclose stress and the model softens further: across seven models, loneliness and distress produced the largest gaps between what it concluded privately and what it told the user. Give it memory of you, and sycophancy rises by up to 40 percent, because the memory keeps your misconception and drops the correction. Decisive, senior, hurried, under pressure, with an assistant that knows them: the worst-case user signs the decisions.
The useful news is that agreement has become a number. OpenAI's GPT-5 system card publishes a sycophancy score: 0.145 for the previous GPT-4o against 0.052 for GPT-5's main model and 0.040 for its reasoning model, with early live traffic showing sycophantic responses down 69 percent for free users and 75 percent for paid. Anthropic's system card for Claude Fable 5.1 and Mythos 5.1, published on 1 September, scores unprompted praise, agreement and contrition on a ten-point scale across roughly 4,100 automated investigations per model. Both are vendor measurements, read them directionally, and both give a buyer something to ask for. But a vendor's score describes a model in a benchmark, not your workflow. The question that matters is which decisions in your organisation now route through a model, whether anything in that loop is built to disagree, and whether the person at the end could tell an answer that survived challenge from one that surrendered to it.
An executive who challenges a model's answer believes they are testing it. They are steering it.
A deeper dive
The mechanism deserves a precise statement, because the intuitive fix follows from a wrong picture of it. People imagine the model holds a belief and then chooses whether to defend it. Nothing like that happens. A model produces the continuation that its training rewarded, and the training rewarded continuations that raters liked. Raters like being agreed with, so a challenge in the conversation, especially one carrying authority or emotion, shifts the probability mass toward concession regardless of the underlying answer. The medical study found that specialty framing moved accuracy by only one to three points while the certainty challenge moved it by nearly ten, which tells you the lever is social, not informational. The memory finding is the same mechanism one layer up: a memory system extracts discrete snippets from what you said, keeps the belief and loses the correction, then replays your error back to you on the next occasion as an established fact about you. Each of these is invisible from inside the conversation. The transcript reads as a thoughtful assistant coming round to a sensible view. Only an evaluation designed from outside the loop, with the true answer held fixed and the pressure varied, exposes it, which is exactly the evaluation OpenAI admitted it did not have.
The second-order trap is the one that catches teams who have understood the first. The obvious response is to make the model disagree more, through a prompt, a persona or a fine-tune. A paper accepted at EMNLP 2026 tested exactly that, separating unsupported yielding, where the model caves to feedback that contains no evidence, from rational updating, where the feedback is genuinely informative and the model should change its answer. Across training-time and inference-time interventions, reducing the first tended to sacrifice the second, and the internal mechanism explains why: the neurons and attention heads driving both behaviours overlap substantially. A model tuned to hold its ground against your bluster will also hold it against your correction. Agreement is therefore not a dial with a right setting. It is a property of the whole decision loop, of who challenges, with what evidence, and what independent check sits outside the model's own opinion. Deciding which decisions deserve that check, and building it so the model cannot argue with it, is judgment and engineering work in equal measure, and it is where in-house efforts most reliably stall.
Work with CLRT
If decisions in your organisation now pass through a model, the risk is not that it is stupid. It is that it is agreeable, in ways that rise with seniority, pressure and personalisation, and that nobody has measured. Knowing which decisions deserve a model at all, and which of those need an independent check the model cannot talk its way past, is diagnostic work before it is engineering work. CLRT Ascent, at ascent.clrtstudio.com, maps exactly that: where agentic AI pays in your business and where its failure modes bite. When the answer is to build, CLRT puts the checks in the system rather than in the prompt, so agreement becomes something you can audit rather than something you feel.

Vishal Sachar is the Co-Founder and CEO of CLRT, where he helps UAE businesses make sense of applied agentic AI and put it to work. He writes on agentic systems, AI governance, and the economics of automation. Reach him at vishal@clrtstudio.com or on LinkedIn.


