The Prompt Course Is the Wrong Purchase
Co-Founder & CEO of CLRT
The request arrives in the same words at most companies. Leadership has decided the workforce should be able to use AI, and the instrument chosen is a course: prompt engineering, AI literacy, a few hours per head, a certificate at the end, the whole company covered by the end of the quarter. It is easy to sign because it looks like every capability programme the organisation has bought before. The trouble is that the thing being bought has now been measured, item by item and head to head, and it does not do what the invoice says. The syllabus does not move accuracy. The certificate does not protect its holder from a wrong answer. And the one intervention that moved the number by twenty percent is not a prompting skill at all. It is something the company already owns and has never written down.
A typical prompt course teaches people to be polite to the model, give it a persona, ask it to think step by step, and sometimes offer it a tip or a threat. Wharton's prompting research group has tested each of these in turn, in four reports across 2025. Politeness sometimes helped and sometimes hurt. Chain-of-thought gave marginal gains at best on models that already reason. Threats and tips had no significant effect. Expert personas did not improve accuracy. One randomised experiment put the course itself on trial. Ma and colleagues, in ACM Transactions on Computer-Human Interaction, trained thirty novices, a small sample, either in conventional prompt engineering or in writing clear, complete requirements. Conventional training produced a 1 percent gain. Requirement training produced 20 percent, a gap automatic prompt optimisation could not close.
That second arm tells you what the twenty was made of. It was not a trick. It was the ability to say precisely what the business needs from an output: the audience, the exclusions, the format the next person can use, the check that decides whether it is done. That is domain knowledge, and the organisation already holds it. The field experiments agree. At Procter and Gamble, 776 professionals in a randomised trial received exactly one hour of prompt training; the gains, a 0.37 standard deviation lift in quality for individuals with AI and 16.4 percent less time, are attributed by the study to the tool inside a designed task; nobody measured the sixty minutes separately. Brynjolfsson, Li and Raymond gave 5,172 support agents the tool inside the workflow, not a course, and they resolved 15 percent more issues an hour, the largest gains landing on the least experienced. The lever in every case was the design of the work.
The literacy course fails on the other side of the transaction too, where a person decides whether to trust what the model said. In a randomised trial in Pakistan, published as a preprint, 44 physicians who had all completed a 20-hour AI-literacy course diagnosed six cases with a model's help. The arm given correct advice scored 84.9 percent. The arm whose advice carried planted errors in three of the six cases scored 73.3 percent, a 14-point fall the course had not prevented. Parasuraman and Manzey's 2010 review in Human Factors had already concluded that automation bias occurs in naive and expert participants alike and cannot be prevented by training or instructions. The evidence base is thinner than the brochures suggest: a systematic review published in August found, among nursing students, three quasi-experiments and no randomised trials of AI-literacy training, 443 students in all, measuring literacy scores and classroom outcomes rather than work output. The European Commission's guidance on the AI Act's literacy duty says there is no need for a certificate and no specific level is required.
We have watched the mis-specified purchase arrive twice in the last month. A professional-services firm of about seventy people asked for prompt-engineering training for all staff, online. We proposed two sessions of two hours, about a week apart. In the first, every participant chooses one recurring task from their own week and leaves with a working prompt for it, to run for real in the days between. The second opens on what broke and closes with a safe-use standard the whole company adopts. A regional leadership team at a multinational consumer group, six or seven leaders already using Claude daily on their own numbers, asked a sharper question: what to check and what to trust. The design they received routes their work by what is at stake, always verify, spot-check, let it run, in a four-hour session where verification is the only block that grew, 45 minutes against 30 in the original full day. Both designs move the value out of the room and into the week between.
The economics agree. Anthropic's meta-analysis of 56 randomised US studies of worker retraining, an analogy rather than a measurement of AI courses, found that a training slot costing about 13,000 dollars raised earnings by roughly 1,000 dollars a year; programmes roughly break even. The exception is the small set of sector programmes that partner with employers and place people directly into the work, which produce gains several times larger, though the source notes replication attempts have often failed. The honest contradiction is BCG's 2025 survey of more than 10,600 employees, vendor research and self-reported, which found regular use sharply higher among people with at least five hours of training and in-person coaching. The survey cannot separate the hours from the coaching that came with them. So the right purchase is not the course. It is the organisation's requirements written down, one recurring workflow per person, a checking norm keyed to stakes, and a week of live practice between two sittings. The course is the thinnest layer of that stack, and the only one the market sells.
A prompt is a sentence. A workflow with a check is a change in how work moves.
A deeper dive
The mechanism is that prompting skill is a depreciating asset while requirement skill is not. Every trick a course teaches was found against a particular generation of models, and larger models are less sensitive to phrasing than smaller ones, He and colleagues finding GPT-4 more robust to changes of format than a smaller model whose score swung by up to 40 percent, which is why the Wharton group keeps finding effects that are small, contingent and impossible to predict in advance for any given question. Whatever residual sensitivity remains is exactly the kind of thing that is automated: Battle and Gollapudi found an automated optimiser the most effective method they tested, ahead of hand-written prompts, and the requirement-training study found that this automation could not close the gap between the two arms, because the missing input was not phrasing but knowledge of what was needed. A course therefore sells the part of the problem that is shrinking and automatable, and leaves untouched the part that is durable and human: knowing what a good output is for this business, this audience, this signature. That knowledge does not transfer from a trainer. It has to be extracted from the people who hold it and written into the workflow, which is slower, less scalable, and the only version that survives the next model release.
The second-order trap is what the certificate does to the checker. An organisation that has trained everyone believes it has done something about risk, and the belief changes behaviour in the wrong direction. The trained employee is more confident, and the automation literature links a favourable attitude to AI with deferring to it; the physicians in the Pakistan trial had twenty hours of literacy behind them and still followed the planted errors. The literacy scorecard, meanwhile, measures the thing that is easiest to measure and least connected to outcomes, self-reported understanding. So the company ends up with a completed programme, a rising literacy score, a workforce that trusts the model more than it did, and no rule anywhere that says which outputs must be checked, by whom, before they carry a number or a name. The failure that follows is not a failed course. It is a wrong figure in a board paper with a trained person's signature on it, months after the certificates were issued, at which point the training budget looks in retrospect like the cheapest way the company found to stop thinking about the question.
Work with CLRT
This is the work CLRT does instead of the course. We take the two or three recurring workflows that carry a number or a signature, write down the requirement each one has never had, build the checking norm that decides what is verified and by whom, and run the whole thing live across two sittings with a week of real work in between, so the capability lands in the organisation and not in a certificate. If you want to know which workflows are worth that treatment before you commit to any of them, that judgment is what CLRT Ascent at ascent.clrtstudio.com is built to reach. Bring us the training request. We will show you what it was actually asking for.

Vishal Sachar is the Co-Founder and CEO of CLRT, where he helps UAE businesses make sense of applied agentic AI and put it to work. He writes on agentic systems, AI governance, and the economics of automation. Reach him at vishal@clrtstudio.com or on LinkedIn.


