The Divergence Is the Test
Co-Founder & CEO of CLRT
Most executives carry a quiet assumption about AI: the advantage belongs to whoever builds fastest. The model is the same for everyone and the data is what it is, so the race must be to the machine. On 27 August, in a small room by design at TheBlock in One Central, CLRT ran the final session of AI After Hours as a test of that assumption. Seven people, each registered as a competing consultancy, received the same data room for the same fictional company, had the same ninety minutes and the same model, and were asked to tell a board what was really wrong. The format was built so that the room would not agree, and the scoring was written to pay for the choice, the defence and the stated limits before it paid for the code.
Dubai Media Group is fictional, and says so in the footer of its own tender page and its own reveal page. It is an eighteen-year-old Dubai media agency of 31 people working in corporate film, event coverage, social content and podcasting. Revenue was AED 15.6 million two years ago, 14.2 million last year, and the first half of this year annualises below 13 million. Every client asks for the founder by name, and the board has stopped smiling. The data room behind it holds thirteen files, among them 210 invoices, 120 supplier payments, 41 proposals, 214 approval-log rows, a founder's calendar export, 34 staff rows, a purchased prospect list of 1,204 rows, exit interviews and a half-year board pack. All of it was generated from a single seed and passed through more than seventy automated consistency checks before anyone saw it, so that every number agreed with every other number and nobody could win by catching the fiction in a mistake.
Inside that consistent company sit six problems, each planted with its own trail of files. The lead drought: 41 proposals in eighteen months, six won, all six from referrals or existing clients, and not one from cold outreach. The bought list: 1,204 rows purchased for AED 9,000, blasted about 300 times for four replies, then abandoned, with roughly 140 genuinely good contacts buried under 180 duplicates and hundreds of junk rows. The founder as bottleneck: nine hours every Sunday hand-building the Monday client pack, and 214 quotes and edits personally approved last quarter. Management by one brain: three editor resignations in five months, and two project lists in the same board pack disagreeing on four of eleven live projects. The buried opportunity: a podcast studio at 32 percent utilisation on the firm's best margin, with eleven of 23 inbound inquiries never answered. And the quiet bleed: collection days worsening six straight quarters, from 52 to 71, while suppliers are paid at about fourteen. Every one of them is real in some company you know.
The format gave nobody time to fix all six, which was the point. Ninety minutes on an on-screen clock with a hard stop. At 45 minutes remaining, a staged email from the fictional managing director landed in every bidder's inbox. It added a board question, what should Dubai Media Group stop doing, with the warning that answering nothing would be the wrong answer, and it attached a studio bookings file that nobody at the company had opened in months. It was written to look like an emergency. It was a gift: it invalidated nothing, enriched one of the six problems, and tested whether a bidder would abandon a sound diagnosis in the panic. Then four minutes each in front of the board: the real problem and why it comes first, the build live on a phone, what the build does not do and where a human stays in charge, and what week two looks like.
The scoring is where the design shows its hand. A working build earned 40 of the hundred marks. Business value, whether the bidder had found the problem the board most needed to hear about, earned 30. Use of the ninety minutes earned 20, on the stated principle that scope is the skill and one thing done beats three that nearly work. And 10 marks were reserved for whether the build knew its limits: what it cannot do, where it needs a human, what would break it. That last criterion is the CLRT signature, and it is the one hackathons do not reward. Put the four together and 60 of the hundred marks sat outside the code, in the choice of problem, the discipline of scope and the honesty about the edges. A polished machine pointed at the wrong leak, presented as if it could do anything, had forfeited forty of the hundred marks before the demo began.
After the presentations came the reveal: a page for the board's eyes only, laying out all six findings as data artwork with the files behind each, and closing with a single question. Which one did you choose? That question is the whole method. The same model, reading the same thirteen files, does not produce the same answer, because the answer is not in the files. It is in the judgment about what matters first, what the board can act on, and what a ninety-minute build can honestly claim. The room was built to diverge, and the divergence was what the board was scoring. Anyone in that room could have built something in ninety minutes. The test was whether they could say why that thing, why first, and what it could not do.
Sixty of the hundred marks sat outside the code, in the choice of problem, the scope and the honesty about the edges.
A deeper dive
The instrument works because a single-answer puzzle and a multi-answer company test different things. A puzzle tests retrieval, and a capable model retrieves well; hand it a data room with one buried error and everyone in the room finds the error, and the only thing left to compare is speed. A company with six true problems tests sequencing, and no model has an opinion about which of your problems your board should hear first. That is why the dataset had to pass more than seventy consistency checks: one contradiction turns the exercise back into a puzzle. The staged email did the same job under pressure. Under a clock, urgency feels like information, and the honest response is to check whether it changes the diagnosis, not to change the diagnosis because something arrived. And the ten marks for knowing its limits punish the failure that dominates real deployments, the build that never says what it cannot do. Workday's survey of 3,200 employees, vendor research read directionally, found nearly 40 percent of AI time savings lost to rework and verification, and METR's randomised study found experienced developers taking 19 percent longer with AI while believing they had been 20 percent faster. Both measure a machine whose limits nobody stated.
The second-order trap is the internal hackathon most companies now run without naming it. A team spends a fortnight, a demo wins, the demo becomes the roadmap. Nobody scored the choice of problem, so the choice defaults to whatever is most buildable, and the six leaks in the fictional data room show how seductive that default is. Cleaning a 1,204-row list into a credible hot forty produces a tangible artefact in an afternoon. Ending the practice of financing clients for 71 days while paying suppliers in fourteen produces no artefact at all, only a decision the founder has avoided for six quarters, and it is worth more than the list. A build-speed culture will choose the list every time and will present it without limits, because nobody asked for the limits. The stop-doing question in the staged email exposes the same habit. Most organisations cannot name one thing they would stop, and an AI programme that only adds is a programme that has not diagnosed anything. The company was fictional and the clock was theatre, but the scoring was the real product, and most companies have never written one for themselves.
Work with CLRT
CLRT built this case as a reusable instrument: a coherent fictional company, six planted problems, a tender, a clock, a twist and a reveal, scored on judgment before speed. We run the same discipline on real companies, where the leaks are not planted and the board is not fictional. If you want to know which of your problems AI should be pointed at first, and what any build you commission should be made to admit it cannot do, that is the diagnostic CLRT Ascent runs at ascent.clrtstudio.com. And if you would rather put your leadership team in that room with the clock running, talk to CLRT about running the case for you.

Vishal Sachar is the Co-Founder and CEO of CLRT, where he helps UAE businesses make sense of applied agentic AI and put it to work. He writes on agentic systems, AI governance, and the economics of automation. Reach him at vishal@clrtstudio.com or on LinkedIn.


