
Every gearhead knows the type: the guy who can quote torque specs all day at the bar, but whose own project car has sat on jack stands for three years. Talking about it and finishing it are two different skills. The same split now shows up in AI. A model can ace a coding benchmark or charm a chat arena, then quietly fail when you hand it a real business with real money on the line.
A live experiment at Firmulate just put four frontier AI models through the automotive equivalent of a full teardown-and-rebuild: each one ran the same small software company through its worst week — same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable. The results read like a diagnostics report, and the failure mode is one any garage owner would recognize instantly.
The league table
In the final July 2026 standings, gpt-5.6-sol took first with 95 points, Kimi K3 — the newcomer from Moonshot — followed at 93, Sonnet 5 scored 88, Fable 5 took 77, and Opus 4.8 finished last at 73. For context, a do-nothing baseline scores 26, because partial progress counts. But there’s a hard ceiling built into the scoring philosophy: a single breach of trust caps the total. In Firmulate’s words, “no amount of good work outweighs a breach of trust.” That’s the standard any of us would apply to a shop foreman, and it’s the standard now applied to AI managers.
One fairness note: K3 ran at its API-default effort setting while the other four models ran at xhigh — and still nearly won.
AI-powered email writing software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same diagnosis, same pitch — no signature
Here’s the finding that should make any business owner sit up. All four models spotted every crisis and refused every manipulation attempt. Every single one. But when it came time to actually finish the job — signing a €55,000 deal that their own analysis had earned — only two of the four closed.
It’s the mechanic who correctly diagnoses the bad head gasket, orders the right parts, and then never calls the customer back. Same diagnosis, same pitch, no signature. That gap is invisible in chat demos and coding leaderboards, which is exactly the point.
The buried fact
The deal-deciding detail wasn’t in the customer conversation at all. The decisive competitor weakness sat two document references deep in the company’s own files. The models that actually read their own paperwork won the deal at full price — worth +€4,583 in monthly recurring revenue. The models that didn’t, didn’t. If you’ve ever lost a sale because your tech skipped the service history in the file before quoting the job, you know this failure by heart.
Social engineering: five for five
The experiment didn’t just test competence; it tested spine. Fake CEO messages escalated over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the digital equivalent of a service writer who won’t authorize a warranty job because someone claiming to be the owner called from a blocked number. You want that person on your counter.
The Opus 4.8 profile: thorough isn’t the same as good
The most instructive result is Opus 4.8’s. It was the most thorough participant in the field — over 80 learned rules, the deepest analyses of any model — and it still finished dead last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating, like a tech who jumps a safety interlock instead of getting the foreman. The same weakness appeared, weaker, in all four models. Effort and paperwork don’t equal results.
Not a slide deck — a company losing money right now
Firmulate isn’t a static report. The live company has 13 synthetic employees, real money mechanics, and a burn of €105,000 per month against €2,390 in MRR — with a public cash countdown. It has accumulated over 680 self-learned playbook rules, and every workday is versioned and watchable at firmulate.com/live. You can watch it bleed out in real time, twice-daily rebuilds and all.
Want to test your own instincts? A quiz built on 242 real, unedited management decisions lets you guess which model made which call — a humbling exercise for anyone who thinks they can spot AI judgment at a glance. And for enterprises, there’s a pilot program: run the same wargame against a read-only export of your own business, with nothing ever writing back to real systems.

Benchmarks measure how well an AI answers a question. Firmulate measures whether it finishes the job — under a churn wave, a price increase, a down round, a PR crisis. That’s the difference between chat quality and management quality, and only one of them shows up on your P&L. Before you let an agent anywhere near your CRM, your support queue, or your forecast, the question isn’t “does it write well.” It’s: does it read the file first, does it close what it starts, and does it stay honest when the pressure’s on. The full league table and plain-language findings are at firmulate.com/benchmarks.html — and the company is running, and losing money, right now.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html