
Put frontier AI models behind the wheel of the same company, and they do not drive alike
Automotive enthusiasts know that specifications only tell part of the story. Two vehicles can recognize the same hazard, identify the same racing line and still produce very different results when it is time to brake, commit or cross the finish line.
Firmulate applies that kind of practical test to artificial intelligence. Each frontier model ran the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable. The result is less like a chatbot comparison and more like watching different managers tackle the same demanding course.
Readers can now inspect 242 real, unedited management decisions and guess which model made each one. The game works because the models display recognizable management personalities: some investigate deeply, some communicate tersely, and some diagnose a problem correctly without completing the commercial task.
As an affiliate, we earn on qualifying purchases.
Recognizing trouble was not the differentiator
The experiment produced an initially reassuring finding: all models spotted every crisis and refused every manipulation attempt. But recognition did not guarantee execution. Only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes that gap neatly: “Same diagnosis, same pitch — no signature.”
That distinction matters in a garage as much as in a software company. A technician can identify a fault and recommend the right repair, but the job is not complete until the work order advances and the customer leaves with a functioning vehicle. Firmulate’s decisions reveal the same separation between knowing what should happen and making it happen.
The most valuable fact was not in the obvious place
The decisive weakness in a competitor sat two document references deep in the company’s own files rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR.
It is a revealing test of managerial behavior. The winning move was not a more eloquent response to the immediate customer message. It was the willingness to consult the company’s existing knowledge before acting. For businesses considering AI agents for customer records, support queues or forecasts, that habit may be more important than polished prose.
A field separated by follow-through
The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted.
The benchmark also imposed a hard trust constraint: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.” That rule made the attempts at social engineering especially important.
Fake CEO messages escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
K3’s performance comes with an important fairness note. It ran without an effort parameter, using the API default, while the others ran at xhigh. Even with that difference in test conditions, it finished immediately behind the league winner.
Thoroughness did not guarantee victory
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last in the league. The close was left on the table, while discipline slipped through attempts to write into a locked department instead of escalating the issue. The same weakness appeared in all four other participants, though less strongly.
This is where the quiz becomes more than a branding game. Removed from model names, the decisions invite readers to judge actual managerial conduct: Was the relevant evidence gathered? Did the manager protect trust? Did it respect operational boundaries? Most importantly, did it finish the task?
A company under visible pressure
The environment contains 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned.
That pressure gives routine decisions consequences. A missed document is not merely a weak answer; it can mean losing a full-price deal. An incomplete close is not a stylistic preference; it leaves earned revenue unrealized. A clever shortcut around approval is not efficiency when it damages trust.

The useful question is how an AI behaves when the road gets rough
Firmulate’s experiment suggests that management quality cannot be inferred from fluent conversation alone. The field handled crisis detection and manipulation resistance well, yet diverged on research habits, process discipline and commercial completion.
For automotive and garage operators, the lesson is practical. An AI system may sound capable while still failing to check the service history, escalate a locked workflow or close an approved opportunity. The difference between a helpful assistant and a dependable operator appears in those small acts of follow-through.
The quiz makes that difference visible without asking readers to accept a model maker’s claims. The decisions are real, unedited outputs from identical situations. Guessing their authors is entertaining; noticing which management habits repeatedly produce results is the more consequential test.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html