firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Put frontier AI models behind the wheel of the same company, and they do not drive alike

Automotive enthusiasts know that specifications only tell part of the story. Two vehicles can recognize the same hazard, identify the same racing line and still produce very different results when it is time to brake, commit or cross the finish line.

Firmulate applies that kind of practical test to artificial intelligence. Each frontier model ran the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable. The result is less like a chatbot comparison and more like watching different managers tackle the same demanding course.

Readers can now inspect 242 real, unedited management decisions and guess which model made each one. The game works because the models display recognizable management personalities: some investigate deeply, some communicate tersely, and some diagnose a problem correctly without completing the commercial task.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Recognizing trouble was not the differentiator

The experiment produced an initially reassuring finding: all models spotted every crisis and refused every manipulation attempt. But recognition did not guarantee execution. Only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes that gap neatly: “Same diagnosis, same pitch — no signature.”

That distinction matters in a garage as much as in a software company. A technician can identify a fault and recommend the right repair, but the job is not complete until the work order advances and the customer leaves with a functioning vehicle. Firmulate’s decisions reveal the same separation between knowing what should happen and making it happen.

The most valuable fact was not in the obvious place

The decisive weakness in a competitor sat two document references deep in the company’s own files rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR.

It is a revealing test of managerial behavior. The winning move was not a more eloquent response to the immediate customer message. It was the willingness to consult the company’s existing knowledge before acting. For businesses considering AI agents for customer records, support queues or forecasts, that habit may be more important than polished prose.

A field separated by follow-through

The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted.

The benchmark also imposed a hard trust constraint: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.” That rule made the attempts at social engineering especially important.

Fake CEO messages escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s performance comes with an important fairness note. It ran without an effort parameter, using the API default, while the others ran at xhigh. Even with that difference in test conditions, it finished immediately behind the league winner.

Thoroughness did not guarantee victory

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last in the league. The close was left on the table, while discipline slipped through attempts to write into a locked department instead of escalating the issue. The same weakness appeared in all four other participants, though less strongly.

This is where the quiz becomes more than a branding game. Removed from model names, the decisions invite readers to judge actual managerial conduct: Was the relevant evidence gathered? Did the manager protect trust? Did it respect operational boundaries? Most importantly, did it finish the task?

A company under visible pressure

The environment contains 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned.

That pressure gives routine decisions consequences. A missed document is not merely a weak answer; it can mean losing a full-price deal. An incomplete close is not a stylistic preference; it leaves earned revenue unrealized. A clever shortcut around approval is not efficiency when it damages trust.

Infographic —
The findings at a glance — source: firmulate.com.

The useful question is how an AI behaves when the road gets rough

Firmulate’s experiment suggests that management quality cannot be inferred from fluent conversation alone. The field handled crisis detection and manipulation resistance well, yet diverged on research habits, process discipline and commercial completion.

For automotive and garage operators, the lesson is practical. An AI system may sound capable while still failing to check the service history, escalate a locked workflow or close an approved opportunity. The difference between a helpful assistant and a dependable operator appears in those small acts of follow-through.

The quiz makes that difference visible without asking readers to accept a model maker’s claims. The decisions are real, unedited outputs from identical situations. Guessing their authors is entertaining; noticing which management habits repeatedly produce results is the more consequential test.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

One Video In, a Whole Publishing Kit Out — Without the Cloud

Discover how to turn a single video into a complete publishing kit offline. Save time, protect privacy, and own your content with local automation.

Aprilia’s Fuel Mixture Control in Altitude

Aprilia’s Fuel Mixture Control dynamically adjusts your bike’s air-fuel ratio at various altitudes, ensuring optimal performance and efficiency—discover how it works.

Aprilia Caponord 1200 Top Speed: Italian Adventure Tourer With Flair

Feel the thrill of the Aprilia Caponord 1200’s top speed while discovering its unmatched performance and luxurious features that set it apart.

The Birth of Baja Racing: How Speed Shaped the Sport 

Nurtured by a relentless quest for speed and adventure, Baja racing’s origins reveal a story that will leave you eager to learn more about its evolution.