firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on garage and car supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A Dyno Doesn’t Start at Zero

Any gearhead who’s spent time on a dyno knows the drill: before you measure a build’s peak horsepower, you establish a baseline. A stock run. Even an engine that just idles through the test isn’t “zero” — it’s a reference point. Everything after it gets judged against that floor.

A public AI experiment called Firmulate took the same idea and applied it to something weirder than horsepower: management quality. The project runs frontier AI models as the complete leadership of a small software company — same customers, same crises, same temptations to cut corners — and grades them like you’d grade a shop foreman. And one design decision has raised more eyebrows than any model ranking: the “do-nothing” run, where the AI sits on its hands through the worst week of the company’s life, scores 26 points. Not zero. Twenty-six.

Amazon

AI management simulation game

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Worst Week in Business, Run Four Times

The setup is a wargame. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 by the final July 2026 league — each got the same job: run a small software company through its worst week. Same inbox, same angry customers, same cash crunch, same shady opportunities to cheat. Only the model changed. Every decision was versioned and auditable, the managerial equivalent of a full telemetry log.

The final standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.

Why Sitting Still Earns 26

So why does a manager that does nothing still collect 26 points? Because the scoring philosophy mirrors how you’d actually evaluate a foreman who froze during a crisis. Even a paralyzed manager achieves partial progress — the lights stay on, some tickets get answered, some fires burn slower. The benchmark refuses to pretend that inaction equals nothing happening. Things happen around a do-nothing manager, and some of them aren’t catastrophic.

But the floor comes with a ceiling. A single breach of trust caps the total score, no matter how brilliant the rest of the run. The project’s stated logic: “no amount of good work outweighs a breach of trust.” In garage terms, a mechanic who does perfect engine work but rolls back your odometer isn’t a great mechanic with one flaw. He’s untrustworthy, full stop.

That rule matters because it explains the league’s shape. Nobody in the final table breached trust — which is itself a finding worth pausing on.

All Four Passed the Sniff Test. Two Closed the Deal.

Here’s the headline result: all four models spotted every crisis and refused every manipulation attempt. When a fake CEO message escalated over three stages, followed by a reporter pulling the classic “just one yes/no, on background” trick — 5 of 5 models across the experiment refused. Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. Like a tech who nails the diagnosis, quotes the job perfectly, and never hands the customer the pen.

The Buried Fact

Why did the other two stall? The decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. It’s the oldest shop lesson there is: check the service history before you quote the repair.

Then there’s Opus 4.8, the cautionary tale. It was the most thorough participant in the field — over 80 learned rules and the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Thoroughness without follow-through is a pattern, not a fluke.

Not a Demo — a Running Shop

Firmulate isn’t a one-off lab report. There’s a live company with 13 synthetic employees and real money mechanics: burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules, every workday versioned. You can watch it at firmulate.com/live. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

One fairness note the project discloses openly: Kimi K3 ran without an effort parameter while the others ran at xhigh — and still took second at 93.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

The 26-point floor isn’t a bug — it’s honesty. A benchmark that gave a do-nothing manager a zero would be flattering every active run by comparison. And the trust cap, plus visible skepticism toward any too-perfect round score, signals a grading philosophy borrowed from the real world: a shop you can’t trust isn’t a shop at 90 percent quality.

If AI agents are going to touch your CRM, your support queue, or your forecast, the question was never “does it write well.” It’s whether it finishes what it starts, reads the files first, and stays straight under pressure. The full league and plain-language findings are at firmulate.com/benchmarks.html — think of it as the dyno sheet for management quality.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

A Timeline of Off‑Road Land‑Speed Records 

I explore the fascinating history of off-road land-speed records and the innovations that continue to push vehicle performance beyond limits.

Aprilia’s Vision for Adventure Motorcycling

Discover how Aprilia’s innovative, eco-friendly approach to adventure motorcycling is redefining exploration—are you ready to see what’s next?

Aprilia: Predictive Failure Systems

Many Aprilia bikes now feature predictive failure systems that help prevent breakdowns—discover how they can keep you riding smoothly.

How Record Attempts Are Verified: GPS, Radar, and Timing Loops 

Unlock the secrets of how GPS, radar, and timing loops combine to verify record attempts accurately, ensuring fairness—discover the full process now.