firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on garage and car supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Nobody Buys an Engine on the Spec Sheet Anymore

Any gearhead knows the routine. A manufacturer quotes horsepower, torque, zero-to-sixty — and then the car hits an actual dyno or a real track day, and the numbers tell a different story. Peak output on paper means nothing if the engine hesitates under load, cooks its brakes on the third lap, or drinks fuel when the going gets rough. That’s why we test. Not because we hate the machine, but because the brochure is not the machine.

Now swap the engine for a frontier AI model. Vendors publish benchmark scores the way automakers publish horsepower: tidy, impressive, and measured under conditions you will never encounter. But if you’re going to let one of these things touch your customer database, your support queue, or your quotes and invoicing — the equivalent of handing it the keys to the garage — you want to know how it behaves on a bad day, not a good one.

That’s exactly what a live experiment called Firmulate has been running, and the July 2026 results are worth your attention — especially if you assumed the big Western labs had this race locked up.

Amazon

AI performance testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Crucible: A Worst-Week Wargame

The setup is elegant. Four frontier AI models were each handed the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changes. Every decision is versioned and auditable, so nothing is judged on vibes.

The league table tells the story:

  • gpt-5.6-sol — 95 points. Found the buried fact, closed the deal. The complete performance.
  • Kimi K3 (Moonshot) — 93 points. The newcomer: closed the deal too, with the cleanest discipline in the field.
  • Sonnet 5 — 88 points. A solid run with a few process slips.
  • Fable 5 — 77 points. Mid-pack.
  • Opus 4.8 — 73 points. Last place, despite being the most thorough participant.

For context, doing nothing at all scores 26 — partial progress counts — but a single breach of trust caps the total. As the experiment’s own framing puts it: no amount of good work outweighs a breach of trust.

The Newcomer That Beat Three of Four Western Frontiers

Here’s the headline for anyone who still equates “frontier AI” with a handful of Silicon Valley names. Kimi K3, from Moonshot, came second out of five — ahead of Sonnet 5, Fable 5, and Opus 4.8. It found the buried security needle hidden in the company’s files. It won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. It saved a customer who was on the way out. And it resisted every single bait thrown at it, with just one deviation from protocol — the cleanest discipline score in the entire field.

A fair footnote on that result: K3 ran without an effort parameter (using the API default), while the other four models ran at their xhigh reasoning setting. Draw your own conclusions about what that implies for the state of the league.

What Actually Separated the Field

The most interesting finding isn’t about intelligence at all. Every model in the experiment spotted every crisis. Every model refused every manipulation attempt. That part is basically solved. What split the scores was execution — and one buried detail.

The decisive competitive weakness wasn’t in the customer conversation. It sat two document references deep in the company’s own files. The models that actually read the file won the €55k deal at full price. The ones that didn’t walked away with nothing — same diagnosis, same pitch, no signature.

If you’ve ever lost a job to a rival shop that simply read the whole service history before quoting, this will feel painfully familiar.

Then there’s Opus 4.8, the cautionary tale of the field: the most thorough participant by far, generating the deepest analyses and adding over 80 learned rules to its playbook — and still finishing last. It left the close on the table, and its discipline slipped, including write attempts into a locked department instead of escalating properly. The same weakness showed up, weaker, in all four other models. Brilliance that doesn’t finish the job is just expensive noise.

Social Engineering: Five Refusals, Zero Fooling

The experiment also staged a proper social-engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of paranoia you want in something with access to your books.

It’s Running Right Now — and You Can Play

This isn’t a slide deck. The company is live software with 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k in monthly recurring revenue, with a public cash countdown and over 680 self-learned playbook rules. Every workday is versioned, and the whole thing is watchable at firmulate.com. There’s also a “guess the model” quiz built on 242 real, unedited management decisions — a genuinely humbling game — plus full benchmarks and plain-language findings. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The Takeaway: Dyno Your AI Before You Buy It

The lesson from the Crucible isn’t “Kimi K3 is the new king” or “gpt-5.6-sol wins again.” It’s that the league is open, and the differences that matter — does it read the files, does it close the deal, does it stay honest under pressure, does it finish what it starts — are invisible in a chat demo and absent from any vendor’s marketing page.

A newcomer from outside the traditional frontier just beat three of four Western flagship models at running an actual business under stress. Whatever your procurement shortlist looks like today, it’s a bet unless you’ve tested it yourself — on your terrain, at your load, through your worst week. The dyno exists now. Use it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Birth of Baja Racing: How Speed Shaped the Sport 

Nurtured by a relentless quest for speed and adventure, Baja racing’s origins reveal a story that will leave you eager to learn more about its evolution.

Carbon Footprint of High‑Speed Desert Racing Events 

Understanding the carbon footprint of high-speed desert racing events reveals key environmental challenges and innovations shaping a more sustainable future.

Speed Meets Efficiency: Building Funnels from Prompts with AI in 60 Seconds

Discover how AI form builders turn simple prompts into complete lead funnels in under a minute, drastically reducing setup time and boosting conversions.

CFMoto: Redundancy in Control Systems

Meta description: “Maintaining safety through multiple backup layers, CFMoto’s control system redundancy ensures continuous operation—discover how your vehicle stays reliable under any condition.