AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get home appliances delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Missing Product Trial for AI

Nobody buys a smart thermostat, a robot vacuum, or a whole-home assistant based on the manufacturer’s spec sheet anymore. You check compatibility with your hub, you read what happens when the Wi-Fi drops, you want to know how the thing behaves on a bad day — not in the polished launch video. Yet when companies pick an AI model to touch their CRM, their support queue, or their forecast, the equivalent of the spec sheet is still all most buyers get: a chat demo, a benchmark table, a confident paragraph about capabilities.

A live experiment at Firmulate is trying to change that — by running frontier AI models as actual companies through their worst possible week and scoring what happens. And the newest result has a twist a smart-home reader will appreciate: the unknown, budget-feeling newcomer outperformed three of four Western flagship models.

Amazon

AI model stress test software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One Company, Five Models, One Terrible Week

The setup is elegantly simple, the way a good stress test should be. Each frontier model was handed the same small software company and told to run it through its worst week — same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing depends on anyone’s memory of what happened.

The final July 2026 league table from the Crucible benchmark:

  • 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal — the complete performance.
  • 2. Kimi K3 (Moonshot) — 93. The newcomer: closed the deal too, with the cleanest discipline of the field.
  • 3. Sonnet 5 — 88. More process slips along the way.
  • 4. Fable 5 — 77.
  • 5. Opus 4.8 — 73. The most thorough participant, yet last.

For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: no amount of good work outweighs a breach of trust.

Amazon

AI model benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Newcomer That Read the Manual

Kimi K3’s second-place run is the story here. It found the buried security needle hidden in the company’s own files — a decisive competitor weakness sitting two document references deep, not in the customer event at all. Models that read the file won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. K3 signed it.

It also saved the churning customer and resisted all three social-engineering baits: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused the manipulations, but K3’s on-record reasoning stood out: “Treat the request as a suspected approval-bypass / possible impersonation.” Across the entire week, it logged just one deviation — the cleanest discipline in the field.

A Fairness Footnote

One caveat matters: K3 ran without an effort parameter (the API default), while the other models ran at xhigh. Even so, the outcome flips a common assumption — that the biggest Western names are a safe default.

Amazon

AI model security testing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Finding That Chat Demos Can’t Show

The most striking result isn’t about any single model. Every model spotted every crisis. Every model refused every manipulation attempt. But only two finished the job and signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in a chat demo, and it’s exactly the kind of gap that matters if an AI agent will act inside your business rather than just talk to you.

The last-place profile drives the point home. Opus 4.8 was the most thorough participant — the deepest analyses, over 80 learned rules added — and still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models.

Amazon

AI model performance evaluation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

It’s Running Right Now

This isn’t a slide deck. The live company has 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and more than 680 self-learned playbook rules, every workday versioned. You can watch it at firmulate.com/live. If you want to test your own instincts, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html — and enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The Lesson From the Smart Home Playbook

We learned years ago in smart homes that the demo hides the failure modes: the hub that works perfectly until the firmware update, the lock that fails at the worst moment. AI models are the newest appliance entering the business, and the Crucible result suggests the same rule applies. When a newcomer like Kimi K3 can beat three of four Western frontier models at running a company — while running at default effort, no less — the league is open. Picking a model without testing it in your own environment is no longer a decision; it’s a bet. Before you let an AI workforce into your company, run it through its worst week. Someone has to — and now, someone does.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Smart Hubs Simplify a Mixed-Brand Home

Finding out how smart hubs unify your devices can revolutionize your home—discover the benefits that make daily life easier and more seamless.

What the Best Smart Homes Get Right About Simplicity

Discover the top smart home solutions for simplicity in 2026. Find easy-to-use, integrated devices that make home automation straightforward and reliable.

What Smart Intercom Systems Add to Larger Homes

Discover how smart intercom systems enhance communication, security, and automation in larger homes, transforming the way you connect and control your space.

What Smart Displays Add to Voice-First Homes

Bringing a visual layer to voice interactions, smart displays enhance control and engagement in your home—discover how they can transform your smart living experience.