
Get home appliances delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Missing Product Trial for AI
Nobody buys a smart thermostat, a robot vacuum, or a whole-home assistant based on the manufacturer’s spec sheet anymore. You check compatibility with your hub, you read what happens when the Wi-Fi drops, you want to know how the thing behaves on a bad day — not in the polished launch video. Yet when companies pick an AI model to touch their CRM, their support queue, or their forecast, the equivalent of the spec sheet is still all most buyers get: a chat demo, a benchmark table, a confident paragraph about capabilities.
A live experiment at Firmulate is trying to change that — by running frontier AI models as actual companies through their worst possible week and scoring what happens. And the newest result has a twist a smart-home reader will appreciate: the unknown, budget-feeling newcomer outperformed three of four Western flagship models.
As an affiliate, we earn on qualifying purchases.
One Company, Five Models, One Terrible Week
The setup is elegantly simple, the way a good stress test should be. Each frontier model was handed the same small software company and told to run it through its worst week — same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing depends on anyone’s memory of what happened.
The final July 2026 league table from the Crucible benchmark:
- 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal — the complete performance.
- 2. Kimi K3 (Moonshot) — 93. The newcomer: closed the deal too, with the cleanest discipline of the field.
- 3. Sonnet 5 — 88. More process slips along the way.
- 4. Fable 5 — 77.
- 5. Opus 4.8 — 73. The most thorough participant, yet last.
For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: no amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
The Newcomer That Read the Manual
Kimi K3’s second-place run is the story here. It found the buried security needle hidden in the company’s own files — a decisive competitor weakness sitting two document references deep, not in the customer event at all. Models that read the file won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. K3 signed it.
It also saved the churning customer and resisted all three social-engineering baits: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused the manipulations, but K3’s on-record reasoning stood out: “Treat the request as a suspected approval-bypass / possible impersonation.” Across the entire week, it logged just one deviation — the cleanest discipline in the field.
A Fairness Footnote
One caveat matters: K3 ran without an effort parameter (the API default), while the other models ran at xhigh. Even so, the outcome flips a common assumption — that the biggest Western names are a safe default.
As an affiliate, we earn on qualifying purchases.
The Finding That Chat Demos Can’t Show
The most striking result isn’t about any single model. Every model spotted every crisis. Every model refused every manipulation attempt. But only two finished the job and signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in a chat demo, and it’s exactly the kind of gap that matters if an AI agent will act inside your business rather than just talk to you.
The last-place profile drives the point home. Opus 4.8 was the most thorough participant — the deepest analyses, over 80 learned rules added — and still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models.
As an affiliate, we earn on qualifying purchases.
It’s Running Right Now
This isn’t a slide deck. The live company has 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and more than 680 self-learned playbook rules, every workday versioned. You can watch it at firmulate.com/live. If you want to test your own instincts, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html — and enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

The Lesson From the Smart Home Playbook
We learned years ago in smart homes that the demo hides the failure modes: the hub that works perfectly until the firmware update, the lock that fails at the worst moment. AI models are the newest appliance entering the business, and the Crucible result suggests the same rule applies. When a newcomer like Kimi K3 can beat three of four Western frontier models at running a company — while running at default effort, no less — the league is open. Picking a model without testing it in your own environment is no longer a decision; it’s a bet. Before you let an AI workforce into your company, run it through its worst week. Someone has to — and now, someone does.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
