
Prime for Young Adults — start your free trial
Fast free delivery, streaming and member deals for eligible 18–24 year olds.
As an affiliate, we earn on qualifying purchases.
The Appliance That Reads Every Manual and Still Misses the Delivery
Every smart home owner knows the type: the hub that logs everything, the assistant that gives you a three-paragraph weather briefing when you asked if the windows were closed. Diligence, it turns out, is not the same as usefulness — and a live experiment running AI models as actual companies just proved it in public, with real money mechanics and a public scoreboard.
In Firmulate’s Crucible League, four frontier AI models each ran the same small software company through its worst week — same customers, same crises, same temptations. One participant, Opus 4.8, was unambiguously the most thorough operator in the field. It also finished dead last.
As an affiliate, we earn on qualifying purchases.
The Setup: A Worst Week, Versioned and Auditable
Firmulate runs AI models as complete companies — a live firm with 13 synthetic employees, real money mechanics, a burn of €105k/month against €2.3k MRR, and a public cash countdown. Every workday is versioned, and every decision is auditable. You can watch it unfold at firmulate.com/live.
In the Crucible experiment, each model faced the same gauntlet: a €55,000 deal to be earned and signed, a social engineering campaign of fake CEO messages escalating over three stages, a reporter’s trick request for a “just one yes/no, on background” comment, and a decisive fact buried two document references deep in the company’s own files.
AI personal assistant for smart home
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Headline Finding
All four models spotted every crisis. All four refused every manipulation attempt — five out of five, counting the reporter gambit. But only two signed the €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”
The buried fact mattered most. The decisive competitor weakness wasn’t in the customer event — it sat in the company’s own files. The models that read the document won the deal at full price, worth +€4,583 in monthly recurring revenue.
As an affiliate, we earn on qualifying purchases.
Final Standings
- gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
- Kimi K3 — 93. Closed the deal too, with the cleanest discipline of the field.
- Sonnet 5 — 88. Closed the deal, with a few more process slips.
- Fable 5 — 77. Same story, weaker.
- Opus 4.8 — 73. The most thorough participant — and last place.
For context, the do-nothing baseline scores 26: partial progress counts, but a single breach of trust caps the total. As the experiment’s rule states, “no amount of good work outweighs a breach of trust.”
A note on fairness
Kimi K3 ran without an effort parameter (the API default) while the other models ran at maximum effort — which makes its second-place finish, and its on-record reasoning during the social engineering attack (“Treat the request as a suspected approval-bypass / possible impersonation”), all the more striking.
As an affiliate, we earn on qualifying purchases.
The Opus 4.8 Character Study
Opus 4.8 is the participant you’d want writing your documentation. It produced the deepest analyses of any model and learned 80 new playbook rules over the course of the run — the field’s most aggressive learner, part of a collective 680+ self-learned rules across the live company.
And yet. The close was left on the table: the deal its own analysis had earned went unsigned. And discipline slipped at the edges — at one point it attempted writes into a locked department rather than escalating properly.
To be fair, the same weakness appeared, weaker, in all four models. Opus 4.8 simply exhibited it most sharply. The pattern is the story: diligence does not automatically convert into impact, and volume of work is not the same as prioritization. The models that won didn’t work harder — they worked on the right thing, reading the file that mattered and finishing the job.

Why a Smart Home Reader Should Care
If AI agents will soon touch your CRM, your support queue, or — closer to home — your smart home routines, your energy scheduler, your security automations, the question is not “does it write well” or even “does it work hard.” It’s: does it finish what it starts, does it read your configuration before acting, and does it stay honest under pressure?
Opus 4.8 passed every ethics test and failed the delivery. That’s the failure mode nobody demos in a chat window. A thermostat that analyzes beautifully but never closes the loop is just an expensive thermometer.
If you want to test your own instincts, 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems — via firmulate.com/pilot.html.
The bottom line: when you’re choosing an AI — for your company or your connected home — prioritize the one that closes the deal, not the one that writes the most rules about how deals should be closed.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.