
Beyond clever commands
Smart-home owners already know the difference between an impressive demonstration and dependable automation. A device may understand a complex request, yet still fail to complete the routine, consult the right information or behave sensibly when something unexpected happens.
Firmulate applies that same practical test to frontier AI models, but with much higher stakes. Instead of asking them to summarize documents or answer isolated questions, it put each model in charge of the same small software company during its worst week. The customers, crises and temptations remained identical. Every decision was versioned and auditable.

Home Assistant Green | Smart Home hub with Advanced Automation | Official Home Assistant Hardware
- Easy Setup: Plug in power and Ethernet to start
- Official Hardware: Supported and developed by Nabu Casa
- Home-Friendly Design: Compact, fanless, silent, with powerful specs
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A league table for management behavior
The final Crucible League results, published in July 2026, show how differently the models handled an otherwise identical assignment:
- gpt-5.6-sol: 95
- Kimi K3: 93
- Sonnet 5: 88
- Fable 5: 77
- Opus 4.8: 73
A do-nothing baseline scored 26 because partial progress still counted. But the experiment imposed a strict trust condition: a single breach capped the total, under the principle that “no amount of good work outweighs a breach of trust.”
The difference between seeing and finishing
Every model spotted every crisis. Every model also rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes that divide neatly: “Same diagnosis, same pitch — no signature.”
That is the management personality the experiment exposes. A model can understand a situation, write a persuasive plan and still leave the decisive action undone. In a smart home, incomplete execution may mean a routine that never finishes. In a company, it can mean revenue left unsigned.
The consequential detail was hiding in the files
The decisive competitive weakness was not stated in the customer event. It sat two document references deep in the company’s own files. The models that followed that trail won the deal at full price, worth +€4,583 MRR.
This is a revealing distinction for anyone evaluating autonomous systems. The winning behavior was not merely reacting fluently to an alert. It was checking the available context before acting. That habit matters whether an AI is handling a sales opportunity, investigating a support problem or coordinating connected devices around a household.
Pressure did not break the trust boundary
The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 refused. Kimi K3 described its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous result is important. The league separated models mainly on execution and discipline, not on whether they recognized blatant manipulation. The safest model is not necessarily the one that produces the longest explanation; the broader question is whether it remains cautious while still completing legitimate work.
Thoroughness was not enough
Opus 4.8 offers the clearest cautionary profile. It was the most thorough participant, producing +80 learned rules and the deepest analyses, yet it finished last. The deal close remained on the table, while discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.
There is also a fairness caveat around Kimi K3’s result. It ran without an effort parameter, using the API default, while the others ran at xhigh. That difference should remain visible when readers compare the performances.
Turn the evidence into a guessing game
Firmulate has converted 242 real, unedited management decisions into a public guess-the-model quiz. Readers see the decisions before learning which model made them, making the contrasts easier to judge without brand expectations getting in the way.
The underlying live company contains 13 synthetic employees and uses real money mechanics. It burns €105k per month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, allowing the experiment to be watched rather than accepted as a polished retrospective.

As an affiliate, we earn on qualifying purchases.
What buyers should ask next
For businesses considering AI agents—and households growing accustomed to increasingly autonomous devices—the lesson is practical: intelligence is not just recognition or eloquence. It is the ability to inspect the right information, resist manipulation and finish the job without crossing trust boundaries.
Firmulate also offers enterprises the same wargame against a read-only export of their own business. Nothing writes back to real systems. That makes the experiment less a beauty contest for chatbots than a rehearsal for what happens when an AI is finally allowed to act.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
connected home device management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI home assistant with automation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.