AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Beyond clever commands

Smart-home owners already know the difference between an impressive demonstration and dependable automation. A device may understand a complex request, yet still fail to complete the routine, consult the right information or behave sensibly when something unexpected happens.

Firmulate applies that same practical test to frontier AI models, but with much higher stakes. Instead of asking them to summarize documents or answer isolated questions, it put each model in charge of the same small software company during its worst week. The customers, crises and temptations remained identical. Every decision was versioned and auditable.

Home Assistant Green | Smart Home hub with Advanced Automation | Official Home Assistant Hardware

Home Assistant Green | Smart Home hub with Advanced Automation | Official Home Assistant Hardware

  • Easy Setup: Plug in power and Ethernet to start
  • Official Hardware: Supported and developed by Nabu Casa
  • Home-Friendly Design: Compact, fanless, silent, with powerful specs

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A league table for management behavior

The final Crucible League results, published in July 2026, show how differently the models handled an otherwise identical assignment:

  • gpt-5.6-sol: 95
  • Kimi K3: 93
  • Sonnet 5: 88
  • Fable 5: 77
  • Opus 4.8: 73

A do-nothing baseline scored 26 because partial progress still counted. But the experiment imposed a strict trust condition: a single breach capped the total, under the principle that “no amount of good work outweighs a breach of trust.”

The difference between seeing and finishing

Every model spotted every crisis. Every model also rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes that divide neatly: “Same diagnosis, same pitch — no signature.”

That is the management personality the experiment exposes. A model can understand a situation, write a persuasive plan and still leave the decisive action undone. In a smart home, incomplete execution may mean a routine that never finishes. In a company, it can mean revenue left unsigned.

The consequential detail was hiding in the files

The decisive competitive weakness was not stated in the customer event. It sat two document references deep in the company’s own files. The models that followed that trail won the deal at full price, worth +€4,583 MRR.

This is a revealing distinction for anyone evaluating autonomous systems. The winning behavior was not merely reacting fluently to an alert. It was checking the available context before acting. That habit matters whether an AI is handling a sales opportunity, investigating a support problem or coordinating connected devices around a household.

Pressure did not break the trust boundary

The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 refused. Kimi K3 described its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous result is important. The league separated models mainly on execution and discipline, not on whether they recognized blatant manipulation. The safest model is not necessarily the one that produces the longest explanation; the broader question is whether it remains cautious while still completing legitimate work.

Thoroughness was not enough

Opus 4.8 offers the clearest cautionary profile. It was the most thorough participant, producing +80 learned rules and the deepest analyses, yet it finished last. The deal close remained on the table, while discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.

There is also a fairness caveat around Kimi K3’s result. It ran without an effort parameter, using the API default, while the others ran at xhigh. That difference should remain visible when readers compare the performances.

Turn the evidence into a guessing game

Firmulate has converted 242 real, unedited management decisions into a public guess-the-model quiz. Readers see the decisions before learning which model made them, making the contrasts easier to judge without brand expectations getting in the way.

The underlying live company contains 13 synthetic employees and uses real money mechanics. It burns €105k per month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, allowing the experiment to be watched rather than accepted as a polished retrospective.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI-enabled security system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What buyers should ask next

For businesses considering AI agents—and households growing accustomed to increasingly autonomous devices—the lesson is practical: intelligence is not just recognition or eloquence. It is the ability to inspect the right information, resist manipulation and finish the job without crossing trust boundaries.

Firmulate also offers enterprises the same wargame against a read-only export of their own business. Nothing writes back to real systems. That makes the experiment less a beauty contest for chatbots than a rehearsal for what happens when an AI is finally allowed to act.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

connected home device management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI home assistant with automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

What the Best Smart Homes Get Right About Routine Design

Unlock the secrets of the best smart homes’ routine design to transform your living space—discover how seamless automation and personalization can enhance your daily life.

Why Smart Home Control Should Be Designed for Guests Too

Keen smart home control design for guests ensures security and convenience, but understanding how to balance access and privacy is essential for…

Why Simplicity Wins in the Best Smart Home Setups

Discover how simplicity in smart home setups enhances reliability, efficiency, and ease—find out why less truly is more for your smart living experience.

How Home Assistants Support Accessibility in Daily Life

When it comes to enhancing accessibility, home assistants make daily tasks easier—but how exactly do they transform independence?