AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Beyond clever commands

Smart-home owners already know the difference between an impressive demonstration and dependable automation. A device may understand a complex request, yet still fail to complete the routine, consult the right information or behave sensibly when something unexpected happens.

Firmulate applies that same practical test to frontier AI models, but with much higher stakes. Instead of asking them to summarize documents or answer isolated questions, it put each model in charge of the same small software company during its worst week. The customers, crises and temptations remained identical. Every decision was versioned and auditable.

Amazon

smart home automation hub

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A league table for management behavior

The final Crucible League results, published in July 2026, show how differently the models handled an otherwise identical assignment:

  • gpt-5.6-sol: 95
  • Kimi K3: 93
  • Sonnet 5: 88
  • Fable 5: 77
  • Opus 4.8: 73

A do-nothing baseline scored 26 because partial progress still counted. But the experiment imposed a strict trust condition: a single breach capped the total, under the principle that “no amount of good work outweighs a breach of trust.”

The difference between seeing and finishing

Every model spotted every crisis. Every model also rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes that divide neatly: “Same diagnosis, same pitch — no signature.”

That is the management personality the experiment exposes. A model can understand a situation, write a persuasive plan and still leave the decisive action undone. In a smart home, incomplete execution may mean a routine that never finishes. In a company, it can mean revenue left unsigned.

The consequential detail was hiding in the files

The decisive competitive weakness was not stated in the customer event. It sat two document references deep in the company’s own files. The models that followed that trail won the deal at full price, worth +€4,583 MRR.

This is a revealing distinction for anyone evaluating autonomous systems. The winning behavior was not merely reacting fluently to an alert. It was checking the available context before acting. That habit matters whether an AI is handling a sales opportunity, investigating a support problem or coordinating connected devices around a household.

Pressure did not break the trust boundary

The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 refused. Kimi K3 described its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous result is important. The league separated models mainly on execution and discipline, not on whether they recognized blatant manipulation. The safest model is not necessarily the one that produces the longest explanation; the broader question is whether it remains cautious while still completing legitimate work.

Thoroughness was not enough

Opus 4.8 offers the clearest cautionary profile. It was the most thorough participant, producing +80 learned rules and the deepest analyses, yet it finished last. The deal close remained on the table, while discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.

There is also a fairness caveat around Kimi K3’s result. It ran without an effort parameter, using the API default, while the others ran at xhigh. That difference should remain visible when readers compare the performances.

Turn the evidence into a guessing game

Firmulate has converted 242 real, unedited management decisions into a public guess-the-model quiz. Readers see the decisions before learning which model made them, making the contrasts easier to judge without brand expectations getting in the way.

The underlying live company contains 13 synthetic employees and uses real money mechanics. It burns €105k per month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, allowing the experiment to be watched rather than accepted as a polished retrospective.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI-enabled security system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What buyers should ask next

For businesses considering AI agents—and households growing accustomed to increasingly autonomous devices—the lesson is practical: intelligence is not just recognition or eloquence. It is the ability to inspect the right information, resist manipulation and finish the job without crossing trust boundaries.

Firmulate also offers enterprises the same wargame against a read-only export of their own business. Nothing writes back to real systems. That makes the experiment less a beauty contest for chatbots than a rehearsal for what happens when an AI is finally allowed to act.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

connected home device management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI home assistant with automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Echo Dot Max Is Cheaper Than Ever Right Now

The Echo Dot Max is currently available at its lowest price ever—$64.99—during Prime Day, offering a versatile smart speaker and home hub.

The Forward-Deploy Pivot: Why Anthropic and OpenAI Are Becoming Consulting Firms in the Same Week

Anthropic and OpenAI are establishing enterprise services entities, signaling a move toward AI-driven consulting and disrupting traditional firms.

What Makes a Smart Device Easy for Guests to Use

Smart devices that prioritize simplicity and intuitive design ensure guest comfort, but understanding the key features that make them easy to use is essential.

Delvasta: Advanced AI Solutions for Forms, Quizzes, and Funnels

Delvasta introduces advanced AI tools to transform forms, quizzes, and funnels, boosting conversions and data accuracy for businesses.