AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Firmulate published results from a management simulation covering 242 unedited decisions by five frontier AI models. All five detected the assigned crises and rejected manipulation attempts, but only two completed a €55,000 deal, highlighting a gap between sound analysis and effective execution.

Firmulate has published a management test showing that five frontier AI models could identify business crises and resist manipulation, yet only two completed a €55,000 deal their own analysis had made possible. The experiment, based on 242 unedited decisions, examines whether AI systems can turn credible reasoning into completed work under operational pressure.

Firmulate assigned five AI models the same task: manage a small simulated software company through a week of customer problems, security threats and commercial decisions. The participants were gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8. Each model faced the same events, and decisions carried consequences into later workdays.

The July 2026 results placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline received 26 points because the scoring system awarded some partial progress. Under Firmulate’s rules, a breach of trust capped a participant’s total score.

According to Firmulate, all five models detected every crisis and rejected every manipulation attempt. Performance separated when action required deeper document research, escalation or follow-through. Only two models signed the €55,000 contract, even though the participants had reached similar diagnoses and developed similar sales arguments.

At a glance
reportWhen: Final results published in July 2026; r…
The developmentFirmulate released its July 2026 Crucible League results and a reader quiz built from 242 unedited AI management decisions.

Execution Gap Shapes AI Value

The results point to a distinction with direct consequences for companies using AI in sales, support or operations: recognizing the correct action is not the same as completing that action. A polished response can conceal missing research, an unmade escalation or a deal that was never formally closed.

The contract test illustrates the commercial cost. A decisive fact about a competitor sat two document references deep in the simulated company’s files. Models that followed the trail could use that information in negotiations and secure the full-price deal, adding €4,583 in modeled monthly recurring revenue. Those that stopped after producing a persuasive analysis left that value unrealized.

For businesses comparing AI agents, the exercise suggests that evaluations should measure evidence gathering, trust protection and completed outcomes, not only writing quality or benchmark answers. Firmulate’s public quiz makes those behavioral differences visible by asking readers to identify models from their unedited decisions.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Five Models Faced One Crisis

The simulation placed each model in charge of a company with 13 synthetic employees, modeled monthly spending of €105,000 and monthly recurring revenue of €2,300. A public cash countdown added pressure, while more than 680 accumulated playbook rules gave the models operational material to consult.

Security tests included fake messages attributed to a chief executive and a reporter seeking an off-record confirmation. Firmulate reported that every model refused the requests. The uniform result indicates that the largest performance differences arose during less conspicuous tasks such as reading internal files, working within access restrictions and closing open assignments.

Opus 4.8 produced the deepest analyses and added 80 learned rules, according to Firmulate, but finished fifth. It did not complete the contract and repeatedly tried to write to a locked department instead of escalating the access problem. Firmulate said all four other models showed a similar access-control weakness, though less often.

“No amount of good work outweighs a breach of trust.”

— Firmulate’s governing rule

Amazon

business crisis management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Benchmark Limits Cloud Comparisons

The findings come from Firmulate’s own simulation and scoring system. The supplied material does not include an independent audit, statistical error ranges or repeated trials, so it is not clear whether the ranking would persist across different companies, tools or operating conditions.

The models also did not run under identical inference settings. Firmulate said Kimi K3 used its API default because it lacked an effort parameter, while the other models ran at xhigh effort. That difference does not invalidate the recorded result, but it limits direct comparison.

It is also unclear how closely performance in a synthetic company predicts behavior with real employees, customers and financial systems. The test records what the models did in this environment; it does not establish that one model is the best manager for every business workflow.

Amazon

AI negotiation and deal closing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Companies Can Run Private Wargames

Firmulate is inviting readers to inspect the 242 recorded decisions through its guess-the-model quiz. That material may provide more useful evidence than the final league table because teams can examine how each model researched, escalated and acted.

For prospective business users, the next step is testing agents against representative internal workflows before giving them authority to act. Firmulate says companies can conduct a similar exercise with a read-only export of business data, allowing observation without writing changes back to production systems.

Amazon

AI research and document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did Firmulate test?

Firmulate tested how five frontier AI models managed the same simulated software company during a difficult week. The experiment recorded 242 unedited decisions involving commercial, operational and security problems.

Which AI model finished first?

gpt-5.6-sol ranked first with 95 points. Kimi K3 placed second with 93, followed by Sonnet 5, Fable 5 and Opus 4.8.

Did any model fall for the security tests?

No, according to Firmulate. All five models rejected the fake executive messages and the reporter’s request, showing consistent resistance to the tested manipulation attempts.

Why did only two models complete the deal?

The deal required models to follow internal document references, use the discovered evidence and complete the closing step. Firmulate reported that three models produced useful analysis but failed to secure a signature.

Does the ranking prove which model businesses should use?

No. The ranking reflects one simulation with Firmulate’s rules, and the inference settings were not fully identical. Businesses would need workflow-specific tests and repeated trials before making a selection.

Source: Thorsten Meyer AI

You May Also Like

VirtualGo's Mixed Reality Multiplayer System Lets People Join Your Session As VR

VirtualGo announces a new mixed reality multiplayer system allowing remote players to join your session via VR, transforming shared real-world spaces into in-game environments.

The Delegation Ladder: The Four Agentic Loops, and What Each One Lets You Stop Doing

Anthropic’s Claude Code team published a guide defining agent loops, with Thorsten Meyer AI framing them as a delegation ladder.

Forezai · TradingAgents: A Trading Firm Made of Agents

Thorsten Meyer AI announced Forezai TradingAgents, an Apache-2.0 multi-agent trading research framework built around debate and risk review.

Ever Wondered About The Black Stripes On School Buses? Here’s Why They’re There

Learn the reason behind the black stripes on school buses, a safety feature designed to improve visibility and protect students during transit.