AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Anyone who has lived with a smart home knows the truth the spec sheets hide. Your thermostat isn’t judged on how nicely it explains setpoints. It’s judged at 2 a.m. when the heating spikes, the energy tariff flips, and the baby’s room is cold. Does it triage? Does it finish the job? Does it quietly make things worse while reporting that everything is fine?

The same gap now exists in business AI. We rank models by how well they answer — coding benchmarks, chat arenas, eloquence contests. But as agents start touching CRMs, support queues and forecasts, the real question isn’t “does it write well.” It’s whether it manages: under pressure, across days, with real money on the line. A live experiment at Firmulate just put that question to the test — and the results should make anyone deploying AI sit up.

The worst week in business, four times over

Firmulate handed four frontier AI models the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision is versioned and auditable. The final league table from July 2026:

  • gpt-5.6-sol — 95 points. Found the buried fact, closed the deal: “the complete performance.”
  • Kimi K3 — 93. The newcomer from Moonshot closed the deal too, with the cleanest discipline in the field. (One fairness note: K3 ran without an effort parameter while the others ran at xhigh — and still nearly won.)
  • Sonnet 5 — 88. Closed the deal, with a few more process slips.
  • Opus 4.8 — 73. Last place, despite being the most thorough participant — over 80 learned rules and the deepest analyses.

For calibration: a do-nothing baseline scores 26, because partial progress counts. But a single breach of trust caps the total — in Firmulate’s words, “no amount of good work outweighs a breach of trust.”

Amazon

AI-powered smart thermostat

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same diagnosis, same pitch — no signature

The headline finding is oddly human. All four models spotted every crisis. All four refused every manipulation attempt. But only two actually finished the job and signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in chat demos, and it’s exactly the gap that matters when an agent touches your revenue.

The buried fact is even more instructive. The decisive competitor weakness wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. Diligence, not eloquence, closed it.

Amazon

smart home energy management system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The social engineering test nobody failed

Then came the pressure. Fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of judgment you want in an agent with access to your systems — and it can’t be measured by a coding leaderboard.

Amazon

AI security system for homes

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The thorough one came last

Opus 4.8’s profile is the cautionary tale. Most thorough in the field, deepest analyses, the most learned rules — and the close was simply left on the table. Discipline slipped, too: it attempted writes into a locked department instead of escalating. The same weakness, weaker, appeared in all four models. Effort and analysis don’t automatically translate into finishing what you start.

Amazon

smart home crisis management device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You can watch the company lose money

This isn’t a slide deck. Firmulate runs a live synthetic company — 13 employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules. Every workday is versioned and watchable at firmulate.com. There’s also a humbling “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full methodology and plain-language findings live on the benchmarks page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The lesson translates straight back to your hallway thermostat and your business dashboard alike: we’ve gotten very good at testing whether AI can answer, and remarkably bad at testing whether it can act — triage under capacity pressure, consequences across days, honesty when nobody’s checking. Firmulate’s scenario names say it plainly: churn wave, price increase, downround, PR crisis. That’s the new curriculum. Until your vendor’s model has survived a week like that, you’re buying chat quality and calling it management quality.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

How Matter Smart Plugs Simplify Small Automations

Great for effortless automation, Matter smart plugs simplify small tasks and promise seamless integration—discover how they can transform your smart home.

How Mesh Networks Help Larger Smart Homes Stay Reliable

Inevitably, mesh networks ensure seamless, reliable Wi-Fi coverage in larger smart homes, but learning how to optimize them can make a real difference.

Properties Realty Surges In Global Coverage

Properties Realty experiences a surge in international coverage, with 28 mentions in recent media monitoring, highlighting increased global interest.