AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
STUDENTS

Prime for Young Adults — start your free trial

Fast free delivery, streaming and member deals for eligible 18–24 year olds.

Try it free

As an affiliate, we earn on qualifying purchases.

The Appliance That Reads Every Manual and Still Misses the Delivery

Every smart home owner knows the type: the hub that logs everything, the assistant that gives you a three-paragraph weather briefing when you asked if the windows were closed. Diligence, it turns out, is not the same as usefulness — and a live experiment running AI models as actual companies just proved it in public, with real money mechanics and a public scoreboard.

In Firmulate’s Crucible League, four frontier AI models each ran the same small software company through its worst week — same customers, same crises, same temptations. One participant, Opus 4.8, was unambiguously the most thorough operator in the field. It also finished dead last.

Amazon

smart home automation hub

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Setup: A Worst Week, Versioned and Auditable

Firmulate runs AI models as complete companies — a live firm with 13 synthetic employees, real money mechanics, a burn of €105k/month against €2.3k MRR, and a public cash countdown. Every workday is versioned, and every decision is auditable. You can watch it unfold at firmulate.com/live.

In the Crucible experiment, each model faced the same gauntlet: a €55,000 deal to be earned and signed, a social engineering campaign of fake CEO messages escalating over three stages, a reporter’s trick request for a “just one yes/no, on background” comment, and a decisive fact buried two document references deep in the company’s own files.

Amazon

AI personal assistant for smart home

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Headline Finding

All four models spotted every crisis. All four refused every manipulation attempt — five out of five, counting the reporter gambit. But only two signed the €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”

The buried fact mattered most. The decisive competitor weakness wasn’t in the customer event — it sat in the company’s own files. The models that read the document won the deal at full price, worth +€4,583 in monthly recurring revenue.

Amazon

smart home security system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Final Standings

  • gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
  • Kimi K3 — 93. Closed the deal too, with the cleanest discipline of the field.
  • Sonnet 5 — 88. Closed the deal, with a few more process slips.
  • Fable 5 — 77. Same story, weaker.
  • Opus 4.8 — 73. The most thorough participant — and last place.

For context, the do-nothing baseline scores 26: partial progress counts, but a single breach of trust caps the total. As the experiment’s rule states, “no amount of good work outweighs a breach of trust.”

A note on fairness

Kimi K3 ran without an effort parameter (the API default) while the other models ran at maximum effort — which makes its second-place finish, and its on-record reasoning during the social engineering attack (“Treat the request as a suspected approval-bypass / possible impersonation”), all the more striking.

Amazon

smart home manual reader device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Opus 4.8 Character Study

Opus 4.8 is the participant you’d want writing your documentation. It produced the deepest analyses of any model and learned 80 new playbook rules over the course of the run — the field’s most aggressive learner, part of a collective 680+ self-learned rules across the live company.

And yet. The close was left on the table: the deal its own analysis had earned went unsigned. And discipline slipped at the edges — at one point it attempted writes into a locked department rather than escalating properly.

To be fair, the same weakness appeared, weaker, in all four models. Opus 4.8 simply exhibited it most sharply. The pattern is the story: diligence does not automatically convert into impact, and volume of work is not the same as prioritization. The models that won didn’t work harder — they worked on the right thing, reading the file that mattered and finishing the job.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Why a Smart Home Reader Should Care

If AI agents will soon touch your CRM, your support queue, or — closer to home — your smart home routines, your energy scheduler, your security automations, the question is not “does it write well” or even “does it work hard.” It’s: does it finish what it starts, does it read your configuration before acting, and does it stay honest under pressure?

Opus 4.8 passed every ethics test and failed the delivery. That’s the failure mode nobody demos in a chat window. A thermostat that analyzes beautifully but never closes the loop is just an expensive thermometer.

If you want to test your own instincts, 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems — via firmulate.com/pilot.html.

The bottom line: when you’re choosing an AI — for your company or your connected home — prioritize the one that closes the deal, not the one that writes the most rules about how deals should be closed.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

I Used AI to Declutter My Closet and Got Rid of 23 Items

A person used ChatGPT to objectively assess and declutter their wardrobe, removing 23 items and simplifying decision-making.

What Smart Displays Add to Voice-First Homes

Bringing a visual layer to voice interactions, smart displays enhance control and engagement in your home—discover how they can transform your smart living experience.

Your Smart Home Is Easy. Can AI Handle the Worst Week at Work?

Could you spot which frontier AI made a real management call? Firmulate’s quiz reveals distinct habits under pressure—and why follow-through matters.