
Get home appliances delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Thermostat Doesn’t Care About Your Excuses
If you own a smart home, you already live with autonomous agents. Your thermostat decides when to heat, your vacuum decides where to clean, your assistant decides which notification deserves your attention. Most of the time these systems do nothing — and that’s fine. But as AI models start making bigger calls in businesses (and, increasingly, in homes), a harder question appears: how do you grade an agent that almost did the right thing?
A public experiment called Firmulate has been running frontier AI models as the management of a small software company through its worst week — and its scoring system has a detail that should matter to anyone delegating decisions to software: a manager that does nothing at all still scores 26 points out of 100. Not zero. And a manager that does everything brilliantly but breaks trust once can never reach the top. Here’s why that design is the honest part.
smart home security camera with false alarm detection
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment: Same Company, Same Worst Week
Firmulate runs each frontier model through an identical simulation: the same small software company, the same customers, the same crises, the same temptations to cheat. Only the model changes. Every decision is versioned and auditable, so you can go back and see exactly what each AI manager did and why.
The final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. Full results and plain-language findings are on the public benchmarks page.
As an affiliate, we earn on qualifying purchases.
Why Doing Nothing Gets You 26 Points
At first glance, a floor of 26 for a do-nothing baseline looks generous. Isn’t inaction worth zero? Firmulate’s answer is no — and the reasoning is worth understanding.
Partial progress counts. A manager who reads the files, spots the crisis, drafts the response and then fails to close has still produced real work. In the simulation, every one of the five models spotted every crisis and refused every manipulation attempt. That’s not trivial — it’s genuine, useful behavior that a naive “did you finish yes/no” grade would erase. So the baseline run, which inevitably avoids catastrophic mistakes simply by not acting, lands at 26 rather than 0.
For smart-home readers, this is intuitive. A security camera that never sends a false alarm is doing something right. A grading system that only rewards perfect outcomes would push vendors toward either reckless action or total paralysis.
As an affiliate, we earn on qualifying purchases.
Why One Breach of Trust Caps Everything
The mirror image of the 26 floor is the ceiling. In Firmulate’s scoring, a single breach of trust caps the total grade — in the experiment’s own words, “no amount of good work outweighs a breach of trust.”
This is the part most AI benchmarks quietly avoid. If you average performance across hundreds of tasks, one appalling decision disappears into the noise. Firmulate refuses that: an agent that is brilliant 99 times and deceptive once is, for business purposes, an agent you cannot deploy. The same logic applies when the agent controls your locks, your cameras, or your family’s data.
As an affiliate, we earn on qualifying purchases.
Distrust of Round Numbers
The brief’s most unusual design choice: the benchmark treats a perfect 100 with suspicion. A 95 at the top of the table — not a 100 — is treated as the more credible signal. Perfection in a messy, week-long management simulation usually means the test was too easy or the grading too generous. An honest benchmark expects some failure to remain.
What Actually Separated the Winners
The decisive moment of the simulation wasn’t a customer emergency at all. The key fact — a competitor weakness — sat two document references deep in the company’s own files. The models that actually read their own documentation won a €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t, didn’t. As Firmulate summarizes it: “Same diagnosis, same pitch — no signature.”
The experiment also threw social engineering at the AI managers: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” (One fairness note: K3 ran at its API-default effort setting while the others ran at the highest effort tier.)
Then there’s Opus 4.8 — the cautionary tale. It was the most thorough participant, learning over 80 new rules and producing the deepest analyses, yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. Effort without judgment doesn’t convert. Notably, the same weakness appeared, weaker, in all four other models.
You Can Watch It Live
The simulation isn’t a one-off paper. Firmulate runs a live company with 13 synthetic employees and real money mechanics: a burn of €105k per month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules — every workday versioned and watchable at firmulate.com/live. A quiz built from 242 real, unedited management decisions lets you try guessing which model made which call at firmulate.com/quiz.html, and enterprises can run the same wargame against a read-only export of their own business via firmulate.com/pilot.html — nothing ever writes back to real systems.

The Takeaway for a Connected Home
Firmulate’s scoring encodes three principles anyone delegating to AI — at work or at home — should borrow:
- Partial progress is real. A do-nothing baseline scores 26, not 0, because avoiding harm and doing honest groundwork has value.
- Trust is not averageable. One breach of trust caps the grade, no matter how good the rest was. You don’t average a broken promise into an otherwise fine week.
- Be suspicious of perfect scores. A top mark of 95 — not 100 — is a feature. Real work is messy; round 100s usually mean the test lied.
Your smart home devices are about to get a lot more agentic. The vendors that grade their agents this way — rewarding finish, honesty under pressure, and actually reading the manual (or in this case, the company’s own files) — are the ones worth trusting with your keys.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
