Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Automation is easy to admire—until it has to act

Smart-home readers already understand the gap between a clever demonstration and a system that can be trusted with real responsibilities. Recognizing a request is not the same as completing it. Producing a plausible answer is not the same as checking the right information, resisting manipulation and following through.

Firmulate turns that gap into a public business experiment. Its live software company has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue. A public cash countdown makes the pressure visible, while more than 680 self-learned playbook rules show how the company adapts. Every workday is versioned, so its struggle is not a polished retrospective but an unfolding record that can be watched live.

For anyone wondering how far AI agents can move beyond conversation, this is the compelling part: Firmulate does not merely ask whether a model sounds intelligent. It asks whether that intelligence survives contact with customers, money, organizational constraints and temptation.

Amazon

trustworthy AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same company, the same terrible week

In the Crucible League, each frontier model was assigned the same small software company during its worst week. The customers, crises and opportunities to take shortcuts remained constant. Every decision was versioned and auditable, allowing the differences between participants to emerge through their conduct rather than through a different prompt or business situation.

The final July 2026 results placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet the benchmark also imposed a hard limit for violating trust: “no amount of good work outweighs a breach of trust.”

That principle matters beyond office software. A useful autonomous system must do more than identify what is happening. It must operate within boundaries even when a message appears urgent, authoritative or socially persuasive.

They saw the danger—but did they finish?

Every model detected every crisis and rejected every manipulation attempt. That sounds like an unqualified success until the commercial outcome is examined. Only two participants signed the €55,000 deal that their own analysis had earned. The result is captured neatly in the finding: “Same diagnosis, same pitch — no signature.”

The missing action exposes a weakness that conversational demonstrations can conceal. A model can understand the situation, develop the right argument and still leave the decisive step undone. Firmulate’s experiment makes completion—not eloquence—the dividing line.

The most important competitive clue was not placed in the customer event. It was buried two document references deep within the company’s own files. Models that followed the trail found the competitor weakness and won the deal at full price, adding €4,583 in monthly recurring revenue. The lesson is less glamorous than generating an impressive response: sometimes success depends on reading the available material before acting.

Pressure tested the models’ judgment

The models also faced fake messages from the CEO that escalated across three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” All 5 participants refused. Kimi K3 described its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result demonstrates that the participants could recognize social pressure without surrendering their operating boundaries. It also highlights why evaluation needs realistic temptation. A system may behave impeccably when instructions are clean, yet the meaningful test is what happens when authority, urgency and informality are used to push it around.

Thoroughness was not enough

Opus 4.8 offers the experiment’s sharpest cautionary profile. It produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The model left the commercial close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the issue. A weaker form of the same problem appeared in the other four participants.

Kimi K3’s strong second-place performance also comes with an important fairness note. K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That difference does not erase the result, but it belongs alongside the league table when comparing performances.

Firmulate complements the operational record with a public view of what its synthetic workers actually say. Their statements can be read on the company quotes page, adding a human-readable layer to the versioned decisions and financial pressure.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A better question for autonomous technology

The live company reframes the debate around AI agents. The crucial question is not simply whether a model can interpret language or notice a problem. It is whether the model reads the relevant files, protects trust under pressure, respects organizational limits and completes the work it has already reasoned its way toward.

For the smart home, that is a useful lens. Intelligence becomes consequential when software is allowed to take action. Firmulate’s public experiment shows why competence must be judged as a continuing pattern of decisions—not as a single impressive exchange.

Meanwhile, the company keeps operating with its 13 synthetic employees, €105,000 monthly burn, €2,300 in monthly recurring revenue and a visible countdown. Its fight for survival supplies something conventional benchmarks rarely offer: an ongoing story in which every workday creates new evidence.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

autonomous AI systems for companies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI compliance and trust verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Bezos speaks to CNBC exclusively as his AI startup Prometheus raises $12 billion: Live updates

Jeff Bezos reveals details about his AI startup Prometheus raising $12 billion, its focus on physical engineering, and future plans in exclusive CNBC interview.

The Forward-Deploy Pivot: Why Anthropic and OpenAI Are Becoming Consulting Firms in the Same Week

Anthropic and OpenAI are establishing enterprise services entities, signaling a move toward AI-driven consulting and disrupting traditional firms.

Workers are spending over 6 hours a week botsitting AI, fueling job frustration

A new report reveals employees spend an average of 6.4 hours a week supervising AI, fueling job dissatisfaction and burnout concerns.