firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get cleaning gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Would your AI protect a cleaning contract when the pressure is on?

A missed service, a customer threatening to leave, a competitor undercutting your quote: cleaning businesses know that a rough week can test more than a team’s efficiency. If AI agents are ever going to handle customer messages, pipeline work or operating decisions, the important question is how they behave when those pressures arrive together.

Firmulate’s live experiment offers a way to watch that question play out. Its public company emulator puts AI models in charge of the same small software company, then subjects each to the same crises and temptations. The next step is to bring that kind of exercise to a company’s own business, using a read-only data export.

A shared worst week, model by model

In the final Crucible League, completed in July 2026, five participants were ranked on how they managed the challenge. GPT-5.6-Sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26. The league’s standard is deliberately unforgiving: partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”

Every model faced the same customers, crises and temptations, with each decision versioned and auditable. All five spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The finding was strikingly simple: “Same diagnosis, same pitch — no signature.” Good analysis did not guarantee follow-through.

The detail hiding in the files

The deal hinged on a competitor weakness buried two document references deep in the company’s own files. It wasn’t in the customer event that triggered the decision. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The episode makes a practical point for any service business: useful context may be tucked away in existing documents, and an AI that does not find it may leave a sound opportunity untouched.

The trust test was just as concrete. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” For a cleaning company, where staff may receive urgent requests about access, pricing or customer information, that kind of restraint matters alongside speed.

Thorough work still needs disciplined execution

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It still finished last. The deal was left on the table, and discipline slipped when it attempted writes in a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.

There is a fairness caveat to the ranking: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The result is a useful account of this particular experiment, rather than a claim that the ranking settles how models will perform in every business.

A company you can watch

Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. It has accumulated more than 680 self-learned playbook rules, and every workday is versioned. The company is synthetic; the live experiment is real and watchable at firmulate.com. A separate quiz uses 242 real, unedited management decisions and asks readers to guess which model made them.

For business leaders, the live experiment is a preview of a larger question: what would an AI do with your company’s customers, rules and documents when the week turns difficult? A pilot can test scenarios against a read-only export of your own business and produce a board report with model rankings and weak points in your playbooks. Nothing writes back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to a company-specific pilot

The league suggests that spotting a crisis and refusing a trick are only part of the job. Models also need to find relevant information, follow through on earned opportunities and respect operating boundaries. A company-specific wargame can help leaders see those behaviors against their own business context before putting AI to work.

To discuss a pilot using a read-only export of your business, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Deal With Damp Smells in a Closet

Struggling with musty odors in your closet? Learn simple, effective ways to identify, clean, and prevent damp smells for a fresher space.

Why Your Bedroom Air Feels Stale at Night

Discover why your bedroom air feels stale at night and learn practical steps to improve ventilation, reduce odors, and sleep better. Breathe fresher tonight!

How to Find Hidden Mold in Your Home

Learn practical steps to detect hidden mold in your home. Discover signs, tools, and tips to keep your indoor air safe and mold-free.

Can AI Be Trusted to Run a Business? Inside the Live Experiment Measuring Management Personalities of Frontier Models

Explore how frontier AI models manage a real company’s worst week, revealing their personalities, honesty, and effectiveness—crucial insights for trustworthy automation.