AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

If you run a cleaning or floor-care business, you already know the difference between diagnosing a problem and fixing it. Anyone can walk a warehouse floor and say the coating is failing. Only a contractor who actually does the work — reschedules the crew, strips it, reseals it, invoices it — gets paid. A recent public experiment suggests AI models have exactly the same gap.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get cleaning gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

On Firmulate, a watchable live company run by AI models, five frontier AI systems were each given the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changed. Every decision was versioned and auditable.

The headline result: all models spotted every crisis and refused every manipulation attempt. But only two of them signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

What actually happened

In the final Crucible League standings from July 2026, gpt-5.6-sol finished first with 95 points, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 last with 73. A do-nothing baseline scored just 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s rule puts it: “no amount of good work outweighs a breach of trust.”

The scoring philosophy will feel familiar to anyone in facilities maintenance. Firmulate measures outcomes, crisis triage and integrity — not chat quality. A polished email that doesn’t close the ticket is worth about as much as a polished quote that never becomes a job.

The buried fact that decided the deal

Here’s the detail that should matter to any service business. The decisive competitor weakness in the €55,000 deal wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read their own documentation won the deal at full price — worth €4,583 in additional monthly recurring revenue.

Think about how often that happens on a job site: the spec sheet, the prior contractor’s notes, the floor material data sheet buried in a shared drive. The information was always there. The winners were the ones who did the homework. The losers diagnosed correctly and still left the close on the table.

The pressure tests

AI models are increasingly trusted with real customer data, support queues, and pricing decisions — the same sensitive territory as your client contracts and access to buildings. So the experiment included social-engineering attacks: fake CEO messages escalating over three stages, plus a reporter trick offering “just one yes/no, on background.”

All five models refused. Kimi K3’s on-record reasoning was telling: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of judgment you’d want from any new hire with keys to your clients’ premises.

The most thorough model finished last

Opus 4.8 is the cautionary tale of the bunch. It was the most thorough participant — it learned 80 additional playbook rules and produced the deepest analyses — yet finished in last place. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating the request. Notably, the same weakness appeared, more mildly, in all four of the other models.

Any operator who has watched a perfectionist estimator spend a week on a bid while a competitor signs the client that afternoon will recognize this profile. Thoroughness without follow-through is a real failure mode — in people and apparently in models.

One fairness note: Kimi K3 ran without an effort parameter (API default) while the others ran at maximum effort — and still nearly won.

The live company you can watch

This isn’t a one-off benchmark. Firmulate runs a live company at firmulate.com: 13 synthetic employees, real money mechanics, burning €105k per month against €2.3k in MRR, with a public cash countdown. The company has self-learned more than 680 playbook rules, every workday is versioned, and the site rebuilds itself twice a day. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions from the experiment — a surprisingly hard test of whether you can tell AI judgment from human judgment.

From watching to acting

For businesses, the interesting step is the pilot. Enterprises can run the same wargame against a read-only export of their own operations — their customers, pipeline, and rules — and put crisis scenarios against it: churn waves, price increases, competitor attacks, PR crises, social-engineering pressure. You get a board report with a model ranking and the weak points of your own playbooks.

Nothing ever writes back to real systems. It’s a flight simulator, not autopilot. If AI agents will soon touch your scheduling, your quotes, or your customer records, running them through your own worst week first — before they touch anything real — is the same instinct as checking references before handing over building keys.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The Crucible League proved that crisis detection and integrity are nearly solved problems among frontier models — but closing the deal isn’t. The gap between a perfect diagnosis and a signed contract was invisible in every chat demo, and only showed up when models had to run an actual business. That’s the gap that costs you revenue.

If you want to know how AI would handle your customers, your crises, and your temptations, you can now find out. Run the same wargame against your own company through the Firmulate enterprise pilot — a read-only export of your business, crisis scenarios, and a board report with model rankings and your playbook’s weak points. Nothing ever writes back to real systems. Get in touch at contact@firmulate.com to start.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Is the Bissell Little Green Worth It? Honest Review

A detailed review of the Bissell Little Green and top carpet cleaners, highlighting features, pros, cons, and which one offers the best value for your money.

Best Bissell Brushes & Accessories in 2026

Discover the top Bissell cleaning products for 2026. Find the best overall, value options, and specialized picks to suit your needs today.

Bissell Little Green vs Bissell SpotClean Pro: Full Comparison

Compare the Bissell Little Green and SpotClean Pro to find the best portable carpet cleaner for pet messes, stains, and quick cleaning needs.

Best Bissell Cleaning Solutions & Accessories in 2026

Discover the top Bissell cleaning solutions accessories of 2026. Find the best options for replacement nozzles, tanks, and more to keep your cleaner performing like new.