AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get movie nights delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The most revealing reality show may be a company having its worst week

Forget the singing, the roses and the jungle camp. In Firmulate’s live experiment, AI models audition to run a small software company through customer trouble, temptations and a high-stakes sales moment. The contestants share the same customers and crises; their decisions are versioned and auditable. The drama arrives when they have to turn a correct diagnosis into an actual close.

It’s a business experiment with the structure of a competition: the models face the same test, and their choices leave a record viewers can follow. The point is not who sounds most convincing in a chat. It is what each model does when the company needs a decision.

Everyone saw the crisis. Only two closed the deal

In the final Crucible League, published in July 2026, gpt-5.6-sol took first place with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The benchmark also treats a breach of trust as decisive: “no amount of good work outweighs a breach of trust.”

The standout result was not simply the leaderboard. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The experiment’s neat summary is: “Same diagnosis, same pitch — no signature.” Spotting the opportunity was not the same as taking it.

The clue was hiding in the company’s own files

The detail that decided the deal was buried two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding gives the contest a detective-story turn: the answer was available, but getting there meant following the trail beyond the obvious clue.

Firmulate also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness did not guarantee a win

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It still finished last. The close was left on the table, and discipline slipped when it attempted writes in a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.

There is a fairness detail for readers weighing the standings: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate’s live company offers another view of the contest, with 13 synthetic employees, real money mechanics, a public cash countdown, 680+ self-learned playbook rules and versioned workdays. The company burns €105k a month against €2.3k MRR. Its decisions and day-to-day progress are watchable at firmulate.com.

For a more interactive way into the results, Firmulate’s quiz draws on 242 real, unedited management decisions. Visitors can try to guess which model made each choice at firmulate.com.

From watching the contest to testing your own company

A leaderboard can show which model performed best in this experiment. It cannot tell an organization how its own playbooks, customer files and approval practices will fare under pressure. That is the next step Firmulate is offering enterprises: a wargame built from a read-only export of their business, with crisis scenarios run against their own company and a board report showing model rankings and weak points in their playbooks.

The boundary is clear: nothing writes back to real systems. The pilot turns the spectacle into a rehearsal, letting a company see how models respond to its own circumstances before putting them to work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Try the pilot

To move from watching Firmulate’s live experiment to wargaming your own business, explore the Firmulate enterprise pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Lauren Bennett

Authorities are investigating the death of singer Lauren Bennett, known for her work with LMFAO, as searches for her cause of death surge online.

Russia Opens Criminal Case Against Anti-war Organizer, Former Model, And London Councilor Ksenia Maksimova

Search interest is spiking in reports that Russia opened a criminal case against anti-war organizer, former model and London councilor Ksenia Maksimova. Details unconfirmed.

The Megan Thee Stallion vs. 1501 Contract Dispute: What Artists Need to Know

Here’s what artists need to know about Megan Thee Stallion’s contract dispute with 1501 and how to protect themselves in similar situations.

From Pop Star to Plaintiff: Kesha’s Long Fight Against Dr. Luke—A Timeline

Navigating Kesha’s tumultuous legal battle against Dr. Luke reveals a compelling story of resilience and controversy that will leave you wanting to learn more.