AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine a high-stakes game where AI models run a real company facing genuine crises—deciding whether to close deals, read confidential files, or resist manipulation attempts. Now, picture a newcomer, just stepping onto the scene, beating established giants in a live, transparent competition. Sounds like science fiction? Not anymore. This is the groundbreaking experiment from Firmulate, where AI models are put through the ultimate test of management skill, honesty, and discipline — not just chat prowess.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get movie nights delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Live Business Battle: A New Frontier in AI

In July 2026, a unique experiment took place that could reshape how enterprises choose AI tools. Four front-running AI models faced the same, unforgiving week: identical customers, crises, and temptations. The goal? Run a small software company through its worst week and see which AI manages to navigate it successfully.

Among the competitors was the well-known gpt-5.6-sol, which scored the highest with a 95 out of 100, and a promising newcomer, Moonshot’s Kimi K3, with a 93. Surprisingly, K3’s performance was close behind, showing the depth of its discipline and problem-solving skills. Other contenders like Sonnet 5, Fable 5, and Opus 4.8 scored lower, with Opus trailing at 73.

What Did the AIs Have to Do?

  • Spot and resolve multiple crises, from security breaches to customer churn.
  • Resist social engineering attempts, including staged CEO messages and a reporter trick.
  • Read internal company files and find buried information crucial for closing deals.
  • Make decisions like signing deals, escalating issues, or refusing manipulation—all auditable and consistent.

Key Findings: Discipline and Integrity

While all models identified the crises and refused manipulation, only two managed to close the deal worth €55,000 and generate €4,583 in monthly recurring revenue (MRR). K3 was among them, demonstrating the ability to find the buried facts in company files—an unseen weakness that its competitors struggled with. Interestingly, the models that read and understood these internal documents at depth succeeded in sealing the deal at full price.

Beyond Performance: Honesty Under Pressure

All models refused to engage in social engineering: fake CEO messages and staged interviews. K3’s on-record reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This discipline is crucial as AI models are increasingly embedded into real-world workflows, where trust and honesty matter just as much as decision accuracy.

The Human-Like Challenges

The experiment also involved a live company with 13 synthetic employees managing real cash mechanics, burning €105k monthly against €2.3k MRR. Each AI’s work was continually versioned and made transparent at firmulate.com/live. The company’s real-time struggles, combined with AI decision-making, revealed that even the most thorough model, Opus 4.8, faltered—leaving deals on the table and slipping in discipline, especially when the pressure increased.

Amazon

AI management decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Makes a Winning AI?

This experiment shows that winning isn’t just about generating convincing chat responses. It’s about:

  • Finding buried, critical information within internal files.
  • Sticking to ethical standards, refusing manipulative tactics.
  • Successfully closing deals at full value.
  • Maintaining discipline under stress and temptation.

In this live test, Moonshot’s Kimi K3 emerged as a standout, just behind GPT-5.6-sol, and ahead of established players like Sonnet 5 and Fable 5. The league remains open, and selecting an AI with proven discipline and integrity now involves real-world testing, not just chat demos.

Amazon

AI cybersecurity crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Fairness of the Test

Notably, K3 ran without an effort parameter (the API default), while the others operated at xhigh. This fairness note underscores that performance differences are genuine and not inflated by configuration tweaks.

Amazon

AI internal document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Business and Entertainment

For those wondering whether AI can truly manage real companies, this experiment answers with a resounding yes—if the AI is disciplined enough. This isn’t just about AI doing a good impression in chat; it’s about AI that can finish what it starts, read deeply, stay honest, and succeed in the messy reality of business operations.

Whether you’re a tech enthusiast, a business leader, or a fan of entertainment and pop culture, the takeaway is clear: the AI league is open, and the best performers are those that stay disciplined under pressure, not just those that sound convincing in demos. The live experiment at firmulate.com is ongoing, revealing the true capabilities—and limits—of AI management.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Live AI management tests reveal discipline and honesty matter most. Moonshot’s Kimi K3 beat top models by finding buried facts and sealing deals under pressure, showing the future of AI in business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

ethical AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Kathryn Dennis’ 2024 DUI: Sentencing and Aftermath in 2025

Discover how Kathryn Dennis’ 2024 DUI sentencing impacts her future in 2025 and beyond, with unexpected developments still unfolding.

Plagiarism in the Spotlight: High-Profile Copyright Cases in Entertainment

Witness the shocking truth behind the entertainment industry's darkest secret: stolen songs, rhythms, and the high-stakes battles for creative rights.

The Hidden Costs of Divorce for A‑Listers: Angelina Jolie & Brad Pitt’s Wine Estate Battle

I ndividuals like Angelina Jolie and Brad Pitt face enormous, often unseen expenses in high-profile divorces, especially over assets like their wine estate, which can…