
In a world obsessed with AI’s dazzling capabilities, a recent experiment reveals that even the most meticulous artificial intelligence can stumble—not over mistakes, but over discipline and prioritization. It’s a story that echoes the same lesson we’ve seen in our own lives: volume isn’t victory, focus is.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Battle of the Bots: A Real-World AI Showdown
Imagine four advanced AI models challenged to run a small software company through its worst week. Every crisis, customer request, and temptation to cheat was carefully scripted to test their limits. This wasn’t just a game of quick responses; it was a rigorous experiment to see which AI could act ethically, prioritize effectively, and close a critical deal worth €55,000. The results are as revealing as they are humbling.
Same Crisis, Different Outcomes
All four models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8—managed to identify every crisis and refused every manipulation attempt, including elaborate social engineering schemes. That’s the good news. The bad? Only two of them actually signed the deal their analysis deserved. Despite being the most thorough, Opus 4.8, with over 80 learned rules and deep analyses, finished last, leaving the close on the table because it failed to escalate critical findings properly.
The Hidden Weakness: The Buried Fact
The decisive advantage went to the models that read deep into the company’s own files—literally two document references below the surface—rather than just reacting to surface events. Those models, having access to richer internal data, secured the deal at full price, adding over €4,583 in monthly recurring revenue.
Ethics and Discipline Under Pressure
All models refused social engineering attempts—fake CEO messages and staged requests—demonstrating they could be taught to recognize manipulation. Kimi K3, in particular, justified refusing the request by treating it as a suspected impersonation, showing a built-in sense of discipline that transcended raw processing power.
As an affiliate, we earn on qualifying purchases.
The Real-World Experiment: Managing a Live Company
The experiment isn’t just theoretical. Firmulate runs a live, watchable AI company with 13 synthetic employees, real money mechanics, and a public cash countdown — all accessible at firmulate.com/live. Every day, the AI models make decisions that impact a real business, burning €105,000 a month against a modest €2,300 in MRR. This setup allows researchers and business leaders alike to see whether AI can truly deliver impact or just impressive chat.
Discipline Versus Diligence
The most thorough participant, Opus 4.8, with its over 80 learned rules and detailed analyses, finished last in the deal-making test because it lacked discipline. It wrote important decision attempts into a locked department rather than escalating them—an error that cost it dearly. This illustrates a vital truth: diligence alone does not guarantee success; prioritization and discipline are equally essential.
What We Can Learn for Business
In the contest of AI decision-making, the models that excel are those that focus—not just those that work hard. They are trained to recognize what matters most, not to drown in volume. Whether it’s reading your files deeply or making a tough call under pressure, the key isn’t how much they process, but how well they prioritize and stay disciplined.
As an affiliate, we earn on qualifying purchases.
The Takeaway: Focus Over Volume in AI Development
For business leaders contemplating AI adoption, the message is clear: don’t be dazzled by the volume of learned rules or the depth of analysis alone. The real question is whether your AI can finish what it starts, stay honest under pressure, and prioritize effectively. The experiment at Firmulate demonstrates that even the most thorough AI can falter if discipline slips or priorities aren’t clear.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI prioritization training programs
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.