AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine managing your outdoor adventure company with an AI that not only plans expeditions but also handles crises, keeps its promises, and reads the hidden clues in your files—without slipping up. As outdoor professionals navigate unpredictable environments, the performance of AI management tools is becoming just as crucial behind the scenes. The latest experiment by Firmulate reveals that some AI models are ready for the challenge, while others still stumble when stakes are high.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get travel and outdoor gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Crucible of AI Management: Testing Under Real-World Pressure

On July 2026, a unique test was conducted on five leading AI models to determine their ability to run a small software company through its most challenging week. This was no ordinary test. Every model faced the same simulated crises—customer churn, security breaches, manipulative social engineering attempts—all designed to mimic real management dilemmas. Each decision was fully documented and auditable, ensuring transparency and fairness in the comparison.

The Results: The Leaders and the Laggards

  • gpt-5.6-sol scored highest at 95, finding a buried security fact deep in internal files, securing a €55,000 deal (+€4,583 MRR), and resisting all manipulative tricks.
  • The newcomer, Kimi K3 by Moonshot, scored just slightly below at 93. Despite being new, K3 demonstrated the cleanest discipline, spotting critical hidden information, and securing the same deal.
  • Sonnate 5 finished with 88, with some process slips, while Fable 5 scored 77, and Opus 4.8 lagged at 73. Interestingly, Opus, despite its thorough analysis, left a close deal on the table due to discipline lapses.

The Hidden Weakness: Reading the Files

Crucially, the decisive advantage for Kimi K3 and gpt-5.6-sol was their ability to delve into internal company files—a step that many models overlook. This buried information was essential to closing the deal at full price, demonstrating that reading and interpreting internal data is a key differentiator.

Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Beyond the Demos: Real Business, Real Risks

The experiment was run on a live, functioning company with 13 synthetic employees and real money mechanics—burning €105,000 monthly against a mere €2,300 MRR. Every day, the models made decisions, which were recorded and analyzed, providing a window into how AI might perform in actual business environments at firmulate.com.

The Human-Like Tests: Social Engineering and Trust

In scenarios designed to test trust, all models refused to act on fake CEO messages or background requests from a reporter—showing they can recognize social engineering attempts. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This discipline under pressure is vital for AI systems in sensitive roles.

The Discipline Difference: A Narrow but Crucial Edge

Interestingly, the most thorough participant, Opus 4.8, with over 80 learned rules and deep analyses, finished last in the league. It left close deals unclosed and shifted decisions into locked departments instead of escalation, showing that thoroughness alone doesn’t guarantee discipline or effectiveness in critical moments.

Amazon

internal data analysis AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Fairness Note and What It Means for Business

It’s worth noting that Kimi K3 was run without an effort parameter (the default API setting), while the others operated at a higher effort level (xhigh). Despite this, K3’s performance was among the best, underscoring its efficiency and reliability.

Amazon

AI cybersecurity and social engineering detection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Does This Mean for Outdoor and Travel Businesses?

For outdoor and adventure companies, the lesson is clear: AI tools that understand your internal data, resist manipulation, and make disciplined decisions are the ones ready to support high-stakes management. Whether it’s coordinating expeditions, managing client relationships, or handling emergencies, choosing an AI that can read between the lines—and stay honest under pressure—is crucial.

Try the Wargame Yourself

If you’re curious, businesses can run their own simulations against a read-only export of their operations. This way, they can test how AI might perform without risking real systems or data. Visit firmulate.com/pilot.html to learn more about running your own management wargame.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The latest AI management test shows that models capable of thorough internal data analysis and disciplined decision-making outperform others in high-pressure scenarios. For outdoor and travel firms, choosing an AI that stays honest, reads deeply, and completes tasks reliably could be the key to smarter, safer operations.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

business AI simulation platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What a Do-Nothing AI Benchmark Reveals About Trust and Performance

Discover what a do-nothing AI score reveals about trust, discipline, and performance in high-pressure situations—less about cleverness, more about integrity.

History and Culture of Africa

Keen explorers will find Africa’s rich history and vibrant cultures full of stories that shape our world—discover what makes this continent truly extraordinary.

AI Models Pass Social Engineering Test, Reinforcing Trust in Automation Security

Top AI models successfully resisted simulated social engineering attacks in a live test, demonstrating their potential to uphold security and integrity before real crises occur.

Cape Verde’s Music of Morna: UNESCO Status and Cultural Revival

Nurturing Cape Verde’s morna through UNESCO recognition reveals its cultural importance and ongoing revival efforts that you won’t want to miss.