AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Imagine an AI travel assistant handling a mountain lodge’s busiest week: a supplier fails, guests threaten to cancel, and a tempting message asks it to bend the rules. A polished itinerary is easy to admire. The harder question is whether an AI can protect customers, make a sound call under pressure and follow through when money is on the line. Firmulate has built a live experiment around that question.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get travel and outdoor gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

From watching to trying it on your business

Travel businesses run on decisions made across reservations, customer support, staffing and cash flow. A disruption in one place can ripple through the rest. Firmulate’s experiment puts AI models through a small software company’s worst week, giving each the same customers, crises and temptations. Every decision is versioned and auditable, and the public live company makes the broader experiment watchable.

The results offer a useful distinction for any business considering AI agents: spotting a problem is not the same as completing the job. Every model in the final July 2026 Crucible League spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.”

The detail buried in the files

The deal turned on a competitor weakness hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding is a reminder that business judgment can depend on connecting scattered information, not merely reacting to the latest customer message.

Firmulate also tested social engineering: fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work still has to land

The league ranking shows how varied the outcomes were: gpt-5.6-sol scored 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. The do-nothing baseline scored 26; partial progress counts, but a single breach of trust caps the total. The rule behind that cap is plain: “no amount of good work outweighs a breach of trust”.

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh—a relevant caveat when comparing results.

The live company has 13 synthetic employees, real money mechanics, burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a versioned record of every workday. A separate quiz draws on 242 real, unedited management decisions and invites visitors to guess which model made them.

A rehearsal using your own business

For a travel operator, the value of a wargame is practical: see how an AI might handle a churn wave, a price increase, a competitor move or a public-relations crisis before relying on it during a real season. Firmulate’s enterprise pilot uses a read-only export of your business to run crisis scenarios and produce a board report with model rankings and weak points in your playbooks. Nothing writes back to real systems.

The public experiment moves the conversation from what AI says to what it does under pressure. A pilot brings that same test to the decisions, data and operating rules of your own company.

To discuss an enterprise pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Firmulate’s live experiment suggests that safe AI behavior is only part of the test: agents also need to find the decisive information and carry a good decision through. Enterprises can run a wargame against a read-only business export, with nothing writing back to real systems. Explore a pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Great Green Wall: Measuring Africa’s Ambitious Anti‑Desertification Project

Will you discover how Africa’s Great Green Wall is transforming landscapes and combating desertification through innovative, community-driven efforts?

South Padre Island, Texas, United States Surges In Global Coverage

South Padre Island, Texas, experiences a surge in worldwide media coverage, with 43 mentions in recent reports, highlighting increased international interest.

Cape Verde’s Music of Morna: UNESCO Status and Cultural Revival

Nurturing Cape Verde’s morna through UNESCO recognition reveals its cultural importance and ongoing revival efforts that you won’t want to miss.

Zion National Park, Utah, United States Surges In Global Coverage

Zion National Park in Utah experiences a significant increase in international media mentions, with 38 reports within a recent timeframe, highlighting rising global interest.