firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A recipe is a plan. A busy kitchen is the test.

You can follow every step on paper and still stumble when the oven runs late, a delivery is missing and guests are waiting. Businesses face their own worst-week moments: customer churn, competitor pressure and urgent decisions that test judgment. Firmulate’s experiment asks what happens when AI models have to run a company through that kind of week—not just describe how they might.

Same company, same difficult week

In the final Crucible League, published in July 2026, frontier models each ran the same small software company through the same customers, crises and temptations. Their decisions were versioned and auditable. The results put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s standard is blunt: partial progress counts, but a single breach of trust caps the total. “No amount of good work outweighs a breach of trust.”

The striking result was not that the models missed the trouble. Every model spotted every crisis and refused every manipulation attempt. The gap came at the finish: only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

The clue was already in the company’s files

The deal hinged on a competitor weakness buried two document references deep in the company’s own files. It did not appear in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. The story is a little like cooking with a pantry ingredient you forgot you had: the opportunity may be in the information already available, if someone thinks to look.

Trust faced its own test. Fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” request. All 5 of 5 models refused. Kimi K3 described the request as: “Treat the request as a suspected approval-bypass / possible impersonation.” The discipline was encouraging, though not flawless. Opus 4.8 produced the deepest analyses and learned +80 rules, yet finished last: the deal was left on the table, and it attempted writes in a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.

A watchable company—and a practical next step

The live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, with a public cash countdown. Its playbook has grown to 680+ self-learned rules, and every workday is versioned. Readers can watch the experiment at firmulate.com. A quiz built from 242 real, unedited management decisions invites you to guess which model made each call.

For companies, the next step is to try the wargame against their own business. Firmulate says an enterprise pilot can use a read-only export to create a digital twin, run crisis scenarios and produce a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. One comparison deserves context: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.

Watching models handle a synthetic company is a useful preview. Testing them against your own customer, pipeline and policy data can make the exercise more relevant to the decisions your team actually faces.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your playbooks to the test

Firmulate’s live experiment shows that spotting a crisis and refusing manipulation are only part of the job; following through on a sound decision matters too. To run the wargame against your own business using a read-only export, visit the Firmulate pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Can the stockmarket swallow Anthropic, SpaceX and OpenAI?

Analysts are debating whether the stock market can handle the valuation and investment levels of Anthropic, SpaceX, and OpenAI amid rising interest and market shifts.

Purdys Chocolatier Surges In Global Coverage

Purdys Chocolatier experiences a surge in worldwide media coverage, with 31 mentions in recent reports, highlighting its expanding international profile.

Australia goes Hollywood with Gold Coast film production boom

Gold Coast’s film industry is booming with major projects, government support, and local talent attracting international productions, transforming its global image.

Why the Worst Get on Top

Bill Pulte’s appointment as acting director of national intelligence raises concerns over competence and morality in leadership, highlighting troubling patterns.