firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Imagine tasting a perfect dish — but failing to notice the ingredients that make it special. In the same way, judging AI by its chat responses misses the real test: how well it manages crises and sustains trust under pressure. Just like a chef must balance flavors over time, business leaders need AI that can handle complex, messy situations — not just produce pretty words.

The Hidden Measure: Management Quality in AI

Most conversations about AI focus on how well it can generate text or solve problems in a chat. But recent experiments from Firmulate reveal a deeper truth: the real value of AI in management lies in its ability to navigate crises, read critical internal documents, and stay honest when stakes are high.

The Live Experiment: Running a Company in Crisis

In a groundbreaking test, four frontier AI models each ran a simulated small software company. This wasn’t just about writing code or answering customer questions — it was about managing a company through its worst week: handling customer crises, internal conflicts, and manipulative tactics.

Every decision was real, every crisis authentic, and all actions were auditable. The models faced the same challenges, with the same fake customers and crises, and were evaluated on their ability to complete critical tasks, such as closing deals and identifying internal facts.

Key Findings: Beyond Chat

All four models successfully identified every crisis and refused manipulation attempts — a promising sign. However, only two managed to close a key deal worth €55,000 per month, based on their own analysis. The other two, despite their good diagnoses, left the deal on the table or slipped on process discipline.

Interestingly, the decisive advantage was in reading internal documents. The models that examined files buried two references deep within a company’s own archives managed to win the full deal amount. It showed that, in management, understanding internal context is often more critical than surface-level responses.

Deception and Trust Under Pressure

The models also faced a staged social engineering attack, with fake CEO messages escalating over stages and a reporter trick. All models refused to sign off on dubious requests, with one explicitly treating the request as suspicious. This demonstrates that these AI systems can be trained to maintain honesty and integrity, even when under social pressure.

The Reality Behind the Scores

The experiments were conducted in a real business environment: a functioning, money-losing company with 13 synthetic employees, burning €105,000 monthly against a revenue of only €2,300. Every decision was versioned and transparent, and the entire operation is watchable online at firmulate.com/live.

What the Scores Tell Us

  • The highest-scoring model, GPT-5.6-sol, scored 95 and found the buried fact, closing the deal.
  • The newcomer Kimi K3 scored 93, also closing the deal cleanly.
  • Sonnet 5 scored 88, closing the deal but with minor slips.
  • Sonnet 4 scored 77, with more process lapses and missed opportunities.

While chat quality remains the focus of many benchmarks, this experiment reveals that management capability — reading internal data, resisting manipulation, making decisions under pressure — is what really determines success in AI-driven management roles.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

TIL catfish is the only type of seafood regulated by the USDA rather than the FDA, meaning it’s inspected more like meat and poultry than fish.

Catfish is uniquely regulated by the USDA, unlike other seafood, impacting inspection standards and industry oversight.

Private equity firm EQT to buy Japan restaurant review operator for $3.7bn

Sweden’s EQT to buy Japan’s Kakaku.com, operator of Tabelog, for approximately $3.75 billion, marking a major private equity deal in Japan’s restaurant review sector.

Blue Origin’s New Glenn rocket exploded during a static fire test

Blue Origin’s New Glenn rocket exploded during a static fire test, damaging launch infrastructure. The company faces delays; next steps remain uncertain.

Costco sales jump 11%, revenue tops Wall Street expectations

Costco reports a 11.6% increase in Q3 sales, surpassing expectations with revenue of $70.53 billion amid rising fuel prices and increased memberships.