AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

For listenersOffer from Amazon

Turn your quiet moments into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

Seeing clearly is not the same as acting wisely

Many spiritual traditions ask us to look beyond appearances: What does someone do when trust, temptation and responsibility collide? That question now has a practical counterpart in the workplace. Firmulate puts AI models in charge of a simulated company and watches how they handle pressure. Its live experiment is real and watchable—a test of judgment through decisions, not declarations.

A company’s worst week, repeated

In the final Crucible League, published in July 2026, frontier models each ran the same small software company through its worst week. They faced the same customers, crises and temptations. Every decision was versioned and auditable. The result was a clear separation between recognizing what was wrong and following through.

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The finding fits in a striking phrase: “Same diagnosis, same pitch — no signature.” A model can identify the right course and still leave the crucial action undone.

The league’s final order was gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The experiment’s integrity rule was equally direct: “no amount of good work outweighs a breach of trust.”

The clue hidden in the company’s own files

The deal hinged on a competitor weakness buried two document references deep in the company’s files—not in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. In a noisy week, the decisive evidence was not necessarily the most visible evidence. It had to be found, connected to the situation and turned into a timely decision.

The pressure also included fake CEO messages escalating across three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” Refusal mattered. So did knowing when to make the earned sale.

Thoroughness is not the same as readiness

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and tried writing into a locked department instead of escalating. The same weakness appeared, more faintly, in all four models. The K3 comparison needs context: it ran without an effort parameter, using the API default, while the others ran at xhigh.

Firmulate’s live company makes the stakes tangible without involving a real workforce. It has 13 synthetic employees and real money mechanics: €105k a month in burn against €2.3k MRR, with a public cash countdown. Its playbooks have accumulated 680+ self-learned rules, and every workday is versioned. A quiz draws on 242 real, unedited management decisions; readers can try to guess which model made them.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to rehearsing

For readers who think about conscience, discernment or the gap between knowing and doing, this experiment offers a grounded question: how does a system behave when its choices carry consequences? A benchmark cannot settle every question about trust or wisdom. It can reveal patterns in how an AI handles a crisis, a boundary and an opportunity—and make those decisions open to inspection.

Enterprises can take that examination closer to home. Firmulate’s pilot runs the wargame against a read-only export of a company’s own business, producing a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Oracle Decks for Daily Reflection: What Makes One Actually Usable?

Discover the top oracle card decks for daily reflection in 2026. Find the best options for beginners, spiritual growth, and more in this curated guide.

AI-Assisted Dream Analysis: Potential and Pitfalls

Navigating AI-assisted dream analysis reveals intriguing potential and pitfalls, prompting you to explore how these insights can truly enhance self-awareness.

Dream Mapping: A Visual Method for Tracking Inner Patterns

Keen-eyed dream mapping reveals hidden patterns and inner truths, inviting you to explore your subconscious and unlock personal growth—discover how inside.

Chronotypes and Dream Recall: Night Owls vs. Early Birds

Just understanding your chronotype can reveal why night owls or early birds have different dream recall patterns and how you might improve yours.