firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your wardrobe favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A collection can be on schedule and still miss the market

In fashion, a difficult week can bring several pressures at once: a key customer wavering, a competitor moving into view, and a commercial offer that needs a decision. An AI assistant may describe the risks clearly. The harder question is whether an AI workforce can carry that judgment through to action, while respecting the rules that protect a business.

Firmulate’s live experiment puts that question into a small software company. The aim is relevant well beyond software: see how AI models handle the same crises and temptations before giving them a role in decisions that matter. The experiment is real and watchable at Firmulate.

One company, one difficult week

In the final Crucible League, published in July 2026, frontier models faced the same customers, crises and temptations. Every decision was versioned and auditable. The final order was gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s stated principle is that partial progress counts, but one breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The result was striking for what the models shared as well as where they differed. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The experiment’s summary captures the gap: “Same diagnosis, same pitch — no signature.” Recognizing the right move is not the same as making it.

The detail hidden in the files

The decisive competitor weakness was not in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. That finding points to a practical challenge for any company considering AI: useful context may be present in its own information, but an agent still has to find and act on it.

For a fashion business, the parallel might be a customer or supplier issue whose explanation sits across documents rather than in the latest message. Firmulate’s test does not claim to measure fashion operations. It does show why it matters to evaluate whether an AI can connect evidence, make a sound recommendation and follow through under pressure.

Trust is part of performance

The experiment also put models through fake CEO messages escalating over three stages, followed by a reporter’s “just one yes/no, on background” request. All five refused. Kimi K3 explained its judgment on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 offers a more complicated picture of performance. It was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and discipline slipped: it attempted writes into a locked department instead of escalating. Firmulate says the same weakness appeared, more weakly, in all four models. A strong analysis alone did not guarantee a strong business outcome.

There is a fairness detail in the league: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The ranking is a snapshot of this experiment, not a universal verdict on which model is best for every company or task.

From watching to testing your own business

Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k a month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules, and every workday versioned. A separate quiz draws on 242 real, unedited management decisions and asks visitors to guess the model. These features make the experiment watchable; a pilot offers a way to examine how the same kind of challenge applies to an enterprise’s own business.

That pilot uses a read-only export to create a digital twin, then runs crisis scenarios against it and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. For leaders in fashion and other industries, that creates a way to examine how AI handles their company’s context before entrusting it with live work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Test judgment before handing over responsibility

Firmulate’s league shows a divide between spotting the right answer and carrying it through, alongside a clear test of whether models respect trust boundaries. A pilot lets enterprises put their own data and playbooks into that kind of wargame using a read-only export. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Powerball Drawing

Latest Powerball drawing completed with no jackpot winner; jackpot resets to $X million. Details on winning numbers and upcoming draw.

Philippine Leroy-Beaulieu Is The New Face Of Haus Labs’ Beloved Foundation

Philippine Leroy-Beaulieu has been announced as the new ambassador for Haus Labs’ popular foundation, marking a strategic brand partnership.

Ralph Lauren Clears 2030 Climate Goal With Lower Production In The Mix

Ralph Lauren reports achieving its 2030 climate target by lowering production levels, marking a shift in its sustainability strategy amid rising industry interest.

German ruling declares Google liable for false answers in AI Overviews

Munich court rules Google responsible for false claims made by its AI-generated search overviews, marking a shift in liability for AI content.