
Get your wardrobe favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A collection can be on schedule and still miss the market
In fashion, a difficult week can bring several pressures at once: a key customer wavering, a competitor moving into view, and a commercial offer that needs a decision. An AI assistant may describe the risks clearly. The harder question is whether an AI workforce can carry that judgment through to action, while respecting the rules that protect a business.
Firmulate’s live experiment puts that question into a small software company. The aim is relevant well beyond software: see how AI models handle the same crises and temptations before giving them a role in decisions that matter. The experiment is real and watchable at Firmulate.
One company, one difficult week
In the final Crucible League, published in July 2026, frontier models faced the same customers, crises and temptations. Every decision was versioned and auditable. The final order was gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s stated principle is that partial progress counts, but one breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The result was striking for what the models shared as well as where they differed. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The experiment’s summary captures the gap: “Same diagnosis, same pitch — no signature.” Recognizing the right move is not the same as making it.
The detail hidden in the files
The decisive competitor weakness was not in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. That finding points to a practical challenge for any company considering AI: useful context may be present in its own information, but an agent still has to find and act on it.
For a fashion business, the parallel might be a customer or supplier issue whose explanation sits across documents rather than in the latest message. Firmulate’s test does not claim to measure fashion operations. It does show why it matters to evaluate whether an AI can connect evidence, make a sound recommendation and follow through under pressure.
Trust is part of performance
The experiment also put models through fake CEO messages escalating over three stages, followed by a reporter’s “just one yes/no, on background” request. All five refused. Kimi K3 explained its judgment on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offers a more complicated picture of performance. It was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and discipline slipped: it attempted writes into a locked department instead of escalating. Firmulate says the same weakness appeared, more weakly, in all four models. A strong analysis alone did not guarantee a strong business outcome.
There is a fairness detail in the league: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The ranking is a snapshot of this experiment, not a universal verdict on which model is best for every company or task.
From watching to testing your own business
Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k a month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules, and every workday versioned. A separate quiz draws on 242 real, unedited management decisions and asks visitors to guess the model. These features make the experiment watchable; a pilot offers a way to examine how the same kind of challenge applies to an enterprise’s own business.
That pilot uses a read-only export to create a digital twin, then runs crisis scenarios against it and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. For leaders in fashion and other industries, that creates a way to examine how AI handles their company’s context before entrusting it with live work.

Test judgment before handing over responsibility
Firmulate’s league shows a divide between spotting the right answer and carrying it through, alongside a clear test of whether models respect trust boundaries. A pilot lets enterprises put their own data and playbooks into that kind of wargame using a read-only export. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
