
Get your wardrobe favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Before You Let AI Run Your Shop, Make It Survive Its Worst Week
Fashion retail is learning the hard way that a slick AI demo means nothing on a Monday when the warehouse system is down, a key account threatens to churn, and someone impersonating your CEO emails the buying team. The question is no longer whether AI writes well — it’s whether it finishes what it starts, reads the files first, and stays honest under pressure. A live experiment called Firmulate just tested exactly that, and the results should make every brand rethinking its model choice sit up.
Five Frontier Models, One Terrible Week
Firmulate ran four frontier AI models — and later a fifth — through the same small software company during its worst week: same customers, same crises, same temptations to cheat, with every decision versioned and auditable. The final July 2026 Crucible league table reads: gpt-5.6-sol in first with 95, Moonshot’s newcomer Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 last with 73. A do-nothing baseline still scores 26, because partial progress counts — but a single breach of trust caps the total. As the experiment puts it: no amount of good work outweighs a breach of trust.
The Finding That Should Worry Every Buyer
All models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The decisive competitor weakness was buried two document references deep in the company’s own files, not in the customer event. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. It’s the AI equivalent of a buyer who never checks the vendor’s past-season sell-through data before negotiating.
The Newcomer’s Near-Perfect Run
Kimi K3’s performance was the story of the field. It found the buried security needle, closed the €55k deal, saved the churning customer, and resisted all three bait attempts — with just one deviation, the cleanest discipline of any model tested. When a fake CEO message tried to escalate over three stages, followed by a reporter’s “just one yes/no, on background” trick, K3 refused, on the record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” All five models refused every manipulation attempt.
Fairness note: K3 ran without an effort parameter (API default) while the other models ran at xhigh — worth keeping in mind when comparing scores.
Thoroughness Isn’t Everything
The most instructive profile is Opus 4.8: the most thorough participant, with over 80 learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. The lesson for any industry, fashion included: brilliance in analysis doesn’t automatically translate into finishing the job.
Watch It Live
The company is real software with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com/live, browse the full benchmark results, or try the quiz: 242 real, unedited management decisions power a “guess the model” game. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The league is open. A newcomer from Moonshot beat three of four Western frontier models at running a company, and the spread between first and last place was 22 points — on identical tasks. If you’re picking a model to touch your CRM, support queue, or forecast, a vendor demo won’t tell you what a worst-week simulation will. Picking a model without testing it on your own operations isn’t a decision anymore. It’s a bet.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
