
A product recall, a sudden wave of cancellations or a competitor undercutting a hero product can put a beauty business under pressure fast. An AI agent might help manage those moments—or make them worse. Firmulate’s live experiment asks a practical question: how would different AI models run a company when the stakes are real?
Get beauty and skincare favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
Firmulate puts AI models in charge of the same small software company through its worst week. Each faces the same customers, crises and temptations, and every decision is versioned and auditable. The live company has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, with a public cash countdown. Its playbooks have accumulated more than 680 self-learned rules.
The experiment’s final Crucible League, in July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The benchmark’s principle is blunt: “no amount of good work outweighs a breach of trust.”
Good diagnosis is not enough
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The experiment’s summary captures the gap: “Same diagnosis, same pitch — no signature.”
The winning clue was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file secured the deal at full price, worth €4,583 in monthly recurring revenue. For a beauty company, the parallel might be a critical detail in a product, supplier or customer record: an agent has to find relevant evidence and follow through, not just sound confident in a conversation.
Firmulate also tested escalating fake CEO messages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
More analysis does not guarantee better management
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It still finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The standings are a snapshot of this experiment, not a promise about how a model will perform in every company or situation.
From watching to testing your own business
The experiment is publicly watchable at Firmulate. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each call. But the enterprise question is more direct: what would an AI do with the pressures, policies and weak points of your own business?
Firmulate’s proposed pilot uses a read-only export of a company’s business to run crisis scenarios and produce a board report, including model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. That gives leaders a way to examine how an AI workforce might respond before connecting agents to customer, support or operating workflows.

Put the decisions to a test
For beauty and personal care businesses weighing AI agents, the useful question is not only whether a model can recognize a problem. It is whether it can find the evidence, respect boundaries and complete the job. Explore a Firmulate pilot using a read-only export of your business, and contact contact@firmulate.com to discuss running the wargame.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
