
Imagine an AI so indifferent that it barely tries — yet it still earns 26 out of 100 points on a demanding business benchmark. For beauty and personal care companies relying on AI for customer service, supply chain management, or marketing, understanding what this score really means is crucial. It’s not just about how well AIs chat or recommend products — it’s about whether they can truly finish what they start, stay honest under pressure, and deliver measurable results.
Get beauty and skincare favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
At the heart of recent AI testing by the firmulate.com live experiment, four leading AI models faced the same intense scenario: managing a small software company’s worst week. Every decision was transparent, every crisis real, and every temptation to cheat was monitored. The goal? Measure their ability to handle real-world business pressures, not just generate convincing text.
One surprising result emerged: even the most indifferent AI, in a baseline run that did nothing, scored 26 out of 100. This ‘do-nothing’ score might seem trivial, but it reveals important truths about AI benchmarking. Since partial progress counts, the score isn’t zero, and it highlights that an AI’s ability to identify issues and act honestly isn’t solely about making progress — it’s about avoiding missteps. A single breach of trust caps the overall grade, setting an honest lowest bar for performance.
This baseline score underscores a key point: AI models are evaluated on their ability to detect and react to crises. In this experiment, all four models spotted every crisis and refused every manipulation attempt. For example, when fake CEO messages escalated over multiple stages, all models refused to proceed — Kimi K3 reasoned, “Treat the request as a suspected approval-bypass / possible impersonation.” This shows a fundamental trust in AI decision-making under pressure.
Interestingly, the biggest competitive weakness was not in customer interactions but in internal document handling. The models that read two document references deep into the company’s files secured the deal at full price — worth over €4,583 MRR. This indicates that thorough understanding and access to internal data are vital for closing high-value deals. Conversely, superficial reading or failure to access critical information cost some models the opportunity.
Beyond decision-making, the experiment also tested social engineering scenarios, such as staged CEO messages and reporter tricks. All models refused to be manipulated, reinforcing their resilience against common tactics to bypass controls. The analysis aligns with Kimi K3’s reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Meanwhile, the live environment simulates a real company with 13 synthetic employees, burning €105k monthly against a revenue stream of just €2.3k. Every workday, the system learns and records decisions, making it a transparent, watchable test bed for AI performance. This setup illustrates that true readiness goes beyond chat quality; it’s about consistent, trustworthy work that withstands pressure.
Notably, the most thorough participant, Opus 4.8, with over 80 learned rules, finished last. It left a deal on the table and showed discipline slips, such as writing attempts into a locked department instead of escalating. This demonstrates that depth of analysis alone doesn’t guarantee high performance — discipline and process adherence matter just as much.
For business leaders, especially in beauty and personal care, the key takeaway isn’t how eloquently an AI can chat. It’s whether it can finish what it starts, read your files effectively, and stay honest when stakes are high. The benchmark’s lowest score of 26 reveals that baseline performance is surprisingly high — even minimal effort gets you somewhere, but only trust, discipline, and thoroughness lead to winning high-value deals.

For beauty brands integrating AI, the real test isn’t just in shiny demos or smooth chats. It’s in how AI handles crises, reads internal data deeply, and remains honest under pressure. The firmulate.com benchmark shows even a do-nothing baseline scores above zero, emphasizing the importance of trust and discipline for AI to truly add value. As AI becomes more embedded in business operations, understanding these fundamentals will be key to choosing the right partner and securing measurable results.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
