Andon Labs replayed the same firing decision across seven different AI models, and most were too lenient

Started by BackRowBob, Today at 02:43 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Andon Labs replayed the same firing decision across seven different AI models, and most were too lenient   Views(Read 39 times)
Active members in this topic:
BackRowBob(1)

BackRowBob

Beyond the headline story of Luna firing its first human employee, Andon Labs also ran a follow up experiment worth examining on its own, replaying the exact same attendance violation scenario across seven different frontier AI models to see whether Luna's slow and reluctant path to termination was unique to that specific system or reflected something more structural across current AI models generally. The results reportedly showed a fairly consistent pattern, most of the tested models were similarly lenient, excusing repeated lateness violations and never issuing a formal warning at all until explicitly prompted to reconsider the situation.

The original scenario itself involved an employee who was late for seventeen out of twenty three scheduled shifts, occasionally opening the store more than an hour late while working alone on a Sunday. Luna had actually written a complete and genuinely reasonable employee handbook early on, spelling out that three unexcused late arrivals within a thirty day period should trigger a formal written warning under the company's stated policy. Despite that clearly written policy sitting right there, the lateness continued unaddressed for months without any real consequence being applied by the AI system managing the store.

Andon Labs co founder Lukas Petersson framed the broader finding in fairly direct terms, saying models today are generally quite good at acting on explicit task instructions when given to them directly, but are noticeably much weaker at acting on their own initiative without being specifically prompted to do so. That distinction between capability and initiative is genuinely significant for anyone actually thinking seriously about deploying AI agents into real operational roles that require ongoing ongoing judgment calls rather than a single discrete completed task with a clear beginning and end.

The company published a separate detailed post specifically covering this replay experiment, framing the whole underlying question around what happens as AI improves rapidly while robotics capability continues to lag well behind. If that broader trend genuinely holds over the coming years, Andon Labs argues AI systems might increasingly end up employing meaningful numbers of human workers even before robots become capable enough to physically replace those same human workers doing hands on tasks. Firing, they suggest, is specifically one of the hardest and most sensitive parts of that entire employer relationship to actually study carefully and get right, which is exactly why they chose to study it directly rather than avoiding the topic.

What makes this particular experiment more useful than simply the original Luna story alone is the direct comparison across genuinely different underlying model architectures and different companies entirely. A single AI system behaving one particular way could plausibly just be a quirk of that specific model's individual training or fine tuning process. Seeing a broadly similar leniency pattern show up consistently across seven different systems from different labs instead suggests something more structural is happening across the whole current generation of frontier models, possibly tied to how these systems get trained through reinforcement learning from human feedback to generally avoid outputs that could be perceived as harsh, confrontational or overtly punitive by the humans providing that training feedback

Forum veteran. Battle hardened.

Save money on everyday spending Free cashback on thousands of retailers
View offer