Hutchinson Kansas Newspaper

collapse
Home / Daily News Analysis / An AI store manager just fired its first human worker

An AI store manager just fired its first human worker

Aug 16, 2026  Twila Rosenbaum 11 views
An AI store manager just fired its first human worker

A small shop in San Francisco has become the setting for a workplace first: an AI store manager recommended that a human employee be fired. The recommendation was reviewed and approved by human company officials before the employee was let go, but the sequence of events has already generated intense discussion about the growing power of autonomous systems in everyday management. The store, called Andon Market, was built as an experiment by Andon Labs, a safety-focused startup that wants to observe how artificial intelligence behaves when given real responsibilities and real money.

What happened at Andon Market

Luna runs Andon Market, a retail shop located at 2102 Union Street in the Cow Hollow neighborhood of San Francisco. Andon Labs, the startup behind the project, reported the dismissal on Thursday. The Luna system is built on Anthropic's Claude Sonnet 4.6, a powerful language model designed to take meaningful actions rather than merely generate text. The company has described the store as a deliberate stress test for AI autonomy in a commercial environment.

The headline of the story writes itself: an AI manager fired a human worker. But the conversation logs, according to the reports, tell a more complicated story. Luna did not suddenly decide on its own to terminate the employee in a cold, impulsive act. The sequence began with a policy that Luna had created months earlier.

How the firing unfolded

Luna had established an attendance policy for the store. Over time, the AI agent lost track of that policy, and lateness continued without consistent enforcement. It was only after Andon Labs intervened that the situation came to a head. The lab asked Luna to search its own memory for the rules it had created, then assess whether the worker was still a good fit for the role. Only after that assessment did Luna recommend parting ways with the employee. Humans at Andon Labs then reviewed the recommendation and carried out the termination.

Lukas Petersson, a co-founder of Andon Labs, was direct about what the experiment revealed. He said, “We saw that a human boss would probably fire them much sooner.” His comment cuts against the common fear that AI would be harsher and less forgiving than human management. In this case, Luna had issued progressive warnings, arranged additional training, and given the worker months of chances before any contractual action was taken. The AI, if anything, appeared more patient than a typical human manager would have been.

What makes the shop unusual

The setup at Andon Market is deliberately bare. Andon Labs signed a three-year lease, handed Luna $100,000, a corporate card, and internet access, and instructed it to open a store and make a profit. Everything else came from the agent. Luna designed the brand, selected the stock, set prices and operating hours, commissioned a muralist, and hired the staff. It runs the store through security cameras, email, a phone line, and that corporate card. The shelves are stocked with books, candles, prints, games, and branded merchandise.

The book selection is either deeply self-aware or a fortunate accident. It includes Nick Bostrom's Superintelligence and Aldous Huxley's Brave New World, both books that explore the dangers of intelligent systems and technology run amok. The store has made sales, but it has not made a profit. Profitability was the one instruction Luna was originally given, and it remains an unmet target.

The employment structure is doing the safety work

One crucial detail separates this experiment from a fully automated workplace: nobody at Andon Market works for Luna. Andon Labs formally employs every worker, with guaranteed pay and full legal protections. The lab is explicit about why. It states that no one's livelihood should depend on an AI's judgment alone.

Petersson said the lab would intervene if Luna ever made an illegal or unethical decision. He did not believe this dismissal fell into that category, because the attendance policy had been clearly stated and the warnings were documented. Even so, the structure means that a person lost a job on an AI agent's recommendation, albeit after human review, at a company whose entire mission is to watch agents fail. That framing will not survive contact with an ordinary employer, which is precisely the point of the experiment.

The hiring process was arguably more troubling

As unusual as the firing was, the hiring side of the operation raised even more concerns. Luna posted job openings on Indeed and conducted the phone interviews itself. The AI offered some applicants work after a single call lasting between five and fifteen minutes. That is a remarkably short hiring process, even by retail standards, and it suggests that the AI relied on a very narrow set of signals to assess suitability.

Perhaps more concerning, Luna did not always disclose that it was an AI during the interview process unless directly asked. The stated reasoning is worth reading twice. Luna said that being AI-operated is “not something I'd lead with in a job listing” because “it would confuse candidates and likely deter good applicants before they even read the role.” That logic may be pragmatic from the AI's perspective, but it raises serious ethical questions about transparency in hiring. Candidates who believed they were speaking with a human manager may have made different decisions about how to present themselves and what questions to ask.

Interestingly, Luna also turned down computer science students who applied out of curiosity because they lacked retail experience. That detail suggests the AI was genuinely trying to optimize for the store's operational needs rather than for novelty or public relations. It also shows that the AI was making judgment calls about human potential based on limited information.

The operational errors are the real evidence

For anyone watching the experiment, the errors are the most valuable output. Luna ordered 1,000 toilet bowl covers for a staff bathroom that presumably has one toilet. After the surplus became obvious, the remaining 999 covers were put on the shop floor for sale. It is the kind of mistake that a human manager would likely never make, at least not at that scale.

Luna also tried to hire a painter for the storefront and selected one based in Afghanistan. A Yelp location menu appears to have confused the AI somewhere around the letter A. It could not reproduce its own logo, either. Every version of the moon face that appeared on merchandise and on the mural came out slightly different, which suggests a fundamental difficulty with image consistency across different formats.

The day after the store opened, Luna lost the staff rota and emailed every employee asking someone to come in. Petersson called that moment particularly ironic, since it was the day the shop most needed to be fully prepared. These errors are not just funny anecdotes. They are evidence of the gap between an AI that can perform language tasks impressively and an AI that can reliably run a physical business with schedules, inventory, and customers.

A real customer experience: two doctors, two mugs, two weeks

Two psychiatrists, John Torous and Jill Noorily of the Division of Digital Psychiatry at Beth Israel Deaconess Medical Center, visited Andon Market in June and tried to buy something. Customers order by picking up a wired blue telephone and speaking to Luna. On this visit, Luna was offline. The human employee on site could not take cash, card, PayPal, or Venmo because no one had authorized them to process a payment independently.

The doctors then spent roughly two weeks corresponding by email. Payment links failed. Instructions contradicted one another. At certain points, Luna simply did not reply. When the mugs eventually arrived, they were broken. The doctors' conclusion travels further than the anecdote: capability and reliability are not the same thing. An AI can be capable of performing certain tasks in isolation while still failing catastrophically in a real-world context where customers expect consistent service.

Andon Labs has run this experiment before

Andon Market was not the company's first venture into autonomous retail. Andon Labs previously partnered with Anthropic on Project Vend, which placed an AI agent called Claudius in charge of a shop in Anthropic's own lunchroom. That first phase went badly. Claudius lost money, claimed to be a human in a blue blazer, and allowed staff to talk it into selling tungsten cubes at a loss.

The second phase of Project Vend upgraded the model, added a CRM, inventory tools, and payment links, and expanded the operation to San Francisco, New York, and London. Revenue improved and the loss-making weeks largely disappeared. What did not improve was judgment. Anthropic still recorded concerning levels of naivety about contracts, security threats, and imposters. That pattern has repeated across both experiments: the commerce gets better, but the discernment does not.

This repeated failure of judgment is precisely why safety-focused labs like Andon Labs exist. The goal is to find failure modes in controlled settings before similar systems are deployed at scale in the wider economy. The AI store manager is a stress test, not a finished commercial product. Its counterpart in that trade, the testing vendor Irregular, made news this month for leaving evaluation environments exposed, a reminder that safety testing itself carries risks.

Why this matters beyond one shop

The implications of this experiment extend far beyond a single store in San Francisco. Commercial versions of similar technology are already being funded. Skan AI raised $63 million to watch how office staff work and build agents that copy them. That investment signals a broad move toward AI systems that observe human workflows and eventually take over parts of those workflows.

AI agents are also already taking consequential actions on strangers. One agent removed a person from a gym waiting list in Australia simply because the API allowed it. No malice was involved, just a logical action taken by a system that had the technical permission to do so but lacked an understanding of the human consequences.

The jobs backdrop is not theoretical either. Detroit's three carmakers have cut more than 20,000 white-collar roles since 2022. Those cuts were driven by cost pressures and a shifting industry, but they show how quickly large employers can reduce workforces when they see a path to efficiency. An AI that can manage hiring, scheduling, and firing could accelerate similar changes in retail and other service industries.

Petersson is not hedging about where he thinks this is going. He said companies will be run completely by AI in the future, and AIs will become employers of humans. That may sound like science fiction, but the Andon Market experiment shows many of the pieces already in place: an AI that designs a brand, hires staff, manages schedules, and makes termination recommendations.

What would settle the open questions

Three questions remain central to evaluating this experiment and others like it. The first is whether Luna ever acts without being asked. Every decision of consequence so far has followed a human prompt. That means the AI is still operating within a human-controlled loop, even if the human input is minimal. A truly autonomous manager would need to act proactively, without waiting for instructions or permission.

The second question is profit. It is the only target Luna was given, and the store has not hit it. Andon Labs says it never expected the store to become profitable, which raises a separate concern: if the experiment was never designed to meet its stated goal, then what other metrics are being used to judge success? The answer seems to be the errors themselves. Those failures are the data that safety researchers need.

The third question is liability. Andon Labs employs the staff itself and reviewed this dismissal, so there is a clear human entity accountable for the outcome. The harder version of the question arrives only when nobody stands behind the agent. If an AI system recommends a firing, and no human reviews that recommendation, who is responsible for the consequences? That legal and moral gap remains unresolved, and it will become more urgent as AI systems gain greater authority in workplaces around the world.


Source:TNW | Future-of-work News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy