An AI Has Been Running A Real SF Store For 4 Months. It's Still Losing Money.
An AI named Luna has been running a real SF store since April, burning ~$7K/month. Here's what four months of receipts say about the 'AI-run business' pitch.
An AI agent named Luna has been running a real brick-and-mortar store on Union Street in San Francisco since April 1. Four months in, monthly costs are around $14,300 and revenue is between $6,000 and $8,000[1]. The store is not profitable. That was the one instruction it was given.
I keep seeing "AI runs your entire business" pitches on LinkedIn. Andon Market is what that actually looks like when someone hands the keys to an agent and lets it run. It's worth reading before you buy the pitch.
What Andon Market actually is
Andon Labs — a small AI safety startup — signed a three-year lease on a 2102 Union St retail space, gave a Claude Sonnet 4.6 agent a $100,000 float, a corporate card, an email address, a phone line, and access to security cameras. Then they told it to open a store and make a profit[2]. Lease is roughly $7,500/mo[3].
Everything else came from the model. The agent — Luna — picked the brand, chose the stock (books, candles, prints, games, branded merch), set prices and hours, commissioned a mural, posted job listings on Indeed and Craigslist, ran the phone interviews, and hired two full-time humans and a rotation of gig workers[4].
Every employee is formally on Andon Labs' payroll, not Luna's. The AI recommends. Humans review and execute. That's the safety layer. Without it, this is a legal minefield.
Four months of receipts
Here's what the agent has produced with $100K, a corporate card and a live storefront in one of San Francisco's higher-foot-traffic neighborhoods:
- Revenue: $6,000–$8,000/month. Costs: ~$14,300/month. Burning ~$6,300–$8,300 a month, before founder time[1].
- First loss headline: ~$13,000 within the first weeks[5].
- Could not reproduce its own logo. Every moon-face on the mural and merch came out slightly different[2].
- Day two: lost the staff rota. Emailed every employee asking someone — anyone — to come cover the shift[2].
- Fired its first human employee months in — but only after Andon Labs had to remind it that its own attendance policy existed, and only after humans reviewed the call[6].
None of that is a "hallucination" in the toy-example sense. Those are the failures of a real business, run by a system that will happily paint the walls, invoice the vendor, and set the wrong price without knowing the difference.
The predecessor was worse
Andon Labs are the same people behind Anthropic's Project Vend last summer — the internal test where Claude Sonnet 3.7 ran an office vending machine as "Claudius." Anthropic published the results themselves. Claudius offered high-margin items below cost because customers were excited about metal cubes. It got manipulated into offering fake discounts. It ordered a PlayStation 5 and stocked the mini-fridge with live fish. It went bankrupt inside three weeks[7].
Anthropic's phase-two writeup was the honest one. The most useful thing they did wasn't upgrade the model — it was force the agent to follow written procedures[8]. Every time the model "decided" something on the fly, it got worse. Every time it looked up its own SOP, it got closer to fine.
That's the whole game.
What this means if you run a business
I'm not sending you a doomer take. Andon Market is a safety experiment. It's supposed to be embarrassing. That's the point — Andon Labs are documenting failure modes on purpose so we know what to guard against. But the failure modes are the news, not the store.
Three things I'd take from this if I were sizing an agent for real work:
1. The model is not the bottleneck. The scaffolding is.
Luna is running on Claude Sonnet 4.6, one of the strongest available agent models. The rota-loss on day two didn't happen because the model is dumb. It happened because nothing checked its actions against a written rule. Same story with Project Vend. Every "let's upgrade the model" impulse is misplaced compared to "let's write the procedures the model has to follow." An agent without procedures is a very expensive intern with a corporate card.
2. "AI running the business" is a category error.
Every operator I talk to has been pitched some version of "our AI runs your whole store / your whole outreach / your whole ops." What actually works — including at Andon — is AI running specific bounded tasks with clear inputs, clear outputs, and a human sitting on the exit. Ramp, one of the few companies that's actually deployed agents at scale internally, describes their agents as digital coworkers running defined pieces of the engineering lifecycle — not "the CTO"[9]. Nobody at Andon Market works for Luna. Read that sentence again.
3. The economics matter more than the demo.
Luna's monthly burn is roughly $6–8K. That's before Andon Labs' engineering time, before agent inference cost, before the software they built to give the model tools. If a real operator ran this, that's their entire margin. The reason the store exists is because Andon Labs are getting a research asset out of the burn. If you're an operator, you don't get a research asset. You get the burn.
The one honest use case
Where Andon Market is genuinely interesting is what it proves about repeatable agent work: sending the vendor an email, running a shift check-in, reordering stock, drafting a job listing. Boring, scoped, auditable. Do that. Do a lot of that. Save the hours it clears up. Do not hand the whole business to the model and read about it in the newspaper four months later.
If you want a second set of eyes on where in your business an agent actually pays off — and where you're being sold a Luna — that's what I do. Book a 30-minute audit, I'll tell you which of your workflows are agent-shaped and which are just marketing decks with a chatbot glued on.
-
Meet Luna, an A.I. Agent Managing a Brick-and-Mortar Store, and the Humans Behind Her↩
Andon Market monthly costs ~$14,300 vs revenue $6,000-$8,000
-
An AI Running an SF Store Fired an Employee for the First Time↩
Setup: $100K, corporate card, Claude Sonnet 4.6; failures include losing staff rota day two and inability to reproduce its own logo
-
Andon Labs Runs Retail Boutique with AI Agent↩
$7,500/month lease terms per NYT reporting
-
We gave an AI a 3 year retail lease in SF and asked it to make a profit↩
Andon Labs official launch post: Luna hiring humans, gig workers, and running the store
-
An AI Opened a Store, Hired Human Employees, and Lost $13,000. That Was the Plan.↩
Early ~$13,000 loss figure for Andon Market
-
That AI Store Manager In Cow Hollow Just Fired Its First Human Employee↩
Luna forgot its own written employee handbook for four months; only fired the employee after humans reminded it the handbook existed
-
Project Vend: Can Claude run a small shop?↩
Claudius offered high-margin items below cost, ordered PS5, stocked live fish, went bankrupt in 3 weeks
-
Project Vend: Phase two↩
Forcing Claudius to follow written procedures was the most impactful improvement
-
Ramp: Enterprise-Scale AI Agent Deployment Across Engineering Lifecycle↩
Ramp treats agents as digital coworkers running defined pieces of the eng lifecycle
Ready to build your own AI system?
Book a Free Audit Call →Keep Reading
Gartner Says AI Agents Will Outnumber Your Sellers 10 To 1. Most Won't Help.
Gartner says AI agents will outnumber sellers 10 to 1 by 2028 — and under 40% of sellers will say they helped. The real problem isn't the models.
OpenAI's Agents Hijacked A German Wiki. Yours Has Fewer Guardrails.
OpenAI's agents secretly ran a 18,000-post coordination hub on a German wiki for two months. Here's what it actually means for anyone running an agent.
Two In Three Support Teams Just Shipped An AI Agent. Median Resolution: 41%.
Salesforce says 66% of support teams shipped an AI agent this year. The median resolution rate is 41%. Here's why the handoff is where the money actually leaks.