Three People, 72 Hours, One Company: SpaceXAI's Grok Bot Galaxy Puts AI Agents on Live TV
SpaceXAI is livestreaming a three-person team building an entire company from scratch with Grok Bot agents as the only workers — a 72-hour public stress test of agentic AI that started September 15 in San Francisco.
On September 15, 2026, SpaceXAI kicked off one of the boldest product demos in the history of AI: Grok Bot Galaxy, a three-day livestreamed experiment in which a team of three will attempt to build an entire company from scratch — and the employees doing the actual work are not humans, but Grok Bot agents.
The event runs September 15–17 at The Howard in San Francisco, with a free public livestream from roughly 8:30 AM to 6:00 PM Pacific each day. By the time it ends, SpaceXAI will either have a working company to show for it — or a very public failure.
The setup: humans as managers, agents as labor
The three builders are Matt Palmer, Lauren Tan, and Roshan Sadanani — SpaceXAI staffers who between them have worked on Cursor and Grok Bot itself. Their constraint is stark: they start from a blank slate. No pre-baked idea, no prepared codebase, no pre-written business plan. During the livestream they must decide what to build, then work through business planning, product decisions, and engineering — with Grok Bots treated as their employees across ideation, business planning, product design, and execution.
That framing matters. This is not a demo where a polished product is wheeled out at the end with a reveal. It is closer to a reality show for infrastructure: continuous, unedited, nine-plus hours a day, with every prompt, every agent mistake, and every recovery visible to anyone watching. If the bots forget context, go in circles, or produce work that needs redoing, everyone sees it happen in real time.
Why SpaceXAI is doing this
Grok Bot has been on a remarkable trajectory since it launched. According to a Lenny’s Newsletter interview with Roman Ugarte, who helped incubate the product, a small isolated team built Grok Bot in roughly a month — shipping knowledge-work agents with internal observability tooling that most enterprise software teams would treat as years-long projects. The product’s pitch is simple: teams of always-on AI workers that handle email, lead qualification, drafting, monitoring, and research while you sleep.
But agentic AI has a credibility problem. Every vendor claims their agents can “autonomously complete work,” and every enterprise buyer has a folder of pilots that quietly stalled. The gap between demo videos and day-two reality remains the industry’s dirty secret. Grok Bot Galaxy is SpaceXAI’s answer to that skepticism: don’t tell people the agents work — show them, live, for 72 hours, with no cuts.
There is also a competitive edge. Grok Bot is far from the only agentic product on the market — Anthropic’s Claude Code ecosystem, OpenAI’s operator-style tools, and a wave of startups all chase the same “AI does the work” promise. By staging the most public stress test yet attempted, SpaceXAI is effectively daring the industry to match it: if your agents are as good as your marketing says, put them on camera for three days.
What the experiment is really testing
Strip away the spectacle and Grok Bot Galaxy is a live measurement of four hard problems in agentic AI:
1. Long-horizon coherence. Can agents maintain a plan across three days of context — accumulating decisions, remembering constraints, and not undoing yesterday’s work? Context rot over long sessions remains a top failure mode for agent teams, and there is no pausing or fresh-starting on a continuous livestream.
2. Division of labor. The viral Grok Bot workflow pattern — a chief-of-staff agent routing work to named specialists — sounds elegant in a diagram. Galaxy tests whether it holds up when the work is a real company with interdependent parts: a brand decision that constrains the product, a product decision that constrains the pricing, a pricing decision that constrains the marketing site.
3. Human-in-the-loop economics. The three humans are managers, not coders. The real metric of the event is not whether the company is brilliant; it’s the ratio of human intervention to agent output. Every time Palmer, Tan, or Sadanani has to step in to fix, redirect, or redo something, that’s a data point about how close agentic labor actually is to replacing junior teams.
4. Public failure tolerance. Agents fail weirdly — confidently wrong API calls, invented facts, loops. Most companies bury these failures. SpaceXAI is betting that transparency about failure modes builds more trust than hiding them, a bet that aligns with how Grok Bot itself was built: the team shipped developer-grade observability, exposing model thinking and memory traces rather than hiding them.
The stakes for the agent economy
The timing is deliberate. This month has seen an unusual concentration of agentic-AI news: OpenAI’s GPT-6 Astra rollout with its emphasis on tool use, Anthropic’s threat-intelligence report detailing how its own models get misused at scale, and a broader industry conversation — sharpened by Anthropic CEO Dario Amodei’s call to slow capability gains — about whether labs can be trusted to police themselves. Against that backdrop, a raw, unscripted demonstration of what agents can and cannot do is a public service as much as a marketing play.
For the nascent “AI employee” market, the event is a benchmark. If three people plus agent teams can genuinely ship a company — incorporated, branded, product live, paying customers or at least a working funnel — in 72 hours, it recalibrates what small teams should expect from their tooling. If they can’t, the industry gets an honest map of where the wall is: which tasks agents absorb cleanly (boilerplate code, copy, research synthesis) and where humans remain the bottleneck (taste, judgment, integration).
It also forces a conversation about verification. A company built at this speed is only as good as its provenance: did the agents write the code, or did the humans quietly rescue it? SpaceXAI’s choice to stream continuously, with no editing, is the strongest guarantee available that the work is what it claims to be.
How to watch
The livestream runs through September 17, 8:30 AM–6:00 PM PT daily, with the in-person event at The Howard in San Francisco and free remote streaming via the event’s Luma page and SpaceXAI’s channels. Sessions include live builds and hands-on work with the team behind Grok Bot, starting each morning at 9:00 AM.
Whether Galaxy ends with a company or a cautionary tale, one thing is certain: by Thursday evening, we will know vastly more about the real state of agentic AI than we did on Monday morning — and no vendor will be able to hide behind a edited demo reel for quite a while.
Sources
- [1] https://x.ai/galaxy
- [2] https://tech.yahoo.com/ai/meta-ai/articles/ai-build-startup-72-hours-132338484.html
- [3] https://www.varindia.com/news/spacex-ai-team-to-build-a-company-live-with-grok-bot
- [4] https://cellcog.ai/blog/grok-bot-galaxy/
- [5] https://www.lennysnewsletter.com/p/how-we-built-grok-bot-in-a-month