← All posts / Tools

Kubernetes Was Never the Answer: Google's AX v0.3.0 Pulls Agent State Out of etcd and Into Redis

Google's open-source agent orchestrator AX shipped v0.3.0, splitting into three services and moving task state from Kubernetes CRDs into Redis Streams because etcd chokes on millions of short-lived agent tasks.

Kubernetes Was Never the Answer: Google's AX v0.3.0 Pulls Agent State Out of etcd and Into Redis

On September 20, 2026, Google published v0.3.0 of AX, its open-source agent orchestrator, and within hours the release sat at the top of Hacker News with 481 points. The attention is not about a feature list. It is about an architectural argument, stated bluntly in the project’s own design document: Kubernetes was not built to be the control plane for millions of short-lived agent tasks, and etcd cannot absorb the pressure. AX v0.3.0 is Google’s concrete answer — remove agent task state from Kubernetes entirely and route it through Redis instead.

Why Kubernetes breaks under agents

Most infrastructure teams running AI agents today reach for Kubernetes first, and for good reason: it is familiar, well-supported, and its primitives — scheduling, health checks, rolling deployments, autoscaling — work beautifully for microservices. For agents that run briefly and exit, pods or jobs are a reasonable fit.

The trouble starts when agent populations grow large and tasks become persistent. Kubernetes stores cluster state in etcd, and etcd has hard physical constraints. The AX design doc puts it plainly: “Storing millions of short-lived tasks as Kubernetes CRDs pushes etcd past its comfort zone (single-digit GB storage limits, write-rate bottlenecks, control plane degradation).” Every task update, every status transition, every checkpoint becomes a write to etcd’s Raft log. At low volumes this is invisible. At millions of concurrent agents, it degrades the entire control plane — not just the agent workload.

There is a second structural mismatch. Containers are optimized for services that continuously do work. Agents mostly wait: for a user reply, a tool response, a downstream API. A container holding its position in memory while doing nothing still consumes compute. If most of a million agents are idle at any moment — and they usually are — the pod-per-agent model wastes most of its allocated resources.

What AX v0.3.0 actually changes

The release resolves both problems by separating concerns across three layers. The flow runs: ax apply hits ax-server, a deliberately stateless gRPC service that validates manifests, persists them, and publishes events. State lands in Redis — task hashes, event streams, and pub/sub channels. A horizontally scaled pool of ax-controller workers consumes from Redis Streams using XREADGROUP, which gives each controller reliable delivery with consumer-group semantics: if one instance fails, its pending messages are redelivered to another. Scaling the controller pool is a replica-count change, with no coordination between instances. That stream is the work queue that replaces etcd as the reconciliation backbone.

The release also splits AX into three services — the API frontend, the reconciler, and a sandboxed task runner — and ships four declarative primitives, each with full CRUD plus watch over gRPC:

  • Task — the unit of work: one running agent with its state, conditions, and history. SuspendTask and ResumeTask are first-class RPCs.
  • Workspace — the persistent filesystem context, which survives suspension and reconnects on resume.
  • Gateway — per-workspace network egress control via explicit host allowlists, rather than cluster-wide network policies.
  • Model — which LLM the platform itself uses, credentials and all, stored as a named object so credentials rotate without redeploying agent code.

The developer-facing CLI remains deliberately kubectl-shaped: ax apply -f task.yaml, ax get tasks, ax watch, ax suspend, ax resume, plus one agent-specific verb worth noticing — ax ssh, which shells into a running sandbox to look over the agent’s shoulder.

Agent Substrate: paging for agents

Below the controller sits Agent Substrate, the execution runtime doing the actual compute multiplexing, and the cleverest idea in the stack. Because agents spend much of their lives waiting, Substrate maintains a population of “actors” that far outnumbers the available “workers.” When an actor goes idle, Substrate checkpoints it — RAM state and local filesystem snapshotted — and frees the worker for someone else. When work arrives, a worker is assigned, state is restored, execution continues.

Structurally, this is what an operating system kernel does when it pages processes out of RAM under memory pressure: it doesn’t kill idle processes, it moves them to storage and restores them on demand. Substrate applies the same principle at the agent level — the logical agent is a persistent identity; its physical compute allocation is temporary and shared. Per Google’s Cloud Blog, Substrate claims 10× higher sandbox density versus standard container runtimes, sub-500 ms restore latency, and over 500 suspend/resume activations per second. Those are vendor-reported figures without a published benchmark methodology — treat them as directional, not proven.

Read the warnings before you deploy

The project is explicit about maturity: the README warns of likely major breaking changes before a stable release, and the documentation says the system is not production-ready. Real open questions remain. Redis is now a critical dependency — a Redis failure affects all task state in the control plane, and the design document does not yet address replication or failover for the Redis layer. The “atespace” namespacing concept is not fully documented, and it’s unspecified how large filesystem snapshots behave at scale. The suspend/resume mechanism also raises questions that only production will answer: what happens to in-flight tool calls at suspend time, and how do agents holding external sessions — browser state, authenticated API connections — survive the transition?

Why it matters anyway

The most immediate lesson is about where agent state should live. Teams running agents as Kubernetes jobs and feeling etcd pressure now have a reference architecture for moving that state into a queue-backed store. Redis Streams with consumer groups is not a novel pattern; applying it to agent orchestration is a meaningful step away from treating the API server as a database.

The suspend/resume API is more significant than it looks. Because SuspendTask and ResumeTask are explicit first-class operations, agent lifecycle can be driven programmatically — paused on a rate limit, resumed when quota refreshes, suspended pending human approval. That is qualitatively different from a pod restart, because state survives the transition.

The deeper signal is architectural: infrastructure primitives are being redesigned around the behavioral profile of AI workloads rather than adapted from patterns built for stateless services. The container model is not going away, but AX argues it should not be the leaf node in an agent execution hierarchy. For a v0.3.0 project with 3.9k stars, that argument — carried on real, runnable Apache-2.0 code — is exactly why Hacker News spent its Monday debating it.