Why can’t my agent fleet just stay up without burning cash or crashing the cluster?
Five coding agents. Or a research swarm grinding trajectories. One sits forty seconds on a model call. Another blocks on human approval. A third spins a tool loop. Deployments keep pods hot and expensive. Batch Jobs die when the process exits. LangGraph and ADK own the reasoning loop – isolation, resume, egress fences, dense packing? Still your problem.
AX (Google’s Open Agentic Orchestrator) topped Hacker News around 21 Sep 2026 – 600+ points, hundreds of comments. Declarative control plane. Agents as their own workload type. Not another prompt framework. Below: what that means in practice, a first run that isn’t the README golang demo, and the alpha failure modes that actually bite.
Where existing stacks fall short
ADK-style kits handle multi-agent graphs and tool calling. Gemini Managed Agents collapse deploy to one API call and hand away the execution layer. Two agents in worktrees on a laptop? Fine. Strict CPU/memory sandboxes, network allowlists, checkpointed idle state, hundreds of short sessions? That setup folds.
AX rides Agent Substrate. Substrate shipped open source with the May 2026 Agent Executor work: many idle actors multiplexed onto fewer warm workers, sub-second (they claim sub-500ms) resume, gVisor/microVM-style isolation. AX is the kubectl-shaped face – four YAML primitives, familiar verbs. Control plane lands in ax-system with Redis; Substrate in ate-system.
How to stand up Google’s Open Agentic Orchestrator
As of late September 2026 this is alpha. README is blunt: major breaking changes before stable. You need a real cluster – Kind is enough to learn; you still won’t skip Substrate or a registry.
- Agent Substrate first (its README). Namespace
ate-system, Control API atapi.ate-system.svc.cluster.local:443. Sanity check:kubectl get svc api -n ate-system. - CLI:
go install github.com/google/ax/cmd/ax@latest. Put$(go env GOPATH)/binon PATH. - Control plane (Redis + ax-server + ax-controller) into
ax-system:make deploy AX_IMAGE_REPO=<your-registry>. Needskoand a registry the nodes can pull.
After the API answers, everything is ax.io/v1alpha1 manifests. Four primitives. That’s the whole core model per the docs on github.com/google/ax and agentexecutor.io.
apiVersion: ax.io/v1alpha1
kind: Model
metadata:
name: research-model
spec:
provider: google
model: gemini-2.0-flash # pin whatever is current as of your install day
# platform bootstrap secret - not a generic Task secretRef
---
apiVersion: ax.io/v1alpha1
kind: Workspace
metadata:
name: traj-ws
spec:
git:
- repo: https://github.com/your-org/eval-harness.git
branch: main
# or a plain-English goal the platform agent uses to bootstrap tools
---
apiVersion: ax.io/v1alpha1
kind: Task
metadata:
name: traj-collector-01
spec:
workspaces:
- name: traj-ws
goal: "Python 3.12 + the eval harness deps ready"
debug: true
# limits, image, command, env, gateway ref
ax apply -f research-task.yaml
ax watch task traj-collector-01
ax ssh traj-collector-01 -- ls -la /workspace
ax suspend task traj-collector-01
ax resume task traj-collector-01
Task = cheap isolatable sandbox. Workspace materializes git / MCP / skills once so runs start warm. Model holds provider config + platform secrets. Gateway (not in the snippet) = egress allowlist and credential injection.
A concrete research trajectory run
Skip the recycled golang README sample. Pull your eval use in the Workspace. Task runs a headless collector: N episodes, JSONL trajectories on the durable /workspace volume, then idle. While it waits on the next model response, Substrate checkpoints and frees the worker. Resume later. Or branch the trajectory for an alternate policy – same durable-execution idea Google described in the May 2026 Agent Executor announcement.
ax watch for live status. Shell only with debug: true. Suspend on an empty queue; resume when seeds show up. Density = many Tasks sharing workers, not one pod per agent. Design talk is “billions of tasks per cluster”; agents as stateful actors that idle on models, tools, or HITL – not microservices, not batch jobs, and not a framework that writes your agent logic.
Pro tip: Don’t trust WorkspaceReady alone.
ax ssh ... -- git rev-parse HEADor a plainls. Empty clones have shown up Ready (issue discussion around #347).
Honest limits you will hit today
Architecture reads clean. Alpha edges cut.
- Egress TLS: Hostname-style Gateway allowlist entries can open inspectable tunnels that fail on TLS passthrough – including hosts you thought you allowed. GatewayReady still goes true. Prefer CIDR entries until hostname behavior lands (community/issue #345).
- Zombie phase: Command exits; runner stays PID 1 for metadata/ssh. Task often sits Phase=Running Ready=True forever. No exit code in TaskStatus yet (#346). Drop your own success file or metric and poll that.
- Secrets / API: Task env is plaintext strings only – no valueFrom/secretRef. Credentials can round-trip in Redis JSON. Model secrets are platform-bootstrap only. AX gRPC dial defaults to insecure; no auth/TLS on the API server in current builds (#348). Treat the control-plane network as fully trusted or front it yourself.
- Empty git Ready: WorkspaceReady/SetupComplete can flip true when the clone brought no commits. Verify before you burn a run.
Community threads and hands-on notes line up on these as of the v0.3.0 window (control plane restructured around ax-server/controller/task-runner + Redis Streams ~20 Sep 2026). AX is self-hosted Apache 2.0, not a managed agent product. Spend caps, governance, and the agent binary stay yours.
Kubernetes tax for a solo laptop afternoon? Usually no. Reproducible eval fleets, or long-running untrusted tool-using agents that must survive preemption – then the primitives match the workload shape instead of fighting Deployments and Jobs.
FAQ
Is AX free and production-ready?
Apache 2.0, no license fee. Cluster, Substrate, Redis, registry: your bill. Early/alpha; expect breaking changes. Not production-ready as of late 2026.
How is this different from ADK or LangGraph?
They build graphs and tool loops. AX orchestrates sandboxes, workspaces, egress fences, suspend/resume under whatever use you already have – ADK, LangGraph, or a raw binary. Mix them. AX does not replace the reasoning loop.
Can I try it without a full production cluster?
Kind or a small GKE Standard cluster is enough to learn. Substrate still has to be installed; you still need a registry the nodes can pull. No pre-built one-click managed AX offering yet. No platform team? Stay on managed sandboxes or local worktrees until the control plane hardens – then clone google/ax, stand up Substrate on a throwaway cluster, apply a Task, and immediately prove suspend/resume plus one real egress rule with a live model call.