Your long-running coding agent just finished hour three of a multi-repo refactor. The pod gets evicted. Normally everything vanishes. With AX – Google’s Open Agentic Orchestrator, you suspend, resume, and the sandbox picks up with files and state still there. That’s the end state this guide gets you to – on your own cluster, with YAML you control.
AX (Agent Executor) showed up open source from Google around mid-May 2026. Still hot. Still early. Community threads latched onto the four-hour crash story; under the hood it’s a declarative control plane on Agent Substrate. Agents become suspendable actors, not fragile Python processes you babysit.
What you’re actually buying with AX
Not a LangGraph clone. The Google Cloud announcement frames it as self-hosted durable execution: isolation, session consistency, connection recovery, trajectory branching. Keep ADK, LangChain, a custom container, MCP tools – AX takes the lifecycle.
If you already think in Kubernetes objects, the muscle memory transfers. API group ax.io/v1alpha1 (early preview – this can move). Four primitives:
| Primitive | Job |
|---|---|
| Task | Isolated sandbox, image/command, CPU/memory limits, workspace mounts, gateway ref |
| Workspace | Git clones, MCP servers/registries, skills, optional plain-English goal for generative setup |
| Gateway | Listeners the task exposes + egress host/port allowlist |
| Model | Named provider/model/params + Kubernetes secret for the API key |
Everything sits in an atespace (default: default). Idle actors checkpoint on Substrate and wake fast – the project site and docs cite sub-second resume (sometimes sub-500ms) and dense multiplexing, so you aren’t paying full compute rent while an agent waits on a tool or a human.
Practical setup: from zero to a running task
Checklist before you touch YAML: Go for the CLI, a Kubernetes cluster, ko, a registry your nodes can pull, Agent Substrate Control API reachable (in-cluster default as of the README: api.ate-system.svc.cluster.local:443). Early preview. Breaking changes are expected.
- CLI:
go install github.com/google/ax/cmd/ax@latest– binary lands in$(go env GOPATH)/bin; put that onPATH. - Control plane from a clone of github.com/google/ax:
make deploy AX_IMAGE_REPO=<your-registry>. As of the current quick start that brings components up inax-system(Redis + AX server/controllers in the documented path). - Model secret, e.g.
kubectl create secret generic gemini-api-secret --from-literal=GEMINI_API_KEY="...". - Apply one multi-doc YAML (Task + Workspace + Gateway + Model). Then
ax get tasks,ax watch task <name>, and – ifdebug: true–ax ssh <name> -- ls -la /workspace.
Starter below skips the stock chalk/golang demos. Two workspaces: app code + shared tools. A generative goal on first boot so the runner can finish the toolchain instead of you baking every image by hand. Model id and runner digest are examples from docs/examples as of mid-2026 preview – pin what your cluster actually pulls; digests rotate.
apiVersion: ax.io/v1alpha1
kind: Workspace
metadata:
name: app-src
atespace: default
spec:
git:
- name: origin
repo: "https://github.com/your-org/your-service.git"
branch: "main"
---
apiVersion: ax.io/v1alpha1
kind: Workspace
metadata:
name: team-tools
atespace: default
spec:
mcp:
servers:
- name: git-tools
endpoint: "http://git-mcp.default.svc.cluster.local:8080"
---
apiVersion: ax.io/v1alpha1
kind: Gateway
metadata:
name: default-gateway
atespace: default
spec:
listeners:
- name: http
port: 8080
protocol: HTTP
- name: grpc
port: 8494
protocol: gRPC
egress:
allowlist:
hosts:
- host: "*"
port: 443
---
apiVersion: ax.io/v1alpha1
kind: Model
metadata:
name: default-model
atespace: default
spec:
provider: google
model: gemini-3.8-flash
secretKey:
name: gemini-api-secret
key: GEMINI_API_KEY
parameters:
temperature: 0.9
---
apiVersion: ax.io/v1alpha1
kind: Task
metadata:
name: refactor-bot
atespace: default
spec:
image: "gcr.io/ax-substrate/ate-images/ax-task-runner@sha256:3a0dea6ad8b55278685db58aca6e37dc4ba04056831d45bef3aaeafdca43cac6"
env:
- name: ENVIRONMENT
value: "staging"
resources:
requests:
cpu: "500m"
memory: "1Gi"
limits:
cpu: "2"
memory: "4Gi"
workspaces:
- name: app-src
path: "/workspace"
goal: "Install the project toolchain and make tests runnable"
- name: team-tools
path: "/workspace/tools"
gateway:
name: default-gateway
debug: true
Apply: ax apply -f task.yaml. Watch WorkspaceReady and Ready conditions – not only the phase string – before you assume the agent can work.
Pro tip: Leave
debug: trueon while learning. Guest services stay off without it, soax sshis dead. Flip it off when you stop poking production-shaped sandboxes.
Day-two ops that matter
Suspend/resume is why most people showed up.
ax suspend task refactor-bot
ax resume task refactor-bot
ax ssh refactor-bot -- cat /workspace/notes.txt
Checkpointed actor state returns. Files you wrote pre-suspend are still on disk. Restart-the-prompt vs continue-the-job – different products pretending to be the same demo.
Generative workspaces are the other lever. Plain-language goal on a binding gets handed to a bootstrap agent on first boot (docs describe an Antigravity-style path). Needs model credentials and a bootstrap timeout – default around ten minutes; override path in sandbox docs is AX_BOOTSTRAP_TIMEOUT as of writing (preview: names can change). Use it when every task would otherwise reinstall the same toolchain from scratch.
Multi-workspace binding splits “code under change” from “shared MCP/skills.” First entry is the command working directory. Full workspace readiness waits until every binding is prepared. Outside traffic? Through Substrate’s atenet router with an ate-target-actor: <atespace>/<task> header – tasks don’t mint their own Service/Ingress.
The catch is managed vs self-host. Want zero cluster ops? Google’s managed agent paths (Agent Engine / Managed Agents) are the other branch. AX is own-the-data-plane, bring-any-use. I lean AX when compliance or idle-sandbox cost matters; I lean managed when the team refuses Redis + controllers + Substrate in the critical path.
Honest limitations (the stuff tutorials skip)
Preview software. README: major breaking changes before stable; external PRs paused while core settles – file Issues for feedback.
- Egress allowlists can backfire. Open issue discussion on google/ax (#345): hostname-style allowlist entries can block all TLS egress – including hosts you thought you permitted. Lab start:
host: "*"on 443, then tighten and verify with a real model call from inside the sandbox. - Secrets into Tasks are awkward. Issue #348 territory: Task env is plain values – no
valueFrom, no cleanmodelRefon the Task itself. Model resources reference secrets for platform/bootstrap use, but your agent process doesn’t get a Kubernetes-native secret mount story for free. Bake carefully or inject via a custom runner image until the API grows. - Ready doesn’t mean “command finished.” Issue #346: runner stays PID 1 after
spec.commandexits so metadata and ssh keep working. Phase can sit Running / Ready=True long after the agent process is gone. Don’t script “wait for Ready then delete” as completion without checking the actual process. - WorkspaceReady can lie about git. Issue #347: empty clones still flipping WorkspaceReady/SetupComplete. Always
ax ssh ... -- ls(or hit metadata) before trusting the condition alone.
AX also doesn’t solve governance, spend caps on model loops, or policy over what the agent may decide. Dependable runtime. Guardrails above? Still yours.
Full Substrate + AX stack for a single laptop agent? Probably not. Fleets of long-lived, untrusted, tool-calling workers that must survive preemption – yes. Few open pieces aim exactly there.
FAQ
Do I need Google Cloud or Gemini to run AX?
No. Self-hosted Apache 2.0. Kubernetes + Agent Substrate. Models: Google, Anthropic, or whatever your runner image talks to – via Model + secrets.
How is this different from just running LangGraph on Kubernetes?
LangGraph (or ADK, CrewAI) still owns the graph. Hour 3.5 OOM without a durable runtime? In-memory graph state dies with the process. AX sits underneath: sandbox isolation, egress fence, workspace materialization, snapshot suspend/resume, atenet routing. Run LangGraph as the Task command if you want – the pitch is the sandbox layer survives when the process doesn’t.
What’s the fastest way to see suspend/resume work?
After make deploy and a good ax apply (your file or examples/task.yaml, debug: true), ssh in and write a file. ax suspend task <name>, then ax resume task <name>, ssh again, confirm the file. Minimal confidence check before a multi-hour job. Resume slow or empty? Substrate worker health first – and whether the task ever left Ready during suspend. Clone github.com/google/ax, deploy on a scratch cluster, apply one Task, deliberately kill the worker, resume. That experiment beats another feature list.