Skip to content

AX Google Open Agentic Orchestrator: Hands-On Guide

AX - Google's Open Agentic Orchestrator just dropped. Run durable agent tasks on Kubernetes with suspend/resume, workspaces, and network fences. Setup walkthrough.

7 min readBeginner

Your long-running coding agent just finished hour three of a multi-repo refactor. The pod gets evicted. Normally everything vanishes. With AX – Google’s Open Agentic Orchestrator, you suspend, resume, and the sandbox picks up with files and state still there. That’s the end state this guide gets you to – on your own cluster, with YAML you control.

AX (Agent Executor) showed up open source from Google around mid-May 2026. Still hot. Still early. Community threads latched onto the four-hour crash story; under the hood it’s a declarative control plane on Agent Substrate. Agents become suspendable actors, not fragile Python processes you babysit.

What you’re actually buying with AX

Not a LangGraph clone. The Google Cloud announcement frames it as self-hosted durable execution: isolation, session consistency, connection recovery, trajectory branching. Keep ADK, LangChain, a custom container, MCP tools – AX takes the lifecycle.

If you already think in Kubernetes objects, the muscle memory transfers. API group ax.io/v1alpha1 (early preview – this can move). Four primitives:

Primitive Job
Task Isolated sandbox, image/command, CPU/memory limits, workspace mounts, gateway ref
Workspace Git clones, MCP servers/registries, skills, optional plain-English goal for generative setup
Gateway Listeners the task exposes + egress host/port allowlist
Model Named provider/model/params + Kubernetes secret for the API key

Everything sits in an atespace (default: default). Idle actors checkpoint on Substrate and wake fast – the project site and docs cite sub-second resume (sometimes sub-500ms) and dense multiplexing, so you aren’t paying full compute rent while an agent waits on a tool or a human.

Practical setup: from zero to a running task

Checklist before you touch YAML: Go for the CLI, a Kubernetes cluster, ko, a registry your nodes can pull, Agent Substrate Control API reachable (in-cluster default as of the README: api.ate-system.svc.cluster.local:443). Early preview. Breaking changes are expected.

  1. CLI: go install github.com/google/ax/cmd/ax@latest – binary lands in $(go env GOPATH)/bin; put that on PATH.
  2. Control plane from a clone of github.com/google/ax: make deploy AX_IMAGE_REPO=<your-registry>. As of the current quick start that brings components up in ax-system (Redis + AX server/controllers in the documented path).
  3. Model secret, e.g. kubectl create secret generic gemini-api-secret --from-literal=GEMINI_API_KEY="...".
  4. Apply one multi-doc YAML (Task + Workspace + Gateway + Model). Then ax get tasks, ax watch task <name>, and – if debug: true – ax ssh <name> -- ls -la /workspace.

Starter below skips the stock chalk/golang demos. Two workspaces: app code + shared tools. A generative goal on first boot so the runner can finish the toolchain instead of you baking every image by hand. Model id and runner digest are examples from docs/examples as of mid-2026 preview – pin what your cluster actually pulls; digests rotate.

apiVersion: ax.io/v1alpha1
kind: Workspace
metadata:
 name: app-src
 atespace: default
spec:
 git:
 - name: origin
 repo: "https://github.com/your-org/your-service.git"
 branch: "main"
---
apiVersion: ax.io/v1alpha1
kind: Workspace
metadata:
 name: team-tools
 atespace: default
spec:
 mcp:
 servers:
 - name: git-tools
 endpoint: "http://git-mcp.default.svc.cluster.local:8080"
---
apiVersion: ax.io/v1alpha1
kind: Gateway
metadata:
 name: default-gateway
 atespace: default
spec:
 listeners:
 - name: http
 port: 8080
 protocol: HTTP
 - name: grpc
 port: 8494
 protocol: gRPC
 egress:
 allowlist:
 hosts:
 - host: "*"
 port: 443
---
apiVersion: ax.io/v1alpha1
kind: Model
metadata:
 name: default-model
 atespace: default
spec:
 provider: google
 model: gemini-3.8-flash
 secretKey:
 name: gemini-api-secret
 key: GEMINI_API_KEY
 parameters:
 temperature: 0.9
---
apiVersion: ax.io/v1alpha1
kind: Task
metadata:
 name: refactor-bot
 atespace: default
spec:
 image: "gcr.io/ax-substrate/ate-images/ax-task-runner@sha256:3a0dea6ad8b55278685db58aca6e37dc4ba04056831d45bef3aaeafdca43cac6"
 env:
 - name: ENVIRONMENT
 value: "staging"
 resources:
 requests:
 cpu: "500m"
 memory: "1Gi"
 limits:
 cpu: "2"
 memory: "4Gi"
 workspaces:
 - name: app-src
 path: "/workspace"
 goal: "Install the project toolchain and make tests runnable"
 - name: team-tools
 path: "/workspace/tools"
 gateway:
 name: default-gateway
 debug: true

Apply: ax apply -f task.yaml. Watch WorkspaceReady and Ready conditions – not only the phase string – before you assume the agent can work.

Pro tip: Leave debug: true on while learning. Guest services stay off without it, so ax ssh is dead. Flip it off when you stop poking production-shaped sandboxes.

Day-two ops that matter

Suspend/resume is why most people showed up.

ax suspend task refactor-bot
ax resume task refactor-bot
ax ssh refactor-bot -- cat /workspace/notes.txt

Checkpointed actor state returns. Files you wrote pre-suspend are still on disk. Restart-the-prompt vs continue-the-job – different products pretending to be the same demo.

Generative workspaces are the other lever. Plain-language goal on a binding gets handed to a bootstrap agent on first boot (docs describe an Antigravity-style path). Needs model credentials and a bootstrap timeout – default around ten minutes; override path in sandbox docs is AX_BOOTSTRAP_TIMEOUT as of writing (preview: names can change). Use it when every task would otherwise reinstall the same toolchain from scratch.

Multi-workspace binding splits “code under change” from “shared MCP/skills.” First entry is the command working directory. Full workspace readiness waits until every binding is prepared. Outside traffic? Through Substrate’s atenet router with an ate-target-actor: <atespace>/<task> header – tasks don’t mint their own Service/Ingress.

The catch is managed vs self-host. Want zero cluster ops? Google’s managed agent paths (Agent Engine / Managed Agents) are the other branch. AX is own-the-data-plane, bring-any-use. I lean AX when compliance or idle-sandbox cost matters; I lean managed when the team refuses Redis + controllers + Substrate in the critical path.

Honest limitations (the stuff tutorials skip)

Preview software. README: major breaking changes before stable; external PRs paused while core settles – file Issues for feedback.

  • Egress allowlists can backfire. Open issue discussion on google/ax (#345): hostname-style allowlist entries can block all TLS egress – including hosts you thought you permitted. Lab start: host: "*" on 443, then tighten and verify with a real model call from inside the sandbox.
  • Secrets into Tasks are awkward. Issue #348 territory: Task env is plain values – no valueFrom, no clean modelRef on the Task itself. Model resources reference secrets for platform/bootstrap use, but your agent process doesn’t get a Kubernetes-native secret mount story for free. Bake carefully or inject via a custom runner image until the API grows.
  • Ready doesn’t mean “command finished.” Issue #346: runner stays PID 1 after spec.command exits so metadata and ssh keep working. Phase can sit Running / Ready=True long after the agent process is gone. Don’t script “wait for Ready then delete” as completion without checking the actual process.
  • WorkspaceReady can lie about git. Issue #347: empty clones still flipping WorkspaceReady/SetupComplete. Always ax ssh ... -- ls (or hit metadata) before trusting the condition alone.

AX also doesn’t solve governance, spend caps on model loops, or policy over what the agent may decide. Dependable runtime. Guardrails above? Still yours.

Full Substrate + AX stack for a single laptop agent? Probably not. Fleets of long-lived, untrusted, tool-calling workers that must survive preemption – yes. Few open pieces aim exactly there.

FAQ

Do I need Google Cloud or Gemini to run AX?

No. Self-hosted Apache 2.0. Kubernetes + Agent Substrate. Models: Google, Anthropic, or whatever your runner image talks to – via Model + secrets.

How is this different from just running LangGraph on Kubernetes?

LangGraph (or ADK, CrewAI) still owns the graph. Hour 3.5 OOM without a durable runtime? In-memory graph state dies with the process. AX sits underneath: sandbox isolation, egress fence, workspace materialization, snapshot suspend/resume, atenet routing. Run LangGraph as the Task command if you want – the pitch is the sandbox layer survives when the process doesn’t.

What’s the fastest way to see suspend/resume work?

After make deploy and a good ax apply (your file or examples/task.yaml, debug: true), ssh in and write a file. ax suspend task <name>, then ax resume task <name>, ssh again, confirm the file. Minimal confidence check before a multi-hour job. Resume slow or empty? Substrate worker health first – and whether the task ever left Ready during suspend. Clone github.com/google/ax, deploy on a scratch cluster, apply one Task, deliberately kill the worker, resume. That experiment beats another feature list.