Skip to content

Deploy Llama 4 Locally: Scout Install Guide

Install Meta's open-weight Llama 4 Scout (109B MoE) locally with Ollama or official weights. Real VRAM numbers, gated download steps, and fixes for OOM.

6 min readIntermediate

Two install paths for Llama 4: pull a ready quant from Ollama’s llama4 library, or take gated Meta/Hugging Face checkpoints and serve them yourself (llama.cpp, vLLM, llama-models). Most installs should stop at Ollama – one binary, local store, OpenAI-compatible API on 11434. Official weights only if you need BF16/FP8 control or training hooks.

Llama 4 Scout shipped April 5, 2025 (Meta AI): 17B active / 109B total MoE, 16 experts, text+image, headline context up to 10M. Maverick is the heavier sibling (~400B total, 128 experts, ~1M context). Below: disk/VRAM, download gates, commands, min config, verify, real error strings, upgrade/uninstall.

System requirements for Llama 4 Scout

The catch is MoE memory. ~17B parameters fire per token, but every expert still loads. Plan RAM/VRAM like a ~109B model, not a dense 17B.

Resource Minimum (usable) Recommended
OS Linux, macOS, Windows 10/11 Linux or macOS (Apple Silicon unified memory)
Disk ~70 GB free (Ollama Q4 Scout ~67 GB as of library listing) 100+ GB (quants + KV headroom)
System RAM 64 GB with heavy expert offload 96-128 GB unified or GPU + host RAM
GPU / memory 24-32 GB VRAM + CPU expert offload, or Unsloth ~1.78-bit (~33.8 GB) One H100 80GB Int4 (Meta), 96 GB+ unified Mac, or multi-GPU
Deps (Ollama path) Current Ollama release NVIDIA/CUDA on Linux GPU; Metal on Apple

Maverick’s Ollama tags sit near ~245 GB – multi-GPU or very large unified memory. Scout is the single-accelerator path; Maverick is not a “just buy one more stick of RAM” bump.

Official download sources

  • Fast path:ollama.com/library/llama4 – llama4 / llama4:scout / llama4:16x17b (~67 GB Scout), llama4:maverick / llama4:128x17b (~245 GB). Sizes can move; check the library page before you free disk.
  • Meta weights: accept terms on llama.com, then the CLI in meta-llama/llama-models.
  • Hugging Face: gated ids such as meta-llama/Llama-4-Scout-17B-16E-Instruct. Accept the Llama 4 Community License on the card first. transformers support starts at v4.51.0 (as of the HF Llama 4 release notes) – pin newer if your env is stale.
  • Tight VRAM quants: Unsloth GGUFs (example: ~1.78-bit Scout ~33.8 GB) for llama.cpp on 24-32 GB cards.

License (as of the April 2025 Llama 4 Community License): not Apache. Commercial use is free under a 700M MAU cap; above that, request a separate grant from Meta. Show “Built with Llama.” The Acceptable Use Policy adds a domicile limit on multimodal rights for individuals/companies in the EU (end users of products that embed the model are carved out) – read USE_POLICY if that is you.

Install Llama 4 with Ollama (recommended)

# macOS / Linux
curl -fsSL https://ollama.com/install.sh | sh

# Windows: winget install Ollama
# or installer from https://ollama.com/download

ollama --version
ollama pull llama4:scout
# aliases: llama4 , llama4:16x17b , llama4:latest → Scout ~67GB

Chat:

ollama run llama4:scout

Maverick only with the hardware budget:

ollama run llama4:maverick

HTTP API (default http://127.0.0.1:11434):

curl http://localhost:11434/api/chat -d '{
 "model": "llama4:scout",
 "messages": [{"role": "user", "content": "Explain MoE routing in two sentences."}]
}'

First-time configuration

Default context is conservative on purpose. Bump it only while memory stays green – skip the 10M marketing number on day one. Community vLLM-style stacks OOM long before 10M unless KV is sized with care.

# One-shot 32k example
OLLAMA_NUM_CTX=32768 ollama run llama4:scout

# Persistent defaults via Modelfile
cat > Modelfile.llama4 <<'EOF'
FROM llama4:scout
PARAMETER temperature 0.6
PARAMETER top_p 0.9
PARAMETER num_ctx 16384
EOF
ollama create llama4-local -f Modelfile.llama4
ollama run llama4-local

Pro tip: Start sampling around temperature 0.6 and top_p 0.9. Change one knob at a time after you have a clean baseline.

Vision smoke test: pass a local image path in a UI/runtime that already wires Llama 4 vision. Ollama’s Scout tags are text+image capable when your build matches the library.

Verify the install works

ollama list
ollama show llama4:scout
ollama run llama4:scout "Reply with exactly: scout-ok"
ollama ps

Expect a ~67 GB-class listing, a reply that contains scout-ok, and GPU/Metal activity in ollama ps if you planned acceleration. 0% GPU with ballooning host RAM means experts spilled to CPU – tokens/sec will crawl.

That’s usually when it clicks: “17B active” was never “17B on disk.”

Still want original BF16 checkpoints, or is the quant enough for the app you’re shipping?

Alternate path: official weights + llama-models / HF

  1. Accept the license on llama.com or the HF model card.
  2. pip install llama-models huggingface_hub
  3. llama-model list then llama-model download --source meta --model-id Llama-4-Scout-17B-16E-Instruct (paste the signed URL when asked; links expire on the order of ~24h).
  4. Or hf auth login and hf download meta-llama/Llama-4-Scout-17B-16E-Instruct --local-dir ./Llama-4-Scout

Full-precision chat scripts in the llama-models README assume multiple GPUs. --quantization-mode int4_mixed aims at one 80GB GPU; fp8_mixed at two. On transformers (≥4.51.0), load Llama4ForConditionalGeneration with device_map="auto".

Common install errors and fixes

  • unknown model architecture: 'llama4' – llama.cpp / llama-cpp-python build predates Llama 4. Upgrade, then retry.
  • HF 401 / gated repo / download refused – accept the card license, hf auth login, retry. Ollama library pulls usually skip this gate.
  • OOM on load – default Q4 Scout on 24GB. Drop to Unsloth lower-bit GGUF, cut num_ctx, or offload experts in llama.cpp: -ot ".ffn_.*_exps.=CPU" (attention on GPU, MoE experts in host RAM).
  • 403 Forbidden on Meta signed URL – expired link or quota; request a fresh URL and pass it again.
  • Loads “fine,” runs like mud – weights mostly on CPU. Check ollama ps / vendor util; shrink context or add VRAM/unified memory.
  • Broken CLI after pip install llama-models – entry points drift across versions. Clean venv, follow current README names (llama-model vs older llama).

Upgrade and uninstall

Ollama: install the new app over the old one (store keeps models). ollama pull llama4:scout refreshes the tag. ollama rm llama4:scout drops weights. Uninstall the app per OS; wipe the models dir for disk (~/.ollama on Linux/macOS, user profile on Windows).

Official CLI:llama-model remove -m MODEL_ID after listing. Delete HF cache dirs if you used hf download.

No Scout→Maverick migration path – new weights, new disk plan. Pull the tag you actually need.

FAQ

Is Llama 4 Scout free for commercial products?

Under the Llama 4 Community License, yes if you stay under 700M MAU, keep “Built with Llama,” and follow the Acceptable Use Policy. Cross 700M MAU and you negotiate with Meta. Building multimodal features from an EU domicile? The AUP trims license rights for the model itself (product end users are generally carved out) – read the policy text before you ship.

Why does a “17B active” model need ~67 GB?

All experts stay resident. Compute ≈17B; memory ≈109B total. Q4 Scout lands near 67 GB. A 24 GB card needs offload or a ~1.78-bit quant.

Ollama or vLLM for production?

Ship the prototype on Ollama: minutes to /v1-style local calls. When you need multi-GPU batching and strict max-model-len, serve the HF checkpoint with vLLM after license access is done – e.g. vllm serve meta-llama/Llama-4-Scout-17B-16E-Instruct ... – and cap context on purpose. People keep pasting 10M into configs and watching KV blow up; set the limit your GPUs can actually hold.

Next: ollama pull llama4:scout, then ollama run llama4:scout "scout-ok". Clean reply? Point the app at http://127.0.0.1:11434/v1 and move on.