Skip to content

Install Flux Dev Locally: Official Guide [2025]

Deploy FLUX.1 [dev] from Black Forest Labs with the official repo, real VRAM numbers, HF gate fix, and CLI verification commands that work today.

6 min readIntermediate

Why bother installing Flux Dev yourself

Cloud Flux endpoints are fine until the bill shows up or the queue stalls mid-client work. Flux Dev (FLUX.1 [dev], Black Forest Labs) is the open-weight, guidance-distilled 12B model just under their Pro tier for quality – runnable on your GPU once you clear the license gate and feed it enough VRAM.

We start from the official black-forest-labs/flux inference repo (CLI + demos), then a short Diffusers check. Not a ComfyUI-only path. Commands track the main-branch README as of late 2025.

System requirements that bite you first

pyproject.toml wants Python ≥ 3.10. The optional torch extra pins torch==2.6.0 (as of the current main branch – this may have changed). NVIDIA CUDA is the realistic route; CPU-only works and feels frozen.

Resource Minimum Comfortable
GPU VRAM (bf16/FP16) ~16 GB with heavy offload 24 GB+
GPU VRAM (FP8 / quant) ~12 GB 16 GB
System RAM 32 GB 64 GB
Disk Weights (~23.8 GB) + HF cache + env headroom Spare room for re-downloads
OS Linux or Windows (WSL2 OK) Linux + recent NVIDIA driver

~24 GB peak for a plain 1024×1024 bf16 load before tricks – that’s the community consensus number people hit on 24 GB cards. NVIDIA’s NIM matrix lists 16 GB GPU / 40 GB RAM as minimal for their optimized build (support matrix). On 8-12 GB iron, assume quantized weights or aggressive offload from day one.

“Enough VRAM” is a moving target anyway. Same card, same weights: 768² can feel fine while 1024² plus a fat T5 load spikes you into swap. If you’ve ever watched a progress bar freeze with zero CUDA error, you already know the feeling.

Download source and clear the HF gate

Downloads die first. Not the clone – the gate. Weights sit at black-forest-labs/FLUX.1-dev under the FLUX.1-dev Non-Commercial License. Repo is gated. Log in, open the model card, click Agree, then:

pip install -U huggingface_hub
huggingface-cli login
# paste a token with read access

Skip Agree and you get “access is restricted” / “not in the authorized list.” I’ve watched people re-clone for an hour before noticing the button.

Pro tip: Accept the license in the same browser profile that owns the HF token. Cross-account tokens still bounce.

cd $HOME
git clone https://github.com/black-forest-labs/flux
cd flux
python3.10 -m venv .venv
source .venv/bin/activate # Windows: .venvScriptsactivate
pip install -e ".[all]"

That pulls Gradio/Streamlit demos plus the pinned torch stack. TensorRT via enroot + NVIDIA PyTorch container is a separate README path – chase it only if you already live there.

First-time configuration

First t2i run auto-fetches checkpoints into checkpoints/. Local files instead:

export FLUX_MODEL=/path/to/flux1-dev.safetensors # or HF cache layout
export FLUX_AE=/path/to/ae.safetensors

Single image, no loop:

python -m flux t2i --name flux-dev 
 --height 1024 --width 1024 
 --prompt "studio product shot of a ceramic mug, soft window light"

Interactive:

python -m flux t2i --name flux-dev --loop

Gradio after [all]:

python demo_gr.py --name flux-dev --device cuda
# --offload if VRAM is tight; --share for a public link

Verify the install works

You want: download bars (first run), sampling steps, a PNG path printed, no traceback, no OOM. Open the file. If the mug looks like a mug, you’re good.

Diffusers smoke test (separate env is fine) – different prompt on purpose so you’re not copy-pasting the HF card demo everyone else ships:

pip install -U diffusers accelerate transformers sentencepiece protobuf
python - <<'PY'
import torch
from diffusers import FluxPipeline
pipe = FluxPipeline.from_pretrained(
 "black-forest-labs/FLUX.1-dev", torch_dtype=torch.bfloat16
)
pipe.enable_model_cpu_offload()
img = pipe(
 "overcast street photo of a red bicycle leaning on a brick wall",
 height=1024, width=1024,
 guidance_scale=3.5, num_inference_steps=50,
 max_sequence_length=512,
 generator=torch.Generator("cpu").manual_seed(0),
).images[0]
img.save("flux-dev-verify.png")
print("ok", img.size)
PY

File on disk means HF auth and the GPU path both work. (Model-card defaults above: guidance_scale=3.5, 50 steps, bf16 + enable_model_cpu_offload().)

Common install errors and fixes

  • Access restricted / not authorized – License not accepted on the model page, or token is from another account. Browser Agree + huggingface-cli login again.
  • CUDA OOM (~21-24 GB spike) – Diffusers: try pipe.enable_sequential_cpu_offload() instead of (or after) enable_model_cpu_offload(). Docs push model offload first; on 24 GB cards sequential often wins. Drop to 768, or move to FP8/GGUF community builds. Official Gradio demo: offload flags.
  • Safetensors / hash / “does not seem to be a safetensors file” – Partial ~23.8 GB download. Delete the stub, resume (huggingface-cli download black-forest-labs/FLUX.1-dev --local-dir ... or wget -c). Check size before load.
  • Torch / CUDA mismatch after pip install -e ".[all]" – Extra wants torch==2.6.0. Don’t quietly upgrade torch in that venv unless the CUDA wheel matches.
  • Freeze mid-load, no error – Usually system RAM while T5-XXL and the transformer map in. Close browsers, add swap, or use an fp8 T5 path in UIs that offer it.

Commercial teams: default weight license is non-commercial for weights and derivatives. BFL sells commercial terms at bfl.ai. README also notes optional usage tracking via BFL_API_KEY on paid commercial deals – only relevant once you’re on that paper.

Upgrade, migrate, uninstall

Inference code tracks main. No classic versioned release tags the way Flux CD does.

cd $HOME/flux
git pull
source .venv/bin/activate
pip install -e ".[all]" --upgrade

Weights refresh only when you re-pull from HF or clear cache. No special checkpoint migration beyond matching the code that expects them.

deactivate
rm -rf $HOME/flux
# optional: wipe this model's HF cache
rm -rf ~/.cache/huggingface/hub/models--black-forest-labs--FLUX.1-dev

Docker/enroot TensorRT stacks die with the container. Host leftovers are whatever you installed yourself.

Clean verify PNG on disk – then what? Wire a tiny script, drop it in a queue, or A/B prompt adherence against FLUX.1 [schnell] on one fixed seed. Which workflow feels honest for your deadlines: official --loop with real product lines, or a node graph you’ll fight for a week?

FAQ

Is Flux Dev free for commercial image sales?

No – not under the default FLUX.1-dev Non-Commercial License for the weights. Get BFL’s commercial license first; read the HF LICENSE text and bfl.ai licensing before you ship.

Can I run Flux Dev on a 12 GB card?

Yes, with bruises. FP8 or GGUF transformer, text-encoder offload, modest resolution. Full bf16 1024² without offload still wants ~24 GB. Scenario that works: Diffusers + sequential CPU offload to prove the seed, then a quantized ComfyUI graph when you’re iterating all afternoon.

Official CLI vs Diffusers vs ComfyUI – which should I install first?

People treat this like a loyalty test. It isn’t. Official repo = minimal BFL-supported surface and a boring python -m flux t2i health check – start there when the goal is “does my stack work.” Diffusers wins when generation lives inside your own Python app. ComfyUI wins for node graphs, ControlNets, LoRAs – different tree entirely (flux1-dev.safetensors under diffusion_models/unet, encoders under text_encoders/clip, ae.safetensors under vae). Keep the official venv for debugging and ComfyUI for daily graphs on the same machine if you want both; neither path cancels the HF gate or the VRAM math above.