Why bother installing Flux Dev yourself
Cloud Flux endpoints are fine until the bill shows up or the queue stalls mid-client work. Flux Dev (FLUX.1 [dev], Black Forest Labs) is the open-weight, guidance-distilled 12B model just under their Pro tier for quality – runnable on your GPU once you clear the license gate and feed it enough VRAM.
We start from the official black-forest-labs/flux inference repo (CLI + demos), then a short Diffusers check. Not a ComfyUI-only path. Commands track the main-branch README as of late 2025.
System requirements that bite you first
pyproject.toml wants Python ≥ 3.10. The optional torch extra pins torch==2.6.0 (as of the current main branch – this may have changed). NVIDIA CUDA is the realistic route; CPU-only works and feels frozen.
| Resource | Minimum | Comfortable |
|---|---|---|
| GPU VRAM (bf16/FP16) | ~16 GB with heavy offload | 24 GB+ |
| GPU VRAM (FP8 / quant) | ~12 GB | 16 GB |
| System RAM | 32 GB | 64 GB |
| Disk | Weights (~23.8 GB) + HF cache + env headroom | Spare room for re-downloads |
| OS | Linux or Windows (WSL2 OK) | Linux + recent NVIDIA driver |
~24 GB peak for a plain 1024×1024 bf16 load before tricks – that’s the community consensus number people hit on 24 GB cards. NVIDIA’s NIM matrix lists 16 GB GPU / 40 GB RAM as minimal for their optimized build (support matrix). On 8-12 GB iron, assume quantized weights or aggressive offload from day one.
“Enough VRAM” is a moving target anyway. Same card, same weights: 768² can feel fine while 1024² plus a fat T5 load spikes you into swap. If you’ve ever watched a progress bar freeze with zero CUDA error, you already know the feeling.
Download source and clear the HF gate
Downloads die first. Not the clone – the gate. Weights sit at black-forest-labs/FLUX.1-dev under the FLUX.1-dev Non-Commercial License. Repo is gated. Log in, open the model card, click Agree, then:
pip install -U huggingface_hub
huggingface-cli login
# paste a token with read access
Skip Agree and you get “access is restricted” / “not in the authorized list.” I’ve watched people re-clone for an hour before noticing the button.
Pro tip: Accept the license in the same browser profile that owns the HF token. Cross-account tokens still bounce.
cd $HOME
git clone https://github.com/black-forest-labs/flux
cd flux
python3.10 -m venv .venv
source .venv/bin/activate # Windows: .venvScriptsactivate
pip install -e ".[all]"
That pulls Gradio/Streamlit demos plus the pinned torch stack. TensorRT via enroot + NVIDIA PyTorch container is a separate README path – chase it only if you already live there.
First-time configuration
First t2i run auto-fetches checkpoints into checkpoints/. Local files instead:
export FLUX_MODEL=/path/to/flux1-dev.safetensors # or HF cache layout
export FLUX_AE=/path/to/ae.safetensors
Single image, no loop:
python -m flux t2i --name flux-dev
--height 1024 --width 1024
--prompt "studio product shot of a ceramic mug, soft window light"
Interactive:
python -m flux t2i --name flux-dev --loop
Gradio after [all]:
python demo_gr.py --name flux-dev --device cuda
# --offload if VRAM is tight; --share for a public link
Verify the install works
You want: download bars (first run), sampling steps, a PNG path printed, no traceback, no OOM. Open the file. If the mug looks like a mug, you’re good.
Diffusers smoke test (separate env is fine) – different prompt on purpose so you’re not copy-pasting the HF card demo everyone else ships:
pip install -U diffusers accelerate transformers sentencepiece protobuf
python - <<'PY'
import torch
from diffusers import FluxPipeline
pipe = FluxPipeline.from_pretrained(
"black-forest-labs/FLUX.1-dev", torch_dtype=torch.bfloat16
)
pipe.enable_model_cpu_offload()
img = pipe(
"overcast street photo of a red bicycle leaning on a brick wall",
height=1024, width=1024,
guidance_scale=3.5, num_inference_steps=50,
max_sequence_length=512,
generator=torch.Generator("cpu").manual_seed(0),
).images[0]
img.save("flux-dev-verify.png")
print("ok", img.size)
PY
File on disk means HF auth and the GPU path both work. (Model-card defaults above: guidance_scale=3.5, 50 steps, bf16 + enable_model_cpu_offload().)
Common install errors and fixes
- Access restricted / not authorized – License not accepted on the model page, or token is from another account. Browser Agree +
huggingface-cli loginagain. - CUDA OOM (~21-24 GB spike) – Diffusers: try
pipe.enable_sequential_cpu_offload()instead of (or after)enable_model_cpu_offload(). Docs push model offload first; on 24 GB cards sequential often wins. Drop to 768, or move to FP8/GGUF community builds. Official Gradio demo: offload flags. - Safetensors / hash / “does not seem to be a safetensors file” – Partial ~23.8 GB download. Delete the stub, resume (
huggingface-cli download black-forest-labs/FLUX.1-dev --local-dir ...orwget -c). Check size before load. - Torch / CUDA mismatch after
pip install -e ".[all]"– Extra wantstorch==2.6.0. Don’t quietly upgrade torch in that venv unless the CUDA wheel matches. - Freeze mid-load, no error – Usually system RAM while T5-XXL and the transformer map in. Close browsers, add swap, or use an fp8 T5 path in UIs that offer it.
Commercial teams: default weight license is non-commercial for weights and derivatives. BFL sells commercial terms at bfl.ai. README also notes optional usage tracking via BFL_API_KEY on paid commercial deals – only relevant once you’re on that paper.
Upgrade, migrate, uninstall
Inference code tracks main. No classic versioned release tags the way Flux CD does.
cd $HOME/flux
git pull
source .venv/bin/activate
pip install -e ".[all]" --upgrade
Weights refresh only when you re-pull from HF or clear cache. No special checkpoint migration beyond matching the code that expects them.
deactivate
rm -rf $HOME/flux
# optional: wipe this model's HF cache
rm -rf ~/.cache/huggingface/hub/models--black-forest-labs--FLUX.1-dev
Docker/enroot TensorRT stacks die with the container. Host leftovers are whatever you installed yourself.
Clean verify PNG on disk – then what? Wire a tiny script, drop it in a queue, or A/B prompt adherence against FLUX.1 [schnell] on one fixed seed. Which workflow feels honest for your deadlines: official --loop with real product lines, or a node graph you’ll fight for a week?
FAQ
Is Flux Dev free for commercial image sales?
No – not under the default FLUX.1-dev Non-Commercial License for the weights. Get BFL’s commercial license first; read the HF LICENSE text and bfl.ai licensing before you ship.
Can I run Flux Dev on a 12 GB card?
Yes, with bruises. FP8 or GGUF transformer, text-encoder offload, modest resolution. Full bf16 1024² without offload still wants ~24 GB. Scenario that works: Diffusers + sequential CPU offload to prove the seed, then a quantized ComfyUI graph when you’re iterating all afternoon.
Official CLI vs Diffusers vs ComfyUI – which should I install first?
People treat this like a loyalty test. It isn’t. Official repo = minimal BFL-supported surface and a boring python -m flux t2i health check – start there when the goal is “does my stack work.” Diffusers wins when generation lives inside your own Python app. ComfyUI wins for node graphs, ControlNets, LoRAs – different tree entirely (flux1-dev.safetensors under diffusion_models/unet, encoders under text_encoders/clip, ae.safetensors under vae). Keep the official venv for debugging and ComfyUI for daily graphs on the same machine if you want both; neither path cancels the HF gate or the VRAM math above.