Skip to content

Run Qwen 3.8 Flash Next 125B on RTX 4090 at 100T/s

Run Qwen 3.8 Flash Next (125B) on consumer hardware like an RTX 4090 near 100 tok/s. Strata + RAM offload setup, real numbers, gotchas.

7 min readBeginner

You don’t need a server rack to run a 125B-class model. You need a memory hierarchy that stops treating VRAM as the only place weights can live.

Feeds this week keep looping the same clip: Qwen 3.8 Flash Next (125B) on one gaming GPU, VRAM barrier “dead.” I wanted the interactive version – near 100 tok/s on something like an RTX 4090 – not a 20 t/s slideshow. Wrong turns first. Then the path that actually moved the counter.

Reader scenario: the 4090 that “shouldn’t” fit 125B

Normal box. RTX 4090 24GB, 64-128GB system RAM, fast SSD. Cloud invoices hurt. Dense 70B was fine. Then Flash-Next shipped August 26, 2026 as Alibaba’s open-weight multimodal MoE: 125B total params, only ~6B active per token, plus a large n-gram table that is happy living off-GPU.

Vendor memory tables still scare people – full BF16 is in the hundreds of GB per Unsloth’s guide. Community llama.cpp threads swung the other way: about 21 tok/s decode with expert offload and a 250k window. Usable. Not “type and the reply keeps pace.” That gap is why this write-up exists.

What Flash-Next actually is (and why speed is possible)

Per the official Qwen model card, you get 125B main weights, ~6B activated (10 routed experts + 1 shared out of 512), ~51B n-gram embeddings, a small MTP head (~4B), native context 262,144 (extensible toward 1M). Hybrid Gated DeltaNet + Qwen Sparse Attention is what keeps long context cheaper than a pure dense stack.

Wrong question: “does 125B fit in 24GB?” Better question: can attention and hot experts stay on the 4090 while cold experts and n-grams sit in DDR? Once that clicks, consumer boxes stop looking undersized.

Unsloth’s GGUF memory table (total RAM+VRAM or unified, as published with the model docs) runs roughly 75GB at 1-bit, ~79GB 2-bit, ~90GB 3-bit, ~96-114GB for solid 4-bit, and ~355GB BF16. Their analysis pegs 1-bit around ~80% top-1 accuracy versus BF16. That’s the real control knob – not hunting a second 24GB card on day one.

Practical setup: high tok/s without compiling all night

Chasing the headline speeds people pin to “100T/s on a 4090-class card”? Start with Strata, not a from-scratch llama.cpp build. README pitch: one GPU (NVIDIA/AMD 12GB+), 32-64GB+ RAM, ~80GB free disk, Windows or Linux, OpenAI/Anthropic-compatible API on localhost:8080.

  1. Current GPU driver only.
  2. Clone or unzip Strata. Windows: START-HERE.bat. Linux: ./setup.sh.
  3. Defaults are fine unless you already know better – Flash-Next family, light quant (Q2_0 when you want max write speed; IQ-style 2-bit class on 64GB RAM), modest context, multimodal bits only if you need them.
  4. Download is large (~80GB disk budget). First cold start can pin tens of GB into RAM and freeze the desktop 1-3 minutes. Looks like a crash. It isn’t. Don’t kill it.
  5. Browser lands on http://127.0.0.1:8080. Point coding agents at the local OpenAI- or Anthropic-shaped endpoints.

~94 tok/s write and multi-thousand tok/s prompt read – that is Strata’s own README figure for an RTX 5070 12GB + 64GB RAM on Q2_0 (short/32K-style runs, as published there). Same doc estimates a 24GB RTX 3090 around 100-140 tok/s write. A 4090 lives in that band when system RAM is full and the quant stays light. Your number tracks host bandwidth and quant harder than the badge on the GPU.

Pro tip: Only 32GB RAM? Enable low-RAM mode and accept SSD expert streaming. Moving to 64GB is usually the biggest speed jump on this model – more than shopping another GPU. Community low-RAM boxes land far slower than full-RAM pins; treat 64GB as the practical floor for “feels fast.”

Prefer Unsloth’s desktop/hub flow? Pull a UD quant of Flash-Next and leave MTP available. Friendly for chat UI. Peak single-stream decode on a lone 4090 still tends to trail a tuned Strata stack in the public numbers people actually post.

Funny thing about local 125B: the moment it feels snappy, you stop caring that the parameter count looks datacenter-shaped. Until the first freeze on load, anyway.

Advanced: llama.cpp, -cmoe, MTP

Custom context, multi-user, research flags – then GGUFs from unsloth/Qwen3.8-Flash-Next-GGUF plus a CUDA llama.cpp (Unsloth documents a qwen4exp/mtp-style branch when mainline lags; verify current docs before you build).

# example shape - paths and branch names change; check Unsloth docs
hf download unsloth/Qwen3.8-Flash-Next-GGUF --include "*UD-Q4_K_XL*" --local-dir ./qwen38

./llama-server 
 -m ./qwen38/UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf 
 -c 131072 -ngl 99 -fa on 
 -b 4096 -ub 4096 -cmoe 
 -md ./qwen38/MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf 
 --spec-type draft-mtp --spec-draft-n-max 3 
 --host 127.0.0.1 --port 8080

-cmoe (and related expert-offload flags) parks experts in system RAM and keeps attention on the GPU – the reason viral footprints show roughly 11-18GB VRAM even at huge context. Community 4090 + ~110GB DDR4 recipes with UD-Q4_K_XL land near 21 t/s decode and about 364 t/s prefill at 250k context. Batch/ubatch bumps (4096 is a common recipe) are what people crank when prefill is the bottleneck.

MTP is the other multiplier. Unsloth’s guide puts the shared draft head around ~2.79GB (Q8_0) and claims roughly 1.3-1.7× faster inference; they’ve shown on the order of ~170 t/s on a single RTX 6000 PRO versus a ~100 t/s baseline. Same flag set does not gift that uplift on every workload – more on that below.

Honest limitations (read before you brag)

Plain llama.cpp on a 4090 with a quality 4-bit is still often ~21-27 t/s decode in public write-ups. One detailed Windows report with 192GB DDR5 averaged 27.13 t/s mean decode across 5×600-token runs – after fixing RAM. The “100T/s” zone is Strata-class engines + lighter quants + enough fast host memory, not the free default.

The catch is RAM speed, not only capacity. That same 4090 write-up sat near the low-20s until BIOS moved DDR5 from 4200 to rated 5600 MT/s, then the mean climbed to ~27 t/s. XMP/EXPO stability testing matters more here than another 24GB of VRAM when experts already live in host memory.

MTP draft acceptance runs hot on code-shaped tokens and much colder on free prose. The --spec-draft-n-max that feels like free speed in a repo agent can break even – or slow you – on chatty writing. Long context is natively huge, but a filled 128K+ session still taxes TTFT and KV; budget the window for the job. Hosted Qwen3.8-Flash (Flash-Next lineage) pricing announced around $0.16/M input and $0.47/M output as of late August 2026 remains the sanity check when the SSD is full and the deadline is tomorrow.

Is local 125B on one gaming card “production”? Solo agent loop, private code – yes. Ten concurrent users – you’re back to multi-GPU or the API.

FAQ

Can an RTX 4090 really hit ~100 tokens/s on Flash-Next?

Yes – with Strata-class expert caching, a light quant (Q2_0 / IQ2-class), and 64GB+ of fast system RAM. Strata’s README bands a 24GB 3090 around 100-140 t/s write; treat a 4090 as the same neighborhood. Heavy 4-bit or plain llama.cpp without those memory tricks? Expect the 20s.

How much RAM do I actually need?

Total memory, not VRAM alone. Unsloth’s table (as of their Flash-Next docs): ~75GB at 1-bit up through ~96-114GB for comfortable 4-bit. Practice: 64GB runs popular Strata sizes well; 32GB forces low-RAM SSD streaming and feels several times slower; 96GB+ is the leave-the-browser-open tier. Pair it with NVMe so first pin and fallback reads don’t crawl.

Strata or llama.cpp – which should I install tonight?

Tonight: Strata. Installer, light quant, wire your OpenAI-compatible client, send a real 4k-token coding prompt, watch your tok/s counter. Tomorrow, if you need custom forks, multi-slot serving, or flag archaeology, build llama.cpp on Unsloth GGUFs, add -cmoe, prove baseline decode, then layer MTP. Mixing both is normal – Strata for daily agents, llama.cpp when you’re measuring.

Clone Strata once. Run the installer. Tune quant and RAM before you buy another card.