RTX 40 machine. Official NR gated to 50-series. You still want pixels.
Official DLSS 5 Neural Rendering is a generative pass: it re-paints lighting and materials on the engine frame, same resolution in and out – not Super Resolution, not Frame Generation. Nvidia’s public write-up frames the real-time stage around RTX 50 Series hardware. Ada owners, researchers, and tool builders get the cold shoulder.
OpenDLSS-NR (MIT, solo dev MAAN / maanHimself, released around 21 Sep 2026 per project NOTICE plus contemporary coverage) is a Vulkan reimplementation of the DLSS-NR 310.8.0 network. Bit-exact at all 75 block boundaries, author claims. No NVIDIA binaries, headers, or weights in the tree. This guide starts at the clone, not the press blurb.
Funny thing about “50-series only” features: once the math is public enough to re-implement, the lock reads more like a product fence than a physics law. Whether that fence should exist is a separate argument. You’re here to compile something you can toggle with a key.
Stack check, then scripts
As of the published README requirements: Windows; NVIDIA Ada or newer; a driver that actually exposes VK_KHR_cooperative_matrix, VK_NV_cooperative_matrix2, VK_EXT_shader_float8, and VK_NV_cuda_kernel_launch. Plus VS 2022 C++ x64, git, Python 3, Node/npm, Pillow if you’ll convert scenes.
- Clone
https://github.com/maanHimself/OpenDLSS-NR. - Once:
powershell -File scriptsfetch_tools.ps1 -Npm– glslang, Vulkan-Headers, volk, CMake, Ninja land intools/. Full Vulkan SDK not required. powershell -File scriptsbuild.ps1→ shaders, PTX,builddlss5vk.exe.- Live demo path:
fetch_filament.ps1→build_filament.ps1→build_demo.ps1. Budget ~15 minutes and ~6 GB under%LOCALAPPDATA%dlss5-vulkan(or your override), per demo/README.md.
builddlss5vk.exe bench --model <your-model-dir> --width 768 --height 768
builddemodlss5-demo.exe --model <your-model-dir>
Env fallbacks: DLSS5VK_MODEL, or modelsnr beside the repo. First green flag is a clean bench minimum at 768² – not a screenshot thread.
The model directory will refuse you
Repo ships neither weights nor a producer. You bring a directory with manifest.json (stages, sha256, tensors array) and packed E4M3 stage files – on the order of ~141 MiB when complete. Loader rule is brutal: block count must be exactly the 71-block graph from DLSS-NR 310.8.0. Anything else → hard refuse at load. Layout notes live in docs/weights.md; how you legally obtain proprietary weights is your problem, not the MIT tree’s.
Who actually owns a clean dump of those stages, and what license covers research use versus shipping a binary? The README will not answer that. Sit with it before you automate a scraper.
Filament window: keys that matter
Demo wires the network into Filament with per-object motion vectors and temporal feedback – the conditioning shape the model was trained for. dlss5vk itself stays single-frame; that matches how reference captures were cut, and it looks flatter once the camera moves.
Keys:N NR on/off. T temporal history. R reset history after a cut. Leave history on unless you’re matching a still fixture.
- WASD / QE, shift speed, left-drag look.
- Styles: off / natural / cinematic, or custom exposure, contrast, gamma, saturation, hue, vibrance, strength – plus the style id the net sees.
- C writes PPM dumps; resize refits the network.
Pipeline: engine HDR → LDR display proxy + three Gaussian noise lanes + reprojected previous output + five conditioning scalars → head emits RGB residual + temporal-blend logit → display is clamp(proxy + rgb/4) blended by sigmoid(logit). No stock scenes. Fox / AnimatedMorphCube glbs work after a view.json; public glTF (Lone Monk, Bistro, …) via the Node/Python helpers.
What you just ran (architecture, without the press tour)
U-net of shifted-window Swin blocks, global ViT at the bottom, 71 blocks across six pooling levels. FP8 E4M3 activations, FP16 accumulation. Same input/output resolution – generative neural renderer with injected noise and style conditioning, not an upscaler. Nvidia’s DLSS 5: Generative Neural Rendering page is the training-side story behind weights you supply; OpenDLSS-NR is the open inference graph.
| Resolution | Min network time (RTX 4070 SUPER) |
|---|---|
| 768×768 | 2.8 ms |
| 1920×1080 | 7.8 ms |
| 2560×1440 | 12.6 ms |
| 3840×2160 | 29.3 ms |
Author table (README Performance): whole-network minima over 40 frames, 241 dispatches per resolution. Sustained load → clocks bounce → medians a few percent higher. Treat those ms as best-case network cost, not free in-game headroom.
WebGPU twin
~73 ms warm at 512×512 on the same 4070 SUPER. Vulkan path on that size: ~2.7 ms. Same published bytes; no Tensor Cores, no native FP8 – E4M3 decoded in integer math, roundings placed by hand.
cd ports/browser-webgpu
npm install
node serve.mjs
Point NR_WEIGHTS at your model dir. Self-test / parity pages share the fixture contract with dlss5vk parity. Interactive demo pays WebGL↔WebGPU every frame (~15 ms readback around 1904×929, plus network), so it’s a correctness lab. Milliseconds → Vulkan+PTX. Inspectable numerics or non-NVIDIA experiments later → WebGPU.
The catch is scope
Research/demo stack. Not a drop-in into shipping titles. No DLSS-SR path here. Fast route is Windows + those Ada vendor extensions. Filament is patched for motion vectors and Vulkan interop – glTF playground, not your engine’s frame graph. Early third-party bit-exact checks of the author’s fixtures looked thin; if you have captures, run parity yourself.
After bench is green: cooperative-matrix ports on other vendors, separate game-side injector experiments, and the Nvidia ADLR report above for why the conditioning channels look the way they do.
FAQ
Does this enable official DLSS 5 inside my games on a 4070?
No. Standalone NR network + Filament demo. Game hooks are a different codebase and a different legal mess.
Black screen / model refused – first checks?
Open manifest.json. You want totals.blockCount: 71 and stage hashes that match the packed E4M3 files. Wrong graph? Loader refuses; no fallback. Next: DLSS5VK_LIST_EXTENSIONS=1. Missing VK_NV_cuda_kernel_launch or float8 almost always means driver or pre-Ada GPU. Demo still dark? Confirm --model / DLSS5VK_MODEL / modelsnr and that buildscenes has a glTF folder with view.json. Build tree present does not imply scene present – nothing ships in-tree.
Vulkan or WebGPU for a first run?
Ada laptop already has VS 2022? Finish native demo first. You get temporal history, motion vectors, and the 7.8 ms-class 1080p path from the author’s minima table. Browser route is for stepping numerics or parity without PTX. Near-1080p in WebGPU feels like a slideshow because WGSL is emulating tensor-core MMA – that’s the design tradeoff, not a broken install.
Clone → fetch_tools → build.ps1 → bench at 768² against a model directory you have rights to. Green minimum? Build the Filament demo and hit N. That’s enough signal you control the graph.