NVIDIA DLSS 5 Modders Test Hybrid FP8 and NVFP4 Model on RTX 50 GPUs, but Performance Gains Remain Small
NVIDIA DLSS 5 continues to attract attention from PC gaming enthusiasts and modders, especially because its Neural Rendering model can be demanding even on the latest GeForce RTX 50 Series graphics cards. A new community experiment has now tested whether a hybrid FP8 and NVFP4 version of the DLSS 5 model can reduce that performance cost on Blackwell GPUs.
The latest work comes through version 0.7.1 of the OptiScaler-DLSSNR-PreSR-Multipass project, an experimental modding effort focused on DLSS 5 Neural Rendering. The update introduced a hybrid model that moves some parts of the workload from FP8 to NVFP4, a lower-precision format that is natively accelerated by NVIDIA’s fifth-generation Tensor Cores in Blackwell GPUs.
On paper, the idea makes sense. NVFP4 is designed to reduce memory requirements compared to FP8 while taking advantage of dedicated hardware acceleration on RTX 50 Series cards. That raised hopes that DLSS 5 could see a noticeable boost by shifting more of its neural rendering workload to NVFP4.
One early user report suggested a dramatic improvement, with performance allegedly rising from around 130 FPS to 165 FPS on a GeForce RTX 5080 while using DLSS 5. However, the project’s official release notes paint a much more cautious picture.
According to the developer, the hybrid FP8 and NVFP4 model reduced the total DLSS 5 Neural Rendering pass time at 4K by only about 1% to 2% compared with the original FP8 model. The improvement on Blackwell GPUs was described as very minor.
That distinction is important. A 1% to 2% reduction in the Neural Rendering pass time does not mean a 1% to 2% increase in total game frame rate. DLSS 5 is only one part of the full rendering pipeline, so a small improvement in that specific pass may translate into an even smaller change during actual gameplay. In some cases, the difference may be difficult to notice or measure consistently.
There are also technical limitations. Not every model shape is supported by the hybrid NVFP4 path, meaning unsupported portions still fall back to FP8. That limits how much performance can be gained from the current approach.
The results highlight a key challenge with low-precision neural network optimization. Simply converting parts of an existing model from FP8 to NVFP4 is not enough to unlock major performance improvements. To make full use of NVFP4, the model needs proper quantization, tuning, and optimization from the ground up. Without that deeper work, the benefits may remain modest.
The developer has continued refining the hybrid implementation. Version 0.7.2 later made the combined NVFP4 hybrid path the recommended option for Blackwell GPUs. In one 120-second Baldur’s Gate 3 test, the hybrid model reached 55.16 rendered FPS, compared with 54.57 FPS using the FP8 model. While that is technically an improvement, the developer warned that the difference is extremely small and has not been proven repeatable across more tests.
For players using RTX 50 Series GPUs, the update is still interesting because it shows that the DLSS 5 modding scene is actively exploring ways to reduce the technology’s performance cost. Even if the current gain is minor, experiments like this can help uncover better optimization paths for future versions.
NVIDIA is also continuing its own DLSS 5 optimization work. The company has previously stated that DLSS 5 has become significantly faster compared with its earliest public demonstrations, and more performance improvements are planned. Official support for GeForce RTX 40 Series GPUs is also expected, which could make DLSS 5 accessible to a wider group of PC gamers.
DLSS 5 remains one of NVIDIA’s most computationally demanding upscaling and neural rendering technologies. Because of that, even small efficiency improvements matter, especially at 4K resolution where GPU workloads are already heavy.
For now, though, this hybrid FP8 and NVFP4 experiment suggests there is no simple shortcut to a massive DLSS 5 performance boost. NVFP4 has clear potential on Blackwell GPUs, but unlocking that potential will likely require deeper model-level optimization rather than a straightforward conversion from FP8.






