A new community fork of OptiScaler DLSS 5 reduced the measured cost of Neural Rendering on a GeForce RTX 4070 from 13.0 ms to 2.9 ms. Furthermore, the author published the test on October 6, 2026, using Control Resonant at 1440p with DLSS Balanced, excluding Frame Generation. As a result, the base rate rose from 34.4 FPS in the implementation used as a reference to 55.3 FPS in the Performance preset.
The project is experimental and does not represent an official NVIDIA implementation. Even so, the result is noteworthy because it combines two strategies: running the neural model before Super Resolution and temporally reusing part of the result between frames.
OptiScaler DLSS 5 drops from 13 ms to 2.9 ms
In the benchmark the author released, DLSS 5 disabled delivered 62.2 FPS. Meanwhile, the path he calls “vanilla”, with the community integration of OptiScaler NR, dropped to 34.4 FPS and consumed 13.0 ms per frame just in the neural stage.
On the other hand, the Quality preset, which executes Neural Rendering before Super Resolution on all frames, scored 46.7 FPS and 5.5 ms. Next, the Performance preset added temporal caching and reached 55.3 FPS with a 2.9 ms cost. Therefore, in the developer's test, the penalty on the base rate dropped from about 28 FPS to approximately 7 FPS.

Official NVIDIA panel shows DLSS 5 model, intensity, and masking controls.
How the new fork reduces the cost of Neural Rendering
First, the fork moves the model to before DLSS Super Resolution. With DLSS Balanced at 1440p, the author claims that the model now works on 1484 × 835 instead of 2560 × 1440. Thus, there are approximately three times fewer pixels to process in this step.
Furthermore, the project implements a temporal cache. Instead of running the full model on every frame, it stores changes as a per-pixel ratio, reprojects this data with motion vectors, and validates the reuse with depth and color. By default, the model runs every two frames, although the interval can be extended.

Experimental implementations help measure the computational cost of different Neural Rendering paths.
At the same time, the fork includes temporal stabilization, despeckling, and smoothing of regional brightness variations to reduce flicker. In addition, there is automatic white point, one to three-pass multipass, and Default, Natural, and Cinematic styles.
Benchmark is still limited to an RTX 4070
These numbers require caution because they reflect an individual test. So far, the author has focused tests on Control Resonant and a single GPU, the RTX 4070. Furthermore, Pre-SR and temporal cache only work in DirectX 12 at this stage.

The gain observed in the fork should not be automatically extrapolated to other GPUs and games.
Therefore, it is important to separate this work from the official technology. The repository does not include nvngx_dlssnr.dll; therefore, the user needs to provide their own DLL. The developer himself classifies the project as unofficial and unrelated to NVIDIA.
What the result changes for the community DLSS 5
In this context, the test reinforces that the order of the steps can significantly alter the cost of Neural Rendering. Furthermore, previous experiments with Pre-SR already pointed to gains when processing a lower internal resolution. However, the new fork adds temporal reuse and measures the combination in its own presets.
For comparison, the Allves Games has already covered the DLSS 5 Swapper 2.2.9 with OptiScaler Pre-SR Multipass. We also showcase Neural Lens 0.6.0, another community approach aimed at reducing the weight of neural processing.
For now, read the 2.9 ms as the author's result in a specific scenario, not as guaranteed performance. Even so, the reduction compared to the 13.0 ms of the reference makes the fork a relevant technical evolution for community experiments with DLSS 5.
Sources: Skynizz/optiscaler-dlss5 on GitHub; author's technical post on r/DLSS.