25 Aug 2026
Signal Headquarters
Vol. I
No. 241
· · 3 min read

NVIDIA is pre-training Nemotron models in FP4, a precision level no public effort had reached before

Bryan Catanzaro disclosed that NVIDIA was pre-training its Nemotron 3 Super and Ultra models using FP4 precision before either model launched publicly. Technical reports for both models now confirm it, marking a genuine first in published large-scale pre-training.

Bryan Catanzaro, a vice president at NVIDIA, disclosed that the company was using FP4 numerical precision to pre-train its Nemotron 3 Super and Ultra models, noting it was something that “hasn’t been done publicly anyway.” That phrasing, hedged only slightly, was a meaningful signal: FP4 pre-training at the scale of a frontier language model had not appeared in any public technical record at the time he said it.

It has since appeared. NVIDIA’s official technical reports for both Nemotron 3 Super, published in April 2026, and Nemotron 3 Ultra, published in June 2026, confirm that both models were pre-trained using NVFP4. The reports describe it as a first for the Nemotron 3 family. The external record, in other words, matches Catanzaro’s claim with precision.

The significance of that match goes beyond a single product line. Pre-training is the most compute-intensive phase of building a large language model. It is the phase where numerical precision matters most, because errors compound across billions of parameter updates. The field has long treated lower-precision formats with caution during pre-training, using them freely for inference, where a model processes inputs, but conservatively during the training runs themselves. FP4 sits below the formats that have become standard even for aggressive mixed-precision training. Deploying it at the pre-training stage, and doing so at the scale of a Super or Ultra model, is not a routine engineering choice.

Um these days, you know, we're uh pre-training our um Neotron 3 um Super and Ultra models using FP4, um which is a thing that, you know, h hasn't been done publicly anyway. Bryan Catanzaro

What makes the Nemotron case worth examining carefully is the sequence. Catanzaro’s disclosure came before either technical report was public. The corroboration arrived after. That ordering matters because it establishes the claim as early and specific, not as a post-hoc reading of a published paper. When a company discloses a non-obvious technical choice before the supporting documentation exists, and the documentation subsequently confirms the choice, the signal is reliable.

NVIDIA’s decision to use NVFP4 for pre-training also reflects the company’s position in the hardware stack. NVIDIA designs the accelerators on which most large models are trained, and it also builds its own models on that hardware. That vertical integration gives NVIDIA an incentive, and an ability, to push numerical precision further than groups that are working with hardware they did not design. Whether NVFP4 pre-training produces quality equivalent to or better than higher-precision alternatives is a question the technical reports address; the broader point is that the threshold has moved.

The pattern that Catanzaro’s disclosure fits is one where hardware-native choices made inside large, vertically integrated companies establish new defaults before the broader research community has had time to validate or replicate them. FP8 pre-training followed a similar arc before it became common. FP4 is now on the same trajectory. The Nemotron 3 reports will be read by teams at other organizations who are deciding whether to attempt FP4 pre-training themselves, and the answer NVIDIA has supplied, through both Catanzaro’s early disclosure and the subsequent technical documentation, is that it is possible and that it has been done.

What remains to be seen is how quickly the approach spreads, and whether the quality outcomes reported for Nemotron 3 Super and Ultra hold at different scales, on different hardware, and in the hands of teams that are not simultaneously the chip designer and the model trainer. Those questions are open. What is not open, following the April and June 2026 technical reports, is whether FP4 pre-training has crossed from speculation into demonstrated practice. It has.

The Editor, for the readers of Signal Headquarters

From the Archive