Research · Updated 6 Oct, 09:30 am IST
LatentQuant preserves VAE policy latents under NVFP4 quantization
Why it matters for readers: It explains why quantizing parts of a model can subtly break downstream behavior and provides a tested fix.
- Direct NVFP4 quantization without compensation degrades policy-facing latents and can collapse control performance despite acceptable reconstruction metrics.1
- LatentQuant uses a two-stage NVFP4 QAT process: first align the quantized encoder to the high-precision encoder, then freeze it and adapt the decoder.1
- On Wan2.1 and Wan2.2, LatentQuant preserves near-baseline control and reconstruction, achieving 95.75% success on LIBERO and 68.8% on RoboTwin.1
- NVFP4 execution on NVIDIA B300 GPUs yields 1.17x–1.26x end-to-end VAE speedups over BF16 cuDNN.1
Get a brief like this every morning
Uzha reads hundreds of sources and gives you the stories that matter for your work, with every source linked. Free.
Get started