← Today's brief

Research · Updated 6 Oct, 09:30 am IST

LatentQuant preserves VAE policy latents under NVFP4 quantization

Why it matters for readers: It explains why quantizing parts of a model can subtly break downstream behavior and provides a tested fix.

  • Direct NVFP4 quantization without compensation degrades policy-facing latents and can collapse control performance despite acceptable reconstruction metrics.1
  • LatentQuant uses a two-stage NVFP4 QAT process: first align the quantized encoder to the high-precision encoder, then freeze it and adapt the decoder.1
  • On Wan2.1 and Wan2.2, LatentQuant preserves near-baseline control and reconstruction, achieving 95.75% success on LIBERO and 68.8% on RoboTwin.1
  • NVFP4 execution on NVIDIA B300 GPUs yields 1.17x–1.26x end-to-end VAE speedups over BF16 cuDNN.1

Get a brief like this every morning

Uzha reads hundreds of sources and gives you the stories that matter for your work, with every source linked. Free.

Get started