Player not loading? Watch on YouTube
This tutorial explains NVIDIA's NVFP4 format, then follows a Qwen 3.8 27B conversion on DGX Spark using ModelOpt through the Hermes Agent desktop app. Hermes uses a hosted GPT model to help prepare the conversion; the converted model runs locally. It connects local LLM memory requirements with the arithmetic supported by the hardware.
The explanation compares evenly spaced INT4 values with E2M1 floating-point values, which have finer spacing near zero. It covers NVFP4's FP8 scale for each group of 16 weights and contrasts it with MXFP4's power-of-two scaling. Including scale overhead, the speaker describes NVFP4 as roughly 4.5 bits per weight. He attributes its compute advantage to Blackwell Tensor Cores and cautions that older GPUs may not deliver a speed increase over INT4.
The hands-on section starts with a downloaded BF16 checkpoint. The workflow pins the ModelOpt source, prepares a Python environment with uv, checks recipe coverage and adds memory monitoring. A conservative MLP-only recipe quantizes those weights and activations while retaining other parts in BF16. Calibration uses 128 text samples before export and validation.
The speaker reports an 18-minute conversion, a checkpoint reduction from 54.56 GB to 28.86 GB, and a 1.76-times speed increase. He reports matching accuracy in a quick GSM8K check, but these limited tests do not establish unchanged quality across tasks. The demo ends with a dedicated Hermes Agent profile using the converted model; the speaker cautions that the unoptimized timings are comparison results.