Qwen3.8-27B: serving with vLLM and SGLang

Compare quantizations and reasoning settings for Qwen3.8-27B, with serving tests on an RTX Pro 6000 GPU with 96 GB VRAM.

Player not loading? Watch on YouTube

This review tests Qwen3.8-27B as a local LLM, comparing model formats, reasoning settings and inference servers. The speaker reports benchmark gains over Qwen3.6-27B, then examines BFloat16, FP8 and NVFP4 options. He also notes MLX versions for Macs and warns that NVFP4 support depends on the GPU.

The demonstrations use an Nvidia RTX Pro 6000 with 96 GB VRAM, with compute sponsored by Dell. Those results do not establish speeds on smaller GPUs. HTML and SVG tasks show why reasoning settings matter: medium is the speaker's preferred balance, while X high spends 17,500 to 22,000 thinking tokens on a website task that fails to finish within a 32K output limit. Medium sometimes uses fewer thinking tokens than low in his tests. He also reports repeated reasoning loops in some obliterated variants.

For serving, he reports roughly 30 tokens per second for BFloat16 without speculative decoding. With built-in multi-token prediction set to three, FP8 reaches about 80 to 120, and Unsloth's quantization reaches around 120. His fastest configuration uses SGLang's Docker image, its specified NVFP4 weights and DSpark speculative decoding with a 262K context window. One long output averages 173 tokens per second; shorter outputs sometimes reach 200 to 220. The setup focuses on Blackwell GPUs, and he suggests testing llama.cpp when VRAM is more limited.