Models
reddit
Single RTX 5090: Qwen3.8-27B NVFP4 at 262K context in vLLM
TLDR
A detailed setup guide documents running Qwen3.8-27B in NVFP4 quantization on a single RTX 5090 under vLLM, achieving 77 tokens/second at short context and 64.7 tokens/second at a real 262K token context window. The post provides reproducible configuration parameters for other 5090 owners targeting long-context local inference.