Claude CodeModelsToolsOpen SourceImage and VideoAudioAgentsAppsResearchPeopleCompanies
Models reddit

Single RTX 5090: Qwen3.8-27B NVFP4 at 262K context in vLLM

TLDR

A detailed setup guide documents running Qwen3.8-27B in NVFP4 quantization on a single RTX 5090 under vLLM, achieving 77 tokens/second at short context and 64.7 tokens/second at a real 262K token context window. The post provides reproducible configuration parameters for other 5090 owners targeting long-context local inference.

Read the original