HomePeopleCompaniesAI ModelsOpen SourceAgentsResearchApps
AllModelsInferenceToolingImage and Video
Open SourceInferenceMeta AI 28 Aug 2026 reddit

llama.cpp support for Qwen3.8-Flash-Next has been merged

TLDR

llama.cpp has merged support for Qwen 3.8 Flash Next, the new efficient Qwen model variant. GGUF quantizations are now available, with one user reporting 55 tokens per second on a four-GPU RTX 3090 setup using Q4 quantization. The merge makes the model accessible through the standard llama.cpp toolchain without additional patching.

Read the original