llama.cpp support for Qwen3.8-Flash-Next has been merged
TLDR
llama.cpp has merged support for Qwen 3.8 Flash Next, the new efficient Qwen model variant. GGUF quantizations are now available, with one user reporting 55 tokens per second on a four-GPU RTX 3090 setup using Q4 quantization. The merge makes the model accessible through the standard llama.cpp toolchain without additional patching.