Confirmed: bolting Q8 NGram into IQ4 Qwen with no speed degradation
The brief
A user confirmed that replacing the NGram layer in Qwen 3.8 Next with a Q8 precision version causes no speed degradation when running on an RTX 5090.
Key points
- The technique targets the 51B NGram layer specifically, boosting its precision while keeping the rest of the model at IQ4.
- Multiple users have replicated the result in the thread.
Sources
- redditreddit.com