Increasing active parameters per token in MoE at runtime reduces reasoning tokens by 8.5%
The brief
A new paper shows that tweaking the MoE router at inference time to activate more experts per token, without any retraining, reduces the number of reasoning tokens needed on Qwen 35B A4B models by 8.5%.
Key points
- The approach requires no fine-tuning and can be applied to any sparse MoE model.
- The thread includes community tests on several compatible local models.
Sources
- redditreddit.com