DeepSeek V4-1 Flash Is Out: 552B MoE with 1M Token Context
The brief
DeepSeek released V4-1 Flash, a multimodal Mixture-of-Experts model with 552 billion backbone parameters and support for up to one million tokens of context.
Key points
- The release follows the earlier V4 Pro and continues DeepSeek's rapid model shipping pace.
Sources
- redditreddit.com