DeepSeek-V4.1-Flash Arrives on SiliconFlow: Faster, Cheaper, and Vision-Ready
Key Info
DeepSeek-V4.1-Flash is now live on SiliconFlow on day one, touted as a flagship performer: a 552B-parameter MoE model with ~8B active parameters during prefill and ~16B during decode, native vision, a 1M context window, and an MIT license.
Highlights
- Faster inference and higher throughput, keeping Flash fast even at production scale.
- Native vision support extends the model to multimodal tasks.
- 1M context window with a roughly 4× smaller KV cache footprint than V4 Flash.
- Production-ready on SiliconFlow today, available to build with immediately.