DeepSeek V4.1 Flash
deepseek-ai/DeepSeek-V4.1-Flash
Open Source · chat · open-weights
Open
Alert me on changes
Context
1M
Max output
384K
Weights
Open
API $/1M
$0.15 / $0.60
Modalities
text · image
Released
10 Sept 2026
License: mit · deepseek-ai/DeepSeek-V4.1-Flash
AI summary
● machine-written
DeepSeek V4.1 Flash: 1M context, multimodal model with cache compression
DeepSeek V4.1 Flash is a Mixture-of-Experts model with 284B total parameters and 13B activated parameters, supporting 1M-token context and 384K max output. It handles text and image inputs and is optimized for fast inference and high-throughput workloads including coding, chat, and agent systems.
What's new
- KV cache compression technology for efficient long-context processing
- Hybrid attention mechanism for improved performance
- Vision capabilities added (image-text-to-text)
- Supports both non-thinking and thinking modes
- Available with 8-bit and fp8 precision options
Best for
Coding assistants and programming tasksChat systems and conversational AIAgent workflows requiring fast responsivenessHigh-throughput, cost-sensitive deployments
Source: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash