Skip to content

DeepSeek V4.1 Flash

deepseek-ai/DeepSeek-V4.1-Flash
Open Source · chat · open-weights
Open Alert me on changes
Context
1M
Max output
384K
Weights
Open
API $/1M
$0.15 / $0.60
Modalities
text · image
Released
10 Sept 2026
Download image Share on X Share on LinkedIn
AI summary
● machine-written

DeepSeek V4.1 Flash: 1M context, multimodal model with cache compression

DeepSeek V4.1 Flash is a Mixture-of-Experts model with 284B total parameters and 13B activated parameters, supporting 1M-token context and 384K max output. It handles text and image inputs and is optimized for fast inference and high-throughput workloads including coding, chat, and agent systems.

What's new
  • KV cache compression technology for efficient long-context processing
  • Hybrid attention mechanism for improved performance
  • Vision capabilities added (image-text-to-text)
  • Supports both non-thinking and thinking modes
  • Available with 8-bit and fp8 precision options
Best for
Coding assistants and programming tasksChat systems and conversational AIAgent workflows requiring fast responsivenessHigh-throughput, cost-sensitive deployments
Sources

Source: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash