DeepSeek V4 Flash Vision Exp
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
Open Source · chat · open-weights
Open
Alert me on changes
Context
1M
Max output
384K
Weights
Open
API $/1M
$0.22 / $0.66
Modalities
text · image
Released
31 Aug 2026
License: mit · deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
AI summary
● machine-written
DeepSeek V4 Flash Vision Exp adds image understanding to efficient MoE model
DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash, adding image understanding capabilities while maintaining text performance in agents, reasoning, and world knowledge. The model is a sparse mixture-of-experts architecture with 13B active parameters out of 284B total and supports a 1M-token context window. It is suited for document and chart understanding, visual question answering, and multimodal agent workflows that interleave text and images.
What's new
- Adds image understanding to V4 Flash base model
- Supports multimodal input (text and images)
- Sparse MoE architecture with 13B active/284B total parameters
- 1M-token context window with up to 384K output tokens
Best for
Document and chart understandingVisual question answeringMultimodal agent workflowsImage-text reasoning tasks
Source: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp