Mage VL
microsoft/Mage-VL
Open Source · chat · open-weights
Open
Alert me on changes
Context
—
Max output
—
Weights
Open
API $/1M
—
Modalities
text · image
Released
25 Jul 2026
License: apache-2.0 · microsoft/Mage-VL
AI summary
● machine-written
Microsoft releases Mage-VL, an efficient multimodal vision-language model
Mage-VL is an open-weights vision-language model from Microsoft designed for image-text-to-text tasks, supporting both image and video understanding. The model uses variable-resolution pretraining to improve performance with different token budgets, achieving over 96.1% accuracy on Food-101 and over 86.3% on ImageNet at 676 tokens. It is available under the Apache 2.0 license via Hugging Face.
What's new
- Variable-resolution pretraining enables performance scaling with token budget
- Supports streaming multimodal inference for image and video content
- Achieves >96.1% Food-101 and >86.3% ImageNet accuracy at 676 tokens
- Compatible with Transformers, vLLM, and SGLang libraries
- Available as open-weights model under Apache 2.0 license
Best for
Image understanding and analysis tasksVideo-based content understandingMultimodal conversational applicationsStreaming inference scenarios
Sources