Skip to content

Mage VL

microsoft/Mage-VL
Open Source · chat · open-weights
Open Alert me on changes
Context
Max output
Weights
Open
API $/1M
Modalities
text · image
Released
25 Jul 2026
License: apache-2.0 · microsoft/Mage-VL
Download image Share on X Share on LinkedIn
AI summary
● machine-written

Microsoft releases Mage-VL, an efficient multimodal vision-language model

Mage-VL is an open-weights vision-language model from Microsoft designed for image-text-to-text tasks, supporting both image and video understanding. The model uses variable-resolution pretraining to improve performance with different token budgets, achieving over 96.1% accuracy on Food-101 and over 86.3% on ImageNet at 676 tokens. It is available under the Apache 2.0 license via Hugging Face.

What's new
  • Variable-resolution pretraining enables performance scaling with token budget
  • Supports streaming multimodal inference for image and video content
  • Achieves >96.1% Food-101 and >86.3% ImageNet accuracy at 676 tokens
  • Compatible with Transformers, vLLM, and SGLang libraries
  • Available as open-weights model under Apache 2.0 license
Best for
Image understanding and analysis tasksVideo-based content understandingMultimodal conversational applicationsStreaming inference scenarios
Sources

Source: https://huggingface.co/microsoft/Mage-VL