NVIDIA Nemotron Parse 2.0
nvidia/NVIDIA-Nemotron-Parse-2.0
Open Source · chat · open-weights
Open
Alert me on changes
Context
—
Max output
—
Weights
Open
API $/1M
—
Modalities
text · image
Released
30 Jun 2026
License: other · nvidia/NVIDIA-Nemotron-Parse-2.0
AI summary
● machine-written
NVIDIA Nemotron Parse 2.0 open-weights multimodal model released
NVIDIA Nemotron Parse 2.0 is an open-weights vision-language model designed for image-text-to-text tasks including optical character recognition and document parsing. The model supports both text and image modalities and is available for local deployment or inference via compatible platforms like Transformers, vLLM, and SGLang.
What's new
- Open-weights release with Safetensors format
- Supports image and text inputs for multimodal tasks
- Compatible with Transformers, vLLM, and SGLang inference servers
- Custom implementation via NemotronParseForConditionalGeneration class
Best for
Document parsing and optical character recognition (OCR)Image captioning and visual understandingMultimodal feature extractionEnterprise perception sub-agent systems
Source: https://huggingface.co/nvidia/NVIDIA-Nemotron-Parse-2.0