Skip to content

NVIDIA Nemotron Parse 2.0

nvidia/NVIDIA-Nemotron-Parse-2.0
Open Source · chat · open-weights
Open Alert me on changes
Context
Max output
Weights
Open
API $/1M
Modalities
text · image
Released
30 Jun 2026
Download image Share on X Share on LinkedIn
AI summary
● machine-written

NVIDIA Nemotron Parse 2.0 open-weights multimodal model released

NVIDIA Nemotron Parse 2.0 is an open-weights vision-language model designed for image-text-to-text tasks including optical character recognition and document parsing. The model supports both text and image modalities and is available for local deployment or inference via compatible platforms like Transformers, vLLM, and SGLang.

What's new
  • Open-weights release with Safetensors format
  • Supports image and text inputs for multimodal tasks
  • Compatible with Transformers, vLLM, and SGLang inference servers
  • Custom implementation via NemotronParseForConditionalGeneration class
Best for
Document parsing and optical character recognition (OCR)Image captioning and visual understandingMultimodal feature extractionEnterprise perception sub-agent systems
Sources

Source: https://huggingface.co/nvidia/NVIDIA-Nemotron-Parse-2.0