Skip to content

NVIDIA Nemotron 3.5 Lightning 30B A3B BF16

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
Open Source · chat · open-weights
Open Alert me on changes
Context
1M
Max output
65.5K
Weights
Open
API $/1M
Modalities
text
Released
01 Aug 2026
Download image Share on X Share on LinkedIn
AI summary
● machine-written

NVIDIA Nemotron 3.5 Lightning 30B A3B released with 1M token context

NVIDIA Nemotron 3.5 Lightning 30B A3B is an open-weights mixture-of-experts model with 3B active parameters out of 30B total. The model supports a 1 million token context window and 65,536 token maximum output. It is designed for high-throughput agentic workloads and specialized tasks.

What's new
  • Open-weights availability for local deployment
  • 1M token context window capacity
  • 3B active parameters with 30B total parameter model
  • Mixture-of-experts architecture optimized for throughput
Best for
High-throughput agentic workloadsDomain-specific task customizationSpecialized agentic AI systems
Sources

Source: https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16