HeadlinesBriefing favicon HeadlinesBriefing.com

NVIDIA Nemotron 3.5 Lightning 30B Model

Hacker News •
×

The NVIDIA family of open models introduces the Nemotron-3.5 Lightning 30B, a large language model released on August 11, 2026. It features a hybrid Mixture-of-Experts architecture combining Mamba-2, MoE, and selective Attention, with 30B total parameters and 3B active. The model supports up to 1M token context and runs on NVIDIA Blackwell, Hopper, and Ampere GPUs, including the DGX Spark and H100.

Training leveraged over 20T tokens and included Multi-Token Prediction layers, with a post‑training quantization step to optimize inference. The model supports English, Spanish, French, German, Italian, and Japanese, and is licensed under the Open MDW‑1.1 agreement. Deployment is streamlined for DGX Spark via DSpark speculative decoding, and it performs strongly on reasoning, coding, and instruction‑following benchmarks.

Designed for developers building autonomous agents, chatbots, and RAG systems, the model excels in long‑running, efficient local inference and can be run on personal hardware with the provided vLLM recipe. The release includes detailed quick‑start guides, benchmark scripts, and support for tool‑calling and multi‑step reasoning.

Overall, the Nemotron-3.5 Lightning delivers high accuracy, scalability, and open‑source flexibility for next‑generation AI applications, with official results available through NeMo Gym and NeMo Evaluator.