Nemotron 3.5 Lightning: coding and recall tests

Explore Nemotron 3.5 Lightning's coding limits and document recall in tests on an RTX Pro 6000, including questions about a 60,000-token text.

Player not loading? Watch on YouTube

This first look tests NVIDIA's Nemotron 3.5 Lightning as a local LLM through OpenCode on an RTX Pro 6000. The speaker describes an open-weight mixture of experts model with 30 billion total parameters, 3 billion active, hybrid Mamba 2 architecture and a 1 million token context window. He also notes optional reasoning, a speculative decoding model, and training data cutoffs of September 2025 for pre-training and May 2026 for post-training. DGX Spark is discussed as a target system, rather than the hardware used for these tests.

The coding results are uneven. Browser OS attempts have little working functionality, and checks through NVIDIA's endpoint on OpenRouter do not resolve the speaker's concerns. A C++ skateboarding game progresses past compilation errors but opens and closes quickly. The watch website produces a 3D watch the speaker finds better than expected, while dashboard edits improve text contrast but appear to reduce functionality. The final GPU rental frontend also has layout problems.

Document handling is more successful in this sample. The model installs tools to read a Word file and count its tokens, then answers questions about a roughly 60,000-token text that the speaker confirms are correct. This does not test the full advertised context window. He sees promise for AI agent tasks involving tools and recall, but reports weak coding performance and an incorrect answer to a niche automotive question.