technologyScore 35/100Watch
Nemotron 3 Ultra uses Mixture of Experts (MoE) architecture with only ~10% of 550B parameters active per token, enabling efficient inference despite massive model size
Andrej Karpathy· Independent· AI· 2026-06-14· about NVIDIA (NVDA)
“550 billion parameters total, but only about 10% of that is active per token. These are specialist mini-brains that are being activated at a time. We call that mixture of experts.”
Why it matters
MoE architecture is becoming the dominant paradigm for scaling open-source models efficiently, reducing compute costs while maintaining capability. This validates NVIDIA's architectural choices and creates demand for specialized inference optimization.
Investment implication
Companies building MoE-optimized inference engines and hardware accelerators (NVIDIA, specialized inference startups) are positioned to capture value from the shift to sparse model activation. This drives capex in specialized inference hardware.