technologyScore 35/100Watch

Nemotron 3 Ultra uses Mixture of Experts (MoE) architecture with only ~10% of 550B parameters active per token, enabling efficient inference despite massive model size

Andrej Karpathy· Independent· AI· 2026-06-14· about NVIDIA (NVDA)
550 billion parameters total, but only about 10% of that is active per token. These are specialist mini-brains that are being activated at a time. We call that mixture of experts.

Why it matters

MoE architecture is becoming the dominant paradigm for scaling open-source models efficiently, reducing compute costs while maintaining capability. This validates NVIDIA's architectural choices and creates demand for specialized inference optimization.

Investment implication

Companies building MoE-optimized inference engines and hardware accelerators (NVIDIA, specialized inference startups) are positioned to capture value from the shift to sparse model activation. This drives capex in specialized inference hardware.

Source

NVIDIA's New Free Al - A Gift To All Of Us (YouTube)
← All signals