technologyScore 40/100Watch
Nemotron 3 Ultra lacks multi-modal vision capabilities, creating market opportunity for smaller specialized vision models (e.g., Gemma) to be composed with larger language models
Andrej Karpathy· Independent· AI· 2026-06-14· about Google (Gemma) (GOOGL)
“I can't add vision capabilities to Nematron 3 Ultra, but I can bolt Gemma 4 to it with a screwdriver. It's like a seeing-eye dog guiding a smarter blind man along.”
Why it matters
The lack of native multi-modality in large open-source models creates demand for smaller, specialized vision models that can be composed with language models via lightweight integration mechanisms.
Investment implication
Smaller, specialized model providers (Google's Gemma, others) and model-composition infrastructure (routing, orchestration layers) benefit from the fragmentation of capabilities across specialized models rather than monolithic multi-modal systems.