technologyScore 40/100Watch

Nemotron 3 Ultra lacks multi-modal vision capabilities, creating market opportunity for smaller specialized vision models (e.g., Gemma) to be composed with larger language models

Andrej Karpathy· Independent· AI· 2026-06-14· about Google (Gemma) (GOOGL)
I can't add vision capabilities to Nematron 3 Ultra, but I can bolt Gemma 4 to it with a screwdriver. It's like a seeing-eye dog guiding a smarter blind man along.

Why it matters

The lack of native multi-modality in large open-source models creates demand for smaller, specialized vision models that can be composed with language models via lightweight integration mechanisms.

Investment implication

Smaller, specialized model providers (Google's Gemma, others) and model-composition infrastructure (routing, orchestration layers) benefit from the fragmentation of capabilities across specialized models rather than monolithic multi-modal systems.

Source

NVIDIA's New Free Al - A Gift To All Of Us (YouTube)
← All signals