bottleneckScore 40/100Watch
LeCun criticizes reconstruction-based video prediction as fundamentally flawed—10 years of failed experiments at FAIR suggest generative video models may not be viable path, forcing shift to alternative architectures.
Yann LeCun· Meta· AI· 2026-07-06· about Meta (FAIR) (META)
“at fair u some colleagues and I have been trying to do this for about 10 years um and you can't you can't really do the same trick as with LLMs because uh you know LLM as I said you can't predict exactly which word is going to follow a sequence of words but you can predict the distribution of words now if you go to video what you would have to do is predict the distribution over all possible frames in a video and we don't really know how to do that properly.”
Why it matters
A decade of failed research on generative video models suggests this approach is not viable. This creates urgency to adopt non-generative architectures (JEPPA, joint embedding) and may render existing generative video model projects obsolete.
Investment implication
Companies investing heavily in generative video models (e.g., for content creation, synthesis) may need to pivot. Infrastructure optimized for generative video training may face technical dead-ends. Conversely, companies adopting joint-embedding and representation-learning-based video training may leapfrog competitors.