materialScore 30/100Research

Karpathy: Current training datasets are only ~0.1% cognition; 99.99% is noise/garbage, suggesting massive opportunity in data curation and synthetic data generation

Andrej Karpathy· Independent· AI· 2026-07-22
the internet data set which is what we're working with the internet is like 0.1% cognition... and like 99.99% of like information is like you know garbage... most of it is not uh useful to the thinking part

Why it matters

If 99.9% of internet training data is noise, significant model efficiency gains are possible through better curation, creating demand for data filtering, quality assessment, and synthetic data generation tools.

Investment implication

Companies specializing in data curation, synthetic data generation (e.g., Scale AI, Snorkel), and training data filtering could see accelerating adoption. This could reduce training costs and improve model efficiency.

Source

Andre Karpathy - Size Isn'tThe Bottleneck Anymore (YouTube)
← All signals