materialScore 30/100Research
Karpathy: Current training datasets are only ~0.1% cognition; 99.99% is noise/garbage, suggesting massive opportunity in data curation and synthetic data generation
Andrej Karpathy· Independent· AI· 2026-07-22
“the internet data set which is what we're working with the internet is like 0.1% cognition... and like 99.99% of like information is like you know garbage... most of it is not uh useful to the thinking part”
Why it matters
If 99.9% of internet training data is noise, significant model efficiency gains are possible through better curation, creating demand for data filtering, quality assessment, and synthetic data generation tools.
Investment implication
Companies specializing in data curation, synthetic data generation (e.g., Scale AI, Snorkel), and training data filtering could see accelerating adoption. This could reduce training costs and improve model efficiency.