Rich Sutton: synthetic data is "just a big mistake"
Summary
Rich Sutton dismisses synthetic data as "just a big mistake" for scaling LLMs, arguing the "big world hypothesis" suggests AI systems should learn from real-world experience instead. This challenges the dominant foundation model scaling strategy of using generated datasets.
Key Takeaways
- Synthetic data generation won't solve LLM scaling because the world contains infinitely many learnable things. Systems trained only on synthetic data miss novel real-world patterns that require continuous experience-based learning.
- The "big world hypothesis" has circulated in Alberta AI labs for 5-10 years as an alternative scaling paradigm. This suggests the synthetic data approach may be a widespread but fundamentally flawed consensus among foundation model labs.
- Removing humans from the learning loop is critical for AGI-capable systems. If AI can learn directly from real experience rather than curated synthetic datasets, it can develop capabilities across any task domain.
- The current scaling paradigm treats the existing internet as "fossil fuel" running out, prompting synthetic data solutions. This framing misses the opportunity to build systems that learn continuously from the infinite real world.
Related topics
Transcript Excerpt
It seems like a lot of what the foundation model labs are working on right now is synthetic data generation in order to kind of get us beyond the fossil fuel that is the existing human internet. Is synthetic data generation that's part of this LLM scaling paradigm? Is that a general method that leverages computation? >> No, that's that's just a big mistake. >> Why? >> [laughter] >> Maybe it's the next the next big lesson. It's been floating around Alberta for 5 or 10 years. >> Okay. >> And we call it the big world perspective or the big world hypothesis. Khuram who eventually wrote it up as a paper. There's a little paper called the big world hypothesis. >> So the big world is that the world is infinitely big. There are infinitely many things to learn. And you can have people generating th…
More from Sequoia Capital
- Parallel’s Parag Agrawal: Building a New Web for AI Agents
- Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again
- Continual Learning: How AI Agents Get Better With Every Use | Arjun Karanam, Trajectory
- When to Build Your Own Agent Harness | Harrison Chase, LangChain
- How Harvey Built a Research Lab on a Budget | Gabe Pereyra