← back
Adaption Labs: Gradient-Free Continual Learning — Sara Hooker, Adaption
Takeaway
Automating the full data-and-model adaptation loop can make specialized AI development more accessible and compute-efficient.
Summary
- Sara Hooker argues that compute concentration and narrow research career paths restrict who can build frontier AI and which problems get addressed.
- Adaption's Auto Scientist automates model customization by jointly optimizing data and training choices across domains and model architectures.
- Controlling data quality alongside model adaptation was necessary for the reported gains; broad hyperparameter exploration exceeded what researchers typically change together.
- The system targets faster, more predictable customization cycles, with interest from medicine, science, law, and coding and plans for adaptive test-time compute.
automated-researchmodel-adaptationdata-quality
Original description
Fewer than five thousand people in the world know how to train a frontier model at scale, by Sara Hooker's estimate, and that knowledge travels like an apprenticeship rather than a literature. Modern computer science is 77 years old, two generations, and in that time the route to contributing at the frontier narrowed into one funnel: the right PhD, the right industry lab, the right problem at the right moment. She calls it the unreasonably narrow path, and notes it got compounded in this field by compute, so that a handful of labs build what everyone else uses and whole regions of the world appear nowhere on the map of where breakthroughs happen. Her case that the funnel is about to widen rests on two things. AutoScientist automates the training of models, optimizing the whole loop together from data through alignment and evolving itself per domain, and it beats research staff partly because people carry priors about particular architectures while the search ranges across sizes and across dense and mixture of experts designs. It only started paying off once they controlled data quality alongside the model rather than leaving that to the agent. One nice piece of honesty: the win rates all sit just above 60% because the budget was set to stop there, and lifting that ceiling let them keep climbing. The second reason is her slow death of scaling argument, that pretraining size is no longer the most rewarding axis. That matters for access, because pretraining compute has to be colocated and enormous while the compute that now pays off is distributable. If no lab is going to quadruple model size again on this architecture, recipes and algorithms start to matter more than hoarded GPUs. Speaker info: https://x.com/sarahookr / sararosehooker https://www.sarahooker.me/ Timestamps: 0:00 - Seventy seven years of computer science 1:15 - From gentleman scientists to professional labs 1:56 - The unreasonably narrow path 3:13 - GPU poor and GPU rich 3:52 - Where breakthroughs come from, and where they do not 4:30 - Why we are ripe for a revolution 5:08 - AutoScientist 5:47 - Why it beats research staff 6:24 - It only worked once they controlled the data 7:02 - The 60% that was a budget stop 7:41 - Where the demand is: medical, legal, science, code 8:59 - Languages from day one, and non verifiable tasks 10:16 - The compute problem that would undo all this 10:53 - The slow death of scaling 11:31 - Smaller models overtaking larger ones 12:09 - A broader action space, and why that opens the field 12:47 - Q&A begins 14:06 - Fewer than five thousand people 14:43 - Why the cost of asking shapes what gets asked 15:57 - Q&A: the safety objection to open frontier AI 17:10 - Q&A: parametric against nonparametric storage 18:28 - Q&A: are large models still needed for distillation 19:04 - Why this architecture has hit its size ceiling