3 Comments
User's avatar
Elias Schmied's avatar

Really useful, thank you!

Maybe scale dependence should've been the default expectation all along because of some "more is different" style intuition (that has already been validated by the success of scaling itself, after all).

"I think various philosophical arguments suggest that there’s a greater diversity of room for algorithmic progress as compute scales up. The human brain is largely a linearly scaled-up primate brain, but individual humans and humans in aggregate can do many calculations no non-human primate can. And there are likely greater gains from organizing 10,000 humans well than from organizing 10 humans well."

This I don't understand - just to clarify, are you using this as an argument in favor of the last model with increasing bucket sizes?

Daniel Kokotajlo's avatar

Thank you! I found this very helpful.

June Jimenez's avatar

Great post! Thank you.

The rate of algorithmic progress is also a function of the effort put into other channels for improvement. The compute scale-up efforts at leading labs have been a channel that much research energy and capital was directed towards. It could have gone elsewhere.

If this had not happened, we would expect progress to have been somewhat less than the cited ~ 21,400x (otherwise the labs, which are approximately optimizing for capabilities speed growth already, would have just set compute procurement to 0), but if the estimate is for algorithmic growth to have been ~ 155x growth since then if you hold compute scaling constant, *and implicitly research effort* constant, I would expect research to have progressed substantially faster than 155x. My relatively unprincipled guess would be somewhere around ~1000-4000x.

I do expect this to have limits at some level; it does seem like scale is a deeply fundamental driver of algorithmic progress. In particular, I find your model that scale could reduce the research taste necessary to find an algorithmic improvement very insightful and highly plausible.

But overall I'd find "if you block the compute channel, the algorithmic channel will speed up somewhat less" plausible for at least a few OOM / maybe 18 research-months' worth of progress; that we see such rapid algorithmic efficiency progress for capabilities ~1 year behind the frontier suggests that it's possible to to train smaller models to a comparable level to bigger models with more research effort.