Happy birthday to me! Well, tomorrow. New preprint on the arxiv, first solo author paper and looking for feedback. Please email me if you have comments! I’ll say more in the thread. scirate.com/arxiv/2307.144…
[4/5]
To reduce tuning overhead further, we introduce a telescoping algorithm:
We introduce a telescoping algorithm, that reduces muP overhead further. It consists of -
Wide sweep on small models
Narrower sweeps as width grows → Telescoping gives us low error estimates of the
I usually use @scirate3 to get my daily quantum computing news --- what is the status quo for machine learning literature? Seems like it's Twitter --- anyone have suggestions for aggregators, or people, to follow?