
Towards architecture-aware optimisation

Towards architecture-aware optimisation

This is the third post in a series summarising work that seeks to provide a theory of generalisation in Deep Neural Networks (DNNs)...

And why is this a very important step in understanding why they work?

Stochastic Gradient Descent approximates Bayesian sampling