Hessian-Free Methods Did Not Vanish
October 24, 2023
Second-order optimization still offers a powerful way to train deep networks. James Martens’s 2010 Hessian-free method used curvature information without forming the full Hessian, and work with Ilya Sutskever applied it to recurrent networks one year later.
Brian Greenforest reconnects that lineage with modern AdamW-era training and hybrid methods explored at NERSC and Lawrence Berkeley National Laboratory.
Use Curvature Without Storing the Hessian
Hessian-vector products let an optimizer probe local curvature at the cost of gradient-like operations. Conjugate-gradient iterations then approximate a useful update inside that curved landscape.
The method can navigate directions that first-order gradients scale poorly, especially in recurrent systems with long dependencies and difficult curvature.
Combine First- and Second-Order Strengths
AdamW offers cheap adaptive steps across immense parameter sets. Hessian-free stages can add targeted curvature when the optimization reaches regions where first-order progress slows.
Optimizer researchers can revisit these combinations on current architectures, document the cost and convergence, and restore Hessian-free methods to the active large-model toolbox.
Read the Modern Hessian-Free Evidence
The linked paper brings Hessian-free optimization into a current deep-learning setting and makes its continuing technical role explicit.
LinkedIn status when archived: Visible to anyone on or off LinkedIn.
In 2010, Martens invented deep Hessian-free methods. In 2011, Sutskever collaborated with Martens to apply it successfully to arbitrary RNNs (not special LSTMs, GRUs etc.)
In 2020, NERSC, Lawrence Berkeley National Laboratory, publish applications of combining the method with AdamW. What happened to the public who are so obsessed with Transformer? Why anyone stopped talking about the method invented by Sutskever and Martens 12 years ago?
https://lnkd.in/gUvuya3S