Hessian-Free Methods Did Not Vanish
October 24, 2023

Second-order optimization still offers a powerful way to train deep networks. James Martens’s 2010 Hessian-free method used curvature information without forming the full Hessian, and work with Ilya Sutskever applied it to recurrent networks one year later.

Brian Greenforest reconnects that lineage with modern AdamW-era training and hybrid methods explored at NERSC and Lawrence Berkeley National Laboratory.

Use Curvature Without Storing the Hessian

Hessian-vector products let an optimizer probe local curvature at the cost of gradient-like operations. Conjugate-gradient iterations then approximate a useful update inside that curved landscape.

The method can navigate directions that first-order gradients scale poorly, especially in recurrent systems with long dependencies and difficult curvature.

Combine First- and Second-Order Strengths

AdamW offers cheap adaptive steps across immense parameter sets. Hessian-free stages can add targeted curvature when the optimization reaches regions where first-order progress slows.

Optimizer researchers can revisit these combinations on current architectures, document the cost and convergence, and restore Hessian-free methods to the active large-model toolbox.

Read the Modern Hessian-Free Evidence

The linked paper brings Hessian-free optimization into a current deep-learning setting and makes its continuing technical role explicit.

https://lnkd.in/gUvuya3S · https://arxiv.org/abs/2006.00719

Originally posted on LinkedIn

Brian Greenforest · (2023-10-24 18:39:14 UTC)

Open the original LinkedIn post · LinkedIn activity 7122652775242481664

LinkedIn status when archived: Visible to anyone on or off LinkedIn.

In 2010, Martens invented deep Hessian-free methods. In 2011, Sutskever collaborated with Martens to apply it successfully to arbitrary RNNs (not special LSTMs, GRUs etc.) In 2020, NERSC, Lawrence Berkeley National Laboratory, publish applications of combining the method with AdamW. What happened to the public who are so obsessed with Transformer? Why anyone stopped talking about the method invented by Sutskever and Martens 12 years ago? https://lnkd.in/gUvuya3S

Original LinkedIn media