An optimization landscape is the shape of the function you are minimising, viewed as a surface over its parameters. For a two-parameter model you could draw it as hills and valleys. For a neural network with a billion parameters you cannot draw it, but its local shape still decides whether training converges, how fast, at what learning rate it blows up, and which of many solutions you end up with.

This article builds the vocabulary from first principles (gradients, Hessians, convexity, conditioning, saddle points, sharpness), runs small tested experiments that show each effect with real numbers, and then connects them to the practical knobs of training: learning rate, warmup, momentum and batch size. If gradient descent itself is new to you, read Gradient Descent, in depth first.

Gradient, Hessian and critical points

Near any point w, a smooth loss looks like its second-order Taylor expansion: f(w + d) ≈ f(w) + g·d + ½ dᵀ H d, where g is the gradient and H is the Hessian, the matrix of second derivatives. The gradient says which way is downhill; the Hessian says how the slope changes as you move, that is, the curvature in every direction.

Because H is symmetric it has real eigenvalues, and its eigenvectors are the principal directions of curvature. A positive eigenvalue means the surface curves up along that direction, a negative one means it curves down, and zero means it is flat to second order. At a critical point, where g = 0, the eigenvalues classify the point: all positive is a local minimum, all negative a local maximum, mixed signs a saddle. Zero eigenvalues leave the test inconclusive, which is common in neural networks because of their symmetries.

Critical points are classified by the signs of the Hessian eigenvaluesMinimumall eigenvalues > 0Saddlemixed signs: an escape directionSharpFlatlarge vs small top eigenvalueConvexevery local min is globalIll-conditionedkappa = L / mu sets the speedNon-convex, high-dimsaddles dominate, many good minimaGradient descent sees only the local slope; curvature decides step size, speed and stability.
The same gradient-free test, read from curvature: minimum, saddle, and sharp versus flat minima. The lower row is the order in which these properties start to matter as problems grow.

Convexity and the condition number

A function is convex if the straight line between any two points on its graph lies on or above the graph; for twice-differentiable functions this means H is positive semidefinite everywhere. Convexity is the reason linear regression, logistic regression and support vector machines are reliable: every local minimum is a global minimum, so gradient descent cannot get stuck somewhere worse, and the theory gives guaranteed rates.

Two constants describe a well-behaved convex problem. L (smoothness) bounds the largest curvature, and mu (strong convexity) bounds the smallest. Their ratio kappa = L / mu is the condition number. Gradient descent is stable only if the step size is below 2 / L, and with the best fixed step the error shrinks by roughly (kappa - 1) / (kappa + 1) per step on a quadratic. When kappa is large that factor is close to 1 and progress crawls along a long narrow valley.

Experiment: conditioning sets the speed

Take the quadratic f(x) = ½ (x1² + kappa·x2²) and run gradient descent with the optimal fixed step 2 / (1 + kappa) from (1, 1) until the loss is below 1e-8:

import numpy as np

def gd_quadratic(kappa, steps=100000):
    h = np.array([1.0, kappa])          # Hessian eigenvalues: mu = 1, L = kappa
    lr = 2.0 / (1.0 + kappa)
    x = np.array([1.0, 1.0])
    for t in range(steps):
        x = x - lr * h * x              # gradient of the quadratic is h * x
        if 0.5 * np.sum(h * x * x) < 1e-8:
            return t + 1
    return steps

for k in (1, 10, 100, 1000):
    print(k, gd_quadratic(k))

The step counts are 1, 51, 559 and 6,160: the work grows almost linearly with the condition number. Now fix L = 50 and vary the step: at lr = 0.039 (below 2 / 50 = 0.04) the iterate is about 3.5e-4 after 200 steps, and at lr = 0.041 it has grown to about 1.7e4. The stability edge is sharp and set entirely by the steepest direction.

That is the whole story of learning-rate tuning in miniature. The steepest direction caps the step size; the flattest direction sets how many steps you need. Momentum improves the dependence to roughly the square root of kappa, and adaptive optimizers such as Adam rescale each coordinate, which helps when ill-conditioning is aligned with the parameter axes. The per-coordinate maths is worked through in SGD vs Adam vs AdamW Optimizer Math.

Saddle points and plateaus

Non-convex problems add new kinds of critical points. In low dimensions people picture bad local minima, but in high dimensions saddle points are the common case: for a random critical point to be a minimum, every one of millions of curvature directions must point up. Dauphin and colleagues argued in 2014 that, for the error surfaces of large networks, critical points with high loss are overwhelmingly saddles, and that plateaus around them, not local minima, are what slow training.

A two-dimensional example shows the mechanics. Take f(x, y) = x² + y⁴/4 − y²/2. The origin is a saddle with Hessian eigenvalues 2 and −1; the minima sit at (0, ±1). Gradient descent started exactly on y = 0 converges to the saddle and stays there, because the gradient never has a y component. A tiny offset grows by a factor of about 1 + lr per step along the escape direction:

import numpy as np

def grad(p):
    x, y = p
    return np.array([2 * x, y**3 - y])

def escape_steps(offset, lr=0.1):
    p = np.array([1.0, offset])
    for t in range(100000):
        p = p - lr * grad(p)
        if abs(p[1]) > 0.5:
            return t + 1

for off in (1e-2, 1e-4, 1e-8, 1e-12):
    print(off, escape_steps(off))

Measured escape times are 43, 91, 188 and 284 steps: each factor of 10,000 closer to the saddle adds roughly the same number of steps, because escape time grows with log(1 / offset). Saddles do not trap gradient descent in practice, since random initialisation and minibatch noise give a nonzero offset, but they do produce plateaus where the loss sits still for a while and then drops. If you have watched a loss curve stall and then fall off a cliff, this is one explanation; Diagnosing Loss Curves covers how to tell it apart from a bad learning rate.

Measuring sharpness without the Hessian

You cannot form the Hessian of a large model; it has as many entries as parameters squared. You do not need to. Power iteration finds the top eigenvalue, the sharpness, using only Hessian-vector products, and an HVP costs about as much as two gradient evaluations. This tested NumPy version uses finite differences of the gradient:

import numpy as np

def hvp(grad_fn, w, v, eps=1e-4):
    return (grad_fn(w + eps * v) - grad_fn(w - eps * v)) / (2 * eps)

def top_eigenvalue(grad_fn, w, iters=500, seed=0):
    rng = np.random.default_rng(seed)
    v = rng.standard_normal(w.shape)
    v /= np.linalg.norm(v)
    lam = 0.0
    for _ in range(iters):
        hv = hvp(grad_fn, w, v)
        lam = float(v @ hv)               # Rayleigh quotient
        v = hv / np.linalg.norm(hv)
    return lam

On a 10-feature logistic regression with 500 random rows it returns 0.29004 against 0.29008 from the exact Hessian. Convergence depends on the gap between the top two eigenvalues: with 100 iterations the same run was still about 1 percent low, so watch the estimate settle rather than trusting a fixed count. In PyTorch the HVP is exact via double backpropagation: compute grads = torch.autograd.grad(loss, params, create_graph=True), then torch.autograd.grad(grads, params, grad_outputs=v). Tracking this one number during training is the cheapest way to see the landscape your optimizer is actually in.

What deep-network landscapes look like

Three findings about deep-network landscapes are worth knowing, each with its caveat.

  • Many minima are good enough. Heavily over-parameterised networks can usually drive training loss close to zero, and different random seeds reach different minima with similar test accuracy. The hard part is less finding a minimum than finding one that generalises and getting there in budget.
  • Minima are connected. Garipov and colleagues and Draxler and colleagues (both 2018) found that two independently trained solutions can be joined by a simple curved path along which the loss stays low. The low-loss region is better pictured as a connected valley floor than as isolated pits.
  • Flatness correlates with generalisation, imperfectly. Keskar and colleagues (2017) linked large-batch training to sharper minima and worse test accuracy. Dinh and colleagues (2017) showed that sharpness can be changed by reparameterising a ReLU network without changing the function, so raw sharpness is not a reliable measure on its own. Treat it as a diagnostic, not a target.

A fourth result changes how to read learning rates. Cohen and colleagues (2021) observed that with full-batch gradient descent the sharpness rises during training until it reaches about 2 / lr and then hovers there, the "edge of stability", while the loss keeps falling non-monotonically. The quadratic stability limit from the conditioning section still bites, but the network adapts its curvature to the step size instead of the other way round.

From geometry to training knobs

Translate the geometry into training decisions:

Landscape factSymptomKnob
Top curvature caps the stepLoss spikes or NaN at high learning rateLower peak rate; gradient clipping
Curvature is high early in trainingInstability in the first few hundred stepsLearning-rate warmup
Ill-conditioningSlow, zig-zag progressMomentum, Adam, normalisation layers
Saddles and plateausLoss flat, then sudden dropPatience; do not cut the rate too early
Sharp versus flat minimaGood train loss, weak test lossSmaller batches, weight decay, decay schedule

Warmup is the clearest example: early in training curvature is often high, so a small step avoids the divergence the quadratic predicts, and the rate can rise as the landscape flattens. Schedules are covered in Learning Rate Schedules. For one-dimensional problems known to be unimodal you can skip gradients entirely; ternary and golden-section search exploit exactly that shape.

Finding the stability edge with a sweep

The cheapest way to find where your own landscape's stability edge lies is to measure it. Run a short training segment at each of a geometric series of learning rates and record the loss; the rate where it stops improving and then diverges is the practical counterpart of 2 / L:

import numpy as np

def lr_sweep(loss_and_grad, w0, lrs, steps=50):
    # Short run per learning rate; final loss, or inf if it diverged.
    results = {}
    for lr in lrs:
        w = w0.copy()
        for _ in range(steps):
            loss, g = loss_and_grad(w)
            if not np.isfinite(loss) or loss > 1e6:
                loss = float("inf")
                break
            w = w - lr * g
        results[lr] = loss_and_grad(w)[0] if np.isfinite(loss) else float("inf")
    return results

h = np.array([1.0, 50.0])                       # sharpness L = 50, so 2 / L = 0.04
f = lambda w: (0.5 * float(np.sum(h * w * w)), h * w)
print(lr_sweep(f, np.ones(2), [0.005, 0.01, 0.02, 0.035, 0.039, 0.045, 0.08]))

On this quadratic the final losses are 0.303, 0.183, 0.0663, 0.0142, 0.157, then infinity at 0.045 and 0.08. Two things show up that also appear on real models. The best rate (0.035) sits just below the edge, and the loss gets worse again at 0.039 before it diverges, because the steep direction oscillates instead of settling. On a network, run the same sweep for a few hundred steps per rate from the same checkpoint, then set the peak rate a safe factor below the point where the loss turns up, since sharpness can rise later in training.

Failure modes

  • Trusting 2D pictures. Landscape plots slice a huge space along one or two directions; random directions make everything look smooth, and the choice of scaling changes how sharp a minimum appears. Use them to compare runs under the same protocol, not to read absolute geometry.
  • Comparing sharpness across parameterisations. Changing normalisation or scale changes the number without changing the model, as Dinh and colleagues showed.
  • Finite-difference noise. Too small an eps in the HVP amplifies float error, too large a one measures more than curvature. Check against an exact Hessian on a small model first.
  • Minibatch Hessians. Sharpness measured on one batch is noisy; average several batches before acting on it.
  • Blaming local minima. A stalled run is far more often a learning-rate, data or numerical problem than a bad minimum. Rule those out before reaching for exotic optimizers.

What to do next

  • Run the conditioning experiment and confirm the step counts, then add momentum and compare.
  • Run the saddle experiment with minibatch-style noise and watch the plateau shorten.
  • Add the power-iteration sharpness estimate to a small training loop and log it every few hundred steps next to the learning rate; check whether it approaches 2 / lr.
  • Find your model's divergence learning rate with a short sweep and set the peak rate safely below it, with warmup.
  • Read the mode-connectivity and edge-of-stability papers named above, then revisit your own loss curves.
Key takeaway: The landscape's curvature decides almost everything gradient descent does: the steepest direction caps the stable step at about 2 / L, the condition number sets how many steps you need, saddles cause plateaus rather than traps, and deep networks have many connected low-loss solutions. Measure sharpness with Hessian-vector products instead of guessing, and use warmup, momentum and schedules as the practical responses to what it shows.