Back to Discover
Lesson

Why Gradient Descent Slows Near Saddle Points

A saddle point stalls gradient descent because the gradient itself nearly vanishes, and the surrounding curvature decides whether the algorithm escapes or gets stuck.

Before you enter

A complete interactive classroom, not just a preview.

Start when you are ready to enter this Stage's 10 scenes and explore, respond, and learn as you go.

10
Scenes
20 min
Estimated
Content language: en-US
Start this Stage
Sign-in may be required to play
What happens inside
  1. 01The Mystery of the Stallslide
    Orientation

    Frame the question: gradient descent is supposed to walk downhill, so why does it sometimes crawl for thousands of steps near what looks like the middle of the landscape?

    • Set up the puzzle: GD slows near saddle points
    • Preview that the answer lies in geometry, not just slope
    • Roadmap for the lesson
  2. 02Explore a 2D Saddleinteractive
    PredictionPredict

    Let learners manipulate a 2D loss surface and watch a gradient-descent particle move toward the saddle from different starting points.

    • Drag a particle on the 2D loss surface
    • Watch gradient descent animate toward the saddle
    • See the path flatten as it nears the center
  3. 03What Is a Saddle Point?slide
    Model building

    Define a saddle point precisely: a critical point where the Hessian has both positive and negative eigenvalues, so the surface curves up in some directions and down in others.

    • Formal definition: critical point with mixed-sign Hessian eigenvalues
    • The classic 'horse saddle' as analogy
    • Why saddles dominate in high-dimensional loss landscapes
  4. 04Where Did the Gradient Go?slide
    Misconception repairExplain

    Show that near a saddle, the gradient shrinks toward zero even though the surface is sloping strongly in some directions — only the flat direction matters for the step.

    • Gradient is the vector sum of slopes in every direction
    • At a critical point, every directional slope is zero
    • Even far from exact zero, the gradient can be tiny along the flat direction
  5. 05Eigenvalues of the Hessianinteractive
    Model buildingObserve

    Visualize the Hessian's eigenvalues around the saddle: one positive, one negative, and how their magnitudes shape the local bowl along each axis.

    • Drag to resize the two curvatures
    • See how the eigenvalues control local steepness
    • Watch gradient descent speed change as the ratio grows
  6. 06The Flat Direction Trapslide
    Misconception repairExplain

    Explain why ill-conditioning (one very small eigenvalue) makes progress per step extremely slow along the escape direction, even when the negative-curvature direction exists.

    • Small positive eigenvalue ⇒ tiny component of gradient along that direction
    • Newton step would be huge, but GD step is only proportional to gradient
    • Many tiny steps needed before escape becomes visible
  7. 07Escape Directions and Curvatureslide
    Model building

    Distinguish the descent direction (where loss decreases) from the escape direction (negative-curvature direction used to leave the saddle).

    • Negative eigenvalues give directions of locally decreasing loss
    • GD follows the negative gradient, which is a blend of all directions
    • A small gradient along the negative-curvature axis is what stalls escape
  8. 08Check Your Intuitionquiz
    Assessment

    Quick multiple-choice check on whether learners can connect Hessian eigenvalues to the speed of escape from a saddle.

    • Predict GD behavior from eigenvalues
    • Identify the cause of slow progress
    • Distinguish saddle from local minimum
  9. 09From Geometry to Practiceslide
    ApplicationApply

    Translate the intuition into practical consequences: stalls in deep nets, the value of momentum and second-order methods, and how to detect saddles during training.

    • Why ill-conditioned problems stall on a real loss curve
    • Momentum and curvature-aware methods accelerate escape
    • Track gradient norms and curvature diagnostics to find saddles
  10. 10Recap: The Saddle-Point Storyslide
    Synthesis

    Tie the whole picture together: gradient shrinks near saddles because the gradient is small in the flat direction, and the Hessian tells you which direction that is.

    • Saddle = critical point with mixed Hessian eigenvalues
    • Gradient → 0 ⇒ step → 0 in the flat direction
    • Escape needs curvature-aware moves; pure GD crawls
Discussion

Discussion threads for a Stage aren't available yet.

Where this leads
Explore more

More in Math & Logic

See all