Learning Rate

Learning Rate

Definition: A hyperparameter that controls how large a step gradient descent takes when updating model weights on each iteration.

How It Works

  • Too high: updates overshoot the minimum, loss oscillates or diverges
  • Too low: training converges extremely slowly, or gets stuck in a shallow local minimum
  • Modern training often uses a schedule — starting higher and decaying over time, or adaptive optimizers like Adam that adjust it automatically
  • Sits directly inside the Gradient Descent update rule as the multiplier on the gradient: w := w - lr * grad
  • Its effective scale interacts with batch size, model depth, and the chosen optimizer, so a learning rate tuned for one setup rarely transfers unchanged to another
  • Distinct from other step-size-like hyperparameters (e.g., momentum’s decay coefficient) even though they interact — the learning rate scales the gradient itself, momentum scales how much of the previous update direction persists

Visualizing Too High, Too Low, and Just Right

The same starting point, the same gradient direction, three different outcomes — the learning rate alone decides whether a parameter overshoots, crawls, or settles:

This is the same intuition as Gradient Descent’s ball-on-a-hill picture, just isolating the one variable — step size — that determines whether the ball settles in the valley, rolls past it, or barely moves at all.

Under the Hood

  • In the basic update rule w := w - lr * dL/dw, the learning rate lr is a scalar (or, in adaptive optimizers, effectively a per-parameter scalar) multiplying the gradient before it’s subtracted
  • Because gradients can vary by orders of magnitude across layers (especially in deep networks), a single global learning rate is a compromise — this is part of why adaptive optimizers like Adam, which rescale the effective learning rate per parameter based on gradient history, tend to be easier to use than plain SGD
  • The “right” learning rate scales roughly with batch size: doubling the batch size roughly halves gradient noise, which allows (and often benefits from) a proportionally larger learning rate — this is the basis for the “linear scaling rule” used when training with very large batches
  • Learning rate interacts with loss curvature: in directions where the loss surface is steep, a large learning rate causes oscillation; in flat directions, the same learning rate makes almost no progress. This curvature mismatch is exactly what momentum and adaptive methods try to compensate for
  • Recent optimization research has found that gradient descent on real neural networks often operates at the “edge of stability” (Cohen, Kaur, Li, Kolter, and Talwalkar, ICLR 2021) — the loss curvature (the largest eigenvalue of the Hessian) rises until it sits just above the theoretical divergence threshold implied by the learning rate, and training proceeds with loss decreasing over the long run despite bouncing non-monotonically step to step, complicating the simpler textbook picture of smooth, monotonic descent

Variants (Schedules)

  • Constant — a single fixed value for the whole run. Simple but rarely optimal; fine for short runs or well-behaved convex problems
  • Step decay — multiply the learning rate by a fixed factor (e.g., 0.1) every N epochs. Easy to reason about, common in older CNN training recipes
  • Exponential decay — continuously multiply the learning rate by a decay factor each step or epoch, giving a smooth downward curve rather than sudden drops
  • Cosine annealing — decay the learning rate following a cosine curve from an initial value down to (near) zero over training. Popular in modern deep learning because it decays smoothly and predictably
  • Warmup — start the learning rate near zero and linearly ramp it up over the first few hundred/thousand steps before switching to the main schedule. Critical for large transformer models, where large early updates on randomly initialized weights can destabilize training
  • Cyclical learning rates — oscillate the learning rate between a lower and upper bound over the course of training, which can help the model escape sharp local minima and sometimes reduces the need for extensive tuning
  • Adaptive (per-parameter) — optimizers like Adam, RMSprop, and Adagrad adjust the effective learning rate for each parameter individually based on that parameter’s gradient history, rather than following an externally imposed schedule at all
  • One-cycle policy — ramps the learning rate up from a low value to a high peak over roughly the first half of training, then back down (often below the starting value) for the second half, frequently paired with an inverse momentum schedule. Popularized as “superconvergence,” letting some models reach a given accuracy in far fewer epochs than a standard schedule
  • Reduce-on-plateau — monitors validation loss and cuts the learning rate by a fixed factor whenever it stops improving for a set number of epochs, adapting the schedule to actual training progress rather than a predetermined step count
  • Polynomial decay — decays the learning rate following lr * (1 - step/total_steps)^power, generalizing linear decay (power=1) to curves that fall off faster or slower near the end of training; common in semantic segmentation and detection training recipes
  • Noam schedule — the warmup-plus-inverse-square-root-decay schedule introduced alongside the original Transformer (Vaswani et al., 2017) and named after coauthor Noam Shazeer; ramps the learning rate up linearly during warmup, then decays it proportionally to the inverse square root of the step number

Visualizing a Warmup, Peak, and Decay Schedule

A schedule isn’t one number, it’s a path the learning rate follows over the course of training. Mermaid has no native line-chart type, but the same warmup-peak-decay shape used by most large-model training recipes maps cleanly onto a sequence of named stages:

Skipping straight to Peak with no Warmup stage is one of the most common causes of early divergence in large-model training (see Common Pitfalls below) — the randomly initialized weights simply aren’t ready for full-sized updates yet.

Why It Matters

  • Frequently the single most impactful hyperparameter to tune — a well-chosen learning rate can be the difference between a model that trains and one that never converges
  • Explains why the same architecture can produce wildly different results across training runs
  • A learning rate that’s even slightly too high can look like a modeling or data problem (loss plateaus or spikes) when the actual fix is a smaller step size or a warmup phase
  • Choosing the learning rate well often matters more than choosing between similar optimizers — a poorly tuned Adam run can underperform a well-tuned SGD run, and vice versa
  • It’s one of the few hyperparameters with a genuinely non-monotonic relationship to performance — both “too high” and “too low” hurt, with a broad but bounded region in between that actually works, unlike hyperparameters where “more” or “less” is at least directionally predictable
  • The learning rate interacts with almost every other design decision in a training run — batch size, optimizer choice, model depth, even numerical precision (mixed-precision training is more sensitive to a too-high learning rate than full precision) — which is why it’s usually the last thing tuned, after the rest of the setup is fixed

Common Pitfalls

  • Using one fixed learning rate for an entire long training run instead of a decay schedule
  • Not doing a learning-rate sweep/warmup, especially for large models, which are highly sensitive early in training
  • Copying a learning rate from a different codebase/paper without accounting for differences in batch size, optimizer, or model scale
  • Mistaking a too-high learning rate’s symptoms (loss spikes, NaN losses) for a data bug or architecture bug
  • Decaying the learning rate too aggressively too early, effectively freezing the model before it has learned useful representations
  • Using the same learning rate for fine-tuning a pretrained model as was used for training it from scratch — fine-tuning usually needs a much smaller value to avoid destroying pretrained weights
  • Using the same learning rate across all layers when fine-tuning, instead of smaller rates for early (more general-purpose) layers and larger rates for later, task-specific layers — a technique called discriminative or layer-wise learning rates
  • Restarting a scheduler’s step count from zero after resuming a checkpoint, which silently replays the warmup phase or resets a decay curve partway through what should be a continuous schedule
  • Tuning the learning rate on a small-scale proxy run and assuming the exact same value transfers cleanly to a much larger model or dataset, without accounting for how loss curvature and gradient noise both change with scale
  • Using a single learning rate across every parameter group when different parts of a model (e.g., a newly-initialized classification head vs. a pretrained backbone) genuinely need different step sizes
  • Not accounting for gradient accumulation when reasoning about the effective learning rate — accumulating gradients over several micro-batches before stepping changes the effective batch size, which interacts with the learning rate the same way an actual larger batch would

Best Practices

  • Run a learning rate range test (gradually increase the learning rate over a short warm-up run and watch where loss starts to blow up) to find a reasonable upper bound before committing to a schedule
  • Pair a warmup phase with a decay schedule (e.g., linear warmup + cosine decay) for training large models from scratch
  • Use a smaller learning rate (often 10-100x smaller) when fine-tuning a pretrained model compared to training from scratch
  • Log the learning rate alongside the loss curve — a sudden loss spike that lines up with a scheduled learning rate change is diagnostic, not coincidental
  • Prefer adaptive optimizers (Adam/AdamW) as a starting point if you don’t have time to tune a schedule by hand; switch to a hand-tuned SGD + schedule setup only when squeezing out extra performance matters
  • Save the learning rate schedule state alongside model and optimizer checkpoints so a resumed run continues the schedule rather than restarting it
  • When fine-tuning, consider freezing early layers entirely or applying discriminative learning rates rather than a single global rate for the whole network
  • Treat the learning rate range test as a cheap, repeatable diagnostic — rerun it after any significant change to batch size, model architecture, or dataset, rather than trusting a value tuned for a previous setup
  • When in doubt between two learning rates, prefer the smaller one paired with more training steps over the larger one — a too-low learning rate wastes compute, but a too-high one can silently damage a run in ways that are hard to detect until much later
  • Visualize the actual learning rate curve produced by your scheduler code before a long training run, not just the formula — off-by-one errors in step counting or a misconfigured total_steps routinely produce a decay curve that doesn’t match what was intended

Worked Example

Tracing the update rule w := w - lr * grad by hand for L(w) = (w - 3)² (so the gradient is grad = 2(w - 3)), starting every run from w = 0, makes the three scenarios concrete rather than just visual:

Steplr = 0.01 (too low)lr = 0.4 (well tuned)lr = 1.05 (too high)
00.00000.00000.0000
10.06002.40006.3000
20.11882.8800-0.6300
30.17642.97606.9930
40.23292.9952-1.3923

All three are chasing the same target, w = 3 (the minimum of L). After only 4 steps, lr = 0.4 is already at 2.9952 — visually indistinguishable from converged. lr = 0.01 has crawled to 0.2329, on pace to need several hundred steps to arrive. lr = 1.05 is oscillating with growing amplitude (6.3 → -0.63 → 6.99 → -1.39 → …) — each swing larger than the last, the hallmark of divergence rather than a slow, controlled overshoot. The runnable version below carries all three forward until they either converge or blow up, rather than stopping at step 4 by hand.

Code Example

import math
import torch

model = torch.nn.Linear(10, 1)
optimizer = torch.optim.AdamW(model.parameters(), lr=1e-3)

# Linear warmup for 500 steps, then cosine decay to near zero
warmup_steps = 500
total_steps = 10_000

def lr_lambda(step):
    if step < warmup_steps:
        return step / warmup_steps
    progress = (step - warmup_steps) / (total_steps - warmup_steps)
    return 0.5 * (1 + math.cos(math.pi * progress))

scheduler = torch.optim.lr_scheduler.LambdaLR(optimizer, lr_lambda)

for step, (x, y) in enumerate(dataloader):
    optimizer.zero_grad()
    loss = loss_fn(model(x), y)
    loss.backward()
    optimizer.step()
    scheduler.step()

A minimal from-scratch illustration of step decay, without any framework scheduler:

base_lr = 0.1
decay_factor = 0.5
decay_every = 10  # epochs

for epoch in range(50):
    current_lr = base_lr * (decay_factor ** (epoch // decay_every))
    for x_batch, y_batch in dataloader:
        # ... forward pass, compute loss ...
        for param in model_params:
            param -= current_lr * param.grad  # manual update using the decayed rate
    print(f"epoch {epoch}: lr={current_lr:.5f}")

Carrying the same three learning rates forward until each one actually converges or diverges, rather than stopping at step 4 by hand:

Comparison

ScheduleBehaviorBest forTuning effort
ConstantFixed value throughoutShort runs, quick prototypingLow
Step decaySudden drops at fixed intervalsClassic CNN training recipesMedium (choose drop points)
Exponential decaySmooth continuous decreaseGeneral-purposeLow
Cosine annealingSmooth curve to near-zeroModern deep learning, long runsLow
Warmup + decayRamp up, then decayLarge transformersMedium
CyclicalOscillates between boundsEscaping sharp minimaMedium
Adaptive (Adam, etc.)Per-parameter, self-adjustingDefault choice, minimal tuningVery low
Polynomial decaySmooth decrease along a power curveSegmentation/detection recipesLow
Noam (warmup + inverse-sqrt decay)Ramp up, then decay proportional to 1/sqrt(step)Original Transformer-style trainingMedium
One-cycle policyRamp up, then back down past the startFast training within a fixed budgetMedium
Reduce-on-plateauCuts the rate when validation loss stallsTraining runs without a preset scheduleLow

History

Early neural network training used a single fixed learning rate, tuned by hand and often left unchanged for an entire run. Robbins and Monro’s 1951 stochastic approximation theory gave the first formal conditions under which a decaying learning rate guarantees convergence for stochastic optimization. Adagrad (Duchi, Hazan, Singer, 2011) introduced per-parameter adaptive rates based on accumulated squared gradients, aimed at sparse features like word embeddings. Adam (Kingma and Ba, 2014) combined that idea with momentum and became the dominant default. Leslie Smith’s 2015-2017 work on cyclical learning rates and the one-cycle policy showed that deliberately oscillating or overshooting the learning rate could train some models to a target accuracy several times faster than a conventional decaying schedule — a result still referenced whenever “superconvergence” comes up.

Selected milestones not covered above:

  • 1958 — Frank Rosenblatt’s Perceptron used a simple fixed update rule, an early ancestor of the learning-rate-scaled update, though “learning rate” as a deliberately tuned hyperparameter became standard terminology later, alongside backpropagation-trained networks
  • 2017 — the original Transformer paper (Vaswani et al.) introduces the warmup-plus-inverse-square-root-decay schedule, later nicknamed the “Noam schedule” after coauthor Noam Shazeer
  • 2021 — Jeremy Cohen, Simran Kaur, Yuanzhi Li, Zico Kolter, and Ameet Talwalkar’s “Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability” (ICLR 2021) empirically shows that real training runs operate right at the boundary the learning rate and loss curvature jointly define, rather than safely within it

Real-World Example

Training recipes for large transformer models commonly specify something like: linear warmup from 0 to a peak learning rate of 6e-4 over the first 2,000 steps, followed by cosine decay down to 10% of the peak value over the remaining steps of training. The peak value itself is chosen based on model size and batch size — larger batches (more gradient averaging, less noise) tolerate proportionally higher peak learning rates. Skipping the warmup on such a model is a well-documented failure mode: the randomly initialized attention layers receive large gradient-scaled updates before they’ve settled into a reasonable configuration, often producing NaN losses or a model that never recovers within the training budget.

Two more specific, named examples of learning rate choices from published training recipes:

  • GPT-3 (Brown et al., 2020) — the 175B-parameter model was trained with a peak learning rate of 6e-5, warmed up linearly over the first 375 million tokens and then decayed via a cosine schedule down to 10% of the peak value over the rest of training
  • BERT (Devlin et al., 2018) — pretraining used Adam with a peak learning rate of 1e-4, linear warmup over the first 10,000 steps, and linear decay for the remainder — a considerably higher peak than GPT-3’s, reflecting differences in model scale, batch size, and architecture between the two setups

FAQ

What’s a good starting learning rate? There’s no universal number — it depends on the optimizer, batch size, and model. As rough starting points: 1e-3 is common for Adam on small-to-medium networks, 1e-4 to 5e-5 for fine-tuning transformers, and 0.01-0.1 for SGD with momentum. Always verify with a small range test rather than trusting a default blindly.

Why does the learning rate need to change during training? Early in training, large steps help the model quickly move away from a poor random initialization. Later, as it approaches a good region of the loss landscape, large steps cause it to bounce around instead of settling — a smaller step size lets it fine-tune into a good minimum.

Do adaptive optimizers eliminate the need to tune the learning rate? No — they reduce sensitivity but don’t eliminate it. Adam still needs a reasonable base learning rate; set it too high and it will still diverge, just less catastrophically than plain SGD would.

Should the learning rate scale with batch size? Generally yes. The “linear scaling rule” suggests multiplying the learning rate by the same factor you multiply the batch size by, since a larger batch produces a less noisy, more reliable gradient estimate that can support a proportionally larger step — though this breaks down at very large batch sizes and usually needs an accompanying warmup to stay stable.

How is learning rate different from momentum? Learning rate controls the step size taken in the direction of the current gradient signal; momentum controls how much of the previous update direction carries over into the current one. They’re complementary — momentum smooths out the path, learning rate sets how fast you move along it — and both interact with the same underlying update, so tuning one without considering the other is easy to get wrong.

Can the learning rate be negative, or larger than 1? Negative learning rates would climb the gradient instead of descending it, so no. Values above 1 are technically legal but almost always cause divergence on any reasonably well-scaled loss, since the update rule w := w - lr * grad overshoots the minimum by an amount proportional to lr — very large learning rates are occasionally used deliberately for a brief warmup-style range test, but never for the bulk of training.

Does the optimal learning rate change over the course of a single training run, or only between runs? Within a run — that’s the entire premise behind schedules. Early in training, the loss surface is usually far from any minimum and can tolerate large steps; later, closer to a good region, the same step size causes oscillation. A schedule is just an explicit acknowledgment that the “right” learning rate is a moving target, not a constant.

Common Interview Questions

  • Why do we warm up the learning rate instead of starting at the peak value? (Randomly initialized weights, especially in attention layers, produce unreliable early gradients; a small initial learning rate avoids destabilizing updates before the model has settled.)
  • What symptoms indicate the learning rate is too high versus too low? (Too high: loss spikes, oscillates, or produces NaNs. Too low: loss decreases but extremely slowly, or plateaus early at a mediocre value.)
  • Why does fine-tuning use a smaller learning rate than training from scratch? (Pretrained weights already encode useful structure; a large learning rate risks overwriting it before the model adapts to the new task.)
  • If two training runs use the same architecture and data but different learning rates, why might their final accuracy differ significantly? (Different step sizes trace different paths through a non-convex loss surface, landing in different local minima with different generalization properties — the learning rate doesn’t just affect speed, it affects which solution is found.)
  • How would you diagnose whether a stalled loss curve is caused by too low a learning rate versus a genuine local minimum? (Try increasing the learning rate briefly, or restart with a higher one — a genuinely stalled optimization from too small a step size will start moving again, while a true local minimum or saddle point typically needs a structural change like momentum or a different initialization.)
  • Why might two different learning rate schedules produce the same final loss but different generalization performance? (Different schedules trace different paths through a non-convex loss surface, landing in different minima; flatter minima — often favored by schedules that don’t decay too aggressively too early — tend to generalize better than sharp ones, even when both fit the training data equally well.)

Example

Dropping the learning rate from 0.01 to 0.001 partway through training often stabilizes a model that was previously oscillating around its best accuracy. Concretely: a model trained with a constant learning rate of 0.01 might reach 85% validation accuracy but keep bouncing between 82-85% every epoch as it overshoots the minimum each step. Switching to 0.001 — either via a manual step-decay schedule or a scheduler like ReduceLROnPlateau triggered when validation loss stops improving — lets the same model settle into the minimum and often gain another 2-3 points of accuracy in the following epochs, purely from taking smaller, more precise steps once it’s already in the right neighborhood.

Dig deeper