3 October 2026 · 6 min read
Weight decay vs L2 regularization: why they differ under Adam
Weight decay and L2 regularization both pull a model's weights toward zero, and with plain gradient descent they are the same thing. This post explains what each one does, why they stop being the same under Adam, and why AdamW is the usual fix.
Weight decay is a training technique that shrinks every weight of a model a little toward zero at each step, so that no weight grows large unless the data keeps pushing it up. L2 regularization reaches for the same goal by adding a penalty on large weights to the loss. With plain gradient descent the two produce identical updates, which is why the names are often used as if they meant one thing. With Adam, the optimizer most deep learning models train with, they give different results, and that difference is the reason AdamW exists.
What problem do they solve?
A model learns by adjusting its weights, the numbers inside it, until its predictions match the training examples. When a model has many weights and few examples, some weights can grow large to fit details that only exist in the training set, such as noise in a measurement or a mislabeled example. Large weights also make the output react strongly to small changes in the input, so a little noise in a new example can swing the prediction. This is called overfitting: low error on the training data and higher error on everything else.
Both techniques push back on large weights, so the model fits the patterns that repeat and ignores more of the noise.
L2 regularization: a penalty in the loss
The loss is the number training tries to make small. L2 regularization adds a second term to it. Call the original loss $L_{\text{data}}$, call the weights $\theta$, and call the penalty strength $\lambda$ (a small number such as $0.01$). The new loss is
$$L_{\text{total}} = L_{\text{data}} + \frac{\lambda}{2} \lVert \theta \rVert^2$$
$\lVert \theta \rVert^2$ is the sum of every weight squared. The total loss now goes up when the weights get large, so training has a reason to keep them small.
Training moves the weights using the gradient, the direction in which the loss increases fastest. The gradient of the penalty term is one extra piece:
$$\nabla_\theta \left( \frac{\lambda}{2} \lVert \theta \rVert^2 \right) = \lambda \theta$$
So the gradient the optimizer receives becomes $g + \lambda \theta$, where $g$ is the gradient of the data loss. Each weight gets an extra push toward zero that is proportional to its own size.
Weight decay: shrinking the weights directly
Weight decay skips the loss and changes the update step instead. Plain gradient descent updates each weight by subtracting the gradient times a step size $\eta$, called the learning rate. Weight decay adds one more subtraction:
$$\theta \leftarrow \theta - \eta g - \eta \lambda \theta$$
The last term shrinks every weight by the same fraction, $\eta \lambda$, at every step.

Now put the L2 gradient into plain gradient descent: $\theta \leftarrow \theta - \eta (g + \lambda \theta)$. Multiply it out and you get exactly the weight decay update above. For plain gradient descent, L2 regularization and weight decay are the same algorithm.
Why Adam breaks the equivalence
Adam does not use the raw gradient as its step. It keeps two running averages for every weight: $\hat m$, an average of recent gradients, and $\hat v$, an average of recent squared gradients. The hats mean both averages are corrected for starting at zero. The update divides one by the square root of the other:
$$\theta \leftarrow \theta - \eta \frac{\hat m}{\sqrt{\hat v} + \epsilon}$$
$\epsilon$ is a tiny constant that prevents division by zero. Dividing by $\sqrt{\hat v}$ gives every weight a step of similar size, whether its gradients are large or small. This is what makes Adam adaptive, and it helps when different weights see gradients of very different scales.
With L2 regularization, the $\lambda \theta$ term gets added to the gradient before Adam sees it, so it goes into $\hat m$ and $\hat v$ and then gets divided by $\sqrt{\hat v}$ like everything else. A weight whose data gradients are large has a large $\hat v$, so its decay term gets divided by a large number and almost disappears. A weight whose data gradients are tiny has a small $\hat v$, so its decay term gets inflated. The pull toward zero ends up depending on gradient history instead of on the size of the weight.
How AdamW fixes it
AdamW, short for Adam with decoupled weight decay, takes the decay out of the gradient and applies it as its own step after the adaptive update:
$$\theta \leftarrow \theta - \eta \frac{\hat m}{\sqrt{\hat v} + \epsilon} - \eta \lambda \theta$$
The adaptive part handles the data gradient. The last term shrinks every weight by the same fraction $\eta \lambda$, no matter what its gradients looked like.

The difference is easy to see in code. The snippet below starts two weights at 1.0, gives one of them large gradients and the other small ones, and makes both gradients flip sign every step so that they average to zero. With no real data signal, only the decay should move the weights.
1import torch
2
3def run(name, make_opt):
4 w = torch.nn.Parameter(torch.tensor([1.0, 1.0]))
5 opt = make_opt([w])
6 scale = torch.tensor([10.0, 0.01]) # weight 0 sees large gradients, weight 1 small ones
7 for step in range(500):
8 opt.zero_grad()
9 w.grad = scale * (-1) ** step # a data gradient that flips sign, so it averages to zero
10 opt.step()
11 print(f"{name:9s} large-gradient weight: {w[0].item():.3f} small-gradient weight: {w[1].item():.3f}")
12
13run("Adam + L2", lambda p: torch.optim.Adam(p, lr=0.01, weight_decay=0.1))
14run("AdamW", lambda p: torch.optim.AdamW(p, lr=0.01, weight_decay=0.1))It prints:
1Adam + L2 large-gradient weight: 0.935 small-gradient weight: 0.000
2AdamW large-gradient weight: 0.596 small-gradient weight: 0.596PyTorch's Adam with weight_decay is the L2 form. The weight with large gradients barely shrank, and the weight with small gradients was driven all the way to zero. AdamW shrank both by the same factor, $0.999^{500} \approx 0.606$ plus the small wobble from the flipping gradient, which is what weight decay is supposed to do.
Weight decay in PyTorch
For most models trained with Adam, the setup is one line:
1optimizer = torch.optim.AdamW(model.parameters(), lr=1e-3, weight_decay=0.01)weight_decay=0.01 is PyTorch's default for AdamW and a common starting point. Values between 0.01 and 0.1 are typical for transformers. If validation loss is much worse than training loss, try a larger value. If the model can no longer fit the training data well, try a smaller one.
It is common to exclude biases and normalization parameters from decay, since shrinking them does little to reduce overfitting. Both are one-dimensional tensors, so a short filter covers them:
1decay = [p for p in model.parameters() if p.dim() > 1]
2no_decay = [p for p in model.parameters() if p.dim() <= 1]
3optimizer = torch.optim.AdamW([
4 {"params": decay, "weight_decay": 0.01},
5 {"params": no_decay, "weight_decay": 0.0},
6], lr=1e-3)
Common mistakes
- Using
Adam(weight_decay=...)and expecting weight decay. That argument adds the L2 term to the gradient, which is the coupled form shown above. Switch toAdamWif you want decoupled decay. - Copying a
weight_decayvalue between optimizers. The shrink per step is $\eta \lambda$, and SGD and AdamW usually run at very different learning rates, so a value tuned for one rarely transfers to the other. - Adding an L2 term to the loss by hand and also setting
weight_decay. This applies the penalty twice.
Related concepts
L1 regularization penalizes the sum of absolute weight values instead of squares, which pushes many weights to exactly zero rather than making all of them small. Dropout reduces overfitting a different way, by randomly zeroing part of a layer's outputs during training. Early stopping ends training when validation loss stops improving, before the weights have had time to fit the noise.
QuiddityML teaches weight decay and AdamW as concepts in its ML Foundation track, and the exercises on them include ordering the lines of an AdamW step, spotting the bug in a weight decay loss, and writing AdamW from scratch.