QuiddityML

3 October 2026 · 5 min read

Label smoothing explained

Label smoothing is a small change to classification training that stops a model from becoming 100% sure of its answers. This post explains what it changes in the targets, why that helps, how to set it in PyTorch, and when to leave it off.

Label smoothing is a technique for training classifiers in which the correct answer is given a target probability slightly below 1 and every wrong answer gets a small probability above 0. It is used to stop a model from becoming fully confident in its predictions, which tends to make the probabilities it outputs more trustworthy and often improves accuracy on new data. Image classifiers and translation models are common places to find it, and in PyTorch it is one argument to the loss function.

What problem does label smoothing solve?

A classifier picks one of $K$ classes, for example cat, dog, bird, car, or horse. For each input it outputs $K$ raw scores called logits, and a function called softmax turns them into probabilities that add up to 1.

Training compares those probabilities to a target. The usual target is one-hot: probability 1.0 on the correct class and 0.0 on every other class. The loss used for classification, cross-entropy, is $-\log$ of the probability the model gives to the correct class. That loss only reaches zero when the probability is exactly 1.0, and softmax only outputs 1.0 when the correct class's logit is infinitely larger than the rest.

So with one-hot targets, training keeps pushing the logits further apart on every example, even ones the model already gets right. On a large network this tends to make it 99.9% sure on many inputs, including its mistakes and including mislabeled examples in the training set that it has memorized. A model that says 99.9% when it is right 80% of the time is called badly calibrated.

The smoothed target

Label smoothing replaces the one-hot target with a softer one. It has one setting, $\epsilon$ (epsilon), usually $0.1$. The smoothed target is

$$y_{\text{smooth}} = (1 - \epsilon) \cdot y_{\text{one-hot}} + \frac{\epsilon}{K}$$

In words: take $\epsilon$ of the probability away from the correct class and spread it evenly over all $K$ classes. With $\epsilon = 0.1$ and $K = 5$:

A one-hot target with 1.0 on cat and 0.0 on dog, bird, car, and horse, next to a smoothed target with 0.92 on cat and 0.02 on each other class

The prediction with the lowest loss now puts 0.92 on the correct class, not 1.0. Once the model reaches that, the loss stops rewarding it for pushing the logits further apart, and predicting more than 0.92 actually raises the loss. Training stops at a finite gap between the logits instead of chasing an infinite one.

Label smoothing in PyTorch

nn.CrossEntropyLoss takes a label_smoothing argument and builds the smoothed target internally:

1import torch
2import torch.nn as nn
3 
4logits = torch.tensor([[4.0, 1.0, 0.5, 0.2, -1.0]])   # one example, 5 classes
5target = torch.tensor([0])                             # the true class is class 0
6 
7for eps in (0.0, 0.1):
8    loss_fn = nn.CrossEntropyLoss(label_smoothing=eps)
9    print(f"label_smoothing={eps}: loss {loss_fn(logits, target).item():.3f}")
10 
11# the smoothed target PyTorch compares against
12K, eps = 5, 0.1
13smooth = torch.full((K,), eps / K)
14smooth[0] += 1 - eps
15print("smoothed target:", smooth)
1label_smoothing=0.0: loss 0.104
2label_smoothing=0.1: loss 0.410
3smoothed target: tensor([0.9200, 0.0200, 0.0200, 0.0200, 0.0200])

The prediction here is already confident in the correct class. Without smoothing the loss is low and would keep falling as the class-0 logit grows. With smoothing the loss is higher, because the target also asks for a little probability on the other four classes, and a larger class-0 logit would make it worse.

The effect on a trained model is easy to measure. The snippet below trains the same small network twice on 200 examples with random labels, where any confidence the model reaches comes from memorizing:

1import torch
2import torch.nn as nn
3 
4torch.manual_seed(0)
5X = torch.randn(200, 20)
6y = torch.randint(0, 5, (200,))                        # random labels, so a confident model is only memorizing
7 
8for eps in (0.0, 0.1):
9    torch.manual_seed(1)
10    model = nn.Sequential(nn.Linear(20, 64), nn.ReLU(), nn.Linear(64, 5))
11    opt = torch.optim.Adam(model.parameters(), lr=1e-2)
12    loss_fn = nn.CrossEntropyLoss(label_smoothing=eps)
13    for _ in range(500):
14        opt.zero_grad()
15        loss_fn(model(X), y).backward()
16        opt.step()
17    probs = model(X).softmax(dim=-1)
18    print(f"label_smoothing={eps}: mean top probability {probs.max(dim=-1).values.mean().item():.3f}")
1label_smoothing=0.0: mean top probability 0.999
2label_smoothing=0.1: mean top probability 0.919

Without smoothing the network is 99.9% sure of labels that were assigned at random. With smoothing its confidence stops near 0.92, the target it was given.

Writing it by hand

The same loss without the built-in argument takes two steps: turn the class indices into smoothed target rows, then compare them to the log of the softmax.

1def label_smoothing_loss(logits, targets, num_classes, eps=0.1):
2    log_probs = torch.log_softmax(logits, dim=-1)
3    smooth = torch.full_like(log_probs, eps / num_classes)
4    smooth.scatter_(1, targets.unsqueeze(1), 1 - eps + eps / num_classes)
5    return -(smooth * log_probs).sum(dim=-1).mean()

torch.log_softmax computes softmax followed by a log in one numerically stable step. scatter_ writes the correct-class value into each row at the position given by targets, which turns a column of class indices into full target rows without a loop.

When to use it, and common mistakes

Label smoothing is a common choice for classification with many classes and a large model, and when the predicted probabilities get used for decisions, such as flagging low-confidence predictions for a human to check. 0.1 is the usual starting value. Values above 0.2 are rare, because the target for the correct class then drops toward the targets for the wrong ones.

Weight decay pulls the weights toward zero, and dropout randomly zeroes part of a layer during training, so both limit how much the model can memorize. Label smoothing leaves the model's capacity alone and changes the target it aims for. Temperature scaling is a different calibration method that divides the logits by a constant after training instead of changing the training targets.

QuiddityML teaches label smoothing as its own concept in the ML Foundation track, and the exercises on it include tracing tensor shapes through a label smoothing loss and writing that loss from scratch.