3 October 2026 · 5 min read
What is dropout and why does it work?
Dropout is a regularization technique that randomly switches off part of a neural network during training so the model does better on new data. This post explains what dropout does on each step, why it reduces overfitting, how it behaves at test time, and how to use it in PyTorch.
Dropout is a technique for training neural networks in which, on every training step, a random fraction of a layer's outputs is set to zero. It is used to reduce overfitting, the situation where a model does very well on the examples it trained on and worse on new ones. Dropout layers are common in fully connected networks and in transformers, and in PyTorch adding one is a single line.
What problem does dropout solve?
A neural network is built from layers of neurons. Each neuron takes the outputs of the neurons in the layer before it, multiplies each by a weight, adds them up, and passes the result on. Training adjusts those weights until the network's predictions match the training examples.
A large network has enough weights to match the training examples very closely, including details that only exist in that particular dataset, such as noise or a few mislabeled examples. One way this shows up inside the network is that neurons come to rely on very specific partners. A neuron in one layer may learn to fire only when one particular neuron in the layer below fires, and that pairing may only be useful on the training data. Dropout is designed to break up those pairings.
What dropout does on each training step
Dropout has one setting, the dropout rate $p$, which is the probability that each output gets zeroed. With $p = 0.5$, on each forward pass every neuron in the layer has a 50% chance of having its output set to zero, chosen independently and at random.
- A different random set of neurons is dropped on every step, so each batch of training data passes through a different smaller network.
- A dropped neuron sends nothing forward on that step, so it also receives no gradient and its incoming weights do not update on that step.
- The neurons are only dropped for that one step. On the next step all of them are back and a new random set is dropped.

A layer with $n$ neurons has $2^n$ possible patterns of dropped and kept neurons, so training samples a different smaller network, called a subnetwork, on almost every step. All of these subnetworks share the same weights.
Why does dropout reduce overfitting?
Any neuron can vanish on any training step, so a neuron in the next layer cannot depend on one particular input neuron always being there. If it relied on a single partner, its output would break whenever that partner was dropped, and the loss would push it to change. Training tends to settle on weights where each useful signal is carried by several neurons at once.
This has three effects:
- Information is spread out. Each pattern the network detects is encoded across several neurons, so losing one neuron does not erase it.
- Features hold up when parts of the input change. A feature that still works after random neurons are removed tends to be less tied to the exact training examples.
- The network behaves like a combination of many models. Combining the predictions of several separately trained models, called an ensemble, usually generalizes better than any one of them. Dropout trains a large number of subnetworks with shared weights and uses all of them together at test time, which gives some of the same benefit for the cost of training one network.

What happens at test time?
At test time you want every neuron to contribute, because random zeroing would make predictions noisy and different on every run. So dropout is switched off.
That creates a mismatch. With $p = 0.5$, during training a neuron receives input from about half of the neurons below it. At test time it receives input from all of them, so the total it adds up would be about twice as large as anything it saw during training.
The usual fix is called inverted dropout: during training, after zeroing, multiply the surviving outputs by $\frac{1}{1-p}$. With $p = 0.5$ that means the survivors are doubled. The average output of the layer then stays the same with or without dropout, so nothing needs to change at test time.
1mask = (torch.rand_like(x) > p).float() # 1 keeps a value, 0 drops it
2x = x * mask / (1 - p) # scale the survivors so the average stays the same
PyTorch's nn.Dropout does this for you. The snippet below passes a vector of ones through a dropout layer in both modes:
1import torch
2import torch.nn as nn
3
4torch.manual_seed(0)
5drop = nn.Dropout(p=0.5)
6x = torch.ones(8)
7
8drop.train()
9print("train:", drop(x))
10print("train mean over 10,000 values:", drop(torch.ones(10_000)).mean().item())
11
12drop.eval()
13print("eval: ", drop(x))1train: tensor([2., 0., 0., 2., 2., 0., 0., 2.])
2train mean over 10,000 values: 0.9950000047683716
3eval: tensor([1., 1., 1., 1., 1., 1., 1., 1.])In training mode about half the values are zeroed and the rest become 2, so the average stays close to 1. In evaluation mode the layer passes its input through unchanged.
Dropout in a PyTorch model
Dropout layers usually go after the activation function of a hidden layer:
1model = nn.Sequential(
2 nn.Linear(784, 256), nn.ReLU(), nn.Dropout(p=0.5),
3 nn.Linear(256, 256), nn.ReLU(), nn.Dropout(p=0.5),
4 nn.Linear(256, 10),
5)
6
7model.train() # dropout on, for training
8# ... training loop ...
9
10model.eval() # dropout off, for validation and inference
11with torch.no_grad(): # stop recording operations for gradients
12 preds = model(x_val)model.eval() and torch.no_grad() do different jobs. model.eval() changes what the layers do, which for dropout means it stops dropping. torch.no_grad() tells PyTorch not to record operations for backpropagation, which saves memory and time but has no effect on dropout. Evaluation usually needs both.
When to use dropout, and common mistakes
Dropout is a common choice when a network has many more weights than training examples and validation loss starts rising while training loss keeps falling. Typical rates are 0.1 to 0.3 in transformers and up to 0.5 in fully connected layers. If the model can no longer fit the training data well, lower $p$. If it still overfits, raise it or combine it with weight decay.
- Forgetting
model.eval(). The model then evaluates with dropout still on, so validation scores come out lower than they should and change from run to run. - Forgetting to switch back with
model.train(). If evaluation runs inside the training loop, the next epoch trains with dropout off unless the mode is set back. - Using dropout on a model that is underfitting. If training loss is still high, the model has not fit the data yet, and dropout makes that harder.
- Expecting training loss to match validation loss. With dropout on, training loss is measured on a weakened network, so it can sit above the validation loss.
Related concepts
Weight decay reduces overfitting by pulling every weight toward zero instead of dropping neurons. Early stopping ends training at the point where validation loss is lowest. Batch normalization rescales the outputs of a layer during training, and some architectures that use it rely less on dropout.
QuiddityML teaches dropout as its own concept in the ML Foundation track, and the exercises on it include writing inverted dropout from scratch, putting the lines of an MLP with dropout in order, and predicting what happens when model.eval() is missing.