QuiddityML

3 October 2026 · 5 min read

What is dropout and why does it work?

Dropout is a regularization technique that randomly switches off part of a neural network during training so the model does better on new data. This post explains what dropout does on each step, why it reduces overfitting, how it behaves at test time, and how to use it in PyTorch.

Dropout is a technique for training neural networks in which, on every training step, a random fraction of a layer's outputs is set to zero. It is used to reduce overfitting, the situation where a model does very well on the examples it trained on and worse on new ones. Dropout layers are common in fully connected networks and in transformers, and in PyTorch adding one is a single line.

What problem does dropout solve?

A neural network is built from layers of neurons. Each neuron takes the outputs of the neurons in the layer before it, multiplies each by a weight, adds them up, and passes the result on. Training adjusts those weights until the network's predictions match the training examples.

A large network has enough weights to match the training examples very closely, including details that only exist in that particular dataset, such as noise or a few mislabeled examples. One way this shows up inside the network is that neurons come to rely on very specific partners. A neuron in one layer may learn to fire only when one particular neuron in the layer below fires, and that pairing may only be useful on the training data. Dropout is designed to break up those pairings.

What dropout does on each training step

Dropout has one setting, the dropout rate $p$, which is the probability that each output gets zeroed. With $p = 0.5$, on each forward pass every neuron in the layer has a 50% chance of having its output set to zero, chosen independently and at random.

Two training steps of the same network with p = 0.5, each with a different set of hidden neurons crossed out as dropped

A layer with $n$ neurons has $2^n$ possible patterns of dropped and kept neurons, so training samples a different smaller network, called a subnetwork, on almost every step. All of these subnetworks share the same weights.

Why does dropout reduce overfitting?

Any neuron can vanish on any training step, so a neuron in the next layer cannot depend on one particular input neuron always being there. If it relied on a single partner, its output would break whenever that partner was dropped, and the loss would push it to change. Training tends to settle on weights where each useful signal is carried by several neurons at once.

This has three effects:

Three panels showing a downstream neuron depending on one upstream neuron, losing its signal when that neuron is dropped, and keeping its signal when it draws on several upstream neurons

What happens at test time?

At test time you want every neuron to contribute, because random zeroing would make predictions noisy and different on every run. So dropout is switched off.

That creates a mismatch. With $p = 0.5$, during training a neuron receives input from about half of the neurons below it. At test time it receives input from all of them, so the total it adds up would be about twice as large as anything it saw during training.

The usual fix is called inverted dropout: during training, after zeroing, multiply the surviving outputs by $\frac{1}{1-p}$. With $p = 0.5$ that means the survivors are doubled. The average output of the layer then stays the same with or without dropout, so nothing needs to change at test time.

1mask = (torch.rand_like(x) > p).float()   # 1 keeps a value, 0 drops it
2x = x * mask / (1 - p)                     # scale the survivors so the average stays the same

During training with p = 0.5, half the inputs are dropped and survivors are scaled by 2, and at test time with model.eval() all inputs are used with no scaling

PyTorch's nn.Dropout does this for you. The snippet below passes a vector of ones through a dropout layer in both modes:

1import torch
2import torch.nn as nn
3 
4torch.manual_seed(0)
5drop = nn.Dropout(p=0.5)
6x = torch.ones(8)
7 
8drop.train()
9print("train:", drop(x))
10print("train mean over 10,000 values:", drop(torch.ones(10_000)).mean().item())
11 
12drop.eval()
13print("eval: ", drop(x))
1train: tensor([2., 0., 0., 2., 2., 0., 0., 2.])
2train mean over 10,000 values: 0.9950000047683716
3eval:  tensor([1., 1., 1., 1., 1., 1., 1., 1.])

In training mode about half the values are zeroed and the rest become 2, so the average stays close to 1. In evaluation mode the layer passes its input through unchanged.

Dropout in a PyTorch model

Dropout layers usually go after the activation function of a hidden layer:

1model = nn.Sequential(
2    nn.Linear(784, 256), nn.ReLU(), nn.Dropout(p=0.5),
3    nn.Linear(256, 256), nn.ReLU(), nn.Dropout(p=0.5),
4    nn.Linear(256, 10),
5)
6 
7model.train()                 # dropout on, for training
8# ... training loop ...
9 
10model.eval()                  # dropout off, for validation and inference
11with torch.no_grad():         # stop recording operations for gradients
12    preds = model(x_val)

model.eval() and torch.no_grad() do different jobs. model.eval() changes what the layers do, which for dropout means it stops dropping. torch.no_grad() tells PyTorch not to record operations for backpropagation, which saves memory and time but has no effect on dropout. Evaluation usually needs both.

When to use dropout, and common mistakes

Dropout is a common choice when a network has many more weights than training examples and validation loss starts rising while training loss keeps falling. Typical rates are 0.1 to 0.3 in transformers and up to 0.5 in fully connected layers. If the model can no longer fit the training data well, lower $p$. If it still overfits, raise it or combine it with weight decay.

Weight decay reduces overfitting by pulling every weight toward zero instead of dropping neurons. Early stopping ends training at the point where validation loss is lowest. Batch normalization rescales the outputs of a layer during training, and some architectures that use it rely less on dropout.

QuiddityML teaches dropout as its own concept in the ML Foundation track, and the exercises on it include writing inverted dropout from scratch, putting the lines of an MLP with dropout in order, and predicting what happens when model.eval() is missing.