QuiddityML

29 September 2026 · 8 min read

Dying ReLU: what dead neurons are and how to fix them

A dead ReLU neuron outputs zero for every input and stops learning, and a network can lose a large share of its neurons this way during training. This post explains why ReLU neurons die, how to measure the dead fraction in PyTorch, and how to prevent it with initialization, the learning rate, and Leaky ReLU.

A dead ReLU neuron is a neuron whose output is zero for every example in the training data, so it sends no gradient back and its weights stop changing. When many neurons die, the network keeps its full size in memory but trains with only part of it, and the loss often stalls higher than it should. This is called the dying ReLU problem.

How a ReLU neuron works

A neuron in a neural network takes its inputs $x_1, \dots, x_n$, multiplies each by a weight, adds them up, and adds a bias $b$. That sum is the neuron's pre-activation, written $z = \mathbf{w} \cdot \mathbf{x} + b$. The neuron's output is an activation function applied to $z$.

ReLU (rectified linear unit) is the most common activation function in hidden layers. It keeps positive values and sets negative ones to zero:

$$\text{ReLU}(z) = \max(0, z)$$

Training updates the weights with gradient descent. It computes a loss (one number that measures how wrong the predictions are), then the gradient of the loss with respect to each weight (how much the loss changes when that weight changes), and moves each weight a small step in the direction that lowers the loss. The gradients are computed backward from the output through each layer, a procedure called backpropagation, and each layer multiplies the gradient it receives by the derivative of its activation function.

What makes a ReLU neuron dead

ReLU's derivative is 1 when $z > 0$ and 0 when $z \leq 0$. For a single example with a negative pre-activation, that is normal: the neuron is off for that input, and it passes no gradient back for it. In a healthy ReLU layer about half the outputs are zero on any given example.

A neuron is dead when $z \leq 0$ for every example. Its output is then always 0, the derivative through ReLU is always 0, and the gradient reaching its weights and bias is 0 on every step.

A dead ReLU neuron: inputs x1, x2, x3 feed z = w·x + b, with z ≤ 0 for every example, so ReLU outputs 0 and the local gradient is 0. Occasional zero outputs are normal, and zero for every example is the warning sign.

Why dead neurons stay dead

Gradient descent updates a weight $w$ with the rule below, where $\eta$ is the learning rate (the step size) and $\partial \mathcal{L} / \partial w$ is the gradient of the loss $\mathcal{L}$ with respect to $w$:

$$w \leftarrow w - \eta \cdot \frac{\partial \mathcal{L}}{\partial w}$$

For a dead neuron the gradient is 0 on every example, so the update is 0 and the weights do not move. Nothing in the neuron's own update can push $z$ back above zero. The neuron can come back only if the inputs feeding it change enough, which happens when earlier layers are still training, so recovery is possible but rare.

Why a dead ReLU neuron stops learning: the pre-activation z ≤ 0 gives a ReLU output of 0, the local gradient ReLU'(z) is 0, so the total gradient and the updates to w and b are 0. The neuron outputs 0 on every input and cannot self-correct.

What kills ReLU neurons

Dead neurons outside ReLU

A neuron can die with any activation function that has a flat region where the gradient is exactly zero. ReLU6, which caps ReLU at 6, is flat both below 0 and above 6, and hard sigmoid and hard tanh are flat at both ends of their range. A neuron whose pre-activation stays in one of those flat regions for every example gets no gradient, the same as a dead ReLU neuron. Sigmoid and tanh never have an exactly zero gradient, but a neuron that is saturated on every example gets a gradient so small that it barely moves.

Regularization can do it too. L1 regularization adds the sum of the absolute values of the weights to the loss, which pushes many weights to exactly zero. If every incoming weight of a neuron is pushed to zero, its output no longer depends on the input, and it contributes nothing, whichever activation function it uses.

Dead neurons are reported most often with ReLU because ReLU is the most widely used activation in hidden layers and because its whole negative half has a gradient of exactly zero, so one bad update is enough to push a neuron there.

Detecting dead neurons in PyTorch

Two measurements help, and they measure different things. fraction_zero is the share of all activations in a batch that are exactly zero, which describes how sparse the layer is. For a healthy ReLU layer it sits near 0.5. fraction_dead checks each neuron down the batch dimension and counts the neurons that are zero for every example in the batch, which is the dead fraction.

1import torch
2import torch.nn as nn
3 
4def fraction_zero(h):
5    """Share of all activations in the batch that are exactly zero."""
6    return (h == 0).float().mean().item()
7 
8def fraction_dead(h):
9    """Share of units that are zero for every example. h has shape (batch, units)."""
10    return (h == 0).all(dim=0).float().mean().item()
11 
12torch.manual_seed(0)
13layer = nn.Linear(20, 64)
14h = torch.relu(layer(torch.randn(512, 20)))
15print(f"zero {fraction_zero(h):.2f}  dead {fraction_dead(h):.2f}")
1zero 0.51  dead 0.00

A freshly initialized layer has about half its outputs at zero and no dead neurons. Some warning signs to watch for, logged per layer every few hundred steps:

Bar chart of the fraction of zero activations per layer for six layers, with layers 4 and 5 above a 0.5 warning line, the warning signs listed below it, and the one-line check (activations == 0).float().mean().

fraction_dead only sees one batch, so a neuron it counts as dead might still fire on examples outside that batch. Checking a few batches, or a larger one, gives a better estimate.

What happens after neurons die

The experiment below trains a small network to learn $y = 3x_1$ from 20 input features. Before training starts, it subtracts 2 from every bias in the hidden layer, which puts the pre-activations of most neurons below zero, the same state one oversized update can leave behind. It then trains for 1,000 steps and checks the dead fraction again. The pre-activation check $z \leq 0$ is used here so the same count works for Leaky ReLU, which never outputs exactly zero.

1import torch
2import torch.nn as nn
3 
4def fraction_dead(z):
5    """Share of units whose pre-activation is <= 0 for every example."""
6    return (z <= 0).all(dim=0).float().mean().item()
7 
8def experiment(act):
9    torch.manual_seed(0)
10    x = torch.randn(512, 20)
11    y = 3 * x[:, :1] + 0.1 * torch.randn(512, 1)
12    model = nn.Sequential(nn.Linear(20, 64), act, nn.Linear(64, 1))
13    opt = torch.optim.SGD(model.parameters(), lr=0.05)
14 
15    with torch.no_grad():
16        model[0].bias -= 2.0  # what one oversized update can do to the biases
17    print(f"  dead after the push: {fraction_dead(model[0](x)):.2f}")
18 
19    for _ in range(1000):
20        loss = ((model(x) - y) ** 2).mean()
21        opt.zero_grad()
22        loss.backward()
23        opt.step()
24    z = model[0](x)
25    print(f"  after 1000 steps: dead {fraction_dead(z):.2f}  loss {loss.item():.3f}")
26 
27print("ReLU")
28experiment(nn.ReLU())
29print("Leaky ReLU")
30experiment(nn.LeakyReLU(0.01))
1ReLU
2  dead after the push: 0.81
3  after 1000 steps: dead 0.86  loss 0.057
4Leaky ReLU
5  dead after the push: 0.81
6  after 1000 steps: dead 0.78  loss 0.009

With ReLU, 81% of the hidden neurons are dead after the push, and 1,000 training steps do not bring them back: the share goes up to 86%, and the network fits the data with the few neurons left, ending at a loss of 0.057. The Leaky ReLU network starts from the same state, and many of its neurons still have negative pre-activations at the end, but those neurons keep passing a small gradient and keep contributing, so it reaches a loss of 0.009, about 6 times lower.

How to prevent and fix dead neurons

Leaky ReLU replaces the flat zero for negative inputs with a small slope $\alpha$, commonly 0.01:

$$\text{LeakyReLU}(z) = \begin{cases} z & z > 0 \ \alpha z & z \leq 0 \end{cases}$$

Its derivative for negative inputs is $\alpha$ instead of 0, so a neuron with negative pre-activations still receives a small gradient and its weights keep updating. GELU, the activation used in many transformers, also has a nonzero gradient for most negative inputs.

To prevent dead neurons in the first place:

  1. Use Kaiming initialization for ReLU layers (nn.init.kaiming_normal_(layer.weight, nonlinearity="relu")), which scales the starting weights so pre-activations stay at a size ReLU handles well, and start biases at zero.
  2. Keep the learning rate reasonable. If the dead fraction climbs during training, lower the learning rate first.
  3. Switch to Leaky ReLU or GELU if lowering the learning rate does not stop it.

If many neurons are dead at step 1, restart with a better initialization, since training will not revive them.

Common mistakes

Activation functions covers ReLU, Leaky ReLU, and GELU side by side, including where each one's gradient is zero. Weight initialization sets the starting scale of the pre-activations, and Kaiming initialization is built for ReLU. Vanishing and exploding gradients are the network-wide version of the same issue: a dead neuron is a gradient that is exactly zero at one unit, and a vanishing gradient is one that shrinks toward zero across many layers. The learning rate is the most common cause of neurons dying during training.

QuiddityML teaches dead neurons as a concept in Unit 2 of the ML Foundation track, and the exercises include ordering the lines of fraction_zero, spotting the version that sums instead of averaging, and tracing its shapes on a small batch (quiddityml.com).