QuiddityML

29 September 2026 · 8 min read

Weight initialization explained: Xavier vs He (Kaiming) vs orthogonal

Weight initialization is how a neural network's weights are set before training starts, and a bad choice can stop a deep network from learning at all. This post shows what goes wrong with zero, too large, and too small weights, and when to use Xavier, He (Kaiming), or orthogonal initialization in PyTorch.

Weight initialization is the choice of starting values for a neural network's weights before the first training step. Training only nudges the weights a little at a time, so where they start decides whether the network can learn at all: with a bad start, a deep network can produce outputs that are all the same, all enormous, or all close to zero, and the training signal barely reaches the early layers.

What a weight is, and why its starting value counts

A layer in a neural network takes a list of input numbers and produces a list of outputs. Each output is a weighted sum of the inputs plus a bias, passed through an activation function (a fixed non-linear function such as tanh or ReLU, where ReLU keeps positive values and sets negative ones to 0). The multipliers in those sums are the weights.

Training adjusts the weights with gradient descent: compute a loss (one number that measures how wrong the predictions are), compute the gradient (how much the loss changes when each weight changes), and move every weight a small step in the direction that lowers the loss. The gradient is computed layer by layer from the output back to the input, a procedure called backpropagation.

Both passes multiply through every layer. The forward pass multiplies the input by each layer's weights, and backpropagation multiplies the gradient by each layer's weights on the way back. In a network with 10 or 50 layers, a small distortion per layer gets multiplied 10 or 50 times, which is why the starting scale of the weights has so much effect.

Three panels showing how initialization fails: symmetry, where all weights are equal and all neurons are identical, too large, where activations explode, saturate, or vanish, and too small, where the signal shrinks toward zero through the layers.

The symmetry problem: why weights cannot all start at zero

If every weight in a layer starts at the same value, every neuron in that layer computes the same output. They then receive the same gradient, get the same update, and are still identical after the step. This repeats at every step, so a hidden layer of 512 neurons behaves like one neuron copied 512 times.

The code below starts every weight and bias of a small network at 0.5 and trains it for 100 steps.

1import torch
2import torch.nn as nn
3 
4torch.manual_seed(0)
5model = nn.Sequential(nn.Linear(3, 4), nn.Tanh(), nn.Linear(4, 1))
6for p in model.parameters():
7    nn.init.constant_(p, 0.5)  # every weight and bias starts at 0.5
8 
9x = torch.randn(8, 3)
10y = torch.randn(8, 1)
11opt = torch.optim.SGD(model.parameters(), lr=0.1)
12for _ in range(100):
13    loss = ((model(x) - y) ** 2).mean()
14    opt.zero_grad()
15    loss.backward()
16    opt.step()
17 
18print(model[0].weight)

The four rows of the first layer's weight matrix, one per hidden neuron, are still identical after training:

1tensor([[0.1739, 0.7184, 0.3469],
2        [0.1739, 0.7184, 0.3469],
3        [0.1739, 0.7184, 0.3469],
4        [0.1739, 0.7184, 0.3469]], requires_grad=True)

The weights moved, and they moved together. Starting the weights at random values fixes this, because each neuron then computes something different from the first step and gets its own gradient.

Top: two inputs of 1.00 connect to h1 and h2 with every weight equal to 0.50, so both outputs are 1.00 and they get the same gradient and update. Bottom: weights of 0.20, -0.10, 0.40, and 0.40 give outputs 0.50 and 0.30, so the neurons learn different things.

Weights too large or too small

Random weights fix symmetry, but the scale of the random values still decides what happens to the signal.

If the weights are too large, each layer multiplies the size of its input up. For tanh or sigmoid, the weighted sums land far from zero where the function is flat, so the outputs sit at the ends of the range and the gradient through them is close to zero. This is called saturation. For ReLU, which has no upper limit, the outputs grow layer after layer until they overflow or training diverges.

Weights too large: the signal grows from 1 to 8 to 64 to 512 across layers 1 to 4 and explodes before reaching the output.

If the weights are too small, each layer multiplies the size of its input down. After a few layers the outputs are close to zero, and the gradients flowing back are close to zero too, so the early layers barely update.

Weights too small: the signal shrinks from 1 to 0.1 to 0.01 to 0.001 across the layers and vanishes before reaching the output.

The goal: keep the variance the same at every layer

The variance of a set of numbers is the average squared distance from their mean, a measure of how spread out they are. Initialization schemes pick the weight variance so that the variance of each layer's outputs stays close to the variance of its inputs.

Write $h$ for one output of a layer, $x$ for its inputs, $w$ for its weights, and $n_{in}$ for the number of inputs (the layer's fan-in). When the weights and inputs are independent with mean zero, and there is no activation yet:

$$\text{Var}(h) = n_{in} \cdot \text{Var}(w) \cdot \text{Var}(x)$$

To keep $\text{Var}(h)$ equal to $\text{Var}(x)$, the weight variance has to be $\text{Var}(w) = 1/n_{in}$. A wider layer sums more inputs, so each weight has to be smaller to keep the total the same size.

Activation variance across 10 layers on a log scale: with weights too large it climbs to 100, with Var(w) = 1/n_in it stays at 1, and with weights too small it drops to 0.01, next to the equations Var(h) = n_in * Var(w) * Var(x) and Var(w) = 1/n_in.

Xavier and Kaiming initialization both start from this equation and adjust it for the activation function.

Xavier (Glorot) initialization for tanh and sigmoid

Xavier initialization, also called Glorot initialization, balances the forward pass and backpropagation. Keeping the forward outputs stable needs $1/n_{in}$, keeping the gradients stable on the way back needs $1/n_{out}$ (where $n_{out}$ is the number of outputs, the fan-out), and Xavier uses a compromise between the two:

$$\text{Var}(w) = \frac{2}{n_{in} + n_{out}}$$

The derivation assumes the activation function is roughly linear and symmetric around zero, which holds for tanh and sigmoid near zero. PyTorch has a normal version, nn.init.xavier_normal_, and a uniform version, nn.init.xavier_uniform_, which draws from a range with the same variance.

Xavier initialization: Var(w) = 2 / (n_in + n_out) for a layer with 4 inputs and 3 outputs, the uniform version of the formula, and a plot where activation variance stays at 1.0 across layers while too-large variance explodes and too-small variance dies.

Kaiming (He) initialization for ReLU

ReLU sets every negative input to zero. When the weighted sums are centered on zero, about half of them are negative, so ReLU zeroes about half the values and the variance after the activation is about half of what went in. Xavier does not account for that, so a ReLU network initialized with Xavier loses signal at every layer.

Kaiming initialization, also called He initialization, doubles the weight variance to make up for the half that ReLU removes:

$$w \sim \mathcal{N}\left(0, \frac{2}{n_{in}}\right)$$

This reads as: draw each weight from a normal distribution with mean 0 and variance $2/n_{in}$. The standard deviation is $\sqrt{2/n_{in}}$, and that is the number PyTorch's normal_ expects, not the variance.

Kaiming initialization: a bell curve of pre-activations split at zero, with ReLU setting the left half to 0, and the formula w ~ Normal(0, 2/n_in), σ = sqrt(2/n_in), and nn.init.kaiming_normal_(layer.weight, nonlinearity='relu').

What each scheme does to a 10-layer network

The experiment below passes 1,000 random inputs through 10 layers of width 512 and reports the standard deviation of the last layer's output (the square root of the variance, in the same units as the values). For tanh it also reports the share of outputs above 0.99 in size, which is how saturated the layer is.

1import torch
2import torch.nn as nn
3 
4torch.manual_seed(0)
5 
6def run(init, act, depth=10, width=512):
7    """Push one batch through `depth` layers, return the last layer's output."""
8    h = torch.randn(1000, width)
9    for _ in range(depth):
10        layer = nn.Linear(width, width, bias=False)
11        init(layer.weight)
12        h = act(layer(h))
13    return h
14 
15inits = {
16    "std 1.0": lambda w: nn.init.normal_(w, std=1.0),
17    "std 0.01": lambda w: nn.init.normal_(w, std=0.01),
18    "Xavier": nn.init.xavier_normal_,
19    "Kaiming": lambda w: nn.init.kaiming_normal_(w, nonlinearity="relu"),
20}
21 
22for name, init in inits.items():
23    h = run(init, torch.tanh)
24    saturated = (h.abs() > 0.99).float().mean().item()
25    print(f"tanh {name:9s} std {h.std().item():.3g}  saturated {saturated:.0%}")
26for name, init in inits.items():
27    h = run(init, torch.relu)
28    print(f"relu {name:9s} std {h.std().item():.3g}")
1tanh std 1.0   std 0.982  saturated 91%
2tanh std 0.01  std 3.31e-07  saturated 0%
3tanh Xavier    std 0.228  saturated 0%
4tanh Kaiming   std 0.552  saturated 0%
5relu std 1.0   std 7.71e+11
6relu std 0.01  std 9.39e-09
7relu Xavier    std 0.0259
8relu Kaiming   std 0.828

With a standard deviation of 1.0, 91% of the tanh outputs are saturated and the ReLU outputs reach about $10^{11}$. With 0.01, both shrink to almost nothing. For ReLU, Kaiming keeps the output at a similar scale to the input after 10 layers, while Xavier shrinks it by about 40 times. For tanh, Xavier avoids saturation, though the signal still shrinks slowly because tanh pulls values toward zero at each layer.

Orthogonal initialization

An orthogonal matrix is a square matrix $W$ with $W^\top W = I$, where $I$ is the identity matrix. Multiplying a vector by an orthogonal matrix rotates or reflects it without changing its length, so a layer with orthogonal weights neither grows nor shrinks its input.

This is most useful in recurrent neural networks (RNNs), which apply the same weight matrix once per step of a sequence, so a small stretch per step compounds over hundreds of steps. PyTorch builds a random orthogonal matrix with nn.init.orthogonal_(layer.weight), and for a non-square weight it makes the rows or the columns orthonormal, whichever fits the shape.

Orthogonal weights preserve length: a random W stretches the unit circle into an ellipse, while an orthogonal W with W transpose W = I only rotates a vector, keeping its length.

Weight initialization in PyTorch

nn.Linear and nn.Conv2d already initialize their weights with a uniform distribution scaled by the fan-in, so a small ReLU network often trains fine without any extra code. That default is not exactly the Kaiming formula above, though. It draws from a range of $\pm 1/\sqrt{n_{in}}$, which gives a smaller variance. When the initialization is part of what you are testing, or the network is deep, call the scheme you want after building the model:

1import torch.nn as nn
2 
3def init_weights(module):
4    """He init for every Linear layer, biases at zero."""
5    if isinstance(module, nn.Linear):
6        nn.init.kaiming_normal_(module.weight, nonlinearity="relu")
7        nn.init.zeros_(module.bias)
8 
9model = nn.Sequential(
10    nn.Linear(784, 256), nn.ReLU(),
11    nn.Linear(256, 256), nn.ReLU(),
12    nn.Linear(256, 10),
13)
14model.apply(init_weights)
15print(model[0].weight.std().item())

This prints about 0.05, close to $\sqrt{2/784} \approx 0.0505$. model.apply calls the function on every submodule, so the same line works for a network of any depth. If training fails right at the start, the first thing to check is whether the output scale of each layer stays near the input scale, as in the experiment above.

Initialization in practice: nn.Linear uses a fan-in-aware default, and the cases to override it are tanh or sigmoid hidden layers with Xavier, ReLU with exact He scaling with Kaiming, RNN recurrent weights with orthogonal, and custom parameters initialized explicitly.

When to use which initialization

Network Initialization PyTorch call
tanh or sigmoid hidden layers Xavier xavier_uniform_
ReLU hidden layers Kaiming kaiming_normal_
RNN recurrent weights orthogonal orthogonal_
custom module parameters pick one on purpose any of the above

Common mistakes

Activation functions decide which initialization fits: tanh and sigmoid pair with Xavier, ReLU and its variants with Kaiming. Backpropagation is where a bad initialization shows up, since it multiplies the gradient by every layer's weights on the way back. Vanishing and exploding gradients are the training-time version of the too-small and too-large problems above, and a common fix for them is gradient clipping. Dead ReLU neurons, units that output zero for every input, are more likely when the starting scale is wrong.

QuiddityML teaches weight initialization in Unit 2 of the ML Foundation track, where the Kaiming exercises include matching the formula to its code, spotting the bug in an initializer that uses $1/n_{in}$ where ReLU needs $2/n_{in}$, and writing kaiming_normal_ from scratch (quiddityml.com).