1 October 2026 · 8 min read
How to make PyTorch training reproducible: seeds, determinism, and what still varies
Reproducible training means running the same code on the same data and getting the same result. This post explains where the randomness in a PyTorch run comes from, how to seed each source, how to seed a DataLoader, and which differences remain after seeding.
Training is reproducible when running the same code on the same data gives the same result each time. By default a PyTorch run is not: the starting weights, the order of the training examples and several other steps use random numbers, so two runs of one script end at different scores. If a change to the model moves the validation accuracy from 0.73 to 0.75 and a plain rerun can do the same, there is no way to tell whether the change helped.
Where the randomness comes from
Random numbers on a computer come from a random number generator, a function that produces a long sequence of numbers that look random. A seed is the integer that picks the starting point of that sequence. The same seed gives the same sequence, so every "random" choice made from it comes out the same.
A training script draws random numbers in several places.
- Weight initialization. Each layer starts from random weights.
- Shuffling. The training examples are usually put in a new random order each epoch, where an epoch is one pass over the training data.
- Dropout. A dropout layer sets a random subset of values to zero at each training step.
- Data augmentation. Random crops, flips or noise applied to each example.
These draws do not come from one generator. Python's random module, NumPy and PyTorch each keep their own, and PyTorch keeps a separate one for each GPU. Seeding one of them leaves the others untouched, so torch.manual_seed(42) alone does not fix an augmentation that calls np.random.
A set_seed function
The usual fix is one function that seeds each generator, called once at the top of the script.
1import random
2import numpy as np
3import torch
4import torch.nn as nn
5
6def set_seed(seed: int):
7 random.seed(seed)
8 np.random.seed(seed)
9 torch.manual_seed(seed)
10 torch.cuda.manual_seed_all(seed) # every GPU; does nothing on a CPU-only machine
11 torch.backends.cudnn.deterministic = True
12 torch.backends.cudnn.benchmark = False
13 torch.use_deterministic_algorithms(True)
14
15data_rng = torch.Generator().manual_seed(0) # the dataset itself is fixed
16X = torch.randn(400, 10, generator=data_rng)
17y = (X[:, 0] * X[:, 1] + 0.5 * torch.randn(400, generator=data_rng) > 0).long()
18X_train, y_train, X_val, y_val = X[:300], y[:300], X[300:], y[300:]
19
20def train(seed=None):
21 if seed is not None:
22 set_seed(seed) # before the model is built
23 model = nn.Sequential(nn.Linear(10, 32), nn.ReLU(), nn.Dropout(0.2), nn.Linear(32, 2))
24 optimizer = torch.optim.Adam(model.parameters(), lr=0.01)
25 for epoch in range(30):
26 model.train()
27 order = torch.randperm(300) # shuffle
28 for i in range(0, 300, 30):
29 idx = order[i:i + 30]
30 loss = nn.functional.cross_entropy(model(X_train[idx]), y_train[idx])
31 optimizer.zero_grad()
32 loss.backward()
33 optimizer.step()
34 model.eval()
35 with torch.no_grad():
36 return (model(X_val).argmax(dim=1) == y_val).float().mean().item()
37
38print("no seed: ", f"{train():.2f}", f"{train():.2f}", f"{train():.2f}")
39print("seed 42: ", f"{train(42):.2f}", f"{train(42):.2f}", f"{train(42):.2f}")
40scores = [train(seed) for seed in range(10)]
41print("seeds 0-9:", " ".join(f"{s:.2f}" for s in scores))
42print(f"mean {np.mean(scores):.3f}, std {np.std(scores):.3f}, min {min(scores):.2f}, max {max(scores):.2f}")1no seed: 0.76 0.74 0.72
2seed 42: 0.73 0.73 0.73
3seeds 0-9: 0.74 0.74 0.70 0.73 0.71 0.78 0.75 0.68 0.72 0.73
4mean 0.728, std 0.026, min 0.68, max 0.78Without a seed, three runs of the same code gave validation accuracies of 0.76, 0.74 and 0.72 here, and that line prints different numbers each time the script runs. With set_seed(42) called before each run, all three give 0.73. The call sits before the model is built because building the model is the first thing that draws random numbers. A seed set after that point leaves the starting weights different from run to run.
When two seeded runs still disagree, move the set_seed call to the first line of the script before changing anything else.
What the three determinism lines do
The first four lines of set_seed seed generators. The last three deal with a different source of variation: some GPU operations can return slightly different results on identical inputs, with no random numbers involved.
torch.backends.cudnn.deterministic = True: cuDNN is the library PyTorch uses for convolutions on NVIDIA GPUs. It has several algorithms for the same convolution, and some of them add numbers up in an order that changes between runs. This setting asks it for the ones with a fixed order.torch.backends.cudnn.benchmark = False: with benchmark mode on, cuDNN times its algorithms at the start of a run and keeps the fastest. Timing varies, so two runs can end up with different algorithms. Turning it off removes that choice.torch.use_deterministic_algorithms(True): the broader switch. It applies to all of PyTorch, and when an operation has no deterministic version it raises an error that names the operation. That error tells you which line is causing the drift.
The reason order changes a result is that floating-point addition rounds after each step. Adding the same numbers in a different order can change the last digit, and thousands of training steps grow that difference into a visibly different loss.
On a CUDA GPU, use_deterministic_algorithms(True) also needs the environment variable CUBLAS_WORKSPACE_CONFIG=:4096:8 set before the script starts, and PyTorch raises an error that says so if it is missing. Deterministic algorithms are often slower, so a common setup turns them on for experiments that will be compared and leaves them off for a long final training run.
Seeding a DataLoader
A DataLoader is the PyTorch object that shuffles a dataset and hands out batches. It has two sources of randomness of its own.
The first is the shuffle. By default the shuffled order is drawn from PyTorch's global generator, the same one that initializes layers. Seeding it does make the order repeatable, but the order then depends on how many random numbers were drawn before the loader was used. The script below seeds with 0 each time and prints the first batch of a 10-item dataset, with and without one extra layer created first.
1import torch
2import torch.nn as nn
3from torch.utils.data import DataLoader, TensorDataset
4
5dataset = TensorDataset(torch.arange(10))
6
7def first_batch(extra_layer, own_generator):
8 torch.manual_seed(0)
9 if extra_layer:
10 nn.Linear(4, 4) # one more layer draws from the global generator
11 g = None
12 if own_generator:
13 g = torch.Generator()
14 g.manual_seed(42)
15 loader = DataLoader(dataset, batch_size=5, shuffle=True, generator=g)
16 return next(iter(loader))[0].tolist()
17
18print("global generator: ", first_batch(False, False))
19print("global, one extra layer: ", first_batch(True, False))
20print("own generator: ", first_batch(False, True))
21print("own, one extra layer: ", first_batch(True, True))1global generator: [6, 7, 1, 4, 2]
2global, one extra layer: [2, 6, 7, 9, 0]
3own generator: [6, 5, 4, 0, 8]
4own, one extra layer: [6, 5, 4, 0, 8]With the global generator, adding one layer to the model changes which examples land in the first batch. That makes a comparison between two architectures unfair, since they also see the data in different orders. Passing a dedicated torch.Generator through the generator argument gives the shuffle its own sequence, and the batch order stays [6, 5, 4, 0, 8] whatever the model does.
The second source is the worker processes. With num_workers above 0, the loader starts that many separate processes to load data in parallel, and each has its own random state. The usual recipe is a worker_init_fn, a function the loader runs once inside each worker, that seeds NumPy and Python's random there.
1import random
2import numpy as np
3import torch
4from torch.utils.data import DataLoader, Dataset
5
6class NoisyDataset(Dataset):
7 """Returns the row index plus NumPy noise, the way a random augmentation would."""
8 def __len__(self):
9 return 8
10 def __getitem__(self, i):
11 return i + float(np.random.rand())
12
13def seed_worker(worker_id):
14 worker_seed = torch.initial_seed() % 2**32
15 np.random.seed(worker_seed)
16 random.seed(worker_seed)
17
18def one_epoch(seeded):
19 g = torch.Generator()
20 g.manual_seed(42)
21 loader = DataLoader(
22 NoisyDataset(),
23 batch_size=4,
24 shuffle=True,
25 num_workers=4,
26 worker_init_fn=seed_worker if seeded else None,
27 generator=g if seeded else None,
28 )
29 return [[round(v, 3) for v in batch.tolist()] for batch in loader]
30
31if __name__ == "__main__":
32 print("seeded run 1: ", one_epoch(True))
33 print("seeded run 2: ", one_epoch(True))
34 print("unseeded run 1:", one_epoch(False))
35 print("unseeded run 2:", one_epoch(False))1seeded run 1: [[4.063, 6.675, 2.705, 3.124], [0.629, 7.306, 1.212, 5.205]]
2seeded run 2: [[4.063, 6.675, 2.705, 3.124], [0.629, 7.306, 1.212, 5.205]]
3unseeded run 1: [[4.389, 6.032, 2.415, 3.446], [1.827, 0.503, 5.224, 7.145]]
4unseeded run 2: [[4.102, 0.335, 5.857, 2.861], [7.406, 1.154, 6.423, 3.547]]The whole part of each number is the row index and the decimals are the NumPy noise. The two seeded runs agree on both. The two unseeded runs differ in the order of the rows and in the noise added to each one, and they print new values on each run of the script. Inside a worker, torch.initial_seed() returns a seed PyTorch derived from the loader's generator and the worker's number, which is why the generator and the worker_init_fn are passed together.

The picture below collects every source from this post in one place, with the call that controls each.

What still varies after seeding
Seeding makes a run repeatable on one machine with one software setup. It does not make results match across setups.
- Different hardware. A CPU and a GPU, or two different GPU models, can round the same arithmetic differently.
- Different library versions. A new PyTorch or CUDA version can change an algorithm or the sequence a generator produces.
- A different number of workers or GPUs. Each changes which process draws which random numbers.
- Operations with no deterministic version. These either raise an error under
use_deterministic_algorithms(True)or keep varying without it.
Recording the seed, the library versions and the hardware next to each result is what lets someone else, or you in six months, get close to the same number.
A seed does not make a result reliable
A fixed seed makes one run repeatable. It says nothing about how much the score would move with a different seed. In the first script, ten seeds gave accuracies from 0.68 to 0.78 on a 100-example validation set, with a mean of 0.728 and a standard deviation of 0.026. A change that lifts one seed from 0.73 to 0.75 is well inside that spread.
When two models are being compared, train each with several seeds, commonly three to five, and compare the means along with the spread. Picking the seed with the highest score and reporting that one selects for luck, the same way picking the best of many models on a validation set does.
Common mistakes
- Seeding only PyTorch. NumPy and Python's
randomkeep their own generators, and augmentation code often uses them. - Calling
set_seedafter the model is built. The starting weights were already drawn. - Leaving the DataLoader on the global generator. The batch order then changes whenever the model's layers change.
- Expecting the same numbers on a different GPU or PyTorch version. Seeds fix the random draws and leave the arithmetic differences in place.
- Reporting one seed. One number hides the run-to-run spread, which on small validation sets can be several points.
- Leaving deterministic mode off when debugging. A bug that shows up on some runs and not others is easier to find once each run is the same.
Related concepts
A train, validation and test split also takes a seed, so that each run is scored on the same held-out examples. Weight initialization is the first place a run draws random numbers. Mini-batch gradient descent explains why the order and makeup of batches change the path training takes. Saving and resuming training is the other half of a repeatable run, since a checkpoint lets it continue from the same state.
QuiddityML teaches reproducibility as a concept in Unit 3 of the ML Foundation track, where the exercises include filling in and then writing set_seed from scratch, tracing what each seed call controls, and predicting what a DataLoader does without a generator (quiddityml.com).