QuiddityML

4 October 2026 · 7 min read

Hyperparameter tuning: grid search vs random search

Picking the learning rate and layer sizes by trial and error wastes training runs. This post compares grid search, random search, and successive halving, with runnable PyTorch code and advice on which to use.

Hyperparameter tuning is the process of choosing the settings a model trains with, such as the learning rate or the size of a layer, by training it several times with different values and keeping the version that scores best on held-out data. Training does not learn these settings, and a poor choice can leave a good model stuck at a much worse result. The methods below are different ways to spend a fixed number of training runs so that a good setting is among the ones you tried.

What is being searched?

Parameters are the weights and biases that training adjusts. Hyperparameters are values you set before training starts: the learning rate (how big a step each weight update takes), the batch size, the weight decay, the dropout rate, the hidden layer width. The parameters vs hyperparameters post covers the split in more detail.

No formula gives the right hyperparameters for a new dataset. The usual approach is to train with one combination of values, called a configuration or a trial, measure the loss on a validation set, repeat for other configurations, and keep the winner. The validation set is a slice of data the model does not train on, so its score shows how the model does on unseen examples. The train, validation, and test split post explains how to set one up.

Each trial is a full training run, so the budget is usually counted in trials.

Grid search picks a short list of candidate values for each hyperparameter and trains every combination. Three learning rates and three hidden widths give $3 \times 3 = 9$ runs.

The cost multiplies with each hyperparameter you add. A third hyperparameter with three values makes it 27 runs, and a fourth makes it 81. Grid search also wastes runs. In many problems one or two hyperparameters move the validation score a lot and the rest barely change it. If the learning rate is the one that counts, a 3 by 3 grid trains nine models but tries three learning rates, because each one is repeated once per width.

Random search draws each configuration at random from a range for each hyperparameter. With the same nine runs, it tries nine different learning rates and nine widths, since no value has to be repeated to fill out a grid.

That is why random search tends to find a better configuration than grid search for the same compute. It does not need to know in advance which hyperparameter is important, because each run tests a new value of each one.

In the picture below, the horizontal axis is a hyperparameter that changes the validation score and the vertical axis is one that does not. The curve along the top is how good the score is at each value of the important one, with the best values in the middle. Both panels spend six runs. The grid lands on three distinct values of the important hyperparameter, and random search lands on six.

Two panels with six dots each: grid search places its dots in 3 columns of 2, testing 3 values of the important parameter, while random search places 6 dots at 6 different horizontal positions under a curve showing which values of the important parameter score best

Sample the learning rate on a log scale

For the learning rate, the range usually spans several powers of ten, and what changes training is the order of magnitude: going from 0.0001 to 0.001 has about as much effect as going from 0.001 to 0.01. If you draw uniformly between $0.0001$ and $0.1$, roughly 90% of the draws land above $0.01$ and the small values barely get tested.

The fix is to draw an exponent $u$ uniformly between $-3$ and $0$ and use

$$\text{lr} = 10^{u}$$

which gives each power of ten the same share of the trials. Dropout rates and the number of layers sit in small ranges where a uniform draw works.

Grid vs random search in PyTorch

The snippet below trains a small network on a made-up classification task. Grid search tries 3 learning rates and 3 widths. Random search draws 9 configurations, with the learning rate on a log scale between 0.001 and 1. Both get 9 runs of 300 training steps, and each run starts from the same random weights so the hyperparameters are the one thing that differs.

1import itertools
2import torch
3import torch.nn as nn
4 
5torch.manual_seed(0)
6X = torch.randn(1200, 10)
7y = (X[:, 0] * X[:, 1] + torch.sin(2 * X[:, 2]) > 0).long()   # a nonlinear rule to learn
8X_train, y_train = X[:800], y[:800]
9X_val, y_val = X[800:], y[800:]
10 
11def train_and_evaluate(lr, width, steps=300):
12    torch.manual_seed(1)                                  # same init for every trial
13    model = nn.Sequential(nn.Linear(10, width), nn.ReLU(), nn.Linear(width, 2))
14    opt = torch.optim.SGD(model.parameters(), lr=lr)
15    loss_fn = nn.CrossEntropyLoss()
16    for _ in range(steps):
17        opt.zero_grad()
18        loss_fn(model(X_train), y_train).backward()
19        opt.step()
20    with torch.no_grad():
21        return loss_fn(model(X_val), y_val).item()
22 
23def search(configs, name):
24    results = [(train_and_evaluate(lr, w), lr, w) for lr, w in configs]
25    best = min(results)
26    n_lrs = len({lr for _, lr, _ in results})
27    print(f"{name}: {len(results)} runs, {n_lrs} distinct learning rates, "
28          f"best val loss {best[0]:.3f} at lr={best[1]:.3g}, width={best[2]}")
29 
30# grid search: 3 learning rates x 3 widths = 9 runs
31grid = list(itertools.product([0.001, 0.01, 0.1], [16, 64, 256]))
32search(grid, "grid")
33 
34# random search: 9 runs, learning rate drawn on a log scale between 1e-3 and 1
35gen = torch.Generator().manual_seed(0)
36widths = [16, 32, 64, 128, 256]
37rand = []
38for _ in range(9):
39    lr = 10 ** (torch.rand(1, generator=gen).item() * 3 - 3)
40    width = widths[torch.randint(len(widths), (1,), generator=gen).item()]
41    rand.append((lr, width))
42search(rand, "random")
1grid: 9 runs, 3 distinct learning rates, best val loss 0.541 at lr=0.1, width=64
2random: 9 runs, 9 distinct learning rates, best val loss 0.529 at lr=0.124, width=128

Random search tried 9 learning rates against the grid's 3 and found a lower validation loss, 0.529 against 0.541. The grid's winner also sits on the largest learning rate in its list, which suggests the list stopped too early. When the best value of any search lands on the edge of its range, widen the range and search again before trusting the result.

The random sampling uses its own torch.Generator, separate from the seed that sets the starting weights. Resetting one global seed at the start of each trial would also reset the sampler and draw the same configuration over and over.

Successive halving

Grid and random search train each configuration to the end, even ones that look bad after a few steps. Successive halving starts many configurations on a small budget, keeps the best fraction, and gives the survivors a larger budget, repeating until one is left. The fraction kept does not have to be a half, and the version below keeps the top third.

1# successive halving: start all 9 random configs on a small budget, keep the best third
2survivors = rand
3for steps in (30, 100, 300):
4    scored = sorted((train_and_evaluate(lr, w, steps), lr, w) for lr, w in survivors)
5    print(f"{steps:>3} steps, configs trained: {len(scored)}, best val loss {scored[0][0]:.3f}")
6    survivors = [(lr, w) for _, lr, w in scored[: max(1, len(scored) // 3)]]
7print(f"winner: lr={survivors[0][0]:.3g}, width={survivors[0][1]}")
1 30 steps, configs trained: 9, best val loss 0.646
2100 steps, configs trained: 3, best val loss 0.594
3300 steps, configs trained: 1, best val loss 0.529
4winner: lr=0.124, width=128

It picked the same configuration as the full random search for 870 training steps ($9 \times 30 + 3 \times 100 + 1 \times 300$) instead of 2,700. The risk is that a configuration which starts slowly and finishes well, such as one with a small learning rate, gets cut in the first round. Hyperband reduces that risk by running successive halving several times with different starting budgets.

Which method to use, and common mistakes

The mistakes people make:

A learning rate sweep is a search over one hyperparameter, and the learning rate post shows how to run one by multiplying by 10 each trial. Early stopping ends one training run once validation loss stops improving, which applies the stop-early idea behind successive halving to a single configuration. Cross-validation changes how each trial is scored, averaging over several validation splits, and works with any of the three search methods.

QuiddityML teaches hyperparameter search as its own concept in the ML Foundation track, and the exercises on it include multiple choice on why grid search falls behind random search, filling in and ordering the lines of a random search loop with a log-scaled learning rate, and then writing that loop from scratch.