28 September 2026 · 6 min read
Parameters vs hyperparameters in machine learning
Parameters are the numbers a model learns from data, and hyperparameters are the settings you choose before training starts. This post explains the difference with a small PyTorch network, shows where each one lives in code, and how to pick hyperparameters with a validation set.
A parameter is a number inside a model that training learns from data, like a weight or a bias. A hyperparameter is a setting you choose before training starts, like the learning rate or the number of layers, and training does not change it. The difference decides what you can leave to the optimizer and what you have to test yourself, and a badly chosen hyperparameter can make training fail even though the parameters are still free to learn.
What counts as a parameter
A neural network layer takes a list of input numbers and produces a list of output numbers. Each output is a weighted sum of the inputs plus one extra number. The multipliers are the weights and the extra numbers are the biases, and together they are the layer's parameters.
Training finds good values for them. The usual method is gradient descent: the model makes predictions, a loss (one number that measures how wrong the predictions are) is computed, and every weight and bias is nudged a small step in the direction that lowers that loss. One nudge is one training step, and a run repeats it thousands of times.
The number of parameters in a layer follows from its size. If a layer has $n_{in}$ inputs and $n_{out}$ outputs, it has one weight for every input-output pair and one bias per output:
$$\text{params} = n_{in} \times n_{out} + n_{out}$$
Take a network that maps 4 input features to one output through two hidden layers of 8 units each, written $4 \to 8 \to 8 \to 1$. The first layer has $4 \times 8 + 8 = 40$ parameters, the second $8 \times 8 + 8 = 72$, and the output layer $8 \times 1 + 1 = 9$, for 121 in total. Gradient descent sets all 121 of those values.
What counts as a hyperparameter
Look at what that count left out: the choice of two hidden layers, the width of 8 for each one, the learning rate, and the batch size. None of those are updated by gradient descent. They are fixed before the first training step, and they are the hyperparameters.
- Number of hidden layers and width of each layer: the shape of the network. These decide how many parameters exist in the first place, so the 121 above is a result of choosing $4 \to 8 \to 8 \to 1$.
- Learning rate: the size of each gradient descent step. Too small and the loss barely moves, too large and it jumps around or grows.
- Batch size: how many training examples are used to compute each step.
- Number of epochs: how many full passes over the training data the run makes.
Parameters are learned from data. Hyperparameters are chosen by the person training the model, usually by running a few experiments and comparing the results.

| Kind | Examples | Set by | When | Chosen with |
|---|---|---|---|---|
| Parameters | weights, biases | gradient descent | every training step | the training set |
| Hyperparameters | learning rate, layer width | the person training | before training starts | the validation set |
Parameters and hyperparameters in PyTorch
In PyTorch, parameters live inside the model. model.parameters() returns every weight and bias tensor, each one has requires_grad=True so PyTorch computes its gradient, and numel() gives how many numbers it holds. Hyperparameters are plain Python variables that you pass in when building the model and the optimizer.
The snippet below builds the $4 \to 8 \to 8 \to 1$ network, prints its parameters, and then runs a small search over one hyperparameter, the learning rate, on a toy regression dataset.
1import torch
2import torch.nn as nn
3
4torch.manual_seed(0)
5
6# Hyperparameters: chosen by hand before training starts
7hidden_width = 8
8batch_size = 32
9epochs = 20
10
11def make_model(width):
12 return nn.Sequential(
13 nn.Linear(4, width), nn.ReLU(),
14 nn.Linear(width, width), nn.ReLU(),
15 nn.Linear(width, 1),
16 )
17
18model = make_model(hidden_width)
19for name, p in model.named_parameters():
20 print(name, tuple(p.shape), p.numel(), p.requires_grad)
21print("total parameters:", sum(p.numel() for p in model.parameters()))
22
23# Toy regression data: 4 features, 1 target, split into train / validation / test
24X = torch.randn(1000, 4)
25y = (X[:, :1] * X[:, 1:2] + torch.sin(X[:, 2:3]) - X[:, 3:4]) + 0.1 * torch.randn(1000, 1)
26X_train, y_train = X[:600], y[:600]
27X_val, y_val = X[600:800], y[600:800]
28X_test, y_test = X[800:], y[800:]
29
30def train(lr):
31 torch.manual_seed(1)
32 model = make_model(hidden_width)
33 opt = torch.optim.SGD(model.parameters(), lr=lr)
34 loss_fn = nn.MSELoss()
35 for _ in range(epochs):
36 perm = torch.randperm(len(X_train))
37 for i in range(0, len(X_train), batch_size):
38 idx = perm[i:i + batch_size]
39 loss = loss_fn(model(X_train[idx]), y_train[idx])
40 opt.zero_grad()
41 loss.backward()
42 opt.step() # updates the parameters, not lr or batch_size
43 return model
44
45results = {}
46for lr in [0.001, 0.01, 0.1, 1.0]:
47 m = train(lr)
48 with torch.no_grad():
49 results[lr] = nn.functional.mse_loss(m(X_val), y_val).item()
50 print(f"lr={lr}: validation MSE {results[lr]:.3f}")
51
52best_lr = min(results, key=results.get)
53final = train(best_lr)
54with torch.no_grad():
55 print("best lr:", best_lr, "test MSE:", round(nn.functional.mse_loss(final(X_test), y_test).item(), 3))
56
57# 0.weight (8, 4) 32 True
58# 0.bias (8,) 8 True
59# 2.weight (8, 8) 64 True
60# 2.bias (8,) 8 True
61# 4.weight (1, 8) 8 True
62# 4.bias (1,) 1 True
63# total parameters: 121
64# lr=0.001: validation MSE 1.936
65# lr=0.01: validation MSE 0.538
66# lr=0.1: validation MSE 0.221
67# lr=1.0: validation MSE 5.157
68# best lr: 0.1 test MSE: 0.171The printout matches the hand count: 32 + 8 weights and biases in the first layer, 64 + 8 in the second, 8 + 1 in the output layer, 121 in total. opt.step() changes those 121 numbers and nothing else, and lr, batch_size, epochs and hidden_width stay exactly as they were written. MSE is mean squared error, the average of the squared gaps between predictions and targets, so lower is better.
The sweep trains the same network four times, changing only the learning rate. With 0.001 the validation error is still 1.936 after 20 epochs, 0.1 gives the lowest error at 0.221, and 1.0 ends at 5.157, worse than any other setting. If every value in a sweep looks bad, widen the range by a factor of 10 in each direction before changing anything else.
Why hyperparameters are chosen on a validation set
A validation set is a slice of data held out from training and used to compare hyperparameter settings. A test set is a second held-out slice, used once at the end to report how well the final model does on data that played no part in any choice.
The two slices do different jobs. Picking the learning rate that scores best on some data is a form of fitting to that data: out of four runs, you keep the one that happened to do best there. If you pick it on the test set, the test score is no longer a fair estimate of performance on new data, because the test set helped make the choice. In the snippet, the test set is touched once, after best_lr is fixed, and that last line is the number you would report.
Picking on the training set does not work either. Training loss measures how well the parameters fit the examples they were trained on, and settings that let a model memorize those examples, like a very wide network, can score well there and poorly on new data.
Common mistakes
- Tuning on the test set. Running a sweep, checking the test score after each run, and keeping the top scorer turns the test set into a second validation set. Keep a separate validation split for choices and look at the test set at the end.
- Changing several hyperparameters at once. If the learning rate and the width both change between two runs, a better score cannot be credited to either one. Sweep one at a time, or use a proper search over combinations.
- Calling the architecture a parameter. The number of layers and their widths are chosen, not learned. They decide how many parameters exist, and gradient descent then fills in their values.
- Treating one sweep as final. A learning rate that works well at one batch size or width can do worse at another. After changing the batch size or the width, run the learning rate sweep again.
Related concepts
The learning rate is usually the first hyperparameter to tune, because a bad value can stop training from making progress at all. Gradient descent is the procedure that sets the parameters, and the learning rate is the size of its step. Overfitting and underfitting are what the validation score catches: a model that does well on training data and badly on the validation set often has too many parameters for the data or trained for too many epochs, and width and epoch count are both hyperparameters.
QuiddityML teaches parameters vs. hyperparameters as a concept in Unit 2 of the ML Foundation track, where the multiple-choice questions ask which values are learned and which are chosen, and which hand-picked choice shaped a 121-parameter network (quiddityml.com).