QuiddityML

1 October 2026 · 10 min read

How to read a loss curve: training vs validation loss

A loss curve is a plot of a model's loss over the course of training. This post explains how to plot training and validation loss together, what a healthy curve looks like, and what it means when validation loss goes up while training loss keeps going down.

A loss curve is a plot of a model's loss against training time. The loss is one number that measures how wrong the model's predictions are, and lower is better. Plotting it after each pass over the data shows whether the model is learning, whether it has stopped, and whether it is learning the training examples in a way that does not carry over to new ones. It is usually the first plot to look at when a training run gives a disappointing score.

Why two curves and not one

Training adjusts the model's weights to push the loss down on the training examples, so the training loss falls almost by construction. A model with enough weights can drive it close to zero by storing which training input goes with which label. That says nothing about inputs it has not seen.

The validation loss is the same loss measured on held-out examples that the model does not train on. It stands in for new data. The distance between the two curves, validation loss minus training loss, is called the generalization gap, and the training curve alone cannot show it.

Recording both takes two lists. After each epoch, one pass over the training data, the loop appends the training loss and the validation loss. model.train() and model.eval() switch layers such as dropout between their training and evaluation behaviour, and in the picture below criterion is the loss function.

A panel of PyTorch-like code titled "Read two loss curves". It creates train_losses and val_losses lists, and for each epoch calls model.train(), computes train_loss with train_one_epoch, appends it, calls model.eval(), computes val_loss with evaluate, and appends it. Below, a plot of loss against epoch shows the training loss falling toward 0 while the validation loss levels off near 0.4, with an arrow between them labeled generalization gap. A caption says training loss alone cannot show memorization.

Three runs on the same data

The script below produces three different curve shapes from one dataset. The data has 20 input features and a yes/no label that depends on a weighted sum of the features plus noise, so some labels cannot be predicted from the inputs. There are 200 training examples and 1,000 validation examples. The loss is cross-entropy, which for two classes is about 0.69 for a model that guesses 50/50.

The three runs are a linear model, a much larger network with two hidden layers of 256 units (an MLP), and that same large network with a learning rate 1,000 times smaller. The learning rate is the number that scales each weight update.

1import torch
2import torch.nn as nn
3 
4torch.manual_seed(0)
5X = torch.randn(1200, 20)
6w_true = torch.randn(20)
7y = ((X @ w_true + 2.0 * torch.randn(1200)) > 0).long()   # labels with noise in them
8X_train, y_train, X_val, y_val = X[:200], y[:200], X[200:], y[200:]
9loss_fn = nn.CrossEntropyLoss()
10 
11def fit(model, lr, num_epochs=200):
12    optimizer = torch.optim.Adam(model.parameters(), lr=lr)
13    history = {"train": [], "val": [], "val_acc": [], "val_conf": []}
14    for epoch in range(num_epochs):
15        model.train()
16        loss = loss_fn(model(X_train), y_train)
17        optimizer.zero_grad()
18        loss.backward()
19        optimizer.step()
20 
21        model.eval()
22        with torch.no_grad():
23            val_logits = model(X_val)
24            history["train"].append(loss_fn(model(X_train), y_train).item())
25            history["val"].append(loss_fn(val_logits, y_val).item())
26            history["val_acc"].append((val_logits.argmax(dim=1) == y_val).float().mean().item())
27            history["val_conf"].append(torch.softmax(val_logits, dim=1).max(dim=1).values.mean().item())
28    return history
29 
30def mlp():
31    return nn.Sequential(nn.Linear(20, 256), nn.ReLU(), nn.Linear(256, 256), nn.ReLU(), nn.Linear(256, 2))
32 
33runs = {
34    "linear model, lr 0.01": (nn.Linear(20, 2), 0.01),
35    "large MLP, lr 0.001": (mlp(), 0.001),
36    "large MLP, lr 0.000001": (mlp(), 0.000001),
37}
38curves = {}
39for name, (model, lr) in runs.items():
40    h = curves[name] = fit(model, lr)
41    best = min(range(200), key=lambda e: h["val"][e])
42    print(name)
43    for e in [0, 9, 24, 49, 99, 199]:
44        print(f"  epoch {e + 1:3d}  train {h['train'][e]:.3f}  val {h['val'][e]:.3f}  gap {h['val'][e] - h['train'][e]:+.3f}")
45    print(f"  lowest val loss {h['val'][best]:.3f} at epoch {best + 1}")
1linear model, lr 0.01
2  epoch   1  train 0.708  val 0.706  gap -0.002
3  epoch  10  train 0.527  val 0.579  gap +0.052
4  epoch  25  train 0.404  val 0.489  gap +0.085
5  epoch  50  train 0.328  val 0.423  gap +0.094
6  epoch 100  train 0.277  val 0.367  gap +0.090
7  epoch 200  train 0.247  val 0.340  gap +0.093
8  lowest val loss 0.340 at epoch 200
9large MLP, lr 0.001
10  epoch   1  train 0.656  val 0.668  gap +0.012
11  epoch  10  train 0.366  val 0.440  gap +0.074
12  epoch  25  train 0.109  val 0.392  gap +0.283
13  epoch  50  train 0.006  val 0.670  gap +0.664
14  epoch 100  train 0.001  val 0.875  gap +0.874
15  epoch 200  train 0.000  val 1.003  gap +1.002
16  lowest val loss 0.353 at epoch 19
17large MLP, lr 0.000001
18  epoch   1  train 0.697  val 0.700  gap +0.003
19  epoch  10  train 0.697  val 0.700  gap +0.003
20  epoch  25  train 0.696  val 0.699  gap +0.003
21  epoch  50  train 0.695  val 0.699  gap +0.004
22  epoch 100  train 0.693  val 0.698  gap +0.005
23  epoch 200  train 0.688  val 0.695  gap +0.007
24  lowest val loss 0.695 at epoch 200

The plot comes from the same curves dictionary, so this block runs after the one above.

1import matplotlib.pyplot as plt
2 
3plt.style.use("dark_background")
4fig, axes = plt.subplots(1, 3, figsize=(15, 4.5), sharey=True)
5for ax, (name, h) in zip(axes, curves.items()):
6    ax.plot(range(1, 201), h["train"], color="#4fd1b5", linewidth=2.5, label="training loss")
7    ax.plot(range(1, 201), h["val"], color="#f0524f", linewidth=2.5, label="validation loss")
8    ax.set_title(name)
9    ax.set_xlabel("Epoch")
10    ax.grid(alpha=0.2)
11axes[0].set_ylabel("Loss")
12axes[0].legend()
13fig.tight_layout()
14fig.savefig("loss_curves.png", dpi=150)

Three plots of loss against epoch from the script. Left, the linear model: training and validation loss both fall and flatten, ending near 0.25 and 0.34. Middle, the large MLP with learning rate 0.001: training loss drops to 0 by epoch 50 while validation loss dips to about 0.35 near epoch 19 and then climbs to 1.0. Right, the large MLP with learning rate 0.000001: both losses stay flat near 0.69.

The next three sections read these three panels from left to right.

What a healthy curve looks like

In the left panel both losses fall quickly, slow down and flatten. The validation loss sits a little above the training loss, and the gap stops growing: it is 0.094 at epoch 50 and 0.093 at epoch 200. The lowest validation loss is at the last epoch, so no earlier point in training was better.

That is the pattern to compare other runs against.

The flat stretch at the end is called a plateau. Two common ways to handle it are early stopping, which ends training once the validation loss has stopped improving and keeps the weights from its lowest point, and ReduceLROnPlateau, a PyTorch learning rate schedule that lowers the learning rate when the validation loss stalls.

A plot titled "Healthy training". Training loss and validation loss both fall steeply and then flatten, with the validation loss slightly above the training loss. A label on the flat part says plateau: early stopping or ReduceLROnPlateau. A caption says both losses fall with a small stable gap, the model is learning and generalizing.

Validation loss goes up while training loss goes down

In the middle panel the training loss falls to 0.006 by epoch 50 and to 0.000 by epoch 200. The validation loss reaches its lowest value, 0.353, at epoch 19 and then rises to 1.003. The gap grows from 0.074 at epoch 10 to 1.002 at the end.

This is overfitting. After epoch 19 the network keeps lowering the training loss by fitting the noise in the 200 training labels, which has no pattern that could hold for the validation examples. Each further epoch makes it better on the training set and worse on held-out data.

The model to keep is the one from the lowest point of the validation curve. Saving the weights each time the validation loss reaches a new low gives a best checkpoint, a file holding the weights from that epoch, and everything to the right of it on the plot is training that made the model worse.

A plot titled "Overfitting". Training loss keeps falling and flattens near the bottom. Validation loss falls, reaches a minimum marked with a dot and a dashed vertical line labeled best checkpoint, then rises. The region to the right of the line is shaded and labeled overfitting zone. A caption says training loss keeps falling while validation loss diverges, the model is memorizing.

Fixes work on the gap: more training data, a smaller model, regularization such as weight decay or dropout, or stopping at the best checkpoint. The linear model in the left panel is the "smaller model" fix applied to this dataset. It ends at a validation loss of 0.340, lower than the best epoch of the large network.

Both losses stay high

In the right panel the training loss moves from 0.697 to 0.688 over 200 epochs and the validation loss stays next to it. Both are close to 0.69, the loss of a 50/50 guess. The gap is tiny, and that is not good news here, because a small gap only says the model does equally badly on both sets.

This is underfitting: the model has not fit the training data. The run above underfits because the learning rate is too small for 200 epochs to move the weights far. The other common causes are a model with too little capacity, meaning too few weights to represent the pattern, and regularization set so strong that it holds the weights back.

A quick test separates "cannot learn" from "has not learned yet". Train on about 10 examples and check that the training loss goes to near zero. A model and training loop that cannot memorize 10 examples have a bug or a setting that is far off, and more data will not fix that.

A plot titled "Underfitting on the curve". Training loss stays near 0.7 and validation loss near 0.78 for 100 epochs, both far above a dashed line at 0.2 labeled healthy target. Below the plot are three causes, too little capacity, training cut short, regularization too strong, and a highlighted box that says first check: overfit 10 examples. A caption says both losses are high with a small gap, the model is not fitting the training data yet.

What the training loss alone tells you about the learning rate

The shape of the training loss on its own says a lot about the learning rate, before the validation curve comes into it. Here the horizontal axis is usually the iteration, one weight update, instead of the epoch.

Three plots of loss against iteration titled "Loss over iterations: learning rate problems". Too high: the loss oscillates and the swings grow from about 1 to over 10. Good: the loss decreases smoothly from about 10 toward 0. Too low: the loss decreases slowly in a straight line from about 9 to 6. A caption says too high overshoots and never settles, too low barely moves, in range drops fast and then flattens.

What is a learning rate? goes through these three cases and how to pick a starting value.

A table of shapes

Shape What it usually means First thing to try
Both fall, small steady gap Healthy Train until the plateau
Train falls, validation rises Overfitting Keep the best checkpoint
Both high and flat Underfitting Overfit 10 examples
Loss jumps or grows Learning rate too high Divide it by 10
Slow straight decline Learning rate too low Multiply it by 10
Validation below training Dropout, or timing Evaluate both in eval mode

The last row has two ordinary causes. Dropout and similar layers are active while the training loss is computed and switched off for validation, so the training loss is measured on a model with part of its units zeroed out. The training loss is also often averaged over the batches of an epoch, while the weights were still improving, and the validation loss is measured at the end of it. The script above avoids both by recomputing the training loss in eval mode after each update. A validation loss far below the training loss is worth a closer look at how the data was split.

What a loss curve does not show

A loss curve averages over all the examples, so some problems do not appear in it.

Accuracy can move differently from the loss. Accuracy counts whether the top prediction is correct. Cross-entropy loss also depends on how much probability the model gave the correct class, and it grows without limit as that probability goes to zero. A model that becomes very sure of its wrong answers can keep nearly the same accuracy while its loss climbs.

Confidence can drift. The model's confidence in a prediction is the highest probability it assigns to any class. Averaging that over a dataset gives one number, the mean max probability, computed with probs.max(dim=1).values.mean(). The highest probability is used instead of the probability of the correct class because it also captures predictions that are confident and wrong.

The large network from the middle panel shows both effects. This block continues from the first script.

1h = curves["large MLP, lr 0.001"]
2for e in [18, 199]:
3    print(f"epoch {e + 1:3d}  val loss {h['val'][e]:.3f}  val accuracy {h['val_acc'][e]:.3f}  mean confidence {h['val_conf'][e]:.3f}")
4 
5model = runs["large MLP, lr 0.001"][0]             # the weights after epoch 200
6with torch.no_grad():
7    logits = model(X_val)
8    per_example = nn.functional.cross_entropy(logits, y_val, reduction="none")
9    wrong = logits.argmax(dim=1) != y_val
10print(f"wrong predictions: {wrong.float().mean():.1%} of examples, {per_example[wrong].sum() / per_example.sum():.1%} of the loss")
1epoch  19  val loss 0.353  val accuracy 0.850  mean confidence 0.849
2epoch 200  val loss 1.003  val accuracy 0.805  mean confidence 0.960

Between epoch 19 and epoch 200 the validation loss nearly triples, while validation accuracy falls by only 4.5 points, from 85.0% to 80.5%. At epoch 19 the model's mean confidence of 0.849 matches its accuracy of 0.850. At epoch 200 it is 96% confident on average and right 80.5% of the time. The last line splits the epoch-200 loss by example: the 19.5% of validation examples the model gets wrong account for 97.8% of the loss, because each of them is a wrong answer given with high confidence.

The picture below shows the related case where the two loss curves look close and the confidence still differs between the sets: 0.97 on the training data against 0.71 on the validation data. Logging the mean max probability on both sets next to the losses catches this.

A two-panel figure titled "The loss hides the confidence gap". Left, "Losses look close": training and validation loss both fall over 7 epochs and stay near each other. Right, "Confidence has drifted": a bar chart of mean max probability, 0.97 on train and 0.71 on val, with the difference labeled gap. A caption says same loss, different confidence.

Common mistakes

Overfitting and underfitting covers the causes and fixes behind the two failing shapes. The bias-variance tradeoff explains why a model that is too small and one that is too large fail in different ways. The train, validation and test split is where the validation curve's data comes from. Saving and resuming training shows how to write the best checkpoint. When the loss does not move at all, the loss-not-decreasing checklist goes through the usual causes.

QuiddityML teaches reading a loss curve as a concept in Unit 4 of the ML Foundation track, with questions such as whether a validation loss that sits 0.05 above the training loss for the whole run is a problem, and the spaced review brings the curve shapes back until they are quick to recognize (quiddityml.com).