QuiddityML

By · 9 October 2026 · 5 min read

The dot product explained, and what it measures in machine learning

How to compute a dot product by hand, what its sign says about two vectors, how it relates to the angle between them, and the three PyTorch calls that compute it.

The dot product takes two vectors of the same length, multiplies their matching entries, and adds the results into a single number. It is also called the inner product. A typical neuron in a neural network computes one, every entry of a matrix multiplication is one, and comparing two embeddings usually starts with one, so it is one of the operations a model runs most often.

How to compute a dot product

A vector here is an ordered list of numbers, such as $(1, 2, 3)$. Write two vectors as $\mathbf{u}$ and $\mathbf{v}$, and write their entries as $u_i$ and $v_i$, where $i$ is the position in the list. The dot product is written $\mathbf{u} \cdot \mathbf{v}$:

$$\mathbf{u} \cdot \mathbf{v} = \sum_i u_i v_i$$

The $\sum_i$ means "add up over every position $i$". Multiply the first entries together, the second entries together, and so on, then add all the products.

For $\mathbf{u} = (1, 2, 3)$ and $\mathbf{v} = (4, 5, 6)$:

$$1\times4 + 2\times5 + 3\times6$$

$$= 4 + 10 + 18 = 32$$

The answer is one number, not a vector. Both vectors need the same number of entries, because every entry of $\mathbf{u}$ needs a partner in $\mathbf{v}$.

What the sign of the dot product says

The dot product measures how much two vectors point in the same direction, weighted by how long they are.

Take $\mathbf{u} = (3, 1)$ and three different choices of $\mathbf{v}$:

Three plots of u = (3, 1) against a second vector v: v = (2, 2) gives a dot product of 8 for similar directions, v = (-1, 3) gives exactly 0 at a right angle, and v = (-2, -2) gives -8 for opposing directions, with the formula and (1, 2, 3) dot (4, 5, 6) = 32 underneath

Two vectors at a right angle give a dot product of exactly zero, however long they are. That case has its own name: the vectors are orthogonal.

The geometric form: lengths and the angle

The same number can also be written with the lengths of the two vectors and the angle between them. The length of a vector, written $\|\mathbf{u}\|$, is its straight-line size, $\sqrt{\sum_i u_i^2}$, so $(3, 4)$ has length 5. The angle between $\mathbf{u}$ and $\mathbf{v}$ is written $\theta$ (theta). Then:

$$\mathbf{u} \cdot \mathbf{v} = \|\mathbf{u}\| \|\mathbf{v}\| \cos(\theta)$$

Both formulas give the same number. The coordinate formula, multiply and add, is what you compute with. The geometric formula is what explains the sign, because the lengths are never negative and $\cos(\theta)$ decides the rest: it is positive when the angle is under 90 degrees, exactly 0 at 90 degrees, and negative above 90 degrees.

The geometric form u dot v = length of u times length of v times cos theta, drawn with the angle theta between u and v, and three small cases below: theta under 90 degrees gives a positive dot product, theta of 90 degrees gives zero, and theta over 90 degrees gives a negative one

The formula also shows the catch with the raw dot product: double the length of $\mathbf{u}$ and the dot product doubles, even though the direction did not change. A large dot product can mean two vectors point the same way, or that one of them is long.

Where the dot product shows up in machine learning

A neuron. A single neuron takes inputs $x_1, \dots, x_n$, multiplies each by a weight $w_i$, adds them up, and adds a bias $b$. That weighted sum, $\sum_i w_i x_i$, is the dot product $\mathbf{w} \cdot \mathbf{x}$ between the weight vector and the input vector.

Matrix multiplication. To multiply two matrices, each entry of the result is computed by taking one row of the first matrix and one column of the second, multiplying matching entries, and adding them up, which is a dot product between that row and that column.

Comparing embeddings. An embedding is a vector a model uses to represent a word, a sentence, or an image. Two embeddings with a large dot product point in similar directions. Because length also inflates the dot product, comparisons that should depend on direction alone divide the length out first, which gives cosine similarity.

The dot product in PyTorch

1import torch
2 
3u = torch.tensor([1.0, 2.0, 3.0])
4v = torch.tensor([4.0, 5.0, 6.0])
5 
6# three ways to write the same dot product
7print((u * v).sum())
8print(torch.dot(u, v))
9print(u @ v)
10 
11# the sign tracks the direction
12a = torch.tensor([3.0, 1.0])
13for b in ([2.0, 2.0], [-1.0, 3.0], [-2.0, -2.0]):
14    print(b, torch.dot(a, torch.tensor(b)).item())
1tensor(32.)
2tensor(32.)
3tensor(32.)
4[2.0, 2.0] 8.0
5[-1.0, 3.0] 0.0
6[-2.0, -2.0] -8.0

u * v multiplies matching entries and .sum() adds them, which is the definition written out. torch.dot does both steps for two 1D tensors, and @ gives the same result for 1D tensors and also works for matrices.

The same operation inside a neuron, checked against PyTorch's Linear layer:

1import torch
2 
3w = torch.tensor([0.5, -1.0, 2.0])   # weights of one neuron
4x = torch.tensor([1.0, 3.0, 0.5])    # one input
5b = 0.25                             # bias
6print(torch.dot(w, x) + b)
7 
8# the same number from a Linear layer with those weights
9layer = torch.nn.Linear(3, 1)
10with torch.no_grad():
11    layer.weight.copy_(w.unsqueeze(0))
12    layer.bias.fill_(b)
13print(layer(x))
1tensor(-1.2500)
2tensor([-1.2500], grad_fn=<ViewBackward0>)

$0.5\times1 - 1\times3 + 2\times0.5 = -1.5$, plus the bias $0.25$ gives $-1.25$, and the Linear layer returns the same number because it computes exactly this dot product plus bias. If torch.dot raises an error, check that both tensors are 1D and the same length, since it accepts nothing else.

Common mistakes

Cosine similarity is the dot product divided by both lengths, so it measures direction alone and always falls between $-1$ and $1$. Norms are the different ways to measure the length $\|\mathbf{u}\|$ that appears in the geometric form. Matrix multiplication is a grid of dot products, one for every row and column pair, covered in matrix multiplication explained.

QuiddityML teaches the dot product as its own concept in the linear algebra part of the Math track, and the exercises include turning $\sum_i u_i v_i$ into code, ordering the lines of a dot_product function, and writing it from scratch.