By Sagi Shaier · 9 October 2026 · 5 min read
The dot product explained, and what it measures in machine learning
How to compute a dot product by hand, what its sign says about two vectors, how it relates to the angle between them, and the three PyTorch calls that compute it.
The dot product takes two vectors of the same length, multiplies their matching entries, and adds the results into a single number. It is also called the inner product. A typical neuron in a neural network computes one, every entry of a matrix multiplication is one, and comparing two embeddings usually starts with one, so it is one of the operations a model runs most often.
How to compute a dot product
A vector here is an ordered list of numbers, such as $(1, 2, 3)$. Write two vectors as $\mathbf{u}$ and $\mathbf{v}$, and write their entries as $u_i$ and $v_i$, where $i$ is the position in the list. The dot product is written $\mathbf{u} \cdot \mathbf{v}$:
$$\mathbf{u} \cdot \mathbf{v} = \sum_i u_i v_i$$
The $\sum_i$ means "add up over every position $i$". Multiply the first entries together, the second entries together, and so on, then add all the products.
For $\mathbf{u} = (1, 2, 3)$ and $\mathbf{v} = (4, 5, 6)$:
$$1\times4 + 2\times5 + 3\times6$$
$$= 4 + 10 + 18 = 32$$
The answer is one number, not a vector. Both vectors need the same number of entries, because every entry of $\mathbf{u}$ needs a partner in $\mathbf{v}$.
What the sign of the dot product says
The dot product measures how much two vectors point in the same direction, weighted by how long they are.
Take $\mathbf{u} = (3, 1)$ and three different choices of $\mathbf{v}$:
- $\mathbf{v} = (2, 2)$ points roughly the same way as $\mathbf{u}$, and $3\times2 + 1\times2 = 8$, a large positive number.
- $\mathbf{v} = (-1, 3)$ sits at a right angle to $\mathbf{u}$, and $3\times(-1) + 1\times3 = 0$.
- $\mathbf{v} = (-2, -2)$ points roughly the opposite way, and $3\times(-2) + 1\times(-2) = -8$, a large negative number.

Two vectors at a right angle give a dot product of exactly zero, however long they are. That case has its own name: the vectors are orthogonal.
The geometric form: lengths and the angle
The same number can also be written with the lengths of the two vectors and the angle between them. The length of a vector, written $\|\mathbf{u}\|$, is its straight-line size, $\sqrt{\sum_i u_i^2}$, so $(3, 4)$ has length 5. The angle between $\mathbf{u}$ and $\mathbf{v}$ is written $\theta$ (theta). Then:
$$\mathbf{u} \cdot \mathbf{v} = \|\mathbf{u}\| \|\mathbf{v}\| \cos(\theta)$$
Both formulas give the same number. The coordinate formula, multiply and add, is what you compute with. The geometric formula is what explains the sign, because the lengths are never negative and $\cos(\theta)$ decides the rest: it is positive when the angle is under 90 degrees, exactly 0 at 90 degrees, and negative above 90 degrees.

The formula also shows the catch with the raw dot product: double the length of $\mathbf{u}$ and the dot product doubles, even though the direction did not change. A large dot product can mean two vectors point the same way, or that one of them is long.
Where the dot product shows up in machine learning
A neuron. A single neuron takes inputs $x_1, \dots, x_n$, multiplies each by a weight $w_i$, adds them up, and adds a bias $b$. That weighted sum, $\sum_i w_i x_i$, is the dot product $\mathbf{w} \cdot \mathbf{x}$ between the weight vector and the input vector.
Matrix multiplication. To multiply two matrices, each entry of the result is computed by taking one row of the first matrix and one column of the second, multiplying matching entries, and adding them up, which is a dot product between that row and that column.
Comparing embeddings. An embedding is a vector a model uses to represent a word, a sentence, or an image. Two embeddings with a large dot product point in similar directions. Because length also inflates the dot product, comparisons that should depend on direction alone divide the length out first, which gives cosine similarity.
The dot product in PyTorch
1import torch
2
3u = torch.tensor([1.0, 2.0, 3.0])
4v = torch.tensor([4.0, 5.0, 6.0])
5
6# three ways to write the same dot product
7print((u * v).sum())
8print(torch.dot(u, v))
9print(u @ v)
10
11# the sign tracks the direction
12a = torch.tensor([3.0, 1.0])
13for b in ([2.0, 2.0], [-1.0, 3.0], [-2.0, -2.0]):
14 print(b, torch.dot(a, torch.tensor(b)).item())1tensor(32.)
2tensor(32.)
3tensor(32.)
4[2.0, 2.0] 8.0
5[-1.0, 3.0] 0.0
6[-2.0, -2.0] -8.0u * v multiplies matching entries and .sum() adds them, which is the definition written out. torch.dot does both steps for two 1D tensors, and @ gives the same result for 1D tensors and also works for matrices.
The same operation inside a neuron, checked against PyTorch's Linear layer:
1import torch
2
3w = torch.tensor([0.5, -1.0, 2.0]) # weights of one neuron
4x = torch.tensor([1.0, 3.0, 0.5]) # one input
5b = 0.25 # bias
6print(torch.dot(w, x) + b)
7
8# the same number from a Linear layer with those weights
9layer = torch.nn.Linear(3, 1)
10with torch.no_grad():
11 layer.weight.copy_(w.unsqueeze(0))
12 layer.bias.fill_(b)
13print(layer(x))1tensor(-1.2500)
2tensor([-1.2500], grad_fn=<ViewBackward0>)$0.5\times1 - 1\times3 + 2\times0.5 = -1.5$, plus the bias $0.25$ gives $-1.25$, and the Linear layer returns the same number because it computes exactly this dot product plus bias. If torch.dot raises an error, check that both tensors are 1D and the same length, since it accepts nothing else.
Common mistakes
- Expecting a vector back. The dot product returns one number. Multiplying matching entries without adding them,
u * v, is the elementwise product and returns a vector. - Mixing up lengths. $(1, 2, 3) \cdot (4, 5)$ has no answer, because the third entry of the first vector has no partner.
- Reading a big dot product as "very similar". A long vector gives large dot products with almost anything it roughly lines up with, so when only direction should count, divide by both lengths first.
- Forgetting that zero means a right angle. A dot product of 0 between two nonzero vectors means they are orthogonal, not that one of them is small.
Related math
Cosine similarity is the dot product divided by both lengths, so it measures direction alone and always falls between $-1$ and $1$. Norms are the different ways to measure the length $\|\mathbf{u}\|$ that appears in the geometric form. Matrix multiplication is a grid of dot products, one for every row and column pair, covered in matrix multiplication explained.
QuiddityML teaches the dot product as its own concept in the linear algebra part of the Math track, and the exercises include turning $\sum_i u_i v_i$ into code, ordering the lines of a dot_product function, and writing it from scratch.