QuiddityML

By · 9 October 2026 · 5 min read

Cosine similarity explained: the angle between two vectors

How to compute cosine similarity by hand and in PyTorch, what scores of 1, 0 and -1 mean, and when to compare embeddings with it instead of the raw dot product.

Cosine similarity is a number between $-1$ and $1$ that says how closely two vectors point in the same direction, ignoring how long they are. It is the usual way to compare word, sentence, and image embeddings, the vectors a model produces to represent a piece of text or a picture, because the question there is whether two things mean something similar, not how large their vectors happen to be.

The problem with the raw dot product

The dot product of two vectors $\mathbf{u}$ and $\mathbf{v}$ multiplies their matching entries and adds the results: $\mathbf{u} \cdot \mathbf{v} = \sum_i u_i v_i$. A large positive dot product means the vectors point in similar directions, zero means they sit at a right angle, and a large negative one means they point in opposite directions.

The dot product also grows with length. Take $\mathbf{u} = (6, 3)$ and $\mathbf{v} = (2, 1)$. Both point in exactly the same direction, since $\mathbf{u}$ is $\mathbf{v}$ times 3, yet their dot product is $6\times2 + 3\times1 = 15$. Make $\mathbf{u}$ longer and that 15 keeps growing, while the directions stay identical. The raw dot product mixes two separate things, how aligned the vectors are and how long they are, into one number.

Dividing the length out

The dot product has a second form that uses the lengths of the vectors and the angle between them. The length of $\mathbf{u}$, written $\|\mathbf{u}\|$, is $\sqrt{\sum_i u_i^2}$, the straight-line size of the vector. The angle between the two vectors is $\theta$ (theta). Then:

$$\mathbf{u} \cdot \mathbf{v} = \|\mathbf{u}\| \|\mathbf{v}\| \cos(\theta)$$

The lengths are the part to get rid of. Divide both sides by $\|\mathbf{u}\| \|\mathbf{v}\|$ and they cancel, leaving only the angle:

$$\cos(\theta) = \frac{\mathbf{u} \cdot \mathbf{v}}{\|\mathbf{u}\| \|\mathbf{v}\|}$$

That right-hand side is cosine similarity: the dot product divided by both lengths, so length no longer affects the result.

For $\mathbf{u} = (6, 3)$ and $\mathbf{v} = (2, 1)$, the lengths are $\|\mathbf{u}\| = \sqrt{45} \approx 6.71$ and $\|\mathbf{v}\| = \sqrt{5} \approx 2.24$, and:

$$\frac{15}{\sqrt{45} \times \sqrt{5}} = \frac{15}{15} = 1$$

A cosine similarity of 1 says the two vectors point in exactly the same direction, which they do.

Left: the long vector u = (6, 3) and the short vector v = (2, 1) point the same way but have a dot product of 15 that depends on their lengths. Right: after dividing each by its length both become (0.894, 0.447) with length 1, and the cosine similarity is 1.00

Dividing a vector by its own length gives a vector of length 1 pointing the same way, called a unit vector, and the picture shows that view: once both vectors are unit vectors, $(0.894, 0.447)$ here, their plain dot product is the cosine similarity.

What the values mean

Cosine similarity always falls between $-1$ and $1$:

Three plots: v twice the length of u in the same direction gives cos theta = 1 at 0 degrees, v at a right angle to u gives cos theta = 0 at 90 degrees, and v pointing opposite to u gives cos theta = -1 at 180 degrees, with the cosine similarity formula and a number line from -1 (opposite) through 0 (perpendicular) to 1 (same direction)

The bound comes from the Cauchy-Schwarz inequality, which says that for any two vectors $|\mathbf{u} \cdot \mathbf{v}| \leq \|\mathbf{u}\| \|\mathbf{v}\|$: the dot product can never be bigger in size than the product of the two lengths. Divide both sides by $\|\mathbf{u}\| \|\mathbf{v}\|$ and the fraction can never leave $[-1, 1]$.

The zero vector is the one input with no cosine similarity, since its length is 0 and the formula would divide by 0.

Cosine similarity in PyTorch

1import torch
2import torch.nn.functional as F
3 
4u = torch.tensor([6.0, 3.0])
5v = torch.tensor([2.0, 1.0])
6 
7print(torch.dot(u, v))                           # raw dot product
8print(torch.dot(u, v) / (u.norm() * v.norm()))   # cosine similarity by hand
9print(F.cosine_similarity(u, v, dim=0))          # the built-in
10 
11# orthogonal and opposite
12e = torch.tensor([1.0, 0.0])
13print(F.cosine_similarity(e, torch.tensor([0.0, 5.0]), dim=0))
14print(F.cosine_similarity(e, torch.tensor([-3.0, 0.0]), dim=0))
1tensor(15.)
2tensor(1.)
3tensor(1.0000)
4tensor(0.)
5tensor(-1.)

u.norm() is the length $\|\mathbf{u}\|$, so the second line is the formula written out. F.cosine_similarity computes the same thing along the dimension given by dim.

With several embeddings at once, the difference between the two scores shows up directly. Here are four made-up 3-number embeddings, one per row, compared with one query vector:

1import torch
2import torch.nn.functional as F
3 
4# 4 made-up sentence embeddings, one per row
5emb = torch.tensor([
6    [0.9, 0.1, 0.3],
7    [1.8, 0.2, 0.6],
8    [0.1, 0.9, 0.2],
9    [-0.9, -0.1, -0.3],
10])
11query = torch.tensor([1.0, 0.1, 0.25])
12 
13print(emb @ query)                                       # raw dot products
14print(F.cosine_similarity(emb, query.unsqueeze(0), dim=1))
15 
16# normalize first, then a plain matrix product gives every pairwise cosine
17unit = F.normalize(emb, dim=1)
18print(unit @ unit.T)
1tensor([ 0.9850,  1.9700,  0.2400, -0.9850])
2tensor([ 0.9970,  0.9970,  0.2499, -0.9970])
3tensor([[ 1.0000,  1.0000,  0.2713, -1.0000],
4        [ 1.0000,  1.0000,  0.2713, -1.0000],
5        [ 0.2713,  0.2713,  1.0000, -0.2713],
6        [-1.0000, -1.0000, -0.2713,  1.0000]])

Row 2 is row 1 times 2. The raw dot product ranks row 2 as twice as close to the query as row 1, purely because it is longer, while cosine similarity gives both 0.997. F.normalize divides each row by its length, and after that a single matrix product unit @ unit.T gives the cosine similarity of every pair of rows. If the scores look wrong, check dim: it has to point at the axis that holds each vector's entries, which is dim=1 when every row is one embedding.

When to use cosine similarity, and when not to

Cosine similarity is a common choice for comparing embeddings, because the question "do these two pieces of text mean similar things" is about direction. It is also handy whenever vectors come out at very different lengths for reasons unrelated to what you want to compare.

The raw dot product is the better tool when length is meant to carry information. If a larger vector really should score higher, dividing the length out throws that information away.

Common mistakes:

The dot product is the numerator of cosine similarity and is covered on its own in the dot product explained. Norms are the ways to measure the vector length in the denominator, and the L2 norm is the one used here. Euclidean distance is another way to compare two embeddings, and it does depend on length.

QuiddityML teaches cosine similarity as its own concept in the linear algebra part of the Math track, and the exercises include turning the formula into code, spotting the bug in a cosine_similarity function, and writing it from scratch.