QuiddityML

By · 7 October 2026 · 4 min read

Span, basis and linear independence explained

The span of a set of vectors is every point you can reach by scaling and adding them, a basis is the smallest set that reaches a whole space, and linear independence says no vector in the set is redundant. This post explains all three with 2D examples, what they mean for features in a dataset, and how to check them in PyTorch.

The span of a set of vectors is every point you can reach by scaling them and adding the results. A basis is the smallest set of vectors whose span is a whole space, and linear independence is the property that no vector in a set can be built from the others. In machine learning these ideas answer a practical question about data: if one feature can be computed from the others, it adds no information, and a model gets nothing new from it.

Linear combinations

A vector is an ordered list of numbers, such as $(1, 2)$. Two operations on vectors build everything in this post. Scaling multiplies every entry by a number, so $3(1, 2) = (3, 6)$. Adding goes entry by entry, so $(1, 2) + (5, -3) = (6, -1)$.

Scale a few vectors and add the results, and you get a linear combination:

$$c_1\mathbf{v}_1 + c_2\mathbf{v}_2 + \cdots + c_k\mathbf{v}_k$$

The numbers $c_1, \ldots, c_k$ are the coefficients, and they can be any real numbers: positive, negative or zero.

Tip-to-tail addition of (1, 2) and (5, minus 3) landing at (6, minus 1), and (1, 2) scaled by 3 to (3, 6)

Span: everything a set of vectors can reach

The examples here live in $\mathbb{R}^n$, the set of every list of $n$ real numbers. $\mathbb{R}^2$ is the 2D plane. Adding two lists in $\mathbb{R}^n$ or scaling one always gives another list in $\mathbb{R}^n$, and a set with that property is called a vector space.

Take one vector, $(1, 2)$. Scaling it by every real number traces out the line through the origin and $(1, 2)$. Take two vectors that point in different directions, $(1, 0)$ and $(0, 1)$, and the combinations $c_1(1, 0) + c_2(0, 1)$ reach every point of the plane. The set of everything reachable this way is the span:

$$\text{span}{\mathbf{v}_1, \ldots, \mathbf{v}_k} = {c_1\mathbf{v}_1 + \cdots + c_k\mathbf{v}_k : c_i \in \mathbb{R}}$$

Since the coefficients can be 0, every span contains the zero vector. Since they can be negative, a span that contains a vector contains the whole line through it, in both directions.

Two panels: (1, 0) and (0, 1) span all of R2 and form a basis, while (1, 2) and (2, 4) span only one line through the origin because the second adds no new direction

Linear independence: no redundant vectors

Add $(2, 4)$ to the set ${(1, 2)}$ and the span does not grow, because $(2, 4) = 2 \times (1, 2)$ already lies on the same line. A vector that can be built from the others in its set is redundant, and a set with a redundant vector is linearly dependent.

A set is linearly independent when none of its vectors can be built from the others. As an equation, ${\mathbf{v}_1, \ldots, \mathbf{v}_n}$ is linearly independent when the only way to get the zero vector

$$c_1\mathbf{v}_1 + c_2\mathbf{v}_2 + \cdots + c_n\mathbf{v}_n = \mathbf{0}$$

is to set every $c_i = 0$. For $(1, 2)$ and $(2, 4)$, the coefficients $c_1 = 2$ and $c_2 = -1$ give $2(1, 2) - (2, 4) = (0, 0)$, so that set is dependent.

Basis and dimension

A basis for a space is a set of vectors that spans the whole space using as few vectors as possible, which makes it linearly independent. $(1, 0)$ and $(0, 1)$ are a basis for $\mathbb{R}^2$: one vector only reaches a line, and two independent ones reach the plane. $\mathbb{R}^3$ needs three basis vectors, and a basis for $\mathbb{R}^n$ has exactly $n$. A space usually has many valid bases, and they all have the same number of vectors.

That number is the dimension of the space, which is why $\mathbb{R}^n$ has dimension $n$.

Linear independence in R2: the basis vectors (1, 0) and (0, 1) spanning the plane, the equation c1 v1 plus up to cn vn equals 0 true only when every ci is 0, and v and 2v as a dependent example

What this means for features

Treat each feature column of a dataset as a vector. If one feature is a linear combination of others, for example a price in euros next to the same price in dollars, it is linearly dependent on them and carries no extra information once the others are present. A basis is the smallest set of directions that captures everything a set of vectors contains, so asking how much distinct information a dataset holds is asking what the dimension of its span is.

Real features are rarely exactly dependent. Two features that are close to dependent, such as height in centimeters and height measured with some noise, cause their own problems, which is the subject of multicollinearity.

Checking it in PyTorch

1import torch
2 
3v1 = torch.tensor([1.0, 2.0])
4v2 = torch.tensor([2.0, 4.0])
5e1 = torch.tensor([1.0, 0.0])
6e2 = torch.tensor([0.0, 1.0])
7 
8# 2 * v1 - 1 * v2 is the zero vector, with coefficients that are not all 0
9print(2 * v1 - 1 * v2)
10 
11# any target in R^2 is a combination of e1 and e2
12target = torch.tensor([3.0, -5.0])
13print(3 * e1 + (-5) * e2, target)
14 
15# matrix_rank counts how many independent directions the columns have
16print(torch.linalg.matrix_rank(torch.stack([v1, v2], dim=1)))
17print(torch.linalg.matrix_rank(torch.stack([e1, e2], dim=1)))
1tensor([0., 0.])
2tensor([ 3., -5.]) tensor([ 3., -5.])
3tensor(1)
4tensor(2)

torch.stack(..., dim=1) puts the vectors side by side as the columns of a matrix, and torch.linalg.matrix_rank returns how many of those columns are linearly independent. $(1, 2)$ and $(2, 4)$ give 1, a single line, and $(1, 0)$ and $(0, 1)$ give 2, the whole plane. On a feature matrix, a result smaller than the number of columns means at least one feature is redundant.

Common mistakes

Vectors and matrices themselves, with shapes, indexing and the transpose, are in vectors and matrices for machine learning. Multicollinearity is what happens when features are close to dependent without being exactly dependent.

QuiddityML teaches span and basis as their own concepts in the linear algebra part of the Math track, and the exercises include turning $c_1\mathbf{e}_1 + c_2\mathbf{e}_2$ into code, ordering the lines of a linear_combination function, and writing it from scratch.