By Sagi Shaier · 10 October 2026 · 5 min read
What is a tensor? Shapes, batches, and tensor operations
How to read a tensor's shape, why models process a whole batch with one matrix multiplication, and when to use reshape, squeeze, sum and permute, with PyTorch code for each and the reshape mistake that silently scrambles attention heads.
A tensor is an array of numbers with any number of dimensions. The inputs, weights, and activations of a PyTorch model are stored as tensors, so reading a tensor's shape and changing it correctly is a large part of writing model code.
Scalars, vectors, matrices, and beyond
A single number is a 0-dimensional tensor, a scalar. A list of numbers is a 1-dimensional tensor, a vector. A grid of rows and columns is a 2-dimensional tensor, a matrix. Past two dimensions there is no special name, and they are all called tensors.
Each dimension is also called an axis, and the shape lists the size of every axis. The number of axes is sometimes called the tensor's rank, which is a different meaning of the word than the rank of a matrix: tensor rank counts axes, while matrix rank counts independent directions.

The axes usually have a meaning, and the shape tells you how the data is laid out:
- A batch of 32 grayscale images, each 64 by 64 pixels, has shape $(32, 64, 64)$.
- Make them color images and a channel axis appears for red, green, and blue: $(32, 3, 64, 64)$, read as batch, channels, height, width.
- A batch of 64 sentences, each padded to 20 words, with every word stored as a list of 300 numbers (its embedding), has shape $(64, 20, 300)$.
The first axis is usually the batch: the number of examples processed together.
Matrix operations on a whole batch
Linear algebra operations extend to tensors by running independently over the extra leading axes. torch.matmul on tensors of shape $(32, 4, 6)$ and $(32, 6, 2)$ treats the leading 32 as a batch axis and does 32 separate multiplications of a $4 \times 6$ matrix by a $6 \times 2$ matrix in one call, returning shape $(32, 4, 2)$. This is how a neural network layer handles a whole batch in one operation instead of a Python loop over the examples.
1import torch
2
3images = torch.zeros(32, 3, 64, 64) # batch, channels, height, width
4sentences = torch.zeros(64, 20, 300) # batch, words, embedding size
5print(images.ndim, images.shape)
6print(sentences.ndim, sentences.shape)
7
8a = torch.randn(32, 4, 6)
9b = torch.randn(32, 6, 2)
10c = torch.matmul(a, b)
11print(c.shape)
12print(torch.allclose(c[5], a[5] @ b[5]))14 torch.Size([32, 3, 64, 64])
23 torch.Size([64, 20, 300])
3torch.Size([32, 4, 2])
4True.ndim is the number of axes. The last line checks that entry 5 of the batched result equals the ordinary matrix product of entry 5 from each input.
Changing the shape without changing the numbers
Several operations rearrange the same numbers into a different shape, usually to make two tensors line up for a matrix operation:
.reshape()gives the tensor a new shape with the same data, as long as the total number of elements stays the same. A $(32, 4, 6)$ tensor holds 768 numbers and can become $(32, 24)$..squeeze()removes axes of size 1, and.unsqueeze(dim)inserts a size-1 axis at positiondim. Turning one example of shape $(5,)$ into a batch of one, $(1, 5)$, isunsqueeze(0).
Reductions like .sum() and .mean() collapse an axis instead. They take a dim argument naming the axis to collapse, and that axis disappears from the result: t.sum(dim=1) on shape $(32, 4, 6)$ gives $(32, 6)$. Negative numbers count from the end, so dim=-1 is the last axis whatever the number of axes.
1import torch
2
3t = torch.randn(32, 4, 6)
4print(t.reshape(32, 24).shape)
5print(t.sum(dim=1).shape)
6print(t.mean(dim=-1).shape)
7
8v = torch.randn(5)
9print(v.unsqueeze(0).shape)
10print(v.unsqueeze(0).squeeze().shape)1torch.Size([32, 24])
2torch.Size([32, 6])
3torch.Size([32, 4])
4torch.Size([1, 5])
5torch.Size([5])Reordering axes with permute
Changing the order of the axes is a different job. .transpose(i, j) swaps two axes, and .permute(...) reorders all of them. t.permute(0, 2, 1, 3) turns shape $(32, 20, 8, 64)$ into $(32, 8, 20, 64)$, because the new order is old axis 0, then 2, then 1, then 3.
Attention code in transformers commonly does this to go from (batch, sequence, heads, head size) to (batch, heads, sequence, head size), so each attention head gets its own sequence to work on.
![Tensor shape operations on a 4D tensor [B, S, H, D] for batch, seq, heads and head_dim: unsqueeze(dim=2) inserts a size-1 axis giving [B, S, 1, H, D], sum(dim=1) removes the seq axis giving [B, H, D], and permute(0, 2, 1, 3) reorders the axes giving [B, H, S, D]. A warning reads: reshape changes shape, permute changes axis order](images/shape_tools_reshape_squeeze_reduce_permute.png)
Reshape is the wrong tool for this. It can return the right shape filled with the wrong numbers, because reshape keeps the numbers in their existing order and only cuts that order into new rows, while permute moves each number to its new position:
1import torch
2
3# batch=1, seq=2, heads=3, head_dim=1, numbered so each entry is easy to follow
4x = torch.arange(6).reshape(1, 2, 3, 1)
5print(x[0, :, :, 0]) # rows are positions, columns are heads
6
7right = x.permute(0, 2, 1, 3) # (1, 3, 2, 1)
8wrong = x.reshape(1, 3, 2, 1) # same shape, different numbers
9print(right[0, :, :, 0]) # rows are heads, columns are positions
10print(wrong[0, :, :, 0])
11print(right.shape == wrong.shape, torch.equal(right, wrong))1tensor([[0, 1, 2],
2 [3, 4, 5]])
3tensor([[0, 3],
4 [1, 4],
5 [2, 5]])
6tensor([[0, 1],
7 [2, 3],
8 [4, 5]])
9True FalseHead 0 holds the values 0 and 3 (one per position). permute puts them together in the first row. reshape puts 0 and 1 together instead, mixing head 0 and head 1, and no error is raised because the shape is correct.
Common mistakes
- Using reshape to reorder axes. If the goal is a different axis order, use
permuteortranspose. A model trained on reshaped heads still runs and learns from scrambled inputs. - Reducing over the wrong axis.
sum(dim=0)on a batch sums across examples, which is rarely the intent. Print the shape after every reduction until the axes are familiar. - Mixing up the two meanings of rank. A $(32, 4)$ tensor has tensor rank 2 whatever its numbers, while its matrix rank depends on how many independent directions its columns span.
- Forgetting the batch axis. Code that expects shape $(\text{batch}, 5)$ and reduces over
dim=0sums the 5 features instead of the batch when it gets a single example of shape $(5,)$.unsqueeze(0)adds the batch axis.
Related math
Combining tensors whose shapes differ, such as adding one bias vector to every row of a batch, is broadcasting, covered in broadcasting in NumPy and PyTorch explained. The matrix product that torch.matmul repeats across the batch is in matrix multiplication explained, and matrix rank is in rank, column space and null space explained.
QuiddityML teaches tensors and tensor operations as a concept in the linear algebra part of the Math track, and the exercises include tracing the shapes through a batched linear layer, spotting the bug in shape code, and writing an add_batch_dim function from scratch.