QuiddityML

By · 11 October 2026 · 6 min read

Derivatives for machine learning: limits, continuity, and the derivative

How the derivative is built from a limit, when a function has no derivative at a point (and why ReLU has a kink at 0), the seven rules that cover most expressions, and how to get the sigmoid's derivative in one line, with Python to check each result.

A derivative is the slope of a function at a single point: how fast the output changes when the input moves a tiny amount. Training a neural network runs on derivatives, because the model changes each weight in the direction that makes the loss go down, and the derivative of the loss with respect to that weight is what gives the direction.

What a limit is

The slope between two points is easy to compute. The slope at one point needs a way to talk about "getting closer and closer" without ever arriving, and that is what a limit does.

A limit asks what value $f(x)$ approaches as $x$ gets closer and closer to some value $a$. It does not ask what happens exactly at $a$, only what the function is heading toward as $x$ comes in from both sides. With $L$ for the value approached, it is written

$$\lim_{x \to a} f(x) = L$$

and read as "as $x$ approaches $a$, $f(x)$ approaches $L$."

Take $f(x) = \frac{x^2 - 4}{x - 2}$ near $x = 2$. Putting $x = 2$ in directly gives $\frac{0}{0}$, which is not a number, so the function has no value at 2. The numerator factors as $x^2 - 4 = (x - 2)(x + 2)$, and for every $x \neq 2$ the two $(x - 2)$ factors cancel, leaving $x + 2$. As $x$ approaches 2, $x + 2$ approaches 4:

$$\lim_{x \to 2} \frac{x^2 - 4}{x - 2} = 4$$

Evaluating the function on both sides of 2 shows the same thing:

1def f(x):
2    return (x**2 - 4) / (x - 2)
3 
4for x in [1.9, 1.99, 1.999, 2.001, 2.01, 2.1]:
5    print(x, round(f(x), 6))
11.9 3.9
21.99 3.99
31.999 3.999
42.001 4.001
52.01 4.01
62.1 4.1

The function approaches 4 from both sides and has no value at 2 itself. That $\frac{0}{0}$ shape is the one that shows up inside the definition of the derivative.

One-sided limits

"From both sides" can fail. The left-hand limit $\lim_{x \to a^-} f(x)$ is what $f$ approaches as $x$ comes up from below $a$. The right-hand limit $\lim_{x \to a^+} f(x)$ is what it approaches coming down from above. The two-sided limit exists when both one-sided limits exist and are equal. A step function that jumps from 0 to 1 at $a$ has a left-hand limit of 0 and a right-hand limit of 1, so it has no two-sided limit at $a$.

Continuity

A function is continuous at $a$ if its graph passes through that point with no jump, hole, or break. In terms of limits, $f$ is continuous at $a$ when

$$\lim_{x \to a} f(x) = f(a)$$

which says the value the function approaches is the value it takes there.

Top: a smooth curve through the point (a, 2), continuous because the limit equals the value. Bottom: a curve that approaches 2 from the left but takes the value 1 at a and continues from 1, not continuous because the two sides disagree, so there is no single slope at a

A slope at a point needs continuity there, since a function that jumps at $a$ has no single slope at the jump. Continuity is not enough on its own, though. $f(x) = |x|$ is continuous everywhere, and at $x = 0$ its slope is $-1$ coming from the left and $+1$ coming from the right, so it has no single slope at 0. A function with no derivative at a point is called not differentiable there.

Left: a step function with left-hand limit 0 and right-hand limit 1 at a, so the two-sided limit does not exist. Right: a continuous curve with a sharp corner at a, continuous there but not differentiable

This shows up in neural networks. ReLU, a common activation function, is $\text{ReLU}(x) = \max(0, x)$: 0 for negative inputs and $x$ for positive ones. It has the same kind of corner at 0, with slope 0 on the left and 1 on the right.

The derivative as a limit

Pick a point $x$ and a second point a distance $h$ away, at $x + h$. The slope of the straight line through the two points on the curve, called a secant line, is rise over run:

$$\frac{f(x+h) - f(x)}{h}$$

As $h$ shrinks toward 0, the second point slides toward the first. The secant line then turns into the tangent line, which matches the curve's direction at $x$. The slope of the tangent line is the derivative:

$$f'(x) = \lim_{h \to 0} \frac{f(x+h) - f(x)}{h}$$

It is also written $\frac{df}{dx}$, and the two notations mean the same thing.

For $f(x) = x^2$ at $x = 1$ with $h = 2$, the two points are $(1, 1)$ and $(3, 9)$, so the secant slope is $\frac{8}{2} = 4$. Shrinking $h$ moves that slope toward the derivative:

The curve f(x) = x squared from x = 0 to 3, with a dashed secant line from the point at x = 1 to the point at x + h = 3, rise 8 and run h = 2, so the secant slope is 8 / 2 = 4, and as h goes to 0 the secant becomes the tangent

1def f(x):
2    return x**2
3 
4x = 1.0
5for h in [2, 1, 0.1, 0.01, 0.001]:
6    slope = (f(x + h) - f(x)) / h
7    print(h, round(slope, 6))
12 4.0
21 3.0
30.1 2.1
40.01 2.01
50.001 2.001

The secant slope heads to 2, so the derivative of $x^2$ at $x = 1$ is 2.

The rules for computing derivatives

Taking the limit by hand every time is slow, and a few rules cover most expressions. Here $c$ is a constant and $n$ is a number.

Those four handle polynomials. Losses and activation functions in machine learning also use the exponential $e^x$ and the natural logarithm $\ln x$, and divisions of one expression by another, which take three more rules:

Worked example: the derivative of the sigmoid

The sigmoid function squashes any number into the range 0 to 1, and it is used as an activation function and to turn a score into a probability:

$$\sigma(x) = \frac{1}{1 + e^{-x}}$$

It is a quotient, with $u = 1$ on top and $v = 1 + e^{-x}$ underneath. The derivative of the constant 1 is $u' = 0$, and the exponential rule gives $v' = -e^{-x}$. The quotient rule then gives

$$\sigma'(x) = \frac{0 \cdot (1 + e^{-x}) - 1 \cdot (-e^{-x})}{(1 + e^{-x})^2} = \frac{e^{-x}}{(1 + e^{-x})^2}$$

Splitting the fraction into two factors shows more:

$$\frac{e^{-x}}{(1 + e^{-x})^2} = \frac{1}{1 + e^{-x}} \cdot \frac{e^{-x}}{1 + e^{-x}}$$

The first factor is $\sigma(x)$. The second is $1 - \sigma(x)$, because $1 - \frac{1}{1 + e^{-x}} = \frac{e^{-x}}{1 + e^{-x}}$. So

$$\sigma'(x) = \sigma(x)\big(1 - \sigma(x)\big)$$

The derivative uses only the sigmoid's own output. A network that already computed $\sigma(x)$ on the way forward gets the derivative with one subtraction and one multiplication, without evaluating the exponential again.

Three rule panels: the quotient rule u'v minus uv' over v squared with the order of the two terms fixed, the exponential rule with the derivative of e to the x equal to e to the x and of e to the minus x equal to minus e to the minus x, and the natural log rule with derivative 1 over x for x greater than 0, then the sigmoid and its derivative sigma times one minus sigma

A secant slope with a very small $h$ checks the result numerically:

1import math
2 
3def sigmoid(x):
4    return 1 / (1 + math.exp(-x))
5 
6x, h = 0.5, 1e-6
7numeric = (sigmoid(x + h) - sigmoid(x)) / h
8formula = sigmoid(x) * (1 - sigmoid(x))
9print(round(numeric, 6), round(formula, 6))
10.235004 0.235004

Common mistakes

The derivative here takes one input. Functions of many inputs, such as a loss that depends on every weight in a network, use partial derivatives and gradients. The exponential and log functions themselves are in exponents and logarithms for machine learning.

QuiddityML teaches limits and derivatives as the first concepts of the calculus part of the Math track, with exercises that find the sigmoid's derivative and spot the bug in a polynomial derivative function. The last one is writing derivative_of_polynomial from scratch.