The derivative of the absolute value function is a classic example that illustrates how a seemingly simple function can present subtle challenges when we apply the rules of differentiation. That said, in this guide we will walk through the reasoning step‑by‑step, examine the special case at zero, and show how the result can be expressed using familiar functions such as the sign function. Understanding how to find derivative of absolute value expressions is essential for calculus students, engineers, and anyone working with optimization, signal processing, or machine learning models that involve |x| terms. By the end you will have a clear, procedural method for differentiating any expression that contains an absolute value, together with intuition about why the derivative behaves the way it does.
Understanding the Absolute Value Function
The absolute value of a real number x, denoted |x|, returns the non‑negative magnitude of x regardless of its sign. Formally,
[ |x| = \begin{cases} \phantom{-}x, & \text{if } x \ge 0,\[4pt] -,x, & \text{if } x < 0. \end{cases} ]
Graphically, |x| is a V‑shaped curve with its vertex at the origin. Because the function is linear on each side of the vertex, its slope is constant there—except exactly at the vertex where the direction changes abruptly. This piecewise linear nature is the key to finding its derivative Small thing, real impact. No workaround needed..
Why the Derivative is Tricky
If we blindly apply the power rule to |x| as if it were x, we would obtain an incorrect result. The difficulty stems from the fact that |x| is not differentiable at x = 0 in the ordinary sense; the left‑hand and right‑hand limits of the difference quotient do not agree. Away from zero, however, the function behaves like a simple line, and ordinary differentiation rules apply without issue. Recognizing where the function is smooth and where it is not prevents common mistakes and leads to a correct derivative expression.
Piecewise Definition Approach
The most straightforward method to differentiate |x| is to replace it with its piecewise definition and differentiate each branch separately.
Derivative for x > 0
When x > 0, the absolute value reduces to |x| = x. Differentiating gives
[ \frac{d}{dx}|x| = \frac{d}{dx}x = 1. ]
Thus, for any positive x, the slope of |x| is +1 Most people skip this — try not to..
Derivative for x < 0
When x < 0, we have |x| = -x. Differentiating yields
[ \frac{d}{dx}|x| = \frac{d}{dx}(-x) = -1. ]
Hence, for any negative x, the slope is -1.
Collecting these results, we can write the derivative as a piecewise function:
[ \frac{d}{dx}|x| = \begin{cases} \phantom{-}1, & x > 0,\[4pt] -1, & x < 0. \end{cases} ]
Derivative at x = 0 (Non‑differentiability)
To examine the behavior at the origin, we compute the left‑hand and right‑hand limits of the difference quotient:
[ \begin{aligned} \lim_{h\to0^{+}} \frac{|0+h|-|0|}{h} &= \lim_{h\to0^{+}} \frac{h}{h} = 1,\[6pt] \lim_{h\to0^{-}} \frac{|0+h|-|0|}{h} &= \lim_{h\to0^{-}} \frac{-h}{h} = -1. \end{aligned} ]
Since the two one‑sided limits differ (‑1 ≠ +1), the ordinary derivative does not exist at x = 0. In calculus language we say that |x| is not differentiable at zero, although it is continuous there. This sharp corner is the hallmark of a point where the derivative jumps That's the part that actually makes a difference..
People argue about this. Here's where I land on it.
Alternative: Using the Sign Function
A compact way to express the derivative for all x ≠ 0 is to invoke the sign function, denoted sgn(x) or sometimes (\operatorname{sign}(x)). It is defined as
[ \operatorname{sgn}(x) = \begin{cases} \phantom{-}1, & x > 0,\[4pt] 0, & x = 0,\[4pt] -1, & x < 0. \end{cases} ]
Notice that for x ≠ 0, sgn(x) matches the piecewise derivative we derived earlier. Therefore we can write
[ \frac{d}{dx}|x| = \operatorname{sgn}(x) \quad \text{for } x \neq 0. ]
At x = 0 the sign function is conventionally defined as 0, but this does not represent the derivative of |x| (which is undefined). Some texts leave the derivative undefined at zero, while others assign the subdifferential value [−1, 1] (see next section) Worth keeping that in mind..
Subdifferential and Generalized Derivative
In optimization and nonsmooth analysis, the concept of a subgradient extends the derivative to points where the classical derivative fails. For the absolute value function, the subdifferential at x = 0 is the interval
[ \partial|0| = [-1,,1]. ]
Any number g in this interval satisfies the subgradient inequality
[ |y| \ge |0| + g(y-0) \quad \text{for all } y\in\mathbb{R}. ]
Thus, while the ordinary derivative does not exist at zero, we can still work with a set‑valued generalization that captures all possible slopes of supporting lines to the V‑shape at the vertex. This idea is frequently used in convex optimization algorithms such as subgradient descent.
Examples and Applications
Example 1: Simple Linear Combination
Find the derivative of f(x) = 3|x| − 2x.
Using linearity and the result d/dx|x| = sgn(x) (for x≠0),
[ f'(x) = 3,\operatorname{sgn}(x) - 2,\qquad x\neq0. ]
At x = 0
At (x=0) the ordinary derivative still fails to exist, because the one‑sided slopes are different:
[ \begin{aligned} \lim_{h\to0^{+}}\frac{f(0+h)-f(0)}{h} &=\lim_{h\to0^{+}}\frac{3h-2h}{h}=1,\[4pt] \lim_{h\to0^{-}}\frac{f(0+h)-f(0)}{h} &=\lim_{h\to0^{-}}\frac{-3h-2h}{h}= -5 . \end{aligned} ]
Since (1\neq-5), (f'(0)) is undefined in the classical sense.
Even so, the subdifferential of a convex function at a point contains every slope of a supporting line. For (f) we obtain
[ \partial f(0)=\bigl[f'-(0),,f'+(0)\bigr]=[-5,,1], ]
so any number (g\in[-5,1]) can serve as a subgradient at the origin. Worth calling out: a subgradient‑descent algorithm could use (g=0) (the “mid‑point” of the interval) as a legitimate direction of steepest decrease at (x=0).
Example 2: A Piecewise‑Defined Function with Absolute Values
Consider
[ g(x)=|x-2|+|x+3|. ]
Because the absolute‑value function is differentiable everywhere except at its argument’s zero, (g) is differentiable for all (x\neq2) and (x\neq-3). Using the sign function we obtain
[ g'(x)=\operatorname{sgn}(x-2)+\operatorname{sgn}(x+3),\qquad x\neq2,,-3 . ]
Explicitly,
[ g'(x)= \begin{cases} -2, & x<-3,\[4pt] 0, & -3<x<2,\[4pt] 2, & x>2 . \end{cases} ]
At the “kinks’’ (x=-3) and (x=2) the one‑sided derivatives jump, so the ordinary derivative does not exist there. The subdifferentials are
[ \partial g(-3)=[-2,0],\qquad \partial g(2)=[0,2], ]
reflecting the flat region ((-3,2)) where the function is linear with slope 0.
Example 3: Composition with a Smooth Function
Let
[ h(x)=\sin(|x|). ]
Using the chain rule together with the sign function (valid for (x\neq0)):
[ h'(x)=\cos(|x|)\cdot\frac{d}{dx}|x| =\cos(|x|),\operatorname{sgn}(x),\qquad x\neq0. ]
Since (\cos(|x|)=\cos x) for all (x), we can write
[ h'(x)=\cos x;\operatorname{sgn}(x),\qquad x\neq0. ]
Again, at (x=0) the derivative is undefined because the sign factor jumps. Even so, the subgradient set at the origin is ({-1,1}) (the two possible slopes of the supporting lines to (\sin|x|) at the vertex), which can be useful in nonsmooth optimization contexts.
Concluding Remarks
The absolute‑value function furnishes a simple yet powerful illustration of how differentiability can break down at a single point while the function remains continuous everywhere. By introducing the sign function, we obtain a compact description of the derivative away from the kink, and by turning to subgradients we extend the notion of a derivative to those problematic points in a way that is both mathematically rigorous and practically useful.
These ideas permeate many areas of mathematics and its applications: in signal
processing, where total‑variation denoising relies on the subdifferential of the absolute value to preserve edges while removing noise; in compressed sensing, where the (\ell_1)-norm (a sum of absolute values) serves as a convex surrogate for the (\ell_0) “counting” norm, enabling sparse signal recovery; and in machine learning, where hinge‑loss and (\ell_1)-regularization both inherit their nonsmooth character from (|x|) and are routinely optimized via subgradient, proximal‑gradient, or alternating‑direction methods Worth knowing..
The same principles extend naturally to higher dimensions. The multivariate absolute value (|x|_1 = \sum_i |x_i|) has a subdifferential at any point (x) given by the Cartesian product of the one‑dimensional subdifferentials, i.,
[
\partial |x|_1 = \bigl{ g \in \mathbb{R}^n \mid g_i \in \partial |x_i| \text{ for all } i \bigr},
]
which is precisely the set of vectors whose (i)th component equals (\operatorname{sgn}(x_i)) when (x_i \neq 0) and lies in ([-1,1]) when (x_i = 0). And e. This structure underpins the optimality conditions for Lasso, basis pursuit, and a host of other sparse‑estimation problems Still holds up..
More broadly, the passage from classical derivatives to subgradients exemplifies a central theme in modern analysis: nonsmoothness is not a pathology to be avoided but a structure to be exploited. The absolute‑value function, in its simplicity, provides the canonical building block for this theory. Whether one is analyzing the convergence of a subgradient method, designing a proximal operator for a composite objective, or proving exact recovery guarantees in compressed sensing, the humble kink at the origin—and the interval of slopes it generates—remains the essential geometric insight No workaround needed..
To keep it short, the derivative of (|x|) fails to exist at a single point, yet this failure is far from a dead end. So by embracing the sign function away from the origin and the subdifferential at the origin, we obtain a complete, computationally tractable description of the function’s local behavior. This dual perspective—classical where possible, generalized where necessary—lies at the heart of contemporary optimization, signal processing, and statistical learning, turning a seemingly elementary piecewise‑linear function into a cornerstone of applied mathematics.