Maximum and Envelope Theorems
Y. Eddie Lu, Summer 2026
ECON 8001 course index · Chapter 4, Sections 4.6–4.7
Orientation
This page uses the correspondence conventions from Topology. Write a parameter as \(q\in U\subseteq\mathbb R^L\), a choice as \(x\in\mathbb R^N\), and a feasibility correspondence as \(\Psi:U\rightrightarrows\mathbb R^N\). A value function is a number; an argmax is usually a set.
Course map
- Combine compactness and continuity of the objective with continuity of the feasible-set correspondence.
- Apply the maximum theorem.
- Obtain continuity of the value and upper hemicontinuity of the argmax.
- Add differentiability assumptions when the question concerns derivatives.
- Apply an envelope theorem to the optimized value.
The maximum theorem is about continuity. The envelope theorem is about derivatives. The first often supplies existence and continuity needed to set up the second, but it does not by itself make a value differentiable.
Before this page: Topology for compactness and hemicontinuity, Differentiation for the chain rule, and Optimization for first-order and KKT conditions. Next: Convexity gives conditions that make an argmax single-valued.
Prerequisite retrieval
Retrieval check. If \(\Psi(q)\) is compact-valued, what does this say about one particular set \(\Psi(q)\)? What does it not say about the union \(\bigcup_{q\in U}\Psi(q)\)?
4.6 The maximum theorem
4.6.1 Set-up and the three conclusions
Let \(U\subseteq\mathbb R^L\) be open. Let \(u:U\times\mathbb R^N\to\mathbb R\) be continuous and let \(\Psi:U\rightrightarrows\mathbb R^N\) be a feasibility correspondence. Define
\[ V(q)=\max_{x\in\Psi(q)}u(q,x). \]
\[ A(q)=\operatorname*{arg\,max}_{x\in\Psi(q)}u(q,x). \]
The notation above presumes that the maximum is attained. Berge’s theorem supplies precisely that fact under its hypotheses.
The version proved in Nachbar’s Theorem of the Maximum holds \(u(q,x)=f(x)\) fixed in \(q\). The displayed version is the usual joint- continuous formulation of Berge’s theorem. It has the same three conclusions.
Why each argmax set is nonempty and compact
For a fixed \(q\), nonemptiness is the extreme-value theorem: \(\Psi(q)\) is nonempty and compact, and \(x\mapsto u(q,x)\) is continuous. Therefore \(u(q,\cdot)\) attains a maximum on \(\Psi(q)\). The set of all such maximizers is nonempty. It is compact because it is a closed subset of the compact set \(\Psi(q)\).
Why the argmax is only upper hemicontinuous
Upper hemicontinuity says that if \(q_j\to q\) and \(x_j\in A(q_j)\), then every convergent subsequence of \((x_j)\) can only limit to a member of \(A(q)\). In the finite-dimensional compact setting, this is the right stability statement: optimizers cannot suddenly acquire a remote limit outside the old argmax set.
Why the value is continuous
For values, the theorem controls both directions of variation. Nearby maximizers cannot produce values much above \(V(q)\), because their subsequential limits are maximizers at \(q\). Lower hemicontinuity lets a maximizer at \(q\) be approximated by feasible choices at nearby parameters, so nearby values cannot fall much below \(V(q)\).
Retrieval check. In the maximum theorem, which property retrieves a nearby feasible approximation to a fixed \(x\in\Psi(q)\)?
Assumption audit for Theorem 4.6.1
- Nonempty values: without them, there may be no feasible optimization problem.
- Compact values: without them, the supremum may not be attained.
- Continuity of the objective: without it, a limit of optimizing choices need not preserve a value comparison.
- Upper hemicontinuity of feasibility: without it, nearby optimizers can acquire limiting points outside the limiting feasible set.
- Lower hemicontinuity of feasibility: without it, a limiting competitor can disappear after an arbitrarily small parameter change.
The theorem has no uniqueness hypothesis. If strict concavity of \(x\mapsto u(q,x)\) on a convex feasible set makes \(A(q)\) a singleton for every \(q\), then Berge’s upper hemicontinuity makes the unique optimizer continuous. This conclusion needs uniqueness throughout the relevant parameter set, not merely at one parameter value.
4.7 The envelope theorem
4.7.1 Unconstrained envelope theorem
The optimizer may move when \(\theta\) changes. The envelope theorem says that, at a parameter with a unique optimizer, the first-order movement in optimized value is found by holding that optimizer fixed and differentiating only the primitive objective.
Here both sides are \(1\times L\) row vectors. Notice what the theorem does not assume: it does not require a differentiable optimizer selection.
This fixed-feasible-set result is stronger than the elementary chain-rule version because differentiability of an optimizer selection is not assumed. Compactness, upper hemicontinuity of the argmax, and uniqueness at \(\theta_0\) control the optimizing choices used in the difference quotient.
4.7.2 Constrained envelope theorem
The constraint term has a sign and a source. We use the maximization convention
\[ \mathcal L(x,\lambda;\theta) =f(x,\theta)-\lambda^{\mathsf T}G(x,\theta). \]
Thus multipliers on active inequalities of the form \(h_j\le0\) are nonnegative. Equality multipliers remain unrestricted.
Beyond Section 4.7: arbitrary choice sets
The baseline theorem assumes a differentiable optimizer. That assumption can be too strong: optimizers can jump, multiple optima can coexist, and an economic choice set may lack convenient topological or convex structure. Milgrom and Segal separate the envelope identity from the task of proving differentiability of the value.
Let \(X\) be any set, let \(f:X\times[0,1]\to\mathbb R\), and define
\[ V(t)=\sup_{x\in X}f(x,t). \]
\[ X^*(t)=\{x\in X:f(x,t)=V(t)\}. \]
The set \(X^*(t)\) can be empty: a supremum need not be attained. Every use of an optimizer below therefore explicitly assumes \(X^*(t)\ne\varnothing\).
This statement is Theorem 1 of Milgrom and Segal (2002). It imposes no topology, convexity, or continuity on \(X\). It also does not say that \(V\) is differentiable. Differentiability of \(V\) is an explicit premise of the equality conclusion.
A useful stronger conclusion
Milgrom and Segal’s Theorem 2 gives conditions for absolute continuity of \(V\). Suppose \(V(0)\in\mathbb R\), every \(f(x,\cdot)\) is absolutely continuous, and there is a function \(b\in L^1([0,1])\) and a null set \(N\) such that, for every \(t\in[0,1]\setminus N\) and every \(x\in X\),
\[ f_t(x,t)\text{ exists and }\lvert f_t(x,t)\rvert\le b(t). \]
Then \(V\) is absolutely continuous. If, in addition, every \(f(x,\cdot)\) is differentiable on \([0,1]\) and \(X^*(t)\ne\varnothing\) for almost every \(t\), then any optimizer selection \(x^*(t)\in X^*(t)\) on that full-measure set gives
\[ V(t)=V(0)+\int_0^t f_t(x^*(s),s)\,ds. \]
The absolute-value bound is uniform over choices: a separate integrable bound depending on \(x\) would not deliver this conclusion.
For this course, keep two failed implications straight:
- Smooth objective functions do not imply a differentiable value function.
- A differentiable value function does not imply an envelope equality when the equality’s other premises fail.
Assumption audit
- Existence: nonempty compact feasibility and a continuous objective give \(A(q)\ne\varnothing\).
- Continuity of value: a continuous objective and a nonempty, compact-valued, continuous feasibility correspondence give continuous \(V\).
- Upper stability of choices: the same maximum-theorem assumptions make \(A\) upper hemicontinuous.
- Theorem 4.7.1: a compact fixed choice set, joint continuity, continuous parameter derivatives, and a unique optimizer at \(\theta_0\) give the envelope derivative there.
- Theorem 4.7.2: KKT regularity, differentiability of the optimizer, and the active parameterized constraints give the multiplier correction.
- Arbitrary-set equality: an optimizer, the partial derivative \(f_t\), and differentiability of \(V\) give \(V'=f_t\).
Proof blueprints
Applying the maximum theorem
- State \(U\), the objective \(u(q,x)\), and \(\Psi(q)\).
- Check that every feasible value is nonempty and compact.
- Check both UHC and LHC of \(\Psi\).
- Check joint continuity of \(u\).
- State separately: existence, continuity of \(V\), and UHC of \(A\).
- Add uniqueness only if strict curvature and convex feasibility justify it.
Applying an envelope theorem
- Define the value function before differentiating it.
- Identify whether the choice is unconstrained, constrained, or arbitrary.
- State exactly what makes \(V\) differentiable. Do not assume it from a smooth objective alone.
- For Theorem 4.7.1, check compactness, uniqueness at the parameter of interest, and continuity of the parameter derivative. A differentiable optimizer is not required.
- For Theorem 4.7.2, check that the optimizer is differentiable at the parameter, write the Lagrangian convention, and retain the multiplier terms.
Exit tickets
- Which maximum-theorem hypothesis prevents a better limiting competitor from disappearing when the parameter changes?
- Give a one-sentence reason that continuity of \(V\) does not imply differentiability of \(V\).
- In Theorem 4.7.1, what replaces the chain-rule argument involving an optimizer derivative?
- State the missing premise in the false sentence: “\(f_t\) exists, so \(V'(t)=f_t(x^*,t)\).”
Mastery checklist
Sources and further reading
- John Nachbar, Theorem of the Maximum, especially the separation of nonempty argmax/UHC from continuity of the value.
- John Nachbar, The Envelope Theorem, for the differentiable optimizer and KKT formulations.
- Paul Milgrom and Ilya Segal, “Envelope Theorems for Arbitrary Choice Sets” (2002), clearly labeled here as the arbitrary-choice-set extension.