Solutions to Bogachev’s Measure Theory, Volume 1
Solutions to the exercises of V.I. Bogachev’s Measure Theory, Volume 1 (Springer, 2007) — exercises 1.12.47–160, 2.12.25–117, 3.10.29–125, 4.7.42–154, and 5.8.37–142. Exercises marked \(\circ\) in the book are the basic ones. The book itself is filed at Measure Theory (Bogachev).
Constructions and Extensions of Measures
Exercises 1.12.47–1.12.53
(\(\circ\)) Suppose we are given a family of open sets in \(\mathbb{R}^n\). Show that this family contains an at most countable subfamily with the same union.
The countable base of balls selects the subfamily. Let \(\mathcal{B}=\{U(q,r):q\in\mathbb{Q}^n,\ r\in\mathbb{Q},\ r>0\}\), let \(\mathcal{B}_0\) consist of those \(B\in\mathcal{B}\) contained in at least one member of the given family \(\{W_\alpha\}_{\alpha\in A}\), and for each \(B\in\mathcal{B}_0\) fix \(\alpha(B)\) with \(B\subset W_{\alpha(B)}\). Then \(\{W_{\alpha(B)}\}_{B\in\mathcal{B}_0}\) is at most countable, and its union is \(W:=\bigcup_{\alpha\in A}W_\alpha\): the inclusion \(\subset\) is trivial, while for \(x\in W_\alpha\) we pick \(\varepsilon>0\) with \(U(x,\varepsilon)\subset W_\alpha\), then \(q\in\mathbb{Q}^n\) with \(|x-q|<\varepsilon/3\) and rational \(r\in(\varepsilon/3,\varepsilon/2)\), so that
\begin{equation*} x\in U(q,r)\subset U(x,\varepsilon)\subset W_\alpha \end{equation*}
(indeed \(|y-q|<r\) forces \(|y-x|<r+\varepsilon/3<\varepsilon\)); hence \(B:=U(q,r)\in\mathcal{B}_0\) and \(x\in B\subset W_{\alpha(B)}\).
(\(\circ\)) Let \(W\) be a nonempty open set in \(\mathbb{R}^n\). Prove that \(W\) is the union of an at most countable collection of open cubes whose edges are parallel to the coordinate axes and have lengths of the form \(p2^{-q}\), where \(p,q\in\mathbb{N}\), and whose centers have coordinates of the form \(m2^{-k}\), where \(m\in\mathbb{Z}\), \(k\in\mathbb{N}\).
Take \(\mathcal{Q}\) to be the family of all cubes of the prescribed form contained in \(W\); it is at most countable, such a cube being determined by its centre in \(\{m2^{-k}\}^n\) and its edge in \(\{p2^{-q}\}\), both countable sets.
Only \(W\subset\bigcup\mathcal{Q}\) needs proof: given \(x\in W\), pick \(\varepsilon>0\) with \(U(x,\varepsilon)\subset W\), then \(k\in\mathbb{N}\) with \(3\sqrt{n}\,2^{-k}<\varepsilon\), put \(m_i=\lfloor2^kx_i\rfloor\) and
\begin{equation*} Q=\prod_{i=1}^{n}\bigl((m_i-1)2^{-k},\ (m_i+2)2^{-k}\bigr). \end{equation*}
This \(Q\) is admissible, its edge being \(3\cdot2^{-k}\) and its centre coordinates \((2m_i+1)2^{-(k+1)}\); it contains \(x\), since \(m_i2^{-k}\le x_i<(m_i+1)2^{-k}\); and \(\operatorname{diam}Q=3\sqrt{n}\,2^{-k}<\varepsilon\) gives \(Q\subset U(x,\varepsilon)\subset W\). Hence \(Q\in\mathcal{Q}\) and \(x\in Q\).
(\(\circ\)) Let \(\mu\) be a nonnegative measure on a ring \(\mathcal{R}\). Prove that the class of all sets \(Z\in\mathcal{R}\) of measure zero is a ring.
Yes: \(\mathcal{N}=\{Z\in\mathcal{R}:\mu(Z)=0\}\) is a ring. Additivity with \(\mu\ge0\) makes \(\mu\) monotone on \(\mathcal{R}\), since \(\mu(B)=\mu(A)+\mu(B\setminus A)\ge\mu(A)\) for \(A\subset B\). Now \(\varnothing\in\mathcal{N}\), and for \(Z_1,Z_2\in\mathcal{N}\) the sets \(Z_1\cap Z_2\), \(Z_1\setminus Z_2\), \(Z_1\cup Z_2\) lie in \(\mathcal{R}\) with
\begin{equation*} \begin{aligned} \mu(Z_1\cap Z_2)&\le\mu(Z_1)=0, &\qquad \mu(Z_1\setminus Z_2)&\le\mu(Z_1)=0,\\ \mu(Z_1\cup Z_2)&=\mu(Z_1)+\mu(Z_2\setminus Z_1)&&\le\mu(Z_1)+\mu(Z_2)=0, \end{aligned} \end{equation*}
the last line by additivity on the disjoint decomposition \(Z_1\cup Z_2=Z_1\cup(Z_2\setminus Z_1)\) and monotonicity for \(Z_2\setminus Z_1\subset Z_2\). As \(\mu\ge0\), all three values vanish, which is Definition 1.2.13(i) for \(\mathcal{N}\).
(\(\circ\)) Let \(\mu\) be an arbitrary finite Borel measure on a closed interval \(I\). Show that there exists a first category set \(E\) (i.e., a countable union of nowhere dense sets) such that \(\mu(I\setminus E)=0\).
Take \(E=\bigcup_{n\ge1}(I\setminus G_n)\), where \(G_n\) is the dense open set of \(\mu\)-measure at most \(2^{-n}\) built below. Here \(I=[a,b]\) with \(a<b\) (for a one-point interval the claim fails for the Dirac measure), and we may assume \(\mu\ge0\), the general case following by applying the result to \(|\mu|\) (Definition 3.1.4) and using \(|\mu(B)|\le|\mu|(B)\).
The atoms \(A=\{x:\mu(\{x\})>0\}=\bigcup_m\{x:\mu(\{x\})\ge1/m\}\) are at most countable, each of the sets in the union being finite by \(\mu(I)<\infty\). Every member of the countable base \(\{(p,q)\cap I:p<q\ \text{rational}\}\) of \(I\) is uncountable, hence meets \(I\setminus A\); picking one point in each gives a countable dense \(S=\{s_j\}\subset I\setminus A\).
Fix \(n\). Since \((s_j-\delta,s_j+\delta)\cap I\downarrow\{s_j\}\) as \(\delta\downarrow0\), continuity from above (Proposition 1.3.3, legitimate as \(\mu(I)<\infty\)) and \(\mu(\{s_j\})=0\) let us fix \(\delta_{n,j}>0\) with
\begin{equation*} \mu(U_{n,j}\cap I)<2^{-j-n},\qquad U_{n,j}:=(s_j-\delta_{n,j},\,s_j+\delta_{n,j}). \end{equation*}
Set \(G_n=\bigcup_{j\ge1}U_{n,j}\), so \(\mu(G_n\cap I)\le\sum_j2^{-j-n}=2^{-n}\) by countable subadditivity, while \(G_n\cap I\) is open in \(I\) and dense (it contains \(S\)), making the closed set \(I\setminus G_n\) nowhere dense. Hence \(E\) is of the first category and
\begin{equation*} \begin{aligned} 0\le\mu(I\setminus E)&=\mu\Bigl(\bigcap_{n\ge1}(G_n\cap I)\Bigr)\\ &\le 2^{-n}\quad\text{for all }n, \end{aligned} \end{equation*}
i.e. \(\mu(I\setminus E)=0\).
(\(\circ\)) Let \(\mathcal{S}\) be some collection of subsets of a set \(X\) such that it is closed with respect to finite unions and finite intersections and contains the empty set (for example, the class of all closed sets or the class of all open sets in \([0,1]\)). Show that the class of all sets of the form \(A\setminus B\), \(A,B\in\mathcal{S}\), \(B\subset A\), is a semiring, and the class of all sets of the form \((A_1\setminus B_1)\cup\cdots\cup(A_n\setminus B_n)\), \(A_i,B_i\in\mathcal{S}\), \(B_i\subset A_i\), \(n\in\mathbb{N}\), is the ring generated by \(\mathcal{S}\).
Both assertions come from the difference identity, valid whenever \(B\subset A\) and \(D\subset C\),
\begin{equation*} \begin{aligned} (A\setminus B)\setminus(C\setminus D) &=\Bigl(A\setminus\bigl[B\cup(A\cap C)\bigr]\Bigr)\\ &\qquad\cup\ \bigl((A\cap D)\setminus(B\cap D)\bigr), \end{aligned} \end{equation*}
whose two right-hand sets are disjoint (the first lies in \(A\setminus C\), the second in \(D\subset C\)). Put \(\mathcal{P}=\{A\setminus B:A,B\in\mathcal{S},\ B\subset A\}\) and let \(\mathcal{R}\) be the class of finite unions of members of \(\mathcal{P}\), i.e. the class described in the statement; \(\mathcal{S}\subset\mathcal{P}\), since \(A=A\setminus\varnothing\).
(i) \(\mathcal{P}\) satisfies Definition 1.2.13(ii). It contains \(\varnothing=\varnothing\setminus\varnothing\) and is closed under finite intersections, since
\begin{equation*} (A\setminus B)\cap(C\setminus D)=(A\cap C)\setminus\bigl[(B\cup D)\cap A\cap C\bigr], \end{equation*}
where both entries lie in \(\mathcal{S}\) and the subtracted one is contained in \(A\cap C\). And the identity above exhibits \(P\setminus Q\) for arbitrary \(P,Q\in\mathcal{P}\) — in particular for \(Q\subset P\) — as a disjoint union of two members of \(\mathcal{P}\), because \(B\cup(A\cap C)\subset A\) and \(B\cap D\subset A\cap D\) all belong to \(\mathcal{S}\).
(ii) Hence Lemma 1.2.14 applies: \(\mathcal{R}\) is a ring, the smallest one containing \(\mathcal{P}\). It contains \(\mathcal{S}\); conversely any ring \(\mathcal{R}^{\prime}\supset\mathcal{S}\) contains every \(A\setminus B\), hence \(\mathcal{P}\), hence, being closed under finite unions, all of \(\mathcal{R}\). So \(\mathcal{R}\) is the ring generated by \(\mathcal{S}\).
(\(\circ\)) Let \(m\) be an additive set function on a ring of sets \(\mathcal{R}\). Prove the following Poincare formula for all \(A_1,\dots,A_n\in\mathcal{R}\):
\begin{equation*} m\Bigl(\bigcup_{i=1}^{n}A_i\Bigr)=\sum_{i=1}^{n}m(A_i)-\sum_{1\le i<j\le n}m(A_i\cap A_j) +\sum_{1\le i<j<k\le n}m(A_i\cap A_j\cap A_k)-\cdots+(-1)^{n+1}m\Bigl(\bigcap_{i=1}^{n}A_i\Bigr). \end{equation*}
Induction on \(n\), driven by the two-set case. Writing \(A_J=\bigcap_{i\in J}A_i\) for \(\varnothing\neq J\subset\{1,\dots,n\}\), the assertion is
\begin{equation*} m\Bigl(\bigcup_{i=1}^{n}A_i\Bigr) =\sum_{\varnothing\neq J\subset\{1,\dots,n\}}(-1)^{|J|+1}m(A_J), \tag{*} \end{equation*}
the terms with \(|J|=r\) being exactly those of the \(r\)-th sum in the statement.
For \(E,F\in\mathcal{R}\), additivity on the disjoint decompositions \(E\cup F=(E\setminus F)\cup(E\cap F)\cup(F\setminus E)\), \(E=(E\setminus F)\cup(E\cap F)\) and \(F=(F\setminus E)\cup(E\cap F)\) gives, \(m\) being real-valued so that subtraction is legitimate,
\begin{equation*} m(E\cup F)=m(E)+m(F)-m(E\cap F), \tag{1} \end{equation*}
which is () for \(n=2\); the case \(n=1\) is trivial. Assume () for \(n\) sets and put \(B=\bigcup_{i\le n}A_i\). Then (1) with \(E=B\), \(F=A_{n+1}\), the distributivity \(B\cap A_{n+1}=\bigcup_{i\le n}(A_i\cap A_{n+1})\), and the induction hypothesis applied both to \(A_1,\dots,A_n\) and to \(A_1\cap A_{n+1},\dots,A_n\cap A_{n+1}\) (where \(\bigcap_{i\in J}(A_i\cap A_{n+1})=A_{J\cup\{n+1\}}\)) yield
\begin{equation*} \begin{aligned} m\Bigl(\bigcup_{i=1}^{n+1}A_i\Bigr) &=\sum_{\varnothing\neq J\subset\{1,\dots,n\}}(-1)^{|J|+1}m(A_J)+m(A_{n+1})\\ &\qquad+\sum_{\varnothing\neq J\subset\{1,\dots,n\}}(-1)^{|J|}m\bigl(A_{J\cup\{n+1\}}\bigr). \end{aligned} \end{equation*}
Each nonempty \(J^{\prime}\subset\{1,\dots,n+1\}\) occurs exactly once on the right with the coefficient \((-1)^{|J^{\prime}|+1}\): in the first sum when \(n+1\notin J^{\prime}\); as the middle term when \(J^{\prime}=\{n+1\}\), where \(|J^{\prime}|=1\); and in the third sum otherwise, where \(J^{\prime}=J\cup\{n+1\}\) gives \((-1)^{|J|}=(-1)^{|J^{\prime}|-1}\). This is (*) for \(n+1\).
(\(\circ\)) Let \(\mathcal{R}_1\) and \(\mathcal{R}_2\) be two semirings of sets. Prove that
\begin{equation*} \mathcal{R}_1\times\mathcal{R}_2=\{R_1\times R_2:\ R_1\in\mathcal{R}_1,\ R_2\in\mathcal{R}_2\} \end{equation*}
is a semiring. Show that \(\mathcal{R}_1\times\mathcal{R}_2\) may not be a ring even if \(\mathcal{R}_1\) and \(\mathcal{R}_2\) are algebras.
The three conditions of Definition 1.2.13(ii) hold, and the diagonal pair \(\{1\}\times\{1\}\), \(\{2\}\times\{2\}\) defeats the ring property.
(i) \(\varnothing=\varnothing\times\varnothing\in\mathcal{R}_1\times\mathcal{R}_2\).
(ii) \((A_1\times A_2)\cap(B_1\times B_2)=(A_1\cap B_1)\times(A_2\cap B_2)\), and semirings admit finite intersections.
(iii) Let \(Q=B_1\times B_2\subset P=A_1\times A_2\). If \(Q=\varnothing\), then \(P\setminus Q=P\); otherwise \(B_1,B_2\neq\varnothing\), so projecting the inclusion onto the factors forces \(B_1\subset A_1\), \(B_2\subset A_2\), and splitting \((x,y)\in P\setminus Q\) by whether \(x\in B_1\) gives
\begin{equation*} \begin{aligned} P\setminus Q=\bigl[(A_1\setminus B_1)\times A_2\bigr]\\ \cup\ \bigl[B_1\times(A_2\setminus B_2)\bigr], \end{aligned} \end{equation*}
a disjoint union, the first coordinates lying in \(A_1\setminus B_1\) and in \(B_1\). Writing \(A_1\setminus B_1=\bigcup_{i\le k}C_i\) and \(A_2\setminus B_2=\bigcup_{j\le l}D_j\) as finite disjoint unions inside the semirings \(\mathcal{R}_1\), \(\mathcal{R}_2\), we obtain
\begin{equation*} P\setminus Q=\bigcup_{i=1}^{k}(C_i\times A_2)\ \cup\ \bigcup_{j=1}^{l}(B_1\times D_j), \end{equation*}
which is a finite pairwise disjoint union of members of \(\mathcal{R}_1\times\mathcal{R}_2\) (across the two groups because \(C_i\cap B_1=\varnothing\)).
Not a ring: for the algebras \(\mathcal{R}_1=\mathcal{R}_2=2^{\{1,2\}}\), were \(\{(1,1),(2,2)\}\) equal to some \(R_1\times R_2\), its projections would force \(R_1=R_2=\{1,2\}\), a set of four points.
Exercises 1.12.54–1.12.60
(\(\circ\)) Let \(\mathcal{F}\) be some collection of sets in a space \(X\). Prove that every set \(A\) in the \(\sigma\)-algebra \(\sigma(\mathcal{F})\) generated by \(\mathcal{F}\) is contained in the \(\sigma\)-algebra generated by an at most countable subcollection \(\{F_n\}\subset\mathcal{F}\).
The class
\begin{equation*} \begin{aligned} \mathcal{G}:=\bigcup\bigl\{\sigma(\mathcal{C}):\ &\mathcal{C}\subset\mathcal{F}\\ &\text{at most countable}\bigr\} \end{aligned} \end{equation*}
is a \(\sigma\)-algebra containing \(\mathcal{F}\), so \(\sigma(\mathcal{F})\subset\mathcal{G}\), which is the assertion.
Indeed \(X\in\sigma(\emptyset)=\{\emptyset,X\}\subset\mathcal{G}\), and \(\mathcal{G}\) is closed under complements because each \(\sigma(\mathcal{C})\) is. For unions, let \(A_n\in\sigma(\mathcal{C}_n)\) with every \(\mathcal{C}_n\subset\mathcal{F}\) at most countable; then \(\mathcal{C}:=\bigcup_n\mathcal{C}_n\subset\mathcal{F}\) is again at most countable and minimality (Definition 1.2.5) gives \(\sigma(\mathcal{C}_n)\subset\sigma(\mathcal{C})\), whence
\begin{equation*} \bigcup_{n=1}^{\infty}A_n\in\sigma(\mathcal{C})\subset\mathcal{G}. \end{equation*}
Finally \(F\in\sigma(\{F\})\subset\mathcal{G}\) for each \(F\in\mathcal{F}\).
(\(\circ\)) (Brown, Freilich) The aim of this exercise is to show that Proposition 1.2.6 may be false if a \(\sigma\)-algebra is defined in the broader sense mentioned in \(\S 1.2\). Suppose we are given a set \(X\) and a collection \(\mathcal{S}\) of its subsets such that the union of all sets in \(\mathcal{S}\) is \(Y\subset X\). Prove that the following conditions are equivalent: (i) \(Y\) is an at most countable union of sets in \(\mathcal{S}\); (ii) there exists a smallest family of sets \(\mathcal{A}\) with the following properties: \(\mathcal{A}\) is a \(\sigma\)-algebra on some subset \(Z\subset X\) (i.e., \(Z\) is the unit of this \(\sigma\)-algebra) and \(\mathcal{S}\subset\mathcal{A}\), where a smallest family is a family that is contained in every other family with the stated properties. Consider the example where \(X=[0,1]\), \(Y=[0,1/2]\), \(\mathcal{S}\) is the class of all at most countable subsets of \(Y\).
The smallest family in (ii), when it exists, is \(\sigma_Y(\mathcal{S})\), the \(\sigma\)-algebra generated by \(\mathcal{S}\) on \(Y\). Call \(\mathcal{A}\) admissible if it is a \(\sigma\)-algebra on some \(Z\subset X\) with \(\mathcal{S}\subset\mathcal{A}\); its unit is \(Z=\bigcup_{A\in\mathcal{A}}A\), and \(\mathcal{S}\subset\mathcal{A}\) forces \(Y\subset Z\). The printed implication (ii) \(\Rightarrow\) (i) needs \(X\setminus Y\neq\emptyset\), which we assume; for \(Y=X\) every admissible family has unit \(Y\), so (ii) holds whatever (i) says.
(i) \(\Rightarrow\) (ii). Let \(Y=\bigcup_n S_n\) with \(S_n\in\mathcal{S}\) and let \(\mathcal{A}\) be admissible with unit \(Z\). Then \(Y\in\mathcal{A}\), being a countable union of members of \(\mathcal{S}\subset\mathcal{A}\), so the trace \(\mathcal{A}_Y:=\{A\in\mathcal{A}:A\subset Y\}\) is a \(\sigma\)-algebra on \(Y\) (for complements, \(Y\setminus A=Y\cap(Z\setminus A)\)) containing \(\mathcal{S}\). Minimality on \(Y\) gives
\begin{equation*} \sigma_Y(\mathcal{S})\subset\mathcal{A}_Y\subset\mathcal{A}, \end{equation*}
and \(\sigma_Y(\mathcal{S})\) is itself admissible, hence the smallest such family.
(ii) \(\Rightarrow\) (i), contrapositive. Suppose \(Y\) is not an at most countable union of members of \(\mathcal{S}\), and put
\begin{equation*} \begin{aligned} \mathcal{P}:=\Bigl\{A\subset Y:\ &A\subset\bigcup_{n=1}^{\infty}S_n\\ &\text{for some}\ S_n\in\mathcal{S}\Bigr\}. \end{aligned} \end{equation*}
Then \(\mathcal{S}\subset\mathcal{P}\), \(\mathcal{P}\) is hereditary and stable under countable unions, and \(Y\notin\mathcal{P}\): members of \(\mathcal{S}\) lie in \(Y\), so a countable cover of \(Y\) by them would be an exact decomposition. Fix \(z\in X\setminus Y\) and put \(W:=Y\cup\{z\}\),
\begin{equation*} \mathcal{E}:=\{E\subset W:\ E\in\mathcal{P}\ \text{or}\ W\setminus E\in\mathcal{P}\}. \end{equation*}
This \(\mathcal{E}\) is a \(\sigma\)-algebra on \(W\): the definition is symmetric in \(E\) and \(W\setminus E\), and for \(E=\bigcup_nE_n\) either every \(E_n\in\mathcal{P}\), whence \(E\in\mathcal{P}\) by countable stability, or some \(W\setminus E_{n_0}\in\mathcal{P}\), whence \(W\setminus E\subset W\setminus E_{n_0}\) lies in \(\mathcal{P}\) by heredity. It is admissible, and \(Y\notin\mathcal{E}\), because \(Y\notin\mathcal{P}\) while \(W\setminus Y=\{z\}\not\subset Y\). A smallest admissible family \(\mathcal{A}\) with unit \(Z\) would satisfy \(Y\subset Z\) and \(\mathcal{A}\subset\sigma_Y(\mathcal{S})\), all of whose members lie in \(Y\), so \(Z\subset Y\); hence \(Y=Z\in\mathcal{A}\subset\mathcal{E}\), a contradiction.
In the example \(X=[0,1]\), \(Y=[0,1/2]\), \(\mathcal{S}=\) the at most countable subsets of \(Y\), we have \(X\setminus Y\neq\emptyset\) while the uncountable \(Y\) is no countable union of countable sets. So (i) fails, and no smallest \(\sigma\)-algebra in the wide sense contains \(\mathcal{S}\).
(Broughton, Huff) Suppose we are given a sequence of \(\sigma\)-algebras \(\mathcal{A}_n\) in a space \(X\) such that \(\mathcal{A}_n\) is strictly contained in \(\mathcal{A}_{n+1}\) for each \(n\). Prove that \(\bigcup_{n=1}^{\infty}\mathcal{A}_n\) is not a \(\sigma\)-algebra.
We exhibit pairwise disjoint sets in \(\mathcal{A}:=\bigcup_{n}\mathcal{A}_n\) whose union is not in \(\mathcal{A}\); since \(\mathcal{A}\) is an algebra (the \(\mathcal{A}_n\) increase), this is the only property that can fail. Write \(D\cap\mathcal{B}:=\{D\cap B:B\in\mathcal{B}\}\) for the trace, a \(\sigma\)-algebra on \(D\) whenever \(\mathcal{B}\) is one on the ambient space.
Splitting lemma. If \(\mathcal{B}\subset\mathcal{C}\) are \(\sigma\)-algebras on \(T\), \(D\in\mathcal{B}\), and both \(D\cap\mathcal{C}=D\cap\mathcal{B}\) and \((T\setminus D)\cap\mathcal{C}=(T\setminus D)\cap\mathcal{B}\), then \(\mathcal{C}=\mathcal{B}\): for \(C\in\mathcal{C}\) we have \(C\cap D=D\cap B\) with \(B\in\mathcal{B}\), so \(C\cap D\in\mathcal{B}\) because \(D\in\mathcal{B}\), likewise \(C\setminus D\in\mathcal{B}\), and \(C\) is their union.
(1) Disjoint witnesses. Call a nonempty \(E\in\mathcal{A}_m\) unstable from \(m\) if \(E\cap\mathcal{A}_n\subsetneq E\cap\mathcal{A}_{n+1}\) for infinitely many \(n\geq m\); \(E_1:=X\) is unstable from \(j_1:=1\), since \(X\cap\mathcal{A}_n=\mathcal{A}_n\) increases strictly. Given \(E_k\in\mathcal{A}_{j_k}\) nonempty and unstable from \(j_k\), choose \(n\geq j_k\) with \(E_k\cap\mathcal{A}_n\subsetneq E_k\cap\mathcal{A}_{n+1}\), put \(j_{k+1}:=n+1\), and pick
\begin{equation*} D\in (E_k\cap\mathcal{A}_{j_{k+1}})\setminus(E_k\cap\mathcal{A}_{n}). \end{equation*}
Then \(D=E_k\cap A\subset E_k\) with \(A\in\mathcal{A}_{j_{k+1}}\), so \(D\in\mathcal{A}_{j_{k+1}}\) (as \(E_k\in\mathcal{A}_{j_k}\subset\mathcal{A}_{j_{k+1}}\)); also \(D\neq\emptyset,E_k\), both of which lie in \(E_k\cap\mathcal{A}_n\); and neither \(D\) nor \(E_k\setminus D\) lies in \(\mathcal{A}_n\), since either would put \(D=E_k\cap D\) into \(E_k\cap\mathcal{A}_n\), using \(E_k\in\mathcal{A}_n\). Apply the lemma inside \(T=E_k\) to \(\mathcal{B}=E_k\cap\mathcal{A}_l\subsetneq\mathcal{C}=E_k\cap\mathcal{A}_{l+1}\) for the infinitely many \(l\geq j_{k+1}\) where this inclusion is strict, with the set \(D\in E_k\cap\mathcal{A}_l\): using \(D\cap(E_k\cap\mathcal{A}_l)=D\cap\mathcal{A}_l\) and the same for \(E_k\setminus D\) (both are subsets of \(E_k\)), the lemma gives, for each such \(l\),
\begin{equation*} D\cap\mathcal{A}_l\subsetneq D\cap\mathcal{A}_{l+1} \quad\text{or}\quad (E_k\setminus D)\cap\mathcal{A}_l\subsetneq (E_k\setminus D)\cap\mathcal{A}_{l+1}. \end{equation*}
By pigeonhole one alternative holds for infinitely many such \(l\); let \(E_{k+1}\) be the corresponding set, unstable from \(j_{k+1}\), and \(F_k\) the other. Then \(E_{k+1}\subset E_k\) and \(F_k\) are nonempty members of \(\mathcal{A}_{j_{k+1}}\) with
\begin{equation*} F_k=E_k\setminus E_{k+1}\in\mathcal{A}_{j_{k+1}}\setminus\mathcal{A}_{j_{k+1}-1}, \end{equation*}
so in particular \(F_k\notin\mathcal{A}_{j_k}\). The \(F_k\) are pairwise disjoint: for \(l>k\) one has \(F_l\subset E_l\subset E_{k+1}\) while \(F_k\cap E_{k+1}=\emptyset\). Put \(\mathcal{C}_k:=\mathcal{A}_{j_k}\); as \(j_k\to\infty\), \(\bigcup_k\mathcal{C}_k=\mathcal{A}\) and \(F_k\in\mathcal{C}_{k+1}\setminus\mathcal{C}_k\).
(2) Transfer to a countable space. Define \(\pi:X\to\mathbb{N}_0\) by \(\pi\equiv k\) on \(F_k\) and \(\pi\equiv0\) on \(G:=X\setminus\bigcup_kF_k\) (well defined by disjointness). By Corollary 1.2.9 the classes \(\mathcal{A}’_n:=\{A\subset\mathbb{N}_0:\pi^{-1}(A)\in\mathcal{C}_n\}\) are \(\sigma\)-algebras on \(\mathbb{N}_0\), increasing in \(n\), and \(\pi^{-1}(\{k\})=F_k\) gives
\begin{equation*} \{k\}\in\mathcal{A}’_{k+1}\setminus\mathcal{A}’_{k}\qquad (k\geq 1). \tag{A} \end{equation*}
(3) Atoms. For \(n\geq1\) put \(B_n:=\bigcap\{A\in\mathcal{A}’_n:n\in A\}\), the smallest member of \(\mathcal{A}’_n\) containing \(n\): choosing for each \(j\notin B_n\) some \(A_j\in\mathcal{A}’_n\) with \(n\in A_j\), \(j\notin A_j\) exhibits \(B_n=\bigcap_{j\notin B_n}A_j\) as an at most countable intersection, so \(B_n\in\mathcal{A}’_n\) — the one place where countability of \(\mathbb{N}_0\) is used, and it is essential. By (A), \(\{j\}\in\mathcal{A}’_{j+1}\subset\mathcal{A}’_n\) for \(1\leq j<n\), so \(B_n\setminus\{j\}\in\mathcal{A}’_n\) contains \(n\) and minimality forces \(j\notin B_n\); hence
\begin{equation*} B_n\subset\{0\}\cup\{j\in\mathbb{N}:\ j\geq n\},\qquad B_n\neq\{n\}, \end{equation*}
the second relation because \(\{n\}\in\mathcal{A}’_n\) would contradict (A). Consequently some \(m\) has the property that \(B_n\) contains an element \(j>n\) for every \(n\geq m\):
(i) if \(\{0\}\in\mathcal{A}’_m\) for some \(m\) (automatic when \(G=\emptyset\), as then \(\pi^{-1}(\{0\})=\emptyset\)), then for \(n\geq m\) minimality excludes \(0\) too, so \(B_n\subset\{j\geq n\}\) and \(B_n\neq\{n\}\);
(ii) if \(\{0\}\notin\mathcal{A}’_n\) for all \(n\), take \(m=1\): \(B_n=\{0,n\}\) would give \(\{0\}=B_n\setminus\{n\}\in\mathcal{A}’_{n+1}\) by (A), excluded.
Moreover, if \(m’\in B_n\) with \(m’>n\), then \(B_n\in\mathcal{A}’_n\subset\mathcal{A}’_{m’}\) contains \(m’\), so minimality yields
\begin{equation*} B_{m’}\subset B_n. \tag{B} \end{equation*}
(4) The contradiction. Set \(n_1:=m\) and let \(n_{k+1}\in B_{n_k}\) with \(n_{k+1}>n_k\), available since \(n_k\geq m\); then \(n_1<n_2<\cdots\) and \(B_{n_1}\supset B_{n_2}\supset\cdots\) by (B). The sets \(F_{n_{2k}}\) lie in \(\mathcal{A}\), but their union
\begin{equation*} H:=\pi^{-1}(E)=\bigcup_{k=1}^{\infty}F_{n_{2k}},\qquad E:=\{n_2,n_4,n_6,\ldots\}, \end{equation*}
does not. Indeed \(H\in\mathcal{A}\) would give \(H\in\mathcal{C}_p\) for some \(p\) by cofinality, i.e. \(E\in\mathcal{A}’_p\); choosing \(k\) with \(n_{2k}\geq p\) we get \(E\in\mathcal{A}’_{n_{2k}}\), and the singletons of \(n_2,\ldots,n_{2k-2}<n_{2k}\) lie in \(\mathcal{A}’_{n_{2k}}\) by (A), so
\begin{equation*} \begin{aligned} E’&:=E\setminus\{n_2,\ldots,n_{2k-2}\}\\ &\ =\{n_{2k},n_{2k+2},\ldots\}\in\mathcal{A}’_{n_{2k}}. \end{aligned} \end{equation*}
Minimality then forces \(B_{n_{2k}}\subset E’\), whereas \(n_{2k+1}\in B_{n_{2k}}\) and \(n_{2k}<n_{2k+1}<n_{2k+2}\) give \(n_{2k+1}\notin E’\).
(\(\circ\)) Show that every set of positive Lebesgue measure contains a nonmeasurable subset.
A Vitali translate cuts one out: if \(\lambda_n^{\ast}(A)>0\), then \(A\cap(V+q)\) is nonmeasurable for some \(q\in\mathbb{Q}^n\), where \(V\) is a Vitali set (Example 1.7.7), i.e. a set meeting each class of the relation \(x\sim y\iff x-y\in\mathbb{Q}^n\) in exactly one point. We prove this for arbitrary \(A\) of positive outer measure, which covers every measurable \(A\) with \(\lambda_n(A)>0\). Note
\begin{equation*} \mathbb{R}^n=\bigcup_{q\in\mathbb{Q}^n}(V+q), \end{equation*}
disjointly: \(v_1+q_1=v_2+q_2\) forces \(v_1\sim v_2\), hence \(v_1=v_2\) and \(q_1=q_2\).
(i) Every measurable \(S\) contained in a translate \(V+q\) is null. Such an \(S\) meets each class at most once, so the translates \(S+r\), \(r\in\mathbb{Q}^n\), are pairwise disjoint by the same computation. If \(\lambda_n(S)>0\), then \(c:=\lambda_n(T)>0\) for \(T:=S\cap[-N,N]^n\) and some \(N\), by continuity from below; the sets \(T+r\) with \(r\in\mathbb{Q}^n\cap[0,1]^n\) are disjoint, measurable, all of measure \(c\) (translation invariance, Theorem 1.7.3(i)), and all inside \([-N,N+1]^n\), so countable additivity over this infinite index set gives
\begin{equation*} (2N+1)^n=\lambda_n\bigl([-N,N+1]^n\bigr)\geq\sum_{r}\lambda_n(T+r)=\infty, \end{equation*}
a contradiction.
(ii) Were every \(A\cap(V+q)\) measurable, (i) would make each null, and countable subadditivity of the outer measure would give
\begin{equation*} \lambda_n^{\ast}(A)\leq\sum_{q\in\mathbb{Q}^n}\lambda_n\bigl(A\cap(V+q)\bigr)=0, \end{equation*}
contradicting \(\lambda_n^{\ast}(A)>0\).
Prove that there exists a sequence of sets \(A_n\subset[0,1]\) such that for all \(n\) one has \(A_{n+1}\subset A_n\), \(\bigcap_{n=1}^{\infty}A_n=\emptyset\) and \(\lambda^{\ast}(A_n)=1\), where \(\lambda\) is Lebesgue measure.
Take \(A_n:=[0,1]\setminus E_n\), where \(E\subset[0,1]\) is a Vitali set for the relation \(x-y\in\mathbb{Q}\) (Example 1.7.7), \(\{r_n\}_{n\geq1}\) enumerates \(\mathbb{Q}\), \(r_0:=0\), and
\begin{equation*} E_n:=\bigl(E\cup(E+r_1)\cup\cdots\cup(E+r_n)\bigr)\cap[0,1]. \end{equation*}
The \(E_n\) increase, so \(A_{n+1}\subset A_n\), and \(\bigcup_nE_n=[0,1]\cap\bigcup_{q\in\mathbb{Q}}(E+q)=[0,1]\) gives \(\bigcap_nA_n=\emptyset\). It remains to show \(\lambda^{\ast}(A_n)=1\), i.e. that the inner measure \(\lambda_{\ast}(E_n)=1-\lambda^{\ast}(A_n)\) of \(\S1.5\) vanishes. For any \(C\subset[0,1]\),
\begin{equation*} \lambda_{\ast}( C)=\sup\{\lambda(B):\ B\subset C,\ B\ \text{measurable}\}, \tag{1} \end{equation*}
since a measurable \(B\subset C\) gives \(\lambda^{\ast}([0,1]\setminus C)\leq 1-\lambda(B)\), i.e. \(\lambda_{\ast}( C)\geq\lambda(B)\), while a measurable envelope \(H\subset[0,1]\) of \([0,1]\setminus C\) (Corollary 1.5.8), with \(\lambda(H)=\lambda^{\ast}([0,1]\setminus C)\), makes \(B:=[0,1]\setminus H\subset C\) measurable of measure \(\lambda_{\ast}( C)\).
So fix \(n\) and a measurable \(B\subset E_n\). Each of the \(n+1\) translates \(E+r_i\) meets a class \(x+\mathbb{Q}\) in exactly one point, so \(B\) meets it in at most \(n+1\) points; hence, for a finite \(Q\subset\mathbb{Q}\cap[0,1]\) of cardinality \(N\), the points \(x-r\) (\(r\in Q\)) being \(N\) distinct members of \(x+\mathbb{Q}\) and \(x\in B+r\) meaning \(x-r\in B\),
\begin{equation*} \sum_{r\in Q}\mathbf{1}_{B+r}(x)\leq n+1\qquad\text{for all}\ x\in\mathbb{R}. \end{equation*}
The sets \(B+r\) are measurable and contained in \([0,2]\), so integrating over \([0,2]\) and using translation invariance (Theorem 1.7.3(i)),
\begin{equation*} \begin{aligned} N\lambda(B)&=\sum_{r\in Q}\lambda(B+r)\\ &=\int_{[0,2]}\sum_{r\in Q}\mathbf{1}_{B+r}(x)\,dx\leq 2(n+1). \end{aligned} \end{equation*}
As \(N\) is arbitrary, \(\lambda(B)=0\); so \(\lambda_{\ast}(E_n)=0\) by (1) and \(\lambda^{\ast}(A_n)=1\).
Show that every nonempty perfect set contains a nonempty perfect subset of Lebesgue measure zero. In particular, every set of positive Lebesgue measure contains a measure zero compact set of cardinality of the continuum.
A Cantor scheme inside the perfect set \(P\) produces the subset. Since \(P\) is closed with no isolated points, \(P\cap U\) is infinite for every open \(U\) meeting \(P\) (repeat the choice of a point of \(P\cap U\) distinct from the finitely many already taken). Define closed intervals \(I_s\) with centres \(x_s\in P\), indexed by words \(s\in\{0,1\}^k\), so that (a) \(\lambda(I_s)\leq4^{-k}\) for \(|s|=k\), and (b) \(I_{s0},I_{s1}\) are disjoint subsets of \(\operatorname{int}I_s\): start with any \(x_{\varnothing}\in P\) and \(I_{\varnothing}=[x_{\varnothing}-1/2,x_{\varnothing}+1/2]\); given \(I_s\) with \(|s|=k\), pick distinct \(x_{s0},x_{s1}\in P\cap\operatorname{int}I_s\) (infinite, as \(x_s\) lies in this open set) and put \(I_{si}=[x_{si}-\rho,x_{si}+\rho]\) with
\begin{equation*} \begin{aligned} 0<\rho<\min\Bigl\{&\tfrac12|x_{s0}-x_{s1}|,\ \tfrac12\cdot4^{-(k+1)},\\ &\operatorname{dist}(x_{si},\mathbb{R}\setminus\operatorname{int}I_s),\ i=0,1\Bigr\}. \end{aligned} \end{equation*}
By (b) and induction on the first index where two words of length \(k\) differ, the \(2^k\) intervals of level \(k\) are pairwise disjoint. Put \(K_k:=\bigcup_{|s|=k}(I_s\cap P)\) and \(K:=\bigcap_kK_k\); each \(I_s\cap P\) is compact and contains \(x_s\), so the \(K_k\) are nonempty compact sets, decreasing by (b), and \(K\subset P\) is nonempty compact by the finite intersection property, with
\begin{equation*} \lambda(K)\leq\lambda(K_k)\leq\sum_{|s|=k}\lambda(I_s)\leq 2^{k}4^{-k}=2^{-k}\to0 . \end{equation*}
For \(\sigma\in\{0,1\}^{\mathbb{N}}\) the sets \(I_{\sigma|k}\cap P\) decrease with diameters tending to \(0\), so their intersection is a single point \(\varphi(\sigma)\in K\); if \(\sigma,\tau\) first differ at \(k\), then \(\varphi(\sigma),\varphi(\tau)\) lie in the disjoint \(I_{\sigma|k},I_{\tau|k}\), so \(\varphi\) is injective and \(K\subset\mathbb{R}\) has exactly the cardinality of the continuum. Finally \(K\) has no isolated points: given \(x\in K\) and \(\varepsilon>0\), \(x\) lies in a unique level-\(k\) interval \(I_{s(k)}\) for each \(k\), with \(s(k)\) nested; choosing \(k\) with \(4^{-k}<\varepsilon\) gives \(I_{s(k)}\subset(x-\varepsilon,x+\varepsilon)\), and the word \(t\) of length \(k+1\) extending \(s(k)\) with \(t\neq s(k+1)\) has \(K\cap I_t\neq\varnothing\) (it contains \(\varphi(\sigma)\) for any \(\sigma\) extending \(t\)) and \(I_t\cap I_{s(k+1)}=\varnothing\), so any \(y\in K\cap I_t\) lies in \(K\cap(x-\varepsilon,x+\varepsilon)\setminus\{x\}\).
For the second assertion, “positive measure” must be read as measurable of positive measure, a Vitali set having positive outer measure but only null measurable subsets (Exercise 1.12.58). Let \(A\) be measurable with \(\lambda(A)>0\); replacing \(A\) by \(A\cap[-N,N]\) we may assume \(A\) bounded. By Corollary 1.5.8 there is a Borel \(B\subset A\) with \(\lambda(B)=\lambda(A)>0\), and Theorem 1.4.8, applied to the finite Borel measure \(\lambda(\,\cdot\,\cap[-N,N])\), gives a compact \(F\subset B\) with \(\lambda(F)\geq\lambda(B)/2>0\). Put
\begin{equation*} P:=\bigl\{x:\ \lambda\bigl(F\cap(x-\delta,x+\delta)\bigr)>0\ \text{for all}\ \delta>0\bigr\}. \end{equation*}
Then (i) \(P\subset F\), since \(F\) is closed and a point off \(F\) has a neighbourhood missing it; (ii) \(P\) is closed, since \(\lambda(F\cap(x-\delta,x+\delta))=0\) puts every \(y\) with \(|y-x|<\delta/2\) outside \(P\); (iii) \(\lambda(F\setminus P)=0\), as \(F\setminus P\) is covered by the countably many rational-endpoint intervals \(J\) with \(\lambda(F\cap J)=0\). So \(P\) is compact with \(\lambda(P)=\lambda(F)>0\), in particular nonempty, and has no isolated points: \(P\cap(x-\delta,x+\delta)=\{x\}\) for some \(x\in P\) would, by (iii), force \(\lambda(P\cap(x-\delta,x+\delta))>0\). Hence \(P\) is perfect, and the first part gives a compact perfect \(K\subset P\subset A\) with \(\lambda(K)=0\) of cardinality of the continuum.
(\(\circ\)) Let \(C\) be the Cantor set in \([0,1]\). Show that
\begin{equation*} C+C:=\{c_1+c_2:\ c_1,c_2\in C\}=[0,2],\qquad C-C:=\{c_1-c_2:\ c_1,c_2\in C\}=[-1,1]. \end{equation*}
Both equalities follow by splitting ternary digits, using the description
\begin{equation*} C=\Bigl\{\sum_{j=1}^{\infty}a_j3^{-j}:\ a_j\in\{0,2\}\ \text{for all}\ j\Bigr\} \tag{1} \end{equation*}
of Example 1.7.5.
The sum. \(C+C\subset[0,2]\) because \(C\subset[0,1]\). Conversely, given \(x\in[0,2]\) expand \(x/2=\sum_{j\geq1}b_j3^{-j}\) with \(b_j\in\{0,1,2\}\) and set
\begin{equation*} a_j:=\begin{cases}0,& b_j=0,\\ 2,& b_j\in\{1,2\},\end{cases} \qquad c_j:=\begin{cases}0,& b_j\in\{0,1\},\\ 2,& b_j=2,\end{cases} \end{equation*}
so that \(a_j,c_j\in\{0,2\}\) and \(a_j+c_j=2b_j\) in each of the three cases. By (1) the numbers \(u:=\sum_ja_j3^{-j}\) and \(v:=\sum_jc_j3^{-j}\) lie in \(C\), and
\begin{equation*} u+v=\sum_{j=1}^{\infty}(a_j+c_j)3^{-j}=2\sum_{j=1}^{\infty}b_j3^{-j}=x . \end{equation*}
The difference. \(1-C=C\): for \(x=\sum_ja_j3^{-j}\) as in (1), \(\sum_{j\geq1}2\cdot3^{-j}=1\) gives \(1-x=\sum_j(2-a_j)3^{-j}\) with \(2-a_j\in\{0,2\}\), so \(1-x\in C\), and \(x\mapsto1-x\) is an involution. Substituting \(c_2=1-c_2’\), a bijection of \(C\) onto itself,
\begin{equation*} C-C=\{c_1+c_2’-1:\ c_1,c_2’\in C\}=(C+C)-1=[-1,1]. \end{equation*}
Exercises 1.12.61–1.12.67
(\(\circ\)) Give an example of two closed sets \(A, B \subset \mathbb{R}\) of Lebesgue measure zero such that the set \(A + B := \{a + b : a \in A, b \in B\}\) is \(\mathbb{R}\).
Take the Cantor set \(C\subset[0,1]\) and
\begin{equation*} A := C, \qquad B := \bigcup_{n \in \mathbb{Z}} (C + n). \end{equation*}
Both are closed: \(C\) is, and \(C+n\subset[n,n+1]\) makes the family \(\{C+n\}_{n\in\mathbb{Z}}\) locally finite, so a sequence in \(B\) converging to \(x\) eventually lies in \([x-1,x+1]\), which meets only finitely many \(C+n\); some closed \(C+n_0\) then contains a subsequence, and \(x\in C+n_0\subset B\). Both are null: \(\lambda( C)=0\) (Example 1.7.5), so \(\lambda(C+n)=0\) by translation invariance and \(\lambda(B)=0\) by countable additivity. Finally \(A+B=\mathbb{R}\): given \(x\in\mathbb{R}\), pick \(n\in\mathbb{Z}\) with \(x-n\in[0,2]\) and write \(x-n=u+v\) with \(u,v\in C\), possible since \(C+C=[0,2]\) by Exercise 1.12.60; then
\begin{equation*} x = u + (v + n) \in A + B . \end{equation*}
(\(\circ\)) (Steinhaus) Let \(A\) be a set of positive Lebesgue measure on the real line. Show that the set \(A - A := \{a_1 - a_2 : a_1, a_2 \in A\}\) contains some interval. Prove an analogous assertion for \(\mathbb{R}^n\) (obtained in Rademacher).
\(A-A\) contains a ball \(U(0,\varepsilon)\) about the origin; we argue at once in \(\mathbb{R}^n\), the case \(n=1\) giving the interval \((-\varepsilon,\varepsilon)\). By inner regularity (Theorem 1.4.8) pick a compact \(K\subset A\) with \(\lambda(K)>0\), after replacing \(A\) by \(A\cap B\) for a closed ball \(B\) with \(\lambda(A\cap B)>0\) if \(\lambda(A)=\infty\); by outer regularity pick an open \(U\supset K\) with \(\lambda(U)<2\lambda(K)\). Then \(\lambda(U)<\infty\), so \(\mathbb{R}^n\setminus U\) is a nonempty closed set disjoint from \(K\) and
\begin{equation*} \varepsilon:=\operatorname{dist}(K,\mathbb{R}^n\setminus U)>0, \end{equation*}
the continuous function \(x\mapsto\operatorname{dist}(x,\mathbb{R}^n\setminus U)\), strictly positive on \(K\), attaining its minimum on the compact set \(K\). For \(|h|<\varepsilon\) every \(x\in K\) then has \(x+h\in U\), so \(K\cup(K+h)\subset U\); were \(K\cap(K+h)=\varnothing\), additivity and translation invariance would give
\begin{equation*} \lambda(U) \ge \lambda(K) + \lambda(K+h) = 2\lambda(K) > \lambda(U). \end{equation*}
Hence \(K\cap(K+h)\neq\varnothing\), i.e. \(h=k_1-k_2\in K-K\subset A-A\) for some \(k_1,k_2\in K\).
(P.L. Ulyanov) Let \(E \subset [0,1]\) be a measurable set of positive measure. (i) Prove that for every sequence \(\{h_n\}\) converging to zero and every \(\varepsilon > 0\), there exist a measurable set \(E_\varepsilon \subset E\) and a subsequence \(\{h_{n_k}\}\) such that \(\lambda(E_\varepsilon) > \lambda(E) - \varepsilon\) and for all \(x \in E_\varepsilon\) we have \(x + h_{n_k} \in E\), \(x - h_{n_k} \in E\) for all \(k\).
(ii) Prove that there exist a measurable set \(E_0 \subset E\) and a sequence of numbers \(h_n > 0\) converging to zero such that \(\lambda(E_0) = \lambda(E)\) and for every \(x \in E_0\), we have \(x + h_n \in E\) for all \(n \ge n(x)\).
Take \(E_\varepsilon:=E\cap\bigcap_{k}\bigl((E+h_{n_k})\cap(E-h_{n_k})\bigr)\) in (i), for the subsequence chosen below; (ii) follows by diagonalising (i). Here \(\triangle\) is symmetric difference, and we use \(A\triangle C\subset(A\triangle B)\cup(B\triangle C)\) and \((A\triangle B)+h=(A+h)\triangle(B+h)\).
Continuity of translation: \(\lambda\bigl(S\triangle(S+h)\bigr)\to0\) as \(h\to0\), for measurable \(S\) with \(\lambda(S)<\infty\). Given \(\delta>0\), outer regularity provides an open \(U\supset S\) with \(\lambda(U\setminus S)<\delta\), \(\lambda(U)<\infty\); writing \(U=\bigcup_jI_j\) as a disjoint union of open intervals with \(\sum_{j>N}\lambda(I_j)<\delta\) and putting \(V:=I_1\cup\cdots\cup I_N\), we get \(\lambda(S\triangle V)<2\delta\), while \(\lambda(I\triangle(I+h))\le2|h|\) for a bounded interval gives \(\lambda(V\triangle(V+h))\le2N|h|\). Hence
\begin{equation*} \begin{aligned} \lambda\bigl(S\triangle(S+h)\bigr)&\le 2\lambda(S\triangle V)\\ &\quad+\lambda\bigl(V\triangle(V+h)\bigr)<4\delta+2N|h| , \end{aligned} \end{equation*}
so \(\limsup_{h\to0}\lambda(S\triangle(S+h))\le4\delta\) for every \(\delta>0\).
(i) Applying this to \(S=E\), choose \(n_1<n_2<\cdots\) with
\begin{equation*} \begin{aligned} \lambda\bigl(E\triangle(E+h_{n_k})\bigr)&\le\varepsilon8^{-k},\\ \lambda\bigl(E\triangle(E-h_{n_k})\bigr)&\le\varepsilon8^{-k}. \end{aligned} \end{equation*}
The set \(E_\varepsilon\) above is then a measurable subset of \(E\) with
\begin{equation*} \begin{aligned} E\setminus E_\varepsilon\subset\bigcup_{k\ge1}\Bigl(&\bigl(E\setminus(E+h_{n_k})\bigr)\\ &\cup\bigl(E\setminus(E-h_{n_k})\bigr)\Bigr), \end{aligned} \end{equation*}
so \(\lambda(E\setminus E_\varepsilon)\le\sum_k2\varepsilon8^{-k}=2\varepsilon/7<\varepsilon\), i.e. \(\lambda(E_\varepsilon)>\lambda(E)-\varepsilon\); and \(x\in E_\varepsilon\) says precisely that \(x\in E+h_{n_k}\) and \(x\in E-h_{n_k}\) for every \(k\), i.e. \(x-h_{n_k}\in E\) and \(x+h_{n_k}\in E\).
(ii) Put \(\sigma_0:=(2^{-m})_{m\ge1}\) and, inductively, apply (i) to \(E\), the sequence \(\sigma_{n-1}\) and \(\varepsilon=2^{-n}\), obtaining a measurable \(F_n\subset E\) with \(\lambda(F_n)>\lambda(E)-2^{-n}\) and a subsequence \(\sigma_n=(h^{(n)}_k)_{k\ge1}\) of \(\sigma_{n-1}\) such that
\begin{equation*} x\pm h^{(n)}_k\in E\qquad\text{for all}\ x\in F_n,\ k\ge1. \tag{*} \end{equation*}
Set \(E_0:=\bigcup_nF_n\subset E\) and \(h_n:=h^{(n)}_n>0\). Then \(\lambda(E_0)\ge\lambda(F_n)>\lambda(E)-2^{-n}\) for every \(n\) forces \(\lambda(E_0)=\lambda(E)\). Writing \(\sigma_n=(2^{-m^{(n)}_k})_{k\ge1}\), a subsequence satisfies \(m^{(n)}_k\ge m^{(n-1)}_k\), so \(m^{(n)}_n\ge m^{(0)}_n=n\) and \(0<h_n\le2^{-n}\to0\). Finally, for \(x\in F_N\) put \(n(x):=N\): for \(n\ge N\) the sequence \(\sigma_n\) is a subsequence of \(\sigma_N\), so \(h_n\) is a term of \(\sigma_N\) and \((*)\) at stage \(N\) gives \(x+h_n\in E\).
Let \(A\) be a set of positive Lebesgue measure in \(\mathbb{R}^n\) and let \(k \in \mathbb{N}\). Prove that there exist a set \(B\) of positive Lebesgue measure and a number \(\delta > 0\) such that the sets \(B_{i_1,\dots,i_n} := B + \delta(i_1,\dots,i_n)\), where \(i_j \in \{1,\dots,k\}\), are disjoint and are contained in \(A\).
Take \(B:=Q^{\prime}\cap\bigcap_{v\in\{1,\dots,k\}^n}(A-\delta v)\), with the corner subcube \(Q^{\prime}\) and the step \(\delta\) chosen as follows. Put \(\eta:=\bigl(2k^n(k+1)^n\bigr)^{-1}\) and \(\theta:=1-\eta\).
Some half-open cube \(R\) of side \(r\) has \(\lambda(R\setminus A)<\eta r^n\). Indeed, fix \(\rho\) with \(0<\lambda(A_0)<\infty\) for \(A_0:=A\cap(-\rho,\rho)^n\) and, by outer regularity, an open \(U\supset A_0\) with \(\lambda(U)<\theta^{-1}\lambda(A_0)<\infty\); as \(\lambda(U)<\infty\), every \(x\in U\) lies in a maximal half-open dyadic cube contained in \(U\), and these maximal cubes \(Q_j\) are pairwise disjoint (two dyadic cubes are nested or disjoint) with \(U=\bigcup_jQ_j\). Were \(\lambda(A_0\cap Q_j)\le\theta\lambda(Q_j)\) for all \(j\), then
\begin{equation*} \lambda(A_0)=\sum_j\lambda(A_0\cap Q_j)\le\theta\lambda(U)<\lambda(A_0). \end{equation*}
So take \(R:=Q_j\) with \(\lambda(A\cap R)>\theta\lambda( R)\), let \(c\) be its lower corner, and set
\begin{equation*} \delta:=\frac{r}{k+1},\qquad Q^{\prime}:=[c_1,c_1+\delta)\times\cdots\times[c_n,c_n+\delta). \end{equation*}
For \(v=(i_1,\dots,i_n)\in\{1,\dots,k\}^n\) the cube \(Q^{\prime}+\delta v\) has \(j\)-th edge \([c_j+\delta i_j,\,c_j+\delta(i_j+1))\) with \(i_j+1\le k+1\), so \(Q^{\prime}+\delta v\subset R\); and \(B+\delta v\subset A\) by construction. These translates are disjoint, since the \(Q^{\prime}+\delta w\), \(w\in\mathbb{Z}^n\), are the cells of the partition of \(\mathbb{R}^n\) into half-open cubes of side \(\delta\) with corners in \(c+\delta\mathbb{Z}^n\). Finally \(Q^{\prime}\setminus B\subset\bigcup_v\bigl(Q^{\prime}\setminus(A-\delta v)\bigr)\), where translation invariance gives \(\lambda(Q^{\prime}\setminus(A-\delta v))=\lambda((Q^{\prime}+\delta v)\setminus A)\le\lambda(R\setminus A)<\eta r^n\), so over the \(k^n\) vectors \(v\),
\begin{equation*} \lambda(Q^{\prime}\setminus B)<k^n\eta r^n=k^n\eta(k+1)^n\delta^n=\tfrac12\delta^n, \end{equation*}
whence \(\lambda(B)>\lambda(Q^{\prime})-\delta^n/2=\delta^n/2>0\).
(Jones) In this exercise, by a Hamel basis we mean a Hamel basis of the space \(\mathbb{R}^1\) over the field of rational numbers.
(i) Let \(M\) be a set in \([0,1]\) and let \(\lambda_*(M - M) > 0\). Prove that \(M\) contains a Hamel basis. Deduce that the Cantor set contains a Hamel basis and that every set of positive measure contains a Hamel basis.
(ii) Prove that there exists a Hamel basis containing a nonempty perfect set.
(iii) Let \(H\) be a Hamel basis and \(DE := \{e_1 - e_2,\ e_1, e_2 \in E,\ e_1 \ge e_2\}\) for any set \(E\). Prove that \(\lambda^*(D^n H) > 0\) for some \(n\) and \(\lambda_*(D^n H) = 0\) for all \(n\), where \(D^n\) is defined inductively.
(iv) Let \(H\) be a Hamel basis and \(TE := \{e_1 + e_2 - e_3,\ e_1, e_2, e_3 \in E\}\) for any set \(E\). Prove that \(\lambda^*(T^n H) > 0\) for some \(n\) and \(\lambda_*(T^n H) = 0\) for all \(n\).
For a fixed Hamel basis \(H\) write \(\operatorname{supp}(x)\subset H\) for the finite set of basis vectors carrying a nonzero rational coefficient in the expansion of \(x\), and
\begin{equation*} W_N := \{x \in \mathbb{R} : |\operatorname{supp}(x)| \le N\} \qquad (N \ge 0), \end{equation*}
so that \(W_N+W_M\subset W_{N+M}\) and \(-W_N=W_N\).
Key Lemma: no \(W_N\) contains an interval. For if an open interval satisfies \(I\subset W_M\) with \(M\ge1\), then each \(x\in I\) with \(F:=\operatorname{supp}(x)\) admits \(h\in H\setminus F\) (\(H\) being infinite) and a nonzero rational \(q\) with \(x+qh\in I\), whence \(\operatorname{supp}(x+qh)=F\cup\{h\}\) forces \(|F|\le M-1\); thus \(I\subset W_{M-1}\), and \(M\) iterations give the absurdity \(I\subset W_0=\{0\}\).
Corollary: \(\lambda_*(S)=0\) whenever \(S\subset W_N\). Otherwise \(S\) contains a compact \(K\) with \(\lambda(K)>0\), so \(K-K\subset W_{2N}\) contains an interval by Exercise 1.12.62.
(i) A \(\mathbb{Q}\)-subspace \(V\subset\mathbb{R}\) with \(\lambda_*(V)>0\) equals \(\mathbb{R}\): it contains a compact \(K\) of positive measure, so \((-\varepsilon,\varepsilon)\subset K-K\subset V-V=V\) for some \(\varepsilon>0\) by Exercise 1.12.62, and every \(x\) has \(x/m\in V\) for large \(m\), hence \(x=m(x/m)\in V\). Now \(M-M\subset V:=\operatorname{span}_{\mathbb{Q}}(M)\) gives \(\lambda_*(V)>0\), so \(V=\mathbb{R}\); a set \(H\subset M\) maximal among the \(\mathbb{Q}\)-independent subsets of \(M\) (Zorn) satisfies \(M\subset\operatorname{span}_{\mathbb{Q}}(H)\), since otherwise \(H\cup\{m\}\) would be independent, so \(\operatorname{span}_{\mathbb{Q}}(H)=\mathbb{R}\) and \(H\) is a Hamel basis inside \(M\). For the Cantor set, \(C-C=[-1,1]\) by Exercise 1.12.60, so \(\lambda_*(C-C)>0\); for \(M\) of positive measure, \(M-M\) contains an interval by Exercise 1.12.62. Neither deduction used \(M\subset[0,1]\).
(ii) It suffices to build a nonempty perfect \(\mathbb{Q}\)-independent set \(P\): a maximal independent \(H\supset P\) is a Hamel basis as in (i). Construct nondegenerate closed intervals \(I_s\), \(s\) a finite binary string, with (a) \(I_{s0},I_{s1}\) disjoint subintervals of \(I_s\); (b) \(|I_s|\le2^{-|s|}\); (c) for all \(n\), all \(m\le n\), all distinct strings \(s_1,\dots,s_m\) of length \(n\) and all integers \(0<|k_i|\le n\),
\begin{equation*} 0 \notin k_1 I_{s_1} + \dots + k_m I_{s_m} . \end{equation*}
Start with \(I_{\varnothing}:=[1,2]\). Given the (finitely many, pairwise disjoint, nondegenerate) level-\(n\) intervals, the conditions (c) at this level form a finite list, treated one at a time: for data \(s_1,\dots,s_m\), \(k_1,\dots,k_m\), the function \(\Phi(x)=\sum_ik_ix_i\) is nonconstant on \(I_{s_1}\times\cdots\times I_{s_m}\) (as \(k_1\neq0\) and \(I_{s_1}\) is nondegenerate), so \(\Phi(x^0)\neq0\) somewhere, and continuity provides nondegenerate \(J_i\subset I_{s_i}\) around \(x^0_i\) with \(\Phi\neq0\) on \(\prod_iJ_i\); replace \(I_{s_i}\) by \(J_i\), which preserves the conditions already arranged, these being stable under shrinking. Then split each level-\(n\) interval into two disjoint nondegenerate closed subintervals of length at most \(2^{-(n+1)}\).
Put \(P:=\bigcap_n\bigcup_{|s|=n}I_s\), a nonempty compact set as a decreasing intersection of such; no point is isolated, since each \(I_s\) meets \(P\) in both of its disjoint children and \(|I_s|\to0\), so \(P\) is perfect. It is independent: given distinct \(x_1,\dots,x_m\in P\) and nonzero rationals, clear denominators to nonzero integers \(k_i\) and take \(n\ge\max\{m,|k_1|,\dots,|k_m|\}\) large enough that the \(x_i\) lie in distinct level-\(n\) intervals; then (c) gives \(\sum_ik_ix_i\neq0\).
(iii) Inner measure: \(D^{n+1}H\subset D^nH-D^nH\) and \(H\subset W_1\) give \(D^nH\subset W_{2^n}\) by induction, so \(\lambda_*(D^nH)=0\) by the Corollary.
Outer measure. As \(0\in DH\) and inductively \(0\in D^nH\), we get \(D^nH\subset D^{n+1}H\) (because \(a=a-0\)); put \(G:=\bigcup_{n\ge1}D^nH\subset[0,\infty)\), closed under nonnegative differences since \(\alpha\ge\beta\) in \(G\) lie in a common \(D^NH\) and \(\alpha-\beta\in D^{N+1}H\). Let \(\Lambda\) be the group generated by \(H-H\),
\begin{equation*} \begin{aligned} \Lambda = \Bigl\{ \sum_{i=1}^{m} c_i h_i :\ &c_i \in \mathbb{Z},\ h_i \in H,\\ &\textstyle\sum_{i=1}^m c_i = 0 \Bigr\}. \end{aligned} \end{equation*}
Claim 1: \(\lambda^*(\Lambda)>0\). Fix \(h_0\in H\); clearing a common denominator in \(x=\sum_iq_ih_i\) gives \(mx=\sum_ic_ih_i\) with \(c_i\in\mathbb{Z}\), so \(mx-ch_0\in\Lambda\) for \(c:=\sum_ic_i\) and \(x\in\frac1m(\Lambda+ch_0)\). Thus
\begin{equation*} \mathbb{R} = \bigcup_{m \ge 1} \bigcup_{c \in \mathbb{Z}} \frac1m (\Lambda + c h_0) \end{equation*}
is a countable union of affine images of \(\Lambda\), all null if \(\lambda^*(\Lambda)=0\).
Claim 2: \(G\supset\Lambda\cap[0,\theta)\), where \(\theta:=g_1-g_0\) for two elements \(g_0<g_1\) of the infinite set \(H\), so that \(\theta\in DH\subset G\) and \(\theta\in\Lambda\). Put \(S:=G\cap[0,\theta)\), identify \([0,\theta)\) with \(\mathbb{R}/\theta\mathbb{Z}\) and let \(\pi\) be the quotient map. Then \(S\) is a subgroup: \(b\in G\cap[0,\theta]\) gives \(\theta-b\in G\); for \(a,b\in S\) with \(a+b<\theta\), the element \(\theta-a\in G\) exceeds \(b\), so \(\theta-(a+b)\in G\cap(0,\theta]\) and one more such step gives \(a+b\in S\); for \(a+b\ge\theta\), \(\theta-b\in G\) and \(a\ge\theta-b\) give \(a+b-\theta\in G\cap[0,\theta)=S\); and the inverse of \(a\in S\) is \(0\) or \(\theta-a\). Next \(\pi(DH)\subset S\): for \(d\in DH\subset G\), repeated subtraction of \(\theta\in G\) stays inside \(G\) (each step is a nonnegative difference) and lands in \(G\cap[0,\theta)=S\). Every element of \(H-H\) lies in \(DH\) or is the negative of one, so \(\pi(\Lambda)\subset S\), \(\Lambda\) being generated by \(H-H\) and \(S\) a subgroup; and \(x\in\Lambda\cap[0,\theta)\) is the representative of \(\pi(x)\), hence \(x\in S\subset G\).
Since \(\theta\in\Lambda\) gives \(\Lambda=\bigcup_{k\in\mathbb{Z}}\bigl(k\theta+(\Lambda\cap[0,\theta))\bigr)\), countable subadditivity and translation invariance of \(\lambda^*\) turn Claim 1 into \(\lambda^*(\Lambda\cap[0,\theta))>0\), so \(\lambda^*(G)>0\) by Claim 2 and, \(G\) being a countable union, \(\lambda^*(D^nH)>0\) for some \(n\).
(iv) Inner measure: \(TE\subset E+E-E\) and \(H\subset W_1\) give \(T^nH\subset W_{3^n}\), so \(\lambda_*(T^nH)=0\) by the Corollary.
Outer measure. Taking \(e_2=e_3\) gives \(E\subset TE\), so \(T^nH\subset T^{n+1}H\); put \(U:=\bigcup_{n\ge1}T^nH\). Every \(h\in H\) is an integer combination of basis vectors with coefficient sum \(1\), a property preserved by \((e_1,e_2,e_3)\mapsto e_1+e_2-e_3\), so
\begin{equation*} \begin{aligned} U \subset \Gamma := \Bigl\{ \sum_{i=1}^m c_i h_i :\ &c_i \in \mathbb{Z},\ h_i \in H,\\ &\textstyle\sum_{i=1}^m c_i = 1 \Bigr\}. \end{aligned} \end{equation*}
Conversely, for \(x=\sum_{i=1}^mc_ih_i\in\Gamma\) with distinct \(h_i\) and \(c_i\neq0\) induct on \(\sum_i|c_i|\): the value \(1\) gives \(x\in H\subset U\); otherwise some \(c_k\le-1\) and some \(c_j\ge1\), and \(y:=x-h_j+h_k\) again has coefficient sum \(1\) with absolute sum \(\sum_i|c_i|-2\), so \(y\in U\) by induction and, \(y,h_j,h_k\) lying in a common \(T^NH\),
\begin{equation*} x = y + h_j - h_k \in T^{N+1}H \subset U . \end{equation*}
Thus \(U=\Gamma=h_0+\Lambda\) for any \(h_0\in H\), so \(\lambda^*(U)=\lambda^*(\Lambda)>0\) by Claim 1 and \(\lambda^*(T^nH)>0\) for some \(n\).
Prove the existence of a nonmeasurable (in the sense of Lebesgue) Hamel basis of \(\mathbb{R}^1\) over \(\mathbb{Q}\) without using the continuum hypothesis (see Example 1.12.21).
Build by transfinite recursion, using only the axiom of choice, a Hamel basis meeting every compact set of positive measure. Let \(\mathfrak c=2^{\aleph_0}\) and let \(\omega_{\mathfrak c}\) be the least ordinal of cardinality \(\mathfrak c\), so that \(|\alpha|<\mathfrak c\) for \(\alpha<\omega_{\mathfrak c}\).
There are exactly \(\mathfrak c\) compact sets of positive measure: a closed \(F\) is determined by the countable family \(\{(p,q):p,q\in\mathbb{Q},\ (p,q)\cap F=\varnothing\}\), giving at most \(2^{\aleph_0}=\mathfrak c\) closed sets, while the intervals \([0,t]\), \(t>0\), give at least \(\mathfrak c\). Enumerate them as \(\{K_\alpha:\alpha<\omega_{\mathfrak c}\}\); each has \(|K_\alpha|=\mathfrak c\) by Exercise 1.12.59.
Choose \(h_\alpha\in K_\alpha\setminus L_\alpha\) recursively, where \(L_\alpha:=\operatorname{span}_{\mathbb{Q}}\{h_\beta:\beta<\alpha\}\); this is possible because the finite subsets of a set of cardinality \(\kappa\) number \(\max(\kappa,\aleph_0)\) and each carries countably many rational coefficient vectors, so
\begin{equation*} |L_\alpha| \le \max(|\alpha|, \aleph_0) < \mathfrak c = |K_\alpha| . \end{equation*}
The family \(\{h_\alpha\}\) is \(\mathbb{Q}\)-independent: in a relation \(\sum_{i\le m}q_ih_{\alpha_i}=0\) with \(\alpha_1<\cdots<\alpha_m\) and \(j\) the largest index with \(q_j\neq0\),
\begin{equation*} h_{\alpha_j} = - q_j^{-1}\sum_{i < j} q_i h_{\alpha_i} \in L_{\alpha_j}, \end{equation*}
against the choice of \(h_{\alpha_j}\).
Extend \(\{h_\alpha\}\) by Zorn to a maximal \(\mathbb{Q}\)-independent \(H\subset\mathbb{R}\), a Hamel basis as in Exercise 1.12.65(i), which by construction meets every compact set of positive measure. Were \(H\) measurable, Lemma 1.12.20 (every Hamel basis has inner measure zero) would force \(\lambda(H)=0\), so \(\lambda([0,1]\setminus H)=1\) and inner regularity would give a compact \(K\subset[0,1]\setminus H\) with \(\lambda(K)>0\) disjoint from \(H\).
Prove that there exists a bounded set \(E\) of measure zero such that \(E + E\) is nonmeasurable.
Take \(E:=E_{N-1}\), where \(E_0:=A:=\{rh:r\in\mathbb{Q}\cap[-1,1],\ h\in H\}\), \(E_{n+1}:=E_n+E_n\), \(H\subset C\) is a Hamel basis inside the Cantor set (Exercise 1.12.65(i)), and \(N\ge1\) is the least index with \(E_N\) nonmeasurable. The book’s hint is used with the rational factors allowed to be negative, which is what makes the \(E_n\) exhaust the line.
Here \(\lambda(H)=0\) by \(\lambda( C)=0\) and completeness of \(\lambda\), so \(A=\bigcup_rrH\subset[-1,1]\) is measurable with \(\lambda(A)=0\), each \(rH\) being a dilation of a null set; and by induction \(E_n\subset[-2^n,2^n]\) with
\begin{equation*} E_n = \Bigl\{ \sum_{j=1}^{2^n} r_j h_j :\ r_j \in \mathbb{Q}\cap[-1,1],\ h_j \in H \Bigr\}. \end{equation*}
(i) \(\lambda_*(E_n)=0\): every element of \(A\) has the form \(rh\), so \(A\subset W_1\) and hence \(E_n\subset W_{2^n}\) in the notation of Exercise 1.12.65, whose Corollary applies.
(ii) \(\bigcup_{n\ge0}E_n=\mathbb{R}\): here \(0=0\cdot h\in A\), while \(x\neq0\) expands as \(x=\sum_{i\le m}q_ih_i\) with \(q_i\in\mathbb{Q}\setminus\{0\}\) and distinct \(h_i\in H\), and splitting each term into \(k_i:=\lceil|q_i|\rceil\) equal parts \((q_i/k_i)h_i\) with \(q_i/k_i\in\mathbb{Q}\cap[-1,1]\) writes \(x\) as a sum of \(k:=k_1+\cdots+k_m\) elements of \(A\); pad with \(0\in A\) and take \(2^n\ge k\).
So not every \(E_n\) is null for \(\lambda^*\), else the line would be null by (ii) and countable subadditivity; for such an \(n\), necessarily \(n\ge1\) since \(\lambda^*(E_0)=\lambda(A)=0\), the set \(E_n\) is nonmeasurable because \(\lambda_*(E_n)=0<\lambda^*(E_n)\) by (i). Hence \(N\) exists, and \(E=E_{N-1}\) is bounded and measurable by minimality of \(N\), so \(\lambda(E)=\lambda_*(E)=0\) by (i), while \(E+E=E_N\) is nonmeasurable.
Exercises 1.12.68–1.12.74
(Ciesielski, Fejzic, Freiling) Show that every set \(E \subset \mathbb{R}\) contains a subset \(A\) with \(\lambda_*(A+A) = 0\) and \(\lambda^*(A+A) = \lambda^*(E+E)\), where \(\lambda\) is Lebesgue measure.
Take \(A:=E\cap G_b\) for a suitable \(b\) in a Hamel basis \(B\) of \(\mathbb{R}\) over \(\mathbb{Q}\), where \(\operatorname{supp}(x)\subset B\) is the finite support of the expansion of \(x\) and
\begin{equation*} \begin{aligned} G_b &:= \operatorname{span}_{\mathbb{Q}}(B \setminus \{b\})\\ &\ = \{x \in \mathbb{R} : b \notin \operatorname{supp}(x)\} . \end{aligned} \end{equation*}
We may assume \(E\neq\emptyset\).
Two facts. A proper subgroup \(G\subset(\mathbb{R},+)\) has \(\lambda_*(G)=0\): otherwise \(G\) contains a compact \(K\) with \(\lambda(K)>0\), so \((-d,d)\subset K-K\subset G-G=G\) for some \(d>0\) by Exercise 1.12.62, and \(G\supset\bigcup_{n\ge1}n(-d,d)=\mathbb{R}\). And \(B\) is uncountable, since a countable \(B\) would make \(\operatorname{span}_{\mathbb{Q}}(B)=\mathbb{R}\) countable; each \(G_b\) is a \(\mathbb{Q}\)-subspace with \(b\notin G_b\), hence a proper subgroup, so \(\lambda_*(G_b)=0\).
By the axiom of choice fix a representation \(v=a_v+b_v\) with \(a_v,b_v\in E\) for every \(v\in E+E\), put \(\sigma_v:=\operatorname{supp}(a_v)\cup\operatorname{supp}(b_v)\), a finite subset of \(B\), and set \(A_b:=E\cap G_b\), \(T_b:=A_b+A_b\). Then \(T_b\subset E+E\) and, \(G_b\) being a group, \(T_b\subset G_b\), so
\begin{equation*} \lambda_*(T_b) = 0 \quad \text{for every } b \in B; \tag{1} \end{equation*}
moreover \(b\notin\sigma_v\) puts \(a_v,b_v\) in \(A_b\), whence
\begin{equation*} \{v \in E+E:\ b \notin \sigma_v\} \subset T_b. \tag{2} \end{equation*}
It remains to find \(b\) with \(\lambda^*(T_b)=\lambda^*(E+E)\).
Step 1. For distinct \(b_1,b_2,\dots\in B\) and \(M\in\mathbb{N}\), writing \(s_M:=\lambda^*((E+E)\cap[-M,M])\le2M\),
\begin{equation*} \sup_{n} \lambda^*\bigl(T_{b_n} \cap [-M,M]\bigr) = s_M . \tag{3} \end{equation*}
Indeed, \(U_N:=\{v\in E+E:\sigma_v\cap\{b_m:m\ge N\}=\emptyset\}\) increases to \(E+E\), each \(\sigma_v\) being finite and the \(b_m\) distinct, so \(\lambda^*(U_N\cap[-M,M])\to s_M\) by continuity of the outer measure from below (Proposition 1.5.12), while \(U_N\subset T_{b_n}\) for \(n\ge N\) by (2); the reverse inequality is clear from \(T_b\subset E+E\).
Step 2. \(B_M:=\{b\in B:\lambda^*(T_b\cap[-M,M])<s_M\}\) is at most countable: for \(\varepsilon>0\) the set \(B_{M,\varepsilon}:=\{b:\lambda^*(T_b\cap[-M,M])\le s_M-\varepsilon\}\) is finite, an infinite one containing distinct \(b_1,b_2,\dots\) against (3) (legitimate as \(s_M<\infty\)), and \(B_M=\bigcup_{k\ge1}B_{M,1/k}\).
Step 3. Since \(B\) is uncountable and \(\bigcup_MB_M\) is not, choose \(b\in B\setminus\bigcup_MB_M\) and put \(A:=A_b\). Then \(\lambda^*(T_b\cap[-M,M])=s_M\) for every \(M\), and Proposition 1.5.12, applied to \(T_b\cap[-M,M]\uparrow T_b\) and to \((E+E)\cap[-M,M]\uparrow E+E\), gives
\begin{equation*} \begin{aligned} \lambda^*(A+A) &= \lim_{M \to \infty} \lambda^*(T_b \cap [-M,M])\\ &= \lim_{M\to\infty} s_M = \lambda^*(E+E), \end{aligned} \end{equation*}
while \(\lambda_*(A+A)=0\) by (1).
(Sodnomov) Let \(E \subset \mathbb{R}^1\) be a set of positive Lebesgue measure. Then there exists a perfect set \(P\) with \(P + P \subset E\).
Take \(P:=\frac c2+\bigl\{\sum_{i\ge1}\varepsilon_ia_i:\varepsilon_i\in\{0,1\}\bigr\}\), for a point \(c\) and numbers \(a_1>a_2>\cdots>0\) with \(a_{n+1}\le a_n/10\), constructed below so that
\begin{equation*} \begin{aligned} &c + K \subset E, \qquad\text{where}\\ &K := \Big\{ \sum_{i=1}^{\infty} \delta_i a_i :\ \delta_i \in \{0,1,2\} \Big\}. \end{aligned} \tag{1} \end{equation*}
Since \(\{\varepsilon+\varepsilon^{\prime}:\varepsilon,\varepsilon^{\prime}\in\{0,1\}\}=\{0,1,2\}\), this gives \(P+P=c+K\subset E\). And \(P\) is perfect: the map \(\varphi(\varepsilon)=c/2+\sum_i\varepsilon_ia_i\) on \(\{0,1\}^{\mathbb{N}}\) is well defined and continuous, the series converging uniformly by \(a_i\le a_110^{-(i-1)}\), so \(P=\varphi(\{0,1\}^{\mathbb{N}})\) is compact, and \(\varphi\) is injective, since for \(\varepsilon,\varepsilon^{\prime}\) first differing at \(n\), say with \(\varepsilon_n=1\),
\begin{equation*} \begin{aligned} \varphi(\varepsilon) - \varphi(\varepsilon^{\prime}) &\ \ge\ a_n - \sum_{i > n} a_i\\ &\ \ge\ a_n\Big(1 - \sum_{j\ge 1} 10^{-j}\Big) = \tfrac{8}{9} a_n > 0 ; \end{aligned} \end{equation*}
a continuous injection of a compact space is a homeomorphism onto its image, so \(P\) is homeomorphic to the Cantor cube, which is nonempty, compact and without isolated points.
Density lemma: a compact \(C\) with \(\lambda( C)>0\) lies \(\tfrac78\)-densely in some closed interval \(I\) of positive length, \(\lambda(I\setminus C)<\tfrac18\lambda(I)\). Choose by outer regularity a bounded open \(U\supset C\) with \(\lambda(U)<\tfrac87\lambda( C)\) and write \(U=\bigcup_kJ_k\) as a disjoint union of open intervals; were \(\lambda(C\cap J_k)\le\tfrac78\lambda(J_k)\) for all \(k\), then
\begin{equation*} \lambda( C) = \sum_k \lambda(C \cap J_k) \le \tfrac78 \lambda(U) < \lambda( C). \end{equation*}
Take \(I:=\overline{J_k}\) for a good \(k\).
Construction. Start from a compact \(C_0\subset E\) with \(\lambda(C_0)>0\) (inner regularity). Given a compact \(C_n\) of positive measure, take \(I_n=[\alpha_n,\beta_n]\) from the lemma, choose \(0<a_{n+1}<\tfrac1{16}\lambda(I_n)\) subject also to \(a_{n+1}\le a_n/10\) for \(n\ge1\), and put
\begin{equation*} \begin{aligned} I_n^{\prime}&:=[\alpha_n,\beta_n-2a_{n+1}],\\ C_{n+1}&:=C_n\cap(C_n-a_{n+1})\cap(C_n-2a_{n+1})\cap I_n^{\prime}, \end{aligned} \end{equation*}
compact sets decreasing in \(n\). Here \(\lambda(C_{n+1})>0\): from \(I_n^{\prime}+ja_{n+1}\subset I_n\) for \(j=0,1,2\) and translation invariance,
\begin{equation*} \begin{aligned} \lambda\bigl(I_n^{\prime}\setminus(C_n-ja_{n+1})\bigr)&\le\lambda(I_n\setminus C_n)\\ &<\tfrac18\lambda(I_n), \end{aligned} \end{equation*}
so the triple intersection misses less than \(\tfrac38\lambda(I_n)\) of \(I_n^{\prime}\), while \(\lambda(I_n^{\prime})=\lambda(I_n)-2a_{n+1}>\tfrac78\lambda(I_n)\), whence \(\lambda(C_{n+1})>\tfrac12\lambda(I_n)\). Pick \(c\in\bigcap_nC_n\), nonempty as a nested intersection of nonempty compacta.
Verification of (1). Downward induction on \(m\) gives \(c+\sum_{i=m+1}^n\delta_ia_i\in C_m\) for \(0\le m\le n\) and all \(\delta_i\in\{0,1,2\}\): the case \(m=n\) is \(c\in C_n\), and \(y\in C_m\) with \(m\ge1\) lies in \(C_{m-1}\cap(C_{m-1}-a_m)\cap(C_{m-1}-2a_m)\), so \(y+\delta_ma_m\in C_{m-1}\). With \(m=0\) all partial sums \(c+\sum_{i\le n}\delta_ia_i\) lie in \(C_0\), and \(C_0\) is closed, so
\begin{equation*} c + \sum_{i=1}^{\infty}\delta_i a_i \in C_0 \subset E . \end{equation*}
Let \(\beta \in (0,1)\). The operation \(T(\beta)\) over a finite family of disjoint intervals \(I_1, \dots, I_n\) of nonzero length consists of deleting from every \(I_j\) the open interval with the same center as \(I_j\) and length \(\beta\lambda(I_j)\). Given a sequence of numbers \(\beta_n \in (0,1)\), let us define inductively compacts \(K_n\) obtained by consequent application of the operations \(T(\beta_1), \dots, T(\beta_n)\), starting with the interval \(I = [0,1]\).
(i) Show that \(\lambda\big(\bigcap_{n=1}^{\infty} K_n\big) = \lim_{n\to\infty} \prod_{i=1}^{n}(1-\beta_i)\). In particular, letting \(\beta_n = 1 - \alpha^{\frac{1}{n(n+1)}}\), where \(\alpha \in (0,1)\), we have \(\lambda\big(\bigcap_{n=1}^{\infty}K_n\big) = \alpha\).
(ii) Show that there exists a sequence of pairwise disjoint nowhere dense compact sets \(A_n\) with the following properties: \(\lambda(A_n) = 2^{-n}\) and the intersection of \(A_{n+1}\) with each interval contiguous to the set \(\bigcup_{j=1}^{n}A_j\) has a positive measure.
(iii) Show that the intersections of the set \(A := \bigcup_{n=1}^{\infty}A_{2n-1}\) and its complement with every interval \(I \subset [0,1]\) have positive measures.
(i) Deleting from a closed interval \(J\) the concentric open interval of length \(\beta\lambda(J)\) leaves two closed intervals of length \((1-\beta)\lambda(J)/2\), so by induction \(K_n\) is a union of \(2^n\) disjoint closed intervals of common length \(\ell_n=2^{-n}\prod_{i\le n}(1-\beta_i)\), whence \(\lambda(K_n)=\prod_{i\le n}(1-\beta_i)\). The \(K_n\) are compact and decrease with \(\lambda(K_0)=1<\infty\), so continuity from above gives
\begin{equation*} \lambda\Big(\bigcap_{n=1}^{\infty}K_n\Big) = \lim_{n\to\infty}\prod_{i=1}^{n}(1-\beta_i), \end{equation*}
the limit existing because the partial products decrease. For \(\beta_n=1-\alpha^{1/n(n+1)}\),
\begin{equation*} \begin{aligned} \prod_{i=1}^{n}(1-\beta_i) &= \alpha^{\sum_{i\le n}(\frac{1}{i} - \frac{1}{i+1})}\\ &= \alpha^{\,1 - \frac{1}{n+1}} \to \alpha . \end{aligned} \end{equation*}
Since \(\ell_n\le2^{-n}\to0\), the compact \(K:=\bigcap_nK_n\) contains no interval, i.e. is nowhere dense, and it contains the never-deleted points \(0,1\). Conjugating by the affine map of \([0,1]\) onto a closed interval \(J\) yields, for each \(\alpha\in(0,1)\), a nowhere dense compact \(K(J,\alpha)\subset J\) containing both endpoints of \(J\), with \(\lambda(K(J,\alpha))=\alpha\lambda(J)\) and with every component of \(J\setminus K(J,\alpha)\) a deleted concentric interval, of length at most \(\beta_1\lambda(J)=(1-\sqrt{\alpha})\lambda(J)\), the interval deleted at step \(n\) having length \(\beta_n\ell_{n-1}\le\beta_1\).
(ii) As printed the \(A_n\) cannot all be compact: for closed nowhere dense \(F\) of positive measure \([0,1]\setminus F\) has infinitely many components (finitely many would make \(F\) a finite union of closed intervals, with interior points or measure zero), whereas a compact subset of an open \(U\) meets only finitely many components of \(U\) (points in distinct components would have a limit point in \(U\), whose component is a neighbourhood of it containing infinitely many of them, though it can contain at most one). We prove the version used in (iii): the \(A_n\) are pairwise disjoint nowhere dense Borel sets, each a countable union of nowhere dense compacta, with \(\lambda(A_n)=2^{-n}\) and the stated intersection property, and with every partial union \(F_n:=\bigcup_{j\le n}A_j\) compact and nowhere dense, which is what makes “interval contiguous to \(F_n\)” meaningful.
Put \(A_1:=K([0,1],1/2)\). Suppose \(F_n\) is compact and nowhere dense with \(\lambda(F_n)=1-2^{-n}\), containing \(0\) and \(1\), and let \(\delta_n\) be the largest length of a component of \(U_n:=[0,1]\setminus F_n\). Let \(\{J_k\}_{k\ge1}\) be those components, so \(\sum_k\lambda(J_k)=2^{-n}\); let \(J_k^{\prime}\subset J_k\) be the concentric closed subinterval with \(\lambda(J_k^{\prime})=\tfrac34\lambda(J_k)\), the two components of \(J_k\setminus J_k^{\prime}\) having length \(\tfrac18\lambda(J_k)\); and put
\begin{equation*} C_k := K\big(J_k^{\prime}, \tfrac23\big), \qquad A_{n+1} := \bigcup_{k\ge 1} C_k . \end{equation*}
Then \(A_{n+1}\subset U_n\) is disjoint from \(A_1,\dots,A_n\), and
\begin{equation*} \lambda(A_{n+1}) = \sum_k \tfrac23 \lambda(J_k^{\prime}) = \tfrac12 \lambda(U_n) = 2^{-(n+1)} , \end{equation*}
with \(\lambda(A_{n+1}\cap J_k)=\tfrac12\lambda(J_k)>0\) for every contiguous interval \(J_k\) — the required property. Next, \(F_{n+1}:=F_n\cup A_{n+1}\) is compact: a limit point of \(A_{n+1}\) lying in \(U_n\) has a whole neighbourhood inside some \(J_k\), hence is a limit point of the compact \(A_{n+1}\cap J_k=C_k\); the others lie in \(F_n\). It is nowhere dense: for an interval \(I\) of positive length, density of \(U_n\) gives \(k\) with \(\lambda(I\cap J_k)>0\), and the nowhere dense \(C_k\) leaves a point of \(I\cap J_k\) outside \(C_k\), hence outside \(F_{n+1}\). Finally the components of \([0,1]\setminus F_{n+1}\) are the two end intervals of \(J_k\setminus J_k^{\prime}\), of length \(\tfrac18\lambda(J_k)\), and the intervals deleted inside \(J_k^{\prime}\), of length at most \((1-\sqrt{2/3})\lambda(J_k^{\prime})<\tfrac15\lambda(J_k)\) — here \(C_k\) contains both endpoints of \(J_k^{\prime}\), so no end interval is glued to a deleted one. Hence
\begin{equation*} \delta_{n+1} \le \tfrac{1}{5}\,\delta_n , \qquad \text{so} \qquad \delta_n \to 0. \tag{1} \end{equation*}
(iii) Fix an interval \(I\subset[0,1]\) of positive length and, by (1), an \(n\) with \(\delta_n<\lambda(I)/3\). The dense open \(U_n\) meets the concentric subinterval \(I_0\subset I\) of length \(\lambda(I)/3\), and a component \(J\) of \(U_n\) meeting \(I_0\) has \(\lambda(J)\le\delta_n<\lambda(I)/3\), the distance from \(I_0\) to each endpoint of \(I\), so \(J\subset I\). By (ii), \(\lambda(A_{n+1}\cap J)>0\); and \(J\) contains components \(J^{\prime}\) of \(U_{n+1}\), each with \(\lambda(A_{n+2}\cap J^{\prime})>0\) by (ii) at the next step, so \(\lambda(A_{n+2}\cap J)>0\). One of \(n+1,n+2\) is odd, the other even, so with \(A=\bigcup_nA_{2n-1}\) and the \(A_j\) pairwise disjoint,
\begin{equation*} \begin{aligned} \lambda(A \cap I) &\ \ge\ \lambda(A_{m}\cap J) > 0,\\ \lambda(I \setminus A) &\ \ge\ \lambda(A_{m^{\prime}}\cap J) > 0, \end{aligned} \end{equation*}
where \(m\) is the odd and \(m^{\prime}\) the even one of \(n+1,n+2\).
(\(\circ\)) Prove that Lebesgue measure of every measurable set \(E \subset \mathbb{R}^n\) equals the infimum of the sums \(\sum_{k=1}^{\infty}\lambda_n(U_k)\) over all sequences of open balls \(U_k\) covering \(E\).
The infimum in question,
\begin{equation*} \begin{aligned} \nu(E) := \inf\Big\{\sum_{k=1}^{\infty}\lambda_n(U_k):\ &U_k \text{ open balls},\\ &E \subset \bigcup_{k}U_k\Big\}, \end{aligned} \end{equation*}
equals \(\lambda_n(E)\). One inequality is countable subadditivity, \(\lambda_n(E)\le\sum_k\lambda_n(U_k)\) for every cover (covers exist: the balls of radius \(k\) about the origin); for the other we may assume \(\lambda_n(E)<\infty\).
Write \(\omega_n=\lambda_n\{|x|<1\}\). A cube \(Q\) of edge \(s\) contains the concentric ball \(B\) of radius \(s/2\), with \(\lambda_n(B)=\kappa_n\lambda_n(Q)\), \(\kappa_n:=2^{-n}\omega_n\in(0,1)\), and is contained in the concentric ball \(W\) of radius \(s\sqrt n\), with \(\lambda_n(W)=\gamma_n\lambda_n(Q)\), \(\gamma_n:=\omega_nn^{n/2}\). Hence a null set \(N\) can be covered by balls \(W_i\) with \(\sum_i\lambda_n(W_i)<\varepsilon\): cover it by cubes of total volume less than \(\varepsilon/\gamma_n\) and pass to the circumscribed balls.
Filling lemma: an open \(U\) with \(\lambda_n(U)<\infty\) contains pairwise disjoint open balls \(V_j\) with \(\lambda_n(U\setminus\bigcup_jV_j)=0\). By Lemma 1.7.2, applied inside each cube of a partition of \(\mathbb{R}^n\) into translates of \((-1,1)^n\) (whose boundaries form a null set), \(U=\bigcup_jQ_j\cup N_0\) with disjoint open cubes \(Q_j\subset U\) and \(\lambda_n(N_0)=0\). The inscribed balls \(B_j\subset Q_j\) are disjoint with \(\sum_j\lambda_n(B_j)=\kappa_n\lambda_n(U)\), so some finite \(J\) has \(\sum_{j\in J}\lambda_n(B_j)\ge\tfrac12\kappa_n\lambda_n(U)\), and
\begin{equation*} U^{(1)} := U \setminus \bigcup_{j\in J}\overline{B_j} \end{equation*}
is open with \(\lambda_n(U^{(1)})\le(1-\tfrac12\kappa_n)\lambda_n(U)\), the spheres \(\partial B_j\) being null. Iterate. The balls of all generations are pairwise disjoint, a new generation lying in \(U^{(m)}\), which is disjoint from the closed balls already chosen, and are contained in \(U\), while \(U\setminus\bigcup_jV_j\) differs from \(U^{(m)}\) by a subset of finitely many spheres, so
\begin{equation*} \begin{aligned} \lambda_n\Big(U\setminus\bigcup_j V_j\Big) &\le \big(1-\tfrac12\kappa_n\big)^m\lambda_n(U)\\ &\text{for all } m . \end{aligned} \end{equation*}
For open \(U\) of finite measure and \(\varepsilon>0\), take such \(V_j\) and cover the null set \(N:=U\setminus\bigcup_jV_j\) by balls \(W_i\) of total measure less than \(\varepsilon\); the \(V_j\) and \(W_i\) cover \(U\), and the \(V_j\) being disjoint subsets of \(U\),
\begin{equation*} \sum_j \lambda_n(V_j) + \sum_i \lambda_n(W_i) \ \le\ \lambda_n(U) + \varepsilon , \end{equation*}
where a finite family is completed to a sequence by further balls of total measure less than \(\varepsilon\). Hence \(\nu(U)\le\lambda_n(U)\). Finally, for measurable \(E\) with \(\lambda_n(E)<\infty\), regularity gives an open \(U\supset E\) with \(\lambda_n(U)\le\lambda_n(E)+\varepsilon\), and every ball cover of \(U\) covers \(E\), so
\begin{equation*} \nu(E) \le \nu(U) \le \lambda_n(U) \le \lambda_n(E) + \varepsilon . \end{equation*}
Suppose that \(\mu\) is a countably additive measure with values in \([0,+\infty]\) on the \(\sigma\)-algebra of Borel sets in \(\mathbb{R}^n\) and is finite on balls, and let \(W\) be a nonempty open set in \(\mathbb{R}^n\). Prove that there exists an at most countable collection of disjoint open cubes \(Q_j\) in \(W\) with edges parallel to the coordinate axes such that
\begin{equation*} \mu\Big(W \setminus \bigcup_{j=1}^{\infty} Q_j\Big) = 0 . \end{equation*}
Shift the dyadic grid by a vector all of whose grid hyperplanes are \(\mu\)-negligible, and run the proof of Lemma 1.7.2 with the resulting cubes.
Write \(H_{i,t}:=\{x\in\mathbb{R}^n:\ x_i=t\}\). Each set \(T_i:=\{t:\ \mu(H_{i,t})>0\}\) is at most countable: with \(B_m\) the ball of radius \(m\) about the origin we have \(\mu(B_m)<\infty\) by hypothesis, the slices \(H_{i,t}\cap B_m\) are disjoint Borel subsets of \(B_m\), so \(\{t:\ \mu(H_{i,t}\cap B_m)\ge 1/k\}\) has at most \(k\,\mu(B_m)\) elements, and \(T_i\) is the union of these finite sets over \(m,k\in\mathbb{N}\) (continuity of \(\mu\) from below). With \(R:=\{m2^{-k}:\ m\in\mathbb{Z},\ k\in\mathbb{N}\}\) the set \(\{t-r:\ t\in T_i,\ r\in R\}\) is countable, so we may choose \(\alpha_i\) outside it; then
\begin{equation*} \mu\bigl(H_{i,\,\alpha_i+r}\bigr)=0\qquad\text{for all } i\le n,\ r\in R. \tag{1} \end{equation*}
Consider the shifted dyadic cubes
\begin{equation*} P_{k,m}:=\prod_{i=1}^{n}\bigl[\alpha_i+m_i2^{-k},\ \alpha_i+(m_i+1)2^{-k}\bigr). \end{equation*}
For fixed \(k\) they partition \(\mathbb{R}^n\); two of them (with possibly different \(k\)) are disjoint or nested; and \(\partial P_{k,m}\) lies in \(2n\) hyperplanes \(H_{i,\alpha_i+r}\) with \(r\in R\), so \(\mu(\partial P_{k,m})=0\) by (1).
Let \(\mathcal{F}_k\) consist of those \(P_{k,m}\) with \(\overline{P_{k,m}}\subset W\) that are contained in no cube of \(\mathcal{F}_1\cup\dots\cup\mathcal{F}_{k-1}\), and put \(\mathcal{F}:=\bigcup_k\mathcal{F}_k=\{P^{(1)},P^{(2)},\dots\}\), an at most countable family. Its members are pairwise disjoint: for equal \(k\) they are distinct cubes of one grid, and for \(k<k^{\prime}\) nesting is excluded by the selection rule. They cover \(W\): given \(x\in W\), the grid cube of step \(2^{-k}\) containing \(x\) has diameter \(\sqrt n\,2^{-k}\to0\), hence closure inside the open set \(W\) for large \(k\), and at the least such \(k\) that cube either lies in \(\mathcal{F}_k\) or is contained in an earlier cube of \(\mathcal{F}\), which then contains \(x\).
Put \(Q_j:=\operatorname{int}P^{(j)}\): disjoint open cubes in \(W\) with edges parallel to the axes. Since \(P^{(j)}\setminus Q_j\subset\partial P^{(j)}\),
\begin{equation*} \mu\Bigl(W\setminus\bigcup_j Q_j\Bigr)\le\sum_j\mu\bigl(\partial P^{(j)}\bigr)=0. \qquad\blacksquare \end{equation*}
(\(\circ\)) Show that a set \(E \subset \mathbb{R}\) is Lebesgue measurable precisely when for every \(\varepsilon > 0\), there exist open sets \(U\) and \(V\) such that \(E \subset U\), \(U \setminus E \subset V\) and \(\lambda(V) < \varepsilon\).
Both implications run off the regularity of outer measure: every \(S\subset\mathbb{R}\) and \(\delta>0\) admit an open \(G\supset S\) with \(\lambda(G)\le\lambda^{*}(S)+\delta\), since in the definition of \(\lambda^{*}\) the covering intervals may be taken open.
Sufficiency. Apply the hypothesis with \(\varepsilon=1/k\) to get open \(U_k\supset E\) and \(V_k\supset U_k\setminus E\) with \(\lambda(V_k)<1/k\), and put \(G:=\bigcap_k U_k\supset E\). Then
\begin{equation*} G\setminus E\subset U_k\setminus E\subset V_k,\qquad \lambda^{*}(G\setminus E)\le\lambda(V_k)<1/k \end{equation*}
for every \(k\), so \(\lambda^{*}(G\setminus E)=0\). A set of outer measure zero is measurable (completeness of \(\lambda\)), hence so is \(E=G\setminus(G\setminus E)\).
Necessity. Let \(E\) be measurable and \(\varepsilon>0\).
(i) \(\lambda(E)<\infty\). Choose an open \(U\supset E\) with \(\lambda(U)\le\lambda(E)+\varepsilon/2\); then \(\lambda(U\setminus E)=\lambda(U)-\lambda(E)\le\varepsilon/2\), the subtraction being legitimate since \(\lambda(E)<\infty\). Regularity applied to \(U\setminus E\) yields an open \(V\supset U\setminus E\) with \(\lambda(V)<\varepsilon\).
(ii) General \(E\). Put \(E_m:=E\cap[m,m+1)\) and, by (i), choose open \(U_m\supset E_m\) with \(\lambda(U_m\setminus E_m)<\varepsilon\,2^{-|m|-3}\). Then \(U:=\bigcup_{m\in\mathbb{Z}}U_m\) is open, \(E\subset U\), and \(U_m\setminus E\subset U_m\setminus E_m\) gives
\begin{equation*} \lambda(U\setminus E)\le\sum_{m\in\mathbb{Z}}\varepsilon\,2^{-|m|-3}<\varepsilon/2 . \end{equation*}
As \(U\setminus E\) is measurable of finite measure, (i) supplies an open \(V\supset U\setminus E\) with \(\lambda(V)<\varepsilon\). \(\blacksquare\)
(\(\circ\)) Let \(\mu\) be a Borel probability measure on the cube \(I = [0,1]^n\) such that \(\mu(A) = \mu(B)\) for any Borel sets \(A, B \subset I\) that are translations of one another. Show that \(\mu\) coincides with Lebesgue measure \(\lambda_n\).
Every slice of \(I\) orthogonal to a coordinate axis is \(\mu\)-negligible, and this reduces the claim to the dyadic cubes. All cubes below have edges parallel to the axes.
Fix \(i\le n\) and put \(H_t:=\{x\in I:\ x_i=t\}\). Any two of these are translates of one another lying in \(I\), so \(\mu(H_t)=:c\) does not depend on \(t\in[0,1]\); were \(c>0\), then \(N>1/c\) distinct slices would be disjoint and give \(1=\mu(I)\ge Nc>1\). Hence \(c=0\), so \(\partial Q\) is \(\mu\)-negligible for every cube \(Q\subset I\) and \[ \mu(Q)=\mu(\operatorname{int}Q)\qquad\text{for every cube }Q\subset I. \tag{1} \]
Split \(I\) into the \(2^{kn}\) closed cubes \(Q_m\) of edge \(2^{-k}\). They are translates of one another inside \(I\), so all \(\mu(Q_m)\) are equal, their interiors are disjoint, and by (1) \[ 1=\mu(I)\ \ge\ \sum_m\mu(\operatorname{int}Q_m)=\sum_m\mu(Q_m)\ \ge\ \mu(I)=1, \] whence \(\mu(Q_m)=2^{-kn}=\lambda_n(Q_m)\). A cube \(Q\subset I\) of binary-rational edge \(q=p\,2^{-k}\) in grid position is the union of \(p^n\) such \(Q_m\), so \(\mu(Q)=p^n2^{-kn}=q^n\); an arbitrary cube of edge \(q\) inside \(I\) is a translate of a grid cube, so by hypothesis and (1)
\begin{equation*} \mu(Q)=q^{n}=\lambda_n(Q)\qquad\text{for every cube }Q\subset I \tag{2} \end{equation*}
with binary-rational edge (closed, open or half-open alike).
Let \(W\subset(0,1)^n\) be open. By the dyadic decomposition of an open set (proof of Lemma 1.7.2, or Exercise 1.12.72 with zero shift) \(W\) is an at most countable disjoint union of half-open dyadic cubes \(P^{(j)}\), so (2) and countable additivity give \(\mu(W)=\lambda_n(W)\). If \(W\) is merely relatively open in \(I\), then \(W\setminus(0,1)^n\subset\partial I\), a finite union of slices, which is null for \(\mu\) and for \(\lambda_n\); hence
\begin{equation*} \mu(W)=\mu\bigl(W\cap(0,1)^n\bigr)=\lambda_n\bigl(W\cap(0,1)^n\bigr)=\lambda_n(W). \end{equation*}
The relatively open subsets of \(I\) form a class stable under finite intersections, containing \(I\) and generating \(\mathcal{B}(I)\), on which the two probability measures \(\mu\) and \(\lambda_n|_I\) agree (here \(\lambda_n(I)=1\)). By Lemma 1.9.4 they agree on \(\mathcal{B}(I)\), i.e. \(\mu=\lambda_n\). \(\blacksquare\)
Exercises 1.12.75–1.12.81
(\(\circ\)) (i) Show that for any countably additive function \(\mu\colon \mathcal{R}\to[0,+\infty)\) on a semiring \(\mathcal{R}\) and any \(A, A_n\in\mathcal{R}\) such that \(A_n\) either increase or decrease to \(A\), one has the equality \(\mu(A)=\lim_{n\to\infty}\mu(A_n)\).
(ii) Give an example showing that the properties indicated in (i) do not imply the countable additivity of a nonnegative additive set function on a semiring.
(i) Convert the monotone sequence into one countable disjoint family inside \(\mathcal{R}\) and apply countable additivity.
First, \(\mu(\emptyset)=0\) (the sets \(C_j=\emptyset\) are disjoint with union \(\emptyset\in\mathcal{R}\), so \(\mu(\emptyset)=\sum_j\mu(\emptyset)\) with \(\mu(\emptyset)\) finite), hence \(\mu\) is finitely additive (pad a finite disjoint family with copies of \(\emptyset\)) and monotone (if \(B\subset C\) in \(\mathcal{R}\), then \(C=B\sqcup\bigsqcup_i D_i\) with \(D_i\in\mathcal{R}\), so \(\mu( C)\ge\mu(B)\)).
(a) \(A_n\uparrow A\). Definition 1.2.13(ii) gives disjoint \(C_{n,1},\dots,C_{n,k_n}\in\mathcal{R}\) with \(A_{n+1}\setminus A_n=\bigsqcup_j C_{n,j}\), so finite additivity applied to \(A_{n+1}=A_n\sqcup\bigsqcup_j C_{n,j}\) yields \(\mu(A_{n+1})-\mu(A_n)=\sum_j\mu(C_{n,j})\). The sets \(A_1\) and \(C_{n,j}\) form a countable disjoint family in \(\mathcal{R}\) with union \(A\in\mathcal{R}\), so countable additivity and telescoping give
\begin{equation*} \mu(A)=\mu(A_1)+\sum_{n=1}^{\infty}\bigl(\mu(A_{n+1})-\mu(A_n)\bigr) =\lim_{N\to\infty}\mu(A_N). \end{equation*}
(b) \(A_n\downarrow A\). The same decomposition of \(A_n\setminus A_{n+1}\) gives \(\mu(A_n)-\mu(A_{n+1})=\sum_j\mu(C_{n,j})\), and since \(\bigcup_n(A_n\setminus A_{n+1})=A_1\setminus A\), the set \(A\) together with all \(C_{n,j}\) is a countable disjoint family in \(\mathcal{R}\) with union \(A_1\); hence \[ \mu(A_1)=\mu(A)+\lim_{N\to\infty}\bigl(\mu(A_1)-\mu(A_N)\bigr), \] and all quantities being finite, \(\mu(A_N)\to\mu(A)\).
(ii) Take \(Q:=\mathbb{Q}\cap[0,1]\), the family \(\mathcal{R}:=\{Q\cap J:\ J\ \text{an interval with endpoints in }[0,1]\}\), and \(\mu(E):=\sup E-\inf E\) for \(E\ne\emptyset\), \(\mu(\emptyset):=0\). Density of the rationals gives \(\mu(Q\cap J)=\lambda(J)\) for every representation, so \(\mu\) is well defined.
\(\mathcal{R}\) is a semiring: \(\emptyset=Q\cap(0,0)\in\mathcal{R}\), \((Q\cap J_1)\cap(Q\cap J_2)=Q\cap(J_1\cap J_2)\), and if \(Q\cap J_1\subset Q\cap J_2\) we may replace \(J_1\) by \(J_1\cap J_2\) (this changes neither the set nor \(\mu\)), whereupon \(J_2\setminus J_1=K_1\sqcup K_2\) is a union of at most two intervals and \(E_2\setminus E_1=(Q\cap K_1)\sqcup(Q\cap K_2)\).
\(\mu\) is additive: let \(E=\bigsqcup_{i\le m}E_i\) with \(E=Q\cap J\), \(E_i=Q\cap J_i\), \(J_i\subset J\). For \(i\ne j\) the interval \(J_i\cap J_j\) contains no rational, hence is \(\lambda\)-null, and so is \(J\setminus\bigcup_iJ_i\); therefore \[ \mu(E)=\lambda(J)=\sum_{i=1}^{m}\lambda(J_i)=\sum_{i=1}^{m}\mu(E_i). \]
\(\mu\) has both continuity properties. If any of the sets involved is empty the monotonicity forces a tail of empty sets and both sides vanish, so assume all \(E_n\ne\emptyset\) and put \(\alpha_n:=\inf E_n\), \(\beta_n:=\sup E_n\), so \(\mu(E_n)=\beta_n-\alpha_n\) and \(\mathbb{Q}\cap(\alpha_n,\beta_n)\subset E_n\subset[\alpha_n,\beta_n]\).
(a) \(E_n\uparrow E\): here \(\alpha_n\) decreases, \(\beta_n\) increases, and each \(q\in E\) lies in some \(E_n\), so \(\lim\alpha_n\le\inf E\) and \(\lim\beta_n\ge\sup E\), while \(E_n\subset E\) gives the opposite inequalities. Hence \(\mu(E_n)\to\mu(E)\).
(b) \(E_n\downarrow E\): put \(\alpha^{*}:=\lim\alpha_n\), \(\beta^{*}:=\lim\beta_n\). From \(E\subset E_n\) we get \(\mu(E)\le\beta^{*}-\alpha^{*}\) (and \(\mu(E)=0\) when \(\alpha^{*}=\beta^{*}\), also for \(E=\emptyset\)); if \(\alpha^{*}<\beta^{*}\), then \(E\supset\mathbb{Q}\cap(\alpha^{*},\beta^{*})\), so \(\mu(E)\ge\beta^{*}-\alpha^{*}\). Hence \(\mu(E_n)\to\mu(E)\).
\(\mu\) is not countably additive: \(Q\) is the disjoint union of the sets \(\{q\}=Q\cap[q,q]\in\mathcal{R}\), \(q\in Q\), and \(Q\in\mathcal{R}\), yet \(\mu(Q)=1\ne 0=\sum_{q\in Q}\mu(\{q\})\). \(\blacksquare\)
(\(\circ\)) Give an example of a nonnegative additive set function \(\mu\) on a semiring \(\mathcal{R}\) such that \(\mu(A)=\lim_{n\to\infty}\mu(A_n)\) whenever \(A, A_n\in\mathcal{R}\) and \(A_n\) either increase or decrease to \(A\), but the additive extension of \(\mu\) to the ring generated by \(\mathcal{R}\) does not possess this property.
Take the semiring \(\mathcal{R}=\{Q\cap J:\ J\ \text{an interval with endpoints in }[0,1]\}\) on \(Q=\mathbb{Q}\cap[0,1]\) and \(\mu(Q\cap J)=\lambda(J)\) from Exercise 1.12.75(ii), where \(\mu\) was shown to be nonnegative, additive, and continuous along monotone sequences in \(\mathcal{R}\).
By Lemma 1.2.14 the ring \(\mathcal{R}^{\prime}\) generated by \(\mathcal{R}\) consists of the finite unions of members of \(\mathcal{R}\), each of which is a finite disjoint union of members of \(\mathcal{R}\); hence the additive extension is forced to be
\begin{equation*} \overline{\mu}(A)=\sum_{i=1}^{m}\mu(E_i)\qquad\text{for }A=\bigsqcup_{i=1}^{m}E_i,\ E_i\in\mathcal{R}, \end{equation*}
which is well defined because two such decompositions are refined by the sets \(E_i\cap F_j\in\mathcal{R}\) and \(\mu\) is additive on \(\mathcal{R}\).
Every singleton \(\{q\}=Q\cap[q,q]\) lies in \(\mathcal{R}\) with \(\mu(\{q\})=0\), so every finite subset of \(Q\) lies in \(\mathcal{R}^{\prime}\) with \(\overline{\mu}\)-measure \(0\). Enumerate \(Q=\{q_1,q_2,\dots\}\) and put \(A_n:=\{q_1,\dots,q_n\}\). Then \(A_n\uparrow Q\in\mathcal{R}^{\prime}\), while \[ \overline{\mu}(A_n)=0\ \text{ for all }n,\qquad \overline{\mu}(Q)=\mu(Q)=1 . \] The decreasing version fails too: \(B_n:=Q\setminus A_n\in\mathcal{R}^{\prime}\) satisfies \(\overline{\mu}(B_n)=1-0=1\) by additivity, yet \(B_n\downarrow\emptyset\) and \(\overline{\mu}(\emptyset)=0\). \(\blacksquare\)
(\(\circ\)) (i) Show that a bounded set \(E\subset\mathbb{R}^n\) is Jordan measurable (see Definition in 1.1) precisely when the boundary of \(E\) (the set of points each neighborhood of which contains points from the set \(E\) and from its complement) has measure zero. (ii) Show that the collection of all Jordan measurable sets in an interval or in a cube is a ring.
(i) Jordan measurability of a bounded \(E\) is exactly \(\lambda_n(\partial E)=0\), with Definition 1.7.8 read as: for every \(\varepsilon>0\) there are finite unions of cubes \(U\subset E\subset V\) with \(\lambda_n(V\setminus U)<\varepsilon\). All cubes may be taken closed, since replacing each cube of \(V\) by its closure adds only finitely many cube boundaries, a null set.
Necessity. Fix such \(U,V\) for \(\varepsilon>0\). As \(V\) is closed, \(\partial E\subset\overline E\subset V\), and \(U\subset E\) gives \(U^{\circ}\subset E^{\circ}\), which is disjoint from \(\partial E\); hence \[ \partial E\subset V\setminus U^{\circ}=(V\setminus U)\cup\partial U . \] Now \(\partial U\) lies in the union of the boundaries of the finitely many cubes constituting \(U\), so \(\lambda_n(\partial U)=0\) and \(\lambda_n(\partial E)<\varepsilon\); let \(\varepsilon\to0\).
Sufficiency. Let \(\lambda_n(\partial E)=0\) and \(\varepsilon>0\). If \(\partial E=\emptyset\), then \(E\) is open, closed and bounded in the connected space \(\mathbb{R}^n\), so \(E=\emptyset\). Otherwise cover the compact set \(\partial E\) by open cubes of total volume \(<\varepsilon\); their union \(W\) satisfies \(\lambda_n(W)<\varepsilon\) and \(W\ne\mathbb{R}^n\), so
\begin{equation*} \delta:=\operatorname{dist}\bigl(\partial E,\ \mathbb{R}^n\setminus W\bigr)>0 . \end{equation*}
Fix \(k\) with \(2^{-k}\sqrt n<\delta\) and let \(U\), resp. \(V\), be the union of the closed dyadic grid cubes of side \(2^{-k}\) contained in \(E\), resp. meeting \(E\); only finitely many meet the bounded set \(E\), so \(U\subset E\subset V\) are finite unions of cubes. A grid cube \(Q\) occurring in \(V\) but not in \(U\) meets both \(E\) and its complement, hence meets \(\partial E\) (otherwise the connected set \(Q\) would lie in one of the disjoint open sets \(E^{\circ}\), \(\mathbb{R}^n\setminus\overline E\)), and then \(\operatorname{diam}Q<\delta\) forces \(Q\subset W\). Therefore \(\lambda_n(V\setminus U)\le\lambda_n(W)<\varepsilon\).
(ii) By (i), \(\mathcal{J}=\{A\subset I:\ \lambda_n(\partial A)=0\}\), and \(\emptyset\in\mathcal{J}\). For arbitrary \(A,B\subset\mathbb{R}^n\),
\begin{equation*} \partial(A\cup B),\ \partial(A\cap B),\ \partial(A\setminus B)\ \subset\ \partial A\cup\partial B . \end{equation*}
(Check! If \(x\notin\partial A\cup\partial B\), then \(x\in A^{\circ}\) or \(x\notin\overline A\), and likewise for \(B\); in each combination some neighbourhood of \(x\) lies in, or misses, the set in question. The third inclusion follows from the second and \(\partial(\mathbb{R}^n\setminus B)=\partial B\).) For \(A,B\in\mathcal{J}\) the right-hand side is null, and the three boundaries are closed, hence measurable of measure zero; by (i) the three sets lie in \(\mathcal{J}\), which is therefore a ring by Definition 1.2.13(i). \(\blacksquare\)
(\(\circ\)) Prove Proposition 1.6.5.
The proposition follows from Theorem 1.11.4 once we know that \(\mu^{*}\) is an outer measure, that \(\mu^{*}=\mu\) on \(\mathcal{A}\), and that \(\mathcal{A}\subset\mathcal{A}_\mu\). For values in \([0,+\infty]\) the class
\begin{equation*} \mathcal{A}_\mu=\{A\subset X:\ \mu^{*}(E)=\mu^{*}(E\cap A)+\mu^{*}(E\setminus A)\ \ \forall E\subset X\} \end{equation*}
must be read in the Carathéodory sense of Definition 1.11.2, as in the derivation of the proposition from Theorem 1.11.8 in the text (for unbounded \(\mu\) the \(\varepsilon\)-approximation class of Definition 1.5.1 is not a \(\sigma\)-algebra: for \(\mu=\lambda\) on the algebra of finite unions of intervals in \(\mathbb{R}\), each \(\bigcup_{k\le N}(2k,2k+1]\) belongs to it while \(\bigcup_{k\ge1}(2k,2k+1]\) does not, its symmetric difference with any member of the algebra having infinite outer measure).
\(\mu^{*}\) is an outer measure. It is defined on all of \(X\) (as \(X\in\mathcal{A}\)), vanishes at \(\emptyset\) and is monotone. For countable subadditivity let \(A\subset\bigcup_nA_n\) with all \(\mu^{*}(A_n)<\infty\) (else nothing to prove), and pick \(B_{n,k}\in\mathcal{A}\) covering \(A_n\) with \(\sum_k\mu(B_{n,k})\le\mu^{*}(A_n)+\varepsilon2^{-n}\); the family \(\{B_{n,k}\}\) covers \(A\), so \(\mu^{*}(A)\le\sum_n\mu^{*}(A_n)+\varepsilon\).
\(\mu^{*}=\mu\) on \(\mathcal{A}\). Covering \(A\in\mathcal{A}\) by \(A,\emptyset,\emptyset,\dots\) gives \(\mu^{*}(A)\le\mu(A)\). Conversely, if \(A\subset\bigcup_nA_n\) with \(A_n\in\mathcal{A}\), the sets \(B_n:=(A\cap A_n)\setminus\bigcup_{k<n}A_k\in\mathcal{A}\) are disjoint with union \(A\in\mathcal{A}\) and \(B_n\subset A_n\), so countable additivity and monotonicity of \(\mu\) (Proposition 1.3.9, whose proof uses only additivity and nonnegativity, hence holds for values in \([0,+\infty]\)) give \[ \mu(A)=\sum_{n=1}^{\infty}\mu(B_n)\le\sum_{n=1}^{\infty}\mu(A_n); \] take the infimum over covers.
\(\mathcal{A}\subset\mathcal{A}_\mu\). By subadditivity it suffices to check \(\mu^{*}(E)\ge\mu^{*}(E\cap A)+\mu^{*}(E\setminus A)\) for \(A\in\mathcal{A}\) and \(E\) with \(\mu^{*}(E)<\infty\). Given \(\varepsilon>0\) choose \(X_n\in\mathcal{A}\) covering \(E\) with \(\sum_n\mu(X_n)<\mu^{*}(E)+\varepsilon\). Since \(X_n\cap A,\ X_n\setminus A\in\mathcal{A}\) and \(\mu(X_n)=\mu(X_n\cap A)+\mu(X_n\setminus A)\) by additivity, while \(E\cap A\subset\bigcup_n(X_n\cap A)\) and \(E\setminus A\subset\bigcup_n(X_n\setminus A)\),
\begin{equation*} \mu^{*}(E\cap A)+\mu^{*}(E\setminus A)\le\sum_n\mu(X_n)<\mu^{*}(E)+\varepsilon ; \end{equation*}
let \(\varepsilon\to0\).
Theorem 1.11.4, applied to the outer measure \(\mu^{*}\), now says that \(\mathcal{A}_\mu\) is a \(\sigma\)-algebra on which \(\mu^{*}\) is a countably additive measure with values in \([0,+\infty]\); together with the two preceding paragraphs this is Proposition 1.6.5. \(\blacksquare\)
(\(\circ\)) Show that a bounded nonnegative measure \(\mu\) on a \(\sigma\)-algebra \(\mathcal{A}\) is complete precisely when \(\mathcal{A}=\mathcal{A}_\mu\). In particular, the Lebesgue extension of any complete measure coincides with the initial measure.
Both implications use measurable hulls: since \(\mathcal{A}\) is a \(\sigma\)-algebra, the infimum in \[ \mu^{*}(S)=\min\{\mu(D):\ S\subset D\in\mathcal{A}\}\qquad (S\subset X) \] is attained (a countable cover of \(S\) by sets of \(\mathcal{A}\) may be replaced by its union, and \(D:=\bigcap_kD_k\) for \(D_k\supset S\) with \(\mu(D_k)\le\mu^{*}(S)+1/k\) realizes the minimum). Theorem 1.5.6 gives \(\mathcal{A}\subset\mathcal{A}_\mu\) and \(\mu^{*}=\mu\) on \(\mathcal{A}\).
Sufficiency. Let \(\mathcal{A}=\mathcal{A}_\mu\), let \(B\in\mathcal{A}\) with \(\mu(B)=0\) and \(A\subset B\). Then \(\mu^{*}(A)=0\), so \(A_\varepsilon:=\emptyset\) satisfies \(\mu^{*}(A\bigtriangleup A_\varepsilon)=0<\varepsilon\) in Definition 1.5.1 and \(A\in\mathcal{A}_\mu=\mathcal{A}\); thus \(\mu\) is complete.
Necessity. Let \(\mu\) be complete and \(A\in\mathcal{A}_\mu\). By Proposition 1.5.11 (applicable: \(\mu\) is a nonnegative countably additive measure on the algebra \(\mathcal{A}\) with \(\mu(X)<\infty\)), \[ \mu^{*}(A)+\mu^{*}(X\setminus A)=\mu(X). \] Choose \(B,C\in\mathcal{A}\) with \(A\subset B\), \(\mu(B)=\mu^{*}(A)\) and \(X\setminus A\subset C\), \(\mu( C)=\mu^{*}(X\setminus A)\), and put \(D:=X\setminus C\in\mathcal{A}\); then \(D\subset A\subset B\) and \(B\cup C=X\), so, \(\mu(X)\) being finite, \[ \mu(B\setminus D)=\mu(B)-\mu(D)=\mu(B)+\mu( C)-\mu(X)=0 . \] Hence \(A\setminus D\subset B\setminus D\) is a subset of a set of \(\mathcal{A}\) of measure zero, so \(A\setminus D\in\mathcal{A}\) by completeness and \(A=D\cup(A\setminus D)\in\mathcal{A}\). Therefore \(\mathcal{A}_\mu=\mathcal{A}\), and since \(\mu^{*}=\mu\) on \(\mathcal{A}\) by Theorem 1.5.6(i), the Lebesgue extension \((X,\mathcal{A}_\mu,\mu^{*})\) of a bounded complete measure is the measure itself. \(\blacksquare\)
(\(\circ\)) Give an example of a \(\sigma\)-finite measure on a \(\sigma\)-algebra that is not \(\sigma\)-finite on some sub-\(\sigma\)-algebra.
Take Lebesgue measure \(\lambda\) on the \(\sigma\)-algebra \(\mathcal{L}\) of Lebesgue measurable subsets of \(\mathbb{R}\) and the sub-\(\sigma\)-algebra
\begin{equation*} \mathcal{A}_0=\{A\subset\mathbb{R}:\ A\ \text{or}\ \mathbb{R}\setminus A\ \text{is at most countable}\}. \end{equation*}
It is a \(\sigma\)-algebra: it is symmetric under complementation, and a countable union of its members is at most countable unless some member has an at most countable complement, in which case so has the union. All its members are Borel (an at most countable set is a countable union of singletons), so \(\mathcal{A}_0\subset\mathcal{L}\) and \(\lambda_0:=\lambda|_{\mathcal{A}_0}\) is countably additive.
Now \(\lambda\) is \(\sigma\)-finite on \(\mathcal{L}\), since \(\mathbb{R}=\bigcup_n[-n,n]\) with \(\lambda([-n,n])<\infty\), whereas
\begin{equation*} \lambda_0(A)=\begin{cases}0,&A\ \text{at most countable},\\ +\infty,&\mathbb{R}\setminus A\ \text{at most countable},\end{cases} \end{equation*}
the two cases being exclusive as \(\mathbb{R}\) is uncountable. So the sets of finite \(\lambda_0\)-measure are exactly the at most countable ones, and \(\mathbb{R}=\bigcup_nX_n\) with all \(X_n\in\mathcal{A}_0\) of finite measure would make \(\mathbb{R}\) a countable union of countable sets. Hence \(\lambda_0\) is not \(\sigma\)-finite. \(\blacksquare\)
(\(\circ\)) Let \(A_n\) be subsets of a space \(X\). Show that
\begin{equation*} \{x:\ x\in A_n\ \text{for infinitely many}\ n\}=\bigcap_{n=1}^{\infty}\bigcup_{k=n}^{\infty}A_k . \end{equation*}
For \(x\in X\) the set \(S(x):=\{k\in\mathbb{N}:\ x\in A_k\}\) is infinite precisely when it is unbounded, i.e. precisely when for every \(n\) there is \(k\ge n\) with \(x\in A_k\), which is membership in \(\bigcap_{n}\bigcup_{k\ge n}A_k\).
\(\subset\): if \(S(x)\) is infinite and \(x\notin A_k\) for all \(k\ge n\), then \(S(x)\subset\{1,\dots,n-1\}\) would be finite; so \(x\in\bigcup_{k\ge n}A_k\) for every \(n\).
\(\supset\): if \(x\) lies in the right-hand side, choose recursively \(k_1<k_2<\cdots\) with \(x\in A_{k_j}\), applying the property with \(n=k_j+1\) at each step; then \(S(x)\supset\{k_1,k_2,\dots\}\) is infinite. \(\blacksquare\)
Exercises 1.12.82–1.12.88
Let \(\mu\) be a probability measure and let \(A_1,\dots,A_n\) be measurable sets with \(\sum_{i=1}^{n}\mu(A_i) > n-1\). Prove that \(\mu\bigl(\bigcap_{i=1}^{n}A_i\bigr) > 0\).
Pass to the complements \(C_i:=X\setminus A_i\): the hypothesis says exactly that \[ \sum_{i=1}^{n}\mu(C_i)=n-\sum_{i=1}^{n}\mu(A_i)<n-(n-1)=1 . \] By de Morgan’s law \(X\setminus\bigcap_{i\le n}A_i=\bigcup_{i\le n}C_i\), so finite subadditivity gives \(\mu(\bigcup_i C_i)<1\) and therefore
\begin{equation*} \mu\Bigl(\bigcap_{i=1}^{n}A_i\Bigr)=1-\mu\Bigl(\bigcup_{i=1}^{n}C_i\Bigr)\ \ge\ \sum_{i=1}^{n}\mu(A_i)-(n-1)\ >\ 0 . \qquad\blacksquare \end{equation*}
(\(\circ\)) (Baire category theorem) Let \(M_j\), \(j\in\mathbb{N}\), be closed sets in \(\mathbb{R}^d\) such that their union is a closed cube. Prove that at least one of the sets \(M_j\) has inner points. Generalize to the case where \(M_j\) are closed sets in a complete metric space \(X\) with \(\bigcup_{j=1}^{\infty}M_j=X\). A set in a metric space is called nowhere dense if its closure has no interior; a countable union of nowhere dense sets is said to be a first category set. The above result can be formulated as follows: a complete nonempty metric space is not a first category set.
We prove the general form: if \((X,d)\) is a nonempty complete metric space and the closed sets \(M_j\) cover \(X\), then some \(M_j\) has an interior point. Write \(U(x,r)\) and \(\overline U(x,r)\) for the open and the closed ball.
Suppose every \(M_j\) has empty interior. Then \(M_1\ne X\) (otherwise every point of \(M_1\) would be interior, \(X\) being open in itself and nonempty), so \(X\setminus M_1\) is open and nonempty; pick \(x_1\) in it and \(r_1\in(0,1)\) with \(\overline U(x_1,r_1)\subset X\setminus M_1\). Given \(x_j,r_j\), the nonempty open ball \(U(x_j,r_j)\) is not contained in \(M_{j+1}\), so \(U(x_j,r_j)\setminus M_{j+1}\) is open and nonempty and we may pick \(x_{j+1}\) in it together with \(r_{j+1}\in(0,r_j/2)\) such that \[ \overline U(x_{j+1},r_{j+1})\subset U(x_j,r_j)\setminus M_{j+1}. \] The closed balls decrease and \(r_j\le 2^{-(j-1)}r_1\to0\); for \(m\ge j\) we have \(x_m\in\overline U(x_j,r_j)\), so \(d(x_m,x_j)\le r_j\) and \((x_j)\) is Cauchy. Its limit \(x\) lies in every closed set \(\overline U(x_j,r_j)\), hence in no \(M_j\), contradicting \(\bigcup_jM_j=X\).
Hence if the \(A_j\) are nowhere dense, the closed sets \(\overline{A_j}\) have empty interior and \(\bigcup_jA_j\subset\bigcup_j\overline{A_j}\ne X\): a nonempty complete metric space is not of first category in itself.
The cube. Let \(Q\subset\mathbb{R}^d\) be a closed cube and \(M_j\) closed in \(\mathbb{R}^d\) with \(\bigcup_jM_j=Q\). Being closed in \(\mathbb{R}^d\), the space \(Q\) is nonempty complete and the sets \(M_j=M_j\cap Q\) are closed in it, so some \(M_{j_0}\) has an interior point relative to \(Q\): there is an open \(V\subset\mathbb{R}^d\) with \(\emptyset\ne V\cap Q\subset M_{j_0}\). Since a closed cube is the closure of its interior, \(V\) being a neighbourhood of a point of \(V\cap Q\) makes \(W:=V\cap\operatorname{int}Q\) nonempty; it is open in \(\mathbb{R}^d\) and \(W\subset M_{j_0}\), so every point of \(W\) is an interior point of \(M_{j_0}\). \(\blacksquare\)
Prove that \(\mathbb{R}^1\) cannot be written as the union of a family of pairwise disjoint nondegenerate closed intervals.
Suppose \(\mathbb{R}=\bigcup_{\alpha\in A}I_\alpha\) with \(I_\alpha=[a_\alpha,b_\alpha]\), \(a_\alpha<b_\alpha\), pairwise disjoint; the set \(F:=\{a_\alpha\}\cup\{b_\alpha\}\) of endpoints will contradict Exercise 1.12.83.
\(A\) is at most countable, since each \(I_\alpha\) contains a rational and the intervals are disjoint; also \(A\) has more than one element, a single bounded interval not being \(\mathbb{R}\). So \(F\) is a nonempty at most countable set.
\(F\) is closed. If \(x\notin F\), then \(x\in I_\alpha\) for some \(\alpha\) and \(x\in(a_\alpha,b_\alpha)\); no \(y\in(a_\alpha,b_\alpha)\) is an endpoint of some \(I_\beta\), for then \(y\in I_\beta\cap I_\alpha\) would force \(\beta=\alpha\) while \(y\ne a_\alpha,b_\alpha\). Thus \((a_\alpha,b_\alpha)\) is a neighbourhood of \(x\) missing \(F\).
\(F\) has no isolated points. Let \(\varepsilon>0\).
(i) Right endpoint \(b:=b_\alpha\). Take \(x\in(b,b+\varepsilon)\), say \(x\in I_\beta\); then \(\beta\ne\alpha\), and \(a_\beta\le b\) would give \(b\in I_\alpha\cap I_\beta\), so \(b<a_\beta\le x<b+\varepsilon\) and \(a_\beta\in F\setminus\{b\}\) with \(|a_\beta-b|<\varepsilon\).
(ii) Left endpoint \(a:=a_\alpha\). Take \(x\in(a-\varepsilon,a)\), say \(x\in I_\beta\) with \(\beta\ne\alpha\); then \(b_\beta\ge a\) would give \(a\in I_\alpha\cap I_\beta\), so \(a-\varepsilon<x\le b_\beta<a\) and \(b_\beta\in F\setminus\{a\}\) with \(|b_\beta-a|<\varepsilon\).
Every point of \(F\) is of one of these two kinds, so no point of \(F\) is isolated in \(F\). But then \(F\), a nonempty closed subset of \(\mathbb{R}\) and hence a nonempty complete metric space, is the at most countable union of the singletons \(\{x_n\}\), each closed in \(F\) and with empty interior in \(F\) – contradicting Exercise 1.12.83. \(\blacksquare\)
Show that \(\mathbb{R}^n\) with \(n>1\) cannot be written as the union of a family of closed balls with pairwise disjoint interiors.
Slice such a packing by a line meeting no tangency point and tangent to no ball, and contradict Exercise 1.12.84. Here a closed ball is \(B(c,r)\) with \(r>0\) (for \(r=0\) the assertion fails). Suppose \(\mathbb{R}^n=\bigcup_{i\in A}B_i\), \(B_i=B(c_i,r_i)\), with the interiors pairwise disjoint.
(1) \(A\) is at most countable: for the indices with \(r_i\ge1/k\) and \(|c_i|\le m\) the disjoint balls \(\{|x-c_i|<1/k\}\subset\operatorname{int}B_i\) lie in \(B(0,m+1)\), so comparing volumes bounds their number by \((m+1)^nk^n\); now take the union over \(m,k\in\mathbb{N}\).
(2) Two distinct \(B_i,B_j\) meet in at most one point. If \(D:=|c_i-c_j|<r_i+r_j\), the interiors meet: for \(D=0\) this is clear, and for \(D>0\) take \(s\) with \(\max(0,D-r_j)<s<\min(r_i,D)\) and \(p:=c_i+s(c_j-c_i)/D\), so \(|p-c_i|=s<r_i\) and \(|p-c_j|=D-s<r_j\). Hence \(D\ge r_i+r_j\), and for \(x\in B_i\cap B_j\) \[ D\le|x-c_i|+|x-c_j|\le r_i+r_j\le D , \] so \(D=r_i+r_j\) and \(x\) is the point of \([c_i,c_j]\) at distance \(r_i\) from \(c_i\). Let \(T\) be the set of these tangency points; by (1) it is at most countable.
(3) A good line. Fix \(p\notin T\) and put \(L_u:=p+\mathbb{R}u\) for \(u\in S^{n-1}\). For \(t\in T\) the set \(N_t:=\{u:\ t\in L_u\}=\{\pm(t-p)/|t-p|\}\) is closed with empty interior, since \(S^{n-1}\) has no isolated points for \(n\ge2\). For \(i\in A\) put \(w:=c_i-p\); the distance from \(c_i\) to \(L_u\) is \(g_i(u)\) with \(g_i(u)^2=|w|^2-\langle w,u\rangle^2\), so \(L_u\) meets \(B_i\) iff \(g_i(u)\le r_i\) and meets \(\operatorname{int}B_i\) iff \(g_i(u)<r_i\), and the tangency set is the closed set \(T_i=\{u:\ g_i(u)=r_i\}\). It is empty if \(w=0\) or \(|w|<r_i\); otherwise it has empty interior. Indeed, for \(u_0\in T_i\) we have \(|\langle w,u_0\rangle|=\kappa:=\sqrt{|w|^2-r_i^2}<|w|\), so \(u_0\) is not parallel to \(w\); choosing a unit \(v\in\operatorname{span}(u_0,w)\) with \(v\perp u_0\) and moving along the great circle \(u(\theta)=\cos\theta\,u_0+\sin\theta\,v\) gives
\begin{equation*} \langle w,u(\theta)\rangle=\varrho\cos(\theta-\theta_0),\qquad \varrho=\sqrt{\langle w,u_0\rangle^2+\langle w,v\rangle^2}>0 \end{equation*}
(\(\varrho=0\) would make \(w\) orthogonal to \(\operatorname{span}(u_0,v)=\operatorname{span}(u_0,w)\ni w\), i.e. \(w=0\)), and \(|\varrho\cos(\theta-\theta_0)|=\kappa\) has at most four solutions per period; so there are \(\theta_m\to0\) with \(u(\theta_m)\to u_0\) and \(u(\theta_m)\notin T_i\). The sphere \(S^{n-1}\) is a nonempty closed bounded subset of \(\mathbb{R}^n\), hence a nonempty complete metric space, so by Exercise 1.12.83 it is not the union of the at most countably many closed nowhere dense sets \(N_t\) \((t\in T)\) and \(T_i\) \((i\in A)\); choose \(u\) outside all of them and set \(L:=L_u\).
(4) Each nonempty \(J_i:=L\cap B_i\) is a compact convex subset of the line, hence a closed interval, and it is nondegenerate because \(u\notin T_i\) makes \(L\cap\operatorname{int}B_i\) a nonempty open subinterval of \(J_i\). For \(i\ne j\) we have \(J_i\cap J_j\subset B_i\cap B_j\subset T\) while \(L\cap T=\emptyset\), so the nonempty \(J_i\) are pairwise disjoint, and \(\bigcup_iJ_i=L\cap\bigcup_iB_i=L\). Identifying \(L\) with \(\mathbb{R}^1\) isometrically contradicts Exercise 1.12.84. \(\blacksquare\)
(\(\circ\)) Show that the \(\sigma\)-algebra \(\mathcal{B}(\mathbb{R}^1)\) of all Borel subsets of the real line is the smallest class of sets that contains all closed sets and admits countable intersections and countable unions.
Call a family \(\mathcal{K}\) of subsets of \(\mathbb{R}\) admissible if it contains all closed sets and is stable under countable unions and countable intersections, and let \(\mathcal{E}\) be the intersection of all admissible families; this is the smallest admissible family, since \(2^{\mathbb{R}}\) is admissible and all three properties pass to intersections.
\(\mathcal{E}\subset\mathcal{B}(\mathbb{R})\), because \(\mathcal{B}(\mathbb{R})\) is itself admissible: a \(\sigma\)-algebra containing all closed sets is stable under countable unions and, by de Morgan, under countable intersections.
For the reverse inclusion put \(\mathcal{D}:=\{A:\ A\in\mathcal{E}\ \text{and}\ \mathbb{R}\setminus A\in\mathcal{E}\}\) and check that \(\mathcal{D}\) is a \(\sigma\)-algebra containing all closed sets.
(i) Let \(F\) be closed, so \(F\in\mathcal{E}\), and put \(G:=\mathbb{R}\setminus F\). If \(G=\mathbb{R}\), then \(G\) is closed; otherwise
\begin{equation*} G=\bigcup_{n=1}^{\infty}F_n,\qquad F_n:=\{x:\ \operatorname{dist}(x,\mathbb{R}\setminus G)\ge 1/n\}, \end{equation*}
where each \(F_n\) is closed by continuity of the distance function, and the union is \(G\) because \(\mathbb{R}\setminus G\) is closed, so that a point lies in \(G\) exactly when its distance to \(\mathbb{R}\setminus G\) is positive. Either way \(G\in\mathcal{E}\), so \(F\in\mathcal{D}\); in particular \(\mathbb{R},\emptyset\in\mathcal{D}\).
(ii) \(\mathcal{D}\) is stable under complementation, its definition being symmetric.
(iii) If \(A_n\in\mathcal{D}\), then \(\bigcup_nA_n\in\mathcal{E}\) and \(\mathbb{R}\setminus\bigcup_nA_n=\bigcap_n(\mathbb{R}\setminus A_n)\in\mathcal{E}\), so \(\bigcup_nA_n\in\mathcal{D}\).
Thus \(\mathcal{D}\) is a \(\sigma\)-algebra containing all closed sets, hence \(\mathcal{B}(\mathbb{R})\subset\mathcal{D}\subset\mathcal{E}\), and \(\mathcal{E}=\mathcal{B}(\mathbb{R})\). \(\blacksquare\)
(i) Prove that the union of an arbitrary family of nondegenerate closed intervals on the real line is measurable.
(ii) Prove that the union of an arbitrary family of nondegenerate rectangles in the plane is measurable.
(iii) Prove that the union of an arbitrary family of nondegenerate triangles in the plane is measurable.
All three parts follow from one density lemma; \(\lambda_n\) is Lebesgue measure, \(\lambda_n^{*}\) its outer measure, \(B(y,r)\) the closed ball, \(\omega_n:=\lambda_n(B(0,1))\), and the members of the families are closed and arbitrarily situated.
Density Lemma. Let \(\{P_\alpha\}_{\alpha\in A}\) be closed sets in \(\mathbb{R}^n\) such that for some \(\rho,c>0\)
\begin{equation*} \lambda_n\bigl(\operatorname{int}P_\alpha\cap B(y,r)\bigr)\ \ge\ c\,r^{n} \qquad (\alpha\in A,\ y\in P_\alpha,\ 0<r\le\rho). \tag{\(\ast\)} \end{equation*}
Then \(S:=\bigcup_\alpha P_\alpha\) is measurable; in fact \(S=U\cup Z\) with \(U:=\bigcup_\alpha\operatorname{int}P_\alpha\) open and \(Z:=S\setminus U\) null.
Proof. It suffices to show \(\lambda_n^{*}(Z)=0\). Otherwise, by monotonicity and countable subadditivity there is \(m\) with \(\lambda_n^{*}(Z_0)>0\) for \(Z_0:=Z\cap B(0,m)\), and \(\lambda_n^{*}(Z_0)<\infty\). Let \(H\) be a measurable hull of \(Z_0\), i.e. measurable with \(Z_0\subset H\) and \(\lambda_n(H)=\lambda_n^{*}(Z_0)\) (intersect measurable \(H_j\supset Z_0\) with \(\lambda_n(H_j)\le\lambda_n^{*}(Z_0)+1/j\)). Every measurable \(W\subset H\) disjoint from \(Z_0\) is then null, since \(Z_0\subset H\setminus W\) gives \(\lambda_n^{*}(Z_0)\le\lambda_n(H)-\lambda_n(W)\) with \(\lambda_n^{*}(Z_0)\) finite. By the Lebesgue density theorem (a consequence of the Vitali covering theorem, Theorem 5.5.2) almost every point of \(H\) is a density point of \(H\), and \(\lambda_n^{*}(Z_0)>0\) forbids \(Z_0\) to lie in the null set of exceptions, so there is \(y\in Z_0\) with
\begin{equation*} \lim_{r\to0}\frac{\lambda_n\bigl(H\cap B(y,r)\bigr)}{\lambda_n\bigl(B(y,r)\bigr)}=1. \tag{\(\ast\ast\)} \end{equation*}
Pick \(\alpha\) with \(y\in P_\alpha\). Then \(\operatorname{int}P_\alpha\subset U\) is disjoint from \(Z\supset Z_0\), so \(W:=\operatorname{int}P_\alpha\cap H\) is null, and \((\ast)\) gives for \(0<r\le\rho\)
\begin{equation*} \lambda_n\bigl(H\cap B(y,r)\bigr)\le\lambda_n\bigl(B(y,r)\bigr) -\lambda_n\bigl(\operatorname{int}P_\alpha\cap B(y,r)\bigr)+\lambda_n(W) \le(\omega_n-c)r^{n}, \end{equation*}
so the ratio in \((\ast\ast)\) stays \(\le 1-c/\omega_n<1\) – a contradiction. \(\square\)
In each part we cut the family by a scale parameter: \(A_k\) denotes the indices satisfying the stated bound and \(S_k\) the corresponding union. Nondegeneracy gives \(A=\bigcup_kA_k\) and \(S=\bigcup_kS_k\), so it suffices that each \(S_k\) be measurable, i.e. that \((\ast)\) hold for \(\{P_\alpha\}_{\alpha\in A_k}\).
(i) Intervals with \(\lambda_1(I_\alpha)\ge1/k\): for \(y\in I_\alpha\) and \(0<r\le1/(2k)\) the set \(\operatorname{int}I_\alpha\cap(y-r,y+r)\) is an interval of length at least \(r\), so \((\ast)\) holds with \(c=1\), \(\rho=1/(2k)\).
(ii) Rectangles whose shorter side has length \(\ge1/k\). Write \(P_\alpha=\{v_0+se_1+te_2:\ s\in[0,a],\ t\in[0,b]\}\) with \(e_1,e_2\) orthonormal and \(a\ge b\ge1/k\), let \(y=v_0+y_1e_1+y_2e_2\in P_\alpha\) and \(0<r\le1/(8k)\), and put \(h:=r/\sqrt2\), so \(h\le r\le b/8\le a/8\). Choose \(\sigma_1=+1\) if \(y_1\le a/2\) and \(-1\) otherwise, and \(\sigma_2\) likewise with respect to \(b/2\). The square
\begin{equation*} Q:=\{v_0+se_1+te_2:\ s\ \text{between}\ y_1,\,y_1+\sigma_1h;\ t\ \text{between}\ y_2,\,y_2+\sigma_2h\} \end{equation*}
lies in \(P_\alpha\) (for \(\sigma_1=+1\): \(y_1+h\le a/2+a/8\le a\); for \(\sigma_1=-1\): \(y_1-h\ge a/2-a/8\ge0\); similarly in \(t\) using \(h\le b/8\)) and in \(B(y,r)\), since \(y\) is a vertex of \(Q\) and \(\operatorname{diam}Q=\sqrt{2h^2}=r\). As \(\lambda_2(\partial P_\alpha)=0\) and \(\lambda_2(Q)=r^2/2\), condition \((\ast)\) holds with \(c=1/2\), \(\rho=1/(8k)\).
(iii) Triangles with all sides \(\ge1/k\) and all angles in \([1/k,\pi-1/k]\). First, if \(P\subset\mathbb{R}^2\) is compact convex with \(B(z,\delta)\subset P\) and \(\operatorname{diam}P\le d\), then for \(y\in P\) and \(0<r\le\delta\)
\begin{equation*} \lambda_2\bigl(\operatorname{int}P\cap B(y,r)\bigr)\ \ge\ \pi\Bigl(\frac{\delta}{d+\delta}\Bigr)^{2}r^{2}, \end{equation*}
because with \(t:=r/(|y-z|+\delta)\le1\) convexity gives \(y+t(B(z,\delta)-y)=B(y+t(z-y),t\delta)\subset P\), this disc lies in \(B(y,r)\) (its points are within \(t|z-y|+t\delta=r\) of \(y\)) and its area is \(\pi t^2\delta^2\ge\pi(\delta/(d+\delta))^2r^2\). (Check!)
Now let \(T=T_\alpha\) have sides \(a,b,c\), angles \(A,B,C\), circumradius \(R\), area \(\Sigma\), inradius \(\delta\), and put \(\eta:=\sin(1/k)\), so \(\sin A,\sin B,\sin C\ge\eta\). The law of sines gives \(2R\eta\le a,b,c\le 2R\) and \(2R\ge a\ge1/k\); the diameter of \(T\) is its longest side, so \(d:=\operatorname{diam}T\le 2R\), while \(\Sigma=abc/(4R)\ge2R^2\eta^3\) and \(s=(a+b+c)/2\le3R\). Hence
\begin{equation*} \delta=\frac{\Sigma}{s}\ \ge\ \frac{2R\eta^{3}}{3}\ \ge\ \frac{\eta^{3}}{3k}, \qquad \frac{\delta}{d+\delta}\ \ge\ \frac{\eta^{3}}{3+\eta^{3}}, \end{equation*}
and the displayed bound applied to the incircle yields \((\ast)\) for \(\{T_\alpha\}_{\alpha\in A_k}\) with \(\rho_k:=\eta^{3}/(3k)\le\delta\) and \(c_k:=\pi(\eta^{3}/(3+\eta^{3}))^{2}>0\). \(\blacksquare\)
(Nikodym) For any sequence of sets \(E_n\) let
\begin{equation*} \limsup_{n\to\infty}E_n:=\bigcap_{n=1}^{\infty}\bigcup_{k=n}^{\infty}E_k,\qquad \liminf_{n\to\infty}E_n:=\bigcup_{n=1}^{\infty}\bigcap_{k=n}^{\infty}E_k . \end{equation*}
Let \((X,\mathcal{A},\mu)\) be a probability space. Prove that a sequence of sets \(A_n\in\mathcal{A}\) converges to a set \(A\in\mathcal{A}\) in the Fréchet-Nikodym metric \(d(B_1,B_2)=\mu(B_1\bigtriangleup B_2)\) precisely when every subsequence in \(\{A_n\}\) contains a further subsequence \(\{E_n\}\) such that
\begin{equation*} A=\limsup_{n\to\infty}E_n=\liminf_{n\to\infty}E_n \end{equation*}
up to a measure zero set.
Convergence in \(d(B_1,B_2)=\mu(B_1\bigtriangleup B_2)\) is equivalent to the stated subsequence property; equality of sets up to measure zero means null symmetric difference, and the two directions rest on the following two facts.
(1) If \(E_n\in\mathcal{A}\) and \(\sum_n\mu(E_n\bigtriangleup A)<\infty\), then \(\mu(A\bigtriangleup\limsup_nE_n)=\mu(A\bigtriangleup\liminf_nE_n)=0\). Indeed \(N:=\limsup_n(E_n\bigtriangleup A)\) satisfies \(\mu(N)\le\sum_{n\ge m}\mu(E_n\bigtriangleup A)\to0\), i.e. \(\mu(N)=0\) (Borel-Cantelli), and for \(x\notin N\) there is \(m\) such that \(x\in E_n\) and \(x\in A\) are equivalent for all \(n\ge m\); hence \[ A\setminus N\subset\liminf_nE_n\subset\limsup_nE_n\subset A\cup N , \] so both symmetric differences with \(A\) are measurable subsets of the null set \(N\).
(2) If \(L:=\liminf_nE_n\) and \(M:=\limsup_nE_n\) satisfy \(\mu(A\bigtriangleup L)=\mu(A\bigtriangleup M)=0\), then \(\mu(E_n\bigtriangleup A)\to0\). Put \(B_m:=\bigcap_{n\ge m}E_n\uparrow L\) and \(C_m:=\bigcup_{n\ge m}E_n\downarrow M\), so \(B_m\subset E_m\subset C_m\) and \[ E_m\bigtriangleup A\subset(C_m\setminus A)\cup(A\setminus B_m). \] As \(\mu\) is finite, continuity from above along \(C_m\setminus A\downarrow M\setminus A\) and \(A\setminus B_m\downarrow A\setminus L\) gives \(\mu(C_m\setminus A)\to\mu(M\setminus A)=0\) and \(\mu(A\setminus B_m)\to\mu(A\setminus L)=0\).
Necessity. Let \(d(A_n,A)\to0\) and let \(\{A_{n_j}\}\) be any subsequence. Choose \(j_1<j_2<\cdots\) with \(\mu(A_{n_{j_i}}\bigtriangleup A)\le2^{-i}\); by (1) the further subsequence \(E_i:=A_{n_{j_i}}\) satisfies \(A=\limsup_iE_i=\liminf_iE_i\) up to a measure zero set.
Sufficiency. If \(d(A_n,A)\not\to0\), there are \(\varepsilon>0\) and a subsequence with \(\mu(A_{n_j}\bigtriangleup A)\ge\varepsilon\) for all \(j\). By hypothesis it contains a further subsequence \(\{E_i\}\) whose upper and lower limits equal \(A\) up to measure zero, so \(\mu(E_i\bigtriangleup A)\to0\) by (2), contradicting \(\mu(E_i\bigtriangleup A)\ge\varepsilon\). \(\blacksquare\)
Exercises 1.12.89–1.12.95
(\(\circ\)) Let \((X, \mathcal{A}, \mu)\) be a space with a probability measure, let \(A_n \in \mathcal{A}_\mu\), and let \[ B := \{x : x \in A_n \text{ for infinitely many } n\}, \] i.e., \(B = \bigcap_{k=1}^{\infty} \bigcup_{n=k}^{\infty} A_n\) according to Exercise 1.12.81.
(i) (Borel–Cantelli lemma) Show that if \(\sum_{n=1}^{\infty} \mu(A_n) < \infty\), then \(\mu(B) = 0\).
(ii) Prove that if \(\mu(A_n) \ge \varepsilon > 0\) for all \(n\), then \(\mu(B) \ge \varepsilon\).
(iii) (Pták) Show that if \(\mu(B) > 0\), then one can find a subsequence \(\{n_k\}\) such that \(\mu\big(\bigcap_{k=1}^{m} A_{n_k}\big) > 0\) for all \(m\).
Put \(B_k:=\bigcup_{n=k}^{\infty}A_n\), a decreasing sequence in the \(\sigma\)-algebra \(\mathcal{A}_\mu\), on which \(\mu\) is countably additive (Theorem 1.5.6); by Exercise 1.12.81 \(B=\bigcap_kB_k\in\mathcal{A}_\mu\), and \(\mu(X)=1<\infty\) gives \(\mu(B)=\lim_k\mu(B_k)\) (Proposition 1.3.3).
(i) By countable subadditivity (Proposition 1.3.9), \[ \mu(B)\le\mu(B_k)\le\sum_{n=k}^{\infty}\mu(A_n)\xrightarrow[k\to\infty]{}0 \] when \(\sum_n\mu(A_n)<\infty\), so \(\mu(B)=0\).
(ii) \(A_k\subset B_k\) gives \(\mu(B_k)\ge\mu(A_k)\ge\varepsilon\) for every \(k\), whence \(\mu(B)=\lim_k\mu(B_k)\ge\varepsilon\).
(iii) The engine is: if \(C\subset B\) is \(\mu\)-measurable with \(\mu( C)>0\) and \(N\in\mathbb{N}\), then \(\mu(C\cap A_n)>0\) for some \(n\ge N\). Indeed each point of \(C\) lies in \(A_n\) for infinitely many \(n\), hence for some \(n\ge N\), so \(C=\bigcup_{n\ge N}(C\cap A_n)\) and countable subadditivity forbids all these sets to be null.
Apply this inductively. With \(C_0:=B\), \(\mu(C_0)>0\), choose \(n_1\) with \(\mu(B\cap A_{n_1})>0\) and put \(C_1:=B\cap A_{n_1}\). Given \(n_1<\cdots<n_m\) with \[ C_m:=B\cap A_{n_1}\cap\cdots\cap A_{n_m},\qquad \mu(C_m)>0, \] the engine with \(N=n_m+1\) produces \(n_{m+1}>n_m\) with \(\mu(C_{m+1})>0\). Since \(C_m\subset\bigcap_{k\le m}A_{n_k}\), we get \(\mu(\bigcap_{k=1}^{m}A_{n_k})\ge\mu(C_m)>0\) for every \(m\). \(\blacksquare\)
(i) Construct a sequence of sets \(E_n \subset [0,1]\) of measure \(\sigma > 0\) such that the intersection of each subsequence in this sequence has measure zero.
(ii) Let \(\mu\) be a probability measure and let \(A_n\) be \(\mu\)-measurable sets such that \(\mu(A_n) \ge \varepsilon > 0\) for all \(n \in \mathbb{N}\). Show that there exists a subsequence \(n_k\) such that \(\bigcap_{k=1}^{\infty} A_{n_k}\) is nonempty.
(iii) (Erdős, Kestelman, Rogers) Let \(A_n\) be Lebesgue measurable sets in \([0,1]\) with \(\lambda(A_n) \ge \varepsilon > 0\) for all \(n \in \mathbb{N}\). Show that there exists a subsequence \(n_k\) such that \(\bigcap_{k=1}^{\infty} A_{n_k}\) is uncountable (see a stronger assertion in Exercise 3.10.107).
(i) Take the nested-partition sets. Fix \(\sigma\in(0,1)\), put \(P_0:=\{[0,1)\}\), and given a finite partition \(P_{n-1}\) of \([0,1)\) into half-open intervals, split each \(I=[a,b)\in P_{n-1}\) into \(I^{-}:=[a,a+\sigma(b-a))\) and \(I^{+}:=[a+\sigma(b-a),b)\), setting
\begin{equation*} E_n:=\bigcup_{I\in P_{n-1}}I^{-},\qquad P_n:=\{I^{-},I^{+}:\ I\in P_{n-1}\}, \end{equation*}
so that \(\lambda(E_n)=\sum_{I\in P_{n-1}}\sigma\lambda(I)=\sigma\) (for \(\sigma=1/2\) these are the sets of the hint, \(E_n=\{x:\ n\text{-th binary digit }0\}\)).
For \(n_1<\cdots<n_m\) we have \(\lambda(E_{n_1}\cap\cdots\cap E_{n_m})=\sigma^m\), by induction on \(m\): each \(E_n\) is a union of members of \(P_k\) for every \(k\ge n\) (as \(P_k\) refines \(P_n\)), so \(D:=E_{n_1}\cap\cdots\cap E_{n_{m-1}}=\bigsqcup_{J\in F}J\) for some \(F\subset P_{n_m-1}\), and
\begin{equation*} \lambda(E_{n_m}\cap D)=\sum_{J\in F}\lambda(J^{-})=\sigma\lambda(D)=\sigma^{m}. \end{equation*}
Hence any infinite subsequence has \(\lambda(\bigcap_{k}E_{n_k})\le\sigma^m\to0\), i.e. measure zero.
(ii) By Exercise 1.12.89(ii) the set \(B\) of points lying in infinitely many \(A_n\) has \(\mu(B)\ge\varepsilon>0\), so \(B\ne\emptyset\); for \(x\in B\) the indices \(n_1<n_2<\cdots\) with \(x\in A_n\) are infinite in number and \(x\in\bigcap_{k}A_{n_k}\).
(iii) One subsequence must now serve uncountably many points at once, so we build a Cantor scheme inside a set on which a limit measure is comparable to \(\lambda\).
Traces. Let \(\mathcal{R}\) be the countable family of finite unions of intervals with rational endpoints in \([0,1]\). Since \(\lambda(A_n\cap R)\in[0,1]\), a diagonal argument gives a subsequence \((A_{m_j})\) for which \(\lim_j\lambda(A_{m_j}\cap R)\) exists for every \(R\in\mathcal{R}\). For measurable \(C\) and \(\delta>0\), regularity of \(\lambda\) (Theorem 1.4.8) gives an open \(U\supset C\) with \(\lambda(U\setminus C)<\delta\); taking finitely many of its component intervals with \(\lambda(U\setminus V)<\delta\) gives \(\lambda(C\bigtriangleup V)<2\delta\), and moving the finitely many endpoints of \(V\) to nearby rationals yields \(R\in\mathcal{R}\) with \(\lambda(C\bigtriangleup R)<3\delta\). As \(|\lambda(A\cap C)-\lambda(A\cap R)|\le\lambda(C\bigtriangleup R)\) for every measurable \(A\), the sequence \((\lambda(A_{m_j}\cap C))_j\) is Cauchy; put \(\nu( C):=\lim_j\lambda(A_{m_j}\cap C)\).
\(\nu\) is a measure with \(\nu\le\lambda\). Finite additivity passes to the limit, and for \(C=\bigsqcup_kC_k\) and every \(K\)
\begin{equation*} \nu( C)=\sum_{k\le K}\nu(C_k)+\nu\Bigl(C\setminus\bigcup_{k\le K}C_k\Bigr) \le\sum_{k\le K}\nu(C_k)+\lambda\Bigl(C\setminus\bigcup_{k\le K}C_k\Bigr), \end{equation*}
where the last term tends to \(0\); with \(\sum_{k\le K}\nu(C_k)\le\nu( C)\) this gives countable additivity. Also \(\nu([0,1])=\lim_j\lambda(A_{m_j})\ge\varepsilon\).
A set where \(\nu\) dominates \(\lambda\). Put \(\eta:=\varepsilon/2\) and \(\mathcal{G}:=\{D\ \text{measurable}:\ \nu(D)<\eta\lambda(D)\}\); every \(D\in\mathcal{G}\) has \(\lambda(D)>0\), and \(\mathcal{G}\) is stable under countable disjoint unions. Choose disjoint \(D_1,D_2,\dots\in\mathcal{G}\) recursively: given \(D_1,\dots,D_k\), let \(\beta_k\) be the supremum of \(\lambda(D)\) over \(D\in\mathcal{G}\) disjoint from \(D_1\cup\cdots\cup D_k\) and pick \(D_{k+1}\) in that family with \(\lambda(D_{k+1})>\beta_k/2\) (stopping if it is empty). Put \(N:=\bigcup_kD_k\) and \(G:=[0,1]\setminus N\). The \(D_k\) are disjoint subsets of \([0,1]\), so \(\lambda(D_k)\to0\) and \(\beta_k\to0\); a measurable \(D\subset G\) with \(\nu(D)<\eta\lambda(D)\) would belong to \(\mathcal{G}\) and avoid every \(D_k\), forcing \(\lambda(D)\le\beta_k\to0\) against \(\lambda(D)>0\). Hence \(\nu(D)\ge\eta\lambda(D)\) for every measurable \(D\subset G\), while \(\nu(N)\le\eta\lambda(N)\le\eta\), so
\begin{equation*} \varepsilon\le\nu(N)+\nu(G)\le\eta+\lambda(G),\qquad \lambda(G)\ge\varepsilon/2>0 . \end{equation*}
The Cantor scheme. We construct \(j_1<j_2<\cdots\) (writing \(n_k:=m_{j_k}\)) and compact sets \(K_s\subset G\) indexed by finite binary strings such that \(\lambda(K_s)>0\), the sets \(K_{s0},K_{s1}\) are disjoint subsets of \(K_s\), and \(K_s\subset A_{n_1}\cap\cdots\cap A_{n_m}\) for \(|s|=m\ge1\). Start with a compact \(K_\emptyset\subset G\) of positive measure (Theorem 1.4.8). Splitting: for compact \(K\) with \(\lambda(K)>0\) the function \(g(t):=\lambda(K\cap[0,t])\) is nondecreasing and \(1\)-Lipschitz with \(g(0)=0\), \(g(1)=\lambda(K)\), so choosing \(t\) with \(0<g(t)<\lambda(K)\) and then \(t^{\prime}<t<t^{\prime\prime}\) with \(g(t^{\prime})>0\), \(\lambda(K)-g(t^{\prime\prime})>0\) makes \(K^{0}:=K\cap[0,t^{\prime}]\) and \(K^{1}:=K\cap[t^{\prime\prime},1]\) disjoint compact subsets of \(K\) of positive measure. Inductive step: split every \(K_s\) with \(|s|=m\); each \(K_s^{i}\subset G\) has \(\lambda(K_s^i)>0\), so
\begin{equation*} \lim_j\lambda\bigl(A_{m_j}\cap K_s^{i}\bigr)=\nu\bigl(K_s^{i}\bigr)\ge\eta\lambda\bigl(K_s^{i}\bigr)>0, \end{equation*}
and some \(j_{m+1}>j_m\) satisfies \(\lambda(A_{m_{j_{m+1}}}\cap K_s^i)>0\) simultaneously for the finitely many pairs \((s,i)\); take \(K_{si}\subset A_{n_{m+1}}\cap K_s^i\) compact of positive measure (Theorem 1.4.8).
For \(b\in\{0,1\}^{\mathbb{N}}\) the sets \(K_{b|m}\) are nonempty compact and decreasing, so they contain a point \(x_b\), and \(x_b\in\bigcap_kA_{n_k}\) by construction; if \(b\ne b^{\prime}\) first differ at \(m\), then \(K_{b|m}\) and \(K_{b^{\prime}|m}\) are disjoint, so \(x_b\ne x_{b^{\prime}}\). Hence \(\bigcap_kA_{n_k}\) has at least \(2^{\aleph_0}\) elements. \(\blacksquare\)
Let a function \(\alpha \colon \mathbb{N} \to [0,+\infty)\) be such that \(\sum_{k=1}^{\infty} \alpha(k) < \infty\). Prove that the set \(E\) of all \(x \in (0,1)\) such that, for infinitely many natural numbers \(q\), there exists a natural number \(p\) such that \(p\) and \(q\) are relatively prime and \(|x - p/q| < \alpha(q)/q\), has measure zero. In Exercise 10.10.57 in Chapter 10 see a converse assertion.
Apply the Borel-Cantelli lemma (Exercise 1.12.89(i)) to
\begin{equation*} E_q:=\{x\in(0,1):\ |x-p/q|<\alpha(q)/q\ \text{for some }p\in\mathbb{N},\ \gcd(p,q)=1\}, \end{equation*}
each a finite union of intervals, for which \(E=\bigcap_{m}\bigcup_{q\ge m}E_q\) by the definition of \(E\).
The bound is \(\lambda(E_q)\le2\alpha(q)\). If \(\alpha(q)\le1\), then \(x\in(0,1)\), \(p\in\mathbb{N}\) and \(|x-p/q|<\alpha(q)/q\) force \(p/q<1+1/q\), i.e. \(1\le p\le q\), so
\begin{equation*} E_q\subset\bigcup_{p=1}^{q}\Bigl(\frac pq-\frac{\alpha(q)}{q},\ \frac pq+\frac{\alpha(q)}{q}\Bigr), \end{equation*}
a union of \(q\) intervals of length \(2\alpha(q)/q\) (coprimality only removes admissible \(p\)); and if \(\alpha(q)>1\), then \(\lambda(E_q)\le1<2\alpha(q)\). Hence \[ \sum_{q=1}^{\infty}\lambda(E_q)\le 2\sum_{q=1}^{\infty}\alpha(q)<\infty , \] and since \(\lambda\) on \((0,1)\) is a probability measure, Exercise 1.12.89(i) gives \(\lambda(E)=0\). \(\blacksquare\)
(Gillis) Let \(E_k \subset [0,1]\) be measurable sets and let \(\lambda(E_k) \ge \alpha\) for all \(k\), where \(\alpha \in (0,1)\). Prove that for all \(p \in \mathbb{N}\) and \(\varepsilon > 0\), there exist \(k_1 < \cdots < k_p\) such that \(\lambda(E_{k_1} \cap \cdots \cap E_{k_p}) > \alpha^p - \varepsilon\).
Average the indicators over the first \(N\) sets and apply Jensen’s inequality. Fix \(p\) and \(\varepsilon>0\) and put \(g_N:=N^{-1}\sum_{k\le N}\mathbf{1}_{E_k}\colon[0,1]\to[0,1]\). Expanding the \(p\)-th power and integrating,
\begin{equation*} \int_0^1g_N^{\,p}\,d\lambda =\frac1{N^p}\sum_{k_1,\dots,k_p\le N}\lambda\bigl(E_{k_1}\cap\cdots\cap E_{k_p}\bigr), \end{equation*}
while Jensen’s inequality for the convex \(t\mapsto t^p\) on the probability space \(([0,1],\lambda)\) and \(\int_0^1g_N\,d\lambda=N^{-1}\sum_{k\le N}\lambda(E_k)\ge\alpha\) give \(\int_0^1g_N^{\,p}\,d\lambda\ge(\int_0^1g_N\,d\lambda)^p\ge\alpha^p\). Hence
\begin{equation*} \frac1{N^p}\sum_{k_1,\dots,k_p\le N}\lambda\bigl(E_{k_1}\cap\cdots\cap E_{k_p}\bigr)\ \ge\ \alpha^{p} \qquad\text{for every }N. \tag{1} \end{equation*}
Now discard the tuples with repeated entries. Those with pairwise distinct entries number \(N(N-1)\cdots(N-p+1)\), so the count \(R_N\) of the remaining ones satisfies
\begin{equation*} \frac{R_N}{N^p}=1-\prod_{i=0}^{p-1}\Bigl(1-\frac iN\Bigr)\xrightarrow[N\to\infty]{}0 . \end{equation*}
Every summand in (1) lies in \([0,1]\), so dropping those tuples costs at most \(R_N/N^p\), and since there are at most \(N^p\) tuples of distinct indices, their maximal intersection measure \(M_N\) satisfies \(M_N\ge\alpha^p-R_N/N^p\). Choose \(N\) with \(R_N/N^p<\varepsilon\): there are pairwise distinct \(k_1,\dots,k_p\le N\) with
\begin{equation*} \lambda\bigl(E_{k_1}\cap\cdots\cap E_{k_p}\bigr)\ \ge\ \alpha^{p}-\frac{R_N}{N^p}\ >\ \alpha^{p}-\varepsilon , \end{equation*}
and the intersection is independent of the order, so we may relabel them as \(k_1<\cdots<k_p\). \(\blacksquare\)
(i) Let \(E \subset [0,1]\) be a set of Lebesgue measure zero. Prove that there exists a convergent series with positive terms \(a_n\) such that, for any \(\varepsilon > 0\), the set \(E\) can be covered by a sequence of intervals \(I_n\) of length at most \(\varepsilon a_n\). (ii) Show that there is no such series that would suit every measure zero set.
(i) Take \(a_i:=2^{-i}+\sum_{m=1}^{\infty}2^{m}|J^m_i|\), where for each \(m\) the intervals \(J^m_1,J^m_2,\dots\) cover \(E\) with \(\sum_i|J^m_i|<4^{-m}\) (possible since \(\lambda(E)=0\); pad a finite family with degenerate intervals). Each \(a_i\) is finite and positive, and summing nonnegative terms in any order, \[ \sum_{i=1}^{\infty}a_i\le 1+\sum_{m=1}^{\infty}2^m4^{-m}=2<\infty . \] Given \(\varepsilon>0\), pick \(m\) with \(2^{-m}\le\varepsilon\) and put \(I_i:=J^m_i\); then \(E\subset\bigcup_iI_i\) and \[ |I_i|=2^{-m}\cdot 2^{m}|J^m_i|\le 2^{-m}a_i\le\varepsilon a_i . \]
(ii) Given a convergent series \(\sum_na_n\) with \(a_n>0\), we construct a compact null set \(E\) and an \(\varepsilon>0\) admitting no cover of \(E\) by intervals with \(|I_n|\le\varepsilon a_n\). The property is invariant under permutations of the terms (a cover \((I_n)\) for \((a_n)\) corresponds to \((I_{\pi(n)})\) for \((a_{\pi(n)})\)), so we may assume \((a_n)\) nonincreasing. Put \[ r_n:=\sum_{k\ge n}a_k\downarrow 0,\qquad W_n:=\max\bigl(1,\ r_n^{-1/2}\bigr), \] so \((W_n)\) is nondecreasing with \(W_n\to\infty\); since
\begin{equation*} \frac{a_n}{\sqrt{r_n}}=\frac{r_n-r_{n+1}}{\sqrt{r_n}}\le 2\bigl(\sqrt{r_n}-\sqrt{r_{n+1}}\bigr), \end{equation*}
telescoping gives \(S:=\sum_na_nW_n\le\sum_na_n+2\sqrt{r_1}<\infty\). Fix \[ \varepsilon:=\min\Bigl(1,\ \frac{1}{4a_1},\ \frac{1}{24S}\Bigr)>0 . \]
Construction of \(E\). Put \(\delta_0:=1\), \(d_k:=\delta_k2^{-k-1}\), with \([0,1/2]\) as the single level-\(0\) interval. Given \(\delta_k\) and the \(2^k\) level-\(k\) intervals of length \(d_k\), set
\begin{equation*} m(k):=\min\{n:\ \varepsilon a_n\le d_k\},\quad \delta_{k+1}:=\frac{1}{W_{m(k)}},\quad d_{k+1}:=\delta_{k+1}2^{-k-2}, \end{equation*}
and let each level-\(k\) interval \([a,a+d_k]\) carry the level-\((k+1)\) intervals \([a,a+d_{k+1}]\) and \([a+d_k-d_{k+1},a+d_k]\). Here \(m(k)\) exists because \(a_n\to0\); the \(d_k\) are nonincreasing, hence the \(m(k)\) nondecreasing, so \(W_{m(k)}\ge W_{m(k-1)}\), \(\delta_{k+1}\le\delta_k\) and \(2d_{k+1}=\delta_{k+1}2^{-k-1}\le d_k\): the two children lie in the parent, and the \(2^k\) level-\(k\) intervals have pairwise disjoint interiors, any two meeting at most in a common endpoint.
Let \(F_k\) be the union of the level-\(k\) intervals and \(E:=\bigcap_{k\ge0}F_k\), a nonempty compact set with \(\lambda(E)\le\lambda(F_k)=2^kd_k=\delta_k/2\). Since \(d_k\le2^{-k-1}\to0\), for fixed \(n\) we have \(\varepsilon a_n>d_k\) for large \(k\), so \(m(k)\to\infty\), \(W_{m(k)}\to\infty\), \(\delta_k\to0\) and \(\lambda(E)=0\).
A measure on \(E\). Index the level-\(k\) intervals by strings \(s\in\{0,1\}^k\) so that \(J_{s0},J_{s1}\) are the left and right children of \(J_s\). For \(x\in[0,1)\) with binary digits \(d_i(x)\) the intervals \(J_{(d_1(x),\dots,d_k(x))}\) are nested with lengths tending to \(0\), so their intersection is a single point \(c(x)\in E\); the map \(c\) is Borel (a pointwise limit of the Borel maps sending \(x\) to the left endpoint of \(J_{(d_1(x),\dots,d_k(x))}\)), and \(\mu:=\lambda\circ c^{-1}\) is a Borel probability measure concentrated on \(E\) with \[ \mu(J_s)=2^{-k}\qquad (|s|=k), \] since \(c^{-1}(J_s)\) contains \(\{x:\ d_i(x)=s_i,\ i\le k\}\) of measure \(2^{-k}\), while any further \(x\) has \(c(x)\) equal to one of the at most two endpoints \(J_s\) shares with a neighbouring level-\(k\) interval, and each point of \(E\) has at most two \(c\)-preimages, so the excess is finite.
Key estimate. Let \(I\) be an interval and \(k\ge0\). The level-\(k\) intervals contained in \(I\) have disjoint interiors and length \(d_k\), so there are at most \(|I|/d_k\) of them; those meeting \(I\) without being contained in it must contain an endpoint of \(I\), and each point lies in at most two of them, so there are at most four. As \(\mu(E\setminus F_k)=0\) and each carries mass \(2^{-k}\),
\begin{equation*} \mu(I)\le\Bigl(\frac{|I|}{d_k}+4\Bigr)2^{-k};\quad\text{in particular}\quad \mu(I)\le 5\cdot 2^{-k}\ \text{ if }|I|\le d_k. \tag{2} \end{equation*}
Conclusion. Suppose \(E\subset\bigcup_nI_n\) with \(|I_n|\le\varepsilon a_n\). Each \(\varepsilon a_n\le\varepsilon a_1\le1/4<d_0\) while \(d_k\to0\), so \(k(n):=\max\{k\ge0:\ \varepsilon a_n\le d_k\}\) is well defined and \(\mu(I_n)\le5\cdot2^{-k(n)}\) by (2). Maximality gives \(\varepsilon a_n>d_{k(n)+1}=\delta_{k(n)+1}2^{-k(n)-2}\), i.e. \[ 2^{-k(n)}<4\varepsilon a_nW_{m(k(n))}, \] and \(\varepsilon a_n\le d_{k(n)}\) means \(n\ge m(k(n))\), so \(W_{m(k(n))}\le W_n\). Hence
\begin{equation*} 1=\mu(E)\le\sum_{n=1}^{\infty}\mu(I_n)<20\varepsilon\sum_{n=1}^{\infty}a_nW_n=20\varepsilon S\le\frac56 , \end{equation*}
a contradiction. So this \(E\) is not served by \(\sum_na_n\), and no series serves every null set. \(\blacksquare\)
(Wesler; Mergelyan for \(n = 2\)) Let \(U_k\) be disjoint open balls of radii \(r_k\) in the unit ball \(U\) in \(\mathbb{R}^n\) such that \(U \setminus \bigcup_{k=1}^{\infty} U_k\) has measure zero. Show that \(\sum_{k=1}^{\infty} r_k^{n-1} = \infty\).
Suppose \(M:=\sum_kr_k^{n-1}<\infty\) and slice the packing by lines parallel to a fixed direction. For \(n=1\) the series is \(\sum_k1\), so let \(n\ge2\); write \(U_k=B(c_k,r_k)\) (with \(r_k<1\), since no single ball of the disjoint family can be \(U\)), points of \(\mathbb{R}^n\) as \((y,t)\) with \(y\in\mathbb{R}^{n-1}\), and \(L_y:=\{(y,t):\ t\in\mathbb{R}\}\).
The residual set \(R:=\overline U\setminus\bigcup_kU_k\) is compact and \(\lambda_n\)-null, being contained in \(\partial U\cup(U\setminus\bigcup_kU_k)\); moreover \(\partial U_k\subset R\), since a point of \(\partial U_k\subset\overline U\) lies in no open \(U_j\) (every neighbourhood of it meets \(U_k\), so \(j=k\) is excluded by openness and \(j\ne k\) by disjointness).
Counting function. For \(\varepsilon>0\) put \(N(y,\varepsilon):=\#\{k:\ \lambda_1(L_y\cap U_k)\ge2\varepsilon\}\). With \(y^{\prime}_k\) the projection of \(c_k\) to \(\mathbb{R}^{n-1}\), the slice \(L_y\cap U_k\) is an interval of length \(2\sqrt{r_k^2-|y-y^{\prime}_k|^2}\) for \(|y-y^{\prime}_k|<r_k\) and empty otherwise, so
\begin{equation*} \{y:\ \lambda_1(L_y\cap U_k)\ge2\varepsilon\} =\bigl\{y:\ |y-y^{\prime}_k|\le\sqrt{(r_k^2-\varepsilon^2)_+}\bigr\} \end{equation*}
is a ball in \(\mathbb{R}^{n-1}\) of measure \(\omega_{n-1}(r_k^2-\varepsilon^2)_+^{(n-1)/2}\le\omega_{n-1}r_k^{n-1}\). As \(N(\cdot,\varepsilon)\) is the sum of the indicators of these balls, it is Borel and monotone convergence for series of nonnegative functions gives
\begin{equation*} \int_{\mathbb{R}^{n-1}}N(y,\varepsilon)\,dy =\sum_{k}\omega_{n-1}(r_k^2-\varepsilon^2)_+^{(n-1)/2}\le\omega_{n-1}M<\infty \tag{3} \end{equation*}
for every \(\varepsilon>0\).
Almost every line meets infinitely many balls. Fix \(y\) with \(|y|<1\), put \(I_y:=L_y\cap U\), a nonempty open interval, and \(G_y:=L_y\cap\bigcup_kU_k\subset I_y\), whose complement in \(I_y\) lies in \(R\). Since \(\lambda_n( R)=0\), Fubini’s theorem gives \(\lambda_1(R\cap L_y)=0\) for almost every \(y\), so for such \(y\) the open set \(G_y\) has full measure in \(I_y\); each \(L_y\cap U_k\) is an interval, so distinct components of \(G_y\) come from distinct balls. Suppose \(G_y\) had only finitely many components \(J_1,\dots,J_m\), numbered from the left. Then \(m\ge1\) (else \(\lambda_1(G_y)=0<\lambda_1(I_y)\)), and their total length being \(\lambda_1(I_y)\), all gaps degenerate to points.
(i) \(m=1\): then \(J_1=L_y\cap U_k\) has the length of \(I_y\) and lies in it, so the two intervals coincide and \(\partial U_k\) meets \(\partial U\) in two points – impossible for \(U_k\subset U\) with \(r_k<1\), whose sphere meets \(\partial U\) in at most one point.
(ii) \(m\ge2\): two consecutive components share an endpoint \(p\in I_y\) lying on the boundaries of two distinct balls; disjoint open balls whose closures meet are tangent with exactly that point in common, so \(p\) belongs to the at most countable set \(T\) of tangency points of pairs \(U_k,U_{k^{\prime}}\), and \(y\) lies in the projection of \(T\), a \(\lambda_{n-1}\)-null set.
Hence for almost every \(y\) with \(|y|<1\) the line \(L_y\) meets infinitely many balls in intervals of positive length; choosing \(m\) of them and \(\varepsilon\) below half the least of their lengths gives \(N(y,\varepsilon)\ge m\), so \(N(y,\varepsilon)\uparrow\infty\) as \(\varepsilon\downarrow0\). Monotone convergence applied to \(y\mapsto N(y,1/j)\) on the set \(\{|y|<1\}\), of positive \(\lambda_{n-1}\)-measure, then forces \(\int_{\{|y|<1\}}N(y,1/j)\,dy\to\infty\), contradicting (3). Therefore \(\sum_kr_k^{n-1}=\infty\). \(\blacksquare\)
(i) Let \(\alpha = n^{-1}\), where \(n \in \mathbb{N}\). Prove that for any sets \(A\) and \(B\) in \([0,1]\) of positive Lebesgue measure, there exist points \(x, y \in [0,1]\) such that \(\lambda(A \cap [x,y]) = \alpha \lambda(A)\) and \(\lambda(B \cap [x,y]) = \alpha \lambda(B)\). (ii) Show that if \(\alpha \in (0,1)\) does not have the form \(n^{-1}\) with \(n \in \mathbb{N}\), then assertion (i) is false.
Both parts reduce, through the normalized distribution functions
\begin{equation*} a(t):=\frac{\lambda(A\cap[0,t])}{\lambda(A)},\qquad b(t):=\frac{\lambda(B\cap[0,t])}{\lambda(B)}, \end{equation*}
to Levy’s horizontal chord lemma. These are nondecreasing and Lipschitz, hence continuous, with \(a(0)=b(0)=0\) and \(a(1)=b(1)=1\), and since points are null,
\begin{equation*} \lambda(A\cap[x,y])=\lambda(A)\bigl(a(y)-a(x)\bigr),\qquad \lambda(B\cap[x,y])=\lambda(B)\bigl(b(y)-b(x)\bigr) \end{equation*}
for \(x\le y\); so the task is to find \(x\le y\) with \(a(y)-a(x)=b(y)-b(x)=\alpha\).
Chord lemma. If \(h\colon[0,1]\to\mathbb{R}\) is continuous with \(h(0)=h(1)\) and \(\alpha=1/n\), then \(h(s+\alpha)=h(s)\) for some \(s\in[0,1-\alpha]\): the values \(\varphi(j/n)\) of \(\varphi(s):=h(s+\alpha)-h(s)\) at \(j=0,\dots,n-1\) sum to \(h(1)-h(0)=0\), so either all vanish (take \(s=0\)) or two have opposite signs and the continuous \(\varphi\) vanishes between them.
(i) Put \(m:=(a+b)/2\), continuous and nondecreasing from \(0\) to \(1\), hence onto \([0,1]\). If \(m(t)=m(t^{\prime})\) with \(t<t^{\prime}\), then \(a\) and \(b\), being nondecreasing, are constant on \([t,t^{\prime}]\); therefore
\begin{equation*} \mathcal{A}(s):=a(t),\qquad \mathcal{B}(s):=b(t)\qquad\text{for any } t \text{ with } m(t)=s \end{equation*}
are unambiguously defined on \([0,1]\). They are nondecreasing (\(s<s^{\prime}\) forces \(t<t^{\prime}\) for corresponding points), and their ranges are those of \(a\) and \(b\), i.e. all of \([0,1]\); a nondecreasing function whose range is an interval has no jumps, so \(\mathcal{A},\mathcal{B}\) are continuous. By the definition of \(m\), \[ \mathcal{A}(s)+\mathcal{B}(s)=2s\qquad (s\in[0,1]). \tag{4} \] The chord lemma applied to \(h(s):=\mathcal{A}(s)-s\), continuous with \(h(0)=0=h(1)\), gives \(s\in[0,1-\alpha]\) with \(\mathcal{A}(s+\alpha)-\mathcal{A}(s)=\alpha\), whence \(\mathcal{B}(s+\alpha)-\mathcal{B}(s)=2\alpha-\alpha=\alpha\) by (4). Choosing \(x,y\) with \(m(x)=s\), \(m(y)=s+\alpha\) (so \(x<y\)) yields \(\lambda(A\cap[x,y])=\alpha\lambda(A)\) and \(\lambda(B\cap[x,y])=\alpha\lambda(B)\).
(ii) Let \(\alpha\in(0,1)\) not be of the form \(1/n\), so \(\sin(\pi/\alpha)\ne0\), and put
\begin{equation*} c:=\frac{1}{2(\pi/\alpha+1)},\qquad g(x):=c\Bigl(\sin^2\frac{\pi x}{\alpha}-x\sin^2\frac{\pi}{\alpha}\Bigr). \end{equation*}
Then \(g(0)=g(1)=0\) and \(|g^{\prime}(x)|\le c(\pi/\alpha+1)=1/2\), so \(g\) is \(1/2\)-Lipschitz, while \(\sin^2(\theta+\pi)=\sin^2\theta\) gives \[ g(x+\alpha)-g(x)=-\kappa,\qquad \kappa:=c\,\alpha\sin^2\frac{\pi}{\alpha}>0, \] for every \(x\in[0,1-\alpha]\). Approximate \(g\) by a zigzag of slopes \(\pm1/2\): fix \(\eta<\kappa/2\) with \(1/\eta\in\mathbb{N}\) and nodes \(t_i:=i\eta\); as \(|g(t_{i+1})-g(t_i)|\le\eta/2\), let \(F\) rise on \([t_i,t_{i+1}]\) from \(g(t_i)\) with slope \(1/2\) for the time \(u_i:=g(t_{i+1})-g(t_i)+\eta/2\in[0,\eta]\) and then fall with slope \(-1/2\), ending at \(g(t_i)+u_i-\eta/2=g(t_{i+1})\). The glued function \(F\) is continuous and piecewise linear with \(F^{\prime}=\pm1/2\) off a finite set and \(F(t_i)=g(t_i)\); on each \([t_i,t_{i+1}]\) both functions are \(1/2\)-Lipschitz and agree at \(t_i\), so \(\|F-g\|_\infty\le\eta\). Hence \(F(0)=F(1)=0\) and \[ F(x+\alpha)-F(x)\le-\kappa+2\eta<0\qquad (x\in[0,1-\alpha]). \tag{5} \]
Put \(A:=\{t:\ F^{\prime}(t)=1/2\}\), a finite union of intervals, and \(B:=[0,1]\setminus A\). Being piecewise linear, \(F\) is absolutely continuous with \(F(0)=0\) and \(F^{\prime}=\mathbf{1}_A-1/2\) off a finite set, so \[ F(x)=\lambda(A\cap[0,x])-\frac x2 , \] and \(F(1)=0\) gives \(\lambda(A)=\lambda(B)=1/2>0\). If \(x,y\) satisfied both required equalities, then \(x\le y\) (otherwise \([x,y]=\emptyset\) while \(\alpha\lambda(A)>0\)), and adding \(\lambda(A\cap[x,y])=\lambda(B\cap[x,y])=\alpha/2\) gives \(y-x=\alpha\); but then \[ F(x+\alpha)-F(x)=\lambda\bigl(A\cap[x,x+\alpha]\bigr)-\frac\alpha2=0, \] contradicting (5). \(\blacksquare\)
Exercises 1.12.96–1.12.102
A set \(S \subset \mathbb{R}^1\) is called a Sierpinski set if \(S \cap Z\) is at most countable for every set \(Z\) of Lebesgue measure zero.
(i) Under the continuum hypothesis show the existence of a Sierpinski set.
(ii) Prove that no Sierpinski set is measurable.
As assertion (ii) requires, a Sierpinski set is understood to be an uncountable \(S\subset\mathbb{R}^1\) meeting every Lebesgue null set in an at most countable set; without uncountability every countable set would qualify and (ii) would fail.
(i) Transfinite recursion through the null \(G_\delta\) sets produces one. Two facts: every null \(Z\) lies in a null \(G_\delta\) set (take \(\bigcap_nU_n\) for open \(U_n\supset Z\) with \(\lambda(U_n)<1/n\)), and there are at most \(\mathfrak{c}=2^{\aleph_0}\) of them, since every open subset of \(\mathbb{R}\) is the union of the rational intervals it contains, leaving at most \(\mathfrak{c}\) open sets and \(\mathfrak{c}^{\aleph_0}=\mathfrak{c}\) countable intersections.
Assume the continuum hypothesis, \(\mathfrak{c}=\aleph_1\), and enumerate the null \(G_\delta\) sets as \(\{Z_\alpha:\ \alpha<\omega_1\}\) (repeating sets if necessary). Recursively, for \(\alpha<\omega_1\) the ordinal \(\alpha\) is countable, so \[ N_\alpha:=\bigcup_{\beta\le\alpha}Z_\beta \] is null, \(\mathbb{R}\setminus N_\alpha\) has full measure and hence cardinality \(\mathfrak{c}\), and removing the at most countable set \(\{x_\beta:\ \beta<\alpha\}\) leaves it nonempty; choose \(x_\alpha\) in the remainder and put \(S:=\{x_\alpha:\ \alpha<\omega_1\}\), of cardinality \(\aleph_1\), hence uncountable. If \(Z\) is null, then \(Z\subset Z_\gamma\) for some \(\gamma<\omega_1\), and for \(\alpha\ge\gamma\) the choice of \(x_\alpha\) excludes \(Z_\gamma\subset N_\alpha\); hence \(S\cap Z\subset\{x_\beta:\ \beta<\gamma\}\) is at most countable.
(ii) Lemma: every compact \(K\subset\mathbb{R}\) with \(\lambda(K)>0\) contains an uncountable null set.
Splitting claim: if \(J=[a,b]\) has \(\lambda(K\cap J)>0\) and \(\delta>0\), then \(J\) contains disjoint closed intervals \(J^{\prime},J^{\prime\prime}\) of length at most \(\delta\) with \(\lambda(K\cap J^{\prime})>0\) and \(\lambda(K\cap J^{\prime\prime})>0\). Indeed \(g(t):=\lambda(K\cap[a,t])\) is nondecreasing and \(1\)-Lipschitz with \(g(a)=0\) and \(g(b)=:c>0\), so \(g(t_0)=c/2\) for some \(t_0\in(a,b)\), whence \(\lambda(K\cap[t_0,b])\ge c/2>0\), while \(g(s)\to c/2>0\) as \(s\uparrow t_0\) gives \(s<t_0\) with \(\lambda(K\cap[a,s])>0\); now partition the disjoint intervals \([a,s]\) and \([t_0,b]\) into finitely many closed intervals of length at most \(\delta\) with disjoint interiors and pick in each a piece meeting \(K\) in positive measure (finite sets being null).
Build closed intervals \(J_s\) indexed by finite \(0\)-\(1\) words: let \(J_\emptyset\supset K\), and given \(\lambda(K\cap J_s)>0\) with \(|s|=n\), apply the claim with \(\delta=4^{-(n+1)}\) to get disjoint \(J_{s0},J_{s1}\subset J_s\) of length at most \(4^{-(n+1)}\) with \(\lambda(K\cap J_{si})>0\). The sets \(F_s:=K\cap J_s\) are nonempty compact with \(F_{s0},F_{s1}\subset F_s\) disjoint; put \[ P:=\bigcap_{n\ge1}\ \bigcup_{|s|=n}F_s\ \subset K . \] For each \(n\) the set \(P\) is covered by \(2^n\) intervals of length at most \(4^{-n}\), so \(\lambda(P)\le2^{-n}\to0\), i.e. \(\lambda(P)=0\). And \(P\) is uncountable: for an infinite word \(\sigma\) the compact sets \(F_{\sigma|n}\) decrease, hence contain a point \(x_\sigma\in P\), and if \(\sigma\ne\tau\) first differ at \(n\), then \(F_{\sigma|n}\cap F_{\tau|n}=\emptyset\) and \(x_\sigma\ne x_\tau\); so \(P\) has at least \(2^{\aleph_0}\) points.
Now let a Sierpinski set \(S\) be measurable. If \(\lambda(S)=0\), then \(Z:=S\) makes \(S=S\cap Z\) at most countable, against uncountability. So \(\lambda(S)>0\) and \(\lambda(S\cap[-n,n])>0\) for some \(n\); a measurable set contains a Borel set of the same measure, so Theorem 1.4.8 applied to \(\lambda\) on \([-n,n]\) yields a compact \(K\subset S\) with \(\lambda(K)>0\). By the lemma \(K\) contains an uncountable null set \(Z\), and then \(S\cap Z=Z\) is uncountable – contradiction. \(\blacksquare\)
Let \(A\) be a set in \(\mathbb{R}^d\) of Lebesgue measure greater than \(1\). Prove that there exist two distinct points \(x, y \in A\) such that the vector \(x - y\) has integer coordinates.
Fold \(A\) into the unit cube \(Q = [0,1)^d\) and apply the pigeonhole principle. The translates \(k + Q\), \(k \in \mathbb{Z}^d\), partition \(\mathbb{R}^d\), so putting \(A_k = A \cap (k+Q)\) and \(B_k = A_k - k \subset Q\), translation invariance of \(\lambda_d\) and countable additivity over the countable index set \(\mathbb{Z}^d\) give
\begin{equation*} \sum_{k} \lambda_d(B_k) = \sum_{k} \lambda_d(A_k) = \lambda_d(A) > 1 = \lambda_d(Q) . \end{equation*}
Were the \(B_k \subset Q\) pairwise disjoint, the left side would be \(\lambda_d(\bigcup_k B_k) \le \lambda_d(Q) = 1\); hence \(z \in B_k \cap B_{k^{\prime}}\) for some \(k \ne k^{\prime}\). Then
\begin{equation*} x = z + k \in A_k \subset A, \qquad y = z + k^{\prime} \in A_{k^{\prime}} \subset A \end{equation*}
satisfy \(x - y = k - k^{\prime} \in \mathbb{Z}^d \setminus \{0\}\), so \(x \ne y\) and \(x-y\) has integer coordinates.
(\(\circ\)) Prove that each convex set in \(\mathbb{R}^d\) is Lebesgue measurable.
The boundary of a convex set is \(\lambda_d\)-null, so \(A = \mathrm{int}\, A \cup (A \cap \partial A)\) is an open set together with a subset of a null set, hence measurable by completeness of \(\lambda_d\). Since \(A = \bigcup_{n \ge 1}(A \cap U_n)\) with \(U_n\) the ball of radius \(n\) about the origin, each piece convex and bounded, we may assume \(A\) bounded.
The one convexity fact used: if \(a \in \mathrm{int}\, A\) and \(y \in \overline{A}\), then \(sa + (1-s)y \in \mathrm{int}\, A\) for \(s \in (0,1]\) – with \(B(a,r) \subset A\) and \(y^{\prime} \in A\) such that \((1-s)|y - y^{\prime}| < sr/2\), convexity gives \((1-s)y^{\prime} + sB(a,r) \subset A\), a ball of radius \(sr\) whose centre lies within \(sr/2\) of \((1-s)y + sa\).
(i) \(\mathrm{int}\, A = \emptyset\): then \(A\) contains no \(d+1\) affinely independent points (they would span a \(d\)-simplex, which has interior), so \(\overline{A}\) lies in an affine hyperplane and \(\lambda_d(\partial A) = 0\).
(ii) \(\mathrm{int}\, A \ne \emptyset\): translate so that \(0 \in \mathrm{int}\, A\); the fact above with \(a = 0\), \(s = 1-t\) gives \(t\overline{A} \subset \mathrm{int}\, A\) for \(t \in (0,1)\), so \(\partial A \subset \overline{A} \setminus t\overline{A}\), and \(\overline{A}\) being compact with \(\lambda_d(t\overline{A}) = t^d \lambda_d(\overline{A})\),
\begin{equation*} \lambda_d(\partial A) \le (1 - t^d)\, \lambda_d(\overline{A}) \xrightarrow[t \to 1^-]{} 0 . \end{equation*}
Let \(A\) be a bounded convex set in \(\mathbb{R}^d\) and let \(A^\varepsilon\) be the set of all points with the distance from \(A\) at most \(\varepsilon\). Prove that \(\lambda_d(A^\varepsilon)\), where \(\lambda_d\) is Lebesgue measure, is a polynomial of degree \(d\) in \(\varepsilon\).
This is Steiner’s formula: \(\lambda_d(A^\varepsilon)\) is a polynomial of degree exactly \(d\) with leading coefficient \(\kappa_d = \lambda_d(B)\), \(B\) the closed unit ball (the nonempty case being meant). Since \(\mathrm{dist}(x,A) = \mathrm{dist}(x,\overline{A})\) and the distance to a compact set is attained, \(A^\varepsilon = K + \varepsilon B\) with \(K := \overline{A}\) compact convex; after a translation \(0 \in K \subset RB\). With \(f(\varepsilon) := \lambda_d(K + \varepsilon B)\), monotonicity between a ball about a point of \(K\) and \((R+\varepsilon)B\) gives, for \(\varepsilon \ge 0\),
\begin{equation*} \kappa_d \varepsilon^d \le f(\varepsilon) \le \kappa_d (R + \varepsilon)^d . \tag{1} \end{equation*}
(i) Polyhedra. Let \(P = \bigcap_{i \le m} \{ \langle a_i, x\rangle \le b_i \}\) be nonempty and bounded, \(I(x)\) the set of constraints active at \(x \in P\), and \(G_I = \{x \in P : I(x) = I\}\): finitely many disjoint convex Borel sets with union \(P\). The normal cone is constant on each \(G_I\): for \(x, y \in G_I\) and \(u \in N_P(x)\), \(x + t(x-y) \in P\) for small \(t > 0\) forces \(\langle u, x-y\rangle \le 0\), while \(y \in P\) gives \(\langle u, y-x\rangle \le 0\), so
\begin{equation*} \langle u, x-y\rangle = 0, \qquad \langle u, w-y\rangle \le 0 \quad (w \in P), \end{equation*}
i.e. \(u \in N_P(y)\); thus \(N_P \equiv N_I\) on \(G_I\) with \(N_I \subset L_I^{\perp}\), where \(L_I\) is the direction space of the affine hull of \(G_I\) and \(j_I = \dim L_I\). The characterization \(\pi(x) = p\) iff \(x - p \in N_P(p)\) of the metric projection onto a closed convex set gives \(\pi^{-1}(G_I) = G_I + N_I\), the sum direct since \(L_I \cap L_I^{\perp} = \{0\}\); as \(\mathrm{dist}(x,P) = |x - \pi(x)|\), the nonempty \(G_I\) yield the disjoint decomposition
\begin{equation*} P + \varepsilon B = \bigsqcup_{I} \bigl( G_I + (N_I \cap \varepsilon B) \bigr). \end{equation*}
Identifying \(\mathbb{R}^d\) isometrically with \(L_I \times L_I^{\perp}\) based at some \(x_I \in G_I\), under which \(\lambda_d = \lambda_{j_I} \otimes \lambda_{d-j_I}\), Fubini and \(N_I \cap \varepsilon B = \varepsilon(N_I \cap B)\) give
\begin{equation*} \begin{aligned} \lambda_d\bigl( G_I + (N_I \cap \varepsilon B) \bigr) &= c_I \, \varepsilon^{\,d - j_I}, \\ c_I &= \lambda_{j_I}(G_I - x_I)\, \lambda_{d-j_I}(N_I \cap B), \end{aligned} \end{equation*}
all factors finite since \(P\) is bounded, with unit mass on \(\{0\}\) in the degenerate cases \(j_I \in \{0,d\}\). Summing, \(\varepsilon \mapsto \lambda_d(P + \varepsilon B)\) is a polynomial of degree at most \(d\).
(ii) Passage to \(K\). For \(\delta \in (0,1/2)\) take a finite \(\delta\)-net \(S_\delta\) of the unit sphere and \(P_\delta = \bigcap_{u \in S_\delta}\{\langle u,x\rangle \le h_K(u)\}\), \(h_K\) the support function of \(K\); then \(K \subset P_\delta \subset 2RB\), since for \(|x| = \rho\) the net vector \(u\) near \(x/|x|\) has \(\langle u,x\rangle \ge (1-\delta)\rho\) while \(h_K(u) \le R\). Also \(P_\delta \subset K + \eta_\delta B\) with \(\eta_\delta := 4R\delta\): if \(|x| \le 2R\) and \(x \notin K + \eta B\) with \(\eta > 3R\delta\), separating the compact convex \(K + \eta B\) gives a unit \(u\) with \(\langle u,x\rangle > h_K(u) + \eta\), and the nearby \(u^{\prime} \in S_\delta\) satisfies, by \(|\langle u-u^{\prime},x\rangle| \le 2R\delta\) and \(|h_K(u)-h_K(u^{\prime})| \le R\delta\),
\begin{equation*} \langle u^{\prime}, x\rangle > h_K(u^{\prime}) + \eta - 3R\delta > h_K(u^{\prime}), \end{equation*}
so \(x \notin P_\delta\). Hence, with \(f_\delta(\varepsilon) := \lambda_d(P_\delta + \varepsilon B)\) and \(\varepsilon \ge 0\),
\begin{equation*} f(\varepsilon) \le f_\delta(\varepsilon) \le f(\varepsilon + \eta_\delta) . \tag{2} \end{equation*}
Fix \(\varepsilon > 0\). As \(C = K + \varepsilon B\) contains \(\overline{B}(0,\varepsilon)\), for \(x \in C\), \(|u| \le \eta_\delta\) and \(t = \varepsilon/(\varepsilon + \eta_\delta)\) convexity gives \(t(x+u) = tx + (1-t)(\varepsilon/\eta_\delta)u \in C\), so \(C + \eta_\delta B \subset t^{-1}C\) and \(f(\varepsilon + \eta_\delta) \le (1 + \eta_\delta/\varepsilon)^{d} f(\varepsilon)\). With (2) and \(\eta_\delta \to 0\), \(f_\delta(\varepsilon) \to f(\varepsilon)\) for each \(\varepsilon > 0\). Fixing nodes \(0 < \varepsilon_0 < \dots < \varepsilon_d\) with Lagrange basis \(L_i\), each \(f_\delta = \sum_i f_\delta(\varepsilon_i)L_i\) by (i), so letting \(\delta \to 0\),
\begin{equation*} f(\varepsilon) = \sum_{i=0}^{d} f(\varepsilon_i)\, L_i(\varepsilon) =: Q(\varepsilon), \quad \varepsilon > 0 , \end{equation*}
and \(K + \varepsilon B \downarrow K\) with finite measures gives \(f(0) = Q(0)\) by continuity from above.
(iii) Degree. Dividing (1) by \(\varepsilon^d\) and letting \(\varepsilon \to +\infty\) gives \(Q(\varepsilon)/\varepsilon^d \to \kappa_d > 0\), so \(\deg Q = d\) with leading coefficient \(\kappa_d\).
(\(\circ\)) Prove Theorem 1.12.1.
Everything follows from one device, the trace class \(\mathcal{F}_A\) below; notation is that of Theorem 1.12.1, \(\mathcal{G}\) is tacitly nonempty (for \(\mathcal{G} = \emptyset\) all closure requirements are vacuous, so \(\mathcal{F}_0 = \emptyset\) while \(\sigma(\mathcal{G}) = \{\emptyset, X\}\)), and \(\mathcal{F}_0 \subset \mathcal{F}\) by minimality.
For \(A \subset X\) and any \(O\) from the list, intersecting an admissible family with \(A\) again yields an admissible family (monotone stays monotone, disjoint stays disjoint, \(B_2 \subset B_1\) gives \(A \cap B_2 \subset A \cap B_1\)), and
\begin{equation*} A \cap O(B_1, B_2, \dots) = O(A \cap B_1, A \cap B_2, \dots) \end{equation*}
by distributivity for the unions, intersections and limits, and by \(A \cap (B_1 \setminus B_2) = (A \cap B_1) \setminus (A \cap B_2)\) for \(\setminus\) and \(-\). Hence
\begin{equation*} \mathcal{F}_A := \{ B \in \mathcal{F}_0 : A \cap B \in \mathcal{F}_0 \} \end{equation*}
is closed with respect to every \(O \in \mathcal{O}\), for any \(A \subset X\).
(i) Fix \(G \in \mathcal{G}\): the hypothesis says \(\mathcal{G} \subset \mathcal{F}_G\), so minimality gives \(\mathcal{F}_0 = \mathcal{F}_G\), i.e. \(G \cap F \in \mathcal{F}_0\) for all \(G \in \mathcal{G}\), \(F \in \mathcal{F}_0\). Fixing now \(F \in \mathcal{F}_0\), this says \(\mathcal{G} \subset \mathcal{F}_F\), so again \(\mathcal{F}_0 = \mathcal{F}_F\) and \(F \cap F^{\prime} \in \mathcal{F}_0\) for all \(F, F^{\prime} \in \mathcal{F}_0\); induct for finite intersections.
(ii) Two cases, according to \(\mathcal{O}\).
Case A: \(\mathcal{O} \subset \{\cup f, \cup c, \lim \uparrow, \cap f, \cap c, \lim \downarrow\}\). For each of these six, complementation carries an admissible family to one admissible for the dual operation and De Morgan gives
\begin{equation*} X \setminus O(B_1, B_2, \dots) = O^d(X \setminus B_1, X \setminus B_2, \dots) \end{equation*}
(for \(\lim \uparrow\): complements of an increasing sequence decrease, and their intersection is the complement of the union). Hence \(\mathcal{F}_1 := \{F \in \mathcal{F}_0 : X \setminus F \in \mathcal{F}_0\}\) is closed with respect to \(\mathcal{O}\), since \(O^d \in \mathcal{O}\); as \(\mathcal{G} \subset \mathcal{F}_1\) by hypothesis, minimality forces \(\mathcal{F}_1 = \mathcal{F}_0\).
Case B: \(\mathcal{O}\) contains one of \(\sqcup f, \sqcup c, \setminus, -\). Then \(\mathcal{F}_0\) is closed with respect to \(-\): directly if \(- \in \mathcal{O}\); because a proper difference is a difference if \(\setminus \in \mathcal{O}\); and because \(- = (\sqcup f)^d = (\sqcup c)^d \in \mathcal{O}\) by duality otherwise. Also \(X \in \mathcal{F}_0\): for \(G \in \mathcal{G}\) both \(G\) and \(X \setminus G\) lie in \(\mathcal{F}_0\) and are disjoint with union \(X\), and \(\mathcal{O}\) contains \(\cup f = (\setminus)^d\) or \(\sqcup f = (-)^d\). Hence \(X \setminus F = X - F \in \mathcal{F}_0\) for every \(F \in \mathcal{F}_0\).
For \(\mathcal{O} = (\cup c, \cap c)\) (self-dual, Case A) the class \(\mathcal{F}_0\) is closed under countable unions and complements and contains \(X = G \cup (X \setminus G)\), so it is a \(\sigma\)-algebra containing \(\mathcal{G}\); minimality in the other direction gives \(\mathcal{F}_0 = \sigma(\mathcal{G})\).
(iii) Under both sets of hypotheses \(\mathcal{F}_0\) is closed under finite intersections and complements; \(\emptyset = G \cap (X \setminus G) \in \mathcal{F}_0\) gives \(X \in \mathcal{F}_0\), and
\begin{equation*} A \cup B = X \setminus \bigl( (X \setminus A) \cap (X \setminus B) \bigr) \in \mathcal{F}_0 , \end{equation*}
so \(\mathcal{F}_0\) is an algebra containing \(\mathcal{G}\), whence the algebra generated by \(\mathcal{G}\) lies in \(\mathcal{F}_0 \subset \mathcal{F}\). For \(\mathcal{O} = (\lim \uparrow, \lim \downarrow)\) (self-dual) this algebra \(\mathcal{F}_0\) is closed under increasing limits, so \(\bigcup_n A_n = \lim \uparrow \bigcup_{k \le n} A_k \in \mathcal{F}_0\) and \(\mathcal{F}_0\) is a \(\sigma\)-algebra; minimality both ways gives \(\mathcal{F}_0 = \sigma(\mathcal{G})\).
(\(\circ\)) Let \((X, \mathcal{A}, \mu)\) be a probability space, \(\mathcal{B}\) a sub-\(\sigma\)-algebra in \(\mathcal{A}\), and let \(\mathcal{B}^\mu\) be the \(\sigma\)-algebra generated by \(\mathcal{B}\) and all sets of measure zero in \(\mathcal{A}_\mu\).
(i) Show that \(E \in \mathcal{B}^\mu\) precisely when there exists a set \(B \in \mathcal{B}\) such that \(E \triangle B \in \mathcal{A}_\mu\) and \(\mu(E \triangle B) = 0\).
(ii) Give an example demonstrating that \(\mathcal{B}^\mu\) may be strictly larger than the \(\sigma\)-algebra \(\mathcal{B}_\mu\) that is the completion of \(\mathcal{B}\) with respect to the measure \(\mu|_{\mathcal{B}}\).
(i) The class \(\mathcal{E}\) of sets \(E\) admitting \(B \in \mathcal{B}\) with \(E \triangle B \in \mathcal{N}\) is exactly \(\mathcal{B}^\mu = \sigma(\mathcal{B} \cup \mathcal{N})\), where \(\mathcal{N}\) is the class of \(\mu\)-null sets of \(\mathcal{A}_\mu\) (hereditary and stable under countable unions). One inclusion: for such \(E\) the pieces \(E \setminus B\), \(B \setminus E\) lie in \(\mathcal{N} \subset \mathcal{B}^\mu\), so
\begin{equation*} E = \bigl( B \setminus (B \setminus E) \bigr) \cup (E \setminus B) \in \mathcal{B}^\mu . \end{equation*}
The other: \(\mathcal{E}\) is a \(\sigma\)-algebra, since \(X \triangle X = \emptyset\), since \((X \setminus E) \triangle (X \setminus B) = E \triangle B\), and since
\begin{equation*} \bigcup_n E_n \ \triangle \ \bigcup_n B_n \ \subset \ \bigcup_n ( E_n \triangle B_n ) , \end{equation*}
the right side being null; and it contains \(\mathcal{B}\) (take \(B = E\)) and \(\mathcal{N}\) (take \(B = \emptyset\)), hence contains \(\sigma(\mathcal{B} \cup \mathcal{N})\). As \(\mathcal{B}^\mu \subset \mathcal{A}_\mu\), the requirement \(E \triangle B \in \mathcal{A}_\mu\) is automatic.
(ii) Take \(X = [0,1]\) with \(\mathcal{A}\) the Lebesgue measurable sets and \(\mu = \lambda\) (complete, so \(\mathcal{A}_\mu = \mathcal{A}\)), and \(\mathcal{B} = \{\emptyset, [0,1]\}\). The only set of \(\mathcal{B}\) of measure zero is \(\emptyset\), so \(\mathcal{B}_\mu = \mathcal{B}\). By (i), the only possible witnesses being \(B = \emptyset\) and \(B = [0,1]\),
\begin{equation*} \mathcal{B}^\mu = \{ E : \lambda(E) = 0 \ \text{ or } \ \lambda([0,1] \setminus E) = 0 \} \end{equation*}
(such \(E\) are measurable by completeness). The Cantor set is null and differs from \(\emptyset\) and \([0,1]\), so it lies in \(\mathcal{B}^\mu \setminus \mathcal{B}_\mu\).
(\(\circ\)) Let \(\mu\) be a probability measure on a \(\sigma\)-algebra \(\mathcal{A}\). Suppose that \(\mathcal{A}\) is countably generated, i.e., is generated by an at most countable family of sets. Show that the measure \(\mu\) is separable. Give an example showing that the converse is false.
The algebra generated by a countable family is itself countable and dense in \((\mathcal{A}/\mu, d)\), \(d(A,B) = \mu(A \triangle B)\), which is separability.
Let \(\mathcal{A} = \sigma(\mathcal{E})\) with \(\mathcal{E}\) at most countable and let \(\mathcal{A}_0\) be the algebra generated by \(\mathcal{E}\). Putting \(\mathcal{C}_0 = \mathcal{E} \cup \{\emptyset, X\}\) and
\begin{equation*} \mathcal{C}_{k+1} = \mathcal{C}_k \cup \{ X \setminus C \} \cup \{ C \cup C^{\prime} \} \quad (C, C^{\prime} \in \mathcal{C}_k), \end{equation*}
each \(\mathcal{C}_k\) is at most countable, the increasing union \(\mathcal{C} = \bigcup_k \mathcal{C}_k\) is an algebra containing \(\mathcal{E}\) and contained in every such algebra, so \(\mathcal{A}_0 = \mathcal{C}\) is at most countable with \(\sigma(\mathcal{A}_0) = \mathcal{A}\). Now \(\mu\) restricted to the algebra \(\mathcal{A}_0\) is countably additive, so by Theorem 1.5.6 the class \((\mathcal{A}_0)_\mu\) of sets \(A\) with \(\inf_{A_0 \in \mathcal{A}_0} \mu^{*}(A \triangle A_0) = 0\) is a \(\sigma\)-algebra and hence contains \(\sigma(\mathcal{A}_0) = \mathcal{A}\). Thus the at most countably many classes \([A_0]\), \(A_0 \in \mathcal{A}_0\), are dense in \(\mathcal{A}/\mu\).
The converse fails: take \(X = [0,1]\),
\begin{equation*} \mathcal{A} = \{ A \subset X : A \ \text{or} \ X \setminus A \ \text{is at most countable} \}, \end{equation*}
and \(\mu(A) = 0\) or \(1\) according to which alternative holds (well defined since \(X\) is uncountable; countably additive since in a disjoint countable family at most one member can be co-countable, the two alternatives then giving \(0 = 0\) and \(1 = 1\)). Every \(A \in \mathcal{A}\) satisfies \(\mu(A \triangle \emptyset) = 0\) or \(\mu(A \triangle X) = 0\), so \(\mathcal{A}/\mu\) has two points and \(\mu\) is separable.
But \(\mathcal{A}\) is not countably generated. Suppose \(\mathcal{A} = \sigma(\{A_n\})\) and let \(S\) be the union of those \(A_n\) that are countable together with those complements \(X \setminus A_n\) that are countable, an at most countable set. Set \(x \sim y\) iff \(x\) and \(y\) lie in the same \(A_n\) for every \(n\); the unions of \(\sim\)-classes form a \(\sigma\)-algebra \(\mathcal{D}\) containing every \(A_n\), hence \(\mathcal{A} \subset \mathcal{D}\). Any two points of \(X \setminus S\) are equivalent (a countable \(A_n\) lies in \(S\) and a co-countable one contains \(X \setminus S\)), and \(X \setminus S\) is uncountable; so for \(x_0 \in X \setminus S\) the set \(\{x_0\} \in \mathcal{A}\) is not a union of classes, contradicting \(\mathcal{A} \subset \mathcal{D}\).
Exercises 1.12.103–1.12.109
Let \((X, \mathcal{A}, \mu)\) be a measure space with a finite nonnegative measure \(\mu\) and let \(\mathcal{A}/\mu\) be the corresponding metric Boolean algebra with the metric \(d\) introduced in 1.12(iii). Prove that the mapping \(A \mapsto X\setminus A\) from \(\mathcal{A}/\mu\) to \(\mathcal{A}/\mu\) and the mappings \((A,B) \mapsto A\cup B\), \((A,B)\mapsto A\cap B\) from \((\mathcal{A}/\mu)^2\) to \(\mathcal{A}/\mu\) are continuous.
All three maps are Lipschitz for \(d(A,B)=\mu(A\,\triangle\,B)\), hence continuous. Everything rests on
\begin{equation*} \begin{aligned} (X\setminus A)\,\triangle\,(X\setminus B) &= A\,\triangle\,B, \\ (A\cup B)\,\triangle\,(A^{\prime}\cup B^{\prime}) &\subset (A\,\triangle\,A^{\prime})\cup(B\,\triangle\,B^{\prime}), \\ (A\cap B)\,\triangle\,(A^{\prime}\cap B^{\prime}) &\subset (A\,\triangle\,A^{\prime})\cup(B\,\triangle\,B^{\prime}) , \end{aligned} \end{equation*}
checked pointwise: complementation reverses membership in both coordinates at once; a point of \((A\cup B)\setminus(A^{\prime}\cup B^{\prime})\) misses both primed sets and lies in \(A\) or in \(B\), hence in \(A\setminus A^{\prime}\) or \(B\setminus B^{\prime}\); a point of \((A\cap B)\setminus(A^{\prime}\cap B^{\prime})\) lies in \(A\cap B\) and misses \(A^{\prime}\) or \(B^{\prime}\), with the same conclusion; the reversed differences follow by exchanging primed and unprimed sets.
Applying \(\mu\) with monotonicity and subadditivity,
\begin{equation*} \begin{aligned} d(X\setminus A,\ X\setminus B) &= d(A,B), \\ d(A\cup B,\ A^{\prime}\cup B^{\prime}) &\le d(A,A^{\prime})+d(B,B^{\prime}), \\ d(A\cap B,\ A^{\prime}\cap B^{\prime}) &\le d(A,A^{\prime})+d(B,B^{\prime}) . \end{aligned} \end{equation*}
The vanishing of the right-hand sides when \(d(A,A^{\prime})=d(B,B^{\prime})=0\) is exactly well-definedness on classes. The first line makes complementation an isometry of \((\mathcal{A}/\mu,d)\); the other two make union and intersection \(1\)-Lipschitz from \(\bigl((\mathcal{A}/\mu)^2,\rho\bigr)\) to \((\mathcal{A}/\mu,d)\), where \(\rho\bigl((A,B),(A^{\prime},B^{\prime})\bigr) = d(A,A^{\prime})+d(B,B^{\prime})\) induces the product topology, being equivalent to the maximum metric.
Let \(\mu\) be a separable probability measure on a \(\sigma\)-algebra \(\mathcal{A}\) and let \(\{X_t\}_{t\in T}\) be an uncountable family of sets of positive measure. Show that there exists a countable subfamily \(\{t_n\}\subset T\) such that \(\mu\big(\bigcap_{n=1}^{\infty} X_{t_n}\big)>0\).
A separable metric space is second countable, so fix a countable base \(\{U_k\}\) of \((\mathcal{A}/\mu, d)\), \(d(A,B)=\mu(A\triangle B)\) (1.12(iii)); write \(\widehat{A}\) for the class of \(A\), on which \(\mu\) is constant since \(|\mu(A)-\mu(B)|\le\mu(A\triangle B)\), and put \(S:=\{\widehat{X_t}: t\in T\}\). Each case below produces \(X^{\prime}=X_{t^{\prime}}\) and distinct indices \(t_1,t_2,\dots\) with \(\mu\bigl(X^{\prime}\setminus\bigcap_n X_{t_n}\bigr)\le\mu(X^{\prime})/2\), and then
\begin{equation*} \begin{aligned} \mu\Bigl(\bigcap_{n=1}^{\infty}X_{t_n}\Bigr) &\ \ge\ \mu\Bigl(X^{\prime}\cap\bigcap_{n}X_{t_n}\Bigr) \\ &\ =\ \mu(X^{\prime})-\mu\Bigl(X^{\prime}\setminus\bigcap_{n}X_{t_n}\Bigr) \ \ge\ \tfrac12\mu(X^{\prime})\ >\ 0 . \end{aligned} \end{equation*}
(i) Some class is attained by uncountably many indices: \(T^{\prime}=\{t: X_t\sim X^{\prime}\}\) is uncountable for \(X^{\prime}=X_{t^{\prime}}\), \(t^{\prime}\in T^{\prime}\). Choose distinct \(t_n\in T^{\prime}\); then \(\mu(X^{\prime}\setminus X_{t_n})\le\mu(X^{\prime}\triangle X_{t_n})=0\), so
\begin{equation*} X^{\prime}\setminus\bigcap_{n=1}^{\infty}X_{t_n} =\bigcup_{n=1}^{\infty}(X^{\prime}\setminus X_{t_n}) \end{equation*}
is a countable union of null sets, hence null.
(ii) Every class is attained by at most countably many indices. Then \(S\) is uncountable, since otherwise \(T=\bigcup_{s\in S}\{t:\widehat{X_t}=s\}\) would be countable. At most countably many points of \(S\) are isolated in \(S\): an isolated \(p\) admits a basic \(U_{k(p)}\) with \(U_{k(p)}\cap S=\{p\}\), and \(p\mapsto k(p)\) is injective. So pick a non-isolated \(\widehat{X^{\prime}}\in S\) with \(X^{\prime}=X_{t^{\prime}}\), whence \(\mu(X^{\prime})>0\); every ball about \(\widehat{X^{\prime}}\) then meets \(S\) in infinitely many points, since a ball meeting \(S\) in finitely many \(p_i\ne\widehat{X^{\prime}}\) could be shrunk to radius \(\min_i d(\widehat{X^{\prime}},p_i)\) and would isolate \(\widehat{X^{\prime}}\). Choose \(t_n\) with the \(\widehat{X_{t_n}}\) pairwise distinct and distinct from \(\widehat{X^{\prime}}\) (so the \(t_n\) are distinct) and
\begin{equation*} \mu\big(X^{\prime}\,\triangle\,X_{t_n}\big) < \mu(X^{\prime})\,2^{-n-1} ; \end{equation*}
since \(\mu(X^{\prime}\setminus X_{t_n})\le\mu(X^{\prime}\triangle X_{t_n})\), countable subadditivity applied to \(X^{\prime}\setminus\bigcap_n X_{t_n}=\bigcup_n(X^{\prime}\setminus X_{t_n})\) bounds it by \(\sum_n\mu(X^{\prime})2^{-n-1}=\mu(X^{\prime})/2\).
(\(\circ\)) Let \(\mathcal{A}\) be the class of all subsets on the real line that are either at most countable or have at most countable complements. If the complement of a set \(A\in\mathcal{A}\) is at most countable, then we set \(\mu(A)=1\), otherwise we set \(\mu(A)=0\). Then \(\mathcal{A}\) is a \(\sigma\)-algebra and \(\mu\) is a probability measure on \(\mathcal{A}\), the collection \(\mathcal{K}\) of all sets with at most countable complements is a compact class, approximating \(\mu\), but there is no class \(\mathcal{K}^{\prime}\subset\mathcal{A}\) approximating \(\mu\) and having the property that every (not necessarily countable) collection in \(\mathcal{K}^{\prime}\) with empty intersection has a finite subcollection with empty intersection.
Every assertion reduces to one fact: a countable union of countable subsets of \(\mathbb{R}\) is countable, whereas \(\mathbb{R}\) is not; “countable” means “at most countable” throughout.
\(\mathcal{A}\) is a \(\sigma\)-algebra, since its defining condition is symmetric in \(A\) and \(\mathbb{R}\setminus A\), and for \(A=\bigcup_n A_n\) either all \(A_n\) are countable, hence \(A\) is, or some \(A_{n_0}\) is co-countable and \(\mathbb{R}\setminus A\subset\mathbb{R}\setminus A_{n_0}\) is countable. No set is both countable and co-countable, so \(\mu\) is unambiguous, with \(\mu(\varnothing)=0\), \(\mu(\mathbb{R})=1\). Among pairwise disjoint \(A_n\) at most one is co-countable, since \(A_n\cap A_m=\varnothing\) would give \(\mathbb{R}=(\mathbb{R}\setminus A_n)\cup(\mathbb{R}\setminus A_m)\); so for \(A=\bigsqcup_n A_n\) either all \(A_n\) are countable and \(\mu(A)=0=\sum_n\mu(A_n)\), or exactly one \(A_{n_0}\) is co-countable and \(\mu(A)=1=\sum_n\mu(A_n)\).
The class \(\mathcal{K}=\{K:\mathbb{R}\setminus K\ \text{countable}\}=\{A\in\mathcal{A}:\mu(A)=1\}\) is compact in the sense of Definition 1.4.1 vacuously: for \(K_n\in\mathcal{K}\),
\begin{equation*} \mathbb{R}\setminus\bigcap_{n=1}^{\infty}K_n =\bigcup_{n=1}^{\infty}(\mathbb{R}\setminus K_n) \end{equation*}
is countable, so \(\bigcap_n K_n\) is co-countable, hence uncountable and nonempty, and no countable subfamily has empty intersection. It approximates \(\mu\) in the form of Theorem 1.4.3,
\begin{equation*} \mu(A)=\sup\{\mu(K):\ K\in\mathcal{K},\ K\subset A\}, \end{equation*}
because for \(\mu(A)=1\) the set \(A\) itself is admissible, while for \(\mu(A)=0\) the countable set \(A\) contains no member of the uncountable class \(\mathcal{K}\) and \(\sup\varnothing=0\); for the literal Definition 1.4.6, adjoin \(\varnothing\) to \(\mathcal{K}\) (still compact by the display) and take \(A_\varepsilon=K_\varepsilon=A\) when \(\mu(A)=1\) and \(A_\varepsilon=K_\varepsilon=\varnothing\) when \(\mu(A)=0\).
No approximating \(\mathcal{K}^{\prime}\subset\mathcal{A}\) has the uncountable finite intersection property. Suppose one had, and for \(x\in\mathbb{R}\) approximate \(A_x^{0}:=\mathbb{R}\setminus\{x\}\), of measure \(1\), with \(\varepsilon=1/2\): there are \(K_x\in\mathcal{K}^{\prime}\), \(A_x\in\mathcal{A}\) with
\begin{equation*} A_x\subset K_x\subset A_x^{0}=\mathbb{R}\setminus\{x\},\qquad |\mu(A_x^{0})-\mu(A_x)|<\tfrac12 , \end{equation*}
so \(\mu(A_x)=1\) as \(\mu\) takes only the values \(0,1\), and \(\mu(K_x)=1\) by monotonicity (\(K_x\in\mathcal{A}\)), i.e. \(\mathbb{R}\setminus K_x\) is countable while \(x\notin K_x\). Then \(\bigcap_{x\in\mathbb{R}}K_x=\varnothing\), so \(\bigcap_{i\le m}K_{x_i}=\varnothing\) for some \(x_1,\dots,x_m\); but its complement \(\bigcup_{i\le m}(\mathbb{R}\setminus K_{x_i})\) is countable, making that intersection uncountable – a contradiction.
(\(\circ\)) Let \(\mu\) be an atomless probability measure on a measurable space \((X,\mathcal{A})\) and let \(\mathcal{F}\subset\mathcal{A}\) be a countable family of sets of positive measure. Show that there exists a set \(A\in\mathcal{A}\) such that \(0<\mu(A\cap F)<\mu(F)\) for all \(F\in\mathcal{F}\).
Take for \(A\) any set whose class in \((\mathcal{A}/\mu,d)\), \(d(A,B)=\mu(A\triangle B)\), avoids all the sets
\begin{equation*} \widehat{\mathcal{F}_n},\qquad \mathcal{F}_n=\{A\in\mathcal{A}:\ \mu(A\cap F_n)\in\{0,\mu(F_n)\}\}, \end{equation*}
where \(\mathcal{F}=\{F_n\}_{n\in\mathbb{N}}\) is an enumeration with repetitions if \(\mathcal{F}\) is finite and \(\widehat{\ }\) denotes passage to classes. Such \(A\) exists by the Baire category theorem, \((\mathcal{A}/\mu,d)\) being a nonempty complete metric space by Theorem 1.12.6(ii) and each \(\widehat{\mathcal{F}_n}\) closed and nowhere dense as shown below; and then \(0\le\mu(A\cap F_n)\le\mu(F_n)\) forces \(0<\mu(A\cap F_n)<\mu(F_n)\) for every \(n\).
Closed: \(\Phi_n(A):=\mu(A\cap F_n)\) obeys \((A\cap F_n)\triangle(B\cap F_n)=(A\triangle B)\cap F_n\), hence \(|\Phi_n(A)-\Phi_n(B)|\le d(A,B)\), so \(\Phi_n\) descends to a continuous function on \(\mathcal{A}/\mu\) and \(\widehat{\mathcal{F}_n}=\widehat{\Phi}_n^{-1}(\{0,\mu(F_n)\})\) is the preimage of a two-point set.
Nowhere dense: it suffices that \(\widehat{\mathcal{F}_n}\) have empty interior. Fix \(A\in\mathcal{F}_n\), \(\varepsilon>0\), and put \(\alpha=\min(\varepsilon/2,\mu(F_n)/2)>0\); since the restriction of \(\mu\) to \(\{E\in\mathcal{A}:E\subset D\}\) is atomless (an atom of it would be an atom of \(\mu\)), Corollary 1.12.10 gives \(C\in\mathcal{A}\), \(C\subset D\), \(\mu( C)=\alpha\), for any \(D\) with \(\mu(D)\ge\alpha\).
(i) \(\mu(A\cap F_n)=0\): take \(D=F_n\setminus A\), of measure \(\mu(F_n)\ge\alpha\), and \(B=A\sqcup C\), so that \(d(A,B)=\mu( C)=\alpha<\varepsilon\) and \(\mu(B\cap F_n)=0+\alpha\in(0,\mu(F_n))\).
(ii) \(\mu(A\cap F_n)=\mu(F_n)\): take \(D=A\cap F_n\), of measure \(\mu(F_n)\ge\alpha\), and \(B=A\setminus C\), so that \(d(A,B)=\alpha<\varepsilon\) and \(\mu(B\cap F_n)=\mu(F_n)-\alpha\in(0,\mu(F_n))\).
Either way \(\widehat{B}\notin\widehat{\mathcal{F}_n}\) lies within \(\varepsilon\) of \(\widehat{A}\), so no point of \(\widehat{\mathcal{F}_n}\) is interior to it.
Let \(\mathbb{Q}\) be the set of all rational numbers equipped with the \(\sigma\)-algebra \(2^{\mathbb{Q}}\) of all subsets and let the measure \(\mu\) on \(2^{\mathbb{Q}}\) with values in \([0,+\infty]\) be defined as the cardinality of a set. Let \(\nu=2\mu\). Show that the distinct measures \(\mu\) and \(\nu\) coincide on all open sets in \(\mathbb{Q}\) (with the induced topology), and on all sets from the algebra that consists of finite disjoint unions of sets of the form \(\mathbb{Q}\cap(a,b]\) and \(\mathbb{Q}\cap(c,+\infty)\), where \(a,b,c\in\mathbb{Q}\) or \(c=-\infty\) (this algebra generates \(2^{\mathbb{Q}}\)).
Agreement is forced by emptiness or infiniteness: \(2t=t\) in \([0,+\infty]\) exactly for \(t\in\{0,+\infty\}\), so
\begin{equation*} \mu(A)=\nu(A)\iff \mu(A)\in\{0,+\infty\}\iff A=\varnothing\ \text{ or }\ A\ \text{is infinite}, \end{equation*}
and every set named in the statement is of one of those two kinds. The two are measures (countable additivity of the counting measure; a constant multiple of a measure is a measure) and they differ, since \(\mu(\{0\})=1\) while \(\nu(\{0\})=2\).
Open sets: a nonempty open \(U\subset\mathbb{Q}\) contains some \(\mathbb{Q}\cap(q-\delta,q+\delta)\), hence the distinct rationals \(q+\delta 2^{-k}\), \(k\in\mathbb{N}\), so \(U\) is infinite.
The algebra \(\mathcal{A}_0\) of finite disjoint unions of
\begin{equation*} I_{a,b}=\mathbb{Q}\cap(a,b],\qquad J_c=\mathbb{Q}\cap(c,+\infty) \end{equation*}
(\(b\in\mathbb{Q}\), \(a,c\in\mathbb{Q}\cup\{-\infty\}\)) is an algebra because these sets and \(\varnothing\) form a semialgebra: intersections stay of the same form, \(\mathbb{Q}\setminus J_c=I_{-\infty,c}\) and \(\mathbb{Q}\setminus I_{a,b}=I_{-\infty,a}\sqcup J_b\) – which is why \(a=-\infty\) must be admitted, as the statement intends. Every nonempty generator is infinite: \(J_c\) equals \(\mathbb{Q}\) for \(c=-\infty\) and contains \(c+k\), \(k\in\mathbb{N}\), otherwise, while \(I_{a,b}\ne\varnothing\) means \(a<b\) and then contains \(b-(b-a)2^{-k}\) for \(a\in\mathbb{Q}\), respectively \(b-k\) for \(a=-\infty\). Hence \(A=\bigsqcup_{i\le m}S_i\in\mathcal{A}_0\) is either empty or contains an infinite \(S_{i_0}\), so \(\mu(A)=\nu(A)\).
Generation: \(\{q\}=\bigcap_{k\ge1}\bigl(\mathbb{Q}\cap(q-1/k,\,q]\bigr)\in\sigma(\mathcal{A}_0)\) and \(\mathbb{Q}\) is countable, so \(\sigma(\mathcal{A}_0)=2^{\mathbb{Q}}\).
Prove that there exists no countably additive measure defined on all subsets of the space \(X=\{0,1\}^{\infty}\) that assumes only two values \(0\) and \(1\) and vanishes on all singletons.
No such \(\mu\) exists: the coordinate half-spaces of full measure intersect in a single point. Suppose \(\mu\) on \(2^{X}\), \(X=\{0,1\}^{\infty}\), were countably additive with values \(0\) and \(1\), both attained, and \(\mu(\{x\})=0\) for all \(x\). Additivity with \(\mu\ge0\) gives monotonicity, hence \(\mu(X)=1\), and, after disjointifying \(\bigcup_n A_n=\bigsqcup_n(A_n\setminus\bigcup_{k<n}A_k)\), countable subadditivity. For each \(n\) the sets \(X_n=\{x: x_n=0\}\) and \(X\setminus X_n\) have measures summing to \(\mu(X)=1\), so exactly one of them has measure \(1\); calling it \(Y_n\) and letting \(y_n\in\{0,1\}\) be the corresponding digit,
\begin{equation*} Y_n=\{x\in X:\ x_n=y_n\},\qquad \mu(X\setminus Y_n)=0 . \end{equation*}
Then \(x\in\bigcap_n Y_n\) iff \(x_n=y_n\) for every \(n\), so \(\bigcap_n Y_n=\{y\}\) with \(y=(y_n)_{n\ge1}\). But \(X\setminus\bigcap_n Y_n=\bigcup_n(X\setminus Y_n)\) is null by countable subadditivity, so
\begin{equation*} \mu\Big(\bigcap_{n=1}^{\infty}Y_n\Big) =\mu(X)-\mu\Big(X\setminus\bigcap_{n=1}^{\infty}Y_n\Big)=1-0=1 , \end{equation*}
contradicting \(\mu(\{y\})=0\).
Prove that for every Borel set \(E\subset\mathbb{R}^n\), there exists a Borel set \(\widehat{E}\) that differs from \(E\) in a measure zero set and has the following property: for every point \(x\) at the boundary \(\partial\widehat{E}\) of the set \(\widehat{E}\) and every \(r>0\), one has
\begin{equation*} 0<\lambda_n\big(\widehat{E}\cap B(x,r)\big)<\omega_n r^{n}, \end{equation*}
where \(B(x,r)\) is the ball centered at \(x\) with the radius \(r\) and \(\omega_n\) is the measure of the unit ball.
Take \(\widehat{E}=(E\cup E_1)\setminus E_0\), where \(\lambda=\lambda_n\), \(B(x,r)\) is the open ball, of measure \(\omega_n r^{n}\), and
\begin{equation*} \begin{aligned} E_0 &= \{x:\ \lambda\big(E\cap B(x,r)\big)=0\ \text{ for some } r>0\}, \\ E_1 &= \{x:\ \lambda\big(B(x,r)\setminus E\big)=0\ \text{ for some } r>0\} \end{aligned} \end{equation*}
(the condition on \(E_1\) being \(\lambda(E\cap B(x,r))=\omega_n r^{n}\), since \(\lambda(B(x,r))<\infty\)).
Both are open: if \(r\) witnesses \(x\in E_0\) and \(|y-x|<r/2\), then \(B(y,r/2)\subset B(x,r)\) forces \(\lambda(E\cap B(y,r/2))=0\), so \(B(x,r/2)\subset E_0\); the same computation with \(\lambda(B(\cdot)\setminus E)\) in place of \(\lambda(E\cap\,\cdot\,)\) gives \(E_1\) open. Hence \(\widehat{E}\) is Borel.
Next, \(\widehat{E}\triangle E\) is null. Covering \(E_0\) by the witnessing balls \(B(x,r_x)\) and extracting a countable subcover, \(E_0\) being Lindelof in the second countable \(\mathbb{R}^n\), puts \(E\cap E_0\) inside \(\bigcup_k\big(E\cap B(x_k,r_{x_k})\big)\), a countable union of null sets; symmetrically \(E_1\setminus E\subset\bigcup_k\big(B(z_k,s_{z_k})\setminus E\big)\) is null. As \(E\subset E\cup E_1\),
\begin{equation*} E\setminus\widehat{E}=E\cap E_0,\qquad \widehat{E}\setminus E\subset(E\cup E_1)\setminus E=E_1\setminus E, \end{equation*}
so \(\lambda(\widehat{E}\triangle E)=0\) and therefore, for every measurable \(M\),
\begin{equation*} \lambda(\widehat{E}\cap M)=\lambda(E\cap M),\qquad \lambda(M\setminus\widehat{E})=\lambda(M\setminus E). \tag{\(\ast\)} \end{equation*}
Also \(\partial\widehat{E}\cap(E_0\cup E_1)=\varnothing\): witnesses \(r,s\) for \(x\in E_0\cap E_1\) would give \(\lambda(B(x,t))=0\) for \(t=\min(r,s)>0\), so \(E_0\cap E_1=\varnothing\); \(E_0\) is open and disjoint from \(\widehat{E}\), hence disjoint from \(\overline{\widehat{E}}\); and \(E_1\subset(E\cup E_1)\setminus E_0=\widehat{E}\) is open, hence lies in \(\mathrm{int}\,\widehat{E}\).
Thus for \(x\in\partial\widehat{E}\) and \(r>0\) no radius witnesses either membership, i.e. \(\lambda(E\cap B(x,r))>0\) and \(\lambda(B(x,r)\setminus E)>0\), so by \((\ast)\)
\begin{equation*} \begin{aligned} 0<\lambda\big(\widehat{E}\cap B(x,r)\big) &=\omega_n r^{n}-\lambda\big(B(x,r)\setminus\widehat{E}\big)<\omega_n r^{n} . \end{aligned} \end{equation*}
Exercises 1.12.110–1.12.116
Prove that every uncountable set \(G \subset \mathbb{R}\) that is the intersection of a sequence of open sets contains a nowhere dense closed set \(Z\) of Lebesgue measure zero that can be continuously mapped onto \([0,1]\).
Take for \(Z\) a Cantor scheme built inside the condensation points of \(G = \bigcap_{n \ge 1}U_n\), \(U_n\) open.
(i) Condensation points. Let \(P\) be the set of \(x \in G\) every neighbourhood of which meets \(G\) in an uncountable set. Each \(x \in D := G \setminus P\) lies in a rational interval \((p,q)\) with \(G \cap (p,q)\) at most countable, so \(D\) sits inside the union of those countably many countable sets and is itself at most countable; hence \(P\) is uncountable. Every point of \(P\) is a condensation point of \(P\), since for a neighbourhood \(V\) of \(x \in P\) the set \((G \cap V)\setminus P \subset D\) is countable while \(G \cap V\) is not.
(ii) The scheme. Construct nondegenerate closed intervals \(I_s\) and points \(x_s \in P \cap \operatorname{int} I_s\), indexed by finite binary strings \(s\) of length \(n = |s|\), with \(|I_s| \le 4^{-n}\), with \(I_{s^{\frown}0}\), \(I_{s^{\frown}1}\) disjoint and inside \(\operatorname{int} I_s\), and with \(I_s \subset U_n\) for \(n \ge 1\). Start from any \(x_{\emptyset} \in P\) interior to an \(I_{\emptyset}\) of length at most \(1\). Given \(I_s\), \(x_s\): by (i) the set \(P \cap \operatorname{int} I_s\) is uncountable, so it contains distinct \(y_0, y_1\), and since \(P \subset G \subset U_{n+1}\) with \(U_{n+1}\) open there are disjoint closed intervals
\begin{equation*} J_i \subset \operatorname{int} I_s \cap U_{n+1}, \qquad |J_i| \le 4^{-(n+1)}, \qquad y_i \in \operatorname{int} J_i ; \end{equation*}
put \(I_{s^{\frown} i} := J_i\), \(x_{s^{\frown} i} := y_i\).
(iii) The set. \(F_n := \bigcup_{|s|=n} I_s\) is a union of \(2^n\) pairwise disjoint closed intervals (induction on the level), \(F_{n+1} \subset F_n\), and \(Z := \bigcap_{n \ge 0}F_n\) is a decreasing intersection of nonempty compacta, hence nonempty and compact. Each \(z \in Z\) lies in some \(I_s \subset U_n\) with \(|s| = n\), so \(Z \subset \bigcap_n U_n = G\); and \(\lambda(Z) \le \lambda(F_n) \le 2^{n}4^{-n} = 2^{-n} \to 0\), so \(Z\) is null, contains no interval, and, being closed with empty interior, is nowhere dense.
(iv) The surjection. For \(\eta \in \{0,1\}^{\mathbb{N}}\) the intervals \(I_{\eta|n}\) are nested with lengths tending to \(0\), so \(\bigcap_n I_{\eta|n} = \{z(\eta)\} \subset Z\); conversely each \(z \in Z\) lies in exactly one level-\(n\) interval, and these strings cohere into an \(\eta(z)\) inverse to \(\eta \mapsto z(\eta)\). Then
\begin{equation*} f(z) = \sum_{n=1}^{\infty} \eta_n(z) \, 2^{-n}, \qquad z \in Z , \end{equation*}
maps \(Z\) onto \([0,1]\), since \(f(z(\eta)) = t\) for any binary expansion \(t = \sum_n \eta_n 2^{-n}\), and is uniformly continuous: the level-\(n\) intervals being disjoint compacta,
\begin{equation*} \delta_n := \min \{ \operatorname{dist}(I_s, I_{s^{\prime}}) : s \neq s^{\prime}, \ |s| = |s^{\prime}| = n \} > 0 , \end{equation*}
and \(|z - z^{\prime}| < \delta_n\) puts \(z, z^{\prime}\) in one level-\(n\) interval, so \(\eta(z)|n = \eta(z^{\prime})|n\) and \(|f(z)-f(z^{\prime})| \le \sum_{k>n}2^{-k} = 2^{-n}\).
Prove that every uncountable set \(G \subset \mathbb{R}\) that is the intersection of a sequence of open sets has cardinality of the continuum.
\(|G| = \mathfrak{c} := |\mathbb{R}|\). Indeed \(|G| \le \mathfrak{c}\) as \(G \subset \mathbb{R}\), while Exercise 1.12.110 supplies a closed \(Z \subset G\) and a continuous surjection \(f \colon Z \to [0,1]\); choosing \(g(t) \in f^{-1}(t)\) for each \(t\) embeds \([0,1]\) into \(Z\), the fibres being pairwise disjoint. Hence
\begin{equation*} \mathfrak{c} \le |Z| \le |G| \le \mathfrak{c} , \end{equation*}
and the Cantor–Schroeder–Bernstein theorem gives \(|G| = \mathfrak{c}\).
(i) Prove that the class of all Souslin subsets of the real line is obtained by applying the A-operation to the collection of all open sets. (ii) Show that in (i) it suffices to take the collection of all intervals with rational endpoints.
Both parts are the single implication “\(\mathcal{E}_1 \subset S(\mathcal{E}_2)\) forces \(S(\mathcal{E}_1) \subset S(\mathcal{E}_2)\)”, applied twice in each direction. Here \(S(\mathcal{E})\) is the class of sets produced by the A-operation (1.10.1) from \(\mathcal{E}\)-schemes, together with \(\emptyset\); \(\mathcal{F}\), \(\mathcal{U}\), \(\mathcal{I}\) are the closed sets, the open sets, and the intervals with rational endpoints; and by Definition 1.10.7 the Souslin subsets of \(\mathbb{R}\) are the members of \(S(\mathcal{F})\). The implication holds because \(S\) is monotone (an \(\mathcal{E}_1\)-scheme is an \(\mathcal{E}_2\)-scheme, with the same value) and idempotent, \(S(S(\mathcal{E})) = S(\mathcal{E})\) being Theorem 1.10.4(i); and by Example 1.10.2 countable unions and countable intersections of members of \(\mathcal{E}\) lie in \(S(\mathcal{E})\).
(i) \(\mathcal{F} \subset S(\mathcal{U})\): a nonempty closed \(F\) is the countable intersection
\begin{equation*} F = \bigcap_{n=1}^{\infty} U_n, \qquad U_n := \{ x \in \mathbb{R} : \operatorname{dist}(x, F) < 1/n \}, \end{equation*}
of open sets, \(\operatorname{dist}(\cdot,F)\) being continuous with zero set \(\overline{F} = F\), while \(\emptyset\) is open. Conversely \(\mathcal{U} \subset S(\mathcal{F})\), an open \(U\) being the union of the countably many closed rational intervals \([p,q] \subset U\). Hence \(S(\mathcal{F}) = S(\mathcal{U})\).
(ii) \(\mathcal{I} \subset S(\mathcal{U})\): every interval with rational endpoints is a countable intersection of open ones, e.g. \([p,q] = \bigcap_n (p-1/n, q+1/n)\) and \([p,q) = \bigcap_n (p-1/n, q)\). Conversely \(\mathcal{U} \subset S(\mathcal{I})\): each \(x\) in an open \(U\) has \((x-\varepsilon,x+\varepsilon) \subset U\) and rationals \(p \in (x-\varepsilon,x)\), \(q \in (x,x+\varepsilon)\) give \(x \in (p,q) \subset U\), so \(U\) is the union of the countably many rational intervals it contains. With (i), \(S(\mathcal{I}) = S(\mathcal{U}) = S(\mathcal{F})\).
Prove that the classes of all Souslin and all Borel sets on the real line (or in the space \(\mathbb{R}^n\)) have cardinality of the continuum.
Both classes have cardinality \(\mathfrak{c} = |\mathbb{R}|\), by \(\mathfrak{c} \le |\mathcal{B}| \le |\mathcal{S}| \le \mathfrak{c}\) and Cantor–Schroeder–Bernstein; here \(\mathcal{B} = \mathcal{B}(\mathbb{R}^n)\) and \(\mathcal{S} = S(\mathcal{F})\) is the Souslin class of Definition 1.10.7, \(\mathcal{F}\) the closed sets.
\(\mathfrak{c} \le |\mathcal{B}|\): the singletons are closed and pairwise distinct.
\(\mathcal{B} \subset \mathcal{S}\): the complement of a closed set is a countable union of closed sets, e.g. of the sets \(\{x \in U : \operatorname{dist}(x, \mathbb{R}^n\setminus U) \ge 1/k, \ |x| \le k\}\), and \(\emptyset \in \mathcal{F}\), which is the hypothesis of Theorem 1.10.4(ii); hence \(\mathcal{B} = \sigma(\mathcal{F}) \subset S(\mathcal{F})\).
\(|\mathcal{S}| \le \mathfrak{c}\): the argument of Exercise 1.12.112 runs verbatim in \(\mathbb{R}^n\) with the countable class \(\mathcal{I}\) of boxes \((p_1,q_1)\times\cdots\times(p_n,q_n)\) with rational endpoints – every open set is a countable union of such boxes, every closed set a countable intersection of open sets – and gives \(\mathcal{S} = S(\mathcal{I})\). A scheme with values in \(\mathcal{I}\) is by Definition 1.10.1 a map \(T \colon \Sigma \to \mathcal{I}\) on the countable set \(\Sigma = \bigcup_{k \ge 1}\mathbb{N}^{k}\), so there are at most
\begin{equation*} |\mathcal{I}|^{|\Sigma|} = \aleph_0^{\aleph_0} \le \bigl(2^{\aleph_0}\bigr)^{\aleph_0} = 2^{\aleph_0} = \mathfrak{c} \end{equation*}
of them, and \(S(\mathcal{I})\) is their image under the A-operation together with \(\emptyset\).
Method (2) for \(|\mathcal{B}| \le \mathfrak{c}\), avoiding Souslin sets: put \(\mathcal{B}_0 := \{\text{open sets}\}\) and, for \(0 < \alpha < \omega_1\),
\begin{equation*} \begin{aligned} \mathcal{C}_{\alpha} &:= \textstyle\bigcup_{\beta<\alpha}\mathcal{B}_{\beta}, \\ \mathcal{B}_{\alpha} &:= \bigl\{\mathbb{R}^n\setminus A : A \in \mathcal{C}_{\alpha}\bigr\} \cup \bigl\{\textstyle\bigcup_{k\ge1}A_k : A_k \in \mathcal{C}_{\alpha}\bigr\} . \end{aligned} \end{equation*}
Then \(|\mathcal{B}_0| \le \mathfrak{c}\), since \(U \mapsto \{I \in \mathcal{I} : I \subset U\}\) injects it into the power set of the countable \(\mathcal{I}\), and inductively \(|\mathcal{C}_{\alpha}| \le \aleph_0\cdot\mathfrak{c} = \mathfrak{c}\) gives \(|\mathcal{B}_{\alpha}| \le \mathfrak{c} + \mathfrak{c}^{\aleph_0} = \mathfrak{c}\). Finally \(\mathcal{B} = \bigcup_{\alpha<\omega_1}\mathcal{B}_{\alpha}\), the right side being a \(\sigma\)-algebra because a countable set of countable ordinals is bounded in \(\omega_1\); so \(|\mathcal{B}| \le \aleph_1\cdot\mathfrak{c} = \mathfrak{c}\).
Let \((X, \mathcal{A}, \mu)\) be a space with a finite nonnegative measure \(\mu\) such that there exists a set \(E\) that is not \(\mu\)-measurable. Prove that there exists \(\varepsilon > 0\) with the following property: if \(A\) and \(B\) are measurable, \(E \subset A\), \(X \setminus E \subset B\), then \(\mu(A \cap B) \ge \varepsilon\).
Take the measurability defect \(\varepsilon := \mu^{*}(E) + \mu^{*}(X\setminus E) - \mu(X)\); measurable here means belonging to the completion \(\mathcal{A}_{\mu}\), on which \(\mu = \mu^{*}\) by Theorem 1.5.6, and \(\mu(X)<\infty\) makes the subtraction legitimate.
\(\varepsilon > 0\): subadditivity of \(\mu^{*}\) (Lemma 1.5.4) applied to \(X = E \cup (X\setminus E)\), together with \(\mu^{*}(X) = \mu(X)\) (Theorem 1.5.6(i)), gives \(\varepsilon \ge 0\), while \(\varepsilon = 0\) would place \(E\) in \(\mathcal{A}_{\mu}\) by Proposition 1.5.11, against the hypothesis.
The estimate: for \(A, B \in \mathcal{A}_{\mu}\) with \(E \subset A\) and \(X\setminus E \subset B\) one has \(A \cup B \supset E \cup (X\setminus E) = X\), while monotonicity of \(\mu^{*}\) gives \(\mu(A) \ge \mu^{*}(E)\) and \(\mu(B) \ge \mu^{*}(X\setminus E)\), so by additivity and finiteness of \(\mu\),
\begin{equation*} \begin{aligned} \mu(A \cap B) &= \mu(A) + \mu(B) - \mu(A\cup B) \\ &\ge \mu^{*}(E) + \mu^{*}(X\setminus E) - \mu(X) = \varepsilon . \end{aligned} \end{equation*}
Method (2): were there no such \(\varepsilon\), choose measurable \(A_n \supset E\), \(B_n \supset X\setminus E\) with \(\mu(A_n \cap B_n) < 1/n\) and put \(A = \bigcap_n A_n\), \(B = \bigcap_n B_n\); then \(\mu(A\cap B) = 0\) and \(A \cup B = X\), so
\begin{equation*} \begin{aligned} \mu^{*}(E) + \mu^{*}(X\setminus E) &\le \mu(A) + \mu(B) \\ &= \mu(A\cup B) + \mu(A\cap B) = \mu(X), \end{aligned} \end{equation*}
which with the reverse subadditivity inequality puts \(E \in \mathcal{A}_{\mu}\) by Proposition 1.5.11 – a contradiction.
Construct an example of a separable probability measure \(\mu\) on a \(\sigma\)-algebra \(\mathcal{A}\) such that, for every countably generated \(\sigma\)-algebra \(\mathcal{E} \subset \mathcal{A}\), the completion of \(\mathcal{E}\) with respect to \(\mu\) is strictly smaller than \(\mathcal{A}\).
Take the countable–cocountable space on \(X = [0,1]\): with “countable” meaning “at most countable”,
\begin{equation*} \mathcal{A} := \{ A \subset X : A \ \text{or} \ X \setminus A \ \text{is countable} \}, \end{equation*}
and \(\mu(A) := 0\) or \(1\) according to which alternative holds.
It is a probability space: \(\mathcal{A}\) is closed under complements, and for \(A = \bigcup_n A_n\) either all \(A_n\) are countable, so \(A\) is, or some \(A_{n_0}\) is co-countable and \(X\setminus A \subset X \setminus A_{n_0}\); \(\mu\) is unambiguous since \(X\) is uncountable, and two co-countable sets meet, so among pairwise disjoint \(A_n\) at most one is co-countable and both sides of countable additivity are \(0\), respectively \(1\). It is separable: \(\mu(A \triangle B) = 0\) holds exactly when \(A, B\) are both countable or both co-countable, so \(\mathcal{A}/\mu\) has two points.
Let now \(\mathcal{E} = \sigma(E_1,E_2,\dots) \subset \mathcal{A}\) be countably generated, put
\begin{equation*} \begin{aligned} N &:= \bigcup \{ E_n : E_n \ \text{countable} \} \ \cup \ \bigcup \{ X \setminus E_n : X \setminus E_n \ \text{countable} \}, \\ \mathcal{G} &:= \{ A \subset X : A \cap (X \setminus N) \in \{\emptyset,\ X \setminus N\} \} . \end{aligned} \end{equation*}
Then \(N\) is countable; \(\mathcal{G}\) is a \(\sigma\)-algebra (a set missing \(X\setminus N\) has complement containing it, and a countable union either misses \(X\setminus N\) throughout or contains it); \(\mathcal{G} \subset \mathcal{A}\), its members lying in \(N\) or containing \(X\setminus N\); and \(E_n \in \mathcal{G}\) by the definition of \(N\), so \(\mathcal{E} \subset \mathcal{G}\).
The completion meant is the Lebesgue completion \(\mathcal{E}_{\mu}\) of \(\mu|_{\mathcal{E}}\) (Definition 1.5.1, Theorem 1.5.6), so \(\mu^{*}\) is generated by \(\mathcal{E}\) alone, not the larger \(\sigma\)-algebra generated by \(\mathcal{E}\) and the \(\mu\)-null sets of \(\mathcal{A}\) of Exercise 1.12.101. Every \(F \in \mathcal{E} \subset \mathcal{G}\) has \(\mu(F) = 0\) iff \(F \subset N\), so a cover \(S \subset \bigcup_k F_k\) with \(\sum_k \mu(F_k) < 1\) forces every \(F_k \subset N\) and hence \(S \subset N\):
\begin{equation*} \mu^{*}(S) \in \{0, 1\} \quad \text{and} \quad \mu^{*}(S) = 0 \implies S \subset N \qquad (S \subset X). \end{equation*}
Pick \(x \in X \setminus N\), possible as \(N\) is countable and \(X\) is not. Then \(\{x\} \in \mathcal{A} \setminus \mathcal{E}_{\mu}\): for \(F \in \mathcal{E}\), either \(x \notin F\), and then \(\{x\} \subset \{x\} \triangle F\) with \(\{x\} \not\subset N\), or \(x \in F\), and then \(F \supset X \setminus N\) as \(F \in \mathcal{G}\), so \(\{x\} \triangle F \supset (X\setminus N)\setminus\{x\}\), uncountable and hence not inside \(N\); either way \(\mu^{*}(\{x\} \triangle F) = 1\). Finally \(\mathcal{E}_{\mu} \subset \mathcal{G}\): for \(S \in \mathcal{E}_{\mu}\) take \(F \in \mathcal{E}\) with \(\mu^{*}(S \triangle F) < 1\), so \(\mu^{*}(S\triangle F) = 0\), \(S \triangle F \subset N\), and \(S \cap (X\setminus N) = F \cap (X \setminus N) \in \{\emptyset, X\setminus N\}\). Hence
\begin{equation*} \mathcal{E}_{\mu} \subset \mathcal{G} \subsetneq \mathcal{A} . \end{equation*}
(Zink) Let \((X, \mathcal{S}, \mu)\) be a measure space with a complete atomless separable probability measure \(\mu\) and let \(\mu^{*}(E) > 0\). Then there exist nonmeasurable sets \(E_1\) and \(E_2\) such that \(E_1 \cap E_2 = \emptyset\), \(E_1 \cup E_2 = E\) and one has \(\mu^{*}(E_1) = \mu^{*}(E_2) = \mu^{*}(E)\).
Zink’s theorem is proved below under two extra hypotheses, (A) and (H); the reduction to them is unconditional. Fix a measurable hull \(H\) of \(E\), i.e. \(H \in \mathcal{S}\) with \(E \subset H\) and \(\mu(H) = \mu^{*}(E)\), such hulls existing by the argument in the proof of Corollary 1.5.8, and write \(h := \mu(H) = \mu^{*}(E) > 0\).
Hull identity: \(\mu^{*}(E \cap B) = \mu(H \cap B)\) for \(B \in \mathcal{S}\). Here \(\le\) is clear, \(H \cap B\) being measurable and containing \(E \cap B\); and if \(<\) held, a hull \(C\) of \(E \cap B\), replaced by \(C \cap H \cap B\), would make \(H^{\prime} := (H \setminus B) \cup C\) measurable and containing \(E\) (a point of \(E\) lies outside \(B\), hence in \(H \setminus B\), or in \(E \cap B \subset C\)) with
\begin{equation*} \begin{aligned} \mu(H^{\prime}) &\le \mu(H \setminus B) + \mu( C) \\ &< \mu(H \setminus B) + \mu(H \cap B) = \mu(H) = \mu^{*}(E), \end{aligned} \end{equation*}
against the definition of \(\mu^{*}(E)\) as an infimum over measurable supersets.
Nonmeasurability comes free: if \(E = E_1 \sqcup E_2\) with \(\mu^{*}(E_1) = \mu^{*}(E_2) = h\) and \(E_1 \in \mathcal{S}\), then \(E_2 \subset H \setminus E_1\) gives \(h \le h - \mu(E_1)\), so \(\mu(E_1) = 0\) and \(\mu^{*}(E_1) = 0 < h\); symmetrically for \(E_2\). Only the splitting has to be produced.
Singletons are null: a hull \(H_x\) of \(\{x\}\) with \(\mu(H_x) = c > 0\) splits, \(\mu\) being atomless, into two measurable sets of positive measure, one of which contains \(x\) and is a superset of \(\{x\}\) of measure \(< c\). Hence countable sets are null and every set of positive outer measure is uncountable.
Reduction to \(E = X\) (unconditional). On the trace \(\mathcal{S}_E := \{B \cap E : B \in \mathcal{S}\}\) put \(\nu(B \cap E) := \mu(B \cap H)\). This is well defined and countably additive by the hull identity: \(B_1 \cap E = B_2 \cap E\) gives \(\mu\bigl(H \cap (B_1 \triangle B_2)\bigr) = \mu^{*}\bigl(E \cap (B_1 \triangle B_2)\bigr) = 0\), and disjoint traces give \(\mu(H \cap B_n \cap B_m) = 0\) for \(n \ne m\); moreover \(\nu(E) = h\) and \(\nu^{*} = \mu^{*}\) on subsets of \(E\), since \(S \subset B \in \mathcal{S}\) yields \(\nu(B \cap E) = \mu(B \cap H) \le \mu(B)\) while \(S \subset B \cap E\) yields \(\mu^{*}(S) \le \mu^{*}(E \cap B) = \nu(B \cap E)\). The space \((E, \mathcal{S}_E, h^{-1}\nu)\) is again complete (a subset of a \(\nu\)-null trace has \(\mu^{*} = 0\), hence lies in the complete \(\mathcal{S}\)), atomless (split \(B \cap H = P \sqcup Q\) into \(\mu\)-positive halves; then \(\nu(P \cap E) = \mu(P) \in (0, \nu(B \cap E))\)), and separable (\(\nu\bigl((B \cap E) \triangle (A_n \cap E)\bigr) \le \mu(B \triangle A_n)\) for a \(\mu\)-dense \(\{A_n\}\)). So it suffices to split a complete atomless separable probability space \((Y, \mathcal{T}, \nu)\) as \(Y = D \sqcup D^{\prime}\) with \(\nu^{*}(D) = \nu^{*}(D^{\prime}) = 1\).
Test sets (unconditional). Let \(\{A_n\}\) be \(\nu\)-dense and \(\mathcal{T}_0 := \sigma(\{A_n\})\); every \(M \in \mathcal{T}\) agrees up to a null set with some \(F \in \mathcal{T}_0\), namely \(F = \bigcap_j \bigcup_{k \ge j} A_{n_k}\) where \(\nu(M \triangle A_{n_k}) < 2^{-k}\). Hence
\begin{equation*} \begin{aligned} \nu^{*}(D \cap F) > 0 \quad\text{and}\quad \nu^{*}\bigl((Y \setminus D) \cap F\bigr) > 0 \\ \text{for every } F \in \mathcal{T}_0 \text{ with } \nu(F) > 0 \end{aligned} \end{equation*}
already forces \(\nu^{*}(D) = \nu^{*}(Y \setminus D) = 1\): otherwise \(Y \setminus D\) contains a measurable \(M\) with \(\nu(M) > 0\), whose partner \(F\) has \(\nu(F) = \nu(M) > 0\) while \(D \cap F \subset F \setminus M\) gives \(\nu^{*}(D \cap F) = 0\).
The construction assumes (A) \(\mathcal{T}\) is the \(\nu\)-completion of the countably generated \(\mathcal{T}_0 = \sigma(\{A_n\})\), and (H) every set of positive \(\nu\)-outer measure has cardinality \(\mathfrak{c}\); both pass to the trace space above, (H) because \(\nu^{*} = \mu^{*}\) and (A) because \(B = F \triangle N\) with \(N\) inside a null \(G \in \mathcal{S}_0\) gives \(B \cap E = (F \cap E) \triangle (N \cap E)\) with \(\nu(G \cap E) = 0\). Under (A) a \(\nu\)-null \(M = F^{\prime} \triangle N\) lies in the null set \(F^{\prime} \cup G^{\prime} \in \mathcal{T}_0\), so \(\nu^{*}(S \cap F) = 0\) iff \(S \cap F \subset G\) for some null \(G \in \mathcal{T}_0\); consequently the displayed criterion holds as soon as \(D\) and \(Y \setminus D\) meet every member of \(\mathcal{B} := \{F \in \mathcal{T}_0 : \nu(F) > 0\}\), \(|\mathcal{B}| \le \mathfrak{c}\), since \(D \cap F \subset G\) makes \(D\) miss \(F \setminus G \in \mathcal{B}\). Enumerating \(\mathcal{B} = \{B_{\xi} : \xi < \kappa\}\), \(\kappa \le \mathfrak{c}\), choose by transfinite recursion
\begin{equation*} \begin{aligned} x_{\xi}, y_{\xi} \in B_{\xi}, \qquad x_{\xi} \ne y_{\xi}, \\ x_{\xi}, y_{\xi} \notin \{ x_{\eta}, y_{\eta} : \eta < \xi \} , \end{aligned} \end{equation*}
possible because at stage \(\xi\) fewer than \(\mathfrak{c}\) points are used while \(|B_{\xi}| = \mathfrak{c}\) by (H). Then \(Q := \{y_{\xi} : \xi < \kappa\}\) and \(D := Y \setminus Q\) both meet every \(B \in \mathcal{B}\), so \(\nu^{*}(D) = \nu^{*}(Q) = 1\); transporting back through the trace space gives \(E = E_1 \sqcup E_2\) with \(\mu^{*}(E_1) = \mu^{*}(E_2) = \mu^{*}(E)\), both nonmeasurable.
Both hypotheses are genuine restrictions: separable probability measures whose \(\sigma\)-algebra is not the completion of a countably generated one exist by Exercise 1.12.115, while (H) is set-theoretic, holding under the continuum hypothesis and, more generally, when \(\mathrm{non}(\mathcal{N}) = \mathfrak{c}\) for the null ideal \(\mathcal{N}\) of \(\nu\). For the completion of an atomless Borel probability measure on a Polish space (A) is automatic, the Borel \(\sigma\)-algebra being countably generated. (A fully rigorous proof of the remaining step is beyond the scope of this page; see the reference given in the book.)
Exercises 1.12.117–1.12.123
(\(\circ\)) Let \(\mathfrak{m}\) be a Caratheodory outer measure on a space \(X\). Prove that a set \(A\) is Caratheodory measurable precisely when for all \(B \subset A\) and \(C \subset X\setminus A\) one has \(\mathfrak{m}(B \cup C) = \mathfrak{m}(B) + \mathfrak{m}( C)\).
Each direction is Definition 1.11.2 – \(\mathfrak{m}(E\cap A) + \mathfrak{m}(E\setminus A) = \mathfrak{m}(E)\) for all \(E \subset X\) – applied to the right test set.
Necessity: for \(A \in \mathfrak{M}_{\mathfrak{m}}\), \(B \subset A\) and \(C \subset X \setminus A\), the test set \(E := B \cup C\) has \(E \cap A = B\) and \(E \setminus A = C\), so
\begin{equation*} \mathfrak{m}(B \cup C) = \mathfrak{m}(E \cap A) + \mathfrak{m}(E\setminus A) = \mathfrak{m}(B) + \mathfrak{m}( C). \end{equation*}
Sufficiency: for arbitrary \(E \subset X\) the sets \(B := E \cap A \subset A\) and \(C := E \setminus A \subset X \setminus A\) have \(B \cup C = E\), so the hypothesis reads \(\mathfrak{m}(E) = \mathfrak{m}(E \cap A) + \mathfrak{m}(E \setminus A)\); as \(E\) was arbitrary, \(A \in \mathfrak{M}_{\mathfrak{m}}\).
(\(\circ\)) Suppose that \(\mathfrak{m}_1\) and \(\mathfrak{m}_2\) are outer measures on a space \(X\). Show that \(\max(\mathfrak{m}_1, \mathfrak{m}_2)\) is an outer measure too.
Yes: \(\mathfrak{m} := \max(\mathfrak{m}_1, \mathfrak{m}_2)\) meets the three conditions of Definition 1.11.1. Its values lie in \([0,+\infty]\) with \(\mathfrak{m}(\emptyset) = \max(0,0) = 0\); it is monotone, since \(A \subset B\) gives \(\mathfrak{m}_i(A) \le \mathfrak{m}_i(B) \le \mathfrak{m}(B)\) for \(i = 1,2\); and for \(A = \bigcup_{n=1}^{\infty}A_n\) each \(i\) obeys
\begin{equation*} \mathfrak{m}_i(A) \le \sum_{n=1}^{\infty} \mathfrak{m}_i(A_n) \le \sum_{n=1}^{\infty} \mathfrak{m}(A_n), \end{equation*}
a bound independent of \(i\), whence \(\mathfrak{m}(A) \le \sum_{n} \mathfrak{m}(A_n)\).
(\(\circ\)) (Young) Let \((X, \mathcal{A}, \mu)\) be a measure space with a finite nonnegative measure \(\mu\). Prove that a set \(A \subset X\) belongs to \(\mathcal{A}_{\mu}\) precisely when for each set \(B\) disjoint with \(A\) one has the equality \(\mu^*(A \cup B) = \mu^*(A) + \mu^*(B)\).
Both directions run through Proposition 1.5.11, which for finite nonnegative \(\mu\) makes \(A \in \mathcal{A}_{\mu}\) equivalent to \(\mu^*(A) + \mu^*(X \setminus A) = \mu(X)\) and to Caratheodory \(\mu^*\)-measurability.
Necessity: \(A \in \mathcal{A}_{\mu}\) is \(\mu^*\)-measurable, so for \(B\) disjoint with \(A\) the criterion of Exercise 1.12.117 with \(\mathfrak{m} = \mu^*\) and test set \(E := A \cup B\), where \(E \cap A = A\) and \(E \setminus A = B\), gives
\begin{equation*} \mu^*(A \cup B) = \mu^*(E \cap A) + \mu^*(E \setminus A) = \mu^*(A) + \mu^*(B). \end{equation*}
Sufficiency: the single test set \(B := X \setminus A\) has \(A \cup B = X\), so, using \(\mu^*(X) = \mu(X)\) from Theorem 1.5.6(i),
\begin{equation*} \mu(X) = \mu^*(X) = \mu^*(A) + \mu^*(X\setminus A), \end{equation*}
which is the second condition of Proposition 1.5.11; hence \(A \in \mathcal{A}_{\mu}\).
(\(\circ\)) Let \(\mathfrak{m}\) be a Caratheodory outer measure on a space \(X\). Prove that for any \(E \subset X\) the function \(\mathfrak{m}_E(B) = \mathfrak{m}(B \cap E)\) is a Caratheodory outer measure and all \(\mathfrak{m}\)-measurable sets are \(\mathfrak{m}_E\)-measurable.
Everything comes from \(\bigl(\bigcup_n B_n\bigr) \cap E = \bigcup_n (B_n \cap E)\) and the identities
\begin{equation*} (B \cap A) \cap E = (B \cap E) \cap A, \qquad (B\setminus A) \cap E = (B \cap E)\setminus A . \end{equation*}
\(\mathfrak{m}_E\) is an outer measure (Definition 1.11.1): its values lie in \([0,+\infty]\), \(\mathfrak{m}_E(\emptyset) = \mathfrak{m}(\emptyset) = 0\), monotonicity follows from \(B_1 \cap E \subset B_2 \cap E\) for \(B_1 \subset B_2\), and by countable subadditivity of \(\mathfrak{m}\),
\begin{equation*} \begin{aligned} \mathfrak{m}_E\Bigl(\bigcup_{n=1}^{\infty} B_n\Bigr) &= \mathfrak{m}\Bigl(\bigcup_{n=1}^{\infty} (B_n\cap E)\Bigr) \le \sum_{n=1}^{\infty} \mathfrak{m}_E(B_n). \end{aligned} \end{equation*}
\(\mathfrak{M}_{\mathfrak{m}} \subset \mathfrak{M}_{\mathfrak{m}_E}\): for \(A \in \mathfrak{M}_{\mathfrak{m}}\) and an arbitrary test set \(B\), applying the \(\mathfrak{m}\)-measurability of \(A\) to the test set \(B \cap E\),
\begin{equation*} \begin{aligned} \mathfrak{m}_E(B \cap A) + \mathfrak{m}_E(B\setminus A) &= \mathfrak{m}\bigl((B\cap E)\cap A\bigr) + \mathfrak{m}\bigl((B\cap E)\setminus A\bigr) \\ &= \mathfrak{m}(B\cap E) = \mathfrak{m}_E(B). \end{aligned} \end{equation*}
Let \(\tau\) be an additive, but not countably additive nonnegative set function that is defined on the class of all subsets of \([0,1]\) and coincides with Lebesgue measure on all Lebesgue measurable sets (see Example 1.12.29). Show that the corresponding outer measure \(\mathfrak{m}\) from Example 1.11.5 is identically zero under the continuum hypothesis.
\(\mathfrak{m} \equiv 0\), by Ulam’s theorem. Method I of Example 1.11.5 produces the Caratheodory outer measure
\begin{equation*} \mathfrak{m}(A) = \inf\Bigl\{ \sum_{n=1}^{\infty} \tau(X_n) : X_n \subset X, \ A \subset \bigcup_{n=1}^{\infty} X_n \Bigr\} \end{equation*}
on \(X = [0,1]\), finite since \(X_1 = X\) is an admissible cover and \(\tau(X) = \lambda([0,1]) = 1\). The class \(\mathfrak{X} = 2^{X}\) is an algebra on which \(\tau\) is nonnegative, additive and finite, so Theorem 1.11.8 makes every subset of \(X\) \(\mathfrak{m}\)-measurable; hence \(\mathfrak{M}_{\mathfrak{m}} = 2^{X}\) and, by Theorem 1.11.4(iii), \(\mathfrak{m}\) is a finite nonnegative countably additive measure on all subsets of \([0,1]\). It vanishes on singletons, \(\{x\}\) being an admissible cover of itself with \(\tau(\{x\}) = \lambda(\{x\}) = 0\). Under the continuum hypothesis \([0,1]\) has cardinality \(\aleph_1\), so Theorem 1.12.40 (Ulam), in the form of Corollary 1.12.41, gives \(\mathfrak{m} \equiv 0\).
Prove that if \(\mathfrak{X} \subset \mathfrak{M}_{\mathfrak{m}}\), then Method I from Example 1.11.5 gives a regular outer measure.
The \(\mathfrak{m}\)-measurable envelope of \(A\) is \(B := \bigcap_{k} B_k\), where \(B_k\) is the union of a \(1/k\)-efficient \(\mathfrak{X}\)-cover of \(A\). Here \(\mathfrak{m} = \tau^*\) is the Method I outer measure of Example 1.11.5, so \(\mathfrak{M}_{\mathfrak{m}}\) is a \(\sigma\)-algebra (Theorem 1.11.4(iii)) and \(\mathfrak{m}(Y) \le \tau(Y)\) for \(Y \in \mathfrak{X}\) via the one-element cover.
(i) \(\mathfrak{m}(A) = \infty\): take \(B := X \in \mathfrak{M}_{\mathfrak{m}}\), and monotonicity forces \(\mathfrak{m}(B) = \infty\).
(ii) \(\mathfrak{m}(A) < \infty\): for each \(k\) pick \(X_{k,n} \in \mathfrak{X}\) with \(A \subset B_k := \bigcup_n X_{k,n}\) and \(\sum_n \tau(X_{k,n}) \le \mathfrak{m}(A) + 1/k\). Then \(\mathfrak{X} \subset \mathfrak{M}_{\mathfrak{m}}\) puts each \(B_k\), hence \(B := \bigcap_k B_k \supset A\), in \(\mathfrak{M}_{\mathfrak{m}}\), and monotonicity with countable subadditivity gives
\begin{equation*} \mathfrak{m}(A) \le \mathfrak{m}(B) \le \mathfrak{m}(B_k) \le \sum_{n=1}^{\infty} \tau(X_{k,n}) \le \mathfrak{m}(A) + \frac{1}{k} \end{equation*}
for every \(k\), so \(\mathfrak{m}(B) = \mathfrak{m}(A)\) and \(\mathfrak{m}\) is regular (Definition 1.11.6).
Let \(\mathcal{S}\) be a collection of subsets of a set \(X\), closed with respect to finite unions and finite intersections and containing the empty set, i.e., a lattice of sets (e.g., the class of all closed sets or the class of all open sets in \([0,1]\)).
(i) Suppose that on \(\mathcal{S}\) we have a modular set function \(m\), i.e., \(m(\emptyset) = 0\) and \(m(A\cup B) + m(A \cap B) = m(A) + m(B)\) for all \(A, B \in \mathcal{S}\). Show that by the equality \(m(A\setminus B) = m(A) - m(B)\), \(A, B \in \mathcal{S}\), \(B \subset A\), the function \(m\) uniquely extends to an additive set function (which, in particular, is well-defined) on the semiring formed by the differences of elements in \(\mathcal{S}\) (see Exercise 1.12.51), and then uniquely extends to an additive set function on the ring generated by \(\mathcal{S}\).
(ii) Give an example showing that in (i) one cannot replace the modularity by the additivity even if \(m\) is nonnegative, monotone and subadditive on \(\mathcal{S}\).
(i) The extension is \(\widetilde{m}(A\setminus B) := m(A) - m(B)\) on the semiring \(\mathcal{R}_0 := \{A\setminus B : A,B \in \mathcal{S},\ B \subset A\}\), whose generated ring is the ring \(\mathcal{R}\) generated by \(\mathcal{S}\) (Exercise 1.12.51); note \(\mathcal{S} \subset \mathcal{R}_0\), as \(A = A \setminus \emptyset\).
Well defined: if \(A_1\setminus A_1’ = A_2\setminus A_2’\) with \(A_i’ \subset A_i\) in \(\mathcal{S}\), then \(A_1\cup A_2’ = A_2\cup A_1’\) and \(A_1\cap A_2’ = A_1’\cap A_2\) (Check!), so modularity applied to \((A_1,A_2’)\) and to \((A_2,A_1’)\) gives
\begin{equation*} m(A_1) + m(A_2’) = m(A_1\cup A_2’) + m(A_1\cap A_2’) = m(A_2) + m(A_1’). \end{equation*}
Taking \(B = \emptyset\) shows \(\widetilde{m}\) extends \(m\), and \(\widetilde{m}(\emptyset) = 0\).
Additivity rests on a localization lemma: given \(A_1,\dots,A_N \in \mathcal{S}\), some finitely additive real \(\nu\) on the algebra \(\mathcal{E}\) they generate agrees with \(m\) on the sublattice \(\mathcal{S}_0\) they generate together with \(\emptyset\). Indeed \(\mathcal{S}_0\) is \(\emptyset\) together with the sets \(\bigcup_{k \le p}\bigcap_{j \in J_k} A_j\) (closed under both operations by distributivity), hence finite, and \(\mathcal{E}\) is finite with atom set \(\Lambda\), the nonempty \(\bigcap_{j \le N} A_j^{\varepsilon_j}\) (\(A^1 := A\), \(A^0 := X\setminus A\)). Since \(A \mapsto \overline{A} := \{\alpha \in \Lambda : \alpha \subset A\}\) is a Boolean isomorphism of \(\mathcal{E}\) onto \(2^{\Lambda}\), it suffices to find weights \(w : \Lambda \to \mathbb{R}\) with \(\overline{m}(T) = \sum_{\alpha \in T} w(\alpha)\) on the finite lattice \(\mathcal{T} := \{\overline{A} : A \in \mathcal{S}_0\}\) carrying the modular \(\overline{m}(\overline{A}) := m(A)\); then \(\nu(A) := \sum_{\alpha \in \overline{A}} w(\alpha)\) serves.
Set \(w := 0\) off \(\Lambda’ := \bigcup_{T \in \mathcal{T}} T\). For \(\alpha \in \Lambda’\) let \(T_{\alpha} := \bigcap\{T \in \mathcal{T} : \alpha \in T\}\), a finite lattice intersection and hence the smallest member of \(\mathcal{T}\) containing \(\alpha\); preorder \(\Lambda’\) by \(\alpha \preceq \beta \iff \alpha \in T_{\beta} \iff T_{\alpha} \subset T_{\beta}\), and call the classes of \(T_{\alpha} = T_{\beta}\) cells. Every \(T \in \mathcal{T}\) is downward closed (\(\alpha \preceq \beta \in T\) gives \(\alpha \in T_\beta \subset T\)), hence a union of cells; conversely every downward closed \(D \subset \Lambda’\) equals \(\bigcup_{\beta \in D} T_{\beta} \in \mathcal{T}\). For a cell \(C\) put \(T_C := T_{\beta}\) (\(\beta \in C\)) and
\begin{equation*} w( C) := \overline{m}(T_C) - \overline{m}(T_C \setminus C), \end{equation*}
legitimate because \(T_C \setminus C\) is downward closed (Check!). Induct on \(\mathrm{card}(T)\) for \(\overline{m}(T) = \sum_{C \subset T} w( C)\), summed over cells: pick a cell \(C \subset T\) maximal for \(\preceq\), so \(T\setminus C\) is downward closed and lies in \(\mathcal{T}\), while \((T\setminus C)\cup T_C = T\) and \((T\setminus C)\cap T_C = T_C\setminus C\); modularity of \(\overline{m}\) then reads
\begin{equation*} \overline{m}(T) + \overline{m}(T_C\setminus C) = \overline{m}(T\setminus C) + \overline{m}(T_C), \end{equation*}
i.e. \(\overline{m}(T) = \overline{m}(T\setminus C) + w( C)\). Finally \(w(\alpha) := w( C)/\mathrm{card}( C)\) for \(\alpha \in C\).
Now let \(D = A\setminus B\) in \(\mathcal{R}_0\) be the disjoint union of \(D_i = A_i\setminus B_i\), \(i \le n\), and apply the lemma to \(A, B, A_1, B_1, \dots, A_n, B_n\): all of \(D, D_i\) lie in \(\mathcal{E}\), and
\begin{equation*} \widetilde{m}(D) = \nu(A) - \nu(B) = \nu(D) = \sum_{i=1}^{n}\nu(D_i) = \sum_{i=1}^{n}\widetilde{m}(D_i). \end{equation*}
So \(\widetilde{m}\) is additive on \(\mathcal{R}_0\) (Definition 1.3.1) and by Proposition 1.3.10 extends uniquely to an additive function on the generated ring, which is \(\mathcal{R}\). Uniqueness is forced already on \(\mathcal{R}_0\): any additive extension \(\nu\) of \(m\) splits \(A\) into \(B\) and \(A\setminus B\), whence \(\nu(A\setminus B) = m(A) - m(B)\).
(ii) Take \(X = \{0,1,2\}\), \(\mathcal{S} = \{\emptyset, \{1\}, \{0,1\}, \{1,2\}, X\}\), with \(m(\emptyset) = 0\), \(m(X) = 2\), and \(m = 1\) on the other three. This \(\mathcal{S}\) is a lattice, the only noncomparable pair \(\{0,1\},\{1,2\}\) having union \(X\) and intersection \(\{1\}\); \(m\) is nonnegative, monotone, subadditive (only \(m(X) = 2 \le 1+1\) is nontrivial), and additive in the sense of Definition 1.3.1 because every nonempty member contains \(1\). It is not modular: \(m(X) + m(\{1\}) = 3 \neq 2 = m(\{0,1\}) + m(\{1,2\})\).
The generated ring contains \(\{0\} = \{0,1\}\setminus\{1\}\) and \(\{2\} = \{1,2\}\setminus\{1\}\), hence equals \(2^X\). An additive \(\nu\) on it extending \(m\) would satisfy \(\nu(\{0\}) = \nu(\{2\}) = 0\) and \(\nu(\{1\}) = 1\), so
\begin{equation*} \nu(X) = 0 + 1 + 0 = 1 \neq 2 = m(X), \end{equation*}
a contradiction.
Exercises 1.12.124–1.12.130
Suppose that \(\mathcal{F}\) is a family of subsets of a set \(X\), \(\emptyset \in \mathcal{F}\). Let \(\tau \colon \mathcal{F} \to [0, +\infty]\) be a set function with \(\tau(\emptyset) = 0\). Let us define \(\tau_*\) on all sets \(A \subset X\) by formula (1.12.8).
(i) Prove that if \(A_1, \dots, A_n \subset X\) are disjoint sets and \(A_1 \cup \cdots \cup A_n \subset A\), then one has \(\tau_*(A) \ge \sum_{j=1}^{n} \tau_*(A_j)\).
(ii) Prove that \(\tau_*\) coincides with \(\tau\) on \(\mathcal{F}\) if and only if, for all pairwise disjoint sets \(F_1, \dots, F_n \in \mathcal{F}\) and all \(F \in \mathcal{F}\) with \(\bigcup_{j=1}^{n} F_j \subset F\), one has \(\tau(F) \ge \sum_{j=1}^{n} \tau(F_j)\).
(iii) Prove that if \(\tau\) satisfies the condition in (ii) and the class \(\mathcal{F}\) is closed with respect to finite unions of disjoint sets, then
\begin{equation*} \tau_*(A) = \sup\{\tau(F),\ F \in \mathcal{F},\ F \subset A\}, \qquad \forall\, A \subset X. \end{equation*}
All three parts run off the finite form (1.12.9) of (1.12.8),
\begin{equation*} \tau_*(A) = \sup\Big\{ \sum_{j=1}^{n} \tau(F_j) : n \in \mathbb{N}, \ F_j \in \mathcal{F}, \ F_j \subset A \text{ disjoint} \Big\}, \end{equation*}
legitimate since padding with copies of \(\emptyset\) changes no sum and a nonnegative series is the supremum of its partial sums; \(\tau_*\) is monotone, with \(\tau_*(F) \ge \tau(F)\) on \(\mathcal{F}\) by the one-element family.
(i) Fix reals \(c_i < \tau_*(A_i)\) and disjoint \(F_{i1},\dots,F_{i\,n(i)} \in \mathcal{F}\) inside \(A_i\) with \(\sum_j \tau(F_{ij}) > c_i\). The doubly indexed family is disjoint (the \(A_i\) being disjoint) and sits inside \(A\), so
\begin{equation*} \tau_*(A) \ge \sum_{i=1}^{n} \sum_{j=1}^{n(i)} \tau(F_{ij}) > \sum_{i=1}^{n} c_i, \end{equation*}
and letting each \(c_i \uparrow \tau_*(A_i)\) gives \(\tau_*(A) \ge \sum_{i \le n} \tau_*(A_i)\), the value \(+\infty\) included.
(ii) If the inequality holds, every disjoint family \(F_1,\dots,F_n \in \mathcal{F}\) inside \(F \in \mathcal{F}\) has \(\sum_j \tau(F_j) \le \tau(F)\), so \(\tau_*(F) \le \tau(F)\); with \(\tau_*(F) \ge \tau(F)\) this is equality. Conversely, if \(\tau_* = \tau\) on \(\mathcal{F}\) and disjoint \(F_1,\dots,F_n\) have union inside \(F \in \mathcal{F}\), they compete in the supremum for \(\tau_*(F)\), so \(\sum_j \tau(F_j) \le \tau_*(F) = \tau(F)\).
(iii) Write \(s(A) := \sup\{\tau(F) : F \in \mathcal{F},\ F \subset A\}\), a supremum over a family containing \(\emptyset\). One-element families give \(\tau_*(A) \ge s(A)\). Conversely, for disjoint \(F_1,\dots,F_n \in \mathcal{F}\) inside \(A\), closedness under finite disjoint unions puts \(E := \bigcup_{j \le n} F_j\) in \(\mathcal{F}\) with \(E \subset A\), so the hypothesis of (ii) with \(F = E\) yields
\begin{equation*} \sum_{j=1}^{n} \tau(F_j) \le \tau(E) \le s(A); \end{equation*}
taking the supremum gives \(\tau_*(A) \le s(A)\), hence \(\tau_*(A) = s(A)\).
Let \(\mathcal{F}\) and \(\tau\) be the same as in the previous exercise.
(i) Prove that the outer measure \(\tau^*\) coincides with \(\tau\) on \(\mathcal{F}\) precisely when \(\tau(F) \le \sum_{n=1}^{\infty} \tau(F_n)\) whenever \(F, F_n \in \mathcal{F}\) and \(F \subset \bigcup_{n=1}^{\infty} F_n\).
(ii) Prove that if the condition in (i) is fulfilled and the class \(\mathcal{F}\) is closed with respect to countable unions, then
\begin{equation*} \tau^*(A) = \inf\{\tau(F),\ F \in \mathcal{F},\ A \subset F\}, \qquad \forall\, A \subset X. \end{equation*}
Both parts are the dual of Exercise 1.12.124, with the Method I outer measure of Example 1.11.5,
\begin{equation*} \tau^*(A) = \inf\Big\{ \sum_{n=1}^{\infty} \tau(F_n) : F_n \in \mathcal{F}, \ A \subset \bigcup_{n=1}^{\infty} F_n \Big\}, \end{equation*}
equal to \(+\infty\) when no such cover exists. The cover \(F, \emptyset, \emptyset, \dots\) gives \(\tau^*(F) \le \tau(F)\) for all \(F \in \mathcal{F}\).
(i) If the condition holds, every \(\mathcal{F}\)-cover of \(F \in \mathcal{F}\) has total mass at least \(\tau(F)\), so the infimum gives \(\tau^*(F) \ge \tau(F)\), hence equality. Conversely, if \(\tau^* = \tau\) on \(\mathcal{F}\) and \(F \subset \bigcup_n F_n\) with \(F, F_n \in \mathcal{F}\), that cover competes in the infimum, so \(\tau(F) = \tau^*(F) \le \sum_n \tau(F_n)\).
(ii) Write \(q(A) := \inf\{\tau(F) : F \in \mathcal{F},\ A \subset F\}\), again \(+\infty\) over an empty family. The cover \(F, \emptyset, \dots\) of any \(F \supset A\) gives \(\tau^*(A) \le q(A)\). Conversely, if \(A\) has any \(\mathcal{F}\)-cover \(\{F_n\}\), then closedness under countable unions puts \(F := \bigcup_n F_n\) in \(\mathcal{F}\) with \(A \subset F\), and (i) applied to \(F \subset \bigcup_n F_n\) yields
\begin{equation*} q(A) \le \tau(F) \le \sum_{n=1}^{\infty} \tau(F_n); \end{equation*}
taking the infimum gives \(q(A) \le \tau^*(A)\). If \(A\) has no cover, both sides are \(+\infty\), since a cover would supply a containing member of \(\mathcal{F}\). Hence \(\tau^*(A) = q(A)\).
Suppose that \(\mathcal{F}\) is a class of subsets of a space \(X\), \(\emptyset \in \mathcal{F}\). Let \(\tau \colon \mathcal{F} \to [0, +\infty]\) be a set function with \(\tau(\emptyset) = 0\). Prove that the following conditions are equivalent:
(i) \(\tau^*\) coincides with \(\tau\) on \(\mathcal{F}\) and \(\mathcal{F} \subset \mathfrak{M}_{\tau^*}\);
(ii) \(\tau(A) = \tau^*(A \cap B) + \tau^*(A \setminus B)\) for all \(A, B \in \mathcal{F}\).
(ii) is just the Caratheodory test (Definition 1.11.2) for \(B\) run on the test set \(A\), once \(\tau^* = \tau\) on \(\mathcal{F}\); here \(\tau^*\) is the Method I outer measure of Example 1.11.5, so subadditivity makes \(\ge\) the only nontrivial direction of that test.
(i) implies (ii): for \(A, B \in \mathcal{F}\), measurability of \(B\) tested on \(E = A\) gives \(\tau^*(A) = \tau^*(A\cap B) + \tau^*(A\setminus B)\), and \(\tau^*(A) = \tau(A)\).
(ii) implies (i): taking \(B = \emptyset\) and using \(\tau^*(\emptyset) = 0\) gives \(\tau(A) = \tau^*(A)\) on \(\mathcal{F}\). Now fix \(F \in \mathcal{F}\) and \(E \subset X\); if \(E\) has no \(\mathcal{F}\)-cover then \(\tau^*(E) = +\infty\), so let \(\{F_j\} \subset \mathcal{F}\) cover \(E\). Applying (ii) to each pair \((F_j, F)\), then countable subadditivity and monotonicity against \(E \cap F \subset \bigcup_j (F_j \cap F)\) and \(E\setminus F \subset \bigcup_j (F_j\setminus F)\),
\begin{equation*} \sum_{j=1}^{\infty} \tau(F_j) = \sum_{j=1}^{\infty} \tau^*(F_j \cap F) + \sum_{j=1}^{\infty} \tau^*(F_j \setminus F) \ge \tau^*(E \cap F) + \tau^*(E \setminus F). \end{equation*}
Taking the infimum over covers yields \(\tau^*(E) \ge \tau^*(E\cap F) + \tau^*(E\setminus F)\), so \(F \in \mathfrak{M}_{\tau^*}\) and \(\mathcal{F} \subset \mathfrak{M}_{\tau^*}\).
Suppose that \(\mathcal{F}\) is a class of subsets of a space \(X\), \(\emptyset \in \mathcal{F}\). Let \(\tau \colon \mathcal{F} \to [0, +\infty]\) be a set function with \(\tau(\emptyset) = 0\). Denote by \(\tau_*\) the corresponding inner measure (see formula (1.12.8)). Prove that the following conditions are equivalent:
(i) \(\tau_*\) coincides with \(\tau\) on \(\mathcal{F}\) and \(\mathcal{F} \subset \mathfrak{M}_{\tau_*}\);
(ii) \(\tau(A) = \tau_*(A \cap B) + \tau_*(A \setminus B)\), \(\forall\, A, B \in \mathcal{F}\).
This is the exact dual of Exercise 1.12.126, with countable covers of \(E\) replaced by finite disjoint families inside \(E\). Work with the finite form (1.12.9) of \(\tau_*\) used in Exercise 1.12.124; then \(\tau_*(\emptyset) = 0\), \(\tau_*\) is monotone, and \(\tau_*(F) \ge \tau(F)\) on \(\mathcal{F}\), so Definition 1.11.2 applies to \(\tau_*\). The structural difference: \(\tau_*\) need not be subadditive, but Exercise 1.12.124(i) with \(A_1 = E\cap F\), \(A_2 = E\setminus F\) gives the superadditive half
\begin{equation*} \tau_*(E) \ge \tau_*(E \cap F) + \tau_*(E \setminus F), \qquad \forall\, E, F \subset X, \tag{\(\ast\)} \end{equation*}
so here \(\le\) is what must be proved.
(i) implies (ii): measurability of \(B \in \mathcal{F}\) tested on \(E = A\) gives \(\tau_*(A) = \tau_*(A\cap B) + \tau_*(A\setminus B)\), and \(\tau_*(A) = \tau(A)\).
(ii) implies (i): \(B = \emptyset\) and \(\tau_*(\emptyset) = 0\) give \(\tau = \tau_*\) on \(\mathcal{F}\). Fix \(F \in \mathcal{F}\), \(E \subset X\), and disjoint \(F_1,\dots,F_n \in \mathcal{F}\) inside \(E\). The traces \(F_j \cap F\) are disjoint inside \(E \cap F\) and the traces \(F_j\setminus F\) are disjoint inside \(E\setminus F\), so (ii) on each pair \((F_j, F)\) plus superadditivity (Exercise 1.12.124(i)) twice gives
\begin{equation*} \sum_{j=1}^{n} \tau(F_j) = \sum_{j=1}^{n} \tau_*(F_j \cap F) + \sum_{j=1}^{n} \tau_*(F_j \setminus F) \le \tau_*(E \cap F) + \tau_*(E \setminus F). \end{equation*}
Taking the supremum gives \(\tau_*(E) \le \tau_*(E\cap F) + \tau_*(E\setminus F)\), which with \((\ast)\) puts \(F \in \mathfrak{M}_{\tau_*}\); hence \(\mathcal{F} \subset \mathfrak{M}_{\tau_*}\).
(i) Show that if in the situation of the previous exercise we have one of the equivalent conditions (i) and (ii), then on the algebra \(\mathcal{A}_{\mathcal{F}}\) generated by \(\mathcal{F}\), there exists an additive set function \(\tau_0\) that coincides with \(\tau\) on \(\mathcal{F}\).
(ii) Show that if, in addition to the hypotheses in (i), it is known that
\begin{equation*} \tau_*(F) \le \sum_{n=1}^{\infty} \tau_*(F_n) \quad \text{whenever } F, F_n \in \mathcal{A}_{\mathcal{F}} \text{ and } F \subset \bigcup_{n=1}^{\infty} F_n, \end{equation*}
then there exists a countably additive measure \(\mu\) on \(\sigma(\mathcal{F})\) that coincides with \(\tau\) on \(\mathcal{F}\).
(i) Take \(\tau_0 := \tau_*|_{\mathcal{A}_{\mathcal{F}}}\). By Exercise 1.12.127 the assumed condition gives \(\tau_*|_{\mathcal{F}} = \tau\) and \(\mathcal{F} \subset \mathfrak{M}_{\tau_*}\); since \(\tau_*\) is a \([0,+\infty]\)-valued set function on all of \(2^X\) vanishing at \(\emptyset\), Theorem 1.11.4(i) makes \(\mathfrak{M}_{\tau_*}\) an algebra on which \(\tau_*\) is additive, so minimality of \(\mathcal{A}_{\mathcal{F}}\) gives \(\mathcal{F} \subset \mathcal{A}_{\mathcal{F}} \subset \mathfrak{M}_{\tau_*}\). Hence \(\tau_0\) inherits additivity from \(\tau_*\) (disjoint \(A, B\) and their union all lie in \(\mathfrak{M}_{\tau_*}\)) and equals \(\tau\) on \(\mathcal{F}\).
(ii) The same \(\tau_0\) is countably additive on \(\mathcal{A}_{\mathcal{F}}\). Being nonnegative and additive there it is monotone, so a disjoint decomposition \(A = \bigcup_n A_n\) inside \(\mathcal{A}_{\mathcal{F}}\) satisfies \(\sum_{n \le N} \tau_0(A_n) = \tau_0(\bigcup_{n \le N} A_n) \le \tau_0(A)\) for each \(N\), whence \(\sum_n \tau_0(A_n) \le \tau_0(A)\); the added hypothesis, which on \(\mathcal{A}_{\mathcal{F}}\) says exactly that \(\tau_0\) is countably subadditive, supplies the reverse inequality.
Now run Method I (Example 1.11.5) on the algebra \(\mathcal{A}_{\mathcal{F}}\) and \(\tau_0\),
\begin{equation*} m(A) = \inf\Big\{ \sum_{n=1}^{\infty} \tau_0(A_n) : A_n \in \mathcal{A}_{\mathcal{F}}, \ A \subset \bigcup_{n=1}^{\infty} A_n \Big\}, \end{equation*}
an outer measure with no vacuous case, as \(X \in \mathcal{A}_{\mathcal{F}}\) covers every \(A\). Theorem 1.11.8, whose hypothesis of countable additivity on an algebra was just verified, gives \(\mathcal{A}_{\mathcal{F}} \subset \mathfrak{M}_m\) and \(m = \tau_0\) on \(\mathcal{A}_{\mathcal{F}}\), while Theorem 1.11.4(iii) makes \(\mathfrak{M}_m\) a \(\sigma\)-algebra carrying the countably additive \(m\). So \(\sigma(\mathcal{F}) \subset \mathfrak{M}_m\), and \(\mu := m|_{\sigma(\mathcal{F})}\) is a measure with \(\mu = \tau_0 = \tau\) on \(\mathcal{F}\).
Let \((X, \mathcal{A}, \mu)\) be a measure space, where \(\mathcal{A}\) is a \(\sigma\)-algebra and \(\mu\) is a countably additive measure with values in \([0, +\infty]\). Denote by \(\mathcal{L}_{\mu}\) the class of all sets \(E \subset X\) for each of which there exist two sets \(A_1, A_2 \in \mathcal{A}\) with \(A_1 \subset E \subset A_2\) and \(\mu(A_2 \setminus A_1) = 0\).
(i) Show that \(\mathcal{L}_{\mu}\) is a \(\sigma\)-algebra, coincides with \(\mathcal{A}_{\mu}\) and belongs to \(\mathfrak{M}_{\mu^*}\).
(ii) Show that if the measure \(\mu\) is \(\sigma\)-finite, then \(\mathcal{L}_{\mu}\) coincides with \(\mathfrak{M}_{\mu^*}\).
(iii) Let \(X = [0,1]\), let \(\mathcal{A}\) be the \(\sigma\)-algebra generated by all singletons, and let the measure \(\mu\) with values in \([0, +\infty]\) be defined as follows: \(\mu(A)\) is the cardinality of \(A\), \(A \in \mathcal{A}\). Show that \(\mathfrak{M}_{\mu^*}\) contains all sets, but \([0, 1/2] \notin \mathcal{L}_{\mu}\).
Two facts about \(\mu^*\), the outer measure generated by \(\mu\) on \(\mathcal{A}\) (an outer measure by Lemma 1.5.4), drive all three parts.
(a) Because \(\mathcal{A}\) is a \(\sigma\)-algebra, \(\mu^*(S) = \inf\{\mu(A) : A \in \mathcal{A},\ S \subset A\}\) (any cover has its union in \(\mathcal{A}\) with no larger measure) and the infimum is attained at the measurable envelope \(D := \bigcap_k D_k\), where \(D_k \supset S\) in \(\mathcal{A}\) has \(\mu(D_k) \le \mu^*(S) + 1/k\).
(b) \(\mathcal{A} \subset \mathfrak{M}_{\mu^*}\): for \(A \in \mathcal{A}\), \(\mu^*(S) < \infty\), and \(B \supset S\) in \(\mathcal{A}\) with \(\mu(B) \le \mu^*(S) + \varepsilon\),
\begin{equation*} \mu^*(S \cap A) + \mu^*(S \setminus A) \le \mu(B \cap A) + \mu(B \setminus A) = \mu(B) \le \mu^*(S) + \varepsilon. \end{equation*}
(i) \(\mathcal{L}_\mu\) is a \(\sigma\)-algebra: complements come from \(X\setminus A_2 \subset X\setminus E \subset X\setminus A_1\) with the same null difference, and countable unions from \(\bigcup_n A_1^{(n)} \subset \bigcup_n E_n \subset \bigcup_n A_2^{(n)}\), whose difference sits inside the null set \(\bigcup_n (A_2^{(n)}\setminus A_1^{(n)})\).
\(\mathcal{L}_\mu \subset \mathcal{A}_\mu\) since \(E \triangle A_1 = E\setminus A_1 \subset A_2\setminus A_1\) has \(\mu^* = 0\) (Definition 1.5.1). Conversely, for \(E \in \mathcal{A}_\mu\) pick \(A_n \in \mathcal{A}\) with \(\mu^*(E \triangle A_n) < 2^{-n}\) and, by (a), \(B_n \in \mathcal{A}\) with \(E \triangle A_n \subset B_n\), \(\mu(B_n) < 2^{-n}\); then \(A_n \setminus B_n \subset E \subset A_n \cup B_n\), so
\begin{equation*} A_1’ := \bigcup_{n} (A_n \setminus B_n) \subset E \subset \bigcap_{n} (A_n \cup B_n) =: A_2’, \end{equation*}
and for each \(n\), \(A_2’ \setminus A_1’ \subset (A_n\cup B_n)\setminus(A_n\setminus B_n) \subset B_n\) (Check!), giving \(\mu(A_2’\setminus A_1’) < 2^{-n}\) for all \(n\), hence \(0\). So \(\mathcal{L}_\mu = \mathcal{A}_\mu\).
Finally \(\mathcal{L}_\mu \subset \mathfrak{M}_{\mu^*}\): \(E\setminus A_1 \subset A_2\setminus A_1\) has \(\mu^* = 0\), hence lies in \(\mathfrak{M}_{\mu^*}\) by completeness (Theorem 1.11.4(iii)), while \(A_1 \in \mathfrak{M}_{\mu^*}\) by (b), so \(E = A_1 \cup (E\setminus A_1) \in \mathfrak{M}_{\mu^*}\).
(ii) Only \(\mathfrak{M}_{\mu^*} \subset \mathcal{L}_\mu\) is left. Disjointify a countable cover of finite measure into \(X = \bigcup_n X_n\), \(X_n \in \mathcal{A}\), \(\mu(X_n) < \infty\). Fix \(E \in \mathfrak{M}_{\mu^*}\) and \(n\), and use (a) for envelopes \(P_n, Q_n \in \mathcal{A}\) of \(E \cap X_n\) and of \(X_n\setminus E\), both taken inside \(X_n\) (intersect with \(X_n\); harmless), so \(P_n \cup Q_n = X_n\). Testing measurability of \(E\) on \(X_n\) gives \(\mu(P_n) + \mu(Q_n) = \mu(X_n)\), while modularity gives
\begin{equation*} \mu(P_n) + \mu(Q_n) = \mu(P_n \cup Q_n) + \mu(P_n \cap Q_n) = \mu(X_n) + \mu(P_n \cap Q_n), \end{equation*}
all terms finite, so \(\mu(P_n \cap Q_n) = 0\). Put \(R_n := X_n \setminus Q_n \in \mathcal{A}\), so \(R_n \subset E \cap X_n \subset P_n\) and \(P_n \setminus R_n = P_n \cap Q_n\) is null. Then \(A_1 := \bigcup_n R_n \subset E \subset \bigcup_n P_n =: A_2\), and disjointness of the \(X_n\) with \(P_n, R_n \subset X_n\) makes \(A_2 \setminus A_1 = \bigcup_n (P_n\setminus R_n)\) null. Hence \(E \in \mathcal{L}_\mu\).
(iii) Here \(\mathcal{A}\) is the countable-cocountable \(\sigma\)-algebra and \(\mu\) is counting measure on it. First \(\mu^*(S) = \operatorname{card} S\) for every \(S \subset X\): finite \(S\) lies in \(\mathcal{A}\), giving \(\le\), while any cover satisfies \(\sum_n \mu(A_n) \ge \mu(\bigcup_n A_n) \ge \operatorname{card} S\), which for infinite \(S\) run over its finite subsets forces \(\mu^*(S) = +\infty\). Consequently, for all \(E, A \subset X\),
\begin{equation*} \mu^*(E) = \operatorname{card}(E \cap A) + \operatorname{card}(E \setminus A) = \mu^*(E \cap A) + \mu^*(E \setminus A) \end{equation*}
(both sides \(+\infty\) when \(E\) is infinite), so \(\mathfrak{M}_{\mu^*} = 2^{[0,1]}\). But \(\mu(A_2\setminus A_1) = 0\) forces \(A_2 = A_1\) here, so \(\mathcal{L}_\mu = \mathcal{A}\); and \([0,1/2]\) is uncountable with uncountable complement, hence \([0,1/2] \notin \mathcal{L}_\mu\).
Let us consider the following modification of Example 1.11.5. Let \(\mathfrak{X}\) be a family of subsets of a set \(X\) such that \(\emptyset \in \mathfrak{X}\). Suppose that we are given a function \(\tau \colon \mathfrak{X} \to [0, +\infty]\) with \(\tau(\emptyset) = 0\). Set
\begin{equation*} \widetilde{\mathfrak{m}}(A) = \inf\Big\{ \sum_{n=1}^{\infty} \tau(X_n) : X_n \in \mathfrak{X},\ A \subset \bigcup_{n=1}^{\infty} X_n \Big\} \end{equation*}
if such sets \(X_n\) exist and otherwise let \(\widetilde{\mathfrak{m}}(A) = \sup \widetilde{\mathfrak{m}}(A^{\prime})\), where sup is taken over all sets \(A^{\prime} \subset A\) that can be covered by a sequence of sets in \(\mathfrak{X}\).
(i) Show that \(\widetilde{\mathfrak{m}}\) is an outer measure.
(ii) Let \(X = [0,1] \times [0,1]\), \(\mathfrak{X} = \big\{ [a,b) \times t,\ a, b, t \in [0,1],\ a \le b \big\}\), \(\tau([a,b) \times t) = b - a\). Let \(\mathfrak{m}\) be given by formula (1.11.5). Show that \(\mathfrak{m}\) and \(\widetilde{\mathfrak{m}}\) do not coincide and that there exists a set \(E \in \mathfrak{M}_{\mathfrak{m}} \cap \mathfrak{M}_{\widetilde{\mathfrak{m}}}\) such that \(\mathfrak{m}(E) \ne \widetilde{\mathfrak{m}}(E)\).
(i) The two clauses of the definition merge into
\begin{equation*} \widetilde{\mathfrak{m}}(A) = \sup\{ \mathfrak{m}(A^{\prime}) : A^{\prime} \subset A, \ A^{\prime} \in \mathcal{C} \}, \qquad \forall\, A \subset X, \tag{\(\dagger\)} \end{equation*}
where \(\mathcal{C}\) is the class of sets coverable by a sequence from \(\mathfrak{X}\) (hereditary, and closed under countable unions by concatenation) and \(\mathfrak{m}\) is the Method I outer measure (1.11.5), equal to \(+\infty\) off \(\mathcal{C}\). Indeed, for \(A \notin \mathcal{C}\) this is the definition; for \(A \in \mathcal{C}\) the set \(A\) itself competes while monotonicity of \(\mathfrak{m}\) caps the supremum at \(\mathfrak{m}(A)\).
The three properties of Definition 1.11.1 now read off \((\dagger)\): \(\widetilde{\mathfrak{m}}(\emptyset) = 0\) since \(\emptyset \in \mathfrak{X} \subset \mathcal{C}\); monotonicity since \(A \subset B\) makes the supremum for \(B\) run over a larger family; and for \(A \subset \bigcup_n A_n\) and coverable \(A^{\prime} \subset A\), each \(A^{\prime} \cap A_n\) is a coverable subset of \(A_n\), so countable subadditivity of \(\mathfrak{m}\) gives
\begin{equation*} \mathfrak{m}(A^{\prime}) \le \sum_{n=1}^{\infty} \mathfrak{m}(A^{\prime} \cap A_n) \le \sum_{n=1}^{\infty} \widetilde{\mathfrak{m}}(A_n), \end{equation*}
and the supremum over \(A^{\prime}\) yields \(\widetilde{\mathfrak{m}}(A) \le \sum_n \widetilde{\mathfrak{m}}(A_n)\).
(ii) Take for \(E\) the diagonal \(\{(t,t) : t \in [0,1]\}\): then \(\mathfrak{m}(E) = +\infty\), \(\widetilde{\mathfrak{m}}(E) = 0\), and \(E \in \mathfrak{M}_{\mathfrak{m}} \cap \mathfrak{M}_{\widetilde{\mathfrak{m}}}\). Each member of \(\mathfrak{X}\) lies in a single horizontal line, so every coverable set meets at most countably many lines; the second projection of \(E\) is all of \([0,1]\), whence \(E \notin \mathcal{C}\) and \(\mathfrak{m}(E) = +\infty\).
Any coverable \(E^{\prime} \subset E\) is therefore at most countable, say \(E^{\prime} = \{(s_k,s_k)\}_{k \ge 1}\), and each of its points lies in some \([a_n,b_n) \times \{t_n\}\), forcing \(s_k = t_n < b_n \le 1\). Given \(\varepsilon > 0\), the sets
\begin{equation*} X_k := \big[ s_k,\ \min(s_k + \varepsilon 2^{-k},\, 1) \big) \times \{ s_k \} \in \mathfrak{X} \end{equation*}
cover \(E^{\prime}\) with \(\sum_k \tau(X_k) \le \varepsilon\), so \(\mathfrak{m}(E^{\prime}) = 0\) and \((\dagger)\) gives \(\widetilde{\mathfrak{m}}(E) = 0\). In particular \(\mathfrak{m} \neq \widetilde{\mathfrak{m}}\).
Measurability: \(\widetilde{\mathfrak{m}}(E) = 0\) puts \(E \in \mathfrak{M}_{\widetilde{\mathfrak{m}}}\) by completeness (Theorem 1.11.4(iii)). For \(\mathfrak{m}\), let \(S \subset X\) with \(\mathfrak{m}(S) < \infty\); then \(S \in \mathcal{C}\) meets only countably many lines, so \(S \cap E\) is at most countable with every point in a member of \(\mathfrak{X}\), hence with first coordinate \(< 1\), and the covering argument above gives \(\mathfrak{m}(S \cap E) = 0\). Thus
\begin{equation*} \mathfrak{m}(S \cap E) + \mathfrak{m}(S \setminus E) = \mathfrak{m}(S \setminus E) \le \mathfrak{m}(S), \end{equation*}
trivially so when \(\mathfrak{m}(S) = +\infty\) as well; with subadditivity, \(E \in \mathfrak{M}_{\mathfrak{m}}\).
Exercises 1.12.131–1.12.137
Let \(\mu\) be a measure with values in \([0,+\infty]\) defined on a measurable space \((X,\mathcal{A})\). The measure \(\mu\) is called decomposable if there exists a partition of \(X\) into pairwise disjoint sets \(X_\alpha \in \mathcal{A}\) of finite measure (indexed by elements \(\alpha\) of some set \(\Lambda\)) with the following properties: (a) if \(E\cap X_\alpha \in \mathcal{A}\) for all \(\alpha\), then \(E \in \mathcal{A}\), (b) \(\mu(E) = \sum_\alpha \mu(E\cap X_\alpha)\) for each set \(E \in \mathcal{A}\), where convergence of the series \(\sum_\alpha c_\alpha\), \(c_\alpha \ge 0\), to a finite number \(s\) means by definition that among the numbers \(c_\alpha\) at most countably many are nonzero and the corresponding series converges to \(s\), and the divergence of such a series to \(+\infty\) means the divergence of some of its countable subseries.
(i) Give an example of a measure that is not decomposable.
(ii) Show that a measure \(\mu\) is decomposable precisely when there exists a partition of \(X\) into disjoint sets \(X_\alpha\) of positive measure having property (a) and property (b’): if \(A \in \mathcal{A}\) and \(\mu(A\cap X_\alpha) = 0\) for all \(\alpha\), then \(\mu(A) = 0\).
(i) Take \(X\) uncountable, \(\mathcal{A} = 2^X\), and
\begin{equation*} \mu(A) = 0 \ \text{ if } A \text{ is at most countable}, \qquad \mu(A) = +\infty \ \text{ otherwise}, \end{equation*}
a measure because a disjoint countable union is at most countable exactly when all its terms are. Every piece of a decomposition has finite, hence zero, measure, so (b) at \(E = X\) gives \(+\infty = \mu(X) = \sum_\alpha 0 = 0\).
Throughout, the statement’s convention makes \(\sum_\alpha c_\alpha\) (\(c_\alpha \ge 0\)) the unordered sum \(\sup\{\sum_{\alpha\in F}c_\alpha : F \subset \Lambda\ \text{finite}\}\): if uncountably many \(c_\alpha\) are nonzero, then \(\{\alpha : c_\alpha > 1/n\}\) is uncountable for some \(n\), a countable subseries diverges and both sides are \(+\infty\); otherwise both are the ordinary sum of the nonzero terms.
(ii) Sufficiency. Let \(\{X_\alpha\}_{\alpha\in\Lambda}\) be a partition into members of \(\mathcal{A}\) of finite positive measure satisfying (a) and (b’) — finiteness of the pieces being part of the notion of a decomposition, without which \(\{X\}\) would satisfy (a) and (b’) for every measure. Fix \(E \in \mathcal{A}\) and put \(s(E) := \sum_\alpha \mu(E\cap X_\alpha)\). For finite \(F\subset\Lambda\) the sets \(E\cap X_\alpha\), \(\alpha\in F\), are disjoint inside \(E\), so
\begin{equation*} \sum_{\alpha\in F}\mu(E\cap X_\alpha) = \mu\Big(E\cap\bigcup_{\alpha\in F}X_\alpha\Big) \le \mu(E), \end{equation*}
whence \(s(E)\le \mu(E)\) and (b) holds when \(s(E) = +\infty\). If \(s(E) < \infty\), at most countably many indices \(\alpha_n\) have \(\mu(E\cap X_{\alpha_n}) > 0\); put \(E_0 := E\cap\bigcup_n X_{\alpha_n}\) and \(A := E\setminus E_0\), both in \(\mathcal{A}\). Countable additivity gives \(\mu(E_0) = s(E)\), while \(\mu(A\cap X_\alpha) = 0\) for every \(\alpha\) (empty for \(\alpha = \alpha_n\), inside the null set \(E\cap X_\alpha\) otherwise), so (b’) gives \(\mu(A) = 0\) and
\begin{equation*} \mu(E) = \mu(E_0) + \mu(A) = s(E). \end{equation*}
Necessity. (b’) is immediate from (b). For positivity of the pieces put \(\Lambda_0 := \{\alpha : \mu(X_\alpha) = 0\}\) and \(N := \bigcup_{\alpha\in\Lambda_0}X_\alpha\); each \(N\cap X_\beta\) is \(X_\beta\) or \(\emptyset\), so \(N\in\mathcal{A}\) by (a) and \(\mu(N) = 0\) by (b). Excluding the degenerate case \(\mu\equiv0\) (where (b) makes \(\{X\}\) a decomposition while no partition into pieces of positive measure exists), fix \(\alpha_0\) with \(\mu(X_{\alpha_0}) > 0\) and set
\begin{equation*} X^{\prime}_{\alpha_0} := X_{\alpha_0}\cup N, \qquad X^{\prime}_\alpha := X_\alpha \ \ (\alpha\notin\Lambda_0,\ \alpha\ne\alpha_0), \end{equation*}
a partition of \(X\) into members of \(\mathcal{A}\) of finite positive measure, as \(\mu(X^{\prime}_{\alpha_0}) = \mu(X_{\alpha_0}) + \mu(N) = \mu(X_{\alpha_0})\). It inherits (a): every old piece \(X_\beta\) either equals \(X^{\prime}_\beta\) or lies in \(X^{\prime}_{\alpha_0}\) (for \(\beta = \alpha_0\) and for \(\beta\in\Lambda_0\)), and in the latter case \(E\cap X_\beta = (E\cap X^{\prime}_{\alpha_0})\cap X_\beta\in\mathcal{A}\), so (a) for the original decomposition gives \(E\in\mathcal{A}\). It inherits (b’) as well: \(\mu(A\cap X^{\prime}_\alpha) = 0\) for all new \(\alpha\) forces \(\mu(A\cap X_\beta) = 0\) for all old \(\beta\) (because \(\mu(X_\beta) = 0\) for \(\beta\in\Lambda_0\), by containment otherwise), so \(\mu(A) = \sum_\beta\mu(A\cap X_\beta) = 0\) by (b).
Let \(\mu\) be a measure with values in \([0,+\infty]\) defined on a measurable space \((X,\mathcal{A})\). The measure \(\mu\) is called semifinite if every set of infinite measure has a subset of finite positive measure.
(i) Give an example of a measure with values in \([0,+\infty]\) that is not semifinite.
(ii) Give an example of a semifinite measure that is not \(\sigma\)-finite.
(iii) Prove that for any measure \(\mu\) with values in \([0,+\infty]\), defined on a \(\sigma\)-algebra \(\mathcal{A}\), the formula \(\mu_0(A) := \sup\{\mu(B)\colon B\subset A,\ B\in\mathcal{A},\ \mu(B) < \infty\}\) defines a semifinite measure with values in \([0,+\infty]\) and \(\mu\) is semifinite precisely when \(\mu = \mu_0\).
(iv) Show that every decomposable measure is semifinite.
(v) Give an example of a semifinite measure \(\mu\) with values in \([0,+\infty]\) that is defined on an algebra \(\mathcal{A}\) and has infinitely many semifinite extensions to \(\sigma(\mathcal{A})\).
(i) Take \(X = \{0\}\), \(\mathcal{A} = 2^X\), \(\mu(X) = +\infty\): the only subsets of the infinite-measure set \(X\) are \(\emptyset\) and \(X\), of measures \(0\) and \(+\infty\).
(ii) Take counting measure on an uncountable \(X\) with \(\mathcal{A} = 2^X\). A set of infinite measure is infinite, hence contains a singleton of measure \(1\); but the sets of finite measure are exactly the finite ones, so no countable union of them is \(X\).
(iii) Clearly \(\mu_0(\emptyset) = 0\), \(\mu_0\) is monotone, \(\mu_0 \le \mu\), and \(\mu_0(A) = \mu(A)\) whenever \(\mu(A) < \infty\) (take \(B = A\)).
Countable additivity: let \(A = \bigcup_{n=1}^\infty A_n\) with pairwise disjoint \(A_n\in\mathcal{A}\). For \(B\subset A\) in \(\mathcal{A}\) with \(\mu(B)<\infty\), each \(B\cap A_n\subset A_n\) has finite measure, so
\begin{equation*} \mu(B) = \sum_{n=1}^\infty \mu(B\cap A_n)\le \sum_{n=1}^\infty \mu_0(A_n), \end{equation*}
and the supremum over \(B\) gives \(\mu_0(A)\le\sum_n\mu_0(A_n)\). Conversely, fix \(N\) and \(c_n<\mu_0(A_n)\) for \(n\le N\) (with \(c_n := 0\) when \(\mu_0(A_n) = 0\)), and choose \(B_n\subset A_n\) in \(\mathcal{A}\) with \(\mu(B_n)<\infty\), \(\mu(B_n)\ge c_n\); then \(B := \bigcup_{n\le N}B_n\subset A\) has \(\mu(B) = \sum_{n\le N}\mu(B_n)<\infty\), whence
\begin{equation*} \mu_0(A)\ge \mu(B) = \sum_{n\le N}\mu(B_n) \ge \sum_{n\le N}c_n , \end{equation*}
and \(c_n\uparrow\mu_0(A_n)\), \(N\to\infty\) give \(\mu_0(A)\ge\sum_n\mu_0(A_n)\).
Semifiniteness of \(\mu_0\): if \(\mu_0(A) = \infty\), the supremum supplies \(B\subset A\) in \(\mathcal{A}\) with \(0<\mu(B)<\infty\), and \(\mu_0(B) = \mu(B)\in(0,\infty)\). So \(\mu = \mu_0\) forces \(\mu\) semifinite. Conversely let \(\mu\) be semifinite; only \(\mu(A) = \infty\) needs proof, so suppose \(c := \mu_0(A)<\infty\) and choose \(B_n\subset A\) in \(\mathcal{A}\) with \(\mu(B_n)<\infty\), \(\mu(B_n)\to c\). Each finite union \(B_1\cup\dots\cup B_n\subset A\) has finite measure, hence measure \(\le c\), so continuity from below gives \(\mu(B) = c\) for \(B := \bigcup_n B_n\). Then \(\mu(A\setminus B) = \mu(A)-\mu(B) = \infty\), and semifiniteness yields \(C\subset A\setminus B\) with \(0<\mu( C)<\infty\); now \(B\cup C\subset A\) has finite measure \(c+\mu( C)>\mu_0(A)\), contradicting the definition. Hence \(\mu_0 = \mu\).
(iv) If \(\{X_\alpha\}\) decomposes \(\mu\) and \(\mu(E) = \infty\), then (b) forces \(\mu(E\cap X_\alpha)>0\) for some \(\alpha\), and \(0<\mu(E\cap X_\alpha)\le\mu(X_\alpha)<\infty\).
(v) Take \(X = \mathbb{R}\), \(\mathcal{A}\) the algebra of all finite sets and their complements, \(\mu(A) := \operatorname{Card}(A\cap\mathbb{Q})\); the extensions are \(\mu_s := \nu + s\,\delta\), \(s\ge0\), where
\begin{equation*} \nu(A) := \operatorname{Card}(A\cap\mathbb{Q}), \qquad \delta(A) := \begin{cases} 0, & A \text{ at most countable},\\ 1, & A \text{ co-countable}.\end{cases} \end{equation*}
Being the restriction to \(\mathcal{A}\) of the counting measure of \(\mathbb{Q}\), \(\mu\) is countably additive; it is semifinite, since \(\mu(A) = \infty\) makes \(A\) contain a rational \(q\) with \(\{q\}\in\mathcal{A}\) of measure \(1\). Here \(\sigma(\mathcal{A})\) consists of the at most countable subsets of \(\mathbb{R}\) and their complements (that family is a \(\sigma\)-algebra containing all finite sets, and every countable set is a countable union of singletons), so exactly one of the two alternatives in \(\delta\) holds for each \(A \in \sigma(\mathcal{A})\), the irrationals being uncountable.
Both summands are countably additive: for \(\nu\) this is clear, and for \(\delta\) two disjoint co-countable sets cannot exist (the union of their countable complements would be \(\mathbb{R}\)) while a disjoint countable union that is co-countable must have a co-countable term, so exactly one term is co-countable when the union is and none is otherwise. Each \(\mu_s\) extends \(\mu\) (on finite \(A\) both are \(\operatorname{Card}(A\cap\mathbb{Q})\); on cofinite \(A\) the set \(A\cap\mathbb{Q}\) is infinite and both are \(+\infty\)) and is semifinite (\(\mu_s(A) = \infty\) forces \(A\cap\mathbb{Q}\) infinite, as \(s<\infty\)). They are pairwise distinct, since \(\mu_s(\mathbb{R}\setminus\mathbb{Q}) = s\).
Let \(\mu\) be a measure with values in \([0,+\infty]\) defined on a measurable space \((X,\mathcal{A})\). A set \(E\) is called locally measurable if \(E\cap A\in\mathcal{A}\) for every \(A\in\mathcal{A}\) with \(\mu(A) <\infty\). The measure \(\mu\) is called saturated if every locally measurable set belongs to \(\mathcal{A}\).
(i) Let \(X = \mathbb{R}\), \(\mathcal{A} = \{\mathbb{R},\emptyset\}\), \(\mu(\mathbb{R}) = +\infty\), \(\mu(\emptyset) = 0\). Show that \(\mu\) is a complete measure with values in \([0,+\infty]\) that is not saturated.
(ii) Show that every \(\sigma\)-finite measure is saturated.
(iii) Show that locally measurable sets form a \(\sigma\)-algebra.
(iv) Show that every measure with values in \([0,+\infty]\) can be extended to a saturated measure on the \(\sigma\)-algebra \(\mathcal{L}\) of all locally measurable sets by the formula \(\overline{\mu}(E) = \mu(E)\) if \(E\in\mathcal{A}\), \(\overline{\mu}(E) = +\infty\) if \(E\notin\mathcal{A}\).
(v) Construct an example showing that \(\overline{\mu}\) may not be a unique saturated extension of \(\mu\) to the \(\sigma\)-algebra \(\mathcal{L}\).
(i) The only \(\mu\)-null set in \(\mathcal{A}\) is \(\emptyset\), whose only subset lies in \(\mathcal{A}\), so \(\mu\) is complete; and \(\emptyset\) is also the only set of finite measure, so \(E\cap A = \emptyset\in\mathcal{A}\) for every \(E\subset\mathbb{R}\) and every \(A\in\mathcal{A}\) of finite measure, i.e. \(\mathcal{L} = 2^{\mathbb{R}}\). Since \(\{0\}\notin\mathcal{A}\), the measure \(\mu\) is not saturated.
(ii) With \(X = \bigcup_{n=1}^\infty X_n\), \(X_n\in\mathcal{A}\), \(\mu(X_n)<\infty\), a locally measurable \(E\) has \(E\cap X_n\in\mathcal{A}\) for every \(n\), whence
\begin{equation*} E = \bigcup_{n=1}^\infty (E\cap X_n)\in\mathcal{A}. \end{equation*}
(iii) Fix \(A\in\mathcal{A}\) with \(\mu(A)<\infty\). Then \(X\cap A = A\in\mathcal{A}\), so \(X\in\mathcal{L}\); for \(E\in\mathcal{L}\),
\begin{equation*} (X\setminus E)\cap A = A\setminus(A\cap E)\in\mathcal{A}, \end{equation*}
so \(X\setminus E\in\mathcal{L}\); and \(\big(\bigcup_n E_n\big)\cap A = \bigcup_n (E_n\cap A)\in\mathcal{A}\) for \(E_n\in\mathcal{L}\). Thus \(\mathcal{L}\) is a \(\sigma\)-algebra containing \(\mathcal{A}\).
(iv) Countable additivity of \(\overline{\mu}\): let \(E = \bigcup_n E_n\) with pairwise disjoint \(E_n\in\mathcal{L}\). If all \(E_n\in\mathcal{A}\), then \(E\in\mathcal{A}\) and
\begin{equation*} \overline{\mu}(E) = \mu(E) = \sum_n \mu(E_n) = \sum_n\overline{\mu}(E_n). \end{equation*}
Otherwise some \(E_{n_0}\notin\mathcal{A}\) and the right-hand side is \(+\infty\); were \(\overline{\mu}(E)<\infty\), then \(E\in\mathcal{A}\) with \(\mu(E)<\infty\), so local measurability would give \(E_{n_0} = E_{n_0}\cap E\in\mathcal{A}\), a contradiction.
Saturation: the sets \(L\in\mathcal{L}\) with \(\overline{\mu}(L)<\infty\) are exactly the \(A\in\mathcal{A}\) with \(\mu(A)<\infty\), so \(E\cap A\in\mathcal{L}\) for all such \(A\) gives \(E\cap A = (E\cap A)\cap A\in\mathcal{A}\), i.e. \(E\in\mathcal{L}\).
(v) In the space of (i) we have \(\mathcal{L} = 2^{\mathbb{R}}\) and \(\overline{\mu}(E) = +\infty\) for every nonempty \(E\). The measure \(\mu_0\) of Exercise 1.12.131(i) on \(2^{\mathbb{R}}\), zero on at most countable sets and \(+\infty\) otherwise, also extends \(\mu\) and is saturated (its domain is already all of \(2^{\mathbb{R}}\)), yet \(\mu_0(\{0\}) = 0 \neq +\infty = \overline{\mu}(\{0\})\).
Let \((X,\mathcal{A},\mu)\) be a measure space, where \(\mu\) takes values in \([0,+\infty]\). The measure \(\mu\) is called Maharam (or localizable) if \(\mu\) is semifinite and each collection \(\mathcal{M}\subset\mathcal{A}\) has the essential supremum in the following sense: there exists a set \(E\in\mathcal{A}\) such that all sets \(M\setminus E\), where \(M\in\mathcal{M}\), have measure zero and if \(E^{\prime}\in\mathcal{A}\) is another set with such a property, then \(E\setminus E^{\prime}\) is a measure zero set.
(i) Prove that every decomposable measure is Maharam.
(ii) Give an example of a complete Maharam measure that is not decomposable.
(i) Let \(\{X_\alpha\}_{\alpha\in\Lambda}\) decompose \((X,\mathcal{A},\mu)\) and let \(\mathcal{M}\subset\mathcal{A}\); the essential supremum is \(E := X\setminus\Psi\), where \(\Psi := \bigcup_{\alpha}(F_\alpha\cap X_\alpha)\) is built as follows, semifiniteness being Exercise 1.12.132(iv). Let \(\mathcal{F}\) be the family of \(F\in\mathcal{A}\) with \(\mu(F\cap M) = 0\) for all \(M\in\mathcal{M}\): it contains \(\emptyset\), is closed under measurable subsets, and admits countable unions, since \(\mu\big(\bigcup_n F_n\cap M\big)\le\sum_n\mu(F_n\cap M) = 0\). Put
\begin{equation*} c_\alpha := \sup\{\mu(F\cap X_\alpha)\colon F\in\mathcal{F}\}\le\mu(X_\alpha)<\infty, \end{equation*}
pick \(F_{\alpha,n}\in\mathcal{F}\) with \(\mu(F_{\alpha,n}\cap X_\alpha)\to c_\alpha\), and set \(F_\alpha := \bigcup_n F_{\alpha,n}\in\mathcal{F}\), so that \(\mu(F_\alpha\cap X_\alpha) = c_\alpha\) (monotonicity gives \(\ge\), the definition of \(c_\alpha\) gives \(\le\)). Disjointness of the \(X_\alpha\) makes \(\Psi\cap X_\beta = F_\beta\cap X_\beta\in\mathcal{A}\), so \(\Psi\in\mathcal{A}\) by (a) and \(E\in\mathcal{A}\).
For \(M\in\mathcal{M}\), property (b) together with \(\mu(M\cap F_\alpha) = 0\) gives
\begin{equation*} \mu(M\setminus E) = \mu(M\cap\Psi) = \sum_\alpha\mu(M\cap F_\alpha\cap X_\alpha) = 0, \end{equation*}
so in particular \(\Psi\in\mathcal{F}\). If \(E^{\prime}\in\mathcal{A}\) also has \(\mu(M\setminus E^{\prime}) = 0\) for all \(M\in\mathcal{M}\), then \((X\setminus E^{\prime})\cap M = M\setminus E^{\prime}\) is null, so \(X\setminus E^{\prime}\in\mathcal{F}\) and \(\Psi^{\prime} := \Psi\cup(X\setminus E^{\prime})\in\mathcal{F}\); now \(\mu(\Psi^{\prime}\cap X_\alpha)\le c_\alpha\) by the definition of \(c_\alpha\) and \(\ge\mu(\Psi\cap X_\alpha) = c_\alpha\) by \(\Psi\subset\Psi^{\prime}\), all values finite, whence
\begin{equation*} \mu\big((\Psi^{\prime}\setminus\Psi)\cap X_\alpha\big) = c_\alpha - c_\alpha = 0 \quad\text{for every } \alpha \end{equation*}
and \(\mu(\Psi^{\prime}\setminus\Psi) = 0\) by (b). Since \(E\setminus E^{\prime} = (X\setminus E^{\prime})\setminus\Psi\subset\Psi^{\prime}\setminus\Psi\), we get \(\mu(E\setminus E^{\prime}) = 0\).
(ii) Let \(I\) be uncountable, \(Z := I\times\{0,1\}\), write \(B_i := \{t\in\{0,1\}\colon (i,t)\in B\}\) and call \(i\) exceptional for \(B\) when \(B_i\) is a one-point set; put (Fremlin [327, \(\S\)216])
\begin{equation*} \mathcal{B} := \{B\subset Z\colon B \text{ has at most countably many exceptional indices}\}, \end{equation*}
with \(\nu(B) := \operatorname{card}\bigl(B\cap(I\times\{0\})\bigr)\), the counting measure of the lower slice. Then \((Z,\mathcal{B},\nu)\) is complete and Maharam but not decomposable.
\(\mathcal{B}\) is a \(\sigma\)-algebra: \((Z\setminus B)_i = \{0,1\}\setminus B_i\) is a one-point set exactly when \(B_i\) is, so \(B\) and \(Z\setminus B\) have the same exceptional indices; and if no \((B^n)_i\) is a one-point set then neither is \(B_i = \bigcup_n(B^n)_i\), so the exceptional indices of \(\bigcup_n B^n\) lie in the countable union of those of the \(B^n\). Every at most countable \(B\subset Z\) thus lies in \(\mathcal{B}\). Also \(\nu\) is a measure, since \(B\mapsto B\cap(I\times\{0\})\) carries disjoint countable unions to disjoint countable unions.
Every \(B\in\mathcal{B}\) with \(\nu(B)<\infty\) is at most countable: off the countable set \(C\) of exceptional indices, \(B_i\ne\emptyset\) forces \(B_i = \{0,1\}\) and hence \((i,0)\in B\), so at most \(\nu(B)\) such indices occur and \(B\) is covered by \(C\times\{0,1\}\) together with finitely many two-point sections.
Completeness: \(\nu(B) = 0\) means \(B\subset I\times\{1\}\), so every nonempty section of \(B\in\mathcal{B}\) is exceptional and \(B\) is at most countable; then every \(B^{\prime}\subset B\) lies in \(\mathcal{B}\) with \(\nu(B^{\prime}) = 0\). Semifiniteness: \(\nu(B) = \infty\) yields \((i,0)\in B\) with \(\{(i,0)\}\in\mathcal{B}\) of measure \(1\).
Maharam: for \(\mathcal{M}\subset\mathcal{B}\) put \(S := \{i\in I\colon (i,0)\in M \text{ for some } M\in\mathcal{M}\}\) and \(E := S\times\{0,1\}\), which lies in \(\mathcal{B}\) as all its sections are \(\{0,1\}\) or \(\emptyset\). Every \((i,0)\in M\) has \(i\in S\), so \((M\setminus E)\cap(I\times\{0\}) = \emptyset\) and \(\nu(M\setminus E) = 0\). If \(E^{\prime}\in\mathcal{B}\) also satisfies \(\nu(M\setminus E^{\prime}) = 0\) for all \(M\in\mathcal{M}\), then each \(i\in S\), witnessed by \((i,0)\in M\), has \((i,0)\in E^{\prime}\); so \(E^{\prime}\supset S\times\{0\}\) and \((E\setminus E^{\prime})\cap(I\times\{0\}) = \emptyset\), i.e. \(\nu(E\setminus E^{\prime}) = 0\).
Not decomposable: every index is exceptional for \(E^* := I\times\{0\}\), so \(E^*\notin\mathcal{B}\); but each piece \(Z_\alpha\) of a decomposition has finite measure, hence is at most countable, so \(E^*\cap Z_\alpha\in\mathcal{B}\) for every \(\alpha\) and property (a) would force \(E^*\in\mathcal{B}\).
A measure with values in \([0,+\infty]\) is called locally determined if it is semifinite and saturated. Let \(\mu\) be a measure with values in \([0,+\infty]\) defined on a measurable space \((X,\mathcal{A})\). Let \(\mathcal{L}_\mu\) be the \(\sigma\)-algebra of locally \(\mathcal{A}_\mu\)-measurable sets, i.e., all sets \(L\) such that \(L\cap A\in\mathcal{A}_\mu\) for all \(A\in\mathcal{A}_\mu\) with \(\mu(A)<\infty\). Let
\begin{equation*} \widetilde{\mu}(L) = \sup\{\mu(L\cap A)\colon A\in\mathcal{A}_\mu,\ \mu(A)<\infty\},\qquad L\in\mathcal{L}_\mu . \end{equation*}
(i) Show that the measure \(\widetilde{\mu}\) is locally determined and complete and that one has \(\widetilde{\mu}(A) = \mu(A)\) whenever \(A\in\mathcal{A}_\mu\) and \(\mu(A)<\infty\).
(ii) Show that if \(\mu\) is decomposable, then so is \(\widetilde{\mu}\) and in this case \(\widetilde{\mu}\) coincides with the completion of \(\mu\).
(iii) Show that if \(\mu\) is Maharam, then so is \(\widetilde{\mu}\).
(iv) Show that the measure \(\mu\) is complete and locally determined precisely when one has \(\mu = \widetilde{\mu}\).
(i) By Exercise 1.12.133(iii) applied to \((X,\mathcal{A}_\mu,\mu)\), the family \(\mathcal{L}_\mu\) is a \(\sigma\)-algebra containing \(\mathcal{A}_\mu\); write \(\mathcal{A}_\mu^f := \{A\in\mathcal{A}_\mu\colon\mu(A)<\infty\}\). Clearly \(\widetilde{\mu}(\emptyset) = 0\) and \(\widetilde{\mu}\) is monotone, while for \(A\in\mathcal{A}_\mu\)
\begin{equation*} \widetilde{\mu}(A) = \sup\{\mu( C)\colon C\in\mathcal{A}_\mu, \ C\subset A,\ \mu( C)<\infty\} = \mu_0(A), \end{equation*}
the semifinite version of \(\mu\) from Exercise 1.12.132(iii) (each \(A\cap B\) is such a \(C\), and each such \(C\) equals \(A\cap C\)); in particular \(\widetilde{\mu}(A) = \mu(A)\) whenever \(\mu(A)<\infty\).
Countable additivity. Let \(L = \bigcup_n L_n\) with pairwise disjoint \(L_n\in\mathcal{L}_\mu\). For \(A\in\mathcal{A}_\mu^f\),
\begin{equation*} \mu(L\cap A) = \sum_n\mu(L_n\cap A)\le\sum_n\widetilde{\mu}(L_n), \end{equation*}
so \(\widetilde{\mu}(L)\le\sum_n\widetilde{\mu}(L_n)\); conversely, for \(A_1,\dots,A_N\in\mathcal{A}_\mu^f\) and \(A := A_1\cup\dots\cup A_N\in\mathcal{A}_\mu^f\), monotonicity gives
\begin{equation*} \widetilde{\mu}(L)\ge\mu(L\cap A)\ge\sum_{n\le N}\mu(L_n\cap A) \ge\sum_{n\le N}\mu(L_n\cap A_n), \end{equation*}
and suprema over \(A_1,\dots,A_N\) followed by \(N\to\infty\) give \(\widetilde{\mu}(L)\ge\sum_n\widetilde{\mu}(L_n)\).
Semifinite: \(\widetilde{\mu}(L) = \infty\) supplies \(A\in\mathcal{A}_\mu^f\) with \(\mu(L\cap A)>0\), and \(L\cap A\subset L\) lies in \(\mathcal{A}_\mu^f\) with \(\widetilde{\mu}(L\cap A) = \mu(L\cap A)\in(0,\infty)\). Saturated: since \(\widetilde{\mu} = \mu<\infty\) on \(\mathcal{A}_\mu^f\subset\mathcal{L}_\mu\), a set \(M\) with \(M\cap A\in\mathcal{L}_\mu\) for all \(A\in\mathcal{A}_\mu^f\) has \(M\cap A = (M\cap A)\cap A\in\mathcal{A}_\mu\), i.e. \(M\in\mathcal{L}_\mu\). So \(\widetilde{\mu}\) is locally determined. Complete: \(\widetilde{\mu}(L) = 0\) and \(Z\subset L\) give \(\mu(L\cap A) = 0\), so completeness of \(\mu\) on \(\mathcal{A}_\mu\) puts \(Z\cap A\in\mathcal{A}_\mu\) with measure \(0\) for every \(A\in\mathcal{A}_\mu^f\), whence \(Z\in\mathcal{L}_\mu\) and \(\widetilde{\mu}(Z) = 0\).
(ii) Let \(\{X_\alpha\}_{\alpha\in\Lambda}\) decompose \((X,\mathcal{A},\mu)\). Then \(\mu\) is semifinite (Exercise 1.12.132(iv)) and so is its completion: \(E\in\mathcal{A}_\mu\) with \(\mu(E) = \infty\) contains \(A\in\mathcal{A}\) with \(\mu(A) = \infty\), which contains a set of finite positive measure. So \(\widetilde{\mu} = \mu_0 = \mu\) on \(\mathcal{A}_\mu\) by Exercise 1.12.132(iii).
Next, \(E\cap X_\alpha\in\mathcal{A}_\mu\) for all \(\alpha\) forces \(E\in\mathcal{A}_\mu\): choose \(B_\alpha\subset E\cap X_\alpha\subset C_\alpha\subset X_\alpha\) in \(\mathcal{A}\) with \(N_\alpha := C_\alpha\setminus B_\alpha\) null (replace \(C_\alpha\) by \(C_\alpha\cap X_\alpha\)), and put \(B := \bigcup_\alpha B_\alpha\), \(N := \bigcup_\alpha N_\alpha\). Disjointness of the \(X_\alpha\) gives \(B\cap X_\beta = B_\beta\) and \(N\cap X_\beta = N_\beta\), so \(B, N\in\mathcal{A}\) by (a) and \(\mu(N) = \sum_\alpha\mu(N_\alpha) = 0\) by (b), while \(B\subset E\subset B\cup N\).
Since the \(X_\alpha\) have finite measure, every \(L\in\mathcal{L}_\mu\) satisfies \(L\cap X_\alpha\in\mathcal{A}_\mu\), so \(\mathcal{L}_\mu = \mathcal{A}_\mu\) and \(\widetilde{\mu}\) is exactly the completion of \(\mu\). Finally \(\{X_\alpha\}\) decomposes \(\widetilde{\mu}\): (a) is the previous paragraph, and for (b) take \(B\in\mathcal{A}\) and \(\mu\)-null \(N\in\mathcal{A}\) with \(B\subset E\subset B\cup N\), so that \(\widetilde{\mu}(E) = \mu(B)\) and \(\mu(E\cap X_\alpha) = \mu(B\cap X_\alpha)\), giving
\begin{equation*} \widetilde{\mu}(E) = \mu(B) = \sum_\alpha\mu(B\cap X_\alpha) = \sum_\alpha\widetilde{\mu}(E\cap X_\alpha). \end{equation*}
(iii) First, the completion of a Maharam \(\mu\) is Maharam. Given \(\mathcal{M}\subset\mathcal{A}_\mu\), pick \(B_M\in\mathcal{A}\) with \(B_M\subset M\), \(\mu(M\setminus B_M) = 0\), and let \(E\in\mathcal{A}\) be an essential supremum of \(\{B_M\}\) in \((X,\mathcal{A},\mu)\); then
\begin{equation*} \mu(M\setminus E)\le\mu(B_M\setminus E) + \mu(M\setminus B_M) = 0 , \end{equation*}
and if \(E^{\prime}\in\mathcal{A}_\mu\) also has \(\mu(M\setminus E^{\prime}) = 0\) for all \(M\), then \(B^{\prime}\in\mathcal{A}\) with \(B^{\prime}\subset E^{\prime}\), \(\mu(E^{\prime}\setminus B^{\prime}) = 0\) satisfies \(\mu(B_M\setminus B^{\prime}) = 0\), so \(\mu(E\setminus E^{\prime})\le\mu(E\setminus B^{\prime}) = 0\). Semifiniteness passes to the completion as in (ii).
Now let \(\mathcal{M}\subset\mathcal{L}_\mu\), put \(\mathcal{M}_0 := \{M\cap A\colon M\in\mathcal{M},\ A\in\mathcal{A}_\mu^f\}\subset\mathcal{A}_\mu\), and let \(E_0\in\mathcal{A}_\mu\) be an essential supremum of \(\mathcal{M}_0\) for \(\mu\). For \(M\in\mathcal{M}\) and \(A\in\mathcal{A}_\mu^f\) the set \((M\setminus E_0)\cap A = (M\cap A)\setminus E_0\) is \(\mu\)-null, so \(\widetilde{\mu}(M\setminus E_0) = 0\). Let \(E^{\prime}\in\mathcal{L}_\mu\) satisfy \(\widetilde{\mu}(M\setminus E^{\prime}) = 0\) for all \(M\in\mathcal{M}\), fix \(A\in\mathcal{A}_\mu^f\), and put \(F := (E^{\prime}\cap A)\cup(X\setminus A)\in\mathcal{A}_\mu\). For \(M_0 = M\cap B\in\mathcal{M}_0\),
\begin{equation*} M_0\setminus F = (M\cap B\cap A)\setminus E^{\prime} \subset (M\setminus E^{\prime})\cap B, \end{equation*}
a \(\mu\)-null set since \(B\in\mathcal{A}_\mu^f\), so \(F\) bounds \(\mathcal{M}_0\) and \(\mu(E_0\setminus F) = 0\). As \(F\supset X\setminus A\),
\begin{equation*} (E_0\setminus E^{\prime})\cap A = (E_0\cap A)\setminus F\subset E_0\setminus F, \end{equation*}
whence \(\mu\big((E_0\setminus E^{\prime})\cap A\big) = 0\) for every \(A\), i.e. \(\widetilde{\mu}(E_0\setminus E^{\prime}) = 0\). With the semifiniteness from (i), \(\widetilde{\mu}\) is Maharam.
(iv) If \(\mu = \widetilde{\mu}\), then \(\mu\) is complete and locally determined by (i). Conversely, completeness gives \(\mathcal{A}_\mu = \mathcal{A}\) and saturation then gives \(\mathcal{L}_\mu = \mathcal{A}\), while semifiniteness with Exercise 1.12.132(iii) gives \(\widetilde{\mu} = \mu_0 = \mu\) on \(\mathcal{A}\).
Let \((X,\mathcal{A})\) be a measurable space and let a measure \(\mu\) on \(\mathcal{A}\) with values in \([0,+\infty]\) be complete and locally determined. Suppose that there exists a family \(\mathcal{D}\) of pairwise disjoint sets of finite measure in \(\mathcal{A}\) such that if \(E\in\mathcal{A}\) and \(\mu(E\cap D) = 0\) for all \(D\in\mathcal{D}\), then \(\mu(E) = 0\). Prove that the measure \(\mu\) is decomposable.
The decomposition is \(\mathcal{D}\cup\{Z\}\) with \(Z := X\setminus\bigcup_{\alpha\in\Lambda}D_\alpha\); only completeness and saturation of \(\mu\) are used. If \(\mathcal{D} = \emptyset\), the hypothesis forces \(\mu\equiv0\) and \(\{X\}\) is a decomposition, so let \(\mathcal{D} = \{D_\alpha\}_{\alpha\in\Lambda}\ne\emptyset\).
Splitting the sets of finite measure: for \(A\in\mathcal{A}\) with \(\mu(A)<\infty\), the sets \(A\cap D_\alpha\) are disjoint inside \(A\), so \(\sum_{\alpha\in F}\mu(A\cap D_\alpha) = \mu\big(A\cap\bigcup_{\alpha\in F}D_\alpha\big)\le\mu(A)\) for finite \(F\subset\Lambda\), whence \(\Lambda_A := \{\alpha\colon\mu(A\cap D_\alpha)>0\}\) is at most countable and
\begin{equation*} A^{\prime} := A\cap\bigcup_{\alpha\in\Lambda_A}D_\alpha\in\mathcal{A},\qquad A^{\prime\prime} := A\setminus A^{\prime}\in\mathcal{A} \end{equation*}
(a countable union). Every \(\alpha\) has \(\mu(A^{\prime\prime}\cap D_\alpha) = 0\) (empty for \(\alpha\in\Lambda_A\), inside the null set \(A\cap D_\alpha\) otherwise), so the hypothesis on \(\mathcal{D}\) gives \(\mu(A^{\prime\prime}) = 0\), and completeness makes every subset of \(A^{\prime\prime}\) measurable and null.
Hence \(Z\cap A\subset A^{\prime\prime}\) is measurable and null for every such \(A\), so \(Z\) is locally measurable, \(Z\in\mathcal{A}\) by saturation, and \(\mu(Z) = 0\) by the hypothesis, \(Z\cap D_\alpha\) being empty. So \(\mathcal{D}\cup\{Z\}\) is a partition of \(X\) into pairwise disjoint members of \(\mathcal{A}\) of finite measure.
(a) Let \(E\cap D_\alpha\in\mathcal{A}\) for all \(\alpha\), the condition \(E\cap Z\in\mathcal{A}\) being automatic by completeness. For \(A\in\mathcal{A}\) with \(\mu(A)<\infty\),
\begin{equation*} E\cap A^{\prime} = \bigcup_{\alpha\in\Lambda_A}\big(E\cap D_\alpha\cap A\big)\in\mathcal{A} \end{equation*}
is a countable union, while \(E\cap A^{\prime\prime}\subset A^{\prime\prime}\) is measurable by completeness; so \(E\cap A\in\mathcal{A}\), \(E\) is locally measurable, and \(E\in\mathcal{A}\) by saturation.
(b) Let \(E\in\mathcal{A}\) and \(s := \sum_\alpha\mu(E\cap D_\alpha)\), which is also the sum over the whole partition since \(\mu(E\cap Z) = 0\). As above \(\sum_{\alpha\in F}\mu(E\cap D_\alpha)\le\mu(E)\) for finite \(F\), so \(s\le\mu(E)\), settling the case \(s = \infty\). If \(s<\infty\), then \(\Lambda_E := \{\alpha\colon\mu(E\cap D_\alpha)>0\}\) is at most countable,
\begin{equation*} E_1 := E\cap\bigcup_{\alpha\in\Lambda_E}D_\alpha\in\mathcal{A},\qquad \mu(E_1) = \sum_{\alpha\in\Lambda_E}\mu(E\cap D_\alpha) = s , \end{equation*}
and \(E_2 := E\setminus E_1\) has \(\mu(E_2\cap D_\alpha) = 0\) for every \(\alpha\) (empty for \(\alpha\in\Lambda_E\), null otherwise), hence \(\mu(E_2) = 0\). Therefore
\begin{equation*} \mu(E) = \mu(E_1)+\mu(E_2) = s = \sum_\alpha\mu(E\cap D_\alpha)+\mu(E\cap Z). \end{equation*}
Let \(X\) be a set of cardinality of the continuum and let \(Y\) be a set of cardinality greater than that of the continuum. For every \(E\subset X\times Y\), the sets \(\{(a,y)\in E\}\) with fixed \(a\in X\) will be called vertical sections of \(E\), and the sets \(\{(x,b)\in E\}\) with fixed \(b\in Y\) will be called horizontal sections of \(E\). Denote by \(\mathcal{A}\) the class of all sets \(A\subset X\times Y\) such that all their horizontal and vertical sections are either at most countable or have at most countable complements in the corresponding sections of \(X\times Y\). Let \(\gamma(A)\) be the number of those horizontal sections of the complement of \(A\) that are at most countable. Similarly, by means of vertical sections we define the function \(v(A)\). Let \(\mu(A) = \gamma(A)+v(A)\).
(i) Prove that \(\mathcal{A}\) is a \(\sigma\)-algebra and that \(\gamma\), \(v\), and \(\mu\) are countably additive measures with values in \([0,+\infty]\).
(ii) Prove that \(\mu\) is semifinite in the sense of Exercise 1.12.132.
(iii) Prove that \(\mu\) is not decomposable in the sense of Exercise 1.12.131.
Write \(A_a := \{y\in Y\colon (a,y)\in A\}\) and \(A^b := \{x\in X\colon (x,b)\in A\}\) for the vertical and horizontal sections, and
\begin{equation*} H(A) := \{b\in Y\colon A^b\ \text{co-countable in}\ X\},\qquad V(A) := \{a\in X\colon A_a\ \text{co-countable in}\ Y\}, \end{equation*}
so that \(\gamma(A) = \operatorname{card} H(A)\) and \(v(A) = \operatorname{card} V(A)\) as elements of \([0,+\infty]\). Since \(X\) and \(Y\) are uncountable, no subset of either is both at most countable and co-countable.
(i) \(\mathcal{A}\) is a \(\sigma\)-algebra: all sections of \(X\times Y\) are co-countable; the sections of a complement are the complements of the sections, and “at most countable or co-countable” is stable under complementation; and \(A^b = \bigcup_n A_n^b\) is at most countable when all \(A_n^b\) are and co-countable when some \(A_{n_0}^b\) is, likewise for vertical sections.
For countable additivity of \(\gamma\), let \(A = \bigcup_n A_n\) with disjoint \(A_n\in\mathcal{A}\) and fix \(b\). Among the disjoint sets \(A_n^b\) at most one is co-countable (two disjoint co-countable subsets of the uncountable \(X\) would have at most countable complements covering \(X\)), and if \(A^b\) is co-countable then at least one is, a countable union of at most countable sets being at most countable. So \(H(A)\) is the disjoint union of the \(H(A_n)\), and cardinality truncated at \(+\infty\) is countably additive on disjoint countable unions; also \(\gamma(\emptyset) = 0\). By symmetry \(v\), hence \(\mu = \gamma + v\), is a measure.
(ii) Let \(\mu(A) = \infty\), say \(\gamma(A) = \infty\) (the case \(v(A) = \infty\) is symmetric), pick \(b_1\in H(A)\) and put \(B := A\cap(X\times\{b_1\})\). Then \(B^{b_1} = A^{b_1}\) is co-countable, \(B^b = \emptyset\) for \(b\ne b_1\), and every \(B_a\subset\{b_1\}\) is at most countable, so \(B\subset A\) lies in \(\mathcal{A}\) with
\begin{equation*} \mu(B) = \gamma(B)+v(B) = 1+0 = 1 . \end{equation*}
(iii) Suppose \(\{E_\alpha\}_{\alpha\in\Lambda}\) were a decomposition; only property (b) is used. Put \(L_b := X\times\{b\}\) and \(C_a := \{a\}\times Y\), both in \(\mathcal{A}\) of measure \(1\) as in (ii).
Fix \(b\in Y\). Each \(L_b\cap E_\alpha\in\mathcal{A}\) has all vertical sections inside \(\{b\}\), so \(v(L_b\cap E_\alpha) = 0\), while \(\gamma(L_b\cap E_\alpha) = 1\) exactly when \(E_\alpha^b\) is co-countable; property (b) at \(L_b\) reads
\begin{equation*} 1 = \mu(L_b) = \sum_\alpha\mu(L_b\cap E_\alpha), \end{equation*}
so exactly one \(\alpha\) has \(b\in H(E_\alpha)\) and the sets \(H(E_\alpha)\) partition \(Y\). Each is finite, since \(\operatorname{card} H(E_\alpha) = \gamma(E_\alpha)\le\mu(E_\alpha)<\infty\), and \(H(E_\alpha) = \emptyset\) for \(\alpha\notin S := \{\alpha\colon E_\alpha\ne\emptyset\}\), so
\begin{equation*} \operatorname{card} Y = \operatorname{card}\bigcup_{\alpha\in S}H(E_\alpha) \le \operatorname{card} S\cdot\aleph_0 = \max(\operatorname{card} S,\aleph_0), \end{equation*}
whence \(\operatorname{card} S\ge\operatorname{card} Y>\mathfrak{c}\).
Fix instead \(a\in X\). Property (b) at \(C_a\), where now \(\gamma(C_a\cap E_\alpha) = 0\) as all horizontal sections lie in \(\{a\}\), gives a unique \(\beta\) with \((E_\beta)_a\) co-countable in \(Y\); for \(\alpha\ne\beta\) disjointness gives \((E_\alpha)_a\subset Y\setminus(E_\beta)_a\), an at most countable set, so at most countably many of the disjoint sets \((E_\alpha)_a\) are nonempty. Every nonempty \(E_\alpha\) has a nonempty section at some \(a\), so
\begin{equation*} \operatorname{card} S = \operatorname{card}\bigcup_{a\in X}\{\alpha\colon (E_\alpha)_a\ne\emptyset\} \le\operatorname{card} X\cdot\aleph_0 = \mathfrak{c}, \end{equation*}
contradicting \(\operatorname{card} S>\mathfrak{c}\).
Exercises 1.12.138–1.12.144
Let \(X = [0,1]\times\{0,1\}\) and let \(\mathcal{A}\) be the class of all sets \(E\subset X\) such that the sections \(E_x := \{y\colon (x,y)\in E\}\) are either empty or coincide with \(\{0,1\}\) for all \(x\), excepting possibly the points of an at most countable set. Show that \(\mathcal{A}\) is a \(\sigma\)-algebra and the function \(\mu\) that to every set \(E\) assigns the cardinality of the intersection of \(E\) with the first coordinate axis, is a complete and semifinite countably additive measure with values in \([0,+\infty]\), but the measure generated by the outer measure \(\mu^*\) is not semifinite.
Throughout write \(L := [0,1]\times\{0\}\) for the first coordinate axis and, for \(A\subset X\),
\begin{equation*} \pi(A) := \{x\in[0,1]\colon A_x\neq\emptyset\},\qquad D(A) := \{x\in[0,1]\colon A_x \text{ is a singleton}\}. \end{equation*}
Since each section \(A_x\) is a subset of \(\{0,1\}\), the condition defining \(\mathcal{A}\) reads \(A\in\mathcal{A}\iff D(A)\) is at most countable, and \(\mu(E) = \operatorname{card}(E\cap L)\in[0,+\infty]\).
\(\mathcal{A}\) is a \(\sigma\)-algebra: \(D(X) = \emptyset\); a subset of \(\{0,1\}\) is a singleton exactly when its complement is, so \(D(X\setminus E) = D(E)\); and
\begin{equation*} D\Big(\bigcup_{n=1}^\infty E_n\Big)\subset \bigcup_{n=1}^\infty D(E_n), \end{equation*}
since off \(\bigcup_n D(E_n)\) every \((E_n)_x\) is \(\emptyset\) or \(\{0,1\}\), hence so is their union. Note \(D(A)\subset\pi(A)\), so every \(A\) with at most countable projection lies in \(\mathcal{A}\) — singletons in particular.
\(\mu\) is countably additive, the sets \(E_n\cap L\) being disjoint with union \(E\cap L\) and truncated cardinality countably additive on such unions. It is complete: \(\mu(E) = 0\) means \(E\cap L = \emptyset\), i.e. \(E_x\subset\{1\}\) for all \(x\), so no section is \(\{0,1\}\) and \(\pi(E)\subset D(E)\) is at most countable; any \(F\subset E\) then has \(\pi(F)\) at most countable, hence \(F\in\mathcal{A}\) and \(\mu(F) = 0\). It is semifinite: \(\mu(E) = \infty\) gives \(p\in E\cap L\) with \(\{p\}\subset E\) in \(\mathcal{A}\) of measure \(1\).
The outer measure is
\begin{equation*} \mu^*(A) = \begin{cases}\operatorname{card}(A\cap L), & \pi(A) \text{ at most countable},\\ +\infty, & \pi(A) \text{ uncountable}.\end{cases} \end{equation*}
Indeed, at most countable \(\pi(A)\) puts \(A\in\mathcal{A}\); and if \(\pi(A)\) is uncountable and \(A\subset E\in\mathcal{A}\), then for \(x\in\pi(A)\setminus D(E)\) the section \(E_x\) is nonempty and not a singleton, so \(E_x = \{0,1\}\) and \((x,0)\in E\cap L\), while \(\pi(A)\setminus D(E)\) is uncountable, so \(\mu(E) = \infty\). In particular \(\mu^*(B)\ge\operatorname{card}(B\cap L)\) for every \(B\).
Every \(A\subset X\) is \(\mu^*\)-measurable. If \(\mu^*(T)<\infty\), then \(\pi(T)\), hence \(\pi(T\cap A)\) and \(\pi(T\setminus A)\), are at most countable, and \(T\cap L\) splits disjointly, so
\begin{equation*} \mu^*(T\cap A)+\mu^*(T\setminus A) = \operatorname{card}(T\cap A\cap L) +\operatorname{card}\bigl((T\setminus A)\cap L\bigr) = \operatorname{card}(T\cap L) = \mu^*(T). \end{equation*}
If \(\mu^*(T) = \infty\), then either \(\pi(T) = \pi(T\cap A)\cup\pi(T\setminus A)\) is uncountable, whence one term is \(+\infty\), or \(T\cap L\) is infinite, whence one of \(T\cap A\cap L\), \((T\setminus A)\cap L\) is infinite and the corresponding term is \(+\infty\) by the bound just noted. So \(\mathfrak{M}_{\mu^*} = 2^X\) and the generated measure is \(\nu = \mu^*\).
Finally \(\nu\) is not semifinite: \(A_1 := [0,1]\times\{1\}\) has uncountable projection, so \(\nu(A_1) = \infty\), while every \(B\subset A_1\) has \(B\cap L = \emptyset\) and hence \(\nu(B) = 0\) or \(+\infty\) according as \(\pi(B)\) is at most countable or not.
(Luther) Let \(\mu\) be a measure with values in \([0,+\infty]\) defined on a ring \(\mathcal{R}\), let \(\overline{\mu}\) be the restriction of \(\mu^*\) to the \(\sigma\)-ring \(\mathcal{S}\) generated by \(\mathcal{R}\), and let \(\mathcal{R}_0\) and \(\mathcal{S}_0\) be the subclasses in \(\mathcal{R}\) and \(\mathcal{S}\) consisting of all sets of finite measure. Set
\begin{equation*} \widetilde{\mu}(E) = \limsup\{\overline{\mu}(P\cap E),\ P\in\mathcal{R}_0\},\qquad E\in\mathcal{S}. \end{equation*}
(i) Prove that the following conditions are equivalent: (a) \(\mu\) is semifinite, (b) \(\widetilde{\mu}\) is an extension of \(\mu\) to \(\mathcal{S}\), (c) any measure \(\nu\) on \(\mathcal{S}\) with values in \([0,+\infty]\) that agrees with \(\mu\) on \(\mathcal{R}_0\) coincides with \(\mu\) on \(\mathcal{R}\).
(ii) Show that any measure \(\nu\) on \(\mathcal{S}\) with values in \([0,+\infty]\) that agrees with \(\mu\) on \(\mathcal{R}_0\), coincides with \(\widetilde{\mu}\) and \(\overline{\mu}\) on \(\mathcal{S}_0\), and that \(\widetilde{\mu}\le\nu\le\overline{\mu}\) on \(\mathcal{S}\).
(iii) Prove that the following conditions are equivalent: (a) \(\overline{\mu}\) is semifinite, (b) \(\mu\) is semifinite and has a unique extension to \(\mathcal{S}\), (c) \(\widetilde{\mu}=\overline{\mu}\), (d) for all \(E\in\mathcal{S}\) one has \(\overline{\mu}(E)=\limsup\{\overline{\mu}(P\cap E),\ P\in\mathcal{R}_0\}\).
(iv) Prove that if the measure \(\mu\) is \(\sigma\)-finite, then \(\mu\) has a unique extension to \(\mathcal{S}\).
(v) Give an example showing that in (iv) it is not sufficient to require the existence of some \(\sigma\)-finite extension of \(\mu\).
Throughout \(\mu^*\) is the outer measure generated by \(\mu\),
\begin{equation*} \mu^*(A)=\inf\Bigl\{\sum_{n=1}^\infty \mu(R_n)\colon R_n\in\mathcal{R}, \ A\subset\bigcup_{n=1}^\infty R_n\Bigr\}, \end{equation*}
with \(\inf\emptyset=+\infty\), and semifiniteness of a measure on a ring means \(\mu( R)=\sup\{\mu(P)\colon P\in\mathcal{R}_0,\ P\subset R\}\) for all \(R\in\mathcal{R}\) — on a general ring this is strictly stronger than the formulation of Exercise 1.12.132, and it is the condition meant here.
On a \(\sigma\)-ring the two agree: if \(\nu(E)=\infty\) while \(s:=\sup\{\nu(F)\colon F\in\mathcal{S},\ F\subset E,\ \nu(F)<\infty\}\) were finite, choose \(F_n\subset E\) in \(\mathcal{S}\) with \(\nu(F_n)\to s\), replace \(F_n\) by \(F_1\cup\dots\cup F_n\) and put \(F=\bigcup_n F_n\in\mathcal{S}\), so \(\nu(F)=s<\infty\) and \(\nu(E\setminus F)=\infty\); any \(G\subset E\setminus F\) with \(0<\nu(G)<\infty\) would give \(\nu(F\cup G)=s+\nu(G)>s\).
Three facts. (0) \(\mathcal{R}_0\) is a ring directed upward by inclusion, since \(\mu(P\cup Q)\le\mu(P)+\mu(Q)<\infty\), and \(P\mapsto\overline{\mu}(P\cap E)\) is nondecreasing, so the upper limit defining \(\widetilde{\mu}\) is a supremum:
\begin{equation*} \widetilde{\mu}(E)=\sup\{\overline{\mu}(P\cap E)\colon P\in\mathcal{R}_0\}, \qquad E\in\mathcal{S}. \end{equation*}
(1) By Theorem 1.11.8, applied to the ring \(\mathcal{R}\) and the countably additive \(\mu\), all sets of \(\mathcal{R}\) are \(\mu^*\)-measurable with \(\mu^*=\mu\) there, while Theorem 1.11.4 makes \(\mathcal{M}_{\mu^*}\supset\mathcal{S}\) a \(\sigma\)-algebra carrying the countably additive \(\mu^*\); so \(\overline{\mu}\) is a measure on \(\mathcal{S}\) extending \(\mu\), and \(\mathcal{R}_0\subset\mathcal{S}_0\). (2) Every \(E\in\mathcal{S}\) lies in a countable union of members of \(\mathcal{R}\), the class of such sets being a \(\sigma\)-ring containing \(\mathcal{R}\).
Lemma 1: \(\widetilde{\mu}\) is a measure on \(\mathcal{S}\), \(\widetilde{\mu}\le\overline{\mu}\), and \(\widetilde{\mu}=\overline{\mu}\) on \(\mathcal{S}_0\). For \(E=\bigcup_n E_n\) with disjoint \(E_n\in\mathcal{S}\) and \(P\in\mathcal{R}_0\) we have \(\overline{\mu}(P\cap E)=\sum_n\overline{\mu}(P\cap E_n)\le\sum_n\widetilde{\mu}(E_n)\), so \(\widetilde{\mu}(E)\le\sum_n\widetilde{\mu}(E_n)\); conversely, for \(c_n<\widetilde{\mu}(E_n)\), \(n\le N\), pick \(P_n\in\mathcal{R}_0\) with \(\overline{\mu}(P_n\cap E_n)>c_n\) and set \(P=\bigcup_{n\le N}P_n\in\mathcal{R}_0\), whence
\begin{equation*} \overline{\mu}(P\cap E)\ge\sum_{n\le N}\overline{\mu}(P\cap E_n) \ge\sum_{n\le N}\overline{\mu}(P_n\cap E_n)>\sum_{n\le N}c_n . \end{equation*}
For \(E\in\mathcal{S}_0\) and \(\varepsilon>0\), choosing \(R_n\in\mathcal{R}\) with \(E\subset\bigcup_n R_n\) and \(\sum_n\mu(R_n)\le\overline{\mu}(E)+\varepsilon<\infty\) puts every \(R_n\), hence \(P_N:=\bigcup_{n\le N}R_n\), in \(\mathcal{R}_0\); the sets \(P_N\cap E\) increase to \(E\), so continuity from below gives \(\widetilde{\mu}(E)\ge\overline{\mu}(E)\). In particular \(\widetilde{\mu}=\overline{\mu}=\mu\) on \(\mathcal{R}_0\).
Lemma 2: \(\widetilde{\mu}( R)=\sup\{\mu(P)\colon P\in\mathcal{R}_0,\ P\subset R\}\) for \(R\in\mathcal{R}\), since every \(P\in\mathcal{R}_0\) has \(P\cap R\in\mathcal{R}_0\) with \(P\cap R\subset R\) and \(\overline{\mu}(P\cap R)=\mu(P\cap R)\), while \(P\subset R\) gives \(P\cap R=P\).
(ii) Let \(\nu\) be a measure on \(\mathcal{S}\) with \(\nu=\mu\) on \(\mathcal{R}_0\). First \(\nu\le\overline{\mu}\): we may assume \(\overline{\mu}(E)<\infty\), and then choosing \(R_n\in\mathcal{R}_0\) as above, countable subadditivity gives \(\nu(E)\le\sum_n\nu(R_n)=\sum_n\mu(R_n)\le\overline{\mu}(E)+\varepsilon\) for every \(\varepsilon>0\). Second \(\widetilde{\mu}\le\nu\): for \(E\in\mathcal{S}\) and \(P\in\mathcal{R}_0\),
\begin{equation*} \nu(P\cap E)+\nu(P\setminus E)=\nu(P)=\mu(P) =\overline{\mu}(P\cap E)+\overline{\mu}(P\setminus E)<\infty , \end{equation*}
so \(\nu(P\setminus E)\le\overline{\mu}(P\setminus E)\) with all terms finite forces \(\nu(P\cap E)\ge\overline{\mu}(P\cap E)\), and \(\nu(E)\ge\overline{\mu}(P\cap E)\) for every \(P\) gives \(\nu\ge\widetilde{\mu}\). On \(\mathcal{S}_0\), Lemma 1 then squeezes \(\nu=\widetilde{\mu}=\overline{\mu}\).
(i) (a)\(\iff\)(b): by Lemma 2 the equality \(\widetilde{\mu}=\mu\) on \(\mathcal{R}\) is literally the semifiniteness of \(\mu\), and \(\widetilde{\mu}\) is a measure on \(\mathcal{S}\) by Lemma 1. (b)\(\Rightarrow\)(c): for \(\nu\) as in (c) and \(R\in\mathcal{R}\), part (ii) gives \(\mu( R)=\widetilde{\mu}( R)\le\nu( R)\le\overline{\mu}( R)=\mu( R)\). (c)\(\Rightarrow\)(b): \(\nu:=\widetilde{\mu}\) agrees with \(\mu\) on \(\mathcal{R}_0\) by Lemma 1, hence on all of \(\mathcal{R}\) by (c).
(iii) (c)\(\iff\)(d) is the definition of \(\widetilde{\mu}\) written out. (c)\(\Rightarrow\)(a): \(\overline{\mu}(E)=\infty\) supplies \(P\in\mathcal{R}_0\) with \(0<\overline{\mu}(P\cap E)\le\mu(P)<\infty\), and \(P\cap E\subset E\) lies in \(\mathcal{S}\). (a)\(\Rightarrow\)(c): by the \(\sigma\)-ring remark, \(\overline{\mu}(E)=\sup\{\overline{\mu}(F)\colon F\in\mathcal{S}_0,\ F\subset E\}\), and each such \(F\) has \(\overline{\mu}(F)=\widetilde{\mu}(F)\le\widetilde{\mu}(E)\) by Lemma 1 and monotonicity. (c)\(\Rightarrow\)(b): \(\widetilde{\mu}=\overline{\mu}\) makes \(\mu\) semifinite by (i), and any extension \(\nu\) agrees with \(\mu\) on \(\mathcal{R}_0\), so (ii) squeezes \(\nu=\overline{\mu}\). (b)\(\Rightarrow\)(c): by (i) both \(\widetilde{\mu}\) and \(\overline{\mu}\) extend \(\mu\), so uniqueness forces equality.
(iv) \(\sigma\)-finiteness makes every \(R\in\mathcal{R}\) a countable union of members of \(\mathcal{R}_0\), so by Fact 2 any \(E\in\mathcal{S}\) satisfies \(E\subset\bigcup_j P_j\) with \(P_j\in\mathcal{R}_0\); then \(Q_N:=\bigcup_{j\le N}P_j\in\mathcal{R}_0\) has \(Q_N\cap E\uparrow E\), so \(\widetilde{\mu}(E)\ge\overline{\mu}(E)\), which with Lemma 1 is (iii)(c), and (iii) gives uniqueness.
(v) Take \(X=\mathbb{R}\), \(\mathcal{R}\) the ring of finite unions of intervals \([a,b)\), and \(\mu(\emptyset)=0\), \(\mu( R)=+\infty\) for \(R\ne\emptyset\), which is countably additive since a disjoint decomposition of a nonempty \(R\) has a nonempty term. Here \(\mathcal{R}_0=\{\emptyset\}\), so \(\mu\) is not \(\sigma\)-finite and \(\widetilde{\mu}\equiv0\); the generated \(\sigma\)-ring \(\mathcal{S}\) is the Borel \(\sigma\)-algebra (it holds all \([a,b)\) and \(\mathbb{R}=\bigcup_n[-n,n)\), and lies inside the Borel class), and every cover of a nonempty set uses a nonempty member of \(\mathcal{R}\), so \(\overline{\mu}(E)=+\infty\) for every nonempty Borel \(E\).
The counting measure of the rationals, \(\nu(E)=\operatorname{card}(E\cap\mathbb{Q})\), is a second extension: \(\nu([a,b))=+\infty\) as every interval holds infinitely many rationals, yet \(\nu(\{0\})=1\ne+\infty=\overline{\mu}(\{0\})\). And \(\nu\) is \(\sigma\)-finite, \(\mathbb{R}\) being the union of the singletons \(\{q\}\), \(q\in\mathbb{Q}\), of measure \(1\) and of the \(\nu\)-null set \(\mathbb{R}\setminus\mathbb{Q}\).
(Luther) Let \(\mu\) be a measure with values in \([0,+\infty]\) defined on a \(\sigma\)-ring \(\mathcal{R}\). Prove that \(\mu=\mu_1+\mu_2\), where \(\mu_1\) is a semifinite measure on \(\mathcal{R}\), the measure \(\mu_2\) can assume only the values \(0\) and \(\infty\), and in every set \(R\in\mathcal{R}\) there exists a subset \(R^{\prime}\in\mathcal{R}\) such that \(\mu_1(R^{\prime})=\mu_1( R)\) and \(\mu_2(R^{\prime})=0\).
Take
\begin{equation*} \mu_1( R):=\sup\{\mu(P)\colon P\in\mathcal{R},\ P\subset R,\ \mu(P)<\infty\}, \end{equation*}
the semifinite part of \(\mu\) (Exercise 1.12.132(iii)), together with \(\mu_2( R):=\infty\) if some \(P\in\mathcal{R}\) with \(P\subset R\) has \(\mu_1(P)=0\) and \(\mu(P)=\infty\), and \(\mu_2( R):=0\) otherwise.
\(\mu_1\) is a semifinite measure with \(\mu_1\le\mu\): monotonicity, \(\mu_1(\emptyset)=0\), \(\mu_1\le\mu\) and \(\mu_1(P)=\mu(P)\) for \(\mu(P)<\infty\) are clear; for \(R=\bigcup_n R_n\) with disjoint \(R_n\in\mathcal{R}\) and \(P\subset R\) in \(\mathcal{R}\) of finite measure, \(\mu(P)=\sum_n\mu(P\cap R_n)\le\sum_n\mu_1(R_n)\), so \(\mu_1( R)\le\sum_n\mu_1(R_n)\); and for \(c_n<\mu_1(R_n)\), \(n\le N\), choosing \(P_n\subset R_n\) in \(\mathcal{R}\) with \(\mu(P_n)<\infty\), \(\mu(P_n)>c_n\) makes \(P:=\bigcup_{n\le N}P_n\subset R\) of finite measure with
\begin{equation*} \mu_1( R)\ge\mu(P)=\sum_{n\le N}\mu(P_n)>\sum_{n\le N}c_n . \end{equation*}
Semifiniteness: \(\mu_1( R)=\infty\) supplies \(P\subset R\) with \(\mu(P)\) finite and positive, and \(\mu_1(P)=\mu(P)\in(0,\infty)\).
\(\mu_2\) is a measure with values in \(\{0,\infty\}\). Let \(R=\bigcup_n R_n\) with disjoint \(R_n\in\mathcal{R}\). If \(\mu_2(R_n)=\infty\) for some \(n\), a witness \(P\subset R_n\) witnesses \(\mu_2( R)=\infty\) too. If all \(\mu_2(R_n)=0\) but \(\mu_2( R)=\infty\) with witness \(P\), then \(\mu(P)=\sum_n\mu(P\cap R_n)=\infty\) gives \(n\) with \(\mu(P\cap R_n)>0\); here \(\mu_1(P\cap R_n)\le\mu_1(P)=0\), so \(\mu(P\cap R_n)<\infty\) would force \(\mu_1(P\cap R_n)=\mu(P\cap R_n)>0\). Hence \(\mu(P\cap R_n)=\infty\) and \(P\cap R_n\) witnesses \(\mu_2(R_n)=\infty\), a contradiction.
\(\mu=\mu_1+\mu_2\). If \(\mu_2( R)=\infty\), then \(R\) contains a set of infinite \(\mu\)-measure, so \(\mu( R)=\infty\). If \(\mu_2( R)=0\) and \(\mu_1( R)<\mu( R)\), then \(\mu( R)=\infty\) (otherwise \(\mu_1( R)=\mu( R)\)) and \(\mu_1( R)<\infty\); choose \(P_n\subset R\) in \(\mathcal{R}\) with \(\mu(P_n)<\infty\), \(\mu(P_n)\to\mu_1( R)\), increasing after replacing \(P_n\) by \(P_1\cup\dots\cup P_n\), and put \(A=\bigcup_n P_n\in\mathcal{R}\), so continuity from below gives \(\mu(A)=\mu_1( R)<\infty\). Then
\begin{equation*} \mu(R\setminus A)=\mu( R)-\mu(A)=\infty,\qquad \mu_1(R\setminus A)=\mu_1( R)-\mu(A)=0 \end{equation*}
(using \(\mu_1(A)=\mu(A)\) and finiteness of every subtracted term), so \(R\setminus A\) witnesses \(\mu_2( R)=\infty\), a contradiction.
The subset: for \(R\in\mathcal{R}\) choose \(P_n\subset R\) in \(\mathcal{R}\) with \(\mu(P_n)<\infty\) and \(\mu(P_n)\to\mu_1( R)\) (taking \(\mu(P_n)\ge n\) when \(\mu_1( R)=\infty\), and \(P_n=\emptyset\) when \(\mu_1( R)=0\)), and put \(R^{\prime}:=\bigcup_n P_n\subset R\) in \(\mathcal{R}\). Then \(\mu_1(R^{\prime})\ge\mu_1(P_n)=\mu(P_n)\) for every \(n\), so monotonicity gives \(\mu_1(R^{\prime})=\mu_1( R)\). And \(R^{\prime}=\bigcup_n Q_n\) disjointly with \(Q_n:=P_n\setminus(P_1\cup\dots\cup P_{n-1})\in\mathcal{R}\) of finite \(\mu\)-measure, so \(\mu_2(Q_n)\le\mu(Q_n)<\infty\) forces \(\mu_2(Q_n)=0\) and \(\mu_2(R^{\prime})=\sum_n\mu_2(Q_n)=0\).
Let \(\mathcal{E}_1\) and \(\mathcal{E}_2\) be two algebras of subsets of \(\Omega\) and let \(\mu_1,\mu_2\) be two additive real functions on \(\mathcal{E}_1\) and \(\mathcal{E}_2\), respectively (or \(\mu_1,\mu_2\) take values in the extended real line and vanish at \(\emptyset\)). (a) Show that the equality \(\mu_1(E)=\mu_2(E)\) for all \(E\in\mathcal{E}_1\cap\mathcal{E}_2\) is necessary and sufficient for the existence of an additive function \(\mu\) that extends \(\mu_1\) and \(\mu_2\) to some algebra \(\mathcal{F}\) containing \(\mathcal{E}_1\) and \(\mathcal{E}_2\). (b) Show that if \(\mu_1,\mu_2\ge0\), then the existence of a common nonnegative extension \(\mu\) is equivalent to the following relations: \(\mu_1( C)\ge\mu_2(D)\) for all \(C\in\mathcal{E}_1\), \(D\in\mathcal{E}_2\) with \(D\subset C\) and \(\mu_1(E)\le\mu_2(F)\) for all \(E\in\mathcal{E}_1\), \(F\in\mathcal{E}_2\) with \(E\subset F\).
The printed parenthetical extended-real form of (a) is false — for \(\Omega=\{1,2,3\}\), \(\mathcal{E}_1=\{\emptyset,\{1\},\{2,3\},\Omega\}\) with \(\mu_1(\{1\})=+\infty\), \(\mu_1(\{2,3\})=0\), and \(\mathcal{E}_2=\{\emptyset,\{1,2\},\{3\},\Omega\}\) with \(\mu_2(\{1,2\})=5\), \(\mu_2(\{3\})=+\infty\), the two agree on \(\mathcal{E}_1\cap\mathcal{E}_2=\{\emptyset,\Omega\}\), yet any common additive extension has \(\{2\}=\{1,2\}\setminus\{1\}\) in its domain and \(\mu(\{1,2\})=\mu(\{1\})+\mu(\{2\})=+\infty\neq5\) — so we prove the statement for finite (real) additive functions.
Let \(\mathcal{F}\) be the algebra generated by \(\mathcal{E}_1\cup\mathcal{E}_2\), let \(V\) be the space of \(\mathcal{F}\)-simple real functions, and let \(S_i\subset V\) be the span of \(\{I_E\colon E\in\mathcal{E}_i\}\). There is a unique linear \(l_i\) on \(S_i\) with \(l_i(I_E)=\mu_i(E)\), and if \(f\in S_i\) takes the distinct values \(c_1,\dots,c_n\) on \(A_1,\dots,A_n\), then \(A_j\in\mathcal{E}_i\) and \(l_i(f)=\sum_j c_j\mu_i(A_j)\): any \(C_1,\dots,C_m\in\mathcal{E}_i\) generate a finite subalgebra whose atoms \(A_1,\dots,A_N\in\mathcal{E}_i\) partition \(\Omega\) with each \(C_k\) a union of atoms, so for \(f=\sum_k a_kI_{C_k}\), constant \(c_j\) on \(A_j\), additivity of \(\mu_i\) gives
\begin{equation*} \sum_k a_k\mu_i(C_k)=\sum_j\Bigl(\sum_{k\colon A_j\subset C_k}a_k\Bigr)\mu_i(A_j) =\sum_j c_j\mu_i(A_j), \end{equation*}
a quantity depending only on \(f\); the level sets of \(f\), being unions of atoms, lie in \(\mathcal{E}_i\).
(a) Necessity: \(\mu_1(E)=\mu(E)=\mu_2(E)\) for \(E\in\mathcal{E}_1\cap\mathcal{E}_2\).
Sufficiency. If \(\mu_1=\mu_2\) on \(\mathcal{E}_1\cap\mathcal{E}_2\), then \(l_1=l_2\) on \(S_1\cap S_2\): for such an \(f\) with distinct values \(c_j\) on \(A_j\), the lemma applied in \(S_1\) and in \(S_2\) puts each \(A_j\) in \(\mathcal{E}_1\cap\mathcal{E}_2\), so
\begin{equation*} l_1(f)=\sum_j c_j\mu_1(A_j)=\sum_j c_j\mu_2(A_j)=l_2(f). \end{equation*}
Hence \(l(f_1+f_2):=l_1(f_1)+l_2(f_2)\) is well defined and linear on \(S_1+S_2\) (from \(f_1+f_2=g_1+g_2\) one gets \(f_1-g_1=g_2-f_2\in S_1\cap S_2\)); extend \(l\) linearly to \(V\) along a Hamel basis and put \(\mu(A):=l(I_A)\), \(A\in\mathcal{F}\). Disjoint \(A,B\) have \(I_{A\cup B}=I_A+I_B\), so \(\mu\) is additive, and \(\mu=\mu_i\) on \(\mathcal{E}_i\).
(b) Necessity: a nonnegative additive function on an algebra is monotone, so \(\mu_1( C)=\mu( C)\ge\mu(D)=\mu_2(D)\) and \(\mu_1(E)=\mu(E)\le\mu(F)=\mu_2(F)\) in the two configurations.
Sufficiency. Taking \(C=D\) and \(E=F\) in \(\mathcal{E}_1\cap\mathcal{E}_2\) gives \(\mu_1=\mu_2\) there, in particular \(M:=\mu_1(\Omega)=\mu_2(\Omega)\).
Key claim: \(f\in S_1\), \(g\in S_2\) and \(g\le f\) imply \(l_2(g)\le l_1(f)\). For \(f,g\ge0\) pick \(0=c_0<c_1<\dots<c_n\) containing all their values, so that
\begin{equation*} f=\sum_{i=1}^n (c_i-c_{i-1})I_{\{f\ge c_i\}},\qquad g=\sum_{i=1}^n (c_i-c_{i-1})I_{\{g\ge c_i\}} \end{equation*}
(Check! at each point); by the lemma \(\{f\ge c_i\}\in\mathcal{E}_1\) and \(\{g\ge c_i\}\in\mathcal{E}_2\), with \(\{g\ge c_i\}\subset\{f\ge c_i\}\), so the first hypothesis of (b) and \(c_i-c_{i-1}>0\) give
\begin{equation*} l_2(g)=\sum_{i=1}^n(c_i-c_{i-1})\mu_2(\{g\ge c_i\}) \le\sum_{i=1}^n(c_i-c_{i-1})\mu_1(\{f\ge c_i\})=l_1(f). \end{equation*}
In general choose \(c>0\) with \(f+cI_\Omega\ge0\), \(g+cI_\Omega\ge0\); since \(\Omega\in\mathcal{E}_1\cap\mathcal{E}_2\), the case just treated gives \(l_2(g)+cM\le l_1(f)+cM\). Here \(\mu_1(\Omega)=\mu_2(\Omega)\) is where the second hypothesis of (b) enters, the first alone giving only \(\mu_1(\Omega)\ge\mu_2(\Omega)\).
Now set \(p(f):=\inf\{l_1(f_1)+l_2(f_2)\colon f_i\in S_i,\ f\le f_1+f_2\}\) for \(f\in V\). The infimum is over a nonempty set, as \(f\le\|f\|_\infty I_\Omega\) with \(I_\Omega\in S_1\), so \(p(f)<+\infty\); and \(f\le f_1+f_2\) gives \(-f_2\le f_1+\|f\|_\infty I_\Omega\in S_1\), so the key claim yields \(-l_2(f_2)\le l_1(f_1)+\|f\|_\infty M\), i.e. \(p(f)\ge-\|f\|_\infty M\). Adding admissible pairs makes \(p\) sublinear, with \(p(0)=0\).
By Hahn–Banach applied to the zero functional on \(\{0\}\) there is a linear \(l\le p\) on \(V\). For \(E\in\mathcal{E}_1\) the admissible pairs \((I_E,0)\) and \((-I_E,0)\) give
\begin{equation*} l(I_E)\le p(I_E)\le\mu_1(E),\qquad -l(I_E)\le p(-I_E)\le-\mu_1(E), \end{equation*}
so \(l(I_E)=\mu_1(E)\), and similarly \(l(I_F)=\mu_2(F)\) on \(\mathcal{E}_2\); and \((0,0)\) is admissible for \(-I_A\le0\), so \(l(I_A)\ge0\) for every \(A\in\mathcal{F}\). As in (a), \(\mu(A):=l(I_A)\) is the required nonnegative additive extension.
Let \((X,\mathcal{A},\mu)\) be a probability space and let \(\mu^*\) be the corresponding outer measure. For a set \(E\subset X\), we denote by \(\mathfrak{m}_E\) the restriction of \(\mu^*\) to the class of all subsets of \(E\). Show that \(\mathfrak{m}_E\) coincides with the outer measure on the space \(E\) generated by the restriction \(\mu_E\) of \(\mu\) to \(E\) in the sense of Definition 1.12.11. In particular, \(\mathfrak{m}_E\) is a regular Caratheodory outer measure.
The two outer measures agree because both equal \(\mu^*\) on subsets of \(E\). Write \(\mathcal{A}_E:=\{A\cap E\colon A\in\mathcal{A}\}\) and, by Definition 1.12.11, \(\mu_E(A\cap E):=\mu(A\cap\widetilde{E})\) with \(\widetilde{E}\) a measurable envelope of \(E\), so that
\begin{equation*} \mu_E^*(B)=\inf\{\mu_E( C)\colon C\in\mathcal{A}_E,\ B\subset C\},\qquad B\subset E, \end{equation*}
while \(\mathfrak{m}_E(B)=\mu^*(B)=\inf\{\mu(A)\colon A\in\mathcal{A},\ B\subset A\}\).
An envelope exists: take \(A_n\supset E\) in \(\mathcal{A}\) with \(\mu(A_n)\to\mu^*(E)\) and put \(\widetilde{E}=\bigcap_n A_n\), so \(\mu(\widetilde{E})=\mu^*(E)\). Everything rests on
\begin{equation*} \mu(A\cap\widetilde{E})=\mu^*(A\cap E)\qquad\text{for every } A\in\mathcal{A} \tag{\(\ast\)} \end{equation*}
(Proposition 1.12.12), which makes \(\mu_E\) well defined, independently of the representation \(S=A\cap E\) and of \(\widetilde{E}\). Here \(\ge\) is clear from \(A\cap E\subset A\cap\widetilde{E}\); conversely, for \(C\in\mathcal{A}\) with \(A\cap E\subset C\) the measurable set \(\widetilde{E}\cap\bigl(C\cup(X\setminus A)\bigr)\) contains \(E\), hence has measure at least \(\mu(\widetilde{E})\), so finiteness of \(\mu\) forces
\begin{equation*} \mu\bigl(\widetilde{E}\cap A\setminus C\bigr)=\mu(\widetilde{E}) -\mu\bigl(\widetilde{E}\cap(C\cup(X\setminus A))\bigr)=0, \end{equation*}
i.e. \(\mu(A\cap\widetilde{E})\le\mu( C)\); take the infimum over \(C\).
\(\mathfrak{m}_E\ge\mu_E^*\): for \(B\subset E\) and \(A\in\mathcal{A}\) with \(B\subset A\) we have \(B\subset A\cap E\in\mathcal{A}_E\) and \(\mu_E(A\cap E)=\mu(A\cap\widetilde{E})\le\mu(A)\), so the infimum over \(A\) gives \(\mu_E^*(B)\le\mu^*(B)\).
\(\mathfrak{m}_E\le\mu_E^*\): given \(\varepsilon>0\), pick \(A_\varepsilon\in\mathcal{A}\) with \(B\subset A_\varepsilon\cap E\) and \(\mu(A_\varepsilon\cap\widetilde{E})=\mu_E(A_\varepsilon\cap E)<\mu_E^*(B)+\varepsilon\); replacing \(A_\varepsilon\) by \(A_\varepsilon\cap\widetilde{E}\), which still contains \(B\) since \(B\subset E\subset\widetilde{E}\) and leaves that value unchanged, gives \(\mu^*(B)\le\mu(A_\varepsilon)<\mu_E^*(B)+\varepsilon\).
Finally \(\mu_E\) is a finite countably additive measure on the \(\sigma\)-algebra \(\mathcal{A}_E\): for disjoint \(S_n=A_n\cap E\) the sets \(A_n^{\prime}:=(A_n\cap\widetilde{E})\setminus\bigcup_{k<n}A_k\) are measurable and disjoint with \(A_n^{\prime}\cap E=S_n\) (a point of \(S_n\) lies in \(A_n\cap\widetilde{E}\) and in no earlier \(A_k\), else it would lie in \(S_k\cap S_n\)), so \((\ast)\) applied to \(\bigcup_n A_n^{\prime}\) gives
\begin{equation*} \mu_E\Bigl(\bigcup_n S_n\Bigr)=\sum_n\mu(A_n^{\prime}\cap\widetilde{E}) =\sum_n\mu_E(S_n). \end{equation*}
Hence Theorem 1.11.8 makes \(\mu_E^*\), and so \(\mathfrak{m}_E\), a regular Caratheodory outer measure on \(E\).
Suppose that \(\mu\) is a measure with values in \([0,+\infty]\) on a measurable space \((X,\mathcal{A})\). Let \(\mu^*\) and \(\mu_*\) be the corresponding outer and inner measures and let \(\mathfrak{m}:=(\mu^*+\mu_*)/2\).
(i) (Caratheodory) Show that \(\mathfrak{m}\) is a Caratheodory outer measure. Denote by \(\nu\) the measure generated by \(\mathfrak{m}\).
(ii) Let \(X=\{0,1\}\), \(\mathcal{A}=\{X,\emptyset\}\), \(\mu(X)=1\). Show that \(\mu\neq\nu\).
(iii) (Fremlin) Prove that if \(\mu\) is Lebesgue measure on \([0,1]\), then \(\mu=\nu\).
Here \(\mu^*(A)=\inf\{\mu(B)\colon A\subset B,\ B\in\mathcal{A}\}\) and \(\mu_*(A)=\sup\{\mu(B)\colon B\subset A,\ B\in\mathcal{A}\}\); the argument works verbatim for the variant \((1.12.7)\) in which \(B\) is additionally required to have finite measure.
(i) \(\mathfrak{m}(\emptyset)=0\) and \(\mathfrak{m}\) is monotone, so only countable subadditivity is at issue. Let \(A=\bigcup_{n\ge1}A_n\), where we may assume \(\mu^*(A_n)<\infty\) for every \(n\); then each \(A_n\) has a measurable envelope \(A_n^*\supset A_n\) with \(\mu(A_n^*)=\mu^*(A_n)\), namely \(\bigcap_j B_j\) for \(B_j\supset A_n\) in \(\mathcal{A}\) with \(\mu(B_j)\to\mu^*(A_n)\). Put \(D_n:=A_n^*\setminus\bigcup_{k<n}A_k^*\), disjoint and measurable with \(\bigcup_n D_n=\bigcup_n A_n^*\supset A\) and \(A_n^*\subset\bigcup_{k\le n}D_k\) (take the least \(k\) with \(x\in A_k^*\)); since \(A_n^*\cap D_n=D_n\),
\begin{equation*} \sum_n\mu^*(A_n)=\sum_n\mu(D_n)+R,\qquad R:=\sum_k\sum_{n>k}\mu(A_n^*\cap D_k). \tag{1} \end{equation*}
From \(A\subset\bigcup_n D_n\),
\begin{equation*} \mu^*(A)\le\sum_n\mu(D_n). \tag{2} \end{equation*}
For \(K\in\mathcal{A}\) with \(K\subset A\) we have \(\mu(K)=\sum_k\mu(K\cap D_k)\), and \(K\cap D_k\) splits into \(K\cap D_k\setminus\bigcup_{n>k}A_n^*\) and \(K\cap D_k\cap\bigcup_{n>k}A_n^*\). The first is measurable and contained in \(A_k\) — its points lie in \(A=\bigcup_m A_m\), lie in \(D_k\) hence in no \(A_j\) with \(j<k\), and in no \(A_n\) with \(n>k\) — so has measure at most \(\mu_*(A_k)\); the second has measure at most \(\sum_{n>k}\mu(A_n^*\cap D_k)\). Hence
\begin{equation*} \mu_*(A)\le\sum_k\mu_*(A_k)+R. \tag{3} \end{equation*}
Adding \((2)\) and \((3)\) and using \((1)\),
\begin{equation*} \mu^*(A)+\mu_*(A)\le\sum_n\mu^*(A_n)+\sum_n\mu_*(A_n), \end{equation*}
i.e. \(\mathfrak{m}(A)\le\sum_n\mathfrak{m}(A_n)\). By Theorem 1.11.4, \(\mathcal{M}_{\mathfrak{m}}\) is a \(\sigma\)-algebra carrying the countably additive \(\nu:=\mathfrak{m}|_{\mathcal{M}_{\mathfrak{m}}}\).
(ii) Here \(\mu^*(\{0\})=\mu^*(\{1\})=1\) and \(\mu_*(\{0\})=\mu_*(\{1\})=0\), so \(\mathfrak{m}(\{0\})=\mathfrak{m}(\{1\})=\tfrac12\) and \(\mathfrak{m}(X)=1\). Both singletons are \(\mathfrak{m}\)-measurable, the only nontrivial test set being \(T=X\), where \(\tfrac12+\tfrac12=1=\mathfrak{m}(X)\); so \(\mathcal{M}_{\mathfrak{m}}=2^X\) and \(\nu\) splits the \(\mu\)-atom \(X\) into two halves. In particular \(\nu\neq\mu\).
(iii) With \(\mu=\lambda\) Lebesgue measure on \([0,1]\), we show \(\mathcal{M}_{\mathfrak{m}}=\mathcal{A}\) and \(\mathfrak{m}=\mu\) there, which is the assertion \(\nu=\mu\).
For \(A\in\mathcal{A}\) we have \(\mu^*(A)=\mu_*(A)=\mu(A)\), so \(\mathfrak{m}(A)=\mu(A)\); and \(A\in\mathcal{M}_{\mathfrak{m}}\), because \(\mu^*(T)=\mu^*(T\cap A)+\mu^*(T\setminus A)\) for all \(T\) (measurable sets are Caratheodory measurable for \(\mu^*\)) and likewise \(\mu_*(T)=\mu_*(T\cap A)+\mu_*(T\setminus A)\) — for \(\ge\), measurable \(K\subset T\) has \(\mu(K)=\mu(K\cap A)+\mu(K\setminus A)\le\mu_*(T\cap A)+\mu_*(T\setminus A)\); for \(\le\), measurable \(K_1\subset T\cap A\) and \(K_2\subset T\setminus A\) are disjoint inside \(T\) — so adding and halving gives the Caratheodory identity for \(\mathfrak{m}\).
Conversely let \(A\subset[0,1]\) be nonmeasurable, with measurable kernel \(A_*\) and envelope \(A^*\), and put \(M:=A^*\setminus A_*\), \(d:=\mu(M)=\mu^*(A)-\mu_*(A)>0\), \(A_0:=A\cap M\), \(B_0:=M\setminus A\). Then
\begin{equation*} \mu_*(A_0)=\mu_*(B_0)=0,\qquad \mu^*(A_0)=\mu^*(B_0)=d , \end{equation*}
the first pair by the defining properties of kernel and envelope, the second because in these outer measures one may restrict to measurable \(C\subset M\) (replace \(C\) by \(C\cap M\)), and then \(M\setminus C\) is a measurable subset of \(B_0\subset A^*\setminus A\) resp. of \(A_0 = A\setminus A_*\), hence null, forcing \(\mu( C)=d\).
Splitting Lemma. If \(S\subset[0,1]\) and \(\lambda^*(S)>0\), then \(S\) partitions into \(S_1,S_2\) with \(\lambda^*(S_1)=\lambda^*(S)\) and \(\lambda^*(S_2)>0\).
Granted it, apply it to \(S=B_0\): it yields \(B_1\subset B_0\) with \(\mu^*(B_1)=d\) and \(\mu^*(B_0\setminus B_1)>0\). Put \(T:=A_0\cup B_1\subset M\), so \(T\cap A=A_0\), \(T\setminus A=B_1\), \(\mu_*(B_1)\le\mu_*(B_0)=0\), and
\begin{equation*} \mathfrak{m}(T\cap A)+\mathfrak{m}(T\setminus A)=\tfrac d2+\tfrac d2=d . \end{equation*}
But \(\mu^*(T)\le\lambda(M)=d\) and \(\mu_*(T)<d\): the supremum defining \(\mu_*(T)\) is attained at \(K:=\bigcup_n K_n\) for measurable \(K_n\subset T\) with \(\lambda(K_n)\to\mu_*(T)\), and \(\mu_*(T)=d\) would give \(\lambda(M\setminus K)=0\) with \(B_0\setminus B_1=M\setminus T\subset M\setminus K\), contradicting \(\mu^*(B_0\setminus B_1)>0\). Hence \(\mathfrak{m}(T)<d\), so \(A\notin\mathcal{M}_{\mathfrak{m}}\) and \(\mathcal{M}_{\mathfrak{m}}=\mathcal{A}\), \(\nu=\mu\).
It remains to justify the Lemma. Suppose some \(S\) with \(c:=\lambda^*(S)>0\) admits no such partition, and fix an envelope \(\widetilde S\) of \(S\), all envelopes of subsets of \(S\) being taken inside \(\widetilde S\). For measurable \(E\subset\widetilde S\) one has \(\lambda^*(S\setminus E)=c-\lambda(E)\) (\(\le\) from \(S\setminus E\subset\widetilde S\setminus E\); for \(\ge\), a measurable \(C\supset S\setminus E\) makes \(\widetilde S\cap(C\cup E)\) a measurable superset of \(S\), of measure at least \(c\)), and for \(B\subset S\) with envelope \(E_B\) the set \(B^{\prime}:=B\cup(S\setminus E_B)\) has \(\lambda^*(B^{\prime})=c\), since a measurable \(C\supset B^{\prime}\) satisfies
\begin{equation*} \lambda( C)\ge\lambda(C\cap E_B)+\lambda\bigl(C\cap(\widetilde S\setminus E_B)\bigr) \ge\lambda(E_B)+\bigl(c-\lambda(E_B)\bigr)=c . \end{equation*}
As \(S=B^{\prime}\cup\bigl((S\cap E_B)\setminus B\bigr)\) is then a partition with \(\lambda^*(B^{\prime})=c\), failure of the Lemma forces:
(a) \(\lambda^*\bigl((S\cap E_B)\setminus B\bigr)=0\) for every \(B\subset S\).
(b) \(m:=\lambda^*|_{\mathcal{P}(S)}\) is a countably additive atomless measure on \(\mathcal{P}(S)\). For disjoint \(B_1,B_2\subset S\), (a) makes \(B_2\cap E_{B_1}\subset(S\cap E_{B_1})\setminus B_1\) \(\lambda^*\)-null, so \(m(B_2)=\lambda^*(B_2\setminus E_{B_1})\) with an envelope \(E_2\subset\widetilde S\setminus E_{B_1}\), and a measurable \(C\supset B_1\cup B_2\) has
\begin{equation*} \lambda( C)\ge\lambda(C\cap E_{B_1})+\lambda(C\cap E_2)\ge m(B_1)+m(B_2); \end{equation*}
subadditivity gives the reverse, and finite additivity with monotonicity plus countable subadditivity gives countable additivity. Atomless: if \(m(B)>0\), split \(E_B\) into measurable halves \(E^{\prime},E^{\prime\prime}\) of measure \(\lambda(E_B)/2\); then \(m(B\cap E^{\prime})+m(B\cap E^{\prime\prime})=m(B)=\lambda(E_B)\) with both terms at most \(\lambda(E_B)/2\), so both equal \(\lambda(E_B)/2>0\). In particular \(m\) vanishes on countable sets.
(c) \(\Phi\colon B\mapsto E_B\) modulo \(\lambda\)-null sets is an isomorphism of the measure algebra of \((S,\mathcal{P}(S),m)\) onto the Lebesgue algebra of \(\widetilde S\), which is separable in the metric \(\lambda(E\triangle F)\); so \(m\) has countable Maharam type. It is well defined and injective on classes: \(m(B_1\triangle B_2)=0\) makes \(E_{B_2}\cup E_{B_1\setminus B_2}\) a measurable superset of \(B_1\), so minimality of envelopes gives \(E_{B_1}\subset E_{B_2}\cup E_{B_1\setminus B_2}\) up to a null set, with \(\lambda(E_{B_1\setminus B_2})=m(B_1\setminus B_2)=0\), and symmetrically. It preserves measure (\(\lambda(E_B)=m(B)\)) and unions (\(E_{B_1}\cup E_{B_2}\) is an envelope of \(B_1\cup B_2\)), and respects complements: by (b), \(m(S\setminus B)=c-\lambda(E_B)=\lambda(\widetilde S\setminus E_B)\), while \(S\setminus B\subset(\widetilde S\setminus E_B)\cup\bigl((S\cap E_B)\setminus B\bigr)\) with the second set \(\lambda^*\)-null by (a). It is surjective, since for measurable \(E\subset\widetilde S\) the set \(S\cap E\) has envelope \(E\): a measurable \(C\supset S\cap E\) makes \((C\cap E)\cup(\widetilde S\setminus E)\) a measurable superset of \(S\), so \(\lambda(C\cap E)\ge\lambda(E)\).
Under the continuum hypothesis this already closes the argument, by Ulam’s theorem that no countably additive \(m\) on \(\mathcal{P}(\omega_1)\) vanishes on countable sets with \(m(\omega_1)>0\): fixing injections \(f_\beta\colon\beta\to\mathbb{N}\) and setting
\begin{equation*} A_\alpha^n:=\{\beta<\omega_1\colon\ \alpha<\beta,\ f_\beta(\alpha)=n\} , \end{equation*}
each \(\bigcup_n A_\alpha^n\) is co-countable, hence of full measure, so \(m(A_\alpha^n)>1/k\) for some pair \((n,k)\), and one pair \((n_0,k_0)\) serves uncountably many \(\alpha\); but the sets \(A_\alpha^{n_0}\) are pairwise disjoint by injectivity of the \(f_\beta\), so at most \(k_0\,m(\omega_1)<\infty\) of them can have measure exceeding \(1/k_0\). Since \(S\) is non-null it is uncountable, and under CH \(|S|\le\mathfrak c=\aleph_1\) forces \(|S|=\aleph_1\), excluded by (b).
In ZFC the gap is closed instead by the theorem of Gitik and Shelah (Israel J. Math. 68 (1989), 129–160): a countably additive atomless measure on the \(\sigma\)-algebra of ALL subsets of a set never has a separable measure algebra. This contradicts (b) and (c), so the Splitting Lemma, and with it \(\mu=\nu\), holds; this is the route of Fremlin, Fund. Math. 139 (1991), 9–15, cited in the exercise. (A fully rigorous proof of the remaining step is beyond the scope of this page; see the reference given in the book.)
Let \(\mathfrak{m}\) be a Caratheodory outer measure on a space \(X\) and let \(\varphi\colon[0,+\infty]\to[0,+\infty)\) be a bounded concave function such that \(\varphi(0)=0\) and \(\varphi(t)>0\) if \(t\neq0\). Let \(d(A,B)=\varphi\bigl(\mathfrak{m}(A\triangle B)\bigr)\), \(A,B\in\mathfrak{M}_{\mathfrak{m}}\). Denote by \(\widetilde{\mathfrak{M}}_\mu\) the factor-space of the space \(\mathfrak{M}_{\mathfrak{m}}\) by the ring of \(\mathfrak{m}\)-zero sets. Show that \((\widetilde{\mathfrak{M}}_\mu,d)\) is a complete metric space.
By Theorem 1.11.4 the class \(\mathfrak{M}_{\mathfrak{m}}\) is a \(\sigma\)-algebra carrying the countably additive \(\mathfrak{m}\), and the \(\mathfrak{m}\)-zero sets form a \(\sigma\)-ideal in it (measurable by completeness, closed under countable unions by subadditivity). Write \(\rho(A,B):=\mathfrak{m}(A\triangle B)\in[0,+\infty]\) and \(A\sim B\) for \(\rho(A,B)=0\), the relation defining \(\widetilde{\mathfrak{M}}_\mu\), transitive since \(A\triangle C\subset(A\triangle B)\cup(B\triangle C)\).
Properties of \(\varphi\). (a) It is nondecreasing: if \(0\le s<t<\infty\) had \(\varphi(t)<\varphi(s)\), then writing \(t=\lambda s+(1-\lambda)u\) with \(u>t\), \(\lambda=(u-t)/(u-s)\), concavity gives \(\varphi(t)\ge\lambda\varphi(s)+(1-\lambda)\varphi(u)\), i.e.
\begin{equation*} \varphi(u)\le\varphi(s)+\frac{\varphi(t)-\varphi(s)}{t-s}(u-s)\to-\infty \quad (u\to\infty), \end{equation*}
contradicting \(\varphi\ge0\); and \(\varphi(+\infty)\ge\lambda\varphi(+\infty)+(1-\lambda)\varphi(s)\) yields \(\varphi(+\infty)\ge\varphi(s)\). (b) It is subadditive: for \(s,t>0\) with \(s+t<\infty\), concavity and \(\varphi(0)=0\) give \(\varphi(s)\ge\frac{s}{s+t}\varphi(s+t)\) and \(\varphi(t)\ge\frac{t}{s+t}\varphi(s+t)\), which add to the claim; \(s=0\) or \(t=0\) is trivial, and if \(s=+\infty\) then \(\varphi(s+t)=\varphi(+\infty)=\varphi(s)\). (c) \(\varphi(t_n)\to0\) forces \(t_n\to0\), since \(t_{n_k}\ge\varepsilon>0\) would give \(\varphi(t_{n_k})\ge\varphi(\varepsilon)>0\) by (a).
\(d\) is a well-defined metric: \(A\sim A^{\prime}\), \(B\sim B^{\prime}\) give \(A\triangle B\subset(A^{\prime}\triangle B^{\prime})\cup(A\triangle A^{\prime})\cup(B\triangle B^{\prime})\), so \(\rho(A,B)=\rho(A^{\prime},B^{\prime})\) by subadditivity and symmetry; \(d\) is finite (as \(\varphi\) maps into \([0,+\infty)\)) and symmetric; \(d(A,B)=0\) forces \(\rho(A,B)=0\) because \(\varphi>0\) off \(0\), and conversely \(\varphi(0)=0\); and \(\rho(A,C)\le\rho(A,B)+\rho(B,C)\) with (a), (b) gives
\begin{equation*} d(A,C)\le\varphi\bigl(\rho(A,B)+\rho(B,C)\bigr)\le d(A,B)+d(B,C). \end{equation*}
Completeness. Let \((A_n)\) be \(d\)-Cauchy, with representatives in \(\mathfrak{M}_{\mathfrak{m}}\). By (c), \(\rho(A_n,A_k)\to0\): given \(\varepsilon>0\) we have \(d(A_n,A_k)<\varphi(\varepsilon)\) for large \(n,k\), while \(\rho(A_n,A_k)\ge\varepsilon\) would force \(d(A_n,A_k)\ge\varphi(\varepsilon)\). Choose \(n_1<n_2<\cdots\) with \(\rho(A_{n_k},A_{n_{k+1}})\le2^{-k}\) and put \(B_j:=\bigcup_{k\ge j}A_{n_k}\), \(A:=\bigcap_{j\ge1}B_j\in\mathfrak{M}_{\mathfrak{m}}\). Since
\begin{equation*} \bigcup_{k> j}\bigl(A_{n_k}\setminus A_{n_j}\bigr) \subset\bigcup_{i\ge j}\bigl(A_{n_{i+1}}\triangle A_{n_i}\bigr) \end{equation*}
(for \(x\in A_{n_k}\setminus A_{n_j}\) with \(k>j\), take the least \(i\in[j,k)\) with \(x\in A_{n_{i+1}}\); then \(x\notin A_{n_i}\)), countable subadditivity gives \(\mathfrak{m}(A\setminus A_{n_j})\le\mathfrak{m}(B_j\setminus A_{n_j})\le\sum_{i\ge j}2^{-i}=2^{-j+1}\). In the other direction \(A_{n_j}\setminus A=\bigcup_{i\ge j}(A_{n_j}\setminus B_i)\) with the sets \(A_{n_j}\setminus B_i\subset A_{n_j}\setminus A_{n_i}\) increasing in \(i\), so continuity from below gives
\begin{equation*} \mathfrak{m}(A_{n_j}\setminus A)=\lim_i\mathfrak{m}(A_{n_j}\setminus B_i)\le2^{-j+1}, \end{equation*}
whence \(\rho(A_{n_j},A)\le2^{-j+2}\to0\), and \(\rho(A_n,A)\le\rho(A_n,A_{n_j})+\rho(A_{n_j},A)\to0\) by the \(\rho\)-Cauchy property.
Let \(c:=\inf_{t>0}\varphi(t)=\lim_{t\downarrow0}\varphi(t)\ge0\). (i) \(c=0\): given \(\varepsilon>0\) pick \(\delta>0\) with \(\varphi(\delta)<\varepsilon\); then \(\rho(A_n,A)<\delta\) for large \(n\) gives \(d(A_n,A)\le\varphi(\delta)<\varepsilon\) by (a). (ii) \(c>0\): \(d(A_n,A_k)<c\) for large \(n,k\) forces \(\rho(A_n,A_k)=0\), so \(\rho(A_n,A)\le\rho(A_k,A)\to0\), giving \(\rho(A_n,A)=0\) and \(d(A_n,A)=0\) for all large \(n\). Either way \(A_n\to A\) in \(d\).
Exercises 1.12.145–1.12.151
(Steinhaus [910]) Let \(E\) be a set of positive measure on the real line. Prove that, for every finite set \(F\), the set \(E\) contains a subset similar to \(F\), i.e., having the form \(c + tF\), where \(t \neq 0\).
Take \(t := \ell/(4n)\) and any \(c \in \bigcap_{i \le n}(A^{\prime} - t f_i)\), with \(A^{\prime}\) and \(\ell\) supplied by the density argument below; then \(c + tF \subset E\). Here \(\lambda\) is Lebesgue measure and \(\lambda(E)>0\).
Normalizing \(F\). Let \(F = \{f_1 < f_2 < \dots < f_n\}\). For \(n = 1\) pick \(e \in E\), nonempty since \(\lambda(E) > 0\), and take \(c = e - f_1\), \(t = 1\). For \(n \ge 2\) put \(d := f_n - f_1 > 0\) and \(G := d^{-1}(F - f_1) = \{0 = g_1 < \dots < g_n = 1\}\); a copy \(c + tG \subset E\) with \(t\ne0\) gives
\begin{equation*} c + tG = \Bigl(c - \frac{t f_1}{d}\Bigr) + \frac{t}{d} F , \end{equation*}
again of the required form, with the nonzero coefficient \(t/d\). So assume \(0 = f_1 < \dots < f_n = 1\).
An interval of high density. Fix \(\varepsilon := 1/(2n)\), choose \(N\) with \(a := \lambda(A) > 0\) for \(A := E \cap [-N,N]\), and by outer regularity take an open \(U \supset A\) with \(\lambda(U) < a/(1-\varepsilon)\). Writing \(U = \bigcup_k I_k\) as an at most countable disjoint union of open intervals, each of finite length since \(\lambda(U)<\infty\), we have \(\sum_k \lambda(A \cap I_k) = a\); so \(\lambda(A\cap I_k)\le(1-\varepsilon)\lambda(I_k)\) for every \(k\) would give \(a \le (1-\varepsilon)\lambda(U) < a\). Hence some \(I := I_k\) has
\begin{equation*} \lambda(A^{\prime}) > (1-\varepsilon)\,\ell, \qquad A^{\prime} := A \cap I \subset E, \quad \ell := \lambda(I) \in (0,\infty). \end{equation*}
A simultaneous shift. For \(t>0\) put \(S_i := A^{\prime} - t f_i\), so that \(c \in S_i\) means \(c + tf_i \in A^{\prime}\). Since \(0 \le f_i \le 1\), the set \(S_i\) is the translate of \(A^{\prime}\subset I\) by \(-tf_i\) with \(tf_i\in[0,t]\), so \(S_i \subset I - tf_i\) and \(\lambda(S_i\cap I)\ge\lambda(A^{\prime}) - t\). For measurable \(T_1,\dots,T_n\subset I\),
\begin{equation*} \lambda\Bigl(I \setminus \bigcap_{i=1}^n T_i\Bigr) \le \sum_{i=1}^n \bigl(\ell - \lambda(T_i)\bigr), \end{equation*}
that is, \(\lambda(\bigcap_i T_i)\ge\sum_i\lambda(T_i)-(n-1)\ell\); with \(T_i = S_i\cap I\) and \(\varepsilon = 1/(2n)\),
\begin{equation*} \lambda\Bigl(\bigcap_{i=1}^n S_i\Bigr) > n\bigl((1-\varepsilon)\ell - t\bigr) - (n-1)\ell = \frac{\ell}{2} - nt . \end{equation*}
At \(t = \ell/(4n) > 0\) the right-hand side is \(\ell/4 > 0\), so the intersection is nonempty and any \(c\) in it satisfies \(c + tf_i \in A^{\prime} \subset E\) for every \(i\).
(i) Let \(\mu\) be an atomless probability measure on a measurable space \((X, \mathcal{A})\). Show that every point \(x \in X\) belongs to \(\mathcal{A}_\mu\) and has \(\mu\)-measure zero.
(ii) (Marczewski [651]) Prove that if a probability measure \(\mu\) on a measurable space \((X, \mathcal{A})\) is atomless, then there exist nonempty sets of \(\mu\)-measure zero.
(i) \(\mu^*(\{x\}) = 0\). Put \(c := \mu^*(\{x\}) = \inf\{\mu(A) : A \in \mathcal{A},\ x \in A\}\in[0,1]\), choose \(A_k \in \mathcal{A}\) with \(x \in A_k\), \(\mu(A_k) < c + 1/k\), and set \(E := \bigcap_{k} A_k \in \mathcal{A}\), so \(x \in E\) and \(\mu(E) = c\). Then \(E\) is a measurable envelope of \(\{x\}\):
\begin{equation*} B \in \mathcal{A}, \quad B \subset E \setminus \{x\} \ \Longrightarrow \ \mu(B) = 0 , \end{equation*}
since \(E\setminus B \in \mathcal{A}\) contains \(x\), forcing \(\mu(E\setminus B)\ge c = \mu(E)\). If \(c > 0\), then \(E\) is not an atom (Definition 1.12.7), so atomlessness supplies \(B \in \mathcal{A}\), \(B \subset E\), with \(0 < \mu(B) < \mu(E)\); both \(B\) and \(E\setminus B\) then have positive measure while \(x\) lies in at most one of them, so the other is a measurable subset of \(E\setminus\{x\}\) of positive measure — contradiction. Hence \(c = 0\), and \(\emptyset \subset \{x\} \subset E\) with \(\mu(E) = 0\) puts \(\{x\}\) in \(\mathcal{A}_\mu\) with \(\mu(\{x\}) = 0\).
(ii) \(\mu(X) = 1\) makes \(X\) nonempty, and any \(\{x\}\) is then a nonempty set of \(\mu\)-measure zero by (i).
The book’s stronger claim that an uncountable \(\mu\)-null set exists is not provable in ZFC: under the continuum hypothesis, enumerating the Borel null sets as \(\{N_\xi\}_{\xi<\omega_1}\) and the Borel sets of positive Lebesgue measure as \(\{P_\xi\}_{\xi<\omega_1}\) and choosing by transfinite recursion
\begin{equation*} x_\xi \in P_\xi \setminus \Bigl( \bigcup_{\eta \le \xi} N_\eta \cup \{x_\eta : \eta < \xi\} \Bigr) \end{equation*}
(possible since \(\lambda(P_\xi) > 0\) while the removed set is a countable union of null sets) yields an uncountable \(S := \{x_\xi : \xi<\omega_1\}\subset[0,1]\) meeting every Lebesgue null set in an at most countable set (as \(S\cap N_\xi\subset\{x_\eta : \eta<\xi\}\) and every Lebesgue null set sits in a Borel one) and every Borel set of positive measure, so that \(X := S\), \(\mathcal{A} := \{B\cap S : B \text{ Borel}\}\), \(\mu(B\cap S) := \lambda(B)\) is a well-defined atomless probability measure (Check!: \(B\triangle B^{\prime}\) disjoint from \(S\) is Borel null) every one of whose null sets is at most countable.
(Kindler [517]) Let \(\mathcal{S}\) be a family of subsets of a set \(\Omega\) with \(\varnothing \in \mathcal{S}\) and let \(\alpha, \beta : \mathcal{S} \to (-\infty, +\infty]\) be two set functions vanishing at \(\varnothing\). Prove that the following conditions are equivalent:
(i) there exists an additive set function \(\mu\) on the set of all subsets of \(\Omega\) taking values in \((-\infty, +\infty]\) and satisfying the condition \(\alpha \le \mu|_{\mathcal{S}} \le \beta\);
(ii) if \(A_i, B_j \in \mathcal{S}\) and \(\sum_{i=1}^n I_{A_i} = \sum_{j=1}^m I_{B_j}\), then \(\sum_{i=1}^n \alpha(A_i) \le \sum_{j=1}^m \beta(B_j)\).
The printed range \((-\infty,+\infty]\) makes (ii)\(\Rightarrow\)(i) false, so we prove the equivalence for \(\alpha,\beta:\mathcal{S}\to\mathbb{R}\); the counterexample is \(\Omega=\{1,2,3\}\), \(\mathcal{S}=\{\varnothing,\{1,2\},\{1,3\},\{2,3\}\}\), \(\alpha=\beta=+\infty\) on \(\{1,2\}\) and \(=0\) on the other two, where the three indicators are linearly independent (so (ii) is vacuous) while \(\mu(\{1,3\})=\mu(\{2,3\})=0\) forces \(\mu(\{1,2\})=-2\mu(\{3\})\) to be finite.
(i)\(\Rightarrow\)(ii). Let \(C_1,\dots,C_N\) be the nonempty atoms of the finite algebra generated by \(A_1,\dots,A_n,B_1,\dots,B_m\); every \(I_{A_i}, I_{B_j}\) is constant on each \(C_k\), so \(p_k:=\#\{i: C_k\subset A_i\}\) and \(q_k:=\#\{j: C_k\subset B_j\}\) satisfy \(p_k=q_k\), and additivity of \(\mu\) along the atoms gives
\begin{equation*} \sum_{i=1}^n \alpha(A_i) \le \sum_{i=1}^n \mu(A_i) = \sum_{k=1}^N p_k \mu(C_k) = \sum_{k=1}^N q_k \mu(C_k) = \sum_{j=1}^m \mu(B_j) \le \sum_{j=1}^m \beta(B_j). \end{equation*}
(ii)\(\Rightarrow\)(i). On the space \(V\) of simple functions on \(\Omega\) put \(M:=\mathrm{span}\{I_S: S\in\mathcal{S}\}\), call
\begin{equation*} f = \sum_{j=1}^m t_j I_{B_j} - \sum_{i=1}^n s_i I_{A_i}, \qquad A_i, B_j \in \mathcal{S},\ t_j, s_i > 0, \end{equation*}
a representation of \(f\in M\) (split \(f\) into positive and negative parts; empty sums allowed), give it the cost \(\sum_j t_j\beta(B_j) - \sum_i s_i\alpha(A_i)\), and let \(p(f)\) be the infimum of the costs of the representations of \(f\). Concatenating and scaling representations give \(p(f+g)\le p(f)+p(g)\) and \(p(cf)=c\,p(f)\) for \(c>0\) (Check!), and the empty representation gives \(p(0)\le 0\).
Also \(p(0)\ge 0\): for a representation of \(0\) with rational \(t_j,s_i\), clearing a common denominator \(N\) and listing each \(B_j\) exactly \(Nt_j\) times and each \(A_i\) exactly \(Ns_i\) times produces two finite families in \(\mathcal{S}\) with equal indicator sums, so (ii) gives cost \(\ge 0\); for real \((t,s)\), evaluation on the atoms of the algebra generated by the \(A_i,B_j\) turns \(\sum_j t_jI_{B_j}=\sum_i s_iI_{A_i}\) into a linear system with \(0\)–\(1\) coefficients, whose solution space \(W\) therefore has \(W\cap\mathbb{Q}^{m+n}\) dense in \(W\), and a negative cost at the strictly positive point \((t,s)\) would be inherited, by continuity of the cost, at a nearby rational point of \(W\) with positive coordinates.
Hence \(p\) is real-valued on \(M\): costs are real, so \(p(f)<+\infty\), and fixing one representation of \(-f\), of cost \(c\), every representation of \(f\) concatenates with it to a representation of \(0\), forcing cost \(\ge -c\) (the only use of finiteness of \(\beta\)). So \(p\) is sublinear with \(p(0)=0\), and the Hahn–Banach theorem applied to the zero functional on \(\{0\}\) — dominated by \(p\) — yields a linear \(\lambda_0\le p\) on \(M\); extend by \(0\) on an algebraic complement of \(M\) in \(V\) to a linear \(\lambda\) on \(V\) and put \(\mu(E):=\lambda(I_E)\). Then \(I_{E\cup F}=I_E+I_F\) for disjoint \(E,F\) makes \(\mu\) additive on all subsets of \(\Omega\), and the one-term representations of \(I_B\) (cost \(\beta(B)\)) and of \(-I_A\) (cost \(-\alpha(A)\)) give
\begin{equation*} \mu(B) \le p(I_B) \le \beta(B), \qquad -\mu(A) \le p(-I_A) \le -\alpha(A), \end{equation*}
that is, \(\alpha \le \mu|_{\mathcal{S}} \le \beta\).
Prove Proposition 1.12.36. Moreover, show that there is a nonnegative additive function \(\alpha\) on the set of all subsets of \(X\) with \(\alpha|_{\mathcal{R}} \le \beta\) and \(\alpha(X) = \beta(X)\).
We produce the nonnegative additive \(\alpha\) on all subsets of \(X\) with \(\alpha|_{\mathcal{R}}\le\beta\) and \(\alpha(X)=\beta(X)\); its restriction to \(\mathcal{R}\) is nonnegative, hence monotone, and additive, hence modular, which is Proposition 1.12.36. For \(\beta(X)=0\) monotonicity gives \(\beta\equiv 0\) and \(\alpha\equiv 0\) works, so normalize \(\beta(X)=1\).
(a) For \(R_1,\dots,R_n\in\mathcal{R}\) there are \(R_1^{\prime}\subset\dots\subset R_n^{\prime}\) in \(\mathcal{R}\) with \(\sum_i I_{R_i}=\sum_i I_{R_i^{\prime}}\) and \(\sum_i\beta(R_i)\ge\sum_i\beta(R_i^{\prime})\). Induct on \(n\), the case \(n=1\) being trivial: given \(R_1,\dots,R_{n+1}\), apply the hypothesis to \(R_1,\dots,R_n\) to get a chain \(S_1\subset\dots\subset S_n\), put \(S_{n+1}:=R_{n+1}\) and
\begin{equation*} S_i^{\prime} := S_i \ (i<n), \quad S_n^{\prime} := S_n \cap S_{n+1}, \quad S_{n+1}^{\prime} := S_n \cup S_{n+1}, \end{equation*}
apply the hypothesis again to \(S_1^{\prime},\dots,S_n^{\prime}\) to get a chain \(R_1^{\prime}\subset\dots\subset R_n^{\prime}\), and set \(R_{n+1}^{\prime}:=S_{n+1}^{\prime}\). Every point of \(R_n^{\prime}\) lies in some \(S_i^{\prime}\) (the two indicator sums agree there) and \(\bigcup_{i\le n}S_i^{\prime}\subset S_n\subset R_{n+1}^{\prime}\), so the chain extends; \(I_{A\cap B}+I_{A\cup B}=I_A+I_B\) gives the indicator identity, and submodularity gives
\begin{equation*} \begin{aligned} \sum_{i=1}^{n+1} \beta(R_i^{\prime}) &\le \sum_{i=1}^{n} \beta(S_i^{\prime}) + \beta(S_{n+1}^{\prime}) \\ &= \sum_{i=1}^{n-1}\beta(S_i) + \beta(S_n \cap S_{n+1}) + \beta(S_n \cup S_{n+1}) \\ &\le \sum_{i=1}^{n}\beta(S_i) + \beta(R_{n+1}) \le \sum_{i=1}^{n+1}\beta(R_i) . \end{aligned} \end{equation*}
(b) If \(\sum_{i=1}^n I_{R_i}\ge m\) on \(X\), then \(\sum_i\beta(R_i)\ge m\): by (a) assume \(R_1\subset\dots\subset R_n\), and for each \(x\) the upward closed set \(\{j: x\in R_j\}\) has at least \(m\) elements, hence contains \(n-m+1\); so \(R_j=X\) for \(j\ge n-m+1\) and \(\sum_i\beta(R_i)\ge m\beta(X)=m\).
(c) On the space \(L\) of real simple functions on \(X\) put
\begin{equation*} p(f) := \inf\Bigl\{ \textstyle\sum_{i=1}^n a_i \beta(R_i) : R_i \in \mathcal{R}, \ a_i \ge 0,\ f \le \sum_{i=1}^n a_i I_{R_i} \Bigr\} . \end{equation*}
As \(f\) is bounded and \(X\in\mathcal{R}\), the infimum is over a nonempty family, so \(0\le p(f)<\infty\); combining and scaling admissible majorants makes \(p\) sublinear (Check!). Moreover \(p(1)=1\): were \(1\le\sum_i a_iI_{R_i}\) with \(\sum_i a_i\beta(R_i)<1\), enlarging each \(a_i\) slightly to a rational \(n_i/m\) preserves both relations, and repeating each \(R_i\) exactly \(n_i\) times puts (b) in force with constant \(m\), giving \(\sum_i n_i\beta(R_i)\ge m\).
Since \(\lambda_0(t\cdot 1):=t\) satisfies \(\lambda_0\le p\) on \(\mathbb{R}\cdot 1\) (equality for \(t\ge 0\); for \(t<0\) use \(p\ge 0\)), Hahn–Banach gives a linear \(\lambda\le p\) on \(L\) with \(\lambda(1)=1\). Put \(\nu(E):=\lambda(I_E)\), so that \(\nu\) is additive with \(\nu(\varnothing)=0\), \(\nu(X)=1\) and
\begin{equation*} E \subset R \in \mathcal{R} \ \Longrightarrow \ \nu(E) \le p(I_E) \le \beta( R) , \end{equation*}
whence \(\nu\le 1\). Set \(\alpha(E):=\sup_{A\subset E}\nu(A)\in[0,1]\). For disjoint \(E_1,E_2\) with union \(E\), any \(A\subset E\) splits as \(A=(A\cap E_1)\cup(A\cap E_2)\), so \(\nu(A)\le\alpha(E_1)+\alpha(E_2)\), while any \(A_i\subset E_i\) give \(\nu(A_1)+\nu(A_2)=\nu(A_1\cup A_2)\le\alpha(E)\); taking suprema, \(\alpha\) is additive. Finally \(\alpha( R)\le\beta( R)\) by the displayed implication, and
\begin{equation*} 1 = \nu(X) \le \alpha(X) \le \beta(X) = 1 . \end{equation*}
Let \((X, \mathcal{A}, \mu)\) be a probability space and let \(\mathcal{S}\) be a family of subsets in \(X\) such that \(\mu_*\bigl(\bigcup_{n=1}^\infty S_n\bigr) = 0\) for every countable collection \(\{S_n\} \subset \mathcal{S}\). Prove that there exists a probability measure \(\widetilde{\mu}\) defined on some \(\sigma\)-algebra \(\widetilde{\mathcal{A}}\) such that \(\mathcal{A}, \mathcal{S} \subset \widetilde{\mathcal{A}}\), \(\widetilde{\mu}\) extends \(\mu\) and vanishes on \(\mathcal{S}\), and for each \(A \in \widetilde{\mathcal{A}}\) there exists \(A^{\prime} \in \mathcal{A}\) with \(\widetilde{\mu}(A \bigtriangleup A^{\prime}) = 0\).
Take the \(\sigma\)-ideal \(\mathcal{Z}\) of all sets covered by countably many members of \(\mathcal{S}\), and
\begin{equation*} \widetilde{\mathcal{A}} := \{ A \bigtriangleup Z : A \in \mathcal{A}, \ Z \in \mathcal{Z}\}, \qquad \widetilde{\mu}(A \bigtriangleup Z) := \mu(A) . \end{equation*}
Here \(\mathcal{Z}\supset\mathcal{S}\) is closed under subsets and countable unions (a countable union of countable covers is one), and \(\mu_*(Z)=0\) for \(Z\in\mathcal{Z}\) by the hypothesis and monotonicity of \(\mu_*\); so every \(Z\in\mathcal{Z}\cap\mathcal{A}\) is \(\mu\)-null.
\(\widetilde{\mathcal{A}}\) contains \(\mathcal{A}\) (\(Z=\varnothing\)) and \(\mathcal{S}\) (\(A=\varnothing\)), is closed under complements since \(X\setminus(A\bigtriangleup Z)=(X\setminus A)\bigtriangleup Z\), and is closed under countable unions: for \(E_k=A_k\bigtriangleup Z_k\) and \(A:=\bigcup_k A_k\in\mathcal{A}\),
\begin{equation*} \Bigl(\bigcup_k E_k\Bigr) \bigtriangleup A \subset \bigcup_k (E_k \bigtriangleup A_k) = \bigcup_k Z_k \in \mathcal{Z}. \end{equation*}
The hypothesis acts through the unambiguity of \(\widetilde\mu\): if \(A\bigtriangleup Z=A^{\prime}\bigtriangleup Z^{\prime}\), adding \(A\bigtriangleup Z^{\prime}\) in the group \((2^X,\bigtriangleup)\) gives \(A\bigtriangleup A^{\prime}=Z\bigtriangleup Z^{\prime}\in\mathcal{A}\cap\mathcal{Z}\), so \(\mu(A)=\mu(A^{\prime})\).
For countable additivity let \(E_k=A_k\bigtriangleup Z_k\) be pairwise disjoint, so \(\widetilde\mu(\bigcup_k E_k)=\mu(\bigcup_k A_k)\) by the previous paragraph. The \(A_k\) are disjoint up to null sets: a point \(x\in A_k\cap A_l\) (\(k\ne l\)) outside \(Z_k\cup Z_l\) would lie in \(E_k\cap E_l=\varnothing\), so \(A_k\cap A_l\in\mathcal{A}\cap\mathcal{Z}\) is null. Hence \(B_k:=A_k\setminus\bigcup_{j<k}A_j\) are disjoint with \(\mu(B_k)=\mu(A_k)\) and the same union, whence
\begin{equation*} \widetilde{\mu}\Bigl(\bigcup_k E_k\Bigr) = \sum_k \mu(B_k) = \sum_k \mu(A_k) = \sum_k \widetilde{\mu}(E_k). \end{equation*}
Thus \(\widetilde\mu\) is a probability measure extending \(\mu\), it vanishes on \(\mathcal{Z}\supset\mathcal{S}\) because \(Z=\varnothing\bigtriangleup Z\), and \(E=A\bigtriangleup Z\) satisfies \(\widetilde\mu(E\bigtriangleup A)=\widetilde\mu(Z)=0\).
(\(\circ\)) Let \(\mu\) be a bounded nonnegative measure on a \(\sigma\)-algebra \(\mathcal{A}\) in a space \(X\). Denote by \(\mathcal{E}\) the class of all sets \(E \subset X\) such that
\begin{equation*} \mu^*(E) = \mu^*(E \setminus A) + \mu^*(E \cap A) \quad \text{for all } A \in \mathcal{A} . \end{equation*}
Is it true that the function \(\mu^*\) is additive on \(\mathcal{E}\)?
No, and Pugachev’s example from the book’s hint makes \(\mathcal{E}\) the family of all subsets of \(X\): take \(X = \{1,-1,i,-i\}\), let \(\mathcal{A}\) be the algebra generated by the partition
\begin{equation*} X_1 := \{1\}, \qquad X_2 := \{-1\}, \qquad X_3 := \{i,-i\}, \end{equation*}
which is a \(\sigma\)-algebra as it is finite, and let \(\mu(A) := \#\{k : X_k \subset A\}\), a bounded nonnegative measure.
Since \(\mathcal{A}\) is finite, the infimum in \(\mu^*(E) = \inf\{\mu(A) : A \in \mathcal{A},\ E \subset A\}\) is attained at the union of the atoms meeting \(E\), so
\begin{equation*} \mu^*(E) = \#\{k \in \{1,2,3\} : E \cap X_k \ne \varnothing\} . \end{equation*}
Every \(A \in \mathcal{A}\) either contains an atom \(X_k\) or is disjoint from it, so for each \(k\) exactly one of \((E\cap A)\cap X_k\) and \((E\setminus A)\cap X_k\) is empty and the other equals \(E \cap X_k\); summing the resulting indicator identities over \(k = 1,2,3\) gives \(\mu^*(E) = \mu^*(E\cap A) + \mu^*(E \setminus A)\) for all \(E \subset X\) and \(A \in \mathcal{A}\). But \(E_1 := \{i\}\) and \(E_2 := \{-i\}\) are disjoint and
\begin{equation*} \mu^*(E_1) + \mu^*(E_2) = 1 + 1 = 2 \ne 1 = \mu^*(E_1 \cup E_2) . \end{equation*}
(Rado, Reichelderfer [777, p. 260]) Let \(\Phi\) be a finite nonnegative set function defined on the family \(\mathcal{U}\) of all open sets in \((0,1)\) such that:
(i) \(\Phi\bigl(\bigcup_{n=1}^\infty U_n\bigr) = \sum_{n=1}^\infty \Phi(U_n)\) for every countable family of pairwise disjoint sets \(U_n \in \mathcal{U}\),
(ii) \(\Phi(U_1) \le \Phi(U_2)\) whenever \(U_1, U_2 \in \mathcal{U}\) and \(U_1 \subset U_2\),
(iii) \(\Phi(U) = \lim_{\varepsilon \to 0} \Phi(U_\varepsilon)\) for every \(U \in \mathcal{U}\), where \(U_\varepsilon\) is the set of all points in \(U\) with distance more than \(\varepsilon\) from the boundary of \(U\).
Is it true that \(\Phi\) has a countably additive extension to the Borel \(\sigma\)-algebra of \((0,1)\)?
No: with \(K := [1/4,1/2]\), the hint’s set function
\begin{equation*} \Phi(U) := \begin{cases} 1, & K \subset U, \\ 0, & K \not\subset U, \end{cases} \qquad U \in \mathcal{U}, \end{equation*}
satisfies (i)–(iii) and has no countably additive Borel extension.
(ii) is immediate from \(K \subset U_1 \subset U_2\). For (i), let the \(U_n\) be disjoint with union \(U\): if \(K \subset U\), the sets \(U_n \cap K\) are disjoint, relatively open, and cover the connected set \(K\), so exactly one is nonempty and then equals \(K\), giving \(\sum_n \Phi(U_n) = 1 = \Phi(U)\); if \(K \not\subset U\), then \(K \not\subset U_n\) for every \(n\) and both sides vanish. For (iii): if \(K \not\subset U\), then \(U_\varepsilon \subset U\) gives \(\Phi(U_\varepsilon) = 0 = \Phi(U)\); if \(K \subset U\), then \(d := \mathrm{dist}(K,\partial U) > 0\) since \(K\) is compact and \(\partial U\) is closed and disjoint from \(U\) (with \(d = +\infty\) when \(\partial U = \varnothing\)), so \(K \subset U_\varepsilon\) and \(\Phi(U_\varepsilon) = 1 = \Phi(U)\) for \(\varepsilon < d\).
Let \(\nu\) be a finite countably additive extension to the Borel \(\sigma\)-algebra. The open sets \((1/8,3/8)\), \((1/8,1/4)\) and \((1/4,3/8)\) all fail to contain \(K\), so
\begin{equation*} \nu(\{1/4\}) = \nu\bigl((1/8,3/8)\bigr) - \nu\bigl((1/8,1/4)\bigr) - \nu\bigl((1/4,3/8)\bigr) = 0, \end{equation*}
and likewise \(\nu(\{1/2\}) = 0\) via \((3/8,5/8)\); since \((1/4,1/2)\) also omits \(K\), this gives \(\nu(K) = 0\). But \(V_n := (1/4 - 1/n,\, 1/2 + 1/n)\cap(0,1)\) are open sets containing \(K\) with \(V_n \downarrow K\), so continuity from above of the finite measure \(\nu\) forces
\begin{equation*} \nu(K) = \lim_{n\to\infty}\nu(V_n) = \lim_{n\to\infty}\Phi(V_n) = 1 . \end{equation*}
Exercises 1.12.152–1.12.158
Let \(\mu\) be a nonnegative \(\sigma\)-finite measure on a measurable space \((X,\mathcal{A})\) and let \(M_0\) be the class of all sets of finite \(\mu\)-measure. Let
\begin{equation*} \sigma_\mu(A,B)=\mu(A\triangle B)/\mu(A\cup B)\ \text{ if }\ \mu(A\cup B)>0,\qquad \sigma_\mu(A,B)=0\ \text{ if }\ \mu(A\cup B)=0. \end{equation*}
(i) (Marczewski, Steinhaus [653]) (a) Show that \(\sigma_\mu\) is a metric on the space of equivalence classes in \(M_0\), where \(A\sim B\) whenever \(\mu(A\triangle B)=0\).
(b) Show that if \(A_n,A\in M_0\) and \(\sigma_\mu(A_n,A)\to 0\), then \(\mu(A_n\triangle A)\to 0\).
(c) Show that if \(\mu(A_n\triangle A)\to 0\) and \(\mu(A)>0\), then \(\sigma_\mu(A_n,A)\to 0\).
(d) Observe that \(\sigma_\mu(\emptyset,B)=1\) if \(\mu(B)>0\) and deduce that in the case of Lebesgue measure on \([0,1]\), the identity mapping \((M_0,d)\to (M_0,\sigma_0)\), where \(d\) is the Fréchet-Nikodym metric, is discontinuous at the point corresponding to \(\emptyset\).
(ii) (Gladysz, Marczewski, Ryll-Nardzewski [359]) For all \(A_1,\dots,A_n\in M_0\) let
\begin{equation*} \sigma_\mu(A_1,\dots,A_n)=\frac{\mu\bigl((A_1\cup\cdots\cup A_n)\setminus (A_1\cap\cdots\cap A_n)\bigr)}{\mu(A_1\cup\cdots\cup A_n)} \end{equation*}
if \(\mu(A_1\cup\cdots\cup A_n)>0\) and \(\sigma_\mu(A_1,\dots,A_n)=0\) if \(\mu(A_1\cup\cdots\cup A_n)=0\). Prove the inequality
\begin{equation*} \sigma_\mu(A_1,\dots,A_n)\le \frac{1}{n-1}\sum_{i<j}\sigma_\mu(A_i,A_j). \end{equation*}
Deduce that if \(\sigma_\mu(A_i,A_j)<2/n\) for all \(1\le i<j\le n\), then \(\mu(A_1\cap\cdots\cap A_n)>0\).
(i)(a) \(\sigma_\mu\) takes values in \([0,1]\) since \(A\triangle B\subset A\cup B\), and it depends only on the equivalence classes because \((A\triangle B)\triangle(A^{\prime}\triangle B^{\prime})=(A\triangle A^{\prime})\triangle(B\triangle B^{\prime})\) and \((A\cup B)\triangle(A^{\prime}\cup B^{\prime})\subset(A\triangle A^{\prime})\cup(B\triangle B^{\prime})\) are null; symmetry is clear, and \(\sigma_\mu(A,B)=0\) exactly when \(\mu(A\triangle B)=0\) (when \(\mu(A\cup B)=0\) use \(A\triangle B\subset A\cup B\)).
For the triangle inequality assume \(p:=\mu(A\cup C)>0\), put \(E:=B\setminus(A\cup C)\) and
\begin{equation*} t:=\mu(A\triangle C),\quad e:=\mu(E),\quad U:=\mu(A\cup B\cup C)=p+e, \end{equation*}
\(s:=\mu(A\triangle B)+\mu(B\triangle C)\). Then \(E\subset(A\triangle B)\cap(B\triangle C)\), while \(A\triangle C\subset(A\triangle B)\cup(B\triangle C)\) is disjoint from \(E\), so
\begin{equation*} s=\mu\bigl((A\triangle B)\cup(B\triangle C)\bigr) +\mu\bigl((A\triangle B)\cap(B\triangle C)\bigr)\ge t+e . \end{equation*}
Since \(t\le p\), this gives \(tU=tp+te\le tp+pe=p(t+e)\le ps\), i.e. \(\sigma_\mu(A,C)=t/p\le s/U\); and \(\mu(A\cup B)\le U\), \(\mu(B\cup C)\le U\) give \(\mu(A\triangle B)/U\le\sigma_\mu(A,B)\) and \(\mu(B\triangle C)/U\le\sigma_\mu(B,C)\), whence \(\sigma_\mu(A,C)\le\sigma_\mu(A,B)+\sigma_\mu(B,C)\).
(b) Write \(\sigma_n:=\sigma_\mu(A_n,A)\). (i) If \(\mu(A)>0\), then \(\mu(A_n\cup A)\le\mu(A)+\mu(A_n\triangle A)\) gives
\begin{equation*} \mu(A_n\triangle A)=\sigma_n\,\mu(A_n\cup A) \le\sigma_n\bigl(\mu(A)+\mu(A_n\triangle A)\bigr), \end{equation*}
so \(\mu(A_n\triangle A)\le 2\sigma_n\mu(A)\to 0\) once \(\sigma_n\le 1/2\) (all terms being finite). (ii) If \(\mu(A)=0\), then \(\mu(A_n\triangle A)=\mu(A_n)=\mu(A_n\cup A)\), so \(\sigma_n=1\) whenever \(\mu(A_n)>0\); hence \(\mu(A_n\triangle A)=0\) for all large \(n\).
(c) \(\mu(A_n\cup A)\ge\mu(A)>0\), so \(\sigma_\mu(A_n,A)\le\mu(A_n\triangle A)/\mu(A)\to 0\).
(d) \(\mu(\emptyset\triangle B)=\mu(B)=\mu(\emptyset\cup B)\) gives \(\sigma_\mu(\emptyset,B)=1\) for \(\mu(B)>0\); so for \(\mu=\lambda\) on \([0,1]\) and \(B_n:=[0,1/n]\) we have \(d(B_n,\emptyset)=1/n\to 0\) while \(\sigma_0(B_n,\emptyset)=1\ne 0=\sigma_0(\emptyset,\emptyset)\), and the identity map \((M_0,d)\to(M_0,\sigma_0)\) is discontinuous at the class of \(\emptyset\).
(ii) Put \(U:=A_1\cup\cdots\cup A_n\), \(I:=A_1\cap\cdots\cap A_n\), \(k(x):=\#\{i\le n: x\in A_i\}\), and assume \(\mu(U)>0\) (else the inequality is trivial). A point \(x\in U\setminus I\) has \(1\le k(x)\le n-1\) and lies in exactly \(k(x)(n-k(x))\ge n-1\) of the sets \(A_i\triangle A_j\), since \(k(n-k)-(n-1)=(k-1)(n-1-k)\ge 0\); integrating \(\sum_{i<j}\mathbf{1}_{A_i\triangle A_j}\ge(n-1)\mathbf{1}_{U\setminus I}\) and using \(\mu(A_i\cup A_j)\le\mu(U)\),
\begin{equation*} \sum_{i<j}\sigma_\mu(A_i,A_j)\ \ge\ \frac{1}{\mu(U)}\sum_{i<j}\mu(A_i\triangle A_j) \ \ge\ (n-1)\frac{\mu(U\setminus I)}{\mu(U)} . \end{equation*}
If now \(\sigma_\mu(A_i,A_j)<2/n\) for all \(i<j\) and \(\mu(U)>0\) (without this proviso all \(A_i\) could be null), then \(\sigma_\mu(A_1,\dots,A_n)<\frac{1}{n-1}\binom{n}{2}\frac{2}{n}=1\), so \(\mu(U\setminus I)<\mu(U)<\infty\) and \(\mu(I)=\mu(U)-\mu(U\setminus I)>0\).
(\(\circ\)) Let \(A_1,\dots,A_n\) be measurable sets in a probability space \((\Omega,\mathcal{A},P)\). Prove that
\begin{equation*} 0\le \sum_{i=1}^{n}P(A_i)-P\Bigl(\bigcup_{i=1}^{n}A_i\Bigr)\le \sum_{1\le i<j\le n}P(A_i\cap A_j). \end{equation*}
Integrate the pointwise inequality: with \(k(\omega):=\sum_{i=1}^n\mathbf{1}_{A_i}(\omega)\) one has \(\mathbf{1}_{\bigcup_i A_i}=\mathbf{1}_{\{k\ge 1\}}\), and \(\sum_{i<j}\mathbf{1}_{A_i\cap A_j}=\binom{k}{2}\) because the pairs \(i<j\) with \(\omega\in A_i\cap A_j\) are the pairs drawn from the \(k(\omega)\)-element set \(\{i:\omega\in A_i\}\), while for every integer \(k\ge 0\)
\begin{equation*} 0\ \le\ k-\mathbf{1}_{\{k\ge 1\}}\ \le\ \binom{k}{2} \end{equation*}
(all terms vanish for \(k\le 1\), and \(0\le k-1\le k(k-1)/2\) for \(k\ge 2\)). All functions here are bounded and measurable and \(P\) is finite, so integrating against \(P\) gives both inequalities.
Method (2): the left inequality is subadditivity of \(P\) (Proposition 1.3.9(ii)); for the right one induct on \(n\), the case \(n=1\) being trivial. With \(B_1:=A_n\setminus\bigcup_{i<n}A_i\) and \(B_2:=A_n\cap\bigcup_{i<n}A_i\) disjoint with union \(A_n\), and \(P(\bigcup_{i\le n}A_i)=P(\bigcup_{i<n}A_i)+P(B_1)\),
\begin{equation*} \sum_{i=1}^{n}P(A_i)-P\Bigl(\bigcup_{i=1}^{n}A_i\Bigr) =\sum_{i=1}^{n-1}P(A_i)-P\Bigl(\bigcup_{i=1}^{n-1}A_i\Bigr)+P(B_2), \end{equation*}
and \(P(B_2)\le\sum_{i<n}P(A_i\cap A_n)\) by subadditivity.
(Darji, Evans [203]) Let \(A\) be a measurable set in the unit cube \(I\) of \(\mathbb{R}^n\), let \(F\subset I\setminus A\) be a finite set, and let \(\varepsilon>0\). Show that there exists a finite set \(S\subset A\) with the following property: for every partition \(\mathcal{P}\) of the cube \(I\) into finitely many parallelepipeds of the form \([a_1,b_1]\times\cdots\times[a_n,b_n]\) with pairwise disjoint interiors, letting \(B:=\bigcup\{P\in\mathcal{P}:\ P\cap F\ne\emptyset,\ P\cap S=\emptyset\}\) we have \(\lambda_n(A\cap B)<\varepsilon\).
Take for \(S\) one point of \(A\cap Q(z,s,g)\) out of each corner box on a fine grid that carries mass \(\ge\delta/2\), where
\begin{equation*} \delta:=\frac{\varepsilon}{2\cdot 4^{n}|F|}>0,\qquad Q(z,s,h):=\prod_{i=1}^{n}J_i,\quad J_i:=\begin{cases}[z_i,z_i+h_i]\cap[0,1], & s_i=1,\\ [z_i-h_i,z_i]\cap[0,1], & s_i=-1,\end{cases} \end{equation*}
for \(z\in I\), \(s\in\{-1,1\}^n\), \(h\in[0,1]^n\); here \(|F|\) is the number of points of \(F\), and we may assume \(F\ne\emptyset\) (otherwise \(B=\emptyset\) always and \(S=\emptyset\) works). Precisely: fix an integer \(m>2n/\delta\), let \(G:=\{0,1/m,\dots,1\}^n\), and for every \(z\in F\), \(s\in\{-1,1\}^n\) and \(g\in G\) with \(\lambda_n(A\cap Q(z,s,g))\ge\delta/2\) choose a point \(p(z,s,g)\in A\cap Q(z,s,g)\) (a set of positive measure is nonempty); let \(S\) be the set of these at most \(|F|2^n(m+1)^n\) points, so \(S\subset A\) is finite.
Three facts. (i) At most \(2^n\) nondegenerate cells of a partition contain a given \(z\in I\): for a nondegenerate \(P=\prod_i[a_i,b_i]\ni z\) define \(\sigma(P)\in\{-1,1\}^n\) by \(\sigma_i:=1\) if \(z_i<b_i\) and \(\sigma_i:=-1\) otherwise (then \(z_i=b_i>a_i\)); for all small \(\eta>0\) the point \(y_i:=z_i+\eta\sigma_i\) lies in \(\operatorname{int}P=\prod_i(a_i,b_i)\), so two distinct such cells with the same signature would have intersecting interiors. (ii) If \(P=\prod_i[a_i,b_i]\subset I\) contains \(z\), then multiplying out \([a_i,b_i]=[a_i,z_i]\cup[z_i,b_i]\) gives
\begin{equation*} P=\bigcup_{s\in\{-1,1\}^n}Q\bigl(z,s,h^{s}\bigr),\qquad h^{s}_i:=\begin{cases}b_i-z_i,& s_i=1,\\ z_i-a_i,& s_i=-1,\end{cases} \end{equation*}
whence \(\lambda_n(A\cap P)\le\sum_{s}\lambda_n(A\cap Q(z,s,h^{s}))\). (iii) If \(z\in F\), \(s\in\{-1,1\}^n\), \(h\in[0,1]^n\) and \(Q(z,s,h)\cap S=\emptyset\), then \(\lambda_n(A\cap Q(z,s,h))<\delta\): rounding down, \(g_i:=\lfloor mh_i\rfloor/m\le h_i\), gives \(Q(z,s,g)\subset Q(z,s,h)\) and
\begin{equation*} Q(z,s,h)\setminus Q(z,s,g)\subset \bigcup_{i=1}^{n}\{x\in Q(z,s,h):\ |x_i-z_i|>g_i\}, \end{equation*}
whose \(i\)-th set has measure at most \(h_i-g_i\le 1/m\), so the difference has measure at most \(n/m<\delta/2\); were \(\lambda_n(A\cap Q(z,s,h))\ge\delta\) we would get \(\lambda_n(A\cap Q(z,s,g))\ge\delta/2\), so \(p(z,s,g)\in S\) would lie in \(Q(z,s,g)\subset Q(z,s,h)\).
Now let \(\mathcal{P}\) be any such partition and \(\mathcal{P}^{\prime}:=\{P\in\mathcal{P}: P\cap F\ne\emptyset,\ P\cap S=\emptyset\}\), so \(B=\bigcup\mathcal{P}^{\prime}\). Degenerate cells (some \(a_i=b_i\)) are \(\lambda_n\)-null, and by (i) at most \(2^n|F|\) nondegenerate cells of \(\mathcal{P}^{\prime}\) survive, since each contains a point of \(F\). For such a \(P\), picking \(z\in P\cap F\), every corner box \(Q(z,s,h^{s})\) lies in \(P\), hence misses \(S\), so (ii) and (iii) give \(\lambda_n(A\cap P)\le 2^n\delta\) and
\begin{equation*} \lambda_n(A\cap B)\le 2^n|F|\cdot 2^n\delta=4^n|F|\delta =\frac{\varepsilon}{2}<\varepsilon . \end{equation*}
(Kahane [479]) Let \(E\) be the set of all points in \([0,1]\) of the form \(x=3\sum_{n=1}^{\infty}\varepsilon_n 4^{-n}\), \(\varepsilon_n\in\{0,1\}\). Show that \(E+\tfrac12 E=[0,3/2]\), but for almost all real \(\lambda\), the set \(E+\lambda E\) has measure zero.
\(E\) is the compact set of points of \([0,1]\) all of whose base-\(4\) digits are \(0\) or \(3\) (the \(3\varepsilon_n\) lie in \(\{0,3\}\)), i.e. the self-similar set of \(x\mapsto x/4\) and \(x\mapsto x/4+3/4\); the first assertion is digit splitting, the second is the Besicovitch projection theorem applied to \(K:=E\times E\).
Digit splitting. For \(x=3\sum_n\varepsilon_n4^{-n}\) and \(y=3\sum_n\eta_n4^{-n}\) with \(\varepsilon_n,\eta_n\in\{0,1\}\),
\begin{equation*} x+\tfrac12 y=3\sum_{n=1}^{\infty}\Bigl(\varepsilon_n+\tfrac{\eta_n}{2}\Bigr)4^{-n} =\frac{3}{2}\sum_{n=1}^{\infty}c_n4^{-n},\qquad c_n:=2\varepsilon_n+\eta_n . \end{equation*}
Since \((\varepsilon,\eta)\mapsto2\varepsilon+\eta\) is a bijection of \(\{0,1\}^2\) onto \(\{0,1,2,3\}\) and the pairs \((\varepsilon_n,\eta_n)\) are chosen independently, \((c_n)\) runs over all sequences of base-\(4\) digits, so \(\sum_nc_n4^{-n}\) runs over all of \([0,1]\) (the value \(1\) from all digits \(3\)) and \(E+\tfrac12E=\tfrac32[0,1]=[0,3/2]\).
Projections. Writing \(\pi_\lambda(x,y):=x+\lambda y=\sqrt{1+\lambda^2}\,\langle(x,y),e_\lambda\rangle\) with \(e_\lambda:=(1,\lambda)/\sqrt{1+\lambda^2}\), the compact set \(E+\lambda E=\pi_\lambda(K)\) is \(\sqrt{1+\lambda^2}\) times the orthogonal projection of \(K\) onto \(\mathbb{R}e_\lambda\); as \(\lambda\mapsto\arctan\lambda\) is a diffeomorphism onto \((-\pi/2,\pi/2)\), null sets of \(\lambda\) and of directions correspond. Since \(E=E/4\cup(E/4+3/4)\), the set \(K\) is the union of four copies of itself scaled by \(1/4\) in the corner squares of the unit square, i.e. the four-corner Cantor set; let \(K_m\) be the union of its \(4^m\) squares of side \(4^{-m}\), so \(K\cap Q\) is a copy of \(K\) scaled by \(4^{-m}\) for each such square \(Q\).
(i) \(0<\mathcal{H}^1(K)<\infty\). The \(4^m\) squares of side \(4^{-m}\) have diameter \(\sqrt2\,4^{-m}\), so \(\mathcal{H}^1(K)\le\sqrt2\); and \(\pi_{1/2}\) is Lipschitz with constant \(\sqrt5/2\), so by the first part
\begin{equation*} \mathcal{H}^1(K)\ \ge\ \frac{2}{\sqrt5}\,\mathcal{H}^1\bigl(\pi_{1/2}(K)\bigr) =\frac{2}{\sqrt5}\cdot\frac32>0, \end{equation*}
\(\mathcal{H}^1\) on \(\mathbb{R}\) being Lebesgue measure.
(ii) \(K\) is purely \(1\)-unrectifiable. Fix \(x\in K\) and \(m\), and let \(Q\) be the \(m\)-th step square containing \(x\). Distinct squares inside one square of the \((m-1)\)-st step are separated in some coordinate by \(4^{-m+1}-2\cdot4^{-m}=2\cdot4^{-m}\), and squares in distinct \((m-1)\)-st step squares are farther apart, so with \(r:=\tfrac32 4^{-m}\) we get \(K\cap B(x,r)\subset Q\) and hence
\begin{equation*} \frac{\mathcal{H}^1\bigl(K\cap B(x,r)\bigr)}{2r} \le \frac{4^{-m}\mathcal{H}^1(K)}{3\cdot 4^{-m}}\le \frac{\sqrt2}{3}<1 , \end{equation*}
so \(\theta^1_*(K,x):=\liminf_{r\to0}\mathcal{H}^1(K\cap B(x,r))/2r\le\sqrt2/3\) at every \(x\in K\). A \(1\)-rectifiable \(\mathcal{H}^1\)-measurable \(R\subset K\) with \(\mathcal{H}^1( R)>0\) (necessarily finite by (i)) would have \(\theta^1(R,x)=1\) a.e. on \(R\) by the Besicovitch density theorem (Mattila, Geometry of Sets and Measures in Euclidean Spaces, Theorem 17.6), forcing \(\theta^1_*(K,x)\ge1\) there.
By (i), (ii) and the Besicovitch projection theorem (Mattila, op. cit., Theorem 18.1; Federer, Geometric Measure Theory, 3.3.13) the projection of \(K\) is one-dimensionally null in almost every direction, so \(\lambda(E+\lambda E)=0\) for almost every \(\lambda\in\mathbb{R}\).
Multivariate distribution functions admit the following characterization. For any vectors \(x=(x_1,\dots,x_n)\), \(y=(y_1,\dots,y_n)\) let
\begin{equation*} [x,y):=[x_1,y_1)\times\cdots\times[x_n,y_n). \end{equation*}
Given a function \(F\) on \(\mathbb{R}^n\) let \(F[x,y):=\sum_u s(u)F(u)\), where the summation is taken over all corner points \(u\) of the set \([x,y)\) and \(s(u)\) equals \(+1\) or \(-1\) depending on whether the number of indices \(k\) with \(u_k=y_k\) is even or odd. Prove that the function \(F\) on \(\mathbb{R}^n\) is the distribution function of some probability measure precisely when the following conditions are fulfilled: 1) \(F[x,y)\ge 0\) whenever \(x<y\) coordinatewise, 2) \(F(x^j)\to F(x)\) whenever the vectors \(x^j\) increase to \(x\), 3) \(F(x)\to 0\) as \(\max_k x_k\to-\infty\) and \(F(x)\to 1\) as \(\min_k x_k\to+\infty\).
Throughout \(F(x):=\mu(S_x)\), \(S_x:=(-\infty,x_1)\times\cdots\times(-\infty,x_n)\), and the sign is normalized so that the corner \(u=y\) enters with \(+1\), i.e. \(s(u):=(-1)^{\#\{k:\,u_k=x_k\}}\); the printed convention differs by the factor \((-1)^n\), which for \(n=1\) would turn condition 1) into reverse monotonicity. The printed condition 3) must likewise be read as
\begin{equation*} 3^{\prime})\qquad F(x)\to 0\ \text{ as }\ \min_k x_k\to-\infty,\qquad F(x)\to 1\ \text{ as }\ \min_k x_k\to+\infty, \end{equation*}
since \(F(x_1,x_2):=G(x_1)\) with \(G\) a one-dimensional distribution function satisfies 1), 2) and the printed 3) while \(F[x,y)\equiv0\) forces any measure representing it to vanish on \(\mathcal{B}(\mathbb{R}^2)\).
Necessity. Multiplying \(\mathbf{1}_{[x_k,y_k)}=\mathbf{1}_{(-\infty,y_k)}-\mathbf{1}_{(-\infty,x_k)}\) over \(k\) and expanding gives \(\mathbf{1}_{[x,y)}=\sum_u s(u)\mathbf{1}_{S_u}\) over the \(2^n\) corners, so integration in \(\mu\) yields
\begin{equation*} \mu\bigl([x,y)\bigr)=\sum_u s(u)F(u)=F[x,y)\ \ge 0 , \end{equation*}
which is 1). If \(x^j\uparrow x\) then \(S_{x^j}\uparrow S_x\) (given \(t_k<x_k\) for all \(k\), finitely many convergences give \(j\) with \(t_k<x^j_k\) for all \(k\)), so 2) is continuity from below, Proposition 1.3.3(iii). For 3’): \(\min_kx_k\le-M\) gives \(S_x\subset\{t_{k_0}<-M\}\) for some \(k_0\), and these sets decrease to \(\varnothing\); \(\min_kx_k\ge M\) gives \(S_x\supset\{t:\ t_k<M\ \forall k\}\), whose measure tends to \(1\), and \(F\le1\).
Sufficiency. First, \(F\) is nondecreasing in each variable with \(0\le F\le1\): applying 1) to the box with corners given by the \(y_j\) and by \(x_j:=-T\) \((j\ne k)\) and letting \(T\to+\infty\), every corner with some \(u_j=-T\) has \(F(u)\to0\) by 3’), leaving the two corners \((y_j)_{j\ne k}\) paired with \(y_k\) (sign \(+1\)) and with \(x_k\) (sign \(-1\)); squeezing \(x^{\prime}\le x\le x^{\prime\prime}\) with all coordinates tending to \(\mp\infty\) then gives \(0\le F\le1\).
Fix \(j\), put \(Q_j:=[-j\mathbf{1},j\mathbf{1})\) and let \(\mathcal{S}_j\) be the boxes \([x,y)\subset Q_j\) with \(-j\le x_k\le y_k\le j\). This is a semialgebra in the sense of Definition 1.2.13 — it contains \(\varnothing, Q_j\), is closed under intersections, and splitting each \([x_k,y_k)\) into \([x_k,x_k^{\prime}),[x_k^{\prime},y_k^{\prime}),[y_k^{\prime},y_k)\) exhibits \([x,y)\setminus[x^{\prime},y^{\prime})\) as a finite disjoint union of boxes — so by Lemma 1.2.14 the finite unions form an algebra \(\mathcal{A}_j\). Put \(m([x,y)):=F[x,y)\), nonnegative by 1) for \(x<y\) and \(=0\) when some \(x_k=y_k\) (the \(2^n\) terms cancel in pairs differing in the \(k\)-th coordinate).
With \((\Delta_k^{a,b}G)(t):=G(t_1,\dots,b,\dots,t_n)-G(t_1,\dots,a,\dots,t_n)\) one has \(F[x,y)=\Delta_1^{x_1,y_1}\cdots\Delta_n^{x_n,y_n}F\), and \(\Delta_k^{a,b}=\sum_{i=1}^r\Delta_k^{t_{i-1},t_i}\) for \(a=t_0<\cdots<t_r=b\); multiplying out gives the grid additivity \(F[x,y)=\sum_\alpha F[C_\alpha]\) over the cells of any coordinatewise subdivision. Given a partition \([x,y)=\bigcup_{i=1}^N[x^i,y^i)\), take for \(T_k\) the finite set of \(x_k,y_k\) and all \(x^i_k,y^i_k\): each resulting cell is, coordinatewise, inside or disjoint from each \([x^i_k,y^i_k)\), hence lies in exactly one \([x^i,y^i)\), and the cells inside \([x^i,y^i)\) are exactly the induced grid there, so
\begin{equation*} F[x,y)=\sum_{\text{all cells}}F[C]=\sum_{i=1}^{N}\ \sum_{C\subset [x^i,y^i)}F[C] =\sum_{i=1}^{N}F[x^i,y^i). \end{equation*}
Thus \(m\) is additive on \(\mathcal{S}_j\) and, by Proposition 1.3.10, extends uniquely to a nonnegative additive set function on \(\mathcal{A}_j\).
Countable additivity comes from Theorem 1.4.3 with the compact class of finite unions of closed parallelepipeds in \(Q_j\) (a compact class by Example 1.4.2): for \(A=\bigcup_{i=1}^N[x^i,y^i)\in\mathcal{A}_j\) with disjoint nonempty boxes and \(\varepsilon>0\), choose \(x^i\le y^i( r)\uparrow y^i\), so that each corner of \([x^i,y^i( r))\) increases to the corresponding corner of \([x^i,y^i)\) and 2) gives \(m([x^i,y^i( r)))\to m([x^i,y^i))\); fixing \(r\) with all these gaps \(<\varepsilon/N\), the sets \(A_\varepsilon:=\bigcup_i[x^i,y^i( r))\) and \(K_\varepsilon:=\bigcup_i\prod_k[x^i_k,y^i_k( r)]\) satisfy \(A_\varepsilon\subset K_\varepsilon\subset A\) and \(m(A\setminus A_\varepsilon)<\varepsilon\).
By Theorem 1.5.6(iii), \(m\) has a unique nonnegative countably additive extension \(\mu_j\) to \(\sigma(\mathcal{A}_j)=\mathcal{B}(Q_j)\); for \(i>j\) the restriction of \(\mu_i\) to \(\mathcal{B}(Q_j)\) is such an extension of \(m|_{\mathcal{A}_j}\), hence equals \(\mu_j\). With \(R_1:=Q_1\), \(R_j:=Q_j\setminus Q_{j-1}\), set
\begin{equation*} \mu(B):=\sum_{j=1}^{\infty}\mu_j(B\cap R_j),\qquad B\in\mathcal{B}(\mathbb{R}^n), \end{equation*}
countably additive since double series with nonnegative terms may be summed in either order; by consistency \(\mu=\mu_i\) on Borel subsets of \(Q_i\), so \(\mu([x,y))=F[x,y)\) for every box.
Finally \(Q_j\uparrow\mathbb{R}^n\) gives \(\mu(\mathbb{R}^n)=\lim_j F[-j\mathbf{1},j\mathbf{1})=1\), since the corner \(j\mathbf{1}\) has sign \(+1\) with \(F(j\mathbf{1})\to1\) while every other corner has a coordinate \(-j\) and contributes \(0\), both by 3’); and for \(j>\max_k|x_k|\) we have \(S_x\cap Q_j=[-j\mathbf{1},x)\uparrow S_x\), whence by the same cancellation
\begin{equation*} \mu(S_x)=\lim_{j\to\infty}F[-j\mathbf{1},x)=F(x). \end{equation*}
Let \(\mathcal{A}\) be a \(\sigma\)-algebra of subsets of \(\mathbb{N}\). Show that \(\mathcal{A}\) is generated by some finite or countable partition of \(\mathbb{N}\) into disjoint sets, so that every element of \(\mathcal{A}\) is an at most countable union of elements of this partition.
The partition is by the classes \(K(k):=\{m:\ m\sim k\}\) of the equivalence relation
\begin{equation*} n\sim m\quad :\Longleftrightarrow\quad \bigl(\ \forall A\in\mathcal{A}:\ n\in A\iff m\in A\ \bigr) \end{equation*}
(reflexive, symmetric and transitive by inspection). Each \(K(k)\) lies in \(\mathcal{A}\): if \(K(k)\ne\mathbb{N}\), then for every \(n\notin K(k)\) some set of \(\mathcal{A}\) separates \(n\) from \(k\), and replacing it by its complement if necessary gives \(A_n\in\mathcal{A}\) with \(k\in A_n\), \(n\notin A_n\), so
\begin{equation*} A:=\bigcap_{n\in\mathbb{N}\setminus K(k)}A_n \in\mathcal{A} \end{equation*}
as an at most countable intersection, and \(A=K(k)\): if \(m\sim k\) then \(m\in A_n\) for every \(n\), while \(l\in A\setminus K(k)\) is itself one of the indices and \(l\notin A_l\supset A\).
The classes are nonempty and pairwise disjoint in \(\mathbb{N}\), hence at most countably many, say \(M_1,M_2,\dots\in\mathcal{A}\). Every \(A\in\mathcal{A}\) is a union of classes, since \(n\in A\) and \(m\sim n\) force \(m\in A\) by the definition of \(\sim\); thus \(A=\bigcup\{M_i: M_i\subset A\}\), an at most countable union of members of the partition, and conversely every union \(\bigcup_{i\in S}M_i\) lies in \(\mathcal{A}\). So \(\mathcal{A}\) consists exactly of the unions of members of \(\{M_i\}\), a family that is a \(\sigma\)-algebra containing each \(M_i\) and contained in every \(\sigma\)-algebra doing so, i.e. \(\mathcal{A}=\sigma(\{M_i\})\).
(i) Let \(\mathcal{A}\) be a \(\sigma\)-algebra of subsets of \(\mathbb{N}\) and let \(\mu\) be a probability measure on \(\mathcal{A}\). Show that \(\mu\) extends to a probability measure on the class of all subsets of \(\mathbb{N}\).
(ii) Let \(\mathcal{A}\) be the \(\sigma\)-algebra generated by singletons of a set \(X\) and let \(\mathcal{A}_0\) be its sub-\(\sigma\)-algebra. Show that any measure \(\mu\) on \(\mathcal{A}_0\) extends to a measure on \(\mathcal{A}\).
(i) Take the purely atomic \(\nu:=\sum_{i\in I}c_i\delta_{n_i}\), i.e.
\begin{equation*} \nu(B):=\sum_{i:\,n_i\in B}c_i,\qquad B\subset\mathbb{N}, \end{equation*}
where \(\{M_i\}_{i\in I}\) is the at most countable partition of \(\mathbb{N}\) into nonempty sets of \(\mathcal{A}\) with \(\mathcal{A}=\{\bigcup_{i\in S}M_i: S\subset I\}\) given by Exercise 1.12.157, \(c_i:=\mu(M_i)\ge0\) and \(n_i\in M_i\) is any point. Then \(\nu\ge0\) and \(\nu(\mathbb{N})=\sum_i c_i=\mu(\mathbb{N})=1\) by countable additivity, and \(\nu\) is countably additive on all subsets: for disjoint \(B_r\) with union \(B\),
\begin{equation*} \nu(B)=\sum_{i\in I}c_i\sum_r \mathbf{1}_{B_r}(n_i) =\sum_r\sum_{i\in I}c_i\mathbf{1}_{B_r}(n_i)=\sum_r\nu(B_r), \end{equation*}
the interchange being legitimate for nonnegative double series. Finally \(n_i\in\bigcup_{i\in S}M_i\) exactly for \(i\in S\), so \(\nu(A)=\sum_{i\in S}\mu(M_i)=\mu(A)\) on \(\mathcal{A}\). The argument uses only that the space is at most countable and, after normalization, applies to any finite nonnegative measure.
(ii) Here \(\mathcal{A}=\{A\subset X: A\ \text{or}\ X\setminus A\ \text{is at most countable}\}\) by Example 1.2.12, and by the Jordan decomposition it suffices to extend a finite nonnegative \(\mu\) (extend the positive and negative parts and subtract). If \(X\) is at most countable, then \(\mathcal{A}=2^X\) and (i) applies, so let \(X\) be uncountable; note \(\mathcal{A}_0\subset\mathcal{A}\), so each of its members is at most countable or co-countable.
Split off the countable part: with \(s:=\sup\{\mu(A): A\in\mathcal{A}_0\ \text{at most countable}\}\) and at most countable \(A_r\in\mathcal{A}_0\) with \(\mu(A_r)\to s\), the set \(A_\infty:=\bigcup_r A_r\in\mathcal{A}_0\) is at most countable with \(\mu(A_\infty)\ge\mu(A_r)\) for all \(r\), hence \(\mu(A_\infty)=s\). Put \(Y:=X\setminus A_\infty\in\mathcal{A}_0\), uncountable. If \(A\in\mathcal{A}_0\), \(A\subset Y\) is at most countable, then \(A_\infty\cup A\) is an admissible competitor, so \(s\ge s+\mu(A)\) and \(\mu(A)=0\); if instead such an \(A\) is co-countable in \(X\), then \(Y\setminus A\in\mathcal{A}_0\) is at most countable inside \(Y\), so \(\mu(A)=\mu(Y)\). Hence, with
\begin{equation*} \rho(B):=\begin{cases}0,& B\cap Y\ \text{at most countable},\\ 1,& Y\setminus B\ \text{at most countable},\end{cases} \end{equation*}
well defined on \(\mathcal{A}\) (for \(B\in\mathcal{A}\) one of \(B\cap Y\), \(Y\setminus B\subset X\setminus B\) is at most countable, and not both as \(Y\) is uncountable), we have the trace identity \(\mu(A)=\mu(Y)\rho(A)\) for every \(A\in\mathcal{A}_0\) with \(A\subset Y\).
\(\rho\) is countably additive on \(\mathcal{A}\): for disjoint \(B_r\in\mathcal{A}\) with union \(B\), if \(\rho(B)=0\) then every \(B_r\cap Y\) is at most countable and all \(\rho(B_r)=0\); if \(\rho(B)=1\) then \(B\cap Y\) is uncountable, so not all \(B_r\cap Y\) are at most countable, and exactly one \(B_{r_0}\) has \(Y\setminus B_{r_0}\) at most countable (two disjoint co-countable subsets of the uncountable \(Y\) cannot exist), while \(B_r\cap Y\subset Y\setminus B_{r_0}\) for \(r\ne r_0\).
By (i) applied on the at most countable set \(A_\infty\) to the finite nonnegative measure \(A\mapsto\mu(A\cap A_\infty)\) on the trace \(\sigma\)-algebra \(\{A\cap A_\infty: A\in\mathcal{A}_0\}\) (normalizing when \(\mu(A_\infty)>0\), and taking the zero measure otherwise), there is a finite nonnegative \(\mu_1\) on all subsets of \(A_\infty\) agreeing with \(\mu\) there. Then
\begin{equation*} \nu(B):=\mu_1(B\cap A_\infty)+\mu(Y)\,\rho(B),\qquad B\in\mathcal{A}, \end{equation*}
is nonnegative and countably additive, and for \(A\in\mathcal{A}_0\), using \(\rho(A)=\rho(A\cap Y)\) and the trace identity,
\begin{equation*} \nu(A)=\mu(A\cap A_\infty)+\mu(Y)\rho(A\cap Y)=\mu(A\cap A_\infty)+\mu(A\cap Y)=\mu(A). \end{equation*}
Exercises 1.12.159–1.12.160
Let \(\mu\) be a countably additive measure with values in \([0,+\infty]\) on a ring \(\mathfrak{X}\) of subsets of a space \(X\).
(i) Suppose that \(\mu\) is \(\sigma\)-finite, i.e., \(X=\bigcup_{n=1}^{\infty} X_n\), where one has \(X_n\in\mathfrak{X}\) and \(\mu(X_n)<\infty\). Show that \(\mu\) has a unique countably additive extension to the \(\sigma\)-ring \(\Sigma(\mathfrak{X})\) generated by \(\mathfrak{X}\).
(ii) Suppose that the measure \(\mathfrak{m}:=\mu^{\ast}\) is \(\sigma\)-finite on \(\mathfrak{X}_{\mathfrak{m}}\). Show that it is a unique extension of \(\mu\) to \(\sigma(\mathfrak{X})\).
In both parts the extension is the restriction of \(\mathfrak{m}:=\mu^{\ast}\), the Method I outer measure of Example 1.11.5,
\begin{equation*} \mathfrak{m}(A)=\inf\Big\{\sum_{n=1}^{\infty}\mu(X_n):\ X_n\in\mathfrak{X}, \ A\subset\bigcup_{n=1}^{\infty}X_n\Big\} \end{equation*}
(\(+\infty\) if no such covering exists): since \(\mathfrak{X}\) is a ring and \(\mu\) is countably additive, Theorem 1.11.8 makes every set of \(\mathfrak{X}\) \(\mathfrak{m}\)-measurable with \(\mathfrak{m}=\mu\) on \(\mathfrak{X}\), and by Theorem 1.11.4(iii) the Caratheodory class \(\mathcal{M}_{\mathfrak{m}}\) is a \(\sigma\)-algebra carrying \(\mathfrak{m}\) countably additively, so \(\Sigma(\mathfrak{X})\subset\sigma(\mathfrak{X})\subset\mathcal{M}_{\mathfrak{m}}\) and existence is Corollary 1.11.9. The content is uniqueness. We use twice that a countably additive \([0,+\infty]\)-valued measure on a \(\sigma\)-ring is continuous along increasing sequences (disjointify by \(D_n:=E_n\setminus E_{n-1}\); Check!).
(i) Replacing \(X_n\) by \(X_1\cup\cdots\cup X_n\), which lies in \(\mathfrak{X}\) and has finite measure by finite subadditivity, assume \(X_n\uparrow X\). Let \(\nu\) be any countably additive \([0,+\infty]\)-valued measure on \(\Sigma(\mathfrak{X})\) with \(\nu=\mu\) on \(\mathfrak{X}\). Fix \(n\) and regard \(X_n\) as the space; then
\begin{equation*} \mathfrak{X}_{(n)}:=\{A\cap X_n:\ A\in\mathfrak{X}\}=\{A\in\mathfrak{X}:\ A\subset X_n\} \end{equation*}
(the descriptions agree as \(\mathfrak{X}\) is closed under intersections) is an algebra of subsets of \(X_n\), being closed under finite unions and, \(\mathfrak{X}\) being a ring, under \(B\mapsto X_n\setminus B\). Put \(\mathcal{A}_n:=\sigma(\mathfrak{X}_{(n)})\) in \(X_n\). Then \(\mathcal{A}_n\subset\Sigma(\mathfrak{X})\), because \(\{B\subset X_n: B\in\Sigma(\mathfrak{X})\}\) contains \(\mathfrak{X}_{(n)}\) and is a \(\sigma\)-algebra in \(X_n\) (complements inside \(X_n\) are differences from \(X_n\in\Sigma(\mathfrak{X})\)); and \(E\cap X_k\in\mathcal{A}_k\) for all \(E\in\Sigma(\mathfrak{X})\) and all \(k\), because
\begin{equation*} \mathcal{E}:=\{E\subset X:\ E\cap X_k\in\mathcal{A}_k\ \text{for all}\ k\} \end{equation*}
contains \(\mathfrak{X}\) and is a \(\sigma\)-ring, as \((E\setminus F)\cap X_k=(E\cap X_k)\setminus(F\cap X_k)\) and \((\bigcup_jE_j)\cap X_k=\bigcup_j(E_j\cap X_k)\).
Hence \(\mathfrak{m}\) and \(\nu\) are both defined, nonnegative and countably additive on \(\mathcal{A}_n\), both agree with \(\mu\) on \(\mathfrak{X}_{(n)}\), and both are finite there: for \(B\in\mathcal{A}_n\) additivity along \(B, X_n\setminus B\in\Sigma(\mathfrak{X})\) gives \(\nu(B)\le\nu(X_n)=\mu(X_n)<\infty\) and likewise for \(\mathfrak{m}\). As \(\mu|_{\mathfrak{X}_{(n)}}\) is a finite nonnegative countably additive set function on the algebra \(\mathfrak{X}_{(n)}\), Theorem 1.5.6(iii) gives it exactly one nonnegative countably additive extension to \(\mathcal{A}_n\), so \(\mathfrak{m}=\nu\) on \(\mathcal{A}_n\). For \(E\in\Sigma(\mathfrak{X})\) the sets \(E\cap X_n\in\mathcal{A}_n\) increase to \(E\) inside \(\Sigma(\mathfrak{X})\), whence by continuity
\begin{equation*} \mathfrak{m}(E)=\lim_{n\to\infty}\mathfrak{m}(E\cap X_n) =\lim_{n\to\infty}\nu(E\cap X_n)=\nu(E). \end{equation*}
As a byproduct \(X=\bigcup_nX_n\in\Sigma(\mathfrak{X})\), and a \(\sigma\)-ring containing the whole space is a \(\sigma\)-algebra, so \(\Sigma(\mathfrak{X})=\sigma(\mathfrak{X})\) in case (i).
(ii) The hypothesis of (ii) implies that of (i): if \(X=\bigcup_nE_n\) with \(\mathfrak{m}(E_n)<\infty\), then by the definition of \(\mathfrak{m}\) each \(E_n\) admits \(X_{n,k}\in\mathfrak{X}\) with \(E_n\subset\bigcup_kX_{n,k}\) and \(\sum_k\mu(X_{n,k})<\infty\), so in particular \(\mu(X_{n,k})<\infty\) and
\begin{equation*} X=\bigcup_{n=1}^{\infty}E_n\subset\bigcup_{n,k}X_{n,k}\subset X \end{equation*}
exhibits \(X\) as a countable union of \(\mathfrak{X}\)-sets of finite measure. (Only \(X=\bigcup_nE_n\) and \(\mathfrak{m}(E_n)<\infty\) are used, so it is immaterial whether the \(E_n\) are taken in \(\mathfrak{X}_{\mathfrak{m}}\) or in \(\mathcal{M}_{\mathfrak{m}}\).) By the byproduct \(\Sigma(\mathfrak{X})=\sigma(\mathfrak{X})\), and by (i) the restriction of \(\mathfrak{m}\) — a countably additive extension of \(\mu\) by Corollary 1.11.9 — is the only one on \(\sigma(\mathfrak{X})\).
Two sets \(A\) and \(B\) on the real line are called metrically separated if, for every \(\varepsilon>0\), there exist open sets \(A_{\varepsilon}\) and \(B_{\varepsilon}\) such that \(A\subset A_{\varepsilon}\) and \(B\subset B_{\varepsilon}\) with \(\lambda(A_{\varepsilon}\cap B_{\varepsilon})<\varepsilon\), where \(\lambda\) is Lebesgue measure.
(i) Show that if sets \(A\) and \(B\) are metrically separated, then there exist Borel sets \(A_0\) and \(B_0\) such that \(A\subset A_0\) and \(B\subset B_0\) with \(\lambda(A_0\cap B_0)=0\).
(ii) Let \(A\) be a Lebesgue measurable set on the real line and let \(A=A_1\cup A_2\), where the sets \(A_1\) and \(A_2\) are metrically separated. Show that \(A_1\) and \(A_2\) are Lebesgue measurable.
(i) Take \(A_0:=\bigcap_{n=1}^{\infty}A_n\) and \(B_0:=\bigcap_{n=1}^{\infty}B_n\), where metric separation with \(\varepsilon=1/n\) supplies open sets \(A_n\supset A\), \(B_n\supset B\) with
\begin{equation*} \lambda(A_n\cap B_n)<\frac{1}{n}. \end{equation*}
These are \(G_\delta\), hence Borel, they contain \(A\) and \(B\), and \(A_0\cap B_0\subset A_n\cap B_n\) gives \(\lambda(A_0\cap B_0)<1/n\) for every \(n\), i.e. \(\lambda(A_0\cap B_0)=0\).
(ii) With \(B_1\supset A_1\), \(B_2\supset A_2\) Borel and \(\lambda(B_1\cap B_2)=0\) from (i), and \(E:=A\cap(B_1\setminus A_1)\),
\begin{equation*} A_1=(A\cap B_1)\setminus E \end{equation*}
is Lebesgue measurable. Indeed \(E\subset B_1\cap B_2\), since \(x\in E\) lies in \(A=A_1\cup A_2\) but not in \(A_1\), hence in \(A_2\subset B_2\); so \(\lambda^{\ast}(E)=0\) and \(E\) is measurable with \(\lambda(E)=0\) by completeness of Lebesgue measure. As \(A_1\subset A\cap B_1\), splitting off \(A_1\) gives \(A\cap B_1=A_1\cup E\) disjointly, which is the displayed formula, and \(A\), \(B_1\), \(E\) all lie in the \(\sigma\)-algebra of Lebesgue measurable sets. Symmetrically \(F:=A\cap(B_2\setminus A_2)\subset B_1\cap B_2\) is null and \(A_2=(A\cap B_2)\setminus F\) is measurable.
The Lebesgue Integral
Exercises 2.12.25–2.12.31
(\(\circ\)) Let \((X,\mathcal{A})\), \((Y,\mathcal{B})\), and \((Z,\mathcal{E})\) be measurable spaces. Suppose that a mapping \(f\colon X \to Y\) is \((\mathcal{A},\mathcal{B})\)-measurable and a mapping \(g\colon Y \to Z\) is \((\mathcal{B},\mathcal{E})\)-measurable. Show that the composition \(g\circ f\colon X \to Z\) is \((\mathcal{A},\mathcal{E})\)-measurable.
For \(E\in\mathcal{E}\) the elementary identity \((g\circ f)^{-1}(E)=f^{-1}\bigl(g^{-1}(E)\bigr)\) gives
\begin{equation*} (g\circ f)^{-1}(E) = f^{-1}\bigl(g^{-1}(E)\bigr) \in \mathcal{A}, \end{equation*}
since \(g^{-1}(E)\in\mathcal{B}\) by measurability of \(g\) and \(f^{-1}(\mathcal{B})\subset\mathcal{A}\) by that of \(f\); by Definition 2.1.3 this is the assertion.
(\(\circ\)) Suppose that measurable functions \(f_n\) on \([0,1]\) converge almost everywhere to zero. Show that there exist numbers \(C_n > 0\) such that \(\lim\limits_{n\to\infty} C_n = \infty\), but the sequence \(C_n f_n\) converges almost everywhere to zero.
Take \(C_n:=2^{k}\) for \(n_k\le n<n_{k+1}\) (and \(C_n:=1\) for \(n<n_1\)), where the blocks come from Egoroff’s theorem as follows. Multiplying each \(f_n\) by the indicator of the complement of the null set where some \(f_n\) is undefined or \(f_n(x)\not\to0\) changes nothing almost everywhere, so assume \(f_n\to0\) everywhere. By Egoroff’s theorem (Theorem 2.2.1) applied with \(\varepsilon=1/k\) — legitimate since \(\lambda\) is finite on \([0,1]\) — there are measurable \(Y_k\) with \(\lambda([0,1]\setminus Y_k)<1/k\) carrying uniform convergence, and
\begin{equation*} X_k := Y_1 \cup \dots \cup Y_k \uparrow X_\infty := \bigcup_{k=1}^{\infty} X_k \end{equation*}
again carries uniform convergence (the supremum over a finite union is the largest of the finitely many suprema), with \(\lambda([0,1]\setminus X_k)<1/k\) and hence \(\lambda([0,1]\setminus X_\infty)=0\). Thus \(a_{n,k}:=\sup_{x\in X_k}|f_n(x)|\to0\) as \(n\to\infty\) for each \(k\), so we may pick \(n_1<n_2<\cdots\) with
\begin{equation*} a_{n,k} \le 4^{-k} \qquad \text{for all } n \ge n_k . \end{equation*}
These \(C_n\) are positive and \(C_n\to\infty\), the block index of \(n\) tending to infinity. For \(x\in X_m\) and \(n\ge n_m\) with block index \(k=k(n)\ge m\) we have \(x\in X_m\subset X_k\) and \(n\ge n_k\), so
\begin{equation*} C_n |f_n(x)| \le 2^{k}a_{n,k} \le 2^{k}4^{-k} = 2^{-k(n)} \longrightarrow 0 , \end{equation*}
i.e. \(C_nf_n\to0\) on the full-measure set \(X_\infty\).
(\(\circ\)) Suppose that measurable functions \(f_n\) on \([0,1]\) converge almost everywhere to zero. Prove that there exist numbers \(\varepsilon_n > 0\) and a measurable finite function \(g\) such that \(\lim\limits_{n\to\infty}\varepsilon_n = 0\) and \(|f_n(x)| \le \varepsilon_n g(x)\) almost everywhere for every \(n\).
Take \(\varepsilon_n:=1/C_n\) and \(g:=\sup_n C_n|f_n|\) off a null set, with \(C_n>0\), \(C_n\to\infty\) and \(C_nf_n\to0\) a.e. supplied by Exercise 2.12.26. Precisely, let \(N\) be the null set where some \(f_n\) is undefined or \(C_nf_n(x)\not\to0\) (measurable by completeness of Lebesgue measure), and put
\begin{equation*} h(x) := \sup_{n \ge 1} C_n |f_n(x)| \in [0,+\infty], \qquad g := h\cdot I_{[0,1]\setminus N} . \end{equation*}
Then \(\varepsilon_n\to0\), and \(h\) is measurable as a countable supremum of measurable functions (Theorem 2.1.5(vi)), so \(g\) is measurable; \(g\) is finite everywhere, since for \(x\notin N\) the convergent sequence \(\bigl(C_n|f_n(x)|\bigr)_n\) is bounded. Finally, for \(x\notin N\) the definition of the supremum gives \(g(x)=h(x)\ge C_n|f_n(x)|\), i.e.
\begin{equation*} |f_n(x)| = \varepsilon_n \cdot C_n |f_n(x)| \le \varepsilon_n\, g(x) , \end{equation*}
which holds for every \(n\) outside the single null set \(N\).
(\(\circ\)) Construct a measurable set in \([0,1]\) such that every function on \([0,1]\) that almost everywhere equals its indicator function is discontinuous almost everywhere (and is not Riemann integrable, see Exercise 2.12.38).
Take \(E:=\bigcup_{n=1}^{\infty}K_n\), a Borel set, where \(K_n,L_n\) are the sets constructed below; it satisfies
\begin{equation*} \lambda(E \cap I) > 0 \quad\text{and}\quad \lambda(I \setminus E) > 0 \qquad \text{for every nondegenerate interval } I \subset [0,1], \tag{\(\ast\)} \end{equation*}
and \((\ast)\) forces every function equal to \(I_E\) a.e. to be discontinuous at every point.
Every nondegenerate \([a,b]\subset[0,1]\) contains a compact nowhere dense set of positive measure: enumerating the rationals of \([a,b]\) as \(q_1,q_2,\dots\) and fixing \(0<\eta<(b-a)/4\), the set
\begin{equation*} K := [a,b]\setminus U, \qquad U := \bigcup_{j=1}^{\infty} \bigl(q_j - \eta 2^{-j},\, q_j + \eta 2^{-j}\bigr), \end{equation*}
is compact with \(\lambda(K)\ge(b-a)-2\eta>0\), and it contains no interval (every interval meets \(U\)), hence is nowhere dense.
Let \(I_1,I_2,\dots\) enumerate the intervals \([p,q]\subset[0,1]\) with rational \(p<q\), and construct \(K_n,L_n\) inductively. Given \(K_1,L_1,\dots,K_{n-1},L_{n-1}\), their union \(F_n\) is closed and nowhere dense, so \(\operatorname{int}(I_n)\setminus F_n\) is open and nonempty and contains a nondegenerate closed interval \(J\); split \(J\) into disjoint nondegenerate closed \(J^{\prime},J^{\prime\prime}\) and use the previous paragraph to take compact nowhere dense \(K_n\subset J^{\prime}\) and \(L_n\subset J^{\prime\prime}\) of positive measure. These lie in \(I_n\) and avoid \(F_n\), so the whole family is pairwise disjoint (Check!).
Then \((\ast)\) holds: a nondegenerate \(I\) contains some \(I_n\), whence \(\lambda(E\cap I)\ge\lambda(K_n)>0\), while \(L_n\subset I_n\subset I\) is disjoint from every \(K_m\) and so \(\lambda(I\setminus E)\ge\lambda(L_n)>0\).
Let \(h=I_E\) off a null set \(Z\) and suppose \(h\) were continuous at some \(x_0\), say \(|h(y)-h(x_0)|<\tfrac12\) on \(I:=(x_0-\delta,x_0+\delta)\cap[0,1]\). Since \(I\) contains a nondegenerate interval, \((\ast)\) and \(\lambda(Z)=0\) allow \(y_1\in(E\cap I)\setminus Z\) and \(y_2\in(I\setminus E)\setminus Z\), for which \(h(y_1)=1\), \(h(y_2)=0\) and
\begin{equation*} 1 = |h(y_1) - h(y_2)| \le |h(y_1) - h(x_0)| + |h(x_0) - h(y_2)| < \tfrac12 + \tfrac12 = 1 . \end{equation*}
So \(h\) is everywhere discontinuous, in particular discontinuous almost everywhere; and \(h\) is not Riemann integrable, being either unbounded or, if bounded, continuous nowhere, whereas Lebesgue’s criterion (Exercise 2.12.38) would require continuity almost everywhere.
(\(\circ\)) Suppose that functions \(f\) and \(g\) are measurable with respect to a \(\sigma\)-algebra \(\mathcal{A}\) and that a function \(\Psi\) on the plane is continuous on the set of values of the mapping \((f,g)\). Show that the function \(\Psi(f,g)\) is measurable with respect to \(\mathcal{A}\).
For \(F:=(f,g)\colon X\to\mathbb{R}^2\) with range \(Y:=F(X)\) and any \(c\in\mathbb{R}^1\),
\begin{equation*} \bigl\{ x : \Psi(f(x), g(x)) < c \bigr\} = F^{-1}(Y\cap U) = F^{-1}(U) \in \mathcal{A}, \end{equation*}
where \(U\subset\mathbb{R}^2\) is open with \(Y\cap U=\{y\in Y:\Psi(y)<c\}\) — such a \(U\) exists because \(\Psi|_Y\) is continuous, so this set is open in the topology induced on \(Y\) — and the middle equality holds since \(F\) takes all its values in \(Y\). By Definition 2.1.1 this is the assertion.
It remains to see that \(F\) is \((\mathcal{A},\mathcal{B}(\mathbb{R}^2))\)-measurable. The family \(\mathcal{E}:=\{B\in\mathcal{B}(\mathbb{R}^2): F^{-1}(B)\in\mathcal{A}\}\) is a \(\sigma\)-algebra, by Lemma 1.2.8, and it contains every open rectangle \(R=(a,b)\times(c,d)\), because
\begin{equation*} F^{-1}( R) = \{a < f < b\} \cap \{c < g < d\} \in \mathcal{A} \end{equation*}
via \(\{a<f<b\}=\{f<b\}\setminus\bigcap_{k}\{f<a+1/k\}\) and likewise for \(g\). Every open \(U\subset\mathbb{R}^2\) is a countable union of rational open rectangles contained in it, so \(\mathcal{E}\) contains all open sets and hence equals \(\mathcal{B}(\mathbb{R}^2)\).
(\(\circ\)) Let \(\mathcal{A}\) be the \(\sigma\)-algebra generated by all singletons in a space \(X\). Prove that a function \(f\) is measurable with respect to \(\mathcal{A}\) if and only if it is constant on the complement of some at most countable set.
Here \(\mathcal{A}=\{A\subset X: A\ \text{or}\ X\setminus A\ \text{is at most countable}\}\) by Example 1.2.12, which makes sufficiency immediate: if \(f\equiv c\) off an at most countable \(S\), then for \(t\le c\) the set \(A_t:=\{f<t\}\) is contained in \(S\), and for \(t>c\) its complement is, so \(A_t\in\mathcal{A}\) for every \(t\) and Definition 2.1.1 applies.
Conversely let \(f\) be \(\mathcal{A}\)-measurable and real-valued; we may assume \(X\) uncountable, else \(S:=X\) serves. Each \(A_t\in\mathcal{A}\) is then at most countable or co-countable and not both, so
\begin{equation*} T := \{ t \in \mathbb{R}^1 : A_t \text{ is at most countable} \} \end{equation*}
is a down-set (\(A_t\subset A_s\) for \(t<s\)). It is nonempty, since otherwise every \(\{f\ge -n\}\) would be at most countable and \(X=\bigcup_n\{f\ge-n\}\) (values being finite) would be too; and \(T\ne\mathbb{R}^1\), since otherwise \(X=\bigcup_n\{f<n\}\) would be at most countable. Hence any point of \(\mathbb{R}^1\setminus T\) bounds \(T\) above, and we may put \(c:=\sup T\), noting that every \(t<c\) lies in \(T\).
Now \(S:=\{f\ne c\}\) is at most countable: choosing \(t_n\uparrow c\) with \(t_n<c\) and \(s_n\downarrow c\) with \(s_n>c\), the sets \(\{f<t_n\}\) are at most countable because \(t_n\in T\), and the sets \(\{f\ge s_n\}=X\setminus A_{s_n}\) are at most countable because \(s_n\notin T\), while
\begin{equation*} \{ f < c \} = \bigcup_{n=1}^{\infty} \{ f < t_n \}, \qquad \{ f > c\} = \bigcup_{n=1}^{\infty} \{ f \ge s_n \} \end{equation*}
(if \(f(x)>c\) then \(f(x)\ge s_n\) for large \(n\)). Thus \(f\equiv c\) on \(X\setminus S\).
(\(\circ\)) Let \(\mu\) be a probability measure, let \(\{c_\alpha\}\) be a family of real numbers, and let \(f\) be a \(\mu\)-measurable function. Show that
\begin{equation*} \mu\bigl\{ x : f(x) \ge \sup_\alpha c_\alpha \bigr\} \ \ge\ \inf_\alpha\, \mu\bigl\{ x : f(x) \ge c_\alpha \bigr\}. \end{equation*}
Extract from the (nonempty) family a nondecreasing sequence \(c_{\alpha_n}\to c:=\sup_\alpha c_\alpha\) and apply continuity from above. Such indices exist: pick \(\beta_n\) with \(c_{\beta_n}>c-1/n\) if \(c<\infty\) and with \(c_{\beta_n}>n\) if \(c=+\infty\), then replace \(\beta_n\) by an index realizing \(\max(c_{\beta_1},\dots,c_{\beta_n})\). The sets \(A_n:=\{f\ge c_{\alpha_n}\}\) are \(\mu\)-measurable and decrease, with
\begin{equation*} \bigcap_{n=1}^{\infty} A_n = \{ x : f(x) \ge c \} , \end{equation*}
since \(c_{\alpha_n}\le c\) gives \(\supset\) and passing to the limit in \(f(x)\ge c_{\alpha_n}\) gives \(\subset\). As \(\mu\) is a probability measure, continuity from above applies, and \(\mu(A_n)\ge\inf_\alpha\mu(\{f\ge c_\alpha\})\) for every \(n\) by the definition of the infimum, so
\begin{equation*} \mu\bigl( \{f \ge c\} \bigr) = \lim_{n\to\infty}\mu(A_n) \ \ge\ \inf_\alpha \mu\bigl(\{f \ge c_\alpha\}\bigr). \end{equation*}
Exercises 2.12.32–2.12.38
(Davies) Let \(\mu\) be a finite nonnegative measure on a space \(X\). Prove that a function \(f\colon X \to \mathbb{R}^1\) is measurable with respect to \(\mu\) precisely when for each \(\mu\)-measurable set \(A\) with \(\mu(A) > 0\) and each \(\varepsilon > 0\), there exists a \(\mu\)-measurable set \(B \subset A\) such that \(\mu(B) > 0\) and
\begin{equation*} \sup_{x,y \in B} |f(x) - f(y)| \le \varepsilon . \end{equation*}
Necessity: given \(\mu(A)>0\) and \(\varepsilon>0\), the \(\mu\)-measurable sets
\begin{equation*} A_n := \{ x \in A \colon\ n\varepsilon \le f(x) < (n+1)\varepsilon \}, \qquad n \in \mathbb{Z}, \end{equation*}
are disjoint with union \(A\) (\(f\) being finite), so \(\sum_n\mu(A_n)=\mu(A)>0\) forces \(\mu(B)>0\) for \(B:=A_{n_0}\) with some \(n_0\), and \(|f(x)-f(y)|<\varepsilon\) on \(B\).
Sufficiency: assume the condition \((D)\); we may assume \(\mu(X)>0\), since otherwise every set lies in the complete \(\sigma\)-algebra \(\mathcal{A}_\mu\) and nothing is to prove. Fix \(\varepsilon>0\) and call \(B\) admissible in \(E\) if \(B\in\mathcal{A}_\mu\), \(B\subset E\), \(\mu(B)>0\) and \(\sup_{x,y\in B}|f(x)-f(y)|\le\varepsilon\); by \((D)\) admissible sets exist in every \(E\) of positive measure. Greedily choose disjoint \(B_1,B_2,\dots\): with \(E_n:=X\setminus(B_1\cup\cdots\cup B_{n-1})\), stop if \(\mu(E_n)=0\) and otherwise put
\begin{equation*} \delta_n := \sup\{ \mu(B) \colon B \text{ admissible in } E_n \}, \end{equation*}
picking \(B_n\) admissible in \(E_n\) with \(\mu(B_n)>\delta_n/2\).
Then \(\mu(N_\varepsilon)=0\) for \(N_\varepsilon:=X\setminus\bigcup_nB_n\). This is the stopping rule if the construction halts; otherwise \(\sum_n\mu(B_n)\le\mu(X)<\infty\) gives \(\mu(B_n)\to0\), hence \(\delta_n<2\mu(B_n)\to0\), while a set \(E\subset N_\varepsilon\) supplied by \((D)\) with \(\delta:=\mu(E)>0\) is admissible in every \(E_k\) (as \(E\subset N_\varepsilon\subset E_k\)), forcing \(\delta\le\delta_k\) for all \(k\).
Picking \(x_n\in B_n\ne\emptyset\), the function \(g_\varepsilon:=f(x_n)\) on \(B_n\) and \(:=0\) on \(N_\varepsilon\) takes countably many values with preimages that are countable unions of the \(B_n\) and possibly \(N_\varepsilon\), so it is \(\mu\)-measurable, and \(x,x_n\in B_n\) gives
\begin{equation*} |g_\varepsilon - f| \le \varepsilon \ \text{ on } X \setminus N_\varepsilon, \qquad \mu(N_\varepsilon)=0 . \end{equation*}
Apply this with \(\varepsilon=1/k\) and put \(N:=\bigcup_kN_{1/k}\), a \(\mu\)-null set; then \(g_{1/k}\to f\) uniformly on \(X\setminus N\), so \(g_{1/k}I_{X\setminus N}\to fI_{X\setminus N}\) pointwise on \(X\) and the limit is \(\mathcal{A}_\mu\)-measurable by Theorem 2.1.5. Finally \(\{f<c\}\) differs from \(\{fI_{X\setminus N}<c\}\cap(X\setminus N)\) by a subset of \(N\), which lies in \(\mathcal{A}_\mu\) by completeness, so \(f\) is \(\mu\)-measurable.
(\(\circ\)) (M. Frechet) Suppose that a sequence of measurable functions \(f_n\) on a probability space \((X,\mathcal{A},\mu)\) converges a.e. to a function \(f\) and, for every \(n\), there is a sequence of measurable functions \(f_{n,m}\) a.e. convergent to \(f_n\). Prove that there exist subsequences \(\{n_k\}\) and \(\{m_k\}\) such that \(f_{n_k,m_k} \to f\) a.e.
Choose \(m_n\) diagonally so that \(h_n:=f_{n,m_n}\to f\) in measure, then extract an a.e. convergent subsequence by the Riesz theorem. We may take \(f\) measurable, replacing it by \(\limsup_nf_n\) on the measurable set \(C\) where \(\{f_n(x)\}\) is Cauchy, and by \(0\) off \(C\); this changes \(f\) only on a null set, since \(\mu( C)=1\). As \(\mu\) is finite, Theorem 2.2.3 turns the two hypotheses into \(f_n\to f\) in measure and \(f_{n,m}\to f_n\) in measure as \(m\to\infty\) for each fixed \(n\).
Hence, as in Remark 2.2.7, we may choose inductively \(m_1<m_2<\cdots\) with \(m_n>\max(m_{n-1},n)\) and
\begin{equation*} \mu\bigl(\{ x \colon |f_{n,m_n}(x) - f_n(x)| \ge 2^{-n}\}\bigr) \le 2^{-n} . \end{equation*}
Fix \(c>0\). Then \(\{|h_n-f|\ge c\}\subset\{|f_{n,m_n}-f_n|\ge c/2\}\cup\{|f_n-f|\ge c/2\}\), and as soon as \(2^{-n}<c/2\) the first set has measure at most \(2^{-n}\), so
\begin{equation*} \mu\bigl(\{ |h_n - f| \ge c\}\bigr) \le 2^{-n} + \mu\bigl(\{|f_n - f| \ge c/2\}\bigr) \xrightarrow[n\to\infty]{} 0 , \end{equation*}
i.e. \(h_n\to f\) in measure. By the Riesz theorem (Theorem 2.2.5) there is a subsequence with \(h_{n_k}\to f\) a.e., and \(m_k^{\prime}:=m_{n_k}\) is strictly increasing, so \(f_{n_k,m_k^{\prime}}=h_{n_k}\to f\) a.e. is the required assertion.
(\(\circ\)) Investigate for which real \(\alpha\) and \(\beta\) the function \(x^\alpha \sin(x^\beta)\) is Lebesgue integrable on (a) \((0,1)\), (b) \((0,+\infty)\), (c) \((1,+\infty)\). Answer the same question for the proper and improper Riemann integrability.
With \(f(x):=x^\alpha\sin(x^\beta)\) the criteria are
\begin{equation*} \begin{aligned} \text{(a) } (0,1):\ &\text{Lebesgue} \iff \alpha+\max(\beta,0)>-1, \quad \text{improper} \iff \alpha>-1-|\beta|, \\ &\text{proper Riemann} \iff \alpha+\max(\beta,0)\ge 0; \\ \text{( c) } (1,\infty):\ &\text{Lebesgue} \iff \alpha+\min(\beta,0)<-1, \quad \text{improper} \iff \alpha<|\beta|-1; \\ \text{(b) } (0,\infty):\ &\text{Lebesgue} \iff \text{(a) and ( c)}, \quad \text{improper} \iff |\alpha+1|<|\beta| , \end{aligned} \end{equation*}
with no proper Riemann integral on the two unbounded intervals. Everything follows from two lemmas and the substitution \(t=x^\beta\), \(f\) being continuous on \((0,+\infty)\), so that only the endpoints matter.
Lemma 1. For \(\gamma\in\mathbb{R}\), \(\int_1^\infty t^\gamma|\sin t|\,dt<\infty\) iff \(\gamma<-1\), and \(\lim_{T\to\infty}\int_1^Tt^\gamma\sin t\,dt\) exists in \(\mathbb{R}\) iff \(\gamma<0\). For \(\gamma<-1\) the first follows from \(\int_1^\infty t^\gamma dt<\infty\); for \(\gamma\ge-1\),
\begin{equation*} \int_{n\pi}^{(n+1)\pi} t^\gamma |\sin t| \, dt \ \ge\ 2\min\bigl((n\pi)^\gamma,((n+1)\pi)^\gamma\bigr) \ \ge\ c(\gamma)\, n^{\gamma} \end{equation*}
with \(c=2\pi^\gamma\) for \(\gamma\ge0\) and \(c=2(2\pi)^\gamma\) for \(-1\le\gamma<0\) (using \(n+1\le2n\)), and \(\sum_nn^\gamma=\infty\). For \(\gamma<0\), integration by parts gives
\begin{equation*} \int_1^T t^\gamma \sin t\,dt = \bigl[-t^\gamma \cos t\bigr]_1^T + \gamma \int_1^T t^{\gamma-1}\cos t \, dt , \end{equation*}
where \(T^\gamma\cos T\to0\) and \(\int_1^\infty t^{\gamma-1}|\cos t|\,dt<\infty\) as \(\gamma-1<-1\); for \(\gamma\ge0\) the Cauchy criterion fails, since \(\int_{2n\pi}^{(2n+1)\pi}t^\gamma\sin t\,dt\ge2(2n\pi)^\gamma\ge2\).
Lemma 2. For \(\gamma\in\mathbb{R}\), \(\int_0^1t^\gamma|\sin t|\,dt<\infty\) iff \(\gamma>-2\), and as the integrand is nonnegative this is also the criterion for the improper integral: \(\tfrac56t\le\sin t\le t\) on \((0,1]\) (from \(\sin t\ge t-t^3/6\)) gives \(\tfrac56t^{\gamma+1}\le t^\gamma\sin t\le t^{\gamma+1}\).
For \(\beta\ne0\) the substitution \(t=x^\beta\) is a \(C^1\) diffeomorphism of \((0,\infty)\) onto itself, increasing for \(\beta>0\) and decreasing for \(\beta<0\), and
\begin{equation*} x^\alpha \sin(x^\beta)\,dx = \frac{1}{\beta}\, t^{\gamma}\sin t \, dt, \qquad \gamma := \frac{\alpha+1}{\beta} - 1 ; \end{equation*}
so for \(\beta>0\) the ranges \((0,1)\) and \((1,\infty)\) in \(x\) correspond to the same ranges in \(t\), and for \(\beta<0\) they are interchanged (the reversed orientation only flips a sign). Applying this to \(\int_\delta^1\) and \(\int_1^R\) and letting \(\delta\to0^+\), \(R\to\infty\):
(a) (i) \(\beta>0\): Lemma 2 gives both criteria as \(\gamma>-2\), i.e. \(\alpha+\beta>-1\). (ii) \(\beta=0\): \(f=\sin(1)x^\alpha\), so both read \(\alpha>-1\). (iii) \(\beta<0\): Lemma 1 gives Lebesgue integrability iff \(\gamma<-1\), i.e. \(\alpha>-1\), and convergence of \(\lim_{\delta\to0^+}\int_\delta^1f\) iff \(\gamma<0\), i.e. \(\alpha>\beta-1\). For proper Riemann integrability the only issue is boundedness, \(f\) being continuous on \((0,1]\) (Exercise 2.12.38(i)): for \(\beta\ge0\) one has \(f(x)/x^{\alpha+\beta}\to1\), so boundedness means \(\alpha+\beta\ge0\); for \(\beta<0\), \(\alpha\ge0\) gives \(|f|\le1\) while \(\alpha<0\) gives \(f(x_n)=x_n^\alpha\to\infty\) along \(x_n:=(\pi/2+2\pi n)^{1/\beta}\to0\).
(c) (i) \(\beta>0\): by Lemma 1, Lebesgue iff \(\gamma<-1\), i.e. \(\alpha<-1\), improper iff \(\gamma<0\), i.e. \(\alpha<\beta-1\). (ii) \(\beta=0\): both read \(\alpha<-1\). (iii) \(\beta<0\): by Lemma 2 both read \(\gamma>-2\), i.e. \(\alpha+\beta<-1\).
(b) Continuity of \(f\) on \((0,\infty)\) makes each property on \((0,+\infty)\) equivalent to its holding on \((0,1)\) and on \((1,+\infty)\) at once; the Lebesgue condition says that \(-1\) lies strictly between \(\alpha\) and \(\alpha+\beta\) (so \(-1-\beta<\alpha<-1\) for \(\beta>0\), \(-1<\alpha<-1-\beta\) for \(\beta<0\), impossible for \(\beta=0\)), and the improper condition is \(-1-|\beta|<\alpha<|\beta|-1\).
(\(\circ\)) (Alekhno, Zabreiko) Let \(\mu\) be a finite nonnegative measure on a measurable space \((X,\mathcal{A})\) and let \(\{f_n\}\) be a sequence of \(\mu\)-measurable functions. Suppose that it is not true that this sequence converges to zero \(\mu\)-a.e. Prove that there exist a subsequence \(\{f_{n_k}\}\) in \(\{f_n\}\), measurable sets \(A_k\) with \(\mu(A_k) > 0\) and \(A_{k+1} \subset A_k\) for all \(k\), and \(\varepsilon > 0\) such that \(|f_{n_k}(x)| \ge \varepsilon\) for all \(x \in A_k\) and all \(k\).
Take \(\varepsilon:=1/j\) with \(\mu(\{g>1/j\})>0\), where \(g:=\limsup_n|f_n|=\lim_m g_m\) and \(g_m:=\sup_{n\ge m}|f_n|\), and set \(A_k:=E\cap E_{n_1}\cap\cdots\cap E_{n_k}\) for the sets and indices below. Each \(g_m\) is \(\mu\)-measurable, since \(\{g_m>c\}=\bigcup_{n\ge m}\{|f_n|>c\}\), and \(\{g_m\}\) is nonincreasing, so \(g\) is \(\mu\)-measurable. Now \(f_n(x)\to0\) exactly when \(g(x)=0\), so the failure of a.e. convergence means that the measurable set \(\{g>0\}=\bigcup_j\{g>1/j\}\) is not null; by countable additivity some \(j\) has \(\mu(\{g>1/j\})>0\), which fixes \(\varepsilon\).
Put
\begin{equation*} E := \bigcap_{m \ge 1} \{ x \colon g_m(x) > \varepsilon \}, \qquad E_n := \{ x \in E \colon |f_n(x)| > \varepsilon \} . \end{equation*}
Since \(g(x)>\varepsilon\) forces \(g_m(x)\ge g(x)>\varepsilon\) for every \(m\), we get \(\{g>\varepsilon\}\subset E\) and \(\mu(E)>0\); and for every \(m\),
\begin{equation*} E = \bigcup_{n \ge m} E_n , \end{equation*}
because \(x\in E\) gives \(\sup_{n\ge m}|f_n(x)|=g_m(x)>\varepsilon\), hence \(|f_n(x)|>\varepsilon\) for some \(n\ge m\).
Choose the indices inductively from \(A_0:=E\), \(n_0:=0\): given \(A_{k-1}\subset E\) with \(\mu(A_{k-1})>0\), the last display with \(m=n_{k-1}+1\) writes \(A_{k-1}\) as the countable union of the sets \(A_{k-1}\cap E_n\), \(n>n_{k-1}\), so countable subadditivity yields \(n_k>n_{k-1}\) with \(\mu(A_k)>0\) for \(A_k:=A_{k-1}\cap E_{n_k}\). Then \(n_1<n_2<\cdots\), the sets \(A_k\) are \(\mu\)-measurable with \(\mu(A_k)>0\) and \(A_{k+1}\subset A_k\), and \(A_k\subset E_{n_k}\) gives \(|f_{n_k}(x)|>\varepsilon\) for all \(x\in A_k\).
(\(\circ\)) Investigate for which real \(\alpha\) and \(\beta\) the function \(x^\alpha (\ln x)^\beta\) is Lebesgue integrable on (a) \((0,1)\), (b) \((0,+\infty)\).
(a) Exactly \(\alpha > -1\) and \(\beta > -1\); (b) no pair \((\alpha,\beta)\). Only \(F(x) := x^\alpha|\ln x|^\beta\) matters, and only the three points \(0, 1, +\infty\).
(i) At \(x = 1\): \(|\ln x|/|x-1| \to 1\), so \(F(x)/|x-1|^\beta \to 1\) and \(F\) is integrable near \(1\) iff \(\beta > -1\).
(ii) At \(x = 0^+\): for \(\alpha > -1\) pick \(\eta > 0\) with \(\alpha - \eta > -1\), so \(|\ln x|^\beta = o(x^{-\eta})\) gives \(F \le C x^{\alpha-\eta}\) on \((0,1/2]\), integrable; for \(\alpha < -1\) pick \(\eta > 0\) with \(\alpha + \eta < -1\), so \(x^\eta|\ln x|^{-\beta} \to 0\) gives \(F \ge c\,x^{\alpha+\eta}\) near \(0\), not integrable; for \(\alpha = -1\) the substitution \(u = -\ln x\) gives
\begin{equation*} \int_0^{1/2} \frac{|\ln x|^\beta}{x}\,dx = \int_{\ln 2}^{\infty} u^\beta\,du , \end{equation*}
finite iff \(\beta < -1\).
(iii) At \(x = +\infty\): \(y = 1/x\) turns \(\int_1^\infty x^\alpha|\ln x|^\beta dx\) into \(\int_0^1 y^{-\alpha-2}|\ln y|^\beta dy\), so (ii) with exponent \(-\alpha-2\) gives integrability iff \(\alpha < -1\), or \(\alpha = -1\) and \(\beta < -1\).
Hence (a) needs (i) and (ii), and \(\beta > -1\) kills the branch \(\alpha = -1,\ \beta < -1\), leaving \(\alpha > -1\) and \(\beta > -1\); (b) needs all three, but (ii) and (iii) together force \(\alpha = -1,\ \beta < -1\), which (i) forbids.
(\(\circ\)) Let \(J_n\) be a sequence of disjoint intervals in \([0,1]\), convergent to the origin, \(|J_n| = 4^{-n}\), and let \(f = n^{-1}/|J_{2n}|\) on \(J_{2n}\), \(f = -n^{-1}/|J_{2n+1}|\) on \(J_{2n+1}\), and let \(f\) be zero at all other points. Show that \(f\) is Riemann integrable in the improper sense, but is not Lebesgue integrable.
The improper integral is \(0\), while \(\int_0^1|f| = +\infty\). The printed statement omits the arrangement of the \(J_n\); as intended we order them along \([0,1]\) by index, \(J_{n+1}\) left of \(J_n\) with \(\max J_n \to 0\) (say \(J_n = [\tfrac23 4^{-n}, \tfrac53 4^{-n}]\)), which the improper half needs, the series below converging only conditionally. With \(c_n := \int_{J_n} f\),
\begin{equation*} c_1 = 0, \qquad c_{2n} = \frac1n, \qquad c_{2n+1} = -\frac1n \qquad (n \ge 1). \end{equation*}
Not Lebesgue integrable: monotone convergence on the partial sums gives
\begin{equation*} \int_0^1 f^{+} = \int_0^1 f^{-} = \sum_{n=1}^\infty \frac1n = +\infty . \end{equation*}
Also \(f\) is unbounded, hence not properly Riemann integrable.
Improperly Riemann integrable: the origin is the only point near which \(f\) is unbounded, and for \(\delta \in (0,1)\) only finitely many \(J_n\) meet \([\delta,1]\), so \(f|_{[\delta,1]}\) is a step function with equal Riemann and Lebesgue integrals. The ordering makes \(\{n \colon J_n \subset [\delta,1]\}\) an initial segment \(n \le N(\delta)\), with only \(J_{N+1}\) cut by \(\delta\), so
\begin{equation*} \int_\delta^1 f = \sum_{n=1}^{N(\delta)} c_n + r(\delta), \qquad |r(\delta)| \le |c_{N+1}| . \end{equation*}
Here \(N(\delta) \to \infty\) and \(|c_m| \to 0\), so \(r(\delta) \to 0\), while \(S_{2k+1} = 0\) and \(S_{2k} = 1/k\) for \(S_N := \sum_{n \le N} c_n\); hence \(\lim_{\delta \to 0^+}\int_\delta^1 f = 0\).
(i) (H. Lebesgue, G. Vitali) Show that a bounded function is Riemann integrable on an interval (or a cube) precisely when the set of its discontinuity points has measure zero.
(ii) Prove that a function \(f\) on \([a,b]\) is Riemann integrable precisely when, for each \(\varepsilon > 0\), there exist step functions \(g\) and \(h\) such that \(|f(x) - g(x)| \le h(x)\) and
\begin{equation*} \int_a^b h(x)\,dx \le \varepsilon . \end{equation*}
(i) The oscillation carries everything: with \(\omega(f,S) := \sup_{x,y\in S}|f(x)-f(y)|\) and \(\omega_f(x) := \inf_{\delta>0}\omega(f,(x-\delta,x+\delta)\cap[a,b])\), continuity at \(x\) means \(\omega_f(x)=0\), so the discontinuity set is
\begin{equation*} D = \bigcup_{k\ge1} D_{1/k}, \qquad D_\sigma := \{x \in [a,b] \colon \omega_f(x) \ge \sigma\} , \end{equation*}
each \(D_\sigma\) compact (if \(\omega_f(x)<\sigma\) then every \(y\) near \(x\) inherits the bound, so the complement is open). Put \(M := \sup|f|\); by the Darboux criterion, integrability means \(\sum_i \omega(f,I_i)|I_i| < \eta\) for some partition \(I_i = [t_{i-1},t_i]\) and each \(\eta>0\).
Necessity. Fix \(\sigma,\eta>0\) and a partition with \(\sum_i\omega(f,I_i)|I_i| < \sigma\eta/2\); the indices \(S\) with \((t_{i-1},t_i)\cap D_\sigma \ne \emptyset\) have \(\omega(f,I_i) \ge \sigma\), so
\begin{equation*} \sigma\sum_{i\in S}|I_i| \le \sum_i \omega(f,I_i)|I_i| < \frac{\sigma\eta}{2} . \end{equation*}
Since \(D_\sigma\) lies in \(\bigcup_{i\in S}I_i\) plus finitely many partition points, \(\lambda^*(D_\sigma) \le \eta/2\); as \(\eta \downarrow 0\), \(\lambda(D_\sigma)=0\) for every \(\sigma>0\) and hence \(\lambda(D)=0\).
Sufficiency. Given \(\lambda(D)=0\) and \(\eta>0\), put \(\sigma := \eta/(2(b-a))\) and cover the compact null set \(D_\sigma\) by open \(U_1,\dots,U_p\) with \(\sum_j|U_j| < \eta/(4M+1)\). On the compact \(K := [a,b]\setminus\bigcup_j U_j\) we have \(\omega_f < \sigma\), so each \(x \in K\) admits \(\delta_x>0\) with \(\omega(f,(x-2\delta_x,x+2\delta_x)\cap[a,b])<\sigma\); a finite subcover by \((x_i-\delta_{x_i},x_i+\delta_{x_i})\) gives \(\rho := \min_i \delta_{x_i}>0\). For mesh \(<\rho\): (a) a subinterval meeting \(K\) at \(z \in (x_i-\delta_{x_i},x_i+\delta_{x_i})\) lies in \((x_i-2\delta_{x_i},x_i+2\delta_{x_i})\), so \(\omega(f,I)<\sigma\); (b) the rest lie in \(\bigcup_j U_j\), of total length \(<\eta/(4M+1)\), with \(\omega(f,I)\le 2M\). Hence
\begin{equation*} U(f,P)-L(f,P) \le \sigma(b-a) + \frac{2M\eta}{4M+1} < \eta . \end{equation*}
On a cube read boxes, volumes, balls and maximal diameter for intervals, lengths, intervals and mesh.
(ii) Necessity. Darboux supplies a partition with \(\sum_i (M_i-m_i)(t_i-t_{i-1}) \le \varepsilon\), where \(m_i, M_i\) are the inf and sup of \(f\) on \([t_{i-1},t_i]\); take \(g := m_i\), \(h := M_i - m_i\) on \([t_{i-1},t_i)\), \(g(b):=f(b)\), \(h(b):=0\). Then \(|f-g|\le h\) (Check!) and \(\int_a^b h \le \varepsilon\).
Sufficiency. With \(\varepsilon=1\), \(|f|\le|g|+h\) is bounded; for each \(k\) take \(g_k,h_k\) for \(\varepsilon=4^{-k}\), so Chebyshev on the step function \(h_k\) gives
\begin{equation*} \lambda(E_k) \le 2^k \int_a^b h_k \le 2^{-k}, \qquad E_k := \{h_k > 2^{-k}\} , \end{equation*}
making \(F := \bigcap_m\bigcup_{k\ge m}E_k\) null, with \(|f-g_k| \le h_k \le 2^{-k}\) for \(k \ge m\) off \(\bigcup_{k\ge m}E_k\). Let \(Z\) be the countable, hence null, set of all division points of the \(g_k,h_k\). For \(x \notin F\cup Z\) and \(\eta>0\) pick \(m\) with \(x \notin \bigcup_{k\ge m}E_k\) and \(k \ge m\) with \(2^{-k}<\eta/3\); on a neighbourhood \(V\) of \(x\) both \(g_k,h_k\) are constant, so for \(y \in V\)
\begin{equation*} |f(y)-f(x)| \le h_k(y) + |g_k(y)-g_k(x)| + h_k(x) \le 2\cdot2^{-k} < \eta . \end{equation*}
Thus \(f\) is bounded with discontinuity set inside the null set \(F\cup Z\), and (i) applies.
Exercises 2.12.39–2.12.45
(\(\circ\)) Suppose that a sequence of \(\mu\)-integrable functions \(f_n\) converges to \(f\) in \(L^1(\mu)\) and a sequence of \(\mu\)-measurable functions \(\varphi_n\) converges to \(\varphi\) \(\mu\)-a.e. and is uniformly bounded. Show that the functions \(\varphi_n f_n\) converge to \(\varphi f\) in \(L^1(\mu)\).
Split \(\varphi_n f_n - \varphi f = \varphi_n(f_n-f) + (\varphi_n-\varphi)f\). With \(C\) a common bound for the \(|\varphi_n|\), also \(|\varphi| \le C\) a.e. (pass to the limit), so all products are integrable by Corollary 2.5.6, and a.e.
\begin{equation*} |\varphi_n f_n - \varphi f| \le C|f_n-f| + |\varphi_n - \varphi|\,|f| . \end{equation*}
Integrating, the first term contributes \(C\|f_n-f\|_{L^1(\mu)} \to 0\) by hypothesis. For the second, \(g_n := |\varphi_n-\varphi|\,|f| \to 0\) a.e. with the integrable majorant \(2C|f|\), so dominated convergence (Theorem 2.8.1, extended to \([0,+\infty]\)-valued measures by Corollary 2.8.6) gives \(\int_X g_n\,d\mu \to 0\). Hence \(\varphi_n f_n \to \varphi f\) in \(L^1(\mu)\).
(\(\circ\)) Let a function \(f \ge 0\) be integrable with respect to a measure \(\mu\). Prove the equality
\begin{equation*} \int f \, d\mu = \lim_{r \downarrow 1} \sum_{n=-\infty}^{\infty} r^n \mu\bigl(\{x :\; r^n \le f(x) < r^{n+1}\}\bigr). \end{equation*}
The sum \(S( r)\) is squeezed between \(r^{-1}\int f\,d\mu\) and \(\int f\,d\mu\), which forces the limit. Discard a null set so that \(0 \le f < \infty\) everywhere, fix \(r>1\), and put
\begin{equation*} A_n^r := \{r^n \le f < r^{n+1}\}, \qquad f_r := \sum_{n\in\mathbb{Z}} r^n I_{A_n^r} . \end{equation*}
The \(A_n^r\) are disjoint with union \(\{f>0\}\) (each \(f(x)\in(0,\infty)\) lies in the single band with \(n = \lfloor\log_r f(x)\rfloor\)) and each has finite measure by Chebyshev (Theorem 2.5.3): \(\mu(A_n^r) \le r^{-n}\int_X f\,d\mu\). By construction
\begin{equation*} f_r \le f \le r f_r \end{equation*}
(both sides vanish on \(\{f=0\}\); on \(A_n^r\), \(f_r = r^n \le f < r^{n+1} = rf_r\)), so \(f_r\) is integrable by Corollary 2.5.6, and the partial sums \(\sum_{|n|\le N} r^n I_{A_n^r}\) increase to \(f_r\) with integrals at most \(\int_X f\,d\mu\), whence monotone convergence (Theorem 2.8.2, extended to infinite measures by Corollary 2.8.6) gives
\begin{equation*} S( r) := \sum_{n\in\mathbb{Z}} r^n \mu(A_n^r) = \int_X f_r\,d\mu . \end{equation*}
Integrating the two-sided bound,
\begin{equation*} \frac1r \int_X f\,d\mu \le S( r) \le \int_X f\,d\mu \qquad (r>1) , \end{equation*}
and \(r \downarrow 1\) gives \(S( r) \to \int_X f\,d\mu\).
(\(\circ\)) (i) Construct a sequence of nonnegative functions \(f_n\) on \([0,1]\) convergent to zero pointwise such that their integrals tend to zero, but the function \(\Phi(x) = \sup_n f_n(x)\) is not integrable. In particular, the functions \(f_n\) have no common integrable majorant.
(ii) Construct a sequence of functions \(f_n \ge 0\) on \([0,1]\) such that their integrals tend to zero, but \(\sup_n f_n(x) = +\infty\) for every \(x\).
(i) Take \(f_n := n\,I_{[(n+1)^{-1},\,n^{-1}]}\) on \([0,1]\). Then
\begin{equation*} \int_0^1 f_n\,d\lambda = n\Bigl(\frac1n - \frac1{n+1}\Bigr) = \frac{1}{n+1} \to 0 , \end{equation*}
and \(f_n \to 0\) everywhere, since \(f_n(x) \ne 0\) forces \(n \le 1/x \le n+1\), true for at most two \(n\) (and \(f_n(0)=0\)). On the disjoint \(J_n := ((n+1)^{-1},n^{-1})\) of length \(1/(n(n+1))\) we have \(\Phi := \sup_n f_n \ge n\), so
\begin{equation*} \int_0^1 \Phi\,d\lambda \ \ge\ \sum_{n=1}^{N} n\,\lambda(J_n) = \sum_{n=1}^{N}\frac{1}{n+1} \to \infty . \end{equation*}
Any a.e. majorant \(g\) of all \(f_n\) dominates \(\Phi\) a.e. off the countable union of the exceptional null sets, so an integrable one would make \(\Phi\) integrable by Corollary 2.5.6.
(ii) Take the rescaled sliding humps of Example 2.2.4: split \([0,1]\) into the \(2^n\) intervals \(I_{n,k}\) of length \(2^{-n}\), set \(g_{n,k} := n\,I_{I_{n,k}}\), and enumerate the pairs \((n,k)\) by increasing \(n\) as \(h_m = g_{n(m),k(m)}\). Each level contributes finitely many terms, so \(n(m) \to \infty\) and
\begin{equation*} \int_0^1 h_m\,d\lambda = n(m)\,2^{-n(m)} \to 0 . \end{equation*}
For each \(x\) and \(n\) the \(I_{n,k}\) cover \([0,1]\), so \(h_m(x) = n\) at the index \(m\) of the interval containing \(x\); hence \(\sup_m h_m(x) = +\infty\).
Let \(\mu\) be a probability measure on a space \(X\) and let \(\{f_n\}\) be a sequence of \(\mu\)-integrable functions that converges \(\mu\)-a.e. to a \(\mu\)-integrable function \(f\) such that the integrals of \(f_n\) converge to the integral of \(f\). Prove that for any \(\varepsilon > 0\) there exist a measurable set \(E\) and a number \(N \in \mathbb{N}\) such that for all \(n \ge N\) one has
\begin{equation*} \int_{X \setminus E} f_n \, d\mu \le \varepsilon \quad \text{and} \quad |f_n(x)| \le |f(x)| + 1 \ \text{ for } x \in E . \end{equation*}
Take for \(E\) an Egoroff set carrying uniform convergence whose complement is small enough for the absolute continuity of \(\int|f|\); this even gives \(\bigl|\int_{X\setminus E}f_n\,d\mu\bigr| \le \varepsilon\). Discard a null set so that all functions are finite and \(f_n \to f\) everywhere.
Absolute continuity of the integral of \(f\) (Theorem 2.5.7) gives \(\delta>0\) with \(\int_A|f|\,d\mu < \varepsilon/3\) whenever \(\mu(A)<\delta\), and Egoroff (Theorem 2.2.1, applicable as \(\mu\) and the limit are finite) gives measurable \(E\) with \(\mu(X\setminus E)<\delta\) and \(f_n \to f\) uniformly on \(E\). Choose \(N\) so large that for \(n \ge N\)
\begin{equation*} \sup_{x\in E}|f_n(x)-f(x)| \le \tfrac13\min(1,\varepsilon), \qquad \Bigl|\int_X (f_n-f)\,d\mu\Bigr| \le \frac{\varepsilon}{3} , \end{equation*}
the second by the hypothesis on the integrals. The first bound gives at once \(|f_n| \le |f| + \tfrac13\min(1,\varepsilon) \le |f|+1\) on \(E\).
For the other bound, all functions are integrable, so integrals split (Theorem 2.5.1) and
\begin{equation*} \int_{X\setminus E} f_n\,d\mu = \int_X (f_n - f)\,d\mu + \int_{X\setminus E} f\,d\mu + \int_E (f-f_n)\,d\mu \end{equation*}
(expand the right side: the \(f\) terms recombine into \(-\int_E f_n\,d\mu\)). The three terms are at most \(\varepsilon/3\) in modulus: the first by the choice of \(N\); the second because \(\mu(X\setminus E)<\delta\); the third because \(\mu(E) \le 1\) and the sup bound holds on \(E\). Adding, \(\bigl|\int_{X\setminus E}f_n\,d\mu\bigr| \le \varepsilon\) for all \(n \ge N\).
Let \(\mu\) be a probability measure on a space \(X\) and let \(f_n\) be \(\mu\)-measurable functions. Prove that the following conditions are equivalent:
(i) there exists a subsequence \(f_{n_k}\) convergent a.e. to \(0\);
(ii) there exists a sequence of numbers \(t_n\) such that
\begin{equation*} \limsup_{n \to \infty} |t_n| > 0 \quad \text{and} \quad \sum_{n=1}^{\infty} t_n f_n(x) \ \text{ converges a.e.}; \end{equation*}
(iii) there exists a sequence of numbers \(t_n\) such that
\begin{equation*} \sum_{n=1}^{\infty} |t_n| = \infty \quad \text{and} \quad \sum_{n=1}^{\infty} |t_n f_n(x)| < \infty \ \text{ a.e.} \end{equation*}
One choice of \(t_n\) — the indicator of a rapidly convergent subsequence — proves (i) \(\Rightarrow\) (ii) and (iii) at once; the converses are Borel–Cantelli (\(\sum_k\mu(B_k)<\infty\) puts a.e. \(x\) in only finitely many \(B_k\)). Modify the \(f_n\) on a null set so that all are everywhere finite and \(\mathcal{A}_\mu\)-measurable.
(i) \(\Rightarrow\) (ii), (iii). Let \(f_{n_k} \to 0\) a.e.; as \(\mu\) is finite this gives convergence in measure (Theorem 2.2.3), so thin to indices \(m_j\) with \(\mu(\{|f_{m_j}|>2^{-j}\}) \le 2^{-j}\), whence Borel–Cantelli gives \(\sum_j |f_{m_j}(x)| < \infty\) a.e. Take \(t_n := 1\) for \(n \in \{m_j\}\) and \(t_n := 0\) otherwise: \(\sum_n t_n f_n = \sum_j f_{m_j}\) converges absolutely a.e., while \(\limsup_n|t_n| = 1\) and \(\sum_n|t_n| = \infty\).
(ii) \(\Rightarrow\) (i). Terms of a convergent series tend to \(0\), so \(t_n f_n \to 0\) a.e. Put \(c := \min(1,\tfrac12\limsup_n|t_n|) > 0\) (the truncation covers \(\limsup = +\infty\)); since \(c < \limsup_n|t_n|\) there are \(n_1<n_2<\cdots\) with \(|t_{n_k}| \ge c\), and a.e.
\begin{equation*} |f_{n_k}| = \frac{|t_{n_k}f_{n_k}|}{|t_{n_k}|} \le \frac1c\,|t_{n_k}f_{n_k}| \to 0 . \end{equation*}
(iii) \(\Rightarrow\) (i). Let \(S := \sum_n |t_n f_n| < \infty\) a.e. and \(X_N := \{S \le N\}\), increasing with \(\mu(X_N) \to 1\). On \(X_N\) the partial sums \(S_M I_{X_N}\) increase to \(SI_{X_N}\) with integrals at most \(N\), so monotone convergence (Theorem 2.8.2) gives, in \([0,+\infty]\) with \(0\cdot(+\infty)=0\),
\begin{equation*} \sum_{n=1}^{\infty}|t_n|\int_{X_N}|f_n|\,d\mu = \int_{X_N} S\,d\mu \le N . \end{equation*}
Hence \(\liminf_n \int_{X_N}|f_n|\,d\mu = 0\): a bound \(\ge \alpha > 0\) for all large \(n\) would make the left-hand side infinite, since \(\sum_n|t_n| = \infty\). Choose inductively \(n_1<n_2<\cdots\) with \(\int_{X_N}|f_{n_N}|\,d\mu \le 4^{-N}\); Chebyshev (Theorem 2.5.3) on \(X_N\) gives
\begin{equation*} \mu(C_N) \le 2^N 4^{-N} = 2^{-N}, \qquad C_N := \{x \in X_N \colon |f_{n_N}(x)| > 2^{-N}\} . \end{equation*}
By Borel–Cantelli a.e. \(x\) lies in only finitely many \(C_N\), and a.e. \(x\) lies in some \(X_{N_0}\), hence in every \(X_N\) with \(N \ge N_0\). So for a.e. \(x\) one has \(|f_{n_N}(x)| \le 2^{-N}\) for all large \(N\), i.e. \(f_{n_N} \to 0\) a.e.
(\(\circ\)) Show that a sequence of measurable functions \(f_n\) on a space with a probability measure \(\mu\) converges almost uniformly (in the sense of Egoroff’s theorem) to a measurable function \(f\) precisely when
\begin{equation*} \lim_{n \to \infty} \mu \Bigl( \bigcup_{m \ge n} \{x :\; |f_m(x) - f_n(x)| \ge \varepsilon \} \Bigr) = 0 . \end{equation*}
The condition is the almost-uniform Cauchy criterion, understood for every \(\varepsilon>0\). Write
\begin{equation*} E_n(\varepsilon) := \bigcup_{m \ge n}\{x \colon |f_m(x)-f_n(x)| \ge \varepsilon\} , \end{equation*}
and modify the \(f_n\) on a null set so all are finite everywhere.
Necessity. Fix \(\varepsilon,\delta>0\) and a measurable \(X_\delta\) with \(\mu(X\setminus X_\delta)<\delta\) carrying uniform convergence, so \(\sup_{X_\delta}|f_m-f| < \varepsilon/2\) for \(m \ge n_0\). Then for \(n \ge n_0\), \(m \ge n\) and \(x \in X_\delta\) the triangle inequality gives \(|f_m(x)-f_n(x)|<\varepsilon\), so \(E_n(\varepsilon) \subset X\setminus X_\delta\) and \(\mu(E_n(\varepsilon)) < \delta\); as \(\delta\) was arbitrary, \(\mu(E_n(\varepsilon)) \to 0\).
Sufficiency. Fix \(\delta>0\) and, using the hypothesis with \(\varepsilon = 2^{-k}\), pick \(n_1<n_2<\cdots\) with \(\mu(E_{n_k}(2^{-k})) < \delta 2^{-k-1}\). Put \(Y_\delta := X \setminus \bigcup_k E_{n_k}(2^{-k})\), so \(\mu(X\setminus Y_\delta) \le \delta/2 < \delta\). For \(x \in Y_\delta\) and \(m,m^{\prime} \ge n_k\),
\begin{equation*} |f_m(x)-f_{m^{\prime}}(x)| \le |f_m(x)-f_{n_k}(x)| + |f_{n_k}(x)-f_{m^{\prime}}(x)| < 2^{-k+1} , \end{equation*}
a bound independent of \(x\), so \(\{f_n\}\) is uniformly Cauchy, hence uniformly convergent, on \(Y_\delta\). Take \(Y_j := Y_{1/j}\) and \(Y := \bigcup_j Y_j\), so \(\mu(X\setminus Y) \le \inf_j 1/j = 0\); set \(f := \lim_n f_n\) on \(Y\) and \(f := 0\) off \(Y\). On each \(Y_j\) the uniform limit is \(f\) and \(\mu(X\setminus Y_j) \le 1/j\), so any \(j > 1/\delta\) exhibits almost uniform convergence to \(f\).
Prove the following analog of Egoroff’s theorem for spaces with infinite measure: let \(\mu\)-measurable functions \(f_n\) converge \(\mu\)-a.e. to a function \(f\) such that \(|f_n| \le g\) \(\mu\)-a.e., where the function \(g\) is integrable with respect to \(\mu\); then, for any \(\varepsilon > 0\), there exists a set \(A_\varepsilon\) such that the functions \(f_n\) converge to \(f\) uniformly on \(A_\varepsilon\), and the complement of \(A_\varepsilon\) has \(\mu\)-measure less than \(\varepsilon\).
Slice \(X\) by the size of the majorant and apply Egoroff on each slice of finite measure; the tail slices need none, \(g\) being uniformly small there. Discarding a null set, assume everywhere all \(f_n, g\) are finite, \(f_n \to f\), \(|f_n| \le g\), hence \(g \ge 0\) and \(|f| \le g\). Put
\begin{equation*} G := \{g>1\}, \quad G_k := \{2^{-k} < g \le 2^{1-k}\}, \quad Z_0 := \{g = 0\} , \end{equation*}
a measurable partition of \(X\). Chebyshev (Theorem 2.5.3) gives \(\mu(G) \le \int_X g\,d\mu < \infty\) and \(\mu(G_k) \le 2^k\int_X g\,d\mu < \infty\) (only \(\mu(Z_0)\) may be infinite). Since \(\mu\) restricted to \(G\), resp. \(G_k\), is finite and the limit \(f\) is finite, Egoroff (Theorem 2.2.1) supplies measurable \(A \subset G\) and \(A_k \subset G_k\) carrying uniform convergence with
\begin{equation*} \mu(G\setminus A) < \frac{\varepsilon}{2}, \qquad \mu(G_k \setminus A_k) < \varepsilon 4^{-k} . \end{equation*}
Set \(A_\varepsilon := Z_0 \cup A \cup \bigcup_k A_k\); its complement lies in \((G\setminus A) \cup \bigcup_k(G_k\setminus A_k)\), so
\begin{equation*} \mu(X\setminus A_\varepsilon) < \frac{\varepsilon}{2} + \varepsilon\sum_{k\ge1}4^{-k} = \frac{5\varepsilon}{6} < \varepsilon . \end{equation*}
For uniformity on this infinite union, given \(\eta>0\) pick \(K\) with \(2^{2-K} \le \eta\) and split: (i) on \(Z_0\), \(g=0\) forces \(f_n \equiv f \equiv 0\); (ii) on \(\bigcup_{k>K}A_k\), \(|f_n|,|f| \le 2^{1-k}\) gives \(|f_n - f| \le 2^{2-k} \le \eta\) for every \(n\); (iii) on the finite union \(A \cup A_1 \cup \cdots \cup A_K\), uniformity on each piece gives \(N_0\) with \(|f_n-f| \le \eta\) for \(n \ge N_0\). Hence \(\sup_{A_\varepsilon}|f_n-f| \le \eta\) for \(n \ge N_0\).
Exercises 2.12.46–2.12.52
(Tolstoff) (i) Let \(f\) be a Borel function on \([0,1]^2\), \(y_0\) a fixed point in \([0,1]\) and \(\lim_{y\to y_0} f(x,y) = f(x,y_0)\) for any \(x\in[0,1]\). Prove that for every \(\varepsilon>0\) there exists a measurable set \(A_\varepsilon\subset[0,1]\) of Lebesgue measure \(\lambda(A_\varepsilon)>1-\varepsilon\) such that \(\lim_{y\to y_0} f(x,y)=f(x,y_0)\) uniformly in \(x\in A_\varepsilon\).
(ii) Construct a bounded Lebesgue measurable function \(f\) on \([0,1]^2\) such that it is Borel in every variable separately and \(\lim_{y\to 0}f(x,y)=0\) for any \(x\in[0,1]\), but on no set of positive measure is convergence uniform.
(i) The good sets are cut out by the modulus of continuity in \(y\): for \(n\in\mathbb{N}\) put
\begin{equation*} \delta_n(x):=\sup\bigl\{\delta>0:\ |f(x,y)-f(x,y_0)|<1/n \text{ for } |y-y_0|<\delta\bigr\}, \end{equation*}
with \(y\) ranging over \([0,1]\) and \(\delta_n(x)=+\infty\) if every \(\delta\) is admissible. Pointwise convergence makes \(\delta_n(x)>0\), and the admissible \(\delta\) form an interval with left endpoint \(0\), so
\begin{equation*} |y-y_0|<\delta_n(x)\ \Longrightarrow\ |f(x,y)-f(x,y_0)|<1/n. \tag{1} \end{equation*}
Each \(\delta_n\) is Lebesgue measurable. Put
\begin{equation*} M(n,C):=\{(x,y)\in[0,1]^2\colon\ |f(x,y)-f(x,y_0)|\ge 1/n,\ |y-y_0|<C\}, \end{equation*}
a Borel set (the map \((x,y)\mapsto f(x,y_0)\) is \(f\) composed with continuous maps), and let \(\pi\) be the projection to the first axis. Then \(\{\delta_n<C\}=\pi(M(n,C))\) for \(C\ge0\): if \(\delta_n(x)<C\) then \(C\) itself is inadmissible, and conversely such a \(y\) caps every admissible \(\delta\) at \(|y-y_0|\). Borel sets are Souslin (Theorem 1.10.4), continuous images of Souslin sets are Souslin (Proposition 1.10.8), and Souslin sets are Lebesgue measurable (Corollary 1.10.9); for \(C\le0\) the set is empty.
Since \(\{\delta_n\ge 1/j\}\uparrow[0,1]\) as \(j\to\infty\), continuity from below supplies \(\gamma_n>0\) with
\begin{equation*} \lambda(A_n)>1-\varepsilon 2^{-n-1},\qquad A_n:=\{x\in[0,1]:\ \delta_n(x)\ge\gamma_n\}, \end{equation*}
so \(A_\varepsilon:=\bigcap_n A_n\) satisfies \(\lambda([0,1]\setminus A_\varepsilon)\le\varepsilon/2\), i.e. \(\lambda(A_\varepsilon)>1-\varepsilon\). Given \(\eta>0\), pick \(n\) with \(1/n<\eta\): for \(x\in A_\varepsilon\) and \(|y-y_0|<\gamma_n\le\delta_n(x)\), (1) gives \(|f(x,y)-f(x,y_0)|<1/n<\eta\), which is uniformity in \(x\in A_\varepsilon\).
(ii) Take \(f(x,y):=1\) when \(x>0\), \(x\in E_n\) and \(y=x/n\) for some \(n\), and \(f:=0\) otherwise, where \(\{E_n\}_{n\ge1}\) partitions \([0,1)\) with \(\lambda^{\ast}(E_n)=1\) for every \(n\) (adjoin the point \(1\) to \(E_1\), which changes no outer measure).
Such a partition exists. Identify \([0,1)\) with \(\mathbb{T}=\mathbb{R}/\mathbb{Z}\). As \(1,\sqrt2,\sqrt3\) are rationally independent, the elements \(m\sqrt2+n\sqrt3 \bmod 1\) are pairwise distinct and form a countable subgroup \(\Gamma\subset\mathbb{T}\) whose subgroup \(\Gamma_1:=\{m\sqrt2 \bmod 1\colon m\in\mathbb{Z}\}\) is dense. Let \(C\) meet each coset of \(\Gamma\) exactly once (axiom of choice) and put \(S_n:=C+n\sqrt3+\Gamma_1\) for \(n\in\mathbb{Z}\); uniqueness of the representation in \(\Gamma\) makes the \(S_n\) pairwise disjoint (Check!) with union \(C+\Gamma=\mathbb{T}\). Each \(S_n\) is \(\Gamma_1\)-invariant, so a measurable envelope \(H_n\supset S_n\) with \(\lambda(H_n)=\lambda^{\ast}(S_n)\) has \(\lambda(H_n\triangle(H_n+\gamma))=0\) for \(\gamma\in\Gamma_1\), since \(H_n\cap(H_n+\gamma)\) is measurable, contains \(S_n\), and has measure at most \(\lambda(H_n)\). Now \(\varphi(t):=\lambda(H_n\cap(H_n+t))\) is continuous (use \(|\varphi(t)-\varphi(s)|\le\lambda(H_n\triangle(H_n+s-t))\) and approximate \(H_n\) by a finite union of intervals) and equals \(\lambda(H_n)\) on the dense set \(\Gamma_1\), hence everywhere, so Fubini gives
\begin{equation*} \lambda(H_n)^2=\int_{\mathbb{T}}\varphi(t)\,dt=\lambda(H_n), \end{equation*}
i.e. \(\lambda^{\ast}(S_n)\in\{0,1\}\). The value \(0\) is impossible: \(S_j=S_n+(j-n)\sqrt3\) would make every \(S_j\), hence \(\mathbb{T}=\bigcup_j S_j\), null. Re-index by \(\mathbb{N}\).
The \(E_n\) being disjoint, \(f\) is well defined with \(0\le f\le1\), and \(\{f\ne0\}\) lies in the countably many lines \(y=x/n\), a plane null set; so \(f=0\) a.e. and \(f\) is Lebesgue measurable by completeness. Separately in each variable \(f\) is the indicator of a single point or of a countable set, hence Borel: for \(x\in E_n\) with \(x>0\) the section is \(I_{\{x/n\}}\) and \(f(0,\cdot)\equiv0\), while for \(y>0\) only \(x\in\{ny\colon n\in\mathbb{N}\}\) can give \(f\ne0\) and \(f(\cdot,0)\equiv0\) because \(y=x/n\) with \(x>0\) forces \(y>0\). The same computation gives the pointwise limit: for \(x\in E_n\), \(x>0\), one has \(f(x,y)=0\) for \(0<y<x/n\), so \(\lim_{y\to0}f(x,y)=0\).
Convergence is uniform on no set \(E\) of positive measure: for each \(n\) some \(x_n\in E\cap E_n\) has \(x_n>0\), since otherwise \(E_n\subset([0,1]\setminus E)\cup\{0\}\) would force \(\lambda^{\ast}(E_n)\le1-\lambda(E)<1\). Then \(y_n:=x_n/n\in(0,1/n]\) and \(f(x_n,y_n)=1\), so every \(\delta>0\) admits \(x\in E\) and \(y\in(0,\delta)\) with \(|f(x,y)-0|=1\).
(Frumkin) Let \(f\) be a function on \([0,1]^2\) such that, for every fixed \(t\), the function \(s\mapsto f(t,s)\) is finite a.e. and measurable. Suppose that \(\lim_{t\to0}f(t,s)=f(0,s)\) for a.e. \(s\). Show that, for each \(\delta_1>0\), there exists a measurable set \(E_{\delta_1}\subset[0,1]\) with the following property: \(\lambda(E_{\delta_1})>1-\delta_1\) and, given \(\varepsilon>0\), one can find \(\delta>0\) such that whenever \(t<\delta\), the inequality \(|f(t,s)-f(0,s)|<\varepsilon\) holds for all \(s\), with the exception of points of some set \(E_t\) of measure zero.
Take for \(E_{\delta_1}\) the complement of a countable essential envelope of the bad sets at scales \(1/k\); the asserted inequality is of course claimed for the points \(s\in E_{\delta_1}\) lying outside \(E_t\), since for \(f(t,s):=I_{\{s<t\}}\) the set \(\{|f(t,\cdot)-f(0,\cdot)|\ge 1/2\}=[0,t)\) has positive measure and no null exceptional set could serve. For \(t\in(0,1]\) and \(\varepsilon>0\) put
\begin{equation*} S(t,\varepsilon):=\{s\in[0,1]:\ |f(t,s)-f(0,s)|\ge\varepsilon\}, \end{equation*}
with \(s\) included whenever \(f(t,s)\) or \(f(0,s)\) is not finite; each \(S(t,\varepsilon)\) is measurable, the sections being measurable and finite off null sets and \(\lambda\) complete.
Lemma. Any nonempty family \(\mathcal{S}\) of measurable subsets of \([0,1]\) has a countable subfamily \(\{A_j\}\) whose union \(U\) satisfies \(\lambda(A\setminus U)=0\) for every \(A\in\mathcal{S}\). Indeed, let \(\alpha\) be the supremum of \(\lambda(\bigcup_j A_j)\) over countable \(\{A_j\}\subset\mathcal{S}\), pick countable \(\mathcal{C}_k\subset\mathcal{S}\) with \(\lambda(\bigcup\mathcal{C}_k)>\alpha-1/k\), and let \(U:=\bigcup_k\bigcup\mathcal{C}_k\), so \(\lambda(U)=\alpha\); for \(A\in\mathcal{S}\) the family \(\mathcal{C}\cup\{A\}\) is countable, so \(\lambda(U\cup A)\le\alpha=\lambda(U)\).
Applying the lemma to \(\{S(t,\varepsilon)\colon 0<t<\delta\}\) gives \(t_1,t_2,\dots\in(0,\delta)\) such that
\begin{equation*} U(\varepsilon,\delta):=\bigcup_i S(t_i,\varepsilon) \end{equation*}
satisfies \(\lambda(S(t,\varepsilon)\setminus U(\varepsilon,\delta))=0\) for all \(t\in(0,\delta)\); in particular, since each \(S(t,\varepsilon)\) with \(t<\delta^{\prime}<\delta\) sits in \(U(\varepsilon,\delta)\) up to a null set, so does the countable union \(U(\varepsilon,\delta^{\prime})\):
\begin{equation*} \lambda\bigl(U(\varepsilon,\delta^{\prime})\setminus U(\varepsilon,\delta)\bigr)=0 \qquad (0<\delta^{\prime}<\delta). \tag{2} \end{equation*}
Main step: \(\lambda(U(\varepsilon,\delta))\to0\) as \(\delta\to0\). By (2) the numbers \(\lambda(U(\varepsilon,1/j))\) do not increase; let \(c\) be their limit and suppose \(c>0\). With \(V_j:=\bigcap_{i\le j}U(\varepsilon,1/i)\), (2) gives \(\lambda(V_j)=\lambda(U(\varepsilon,1/j))\), and the \(V_j\) decrease, so \(W:=\bigcap_j U(\varepsilon,1/j)\) has \(\lambda(W)=c>0\). Let \(N\) be the union of the null set where \(f(0,\cdot)\) is not finite, the countably many null sets where some \(f(t_{j,i},\cdot)\) is not finite (the \(t_{j,i}\in(0,1/j)\) being the selected numbers for \(U(\varepsilon,1/j)\)), and the null set off which \(f(t,s)\to f(0,s)\). Then \(\lambda(W\setminus N)=c>0\), so pick \(s_0\in W\setminus N\); for each \(j\) it lies in some \(S(t_{j,i},\varepsilon)\), whence
\begin{equation*} |f(t_{j,i},s_0)-f(0,s_0)|\ge\varepsilon,\qquad 0<t_{j,i}<1/j , \end{equation*}
contradicting \(f(t,s_0)\to f(0,s_0)\). So \(c=0\).
Hence for each \(k\) there is \(\rho_k>0\) with \(\lambda(U(1/k,\rho_k))<\delta_1 2^{-k-1}\), and by (2) we may replace \(\rho_k\) by \(\min(\rho_1,\dots,\rho_k,1/k)\), so that \(\rho_k\) decreases with \(\rho_k\le1/k\). Put
\begin{equation*} E_{\delta_1}:=[0,1]\setminus\bigcup_{k=1}^{\infty}U(1/k,\rho_k), \end{equation*}
a measurable set with \(\lambda([0,1]\setminus E_{\delta_1})\le\sum_k\delta_1 2^{-k-1}=\delta_1/2\), so \(\lambda(E_{\delta_1})>1-\delta_1\).
Given \(\varepsilon>0\), choose \(k\) with \(1/k<\varepsilon\) and set \(\delta:=\rho_k\). For \(0<t<\delta\) the set
\begin{equation*} N_t^{k}:=E_{\delta_1}\cap S(t,1/k)\subset S(t,1/k)\setminus U(1/k,\rho_k) \end{equation*}
is null, because \(E_{\delta_1}\) misses \(U(1/k,\rho_k)\) and \(\lambda(S(t,1/k)\setminus U(1/k,\rho_k))=0\); and \(|f(t,s)-f(0,s)|<1/k<\varepsilon\) for \(s\in E_{\delta_1}\setminus N_t^k\). Finally \(E_t:=\bigcup_{k\colon t<\rho_k}N_t^{k}\) is a null set depending on \(t\) alone and contains each such \(N_t^k\), as required.
(Stampacchia) Suppose we are given a sequence of functions \(f_n\) on \([0,1]\times[0,1]\) measurable in \(x\) and continuous in \(y\). Assume that for every \(y\in[0,1]\) the sequence \(\{f_n(x,y)\}\) converges for a.e. \(x\) and that for a.e. \(x\) the sequence of functions \(y\mapsto f_n(x,y)\) is equicontinuous. Prove that for every \(\varepsilon>0\) there exists a measurable set \(E_\varepsilon\subset[0,1]\) of Lebesgue measure at least \(1-\varepsilon\) such that the sequence \(\{f_n(x,y)\}\) converges uniformly on the set \(E_\varepsilon\times[0,1]\).
Egoroff applied to the decreasing majorants
\begin{equation*} g_n(x):=\sup_{m,k\ge n}\ \sup_{y\in Q}\ |f_m(x,y)-f_k(x,y)| \in[0,+\infty] \end{equation*}
delivers \(E_\varepsilon\), where \(Q:=\mathbb{Q}\cap[0,1]=\{y_1,y_2,\dots\}\). Let \(X_j\) be a full-measure set carrying the convergence of \(\{f_n(\cdot,y_j)\}\) and \(X^{\prime}\) one carrying the equicontinuity of \(\{f_n(x,\cdot)\}\), both furnished by the hypotheses, and put \(X_0:=X^{\prime}\cap\bigcap_j X_j\), of full measure.
For \(x\in X_0\) the sequence \(f_n(x,\cdot)\) converges uniformly on \([0,1]\). Indeed, equicontinuity on a compact interval is uniform equicontinuity, so given \(\eta>0\) there is \(\rho>0\) with \(|f_n(x,y)-f_n(x,y^{\prime})|<\eta/3\) for all \(n\) whenever \(|y-y^{\prime}|<\rho\); take a \(\rho\)-net \(y_{j_1},\dots,y_{j_N}\in Q\) (possible as \(Q\) is dense) and \(M\) with \(|f_n(x,y_{j_i})-f_m(x,y_{j_i})|<\eta/3\) for \(n,m\ge M\) and all \(i\le N\). For \(y\in[0,1]\) and a net point \(y_{j_i}\) within \(\rho\), the three-term estimate gives
\begin{equation*} |f_n(x,y)-f_m(x,y)|<\eta\qquad (n,m\ge M), \end{equation*}
so \(f_n(x,\cdot)\) is uniformly Cauchy.
Hence \(g_1\ge g_2\ge\cdots\ge0\) with \(g_n\downarrow0\) on \(X_0\), that is a.e.; each \(g_n\) is measurable as a countable supremum of measurable functions, and is finite on \(X_0\), a uniformly convergent sequence of functions continuous on \([0,1]\) being uniformly bounded. Continuity of the \(f_m(x,\cdot)\) and density of \(Q\) also give
\begin{equation*} g_n(x)=\sup_{m,k\ge n}\ \sup_{y\in[0,1]}|f_m(x,y)-f_k(x,y)| . \end{equation*}
Egoroff (Theorem 2.2.1; \(\lambda\) is a probability measure and the limit \(0\) is finite) yields a measurable \(E_\varepsilon\) with \(\lambda([0,1]\setminus E_\varepsilon)\le\varepsilon\) on which \(g_n\to0\) uniformly. Given \(\eta>0\), an \(M\) with \(g_M\le\eta\) on \(E_\varepsilon\) forces \(|f_m(x,y)-f_l(x,y)|\le\eta\) for all \(m,l\ge M\), \(x\in E_\varepsilon\) and \(y\in[0,1]\), so \(\{f_n\}\) is uniformly Cauchy, hence uniformly convergent, on \(E_\varepsilon\times[0,1]\).
Suppose we are given a sequence of numbers \(\gamma=\{\gamma_k\}\). For \(x\in[0,1]\) let \(f_\gamma(x)=0\) if \(x\) is irrational, \(f_\gamma(0)=1\), and \(f_\gamma(x)=\gamma_k\) if \(x=m/k\) is an irreducible fraction. Prove that the function \(f_\gamma\) is Riemann integrable precisely when \(\lim_{k\to\infty}\gamma_k=0\).
Riemann integrable exactly when \(\gamma_k\to0\), and then \(\int_0^1 f_\gamma=0\). Write \(f:=f_\gamma\), and for a partition \(P\colon 0=x_0<\cdots<x_N=1\) put \(I_j:=[x_{j-1},x_j]\), \(\ell_j:=x_j-x_{j-1}\). Since the irrationals are dense and \(f\) vanishes there, every subinterval carries a zero of \(f\), so
\begin{equation*} \inf_{I_j}f\le 0\le\sup_{I_j}f,\qquad\text{hence}\qquad L(f,P)\le 0\le U(f,P) \tag{3} \end{equation*}
for every \(P\).
Sufficiency. Let \(\gamma_k\to0\), so \(|f|\le M_0:=\max(\sup_k|\gamma_k|,1)<\infty\) and \(f\) is bounded. Given \(\eta>0\), choose \(K\) with \(|\gamma_k|\le\eta\) for \(k>K\); the set
\begin{equation*} F:=\{0\}\cup\{m/k\in[0,1]:\ 1\le k\le K,\ \gcd(m,k)=1\} \end{equation*}
is finite, say with \(N_F\) points, and \(|f|\le\eta\) off \(F\) (an irrational gives \(0\), a reduced \(m/k\) with \(k>K\) gives \(|\gamma_k|\le\eta\)). Take \(\operatorname{mesh}(P)<\eta/(1+2N_FM_0)\). Each point of \(F\) lies in at most two of the closed \(I_j\), so at most \(2N_F\) subintervals meet \(F\) and on the rest \(|\sup_{I_j}f|,|\inf_{I_j}f|\le\eta\), whence
\begin{equation*} U(f,P)\le \eta+M_0\cdot 2N_F\cdot\operatorname{mesh}(P)\le 2\eta,\qquad L(f,P)\ge-2\eta . \end{equation*}
With (3) both Darboux integrals lie in \([-2\eta,2\eta]\) for every \(\eta>0\), so they equal \(0\).
Necessity. Riemann integrability forces \(f\), hence \(\{\gamma_k\}\), bounded. Suppose \(|\gamma_k|\ge\eta\) for infinitely many \(k\); passing to a subsequence, either \(\gamma_k\ge\eta\) infinitely often or \(\gamma_k\le-\eta\) infinitely often. The following density fact is the crux: for every \(\rho\in(0,1)\) there is \(K(\rho)\) such that for \(k\ge K(\rho)\) every interval \(I\subset[0,1]\) of length \(\rho\) contains an irreducible fraction with denominator exactly \(k\). Indeed, \(J:=\{m\in\mathbb{Z}\colon m/k\in I\}\) is a block of \(\#J\ge\rho k-1\) consecutive integers, and inclusion–exclusion over the squarefree divisors of \(k\) gives
\begin{equation*} \#\{m\in J:\ \gcd(m,k)=1\} =\sum_{d\mid k,\ d\ \text{squarefree}}\mu(d)\Bigl(\frac{\#J}{d}+\theta_d\Bigr) \end{equation*}
with \(|\theta_d|\le1\); the main term is \(\#J\,\varphi(k)/k\) and the error is at most the number \(d(k)\) of divisors, so
\begin{equation*} \#\{m\in J:\ \gcd(m,k)=1\}\ \ge\ \rho\,\varphi(k)-1-d(k)\ \ge\ \rho\sqrt{k/2}-1-Ck^{1/4}, \end{equation*}
by the elementary estimates \(\varphi(k)\ge\sqrt{k/2}\) and \(d(k)=O(k^{1/4})\), and the right side is positive for large \(k\).
Now let \(P\) be arbitrary and \(\rho:=\tfrac12\min_j\ell_j\), so every \(I_j\) contains an interval of length \(\rho\). In the first case pick \(k\ge K(\rho)\) with \(\gamma_k\ge\eta\): each \(I_j\) carries an irreducible \(m/k\), where \(f=\gamma_k\ge\eta\), so \(\sup_{I_j}f\ge\eta\) and \(U(f,P)\ge\eta\sum_j\ell_j=\eta\). Thus the upper Darboux integral is at least \(\eta\) while by (3) the lower one is at most \(0\) — not integrable. The second case is symmetric, with \(L(f,P)\le-\eta\) for every \(P\). Hence \(\gamma_k\to0\).
Method (2) for necessity: by the density fact, \(|\gamma_k|\ge\eta\) infinitely often makes \(f\) discontinuous at every irrational, so the discontinuity set has full measure and Exercise 2.12.38(i) applies.
(\(\circ\)) Let a function \(f\) on the real line be periodic with a period \(T>0\) and integrable on intervals. Show that the integrals of \(f\) over \([0,T]\) and \([a,a+T]\) coincide for all \(a\).
The integral over every interval of length \(T\) equals \(\int_0^T f\). Translation invariance of Lebesgue measure upgrades to translation invariance of the integral, \[\int_c^d g(x)\,dx=\int_{c-h}^{d-h}g(x+h)\,dx \qquad (h\in\mathbb{R}),\] for \(g\) integrable on a bounded \([c,d]\): it holds for \(g=I_A\) because \(I_A(x+h)=I_{A-h}(x)\) and \(\lambda(A-h)=\lambda(A)\), hence for simple \(g\) by linearity, for \(g\ge0\) by monotone convergence along simple approximations, and in general via \(g=g^{+}-g^{-}\).
Let \(0\le a\le T\). Additivity together with this identity for \(h=T\) and \(f(x+T)=f(x)\) gives \[\int_a^{a+T}f=\int_a^Tf+\int_T^{a+T}f=\int_a^Tf+\int_0^af(x+T)\,dx=\int_0^Tf .\] For arbitrary \(a\) write \(a=nT+b\) with \(n\in\mathbb{Z}\), \(b\in[0,T)\); periodicity iterates to \(f(x+nT)=f(x)\) for all \(n\in\mathbb{Z}\), \(f\) is integrable on bounded intervals by hypothesis, and the identity with \(h=nT\) reduces to the case just treated: \[\int_a^{a+T}f=\int_b^{b+T}f(x+nT)\,dx=\int_b^{b+T}f=\int_0^Tf .\]
Construct a set \(E\subset[0,1]\) with Lebesgue measure \(\alpha\in(0,1)\) such that the integral of the function \(|x-c|^{-1}\) over \(E\) is infinite for all \(c\in[0,1]\setminus E\).
Take for \(E\) the complement of a fat Cantor set whose gap ratios shrink only harmonically. The criterion is this: for measurable \(E\subset[0,1]\), \(c\in[0,1]\) and \(m(s):=\lambda(E\cap[c-s,c+s])\), the layer cake formula together with \(|x-c|^{-1}\ge1\) on \(E\) and the substitution \(s=1/u\) gives
\begin{equation*} \int_E\frac{dx}{|x-c|}\ \ge\ \int_1^\infty \lambda\bigl(E\cap(c-1/u,c+1/u)\bigr)\,du =\int_0^1\frac{m(s)}{s^2}\,ds, \tag{4} \end{equation*}
an open and the corresponding closed interval differing by two points.
Construction. Put \(\beta:=1-\alpha\in(0,1)\) and
\begin{equation*} \varepsilon_n:=\frac{\alpha}{\beta(n+1)},\qquad \ell_n:=\beta(1+\varepsilon_n)2^{-n}, \qquad n\ge0 , \end{equation*}
so \(\ell_0=1\), \(\varepsilon_n\downarrow0\) strictly, and \(\ell_n-2\ell_{n+1}=\beta2^{-n}(\varepsilon_n-\varepsilon_{n+1})>0\). Let \(F_0:=[0,1]\), and obtain \(F_{n+1}\) from the \(2^n\) disjoint closed intervals of length \(\ell_n\) constituting \(F_n\) by deleting from each the open middle interval of length \(\ell_n-2\ell_{n+1}\), leaving \(2^{n+1}\) closed intervals of length \(\ell_{n+1}\). Put \(F:=\bigcap_{n\ge0}F_n\) and \(E:=[0,1]\setminus F\). Since \(\lambda(F_n)=2^n\ell_n=\beta(1+\varepsilon_n)\to\beta\),
\begin{equation*} \lambda(F)=\beta,\qquad \lambda(E)=\alpha,\qquad [0,1]\setminus E=F . \end{equation*}
Divergence at every \(c\in F\). Let \(I_n\) be the interval of \(F_n\) containing \(c\), of length \(\ell_n\). For \(m\ge n\) the set \(F_m\cap I_n\) consists of \(2^{m-n}\) intervals of length \(\ell_m\), so \(\lambda(F\cap I_n)=\lim_m 2^{m-n}\ell_m=\beta2^{-n}\) and hence \(\lambda(E\cap I_n)=\ell_n-\beta2^{-n}=\beta\varepsilon_n2^{-n}\). As \(I_n\subset[c-\ell_n,c+\ell_n]\),
\begin{equation*} m(\ell_n)\ \ge\ \lambda(E\cap I_n)=\beta\varepsilon_n2^{-n}. \tag{5} \end{equation*}
The \(\ell_n\) decrease from \(1\) to \(0\) with \(\ell_{n+1}\le\ell_n/2\), so \((0,1]\) is the disjoint union of the \((\ell_{n+1},\ell_n]\) and, \(m\) being nondecreasing and \(\ell_{n+1}^{-1}-\ell_n^{-1}\ge(2\ell_{n+1})^{-1}\),
\begin{equation*} \int_0^1\frac{m(s)}{s^2}ds \ \ge\ \sum_{n\ge0}m(\ell_{n+1})\Bigl(\frac{1}{\ell_{n+1}}-\frac{1}{\ell_n}\Bigr) \ \ge\ \sum_{n\ge0}\frac{m(\ell_{n+1})}{2\ell_{n+1}} . \end{equation*}
By (5) and the definition of \(\ell_{n+1}\), each term is at least
\begin{equation*} \frac{\beta\varepsilon_{n+1}2^{-n-1}}{2\beta(1+\varepsilon_{n+1})2^{-n-1}} =\frac{\varepsilon_{n+1}}{2(1+\varepsilon_{n+1})} \ \ge\ \frac{\varepsilon_{n+1}}{2(1+\varepsilon_0)} , \end{equation*}
and \(\varepsilon_n=\alpha/(\beta(n+1))\) is harmonic, so the sum diverges. By (4), \(\int_E|x-c|^{-1}dx=\infty\) for every \(c\in[0,1]\setminus E\), while \(\lambda(E)=\alpha\).
(M.K. Gowurin) Let a function \(f\) be Lebesgue integrable on \([0,1]\) and let \(\alpha\in(0,1)\). Suppose that the integral of \(f\) over every set of measure \(\alpha\) is zero. Prove that \(f=0\) almost everywhere.
The whole integral vanishes first, and then atomlessness of \(\lambda\) forces \(f=0\) a.e.
(i) \(\int_0^1f=0\). Let \(F\) be the \(1\)-periodic extension of \(f|_{[0,1)}\), integrable on bounded intervals; Exercise 2.12.50 gives \(\int_k^{k+1}F=\int_0^1f\) for \(k\in\mathbb{Z}\), so
\begin{equation*} \int_0^{n}F(x)\,dx=n\int_0^1f(x)\,dx,\qquad n\in\mathbb{N}. \tag{6} \end{equation*}
Every interval of length \(\alpha\) has \(\int F=0\): by periodicity take \(a\in[0,1)\), and if \(a+\alpha\le1\) then \(\int_a^{a+\alpha}F=\int_a^{a+\alpha}f=0\) by hypothesis, while if \(a+\alpha>1\) then
\begin{equation*} \int_a^{a+\alpha}F=\int_{A_a}f=0,\qquad A_a:=[a,1)\cup[0,a+\alpha-1), \end{equation*}
since \(\lambda(A_a)=(1-a)+(a+\alpha-1)=\alpha\). Additivity then gives
\begin{equation*} \int_0^{m\alpha}F(x)\,dx=\sum_{j=0}^{m-1}\int_{j\alpha}^{(j+1)\alpha}F(x)\,dx=0, \qquad m\in\mathbb{N}. \tag{7} \end{equation*}
Given \(\eta>0\), absolute continuity of the integral (Theorem 2.5.7) supplies \(\varepsilon\in(0,1)\) with \(\int_D|f|\,d\lambda<\eta\) whenever \(\lambda(D)<\varepsilon\), and there are \(n,m\in\mathbb{N}\) with \(0\le n-m\alpha\le\varepsilon\): for \(\alpha=p/q\) take \(m=q\), \(n=p\); for irrational \(\alpha\) the fractional parts \(\{m\alpha\}\) are dense (Kronecker), so pick \(m\) with \(\{m\alpha\}>1-\varepsilon\) and \(n=\lceil m\alpha\rceil\). By (6) and (7),
\begin{equation*} n\int_0^1f(x)\,dx=\int_0^nF(x)\,dx=\int_{m\alpha}^{n}F(x)\,dx , \end{equation*}
and \([m\alpha,n]\) has length at most \(\varepsilon<1\), so its image modulo \(1\) is a set \(D\subset[0,1)\) with \(\lambda(D)\le\varepsilon\) and \(|\int_{m\alpha}^nF|\le\int_D|f|<\eta\). As \(n\ge1\) and \(\eta\) was arbitrary, \(\int_0^1f=0\).
(ii) We may assume \(\alpha\le1/2\), since \(\min(\alpha,1-\alpha)\le1/2\) and the hypothesis for \(\alpha\) implies it for \(1-\alpha\): if \(\lambda(A)=1-\alpha\), then \(\lambda([0,1]\setminus A)=\alpha\), so \(\int_Af=\int_0^1f-\int_{[0,1]\setminus A}f=0\) by (i).
(iii) If \(\lambda(\{f\ge0\})\ge\alpha\), then \(\lambda(P)=0\) for \(P:=\{f>0\}\). Suppose \(\lambda(P)>0\) and put \(Z:=\{f=0\}\), so \(\lambda(P)+\lambda(Z)\ge\alpha\). Lebesgue measure is atomless, so every measurable set contains subsets of any prescribed smaller measure (Example 1.12.8). If \(\lambda(P)\ge\alpha\), a set \(B\subset P\) with \(\lambda(B)=\alpha\) has \(\int_Bf=0\) by hypothesis and \(\int_Bf>0\) since \(f>0\) on \(B\) with \(\lambda(B)>0\) — impossible. If \(\lambda(P)<\alpha\), then \(\lambda(Z)\ge\alpha-\lambda(P)\), so some \(B^{\prime}\subset Z\) has \(\lambda(B^{\prime})=\alpha-\lambda(P)\) and \(B:=P\cup B^{\prime}\) has measure \(\alpha\), whence
\begin{equation*} 0=\int_Bf=\int_Pf+\int_{B^{\prime}}f=\int_Pf>0 , \end{equation*}
again impossible. Applied to \(-f\), which satisfies the same hypothesis, this gives: \(\lambda(\{f\le0\})\ge\alpha\) implies \(\lambda(\{f<0\})=0\).
(iv) One of \(\{f\ge0\}\), \(\{f\le0\}\) has measure at least \(1/2\ge\alpha\); say the first. Then (iii) gives \(f\le0\) a.e., so \(\lambda(\{f\le0\})=1\ge\alpha\) and the symmetric statement gives \(f\ge0\) a.e. Hence \(f=0\) a.e.
Exercises 2.12.53–2.12.59
Suppose that a function \(f\) is integrable on \([0,1]\) and \(f(x) > 0\) for all \(x\). Show that for each \(\varepsilon > 0\) there exists \(\delta > 0\) such that
\begin{equation*} \int_A f(x)\,dx \ge \delta \end{equation*}
for every set \(A\) with measure at least \(\varepsilon\).
Take \(\delta:=c\varepsilon/2\) with \(c\) chosen as follows (the claim is vacuous unless \(\varepsilon\le1\)). Since \(f>0\) at every point, the measurable sets \(B_c:=\{x\in[0,1]\colon f(x)\ge c\}\) satisfy \(B_{1/n}\uparrow[0,1]\), so \(\lambda(B_{1/n})\to1\) by continuity from below; fix \(c:=1/n\) with \[\lambda\bigl([0,1]\setminus B_c\bigr)<\varepsilon/2 .\] For measurable \(A\) with \(\lambda(A)\ge\varepsilon\), the inclusion \(A\setminus B_c\subset[0,1]\setminus B_c\) gives \(\lambda(A\cap B_c)=\lambda(A)-\lambda(A\setminus B_c)>\varepsilon/2\), so additivity over \(A=(A\cap B_c)\cup(A\setminus B_c)\) with \(f\ge0\), and \(f\ge c\) on \(B_c\) (Theorem 2.5.1(iv),(v)), give
\begin{equation*} \int_A f\,dx\ \ge\ \int_{A\cap B_c}f\,dx \ \ge\ c\,\lambda(A\cap B_c)\ \ge\ \frac{c\varepsilon}{2}=\delta . \end{equation*}
Let \(E \subset [0, 2\pi]\) be a set of Lebesgue measure \(d\) and let \(n \in \mathbb{N}\). Prove the inequality
\begin{equation*} \int_E |\cos(nx)|\,dx \ \ge\ \frac{d}{2}\,\sin\frac{d}{8}. \end{equation*}
Delete from \(E\) small intervals around the zeros of \(\cos(nx)\); on the rest \(|\cos(nx)|\ge\sin(d/8)\). Assume \(d>0\), so \(0<d\le2\pi\). The zeros in \([0,2\pi]\) are exactly
\begin{equation*} x_k:=\frac{\pi}{2n}+\frac{k\pi}{n},\qquad k=0,1,\dots,2n-1 , \end{equation*}
since \(\cos(nx)=0\) means \(nx=\pi/2+k\pi\). The numbers \(\pi/2+k\pi\), \(k=0,\dots,2n-1\), run from \(\pi/2\) to \(2\pi n-\pi/2\) in steps of \(\pi\), so every \(t=nx\in[0,2\pi n]\) is within \(\pi/2\) of one of them (the interior by the step size, the two end stretches by the first and last terms); that is, every \(x\in[0,2\pi]\) has some \(k\) with \(|x-x_k|\le\pi/(2n)\).
Let \(J\) be the union of the \(2n\) intervals of length \(d/(4n)\) centred at the \(x_k\), so \(\lambda(J)\le2n\cdot d/(4n)=d/2\). For \(x\in[0,2\pi]\setminus J\) and such a \(k\), put \(s:=n(x-x_k)\); then \(|s|\le\pi/2\) and \(|s|\ge d/8\), and
\begin{equation*} \cos(nx)=\cos\Bigl(\frac{\pi}{2}+k\pi+s\Bigr)=(-1)^{k+1}\sin s , \end{equation*}
so \(|\cos(nx)|=\sin|s|\ge\sin(d/8)\), the sine being nondecreasing on \([d/8,\pi/2]\) because \(d/8\le\pi/4\). Since \(\lambda(E\setminus J)\ge\lambda(E)-\lambda(J)\ge d/2\) and the integrand is nonnegative,
\begin{equation*} \int_E|\cos(nx)|\,dx\ \ge\ \int_{E\setminus J}|\cos(nx)|\,dx \ \ge\ \sin\frac{d}{8}\cdot\lambda(E\setminus J)\ \ge\ \frac{d}{2}\sin\frac{d}{8}. \end{equation*}
Let \(E \subset \mathbb{R}\) be a set of finite Lebesgue measure. Evaluate the limit
\begin{equation*} \lim_{k\to\infty}\int_E (2 - \sin kx)^{-1}\,dx . \end{equation*}
The limit is \(\lambda(E)/\sqrt3\). Put \(g(u):=(2-\sin u)^{-1}\), continuous, \(2\pi\)-periodic, with \(1/3\le g\le1\), so every integral below is finite for \(\lambda(E)<\infty\).
Mean value. By periodicity and the Weierstrass substitution \(t=\tan(u/2)\), under which \(\sin u=2t/(1+t^2)\) and \(du=2\,dt/(1+t^2)\), so that \(2-\sin u=2(t^2-t+1)/(1+t^2)\),
\begin{equation*} \int_0^{2\pi} g(u)\,du=\int_{-\infty}^{\infty}\frac{dt}{t^2-t+1} =\int_{-\infty}^{\infty}\frac{dt}{(t-\tfrac12)^2+\tfrac34}=\frac{2\pi}{\sqrt3}, \end{equation*}
by \(\int_{-\infty}^{\infty}(s^2+a^2)^{-1}ds=\pi/a\) with \(a=\sqrt3/2\); hence the mean is \(m=1/\sqrt3\).
Primitive. For \(G(T):=\int_0^T g\), write \(T=2\pi N+r\) with \(N\in\mathbb{Z}\), \(r\in[0,2\pi)\); periodicity gives \(\int_0^{2\pi N}g=2\pi mN\), so \(G(T)-mT=\int_0^r(g-m)\) and, since \(|g-m|\le1\),
\begin{equation*} |G(T)-mT|\le 2\pi\qquad\text{for all } T\in\mathbb{R}. \end{equation*}
Intervals. Substituting \(u=kx\), \(\int_a^b g(kx)\,dx=(G(kb)-G(ka))/k\), so subtracting \(m(b-a)\) and applying the previous bound twice gives
\begin{equation*} \Bigl|\int_a^b g(kx)\,dx-m(b-a)\Bigr|\le\frac{4\pi}{k}\ \xrightarrow[k\to\infty]{}\ 0 , \end{equation*}
and by additivity \(\int_U g(kx)\,dx\to m\lambda(U)\) for every finite disjoint union \(U\) of bounded intervals.
General \(E\). Given \(\eta>0\), regularity supplies an open \(G_0\supset E\) with \(\lambda(G_0\setminus E)<\eta/2\), hence \(\lambda(G_0)<\infty\); writing \(G_0=\bigcup_{j\ge1}(a_j,b_j)\) as a disjoint union of intervals with \(\sum_j(b_j-a_j)=\lambda(G_0)\), choose \(N\) with \(\sum_{j>N}(b_j-a_j)<\eta/2\) and put \(U:=\bigcup_{j\le N}(a_j,b_j)\), so that \(\lambda(E\triangle U)\le\lambda(G_0\setminus U)+\lambda(G_0\setminus E)<\eta\). Since \(0\le g\le1\),
\begin{equation*} \Bigl|\int_E g(kx)\,dx-\int_U g(kx)\,dx\Bigr|\le\lambda(E\triangle U)<\eta , \end{equation*}
and likewise \(|m\lambda(E)-m\lambda(U)|<\eta\); letting \(k\to\infty\) in the interval case gives
\begin{equation*} \limsup_{k\to\infty}\Bigl|\int_E g(kx)\,dx-m\lambda(E)\Bigr|\le2\eta . \end{equation*}
As \(\eta>0\) was arbitrary, \(\int_E(2-\sin kx)^{-1}dx\to\lambda(E)/\sqrt3\).
(\(\circ\)) Let \(\mu\) be a bounded nonnegative measure on a \(\sigma\)-algebra \(\mathcal{A}\). Prove that the definition of the Lebesgue integral given in the text is equivalent to the following definition. For simple functions we keep the same definition; for bounded measurable \(f\) we set
\begin{equation*} \int_X f\,d\mu = \lim_{n\to\infty}\int_X f_n\,d\mu , \end{equation*}
where \(\{f_n\}\) is an arbitrary sequence of simple functions uniformly convergent to \(f\); for nonnegative measurable functions \(f\) we set
\begin{equation*} \int_X f\,d\mu = \lim_{n\to\infty}\int_X \min(f,n)\,d\mu , \end{equation*}
and in the general case we declare \(f\) to be integrable if both functions \(f^{+} = \max(f,0)\) and \(f^{-} = -\min(f,0)\) are integrable, and we set
\begin{equation*} \int_X f\,d\mu = \int_X f^{+}\,d\mu - \int_X f^{-}\,d\mu . \end{equation*}
The two definitions produce the same class and the same values; write \(J(f)\) for the new three-clause quantity. Fix an everywhere defined \(\mathcal{A}_\mu\)-measurable \(f\) and read simple as \(\mathcal{A}_\mu\)-simple throughout (Lemma 2.1.8 is applied to \(\mathcal{A}_\mu\), as in the proof of Theorem 2.5.1(iii)); one cannot pass to an a.e. version here, since uniform convergence to \(f\) need not survive modification on a null set.
Clause 1. For bounded measurable \(f\), Lemma 2.1.8 supplies simple \(f_n\to f\) uniformly, so the clause is not vacuous; with \(\varepsilon_n:=\sup_X|f_n-f|\to0\) and \(|f_n-f_m|\le\varepsilon_n+\varepsilon_m\), Lemma 2.3.2 and Corollary 2.3.3 give
\begin{equation*} \Bigl|\int_X f_n\,d\mu-\int_X f_m\,d\mu\Bigr|\le(\varepsilon_n+\varepsilon_m)\mu(X) \ \xrightarrow[n,m\to\infty]{}\ 0 . \end{equation*}
So the integrals converge; interlacing two admissible sequences shows the limit is the same for both, and \(J(f)\) is well defined. The same display says \(\{f_n\}\) is fundamental in the mean with \(f_n\to f\) everywhere, so Definition 2.4.1 makes \(f\) Lebesgue integrable with \(\int_X f\,d\mu=J(f)\). For simple \(f\) take \(f_n\equiv f\), so clause 1 agrees with the retained definition.
Clause 2. For measurable \(f\ge0\) each \(\min(f,n)\) is bounded measurable, hence covered by clause 1, and the integrals \(\int_X\min(f,n)\,d\mu\) increase (Theorem 2.5.1(v)), so their limit exists in \([0,+\infty]\) and is finite exactly when \(\sup_n\int_X\min(f,n)\,d\mu<\infty\) — precisely Lebesgue integrability, by Proposition 2.5.5. When \(f\) is integrable, \(0\le\min(f,n)\le f\) with \(\min(f,n)\to f\) pointwise, so dominated convergence (Theorem 2.8.1 with majorant \(f\)) gives \(J(f)=\int_X f\,d\mu\). If \(\mu(\{f=+\infty\})>0\), then \(\int_X\min(f,n)\,d\mu\ge n\mu(\{f=+\infty\})\to\infty\) and \(f\) is integrable in neither sense. For bounded \(f\ge0\), \(\min(f,n)=f\) once \(n\ge\sup f\), so clauses 1 and 2 agree.
Clause 3. Here \(f^{\pm}\) are nonnegative and measurable with \(f=f^{+}-f^{-}\), \(|f|=f^{+}+f^{-}\). If \(f\) is integrable in the new sense, clause 2 makes \(f^{\pm}\) Lebesgue integrable with \(J(f^{\pm})=\int_X f^{\pm}d\mu\), so linearity (Theorem 2.5.1(iv)) gives
\begin{equation*} \int_X f\,d\mu=\int_X f^{+}d\mu-\int_X f^{-}d\mu=J(f). \end{equation*}
Conversely, if \(f\) is Lebesgue integrable, then \(|f|\) is (Theorem 2.5.1(ii)) and \(0\le f^{\pm}\le|f|\) with \(f^{\pm}\) measurable, so Corollary 2.5.6 makes \(f^{\pm}\) integrable, hence \(f\) integrable in the new sense with the same value. Clause 3 returns \(J(f)\) for \(f\ge0\) and the Lebesgue integral for bounded measurable \(f\), so it clashes with neither earlier clause.
(\(\circ\)) The purpose of this exercise is to show that our definition of the Lebesgue integral is equivalent to the following definition due to Lebesgue himself. Let \(\mu\) be a bounded nonnegative measure on a \(\sigma\)-algebra \(\mathcal{A}\) and let \(f\) be a measurable function. Let us fix \(\varepsilon > 0\) and consider the partition \(P\) of the real line into intervals \([y_i, y_{i+1})\), \(i \in \mathbb{Z}\), \(y_i < y_{i+1}\), of lengths not bigger than \(\varepsilon\). Let \(\delta(P) = \sup|y_{i+1} - y_i|\). Set
\begin{equation*} I(P) := \sum_{i=-\infty}^{+\infty} y_i\,\mu\big(\{x:\ y_i \le f(x) < y_{i+1}\}\big). \end{equation*}
Suppose that for some \(\varepsilon\) and \(P\) such a series converges (i.e., the series in positive and negative \(i\) converge separately). Show that this series converges for any partition and that, for any sequence of partitions \(P_k\) with \(\delta(P_k)\to 0\), there exists a finite limit \(\lim_{k\to\infty} I(P_k)\) independent of our choice of the sequence of partitions, moreover, the function \(f\) is integrable in the sense of our definition and its integral equals the above limit. Show that it suffices to consider points \(y_i = \varepsilon i\) or \(y_i = i/n\), \(n \in \mathbb{N}\).
Everything follows from the single estimate \(|\int_X f\,d\mu-I(P)|\le\delta(P)\mu(X)\). Assume, as the statement does, that \(f\) is finite everywhere and \(\mu\)-measurable, i.e. \(\mathcal{A}_\mu\)-measurable; setting \(f:=0\) on the exceptional null set of an a.e. finite function changes no term of \(I(P)\) and neither integrability nor the integral. A partition \(P\) is a strictly increasing \((y_i)_{i\in\mathbb{Z}}\) with \(y_i\to\pm\infty\) and \(\delta(P):=\sup_i(y_{i+1}-y_i)<\infty\); put
\begin{equation*} X_i^{P}:=\{x\colon y_i\le f(x)<y_{i+1}\}\in\mathcal{A}_\mu, \qquad g_P:=y_i \ \text{ on } X_i^{P}, \end{equation*}
so that the \(X_i^P\) are disjoint with union \(X\), \(g_P\) takes at most countably many values, and
\begin{equation*} 0\le f(x)-g_P(x)<\delta(P)\qquad\text{for all } x\in X. \tag{1} \end{equation*}
Convergence of \(I(P)=\sum_i y_i\mu(X_i^P)\) is automatically absolute: there is a unique \(i_0\) with \(y_i<0\) for \(i<i_0\) and \(y_i\ge0\) for \(i\ge i_0\), so each of the two halves has terms of constant sign and converges iff it converges absolutely; separate convergence is thus
\begin{equation*} \sum_{i\in\mathbb{Z}}|y_i|\,\mu(X_i^{P})<\infty . \tag{2} \end{equation*}
By Example 2.5.8 applied to \(g_P\), condition (2) says exactly that \(g_P\) is \(\mu\)-integrable, and then \(\int_X g_P\,d\mu=I(P)\).
Suppose \(I(P)\) converges for one partition \(P\). Then \(g_P\) is integrable, \(|f|\le|g_P|+\delta(P)\) by (1), constants are integrable as \(\mu\) is bounded (Theorem 2.5.1(iii)), and \(f\) is measurable, so Corollary 2.5.6 makes \(f\) integrable in the sense of Definition 2.4.1. Conversely, if \(f\) is integrable then \(|f|\) is (Theorem 2.5.1(ii)) and \(|g_Q|\le|f|+\delta(Q)\) for every partition \(Q\), so \(g_Q\) is integrable by Corollary 2.5.6 and Example 2.5.8 returns absolute convergence of \(I(Q)\). Hence convergence for one partition gives convergence for all.
In that case both \(f\) and \(g_P\) are integrable, so linearity with (1) and Theorem 2.5.1(ii),(iii) gives
\begin{equation*} \Bigl|\int_X f\,d\mu-I(P)\Bigr|=\Bigl|\int_X (f-g_P)\,d\mu\Bigr| \le\delta(P)\,\mu(X). \tag{3} \end{equation*}
Thus \(I(P_k)\to\int_X f\,d\mu\) for every sequence of partitions with \(\delta(P_k)\to0\), a finite limit independent of the sequence.
Uniform partitions suffice: if the series converges for \(y_i=\varepsilon i\) (mesh \(\varepsilon\)), the paragraph above makes \(f\) integrable, hence the series converges for every partition, and with \(y_i=i/n\),
\begin{equation*} \int_X f\,d\mu=\lim_{n\to\infty}\ \sum_{i=-\infty}^{+\infty}\frac{i}{n}\, \mu\Bigl(\Bigl\{x\colon \frac{i}{n}\le f(x)<\frac{i+1}{n}\Bigr\}\Bigr), \end{equation*}
the error at stage \(n\) being at most \(\mu(X)/n\) by (3).
(\(\circ\)) Let \(f\) be a bounded function on a space \(X\) with a bounded nonnegative measure \(\mu\). For every partition of \(X\) into disjoint measurable parts \(X_1,\dots,X_n\) we set
\begin{equation*} L(\{X_i\}) = \sum_{i=1}^{n}\ \inf_{x\in X_i} f(x)\,\mu(X_i), \qquad U(\{X_i\}) = \sum_{i=1}^{n}\ \sup_{x\in X_i} f(x)\,\mu(X_i). \end{equation*}
The lower integral \(I_{\ast}\) of the function \(f\) equals the supremum of the sums \(L(\{X_i\})\) over all possible finite partitions, and the upper integral \(I^{\ast}\) of \(f\) equals the infimum of the sums \(U(\{X_i\})\) over all possible finite partitions. The function \(f\) will be called integrable if \(I_{\ast} = I^{\ast}\). Prove that any function integrable in this sense is \(\mu\)-measurable and its Lebesgue integral equals \(I_{\ast} = I^{\ast}\). In addition, show that any bounded and \(\mu\)-measurable function \(f\) is integrable in the indicated sense.
For bounded \(f\), Darboux integrability in this sense is exactly \(\mu\)-measurability, and the two integrals agree. A partition means a finite family of disjoint nonempty sets of \(\mathcal{A}_\mu\) covering \(X\); put \(M:=\sup_X|f|<\infty\).
Four facts. (a) \(\inf_{X_i}f\le\sup_{X_i}f\) gives \(L\le U\), and \(|L|,|U|\le M\mu(X)\), so \(I_{\ast},I^{\ast}\in\mathbb{R}\). (b) The simple functions
\begin{equation*} \varphi=\sum_i\Bigl(\inf_{X_i}f\Bigr)I_{X_i}, \qquad \psi=\sum_i\Bigl(\sup_{X_i}f\Bigr)I_{X_i} \end{equation*}
satisfy \(\varphi\le f\le\psi\) with \(\int_X\varphi\,d\mu=L(\{X_i\})\) and \(\int_X\psi\,d\mu=U(\{X_i\})\). (c) If \(\{Y_j\}\) refines \(\{X_i\}\) then \(\inf_{Y_j}f\ge\inf_{X_i}f\) for \(Y_j\subset X_i\), so summing \(\mu(Y_j)\) over \(j\in J_i\) gives \(L(\{Y_j\})\ge L(\{X_i\})\), and symmetrically \(U(\{Y_j\})\le U(\{X_i\})\). (d) Any two partitions have the common refinement \(\{X_i\cap Z_j\}\), so with (a) and (c) every lower sum is at most every upper sum and \(I_{\ast}\le I^{\ast}\).
Bounded \(\mu\)-measurable \(f\) is Darboux integrable. Such \(f\) is Lebesgue integrable (Theorem 2.5.1(iii)). Fix \(n\), take \(a_0<\cdots<a_m\) with \(a_{k+1}-a_k=1/n\) and \([a_0,a_m)\supset f(X)\), and let \(\mathcal{P}_n\) consist of the nonempty \(X_k:=f^{-1}([a_k,a_{k+1}))\), all in \(\mathcal{A}_\mu\). Oscillation at most \(1/n\) on each part gives
\begin{equation*} 0\le U(\mathcal{P}_n)-L(\mathcal{P}_n)=\sum_k\Bigl(\sup_{X_k}f-\inf_{X_k}f\Bigr)\mu(X_k) \le\frac{\mu(X)}{n}, \tag{2} \end{equation*}
and \(L(\mathcal{P}_n)\le I_{\ast}\le I^{\ast}\le U(\mathcal{P}_n)\) by (d) forces \(I_{\ast}=I^{\ast}\). Moreover monotonicity (Theorem 2.5.1(v)) with (b) puts \(\int_X f\,d\mu\) in the same interval \([L(\mathcal{P}_n),U(\mathcal{P}_n)]\), so \(|\int_X f\,d\mu-I_{\ast}|\le\mu(X)/n\) for all \(n\) and \(I_{\ast}=I^{\ast}=\int_X f\,d\mu\).
Conversely, let \(I_{\ast}=I^{\ast}\). Choose partitions with \(L(\mathcal{Q}_n^{\prime})\ge I_{\ast}-1/(2n)\) and \(U(\mathcal{Q}_n^{\prime\prime})\le I^{\ast}+1/(2n)\), and let \(\mathcal{P}_n\) be the common refinement of \(\mathcal{Q}_1^{\prime},\mathcal{Q}_1^{\prime\prime},\dots,\mathcal{Q}_n^{\prime},\mathcal{Q}_n^{\prime\prime}\), so that \(\mathcal{P}_{n+1}\) refines \(\mathcal{P}_n\) and, by (c),
\begin{equation*} 0\le U(\mathcal{P}_n)-L(\mathcal{P}_n)\le\frac1n . \tag{3} \end{equation*}
Refinement makes the attached \(\varphi_n\) increase and the \(\psi_n\) decrease (on a part \(Y\subset X_i\), \(\varphi_{n+1}=\inf_Y f\ge\inf_{X_i}f=\varphi_n\)), all bounded by \(M\), so \(\varphi:=\lim_n\varphi_n\) and \(\psi:=\lim_n\psi_n\) exist, are \(\mathcal{A}_\mu\)-measurable, and satisfy \(\varphi\le f\le\psi\). Dominated convergence (Theorem 2.8.1, majorant \(2M\)) with (b) and (3) gives
\begin{equation*} \int_X(\psi-\varphi)\,d\mu =\lim_{n\to\infty}\bigl(U(\mathcal{P}_n)-L(\mathcal{P}_n)\bigr)=0 , \end{equation*}
so \(\psi=\varphi\) a.e. by Corollary 2.5.4; on the complement of \(N:=\{\varphi\ne\psi\}\), which is \(\mu\)-null, the squeeze forces \(f=\varphi\). Since \(\mathcal{A}_\mu\) is complete, every subset of \(N\) is measurable, so
\begin{equation*} \{f<c\}=\bigl(\{\varphi<c\}\setminus N\bigr)\cup\bigl(\{f<c\}\cap N\bigr)\in\mathcal{A}_\mu \end{equation*}
for every \(c\), i.e. \(f\) is \(\mu\)-measurable; being bounded it is Lebesgue integrable (Theorem 2.5.1(iii)). Finally
\begin{equation*} \int_X f\,d\mu=\int_X\varphi\,d\mu=\lim_{n\to\infty}\int_X\varphi_n\,d\mu =\lim_{n\to\infty}L(\mathcal{P}_n)=I_{\ast}, \end{equation*}
the first equality because \(f=\varphi\) a.e., the second by dominated convergence with majorant \(M\), the third by (b), and the last from \(I_{\ast}-1/(2n)\le L(\mathcal{P}_n)\le I_{\ast}\).
(MacNeille [642], Mikusinski [690]) Let \(\mathcal{R}\) be an algebra (or semialgebra) of sets in a space \(X\) and let \(\mu\) be a probability measure on \(\mathcal{A} = \sigma(\mathcal{R})\). Prove that the function \(f\) is integrable with respect to \(\mu\) precisely when there exists a sequence of \(\mathcal{R}\)-simple functions \(\psi_k\) (i.e., finite linear combinations of indicators of sets in \(\mathcal{R}\)) such that
\begin{equation*} \sum_{k=1}^{\infty}\int_X |\psi_k|\,d\mu < \infty \end{equation*}
and \(f(x) = \sum_{k=1}^{\infty}\psi_k(x)\) for every \(x\) such that the above series converges absolutely. In addition,
\begin{equation*} \int_X f\,d\mu = \sum_{k=1}^{\infty}\int_X \psi_k\,d\mu . \end{equation*}
The series \(\sum_k\psi_k\) is a telescoped \(L^1\)-approximation by \(\mathcal{R}\)-simple functions, padded by \(\pm\) indicators of a covering of the exceptional null set. If \(\mathcal{R}\) is a semialgebra, Lemma 1.2.14 makes every set of the generated algebra a finite disjoint union of sets of \(\mathcal{R}\), so the two classes of simple functions and the generated \(\sigma\)-algebras coincide; assume \(\mathcal{R}\) is an algebra. Let \(\mu^{\ast}\) be the outer measure generated by \(\mu|_{\mathcal{R}}\). By Theorem 1.5.6 that restriction has a unique countably additive extension to \(\sigma(\mathcal{R})\), so
\begin{equation*} \mu^{\ast}=\mu\ \text{ on }\mathcal{A},\qquad \mathcal{A}\subset\mathcal{R}_\mu ; \tag{0} \end{equation*}
by Definition 1.5.1 this means every \(A\in\mathcal{A}\) admits \(R\in\mathcal{R}\) with \(\mu(A\triangle R)<\varepsilon\), and every \(\mathcal{A}\)-set of measure zero is covered, for each \(j\), by countably many sets of \(\mathcal{R}\) of total measure below \(2^{-j}\).
Sufficiency. Let \(M:=\sum_k\int_X|\psi_k|\,d\mu<\infty\) and \(S:=\sum_k|\psi_k|\). The simple partial sums \(S_n\) increase with \(\int_X S_n\,d\mu\le M\), so Chebyshev (Theorem 2.5.3) gives \(\mu(\{S_n>R\})\le M/R\) and hence \(\mu(\{S>R\})\le M/R\) for the increasing union; letting \(R\to\infty\), the set \(N:=\{S=\infty\}\in\mathcal{A}\) is null. Off \(N\) the series converges absolutely, so \(f=\sum_k\psi_k\) there by hypothesis; the simple functions \(T_n:=\sum_{k\le n}\psi_k\) therefore converge to \(f\) a.e., and
\begin{equation*} \int_X|T_n-T_m|\,d\mu\le\sum_{k>m}\int_X|\psi_k|\,d\mu\ \xrightarrow[m\to\infty]{}\ 0 \end{equation*}
makes \(\{T_n\}\) fundamental in the mean. By Definition 2.4.1, \(f\) is integrable with \(\int_X f\,d\mu=\lim_n\int_X T_n\,d\mu=\sum_k\int_X\psi_k\,d\mu\), the last series absolutely convergent since \(|\int_X\psi_k\,d\mu|\le\int_X|\psi_k|\,d\mu\).
Density. The \(\mathcal{R}\)-simple functions are dense in \(L^1(\mu)\). First the \(\mathcal{A}\)-simple ones are: for integrable \(f\) take \(\{g_n\}\) as in Definition 2.4.1; for fixed \(m\) the simple functions \(|g_n-g_m|\) are fundamental in the mean (by \(||g_n-g_m|-|g_j-g_m||\le|g_n-g_j|\)) and converge a.e. to \(|f-g_m|\), so Definition 2.4.1 gives
\begin{equation*} \int_X|f-g_m|\,d\mu=\lim_{n\to\infty}\int_X|g_n-g_m|\,d\mu\le\varepsilon\qquad (m\ge N), \end{equation*}
where \(N\) is chosen from the fundamentality of \(\{g_n\}\). Second, for \(\varphi=\sum_{i\le p}c_iI_{A_i}\) pick \(R_i\in\mathcal{R}\) with \(\mu(A_i\triangle R_i)<\eta\) by (0); then \(\chi:=\sum_{i\le p}c_iI_{R_i}\) is \(\mathcal{R}\)-simple with \(\int_X|\varphi-\chi|\,d\mu<\eta\sum_i|c_i|\).
Necessity. Let \(f\) be integrable, choose \(\mathcal{R}\)-simple \(\varphi_k\) with \(\|f-\varphi_k\|_{L^1(\mu)}<2^{-k-1}\), and set \(\varphi_0:=0\), \(g_k:=\varphi_k-\varphi_{k-1}\), so \(\sum_{k\le n}g_k=\varphi_n\) and
\begin{equation*} \sum_{k=1}^{\infty}\int_X|g_k|\,d\mu\le\|f\|_{L^1(\mu)}+2^{-2} +\sum_{k\ge2}(2^{-k-1}+2^{-k})<\infty . \tag{1} \end{equation*}
The Chebyshev argument above makes \(N:=\{\sum_k|g_k|=\infty\}\in\mathcal{A}\) null, and applied to \(\sum_k\|f-\varphi_k\|_{L^1(\mu)}<\infty\) it gives \(\varphi_k\to f\) a.e., i.e. \(\sum_k g_k=f\) a.e. Hence
\begin{equation*} E:=\Bigl\{x\colon \sum_k|g_k(x)|<\infty \ \text{ and }\ f(x)\ \text{is undefined or}\ f(x)\ne\textstyle\sum_k g_k(x)\Bigr\} \end{equation*}
is contained in an \(\mathcal{A}_\mu\)-null set, hence in some \(E^{\prime}\in\mathcal{A}\) with \(\mu(E^{\prime})=0\), by the covering property in (0) intersected over \(j\). If \(E^{\prime}=\varnothing\), take \(\psi_k:=g_k\) and we are done by (1) and the definition of \(E\).
Otherwise cover \(E^{\prime}\), for each \(j\), by \(R_{j,m}\in\mathcal{R}\) with \(\sum_m\mu(R_{j,m})<2^{-j}\), and arrange the double family into one sequence \(R_1,R_2,\dots\), so that \(\sum_k\mu(R_k)\le1\) and every point of \(E^{\prime}\) lies in infinitely many \(R_k\). Put
\begin{equation*} \psi_{3k-2}:=g_k,\qquad \psi_{3k-1}:=I_{R_k},\qquad \psi_{3k}:=-I_{R_k}, \end{equation*}
all \(\mathcal{R}\)-simple, with \(\sum_k\int_X|\psi_k|\,d\mu=\sum_k\int_X|g_k|\,d\mu+2\sum_k\mu(R_k)<\infty\). If \(\sum_k|\psi_k(x)|<\infty\), then \(x\notin E^{\prime}\) (a point of \(E^{\prime}\) contributes infinitely many terms equal to \(1\)), hence \(x\notin E\), while \(\sum_k|g_k(x)|<\infty\); so \(f(x)=\sum_k g_k(x)\), and grouping the absolutely convergent series in blocks of three gives \(\sum_k\psi_k(x)=\sum_k g_k(x)=f(x)\). The same grouping gives
\begin{equation*} \sum_{k=1}^{\infty}\int_X\psi_k\,d\mu=\sum_{k=1}^{\infty}\int_X g_k\,d\mu =\lim_{n\to\infty}\int_X\varphi_n\,d\mu=\int_X f\,d\mu , \end{equation*}
the last step from \(|\int_X f\,d\mu-\int_X\varphi_n\,d\mu|\le\|f-\varphi_n\|_{L^1(\mu)}\to0\).
Exercises 2.12.60–2.12.66
(F. Riesz) Denote by \(C_0\) the class of all step functions on \([0,1]\), i.e., functions that are constant on intervals from certain finite partitions of \([0,1]\). Let \(C_1\) denote the class of all functions \(f\) on \([0,1]\) for which there exists an increasing sequence of functions \(f_n \in C_0\) such that \(f_n(x) \to f(x)\) a.e. and the Riemann integrals of \(f_n\) are uniformly bounded. The limit of the Riemann integrals of \(f_n\) is denoted by \(L(f)\). Finally, let \(C_2\) denote the class of all differences \(f = f_1 - f_2\) with \(f_1, f_2 \in C_1\) and let \(L(f) = L(f_1) - L(f_2)\). Prove that the class \(C_2\) coincides with the class of Lebesgue integrable functions and that \(L(f)\) is the Lebesgue integral of \(f\).
\(C_2=\mathcal{L}^1[0,1]\) and \(L(f)=\int f\), where \(\int u\) is the Lebesgue integral over \([0,1]\), equal to the Riemann integral on \(C_0\).
\(C_1\subset\mathcal{L}^1\) with \(L(f)=\int f\). Let \(f_n\in C_0\) increase to \(f\) a.e. with \(\int f_n\le C\). Then \(f\) equals the Borel function \(\limsup_n f_n\) off a null set, hence is measurable by completeness, and the nonnegative \(f_n-f_1\) increase a.e. to \(f-f_1\), so monotone convergence (Theorem 2.8.2) gives
\begin{equation*} \int (f-f_1)=\lim_{n\to\infty}\int(f_n-f_1)\le C-\int f_1<\infty . \end{equation*}
As \(f_1\) is a bounded step function, \(f\) is integrable with \(\int f=\lim_n\int f_n=L(f)\); in particular the limit is independent of the approximating sequence, so \(L\) is well defined on \(C_1\). Consequently, for \(f=f_1-f_2\) with \(f_i\in C_1\) the function \(f\) is integrable with \(\int f=L(f_1)-L(f_2)\), so that value depends only on \(f\): \(C_2\subset\mathcal{L}^1\) and \(L=\int\) on \(C_2\). Only \(\mathcal{L}^1\subset C_2\) remains.
Closure properties. (a) \(C_0\subset C_1\) (constant sequences). (b) \(f,g\in C_1\) and \(c\ge0\) give \(f+g, cf\in C_1\) (add or scale the approximating sequences), so \(C_2=C_1-C_1\) is a vector space, using \(c(f_1-f_2)=(-c)f_2-(-c)f_1\) for \(c<0\). (c) \(u=0\) a.e. lies in \(C_1\) with \(L(u)=0\) (take \(f_n\equiv0\)); hence \(g=f\) a.e. with \(f\in C_2\) gives \(g\in C_2\) and \(L(g)=L(f)\). (d) For \(U\subset[0,1]\) open, \(I_U\in C_1\): write \(U\) as an at most countable disjoint union of intervals \(J_j\) and let \(\varphi_n:=\sum_{j\le n}I_{J_j}\in C_0\) increase to \(I_U\), with \(\int\varphi_n\le1\). (e) If \(b\in C_1\) and \(\varphi\in C_0\), then \(b-\varphi\in C_1\) (subtract \(\varphi\) from the approximating sequence).
Monotone limits stay in \(C_1\): if \(g_k\in C_1\) increase a.e. to \(g\) with \(L(g_k)\le C\), then \(g\in C_1\) and \(L(g)=\lim_k L(g_k)\). Indeed, pick \(\varphi_{k,n}\in C_0\) increasing in \(n\) with \(\varphi_{k,n}\to g_k\) a.e., and set \(h_n:=\max_{1\le k\le n}\varphi_{k,n}\in C_0\), which increases everywhere. Off the union \(N\) of the countably many exceptional null sets, \(\varphi_{k,n}\le g_k\le g_n\) for \(k\le n\) gives \(h_n\le g_n\) and so \(\lim_n h_n\le g\), while \(h_n\ge\varphi_{k,n}\) for \(n\ge k\) gives \(\lim_n h_n\ge g_k\) for every \(k\), hence \(\lim_n h_n\ge g\); thus \(h_n\to g\) a.e. with \(\int h_n\le\int g_n=L(g_n)\le C\), so \(g\in C_1\) and \(L(g)=\lim_n\int h_n=\int g=\lim_k L(g_k)\) by monotone convergence.
Splitting. If \(v\in C_2\), \(v\ge0\) a.e. and \(\varepsilon>0\), then \(v=p-q\) with \(p,q\in C_1\), \(q\ge0\) a.e. and \(L(q)<\varepsilon\): write \(v=a-b\) with \(a,b\in C_1\), take \(b_n\in C_0\) increasing to \(b\) a.e., fix \(m\) with \(L(b)-\int b_m<\varepsilon\), and put \(q:=b-b_m\), \(p:=a-b_m\), both in \(C_1\) by (e), with \(q\ge0\) a.e. since \(b_m\le b_n\to b\).
Monotone limits stay in \(C_2\): if \(u_n\in C_2\) increase a.e. with \(L(u_n)\le C\), then \(u_n\to u\) a.e. with \(u\) a.e. finite, any function equal a.e. to \(u\) lies in \(C_2\), and \(L(u)=\lim_n L(u_n)\). Indeed \(u_n-u_1\ge0\) increases with integrals at most \(C-L(u_1)\), so monotone convergence makes the limit integrable, hence a.e. finite. Split \(v_k:=u_{k+1}-u_k\ge0\) as \(p_k-q_k\) with \(q_k\ge0\) a.e., \(L(q_k)<2^{-k}\), so \(p_k=v_k+q_k\ge0\) a.e.; the partial sums \(Q_n,P_n\in C_1\) increase a.e. with \(L(Q_n)<1\) and
\begin{equation*} L(P_n)=L(Q_n)+L(u_{n+1})-L(u_1)\le 1+C-L(u_1), \end{equation*}
so \(Q:=\lim_n Q_n\) and \(P:=\lim_n P_n\) lie in \(C_1\) by the previous paragraph (their a.e. limits being finite by monotone convergence). Since \(P_n-Q_n=u_{n+1}-u_1\), we get \(u=u_1+P-Q\) a.e., so \(u\in C_2\) by (b),(c), and \(L(u)=\int u=\lim_n\int u_n=\lim_n L(u_n)\). Applying this to \(-u_n\) covers decreasing sequences with \(L(u_n)\) bounded below.
Every integrable function lies in \(C_2\). (i) For \(G=\bigcap_n U_n\) with \(U_n\) open, replacing \(U_n\) by \(U_1\cap\cdots\cap U_n\) makes \(I_{U_n}\in C_1\) decrease to \(I_G\) with \(L(I_{U_n})=\lambda(U_n)\ge0\), so \(I_G\in C_2\). (ii) For measurable \(A\), regularity (Theorem 1.4.8) gives open \(U_k\supset A\) with \(\lambda(U_k)<\lambda(A)+1/k\), so \(G:=\bigcap_k(U_k\cap[0,1])\) is a \(G_\delta\) with \(\lambda(G\setminus A)=0\) and \(I_A=I_G-I_{G\setminus A}\in C_2\) by (i) and (c). (iii) Hence every measurable simple function lies in \(C_2\), by (b). (iv) For integrable \(f\ge0\), altered on a null set to be everywhere finite and measurable, the simple functions \(s_n:=\min(n,2^{-n}\lfloor2^nf\rfloor)\) increase to \(f\) pointwise (cf. Corollary 2.1.9) with \(L(s_n)=\int s_n\le\int f\), so \(f\in C_2\). (v) In general \(f=f^{+}-f^{-}\in C_2\) by (iv) and (b).
(\(\circ\)) Let us define the integral of a bounded measurable function \(f\) on \([0,1]\) as follows. First we define the integral of a continuous function \(g\) over a closed set \(E\) as the difference between the integral of \(g\) over \([0,1]\) and the sum of the series of the integrals of \(g\) over finitely or countably many disjoint intervals forming \([0,1]\setminus E\). Given a closed set \(E\), the integral over \(E\) of any function \(\varphi\) that is continuous on \(E\) is defined as the integral over \(E\) of its arbitrary continuous extension to \([0,1]\) (it is easily seen that this integral is independent of our choice of extension). Next we take a sequence of closed sets \(E_n\) with \(\lambda(E_n) \to 1\) such that on each of them \(f\) is continuous, and define the integral of \(f\) over \([0,1]\) as the limit of the integrals of \(f\) over the sets \(E_n\). Prove that this limit exists and equals the Lebesgue integral of \(f\).
For closed \(E \subset [0,1]\) and continuous \(g\) the quantity defined in the exercise is the Lebesgue integral over \(E\):
\begin{equation*} J_E(g) := \int_0^1 g\,dx - \sum_j \int_{I_j} g\,dx = \int_E g\,d\lambda , \end{equation*}
where \([0,1]\setminus E = \bigcup_j I_j\) is the decomposition into disjoint intervals: countable additivity splits \(\int_0^1 g\,d\lambda\) along the disjoint sets \(E,I_1,I_2,\dots\) (absolutely, since \(|\int_{I_j}g|\le\sup|g|\,\lambda(I_j)\) and \(\sum_j\lambda(I_j)\le1\)), and on an interval the Riemann and Lebesgue integrals of a continuous function agree. Consequently, for \(\varphi\) continuous on \(E\),
\begin{equation*} J_E(\varphi) = \int_E \varphi \, d\lambda \tag{1} \end{equation*}
for every continuous extension of \(\varphi\) (one exists by Lemma 2.2.8, \(E\) being compact, and two of them \(g_1,g_2\) give \(J_E(g_1)-J_E(g_2)=\int_E(g_1-g_2)\,d\lambda=0\)). Sequences \(\{E_n\}\) of the required kind exist by Lusin’s theorem (Theorem 2.2.10, \(f\) measurable on the finite-measure \([0,1]\)): compact \(E_n\) with \(\lambda([0,1]\setminus E_n)<1/n\) and \(f|_{E_n}\) continuous. For any such sequence, with \(|f|\le M\), (1) gives
\begin{equation*} \Bigl| \int_0^1 f \, d\lambda - J_{E_n}(f|_{E_n}) \Bigr| = \Bigl| \int_{[0,1]\setminus E_n} f \, d\lambda \Bigr| \le M\bigl(1 - \lambda(E_n)\bigr) \to 0 . \end{equation*}
So the limit exists, equals \(\int_0^1 f\,d\lambda\), and is independent of the choice of \(\{E_n\}\).
A function \(g\) on \(\mathbb{R}^d\) with values in \([-\infty, +\infty]\) is called lower semicontinuous if, for every \(c \in [-\infty, +\infty]\), the set \(\{x : g(x) > c\}\) is open. Let \(E \subset \mathbb{R}^d\) be a measurable set and let a function \(f \colon E \to \mathbb{R}^1\) be integrable. Prove that, for any \(\varepsilon > 0\), there exists a lower semicontinuous function \(g\) on \(\mathbb{R}^d\) such that \(g(x) \ge f(x)\) for all \(x \in E\), \(g|_E\) is integrable and the integral of \(g - f\) over \(E\) does not exceed \(\varepsilon\).
Take \(g(x) := \sup\{q_n : x \in G_n\}\) (with \(\sup\emptyset = -\infty\)), where \(\{q_n\}\) enumerates \(\mathbb{Q}\), \(B_n\) is the ball of radius \(n\) about the origin,
\begin{equation*} E_n := B_n \cap \{x \in E : f(x) \ge q_n\}, \end{equation*}
and \(G_n \supset E_n\) is open with \(\lambda(G_n \setminus E_n) < \delta_n\), available by outer regularity of \(\lambda\) (Theorem 1.4.8; \(\lambda(E_n)\le\lambda(B_n)<\infty\)). Here \(\delta>0\) comes from absolute continuity of the integral of the integrable \(f\) (Theorem 2.5.7), so that \(\int_A|f|\,d\lambda<\varepsilon/2\) for measurable \(A\subset E\) with \(\lambda(A)<\delta\), and
\begin{equation*} \delta_n := 2^{-n-1}\min\Bigl(\delta,\ \frac{\varepsilon}{1+|q_n|}\Bigr), \qquad \sum_n \delta_n \le \frac{\delta}{2}, \qquad \sum_n \delta_n |q_n| \le \frac{\varepsilon}{2} . \end{equation*}
Then \(\{g>c\}=\bigcup_{n:\,q_n>c}G_n\) for every \(c\) (both inclusions read off the supremum), so \(g\) is lower semicontinuous; and \(g\ge f\) on \(E\) since for \(x\in E\), \(r>0\) any \(n\) with \(q_n\in[f(x)-r,f(x)]\) and \(n>|x|\) has \(x\in E_n\subset G_n\), whence \(g(x)\ge q_n\ge f(x)-r\).
For the estimate put
\begin{equation*} D := \bigcup_n \bigl((E\cap G_n)\setminus E_n\bigr), \qquad h := \sum_n |q_n|\, I_{(E\cap G_n)\setminus E_n} . \end{equation*}
Then \(\lambda(D) \le \sum_n \delta_n < \delta\) with \(D\subset E\), so \(\int_D|f|\,d\lambda<\varepsilon/2\), while monotone convergence applied to the partial sums gives
\begin{equation*} \int_E h \, d\lambda = \sum_n |q_n|\,\lambda\bigl((E\cap G_n)\setminus E_n\bigr) \le \sum_n |q_n|\delta_n \le \varepsilon/2 . \end{equation*}
Write \(F := f + h + |f| I_D \ge f\) on \(E\) and fix \(x\in E\) and \(n\) with \(x\in G_n\). (i) If \(x\in E_n\) then \(q_n \le f(x)\le F(x)\). (ii) If \(x\notin E_n\) then \(x \in D\) and \(h(x)\ge|q_n|\ge q_n\), so \(F(x)=f(x)+|f(x)|+h(x)\ge h(x)\ge q_n\). Taking the supremum over such \(n\) gives \(g \le F\) on \(E\), i.e.
\begin{equation*} 0 \le g - f \le h + |f| I_D \qquad \text{on } E . \end{equation*}
The majorant is integrable with integral at most \(\varepsilon\), so \(g|_E = f + (g-f)\) is integrable and \(\int_E (g-f)\,d\lambda \le \varepsilon\).
(Hahn) Let \(f \in L^1[0,1]\), let \(I\) be the integral of \(f\), and let \(\{\Pi_n\}\) be a decreasing sequence of finite partitions of \([0,1]\) into intervals \(J_{n,k}\) (\(k \le N_n\)) with \(\lambda(J_{n,k}) \le \delta_n \to 0\), where \(\lambda\) is Lebesgue measure. Show that there exist points \(\xi_{n,k} \in J_{n,k}\) such that
\begin{equation*} \Big| \sum_{k=1}^{N_n} f(\xi_{n,k}) \lambda(J_{n,k}) - I \Big| \to 0 \quad \text{as } n \to \infty . \end{equation*}
Take \(\xi_{n,k}\) in \(J_{n,k} \cap \{f_{l(n)} = f\}\) whenever that set is nonempty, where \(f_l\) are continuous functions with \(\|f_l-f\|_{L^1}\to 0\) and \(\lambda(A_l)\to 0\) for \(A_l:=\{f_l \ne f\}\), and \(l(n)\) is specified below. Fix a real representative of \(f\) (Riemann sums depend on it) and assume \(\delta_n \downarrow 0\) (replace \(\delta_n\) by \(\sup_{m\ge n}\delta_m\)); nestedness of the \(\Pi_n\) is not used.
Such \(f_l\) exist: since \(|f|\) is integrable, choose \(M_l\) with \(\int_{\{|f|>M_l\}}|f|\,d\lambda<2^{-l}\) and \(\lambda(|f|>M_l)<2^{-l}\), put \(g_l:=\max(-M_l,\min(f,M_l))\), and apply Lusin’s theorem (Theorem 2.2.10, \(g_l\) measurable on the finite-measure \([0,1]\)) to get a compact \(K_l\) with \(\lambda([0,1]\setminus K_l)<\eta_l:=2^{-l}(1+2M_l)^{-1}\) and a continuous \(f_l=g_l\) on \(K_l\), truncated so that \(|f_l|\le M_l\). Then
\begin{equation*} \|f_l - f\|_{L^1} \le 2M_l\,\lambda([0,1]\setminus K_l) + 2^{-l} < 2^{-l+1}, \end{equation*}
and \(A_l \subset ([0,1]\setminus K_l)\cup\{|f|>M_l\}\) gives \(\lambda(A_l) < 2^{-l+1}\).
By uniform continuity pick \(\rho_l>0\) with \(|f_l(t)-f_l(s)|\le 1/l\) for \(|t-s|\le\rho_l\), then \(p_1=1<p_2<\cdots\) with \(\delta_{p_l}\le\rho_l\); as \(\delta_n\downarrow\), every interval of \(\Pi_n\) with \(n\ge p_l\) carries \(f_l\)-oscillation at most \(1/l\). Let \(l=l(n)\) be defined by \(p_l \le n < p_{l+1}\). Call \(k\) bad if \(J_{n,k}\cap\{f_l=f\}=\emptyset\), i.e. \(J_{n,k}\subset A_l\); for such \(k\) take \(\xi_{n,k}\) with \(|f(\xi_{n,k})|\le\inf_{J_{n,k}}|f|+1\) (the infimum is finite as \(\int_{J_{n,k}}|f|\,d\lambda<\infty\)).
With \(\sigma_n(u):=\sum_k u(\xi_{n,k})\lambda(J_{n,k})\),
\begin{equation*} |\sigma_n(f) - I| \le |\sigma_n(f)-\sigma_n(f_l)| + \Bigl|\sigma_n(f_l) - \int_0^1 f_l\,d\lambda\Bigr| + \|f_l-f\|_{L^1} . \end{equation*}
The middle term equals \(|\sum_k \int_{J_{n,k}}(f_l(\xi_{n,k})-f_l(t))\,dt| \le 1/l\) by that oscillation bound. In the first, non-bad indices contribute \(0\) since there \(f(\xi_{n,k})=f_l(\xi_{n,k})\), while for bad \(k\)
\begin{equation*} |f(\xi_{n,k})|\lambda(J_{n,k}) \le \int_{J_{n,k}}|f|\,d\lambda + \lambda(J_{n,k}), \quad |f_l(\xi_{n,k})|\lambda(J_{n,k}) \le \int_{J_{n,k}}|f_l|\,d\lambda + \frac{\lambda(J_{n,k})}{l}, \end{equation*}
the second again by the oscillation bound. The bad intervals are nonoverlapping inside \(A_l\), so summing and using \(\int_{A_l}|f_l|\le\int_{A_l}|f|+\|f_l-f\|_{L^1}\),
\begin{equation*} |\sigma_n(f)-\sigma_n(f_l)| \le 2\int_{A_l}|f|\,d\lambda + \|f_l-f\|_{L^1} + \lambda(A_l) + \frac{1}{l} . \end{equation*}
Hence \(|\sigma_n(f)-I| \le \varepsilon_l := 2\int_{A_l}|f|\,d\lambda + 2\|f_l-f\|_{L^1} + \lambda(A_l) + 2/l\) for \(p_l\le n<p_{l+1}\). Here \(\int_{A_l}|f|\,d\lambda\to 0\) by absolute continuity of the integral (Theorem 2.5.7) and \(\lambda(A_l)\to 0\); so \(\varepsilon_l\to 0\), and \(l(n)\to\infty\) gives \(|\sigma_n(f)-I| \le \varepsilon_{l(n)}\to 0\).
(Darji, Evans) Let a function \(f\) be integrable on the unit cube \(I \subset \mathbb{R}^n\). Show that there exists a sequence \(\{x_k\}\) that is everywhere dense in \(I\) and has the following property: for every \(\varepsilon > 0\), there exists \(\delta > 0\) such that for every partition \(\mathcal{P}\) of the cube \(I\) into finitely many parallelepipeds of the form \([a_1,b_1]\times \cdots \times [a_n,b_n]\) with pairwise disjoint interiors and \(|b_i - a_i| < \delta\), one has
\begin{equation*} \Big| \sum_{P \in \mathcal{P}} f\big(r(P)\big) \lambda_n(P) - \int_I f(x) \, dx \Big| < \varepsilon , \end{equation*}
where \(r(P)\) is the first element in \(\{x_k\}\) belonging to \(P\).
This is a theorem of Darji and Evans, and the solution given here is partial: proved below are two structural lemmas, the full assertion for Riemann integrable \(f\) with an arbitrary dense sequence, a sufficiency criterion (Lemma C), and Remark D showing that Lemma C cannot reach a general integrable \(f\). Write \(\lambda:=\lambda_n\); a partition is a finite family of boxes \(\prod_i[a_i,b_i]\subset I\) with disjoint interiors and union \(I\), admissible for \(\delta\) if all edges are shorter than \(\delta\); \(\partial\)-sets are null, so \(\sum_{P}\lambda(P)=1\) and \(\int_I u = \sum_P \int_P u\).
Lemma A. A point \(y\) lies in at most \(2^n\) members of a partition \(\mathcal{P}\): for \(P\ni y\) let \(Q(P)\) be the set of open orthants at \(y\) meeting \(\operatorname{int}P\), which is nonempty and near \(y\) recovers \(\operatorname{int}P\), so disjoint interiors force disjoint \(Q(P)\), and there are \(2^n\) orthants. Hence for \(F=\{x_1,\dots,x_N\}\) and \(\mathcal{P}\) admissible for \(\delta\), the union \(U_F\) of the members tagged by \(F\) has
\begin{equation*} \lambda(U_F) \le 2^n N \delta^n \longrightarrow 0 \quad (\delta \to 0), \end{equation*}
so no finite initial segment of \(\{x_k\}\) influences the sums for small \(\delta\).
Lemma B. If \(\mathcal{R}\) is a partition of \(I\) into boxes and \(\Gamma:=\bigcup_{R}\partial R\), then any member \(P\) of a \(\delta\)-admissible \(\mathcal{P}\) not contained in one member of \(\mathcal{R}\) meets \(\Gamma\) (else the connected set \(P\) would lie in the disjoint open union \(\bigcup_R \operatorname{int}R\), hence in a single \(\operatorname{int}R\)); since \(\operatorname{diam}P\le\delta\sqrt n\), all such members lie in the closed \(\delta\sqrt n\)-neighbourhood \(\Gamma_\delta\), and \(\Gamma_\delta \downarrow \Gamma\) gives \(\lambda(\Gamma_\delta)\to 0\), so \(\int_{\Gamma_\delta}|f|\,d\lambda \to 0\) by Theorem 2.5.7.
Riemann integrable \(f\) (complete): every dense \(\{x_k\}\) works. Given \(\varepsilon>0\), the Darboux criterion supplies a partition \(\mathcal{R}\) with \(\sum_R \operatorname{osc}_R(f)\lambda( R)<\varepsilon/2\); with \(|f|\le M\) choose \(\delta\) with \(2M\lambda(\Gamma_\delta)<\varepsilon/2\). For \(\mathcal{P}\) admissible for \(\delta\),
\begin{equation*} \sum_{P} f(r(P))\lambda(P) - \int_I f\,d\lambda = \sum_{P} \Bigl( f(r(P))\lambda(P) - \int_P f\,d\lambda \Bigr), \end{equation*}
where a term with \(P\subset R\) is at most \(\operatorname{osc}_R(f)\lambda(P)\), summing to \(<\varepsilon/2\), and the remaining members lie in \(\Gamma_\delta\) by Lemma B, contributing at most \(2M\lambda(\Gamma_\delta)<\varepsilon/2\).
Lemma C. Suppose there are partitions \(\mathcal{R}_m\), constants \(c_R\) with \(C_m:=\max_{R\in\mathcal{R}_m}|c_R|<\infty\), numbers \(\eta_m\to 0\) with \(|\sum_{R\in\mathcal{R}_m}c_R\lambda( R)-\int_I f\,d\lambda|\to 0\), and a dense \(\{x_k\}\) with integers \(N_m\) such that for \(k>N_m\) the point \(x_k\) lies in some \(\operatorname{int}R\), \(R\in\mathcal{R}_m\), with \(|f(x_k)-c_R|\le\eta_m\). Then \(\{x_k\}\) has the required property.
Proof. Given \(\varepsilon>0\), fix \(m\) with \(\eta_m<\varepsilon/4\) and \(|\sum_R c_R\lambda( R)-\int_I f\,d\lambda|<\varepsilon/4\), put \(N:=N_m\), \(F:=\{x_1,\dots,x_N\}\), and choose \(\delta\) with
\begin{equation*} 2^n N \delta^n\Bigl(\max_{x\in F}|f(x)| + C_m\Bigr) < \varepsilon/8, \qquad (2C_m+\eta_m)\lambda(\Gamma_\delta) < \varepsilon/8 . \end{equation*}
For \(\mathcal{P}\) admissible for \(\delta\) set \(\gamma(P):=\sum_{R}c_R\lambda(P\cap R)\), so \(|\gamma(P)|\le C_m\lambda(P)\) and \(\sum_P \gamma(P)=\sum_R c_R\lambda( R)\). Split \(\mathcal{P}\): (a) members tagged in \(F\); (b) the rest not inside a single \(R\); (c) all others. In (c), \(P\subset R\) uniquely, \(\gamma(P)=c_R\lambda(P)\), and the tag has index \(>N\), so it lies in \(\operatorname{int}R^{\prime}\) with \(|f(r(P))-c_{R^{\prime}}|\le\eta_m\); as \(\operatorname{int}R^{\prime}\) meets \(R=\overline{\operatorname{int}R}\), \(R^{\prime}=R\), giving total \(\le \eta_m<\varepsilon/4\). Group (a) has total volume \(\le 2^n N\delta^n\) by Lemma A, hence total \(<\varepsilon/8\); group (b) lies in \(\Gamma_\delta\) by Lemma B with \(|f(r(P))|\le C_m+\eta_m\), hence total \(<\varepsilon/8\). Adding and using the choice of \(m\) gives \(|\sum_P f(r(P))\lambda(P)-\int_I f\,d\lambda|<\varepsilon\).
Remark D. Lemma C’s hypothesis can be unsatisfiable. Let \(A_0\subset[0,1]\) be compact with empty interior and \(\lambda_1(A_0)\ge 1/2\) (Example 1.7.6), \(A:=A_0\times[0,1]^{n-1}\), \(f:=I_A\). If Lemma C data existed, fix \(m\) with \(\eta_m<1/4\) and \(|\sum_R c_R\lambda( R)-\lambda(A)|<1/8\). For each \(R\in\mathcal{R}_m\) the open set \(\operatorname{int}R\setminus A\) is nonempty (\(A\) nowhere dense) and so contains some \(x_k\) with \(k>N_m\); its cell is \(R\), whence \(|c_R|=|f(x_k)-c_R|\le\eta_m<1/4\). Then \(\sum_R c_R\lambda( R)<1/4\), contradicting \(>\lambda(A)-1/8\ge 3/8\). So any proof must use the priority ordering in \(r(P)\) (early terms placed inside \(A\) tag the heavy boxes, and later terms in the complement are never consulted there), which Lemma C discards.
What remains: for general integrable \(f\) one builds \(\{x_k\}\) together with a priority scheme, enumerating finite nets of compacta \(K_j\subset\{|f|\le j\}\) with \(f|_{K_j}\) continuous and \(j\,\lambda(I\setminus K_j)\to 0\) (Lusin plus \(j\,\lambda(|f|>j)\le\int_{\{|f|>j\}}|f|\,d\lambda\to 0\)) ahead of the complementary points, uniformly over admissible partitions of unbounded eccentricity; this is the content of Darji and Evans [203]. (A fully rigorous proof of the remaining step is beyond the scope of this page; see the reference given in the book.)
Show that there exists a Borel set in \([0,1]\) such that its indicator function cannot coincide a.e. with the limit of an increasing sequence of nonnegative step functions.
Take \(E := \bigcup_{m} K_m\), where \(\{J_m\}\) enumerates the intervals of \([0,1]\) with rational endpoints and positive length and \(K_m, K_m^{\prime}\) are pairwise disjoint nowhere dense compacta of positive measure with \(K_m\cup K_m^{\prime}\subset J_m\); this \(E\) is \(F_\sigma\), hence Borel, and
\begin{equation*} 0 < \lambda(J \cap E) < \lambda(J) \qquad \text{for every nondegenerate interval } J \subset [0,1]. \tag{3} \end{equation*}
The sets are chosen recursively: the union \(C\) of those already built is compact and nowhere dense, so \(J_m\setminus C\) contains an interval, which we split into two and fill with fat Cantor sets (Example 1.7.6). For (3), any such \(J\) contains some \(J_m\), so \(\lambda(J\cap E)\ge\lambda(K_m)>0\) while \(\lambda(J\setminus E)\ge\lambda(K_m^{\prime})>0\), since \(K_m^{\prime}\) misses every \(K_i\).
Suppose now \(f_n \le f_{n+1}\) are nonnegative step functions with \(f_n \to I_E\) a.e., let \(\Pi_n\) be a finite partition into intervals on which \(f_n\) is constant, and let \(D\) be the set of all endpoints of all \(\Pi_n\), a countable, hence null, set. Since \(\lambda(E)>0\) by (3), pick \(x_0 \in E\setminus D\) with \(f_n(x_0)\to I_E(x_0)=1\), and \(n_1\) with \(f_{n_1}(x_0)>1/2\). As \(x_0\notin D\) it lies in the interior \(J_0\) of an interval of \(\Pi_{n_1}\), where \(f_{n_1}\equiv f_{n_1}(x_0)>1/2\); by monotonicity \(\lim_n f_n \ge 1/2\) on \(J_0\), so \(I_E\ge 1/2\) a.e. on \(J_0\) and hence
\begin{equation*} \lambda(J_0 \setminus E) = 0 , \end{equation*}
contradicting (3). Finally, if \(g=\lim_n f_n\) a.e. and \(g=I_E\) a.e., then \(f_n\to I_E\) a.e., which has just been excluded.
(\(\circ\)) Let \(f\) be a measurable function on the real line vanishing outside some interval. Show that if \(\varepsilon_n \to 0\), then the functions \(x \mapsto f(x+\varepsilon_n)\) converge to \(f\) in measure.
Outside a set of measure \(<\eta\) the function \(f\) agrees with a uniformly continuous one, and translation moves that set without changing its measure. Let \(f\) vanish outside \([a,b]\), write \(f_n(x):=f(x+\varepsilon_n)\), and discard finitely many terms so that \(|\varepsilon_n|\le 1\); then all these functions vanish outside \(T:=[a-1,b+1]\). Fix \(c,\eta>0\).
Lusin’s theorem (Theorem 2.2.10, applied to \(f\) on the bounded interval \(T\)) gives a compact \(K\subset T\) with \(\lambda(T\setminus K)<\eta\) and a continuous \(g_0\) on \(T\) with \(g_0=f\) on \(K\). Put \(g:=\chi g_0\) on \(T\) and \(g:=0\) off \(T\), where \(\chi\) is continuous, \(0\le\chi\le1\), \(\chi=1\) on \([a,b]\), \(\chi=0\) off \((a-1/2,b+1/2)\). Then \(g\) is continuous on \(\mathbb{R}\), vanishes off \(T\), and \(g=f\) on \(K\) (on \(K\cap[a,b]\) since \(\chi=1\) there; on \(K\setminus[a,b]\) since both sides vanish). Hence
\begin{equation*} A := \{x : f(x) \ne g(x)\} \subset T \setminus K, \qquad \lambda(A) < \eta . \end{equation*}
Being continuous and compactly supported, \(g\) is uniformly continuous, so there is \(n_0\) with \(\sup_x |g(x+\varepsilon_n)-g(x)| \le c\) for \(n \ge n_0\). From
\begin{equation*} |f_n(x) - f(x)| \le |f(x+\varepsilon_n)-g(x+\varepsilon_n)| + |g(x+\varepsilon_n)-g(x)| + |g(x)-f(x)| , \end{equation*}
whose first term vanishes off \(A-\varepsilon_n\) and third off \(A\), we get for \(n \ge n_0\)
\begin{equation*} \{x : |f_n(x)-f(x)| > c\} \subset (A-\varepsilon_n) \cup A, \qquad \lambda\bigl(\{|f_n-f|>c\}\bigr) < 2\eta , \end{equation*}
using \(\lambda(A-\varepsilon_n)=\lambda(A)\). As \(\eta>0\) was arbitrary, \(f_n \to f\) in measure.
Exercises 2.12.67–2.12.73
Let \(f\) be a bounded measurable function on the real line.
(i) Is it true that \(f(x + n^{-1}) \to f(x)\) for a.e. \(x\)?
(ii) Show that there exists a subsequence \(n_k \to \infty\) such that \(f(x + n_k^{-1}) \to f(x)\) for a.e. \(x\).
(i) No: take \(f = \mathbf{1}_K\) with the compact \(K \subset [0,1]\), \(\lambda(K)>2/3\), constructed below, for which every \(x \in K\cap[0,1)\) admits arbitrarily large \(m\) with \(x+m^{-1}\notin K\); then \(f(x+m^{-1})=0\) infinitely often while \(f(x)=1\), so the convergence fails on a set of positive measure.
Lemma. For \(0<\varepsilon\le 1/2\) and \(d \in [\varepsilon^2,\varepsilon]\), the interval \((d-\varepsilon^2, d)\) contains some \(1/m\), \(m\in\mathbb{N}\). Let \(k\) be least with \(k^{-1}\le\varepsilon\); then \(k\ge 2\) and \((k-1)^{-1}>\varepsilon\), so \((k-1)\varepsilon<1\) gives \(\varepsilon-k^{-1}<\varepsilon/k\le\varepsilon^2\) (as \(k^{-1}\le\varepsilon\)), i.e. \(\varepsilon-\varepsilon^2<k^{-1}\). (a) If \(d>k^{-1}\), then \(d-\varepsilon^2\le\varepsilon-\varepsilon^2<k^{-1}<d\). (b) If \(d\le k^{-1}\), let \(m\) be least with \(m^{-1}<d\); then \(m>k\) and \((m-1)^{-1}\ge d\) by minimality, so
\begin{equation*} \frac1m = \frac{1}{m-1} - \frac{1}{m(m-1)} > d - \frac{1}{(m-1)^2} \ge d - \frac{1}{k^2} \ge d - \varepsilon^2 . \end{equation*}
The set: with \(\varepsilon_n := 2^{-2^n}\), split \([0,1]\) into the intervals \(I_{n,j}=[(j-1)\varepsilon_n, j\varepsilon_n)\), \(j\le 2^{2^n}\) (the last closed), delete from each the open \(U_{n,j}=(j\varepsilon_n-\varepsilon_n^2, j\varepsilon_n)\), and put \(K_n:=[0,1]\setminus\bigcup_j U_{n,j}\), \(K:=\bigcap_n K_n\), an intersection of closed sets, hence compact. Since \(\lambda([0,1]\setminus K_n)=2^{2^n}\varepsilon_n^2=\varepsilon_n\),
\begin{equation*} \lambda\bigl([0,1]\setminus K\bigr) \le \sum_{n\ge1} 2^{-2^{n}} \le \frac14+\frac1{16}+\sum_{i\ge 8}2^{-i} < \frac13 . \end{equation*}
For the key property fix \(x\in K\cap[0,1)\) and \(n\), let \(x\in I_{n,j}\) and \(d:=j\varepsilon_n-x\in(0,\varepsilon_n]\); as \(x\in K_n\) means \(x\notin U_{n,j}\), also \(d\ge\varepsilon_n^2\), so the Lemma (with \(\varepsilon=\varepsilon_n\le 1/4\)) gives \(m\) with \(m^{-1}\in(d-\varepsilon_n^2,d)\), i.e. \(x+m^{-1}\in U_{n,j}\) and hence \(x+m^{-1}\notin K\); moreover \(m^{-1}<d\le\varepsilon_n\) forces \(m>2^{2^n}\to\infty\).
(ii) Apply Exercise 2.12.66 to \(g_N := f\cdot\mathbf{1}_{[-N-1,N+1]}\) (measurable, vanishing outside an interval): \(g_N(\cdot+1/n)\to g_N\) in measure. For \(x\in[-N,N]\) and \(n\ge 1\) both \(x\) and \(x+1/n\) lie in \([-N-1,N+1]\), so \(g_N\) agrees with \(f\) at both, and therefore \(f(\cdot+1/n)\to f\) in measure on \([-N,N]\). Put \(S_0=\mathbb{N}\); given an infinite \(S_{N-1}\), the Riesz theorem (Theorem 2.2.5(i), \(\lambda\) being finite on \([-N,N]\)) yields an infinite \(S_N\subset S_{N-1}\) with
\begin{equation*} f(x+1/n) \to f(x) \quad \text{for a.e. } x\in[-N,N], \ n\to\infty, \ n\in S_N . \end{equation*}
Choose \(n_1<n_2<\cdots\) with \(n_k\in S_k\). For fixed \(N\) the tail \(\{n_k\}_{k\ge N}\) lies in \(S_N\), so \(f(x+n_k^{-1})\to f(x)\) off a null set \(Z_N\subset[-N,N]\); then \(Z=\bigcup_N Z_N\) is null and the convergence holds on \(\mathbb{R}\setminus Z\).
(\(\circ\)) Let \((X,\mathcal{A},\mu)\) be a space with a nonnegative measure and let a function \(f\colon X\times(a,b) \to \mathbb{R}^1\) be integrable in \(x\) for every \(t\) and differentiable in \(t\) at a fixed point \(t_0 \in (a,b)\) for every \(x\). Suppose that there exists a \(\mu\)-integrable function \(\Phi\) such that, for each \(t\), there exists a set \(Z_t\) such that \(\mu(Z_t) = 0\) and
\begin{equation*} |f(x,t) - f(x,t_0)| \le \Phi(x)\,|t - t_0| \qquad \text{if } x \notin Z_t . \end{equation*}
Show that the integral of \(f(x,t)\) with respect to the measure \(\mu\) is differentiable in \(t\) at the point \(t_0\) and
\begin{equation*} \frac{d}{dt}\int_X f(x,t)\,\mu(dx) = \int_X \frac{\partial f(x,t_0)}{\partial t}\,\mu(dx). \end{equation*}
Dominated convergence along sequences does it, the point being that only countably many of the sets \(Z_t\) are then involved. Write \(J(t):=\int_X f(x,t)\,\mu(dx)\), take an arbitrary sequence \(t_n \to t_0\) in \((a,b)\setminus\{t_0\}\), and put
\begin{equation*} g_n(x) := \frac{f(x,t_n) - f(x,t_0)}{t_n - t_0}, \qquad Z := \bigcup_{n=1}^{\infty} Z_{t_n} . \end{equation*}
Then \(\mu(Z)=0\), each \(g_n\) is integrable as a combination of the integrable \(f(\cdot,t_n), f(\cdot,t_0)\), and the hypothesis gives \(|g_n| \le \Phi\) off \(Z\). Since \(t \mapsto f(x,t)\) is differentiable at \(t_0\) for every \(x\) without exception,
\begin{equation*} g_n(x) \longrightarrow \frac{\partial f(x,t_0)}{\partial t} \qquad \text{for every } x \in X , \end{equation*}
so this limit is measurable (Theorem 2.1.5(v) on the completion of \(\mathcal{A}\)) and, being dominated by \(\Phi\) a.e., integrable. Dominated convergence, valid for measures with values in \([0,+\infty]\) by Corollary 2.8.6, together with linearity of the integral yields
\begin{equation*} \lim_{n\to\infty}\frac{J(t_n) - J(t_0)}{t_n - t_0} = \lim_{n\to\infty}\int_X g_n \,d\mu = \int_X \frac{\partial f(x,t_0)}{\partial t}\,\mu(dx) =: L . \end{equation*}
The value \(L\) is independent of the sequence \(\{t_n\}\), so the difference quotient of \(J\) tends to \(L\) as \(t \to t_0\), i.e. \(J^{\prime}(t_0) = L\), as required.
Prove that an arbitrary function \(f\colon [0,1] \to \mathbb{R}\) can be written in the form \(f(x) = \psi\bigl(\varphi(x)\bigr)\), where \(\varphi\colon [0,1] \to [0,1]\) is a Borel function and \(\psi\colon [0,1] \to \mathbb{R}\) is measurable with respect to Lebesgue measure.
Take the binary-to-ternary map
\begin{equation*} \varphi(x) := 2\sum_{n=1}^{\infty} x_n 3^{-n}, \qquad x_n := \lfloor 2^{n}x\rfloor - 2\lfloor 2^{n-1}x\rfloor \in \{0,1\}, \end{equation*}
for \(x \in [0,1)\), with \(x_n := 1\) for all \(n\) when \(x = 1\); and \(\psi := f \circ \varphi^{-1}\) on \(E := \varphi([0,1])\), \(\psi := 0\) on \([0,1]\setminus E\). The whole point is that \(E\) sits inside the Cantor set, hence is null, so completeness of Lebesgue measure makes such a \(\psi\) measurable for an arbitrary \(f\).
Telescoping gives \(\sum_{n\le N}x_n2^{-n} = 2^{-N}\lfloor 2^{N}x\rfloor \to x\), so the \(x_n\) are the binary digits of \(x\) and \(x=\sum_n x_n2^{-n}\) in all cases. Each \(x \mapsto x_n\) is a step function, hence Borel, the partial sums \(2\sum_{n\le N}x_n3^{-n}\) converge to \(\varphi\) uniformly (tail at most \(3^{-N}\)), and \(2\sum_{n\ge1}3^{-n}=1\); thus \(\varphi\colon[0,1]\to[0,1]\) is Borel.
Ternary expansions with digits in \(\{0,2\}\) are unique: two of them first differing at index \(m\), with \(c_m=2\), \(c_m^{\prime}=0\), satisfy
\begin{equation*} \sum_n c_n 3^{-n} - \sum_n c_n^{\prime}3^{-n} \ge 2\cdot 3^{-m} - \sum_{n>m}2\cdot 3^{-n} = 3^{-m} > 0 . \end{equation*}
So \(\varphi(x)=\varphi(y)\) forces \(x_n=y_n\) for all \(n\), hence \(x=y\); \(\varphi\) is injective, and \(E\) is contained in the Cantor set, so \(\lambda(E)=0\) and every subset of \(E\) is Lebesgue measurable.
Hence \(\psi\) is Lebesgue measurable: for Borel \(B\), the set \(\psi^{-1}(B)\cap E\) is a subset of a null set, while \(\psi^{-1}(B)\setminus E\) is \([0,1]\setminus E\) if \(0\in B\) and empty otherwise. Finally \(\psi(\varphi(x))=f(x)\) for every \(x\), by injectivity of \(\varphi\).
Show that almost everywhere convergence on the interval \(I = [0,1]\) with Lebesgue measure cannot be defined by a topology, i.e., there exists no topology on the set of all measurable functions on \(I\) (or on the set of all continuous functions on \(I\)) such that a sequence of functions is convergent in this topology precisely when it converges almost everywhere.
No such topology exists, because a.e. convergence violates the following purely topological property of sequences: if every subsequence of \(\{u_n\}\) in a topological space contains a further subsequence converging to \(u\), then \(u_n \to u\). (Else some open \(U \ni u\) is avoided by a whole subsequence, no subsequence of which can be eventually in \(U\).)
The counterexample is a continuous typewriter sequence: for \(n \ge 2\), \(1 \le k \le 2^{n}\) put \(I_{n,k} := [(k-1)2^{-n}, k2^{-n}]\),
\begin{equation*} g_{n,k}(x) := \max\bigl\{0,\ 1 - 2^{\,n+1}\operatorname{dist}(x, I_{n,k})\bigr\}, \end{equation*}
and arrange these in one sequence \(\{f_j\}\), all of level \(n\) before those of level \(n+1\). Each \(g_{n,k}\) is continuous with values in \([0,1]\), equals \(1\) on \(I_{n,k}\), and vanishes at distance \(\ge 2^{-n-1}\) from \(I_{n,k}\), so \(\lambda(g_{n,k}\ne 0) \le 2^{-n}+2\cdot 2^{-n-1} = 2^{-n+1}\). Hence:
(i) \(f_j \to 0\) in measure, since \(\{|g_{n,k}| \ge c\} \subset \{g_{n,k}\ne 0\}\) for \(c>0\) and the level of \(f_j\) tends to infinity with \(j\).
(ii) \(\{f_j(x)\}\) diverges at every \(x \in [0,1]\): at each level \(n\) some \(k\) has \(x \in I_{n,k}\), so \(g_{n,k}(x)=1\), while any of the \(2^{n}\ge 4\) indices \(j\) with \(|j-k|\ge 2\) has \(\operatorname{dist}(x,I_{n,j}) \ge (|j-k|-1)2^{-n} > 2^{-n-1}\), so \(g_{n,j}(x)=0\); thus infinitely many terms equal \(1\) and infinitely many equal \(0\).
Suppose \(\tau\) were a topology, on the measurable functions on \(I\) or on \(C[0,1]\) (all \(f_j\) and the limit \(0\) are continuous, so either setting applies), whose convergent sequences are exactly the a.e. convergent ones. Every subsequence of \(\{f_j\}\) converges to \(0\) in measure, hence by the Riesz theorem (Theorem 2.2.5(i), \(\lambda\) finite) contains a further subsequence converging to \(0\) a.e., i.e. in \(\tau\). By the property above \(f_j \to 0\) in \(\tau\), so \(\{f_j\}\) converges a.e., contradicting (ii).
(Marczewski) Let \(\mu\) be a probability measure such that convergence in measure for sequences of measurable functions is equivalent to convergence almost everywhere. Prove that the measure \(\mu\) is purely atomic.
Were \(\mu\) not purely atomic, a typewriter sequence on its atomless part would converge in measure but at no point of it. Fix a maximal family \(\{A_n\}\) of pairwise non-equivalent atoms (Definition 1.12.7; at most countably many, since non-equivalent atoms are disjoint up to null sets and \(\mu(X)=1\)) and put \(Y := X\setminus\bigcup_n A_n\), so that \(c:=\mu(Y)>0\) by assumption.
Then \(\mu|_Y\) is atomless: an atom \(B\subset Y\) is disjoint from every \(A_n\), so \(\mu(B\triangle A_n)=\mu(B)+\mu(A_n)>0\), i.e. \(B\) is equivalent to no \(A_n\), contradicting maximality. Consequently, for every \(n\ge 0\) there is a partition
\begin{equation*} Y = \bigsqcup_{k=1}^{2^{n}} A_{n,k},\qquad \mu(A_{n,k}) = 2^{-n}c , \end{equation*}
by induction on \(n\): the restriction of \(\mu\) to any \(A_{n,k}\) is again atomless, so Corollary 1.12.10 gives a measurable \(B\subset A_{n,k}\) with \(\mu(B)=2^{-n-1}c\), and then \(\mu(A_{n,k}\setminus B)=2^{-n-1}c\) too.
Let \(\{h_j\}\) list the indicators \(\mathbf{1}_{A_{n,k}}\), \(n\ge 1\), \(k \le 2^{n}\), all of level \(n\) before those of level \(n+1\). For \(\varepsilon\in(0,1)\) we have \(\mu(|h_j|\ge\varepsilon)=2^{-n}c\) at level \(n\), and the level tends to infinity with \(j\), so \(h_j\to 0\) in measure. But at each \(x\in Y\) and each level exactly one \(k\) has \(x \in A_{n,k}\), giving \(h_j(x)=1\), while some \(k^{\prime}\) (there are \(2^{n}\ge 2\)) gives \(0\); so \(\{h_j(x)\}\) has infinitely many \(1\)’s and \(0\)’s and diverges throughout \(Y\). By hypothesis \(h_j\to 0\) a.e., contradicting \(\mu(Y)=c>0\). Hence \(\mu\) is purely atomic.
(\(\circ\)) Prove that a function \(f\) on an interval \([a,b]\) is continuous at a point \(x\) precisely when its oscillation at \(x\) is zero, where the oscillation at \(x\) is defined by the formula
\begin{equation*} \omega_f(x) := \lim_{\varepsilon\to 0}\ \sup\bigl\{|f(z) - f(y)|\ :\ |z - x| < \varepsilon,\ |y-x| < \varepsilon\bigr\}. \end{equation*}
Both directions are the triangle inequality applied to the quantity
\begin{equation*} S_x(\varepsilon) := \sup\bigl\{|f(z)-f(y)| :\ y,z\in[a,b],\ |z-x|<\varepsilon,\ |y-x|<\varepsilon\bigr\}, \end{equation*}
which is nondecreasing in \(\varepsilon\), so that the limit defining \(\omega_f(x)\) exists in \([0,+\infty]\) and \(\omega_f(x)=\inf_{\varepsilon>0}S_x(\varepsilon)\).
(i) If \(f\) is continuous at \(x\), fix \(\delta>0\) and \(\varepsilon>0\) with \(|f(y)-f(x)|<\delta/2\) for \(y\in[a,b]\), \(|y-x|<\varepsilon\); then
\begin{equation*} |f(z)-f(y)| \le |f(z)-f(x)| + |f(x)-f(y)| < \delta \end{equation*}
for all admissible \(y,z\), so \(\omega_f(x)\le S_x(\varepsilon)\le\delta\) for every \(\delta>0\), i.e. \(\omega_f(x)=0\).
(ii) If \(\omega_f(x)=0\), fix \(\delta>0\) and \(\varepsilon>0\) with \(S_x(\varepsilon)<\delta\); as \(z=x\) is admissible, \(|f(y)-f(x)|\le S_x(\varepsilon)<\delta\) whenever \(y\in[a,b]\), \(|y-x|<\varepsilon\), which is continuity at \(x\).
(Baire’s theorem) Let \(f_n\) be continuous functions on \([a,b]\) such that for every \(x \in [a,b]\) there exists a finite limit \(f(x) = \lim_{n\to\infty} f_n(x)\). Prove that the set of points of continuity of \(f\) is everywhere dense in \([a,b]\).
The set of discontinuity points is a countable union of closed nowhere dense sets, so its complement is dense by the Baire category theorem. With \(\omega_f\) and \(S_x\) as in Exercise 2.12.72, that exercise gives
\begin{equation*} D := \{x : \omega_f(x) > 0\} = \bigcup_{j=1}^{\infty} F_j , \qquad F_j := \bigl\{x \in [a,b] : \omega_f(x) \ge j^{-1}\bigr\} \end{equation*}
for the set \(D\) of discontinuity points, and each \(F_j\) is closed: if \(\omega_f(x)<c\), pick \(\varepsilon\) with \(S_x(\varepsilon)<c\), and then \(S_{x^{\prime}}(\varepsilon/2)\le S_x(\varepsilon)<c\) for \(|x^{\prime}-x|<\varepsilon/2\), so \(\{\omega_f<c\}\) is open.
Each \(F_j\) has empty interior. Suppose \((c,d)\subset F_j\), put \(\eta := 1/(5j)\), and set
\begin{equation*} E_m := \bigcap_{p,q\ge m}\bigl\{x\in[c,d] : |f_p(x)-f_q(x)| \le \eta\bigr\}, \end{equation*}
closed by continuity of the \(f_p-f_q\), with \(\bigcup_m E_m = [c,d]\) since every numerical sequence \(\{f_n(x)\}\) converges and so is fundamental. As \([c,d]\) is complete, the Baire category theorem makes some \(E_m\) not nowhere dense, hence, being closed, it contains an interval \((u,v)\), which we may shrink to lie inside \((c,d)\). Letting \(q\to\infty\) in \(|f_m(x)-f_q(x)|\le\eta\) gives \(|f_m(x)-f(x)|\le\eta\) on \((u,v)\); fix \(x_0\in(u,v)\) and, by continuity of \(f_m\), a \(\rho>0\) with \((x_0-\rho,x_0+\rho)\subset(u,v)\) and \(|f_m(y)-f_m(x_0)|<\eta\) for \(|y-x_0|<\rho\). Then for \(|y-x_0|<\rho\) and \(|z-x_0|<\rho\),
\begin{equation*} \begin{aligned} |f(y) - f(z)| &\le |f(y)-f_m(y)| + |f_m(y)-f_m(x_0)| \\ &\quad + |f_m(x_0)-f_m(z)| + |f_m(z)-f(z)| \le 4\eta = \frac{4}{5j}, \end{aligned} \end{equation*}
so \(\omega_f(x_0) \le S_{x_0}(\rho) \le 4/(5j) < 1/j\), contradicting \(x_0 \in (c,d) \subset F_j\).
Hence the sets \(G_j := [a,b]\setminus F_j\) are open and dense, and the Baire category theorem in the complete space \([a,b]\) makes \(\bigcap_j G_j = [a,b]\setminus D\) dense; by Exercise 2.12.72 this is exactly the set of continuity points of \(f\).
Exercises 2.12.74–2.12.80
(i) Construct an example of a sequence of continuous functions \(f_n\) on \([0,1]\) such that, for every \(x \in [0,1]\), there exists a finite limit \(f(x) = \lim_{n\to\infty} f_n(x)\), but the set of points of discontinuity of \(f\) is everywhere dense in \([0,1]\).
(ii) Construct an example showing that the function \(f\) in (i) may be discontinuous almost everywhere.
(i) Take, with \(\{r_k\}\) an enumeration of \(\mathbb{Q}\cap[0,1)\),
\begin{equation*} f_n := \sum_{k=1}^{n} 2^{-k} h_{k,n}, \qquad h_{k,n}(x) := \min\bigl\{1, \max\{0,\, n(x-r_k)\}\bigr\} . \end{equation*}
Each \(h_{k,n}\) is continuous with values in \([0,1]\), vanishes for \(x\le r_k\) and equals \(1\) for \(x \ge r_k+1/n\), so \(h_{k,n} \to \chi_{(r_k,1]}\) pointwise as \(n\to\infty\); each \(f_n\) is therefore continuous, and splitting at an index \(K\) with \(2^{1-K}<\varepsilon/2\) (the terms with \(k>K\) contribute at most \(2\sum_{k>K}2^{-k}=2^{1-K}\) from \(f_n\) and \(f\) together, the \(K\) remaining ones tend to \(0\)) gives
\begin{equation*} f_n(x) \longrightarrow f(x) := \sum_{k=1}^{\infty} 2^{-k}\chi_{(r_k,1]}(x) \qquad \text{for every } x \in [0,1]. \end{equation*}
All summands are nondecreasing, so for \(y>r_k\) the \(k\)th term alone contributes a full jump and \(f(y)-f(r_k)\ge 2^{-k}\); thus \(f\) is discontinuous at every point of the dense set \(\mathbb{Q}\cap[0,1)\).
(ii) Let \(\{r_k\}\) now enumerate \(\mathbb{Q}\cap[0,1]\) and put
\begin{equation*} G_n := [0,1] \cap \bigcup_{k=1}^{\infty}\bigl(r_k - 2^{-k-n-1},\, r_k+2^{-k-n-1}\bigr), \qquad F_n := [0,1]\setminus G_n , \end{equation*}
so that each \(F_n\) is closed and nowhere dense (\(G_n\) is open and contains all rationals), \(\lambda(G_n)\le 2^{-n}\), and \(F_n\subset F_{n+1}\) because \(G_{n+1}\subset G_n\); hence \(D:=\bigcap_n G_n\) is dense and null. Take
\begin{equation*} f := \sum_{n=1}^{\infty} 2^{-n}\chi_{F_n} = 2^{1-N(x)} \ \ (x\notin D), \qquad f(x) = 0 \ \ (x \in D), \end{equation*}
where \(N(x):=\min\{n : x\in F_n\}\), the \(F_n\) being increasing. This \(f\) is the pointwise limit of the continuous \(f_j := \sum_{n\le j}2^{-n}g_{n,j}\) with \(g_{n,j}(x) := \max\{0,\,1-j\operatorname{dist}(x,F_n)\}\), by the two-block estimate of (i) together with \(g_{n,j}\to\chi_{F_n}\) pointwise (\(F_n\) is closed, and \(\operatorname{dist}(\cdot,F_n)\) is \(1\)-Lipschitz).
Its continuity points are exactly \(D\). (a) For \(x\notin D\) put \(N:=N(x)<\infty\), so \(f(x)=2^{1-N}\); every neighbourhood of \(x\) meets \([0,1]\setminus F_N\), which is nonempty in it since \(F_N\) is nowhere dense, and there \(N(\cdot)>N\), so \(f\le 2^{-N}=f(x)/2\) and \(f\) is discontinuous at \(x\). (b) For \(x\in D\) and \(\varepsilon>0\), choose \(N\) with \(2^{1-N}<\varepsilon\); as \(x\in G_N\) with \(G_N\) open, some neighbourhood of \(x\) lies in \(G_N\), where \(N(\cdot)>N\) and hence \(f\le 2^{-N}<\varepsilon = f(x)+\varepsilon\). So \(f\) is continuous precisely on the null set \(D\), i.e. discontinuous a.e.
Prove that the uniform limit of a sequence of functions of Baire class \(\alpha\) or less is also of Baire class \(\alpha\) or less.
Telescope the uniform limit into a series of uniformly small increments and truncate the approximants. Write \(\mathcal{B}^{\alpha}:=\bigcup_{\beta\le\alpha}B_\beta\) for the functions of Baire class \(\alpha\) or less on a fixed metric space \(X\), the classes \(B_\beta\) of 2.12(iv) being pairwise disjoint.
Lemma 1. For \(\alpha\ge1\), a function \(f\) lies in \(\mathcal{B}^{\alpha}\) if and only if \(f_j \to f\) pointwise for some \(f_j\in\mathcal{B}^{\beta_j}\) with \(\beta_j<\alpha\). Indeed, given such \(f_j\), either \(f\in B_\beta\) for some \(\beta<\alpha\), or \(f\) lies in no such class and then, each \(f_j\) belonging to some \(B_{\gamma_j}\) with \(\gamma_j\le\beta_j<\alpha\), the definition of \(B_\alpha\) puts \(f\in B_\alpha\). Conversely, for \(f\in B_\gamma\) with \(\gamma\le\alpha\) take \(f_j\equiv f\) if \(\gamma=0\) and otherwise the sequence supplied by the definition of \(B_\gamma\).
Lemma 2. Each \(\mathcal{B}^{\alpha}\) is a linear space. Transfinite induction on \(\alpha\): for \(\alpha=0\) this is linearity of \(C(X)\); for \(\alpha\ge1\) take \(u_j\to u\), \(v_j\to v\) as in Lemma 1 and put \(\delta_j:=\max(\beta_j,\gamma_j)<\alpha\), so that \(u_j+cv_j\in\mathcal{B}^{\delta_j}\) by the inductive hypothesis and \(u+cv\in\mathcal{B}^{\alpha}\) by Lemma 1.
Now let \(f_n\in\mathcal{B}^{\alpha}\) with \(f_n\to f\) uniformly on \(X\); for \(\alpha=0\) this is the classical theorem, so let \(\alpha\ge1\). Pick \(n_1<n_2<\cdots\) with \(\sup_X|f_{n_k}-f|\le 2^{-k-1}\) and put \(h_1:=f_{n_1}\), \(h_k:=f_{n_k}-f_{n_{k-1}}\), which lie in \(\mathcal{B}^{\alpha}\) by Lemma 2 and satisfy
\begin{equation*} \sup_X |h_k| \le 2^{-k-1}+2^{-k} \le c_k := 2^{1-k} \quad (k \ge 2), \qquad f = \sum_{k=1}^{\infty} h_k \ \text{ pointwise} \end{equation*}
(the partial sums being \(f_{n_m}\)). By Lemma 1 choose \(h_{k,j}\in\mathcal{B}^{\beta_{k,j}}\), \(\beta_{k,j}<\alpha\), with \(h_{k,j}\to h_k\) pointwise, and truncate at height \(c_k\): with the continuous \(T_k(t):=\max\{-c_k,\min\{c_k,t\}\}\) put \(\widetilde h_{k,j}:=T_k\circ h_{k,j}\) for \(k\ge2\) and \(\widetilde h_{1,j}:=h_{1,j}\). Then \(\widetilde h_{k,j}\in\mathcal{B}^{\beta_{k,j}}\) by Exercise 2.12.76, \(|\widetilde h_{k,j}|\le c_k\), and \(\widetilde h_{k,j}\to T_k(h_k)=h_k\) pointwise, since \(|h_k|\le c_k\) and \(T_k\) is continuous. Put \(u_j:=\sum_{k\le j}\widetilde h_{k,j}\); the finitely many \(\beta_{k,j}\) involved have a maximum \(\gamma_j<\alpha\), so \(u_j\in\mathcal{B}^{\gamma_j}\) by Lemma 2. Finally, for \(K\ge2\) and \(j>K\),
\begin{equation*} |u_j(x)-f(x)| \le \sum_{k=1}^{K}\bigl|\widetilde h_{k,j}(x)-h_k(x)\bigr| + \sum_{k>K} 2c_k , \qquad \sum_{k>K}2c_k = 2^{2-K}, \end{equation*}
the tail bound using \(|\widetilde h_{k,j}|\le c_k\) and \(|h_k|\le c_k\) for \(k\ge2\); the first sum has \(K\) terms each tending to \(0\) as \(j\to\infty\). Hence \(u_j\to f\) pointwise, and Lemma 1 gives \(f\in\mathcal{B}^{\alpha}\).
Prove that if a function \(\varphi\) is continuous on the real line and a function \(f\) is of Baire class \(\alpha\) or less, then so is the function \(\varphi \circ f\).
Transfinite induction on \(\alpha\), with \(\mathcal{B}^{\alpha}=\bigcup_{\beta\le\alpha}B_\beta\) and Lemma 1 of Exercise 2.12.75: for \(\alpha\ge1\), a function lies in \(\mathcal{B}^{\alpha}\) exactly when it is the pointwise limit of functions in classes \(\mathcal{B}^{\beta_j}\) with \(\beta_j<\alpha\).
For \(\alpha=0\) a composition of continuous maps is continuous. For \(\alpha\ge1\), assume the claim for all \(\beta<\alpha\), and let \(f\in\mathcal{B}^{\alpha}\) with \(f_j\to f\) pointwise, \(f_j\in\mathcal{B}^{\beta_j}\), \(\beta_j<\alpha\). Then \(\varphi\circ f_j\in\mathcal{B}^{\beta_j}\) by the inductive hypothesis, and continuity of \(\varphi\) gives
\begin{equation*} \varphi\bigl(f_j(x)\bigr) \longrightarrow \varphi\bigl(f(x)\bigr) \qquad \text{for every } x , \end{equation*}
so \(\varphi\circ f\in\mathcal{B}^{\alpha}\) by Lemma 1.
Prove that if a function \(f\) is of Baire class \(\alpha\) or less on the plane, then the function \(\varphi(x) = f(x,x)\) is of Baire class \(\alpha\) or less on the real line.
Substitution along a continuous map preserves Baire class: writing \(\mathcal{B}^{\alpha}(X)=\bigcup_{\beta\le\alpha}B_\beta(X)\) as in the two previous exercises, transfinite induction on \(\alpha\) gives, for metric spaces \(X,Y\),
\begin{equation*} f \in \mathcal{B}^{\alpha}(X), \ \ T\colon Y\to X \text{ continuous} \ \Longrightarrow \ f\circ T \in \mathcal{B}^{\alpha}(Y) , \end{equation*}
and \(\varphi = f\circ\delta\) with the continuous diagonal \(\delta(x)=(x,x)\).
For \(\alpha=0\) a composition of continuous maps is continuous. For \(\alpha\ge1\), Lemma 1 of Exercise 2.12.75 supplies \(f_j\in\mathcal{B}^{\beta_j}(X)\), \(\beta_j<\alpha\), with \(f_j\to f\) pointwise; the inductive hypothesis gives \(f_j\circ T\in\mathcal{B}^{\beta_j}(Y)\), and applying the convergence at \(z=T(y)\) gives \((f_j\circ T)(y)\to(f\circ T)(y)\) for every \(y\in Y\), so \(f\circ T\in\mathcal{B}^{\alpha}(Y)\) by Lemma 1.
Prove that the Dirichlet function (the indicator of the set of rational numbers) belongs to the second Baire class, but not to the first one.
The Dirichlet function \(d=\chi_{\mathbb{Q}}\) lies in \(\mathcal{B}^{2}=B_0\cup B_1\cup B_2\) but not in \(\mathcal{B}^{1}=B_0\cup B_1\), hence, the classes being disjoint, in \(B_2\). We work on \([0,1]\), the real line being identical.
For the upper bound let \(\{r_k\}\) enumerate \(\mathbb{Q}\cap[0,1]\) and put
\begin{equation*} \psi_m := \chi_{\{r_1,\ldots,r_m\}}, \qquad \psi_{m,j}(x) := \max_{1\le k\le m}\max\bigl\{0,\ 1-j|x-r_k|\bigr\} . \end{equation*}
Each \(\psi_{m,j}\) is continuous, and for fixed \(m\), \(\psi_{m,j}(x)=1\) for all \(j\) on \(\{r_1,\dots,r_m\}\) while \(\psi_{m,j}(x)=0\) once \(j>(\min_{k\le m}|x-r_k|)^{-1}\) off it; so \(\psi_{m,j}\to\psi_m\) pointwise and \(\psi_m\in\mathcal{B}^{1}\). Also \(\psi_m\to d\) pointwise (\(\psi_m(x)=0=d(x)\) at irrational \(x\), \(\psi_m(r_k)=1\) for \(m\ge k\)), so \(d\in\mathcal{B}^{2}\) by Lemma 1 of Exercise 2.12.75.
For the lower bound, \(d\) is continuous at no point, every interval containing rational and irrational points; but a member of \(\mathcal{B}^{1}\) is continuous everywhere or else a pointwise limit of continuous functions, whose continuity points are dense by Exercise 2.12.73. Hence \(d\notin\mathcal{B}^{1}\).
Construct a measurable function on \([0,1]\) that cannot be redefined on a set of measure zero to obtain a function from the first Baire class.
Take \(f := \chi_E\) with \(E := \bigcup_n K_n\), where \(\{I_n\}\) enumerates the intervals of \([0,1]\) with rational endpoints and nonempty interior and \(K_n, L_n \subset I_n\) are pairwise disjoint nowhere dense compacta of positive measure; every modification of \(f\) on a null set is then everywhere discontinuous, while a function of \(\mathcal{B}^{1}\) is either continuous or a pointwise limit of continuous functions, whose continuity points are dense by Exercise 2.12.73. (A single compact \(K\) with empty interior would not do: \(\chi_K\) is already of the first Baire class, being the pointwise limit of the continuous \(\max\{0,1-j\operatorname{dist}(\cdot,K)\}\).)
The sets are chosen recursively: \(S := \bigcup_{m<n}(K_m\cup L_m)\) is compact and nowhere dense, so \(\operatorname{int}I_n\setminus S\) contains a nondegenerate interval, which we split into two and fill with fat Cantor sets (Example 1.7.6). Hence \(E\) is \(F_\sigma\), so \(f\) is Borel measurable, and for every nondegenerate interval \(I\subset[0,1]\), choosing \(I_n\subset I\),
\begin{equation*} \lambda(I\cap E) \ge \lambda(K_n) > 0 , \qquad \lambda(I\setminus E) \ge \lambda(L_n) > 0 , \end{equation*}
the second because the family \(\{K_m\}\cup\{L_m\}\) is disjoint, so \(L_n\cap E=\emptyset\).
Consequently, if \(g=f\) off a null set \(N\), then any nondegenerate \(I\) contains points \(y\in(I\cap E)\setminus N\) and \(z\in(I\setminus E)\setminus N\), since neither \(I\cap E\) nor \(I\setminus E\) is null and hence neither lies in \(N\); thus \(g(y)=1\), \(g(z)=0\), so \(g\) has oscillation \(1\) at every point of \([0,1]\) and is nowhere continuous. By the citation above, no such \(g\) lies in \(\mathcal{B}^{1}\).
Let a function \(f\) on the plane be continuous in every variable separately. Show that at some point \(f\) is continuous as a function on the plane.
The set of such points is even a dense \(G_\delta\), because \(f\) is a pointwise limit of continuous functions on the plane and Baire’s theorem (Exercise 2.12.73) then applies.
For the approximation interpolate linearly in the first variable along the grid \(\{k/n\}\): for \(k/n \le x \le (k+1)/n\),
\begin{equation*} f_n(x,y) := (k+1-nx)\, f\bigl(k/n,\,y\bigr) + (nx-k)\, f\bigl((k+1)/n,\,y\bigr), \end{equation*}
consistently at the grid points, both formulas giving \(f(k/n,y)\) there. On the closed strip \(S_k := \{k/n \le x \le (k+1)/n\}\) this is \(\alpha(x)u(y)+\beta(x)v(y)\) with \(\alpha,\beta\) affine and \(u:=f(k/n,\cdot)\), \(v:=f((k+1)/n,\cdot)\) continuous by separate continuity in the second variable, hence continuous there; the \(S_k\) form a locally finite closed cover whose restrictions agree on the overlaps \(\{x=k/n\}\), so \(f_n\in C(\mathbb{R}^2)\). Fixing \((x,y)\) and \(k_n\) with \(k_n/n \le x \le (k_n+1)/n\), the value \(f_n(x,y)\) is a convex combination of \(f(k_n/n,y)\) and \(f((k_n+1)/n,y)\), both tending to \(f(x,y)\) by separate continuity in the first variable, so
\begin{equation*} f_n(x,y) \longrightarrow f(x,y) \qquad \text{for every } (x,y) \in \mathbb{R}^2 . \end{equation*}
The proof of Exercise 2.12.73 used only completeness of the domain and its closed balls, so it applies verbatim on \(\mathbb{R}^2\) with the oscillation \(\omega_f(z):=\inf\{\operatorname{diam}f(U) : U \ni z \text{ open}\}\): each \(G_\varepsilon := \{\omega_f<\varepsilon\}\) is open (if \(\operatorname{diam}f(U)<\varepsilon\) with \(U\ni z\) open, then \(\omega_f<\varepsilon\) throughout \(U\)) and dense, so
\begin{equation*} C = \{z : \omega_f(z) = 0\} = \bigcap_{j=1}^{\infty} G_{1/j} \end{equation*}
is a dense \(G_\delta\) by the Baire category theorem; and \(f\) is continuous at \(z\) precisely when \(\omega_f(z)=0\).
Exercises 2.12.81–2.12.87
Let \(f\) be a measurable real function on a measure space \((X,\mathcal{A},\mu)\) with a positive measure \(\mu\). Prove that there exists a number \(y\) such that
\begin{equation*} \int_X \frac{1}{|f(x)-y|}\,\mu(dx) = +\infty . \end{equation*}
Take for \(y\) the point produced by the bisection below; the integrand \(|f-y|^{-1}\) is read as \(+\infty\) where \(f=y\), and \(\int=+\infty\) means non-integrability.
Reduce first. As \(f\) is real-valued, \(X=\bigcup_k\{|f|\le k\}\), so subadditivity gives some \(k\ge1\) with \(\mu(X_0)>0\) for \(X_0:=\{|f|\le k\}\). If \(\mu(X_0)=+\infty\), take \(y:=2k\): then \(|f-y|\le 3k\) on \(X_0\), so \(|f-y|^{-1}\ge (3k)^{-1}I_{X_0}\), which is not integrable. So assume \(0<\mu(X_0)<\infty\), pass to the probability measure \(\nu:=\mu|_{X_0}/\mu(X_0)\) and to \(g:=(f|_{X_0}+k)/(2k) \in [0,1]\); for \(y^{\prime}\in[0,1]\) and \(y:=2ky^{\prime}-k\) we have \(|f-y|=2k|g-y^{\prime}|\) on \(X_0\), so
\begin{equation*} \int_X\frac{\mu(dx)}{|f(x)-y|} \ \ge\ \frac{\mu(X_0)}{2k}\int_{X_0}\frac{\nu(dx)}{|g(x)-y^{\prime}|} , \end{equation*}
and it suffices to make the right-hand integral infinite.
Bisect: put \(I_0:=[0,1]\), \(E_n:=g^{-1}(I_n)\), and, given a closed \(I_n\) of length \(2^{-n}\) with \(\nu(E_n)\ge 2^{-n}\), let \(I_{n+1}\) be a closed half of \(I_n\) with \(\nu(g^{-1}(I_{n+1})) \ge \frac12\nu(E_n) \ge 2^{-n-1}\), which exists since the two halves cover \(I_n\). The \(I_n\) decrease with \(|I_n|\to0\), so \(\bigcap_n I_n\) is a single point \(y^{\prime}\in[0,1]\); the \(E_n\) decrease, and on \(E_n\) both \(g\) and \(y^{\prime}\) lie in \(I_n\), so \(h:=|g-y^{\prime}|^{-1}\ge 2^{n}\) there. Hence \(0 \le h_N := \sum_{n\le N}2^{n-1}I_{E_n} \le h\) for every \(N\): off \(E_1\) we have \(h_N=0\), and at a point whose largest index \(n \le N\) with \(x\in E_n\) is \(n^{*}\),
\begin{equation*} h_N(x) = \sum_{n=1}^{n^{*}}2^{n-1} = 2^{n^{*}}-1 < 2^{n^{*}} \le h(x) . \end{equation*}
Therefore \(\int_{X_0}h_N\,d\nu = \sum_{n\le N}2^{n-1}\nu(E_n) \ge N/2\) for every \(N\), so by monotonicity \(h\) is not \(\nu\)-integrable; were \(|f-y|^{-1}\) integrable on \(X\), it would be \(\mu\)-integrable on \(X_0\), hence \(\nu\)-integrable there, making \(h\) integrable.
(\(\circ\)) Let \((X,\mathcal{A},\mu)\) be a measurable space with a finite positive measure \(\mu\) and let \(f\) be a \(\mu\)-measurable function with values in \(\mathbb{R}\) or in \(\mathbb{C}\). A point \(y\) is called an essential value of \(f\) if \(\mu\bigl(x:\ |f(x)-y|<\varepsilon\bigr)>0\) for each \(\varepsilon>0\).
(i) Show that a function \(f\) need not assume every essential value and that not every actual value of \(f\) is essential.
(ii) Show that the set of all essential values of \(f\) has a nonempty intersection with \(f(X)\).
(iii) Show that the set of all essential values is closed and coincides with the intersection of the closures of the sets \(\tilde f(X)\) over all functions \(\tilde f\) a.e. equal to \(f\).
Write \(E\) for the set of essential values of \(f\), \(Y\in\{\mathbb{R},\mathbb{C}\}\) for the range space and \(B(y,\varepsilon)\) for its balls, so that \(\{|f-y|<\varepsilon\}=f^{-1}(B(y,\varepsilon))\). Throughout, \(E\) depends only on the a.e.-class of \(f\): if \(\tilde f=f\) a.e., the two preimages differ inside a null set and so have equal measure. Also \(\mu(X)>0\), which is what positivity of \(\mu\) means, and \(Y\) is separable, hence has a countable base \(\mathcal{B}\) of balls.
(i) On \(X=[0,1]\) with Lebesgue measure take \(f(x):=x\) for \(x\ne 1/2\) and \(f(1/2):=2\). Since \(f\) is the identity a.e.,
\begin{equation*} \lambda\bigl(x:\ |f(x)-y|<\varepsilon\bigr) = \lambda\bigl((y-\varepsilon,y+\varepsilon)\cap[0,1]\bigr), \end{equation*}
positive for all \(\varepsilon>0\) exactly when \(y\in[0,1]\); so \(E=[0,1]\), while \(f(X)=([0,1]\setminus\{1/2\})\cup\{2\}\). Thus \(1/2\) is essential but not attained, and \(2\) is attained but not essential.
(ii) In fact \(f(x)\in E\) for a.e. \(x\). Put \(\mathcal{B}_0:=\{B\in\mathcal{B}:\ \mu(f^{-1}(B))=0\}\), a countable family. Any \(y\notin E\) has \(\varepsilon>0\) with \(\mu(f^{-1}(B(y,\varepsilon)))=0\) and, \(\mathcal{B}\) being a base, some \(B\in\mathcal{B}\) with \(y\in B\subset B(y,\varepsilon)\), so \(B\in\mathcal{B}_0\); hence \(Y\setminus E\subset\bigcup_{B\in\mathcal{B}_0}B\) and countable subadditivity gives
\begin{equation*} \mu\bigl(f^{-1}(Y\setminus E)\bigr) \le \sum_{B\in\mathcal{B}_0}\mu\bigl(f^{-1}(B)\bigr) = 0 , \end{equation*}
the left side being measurable since \(Y\setminus E\) is open by (iii). So \(\mu(f^{-1}(E))=\mu(X)>0\), and any \(x\) there gives \(f(x)\in E\cap f(X)\).
(iii) \(E\) is closed: if \(\mu(f^{-1}(B(y,\varepsilon)))=0\) and \(|z-y|<\varepsilon/2\), then \(B(z,\varepsilon/2)\subset B(y,\varepsilon)\), so \(z\notin E\); i.e. \(B(y,\varepsilon/2)\) misses \(E\). With \(D:=\bigcap\overline{\tilde f(X)}\) over all \(\tilde f=f\) a.e. (all of them \(\mu\)-measurable, \(\mu\)-measurability being measurability for the complete Lebesgue completion):
(a) \(E\subset D\), since for \(y\in E\) and any such \(\tilde f\) the set \(\tilde f^{-1}(B(y,\varepsilon))\) has the same positive measure as \(f^{-1}(B(y,\varepsilon))\), hence is nonempty, for every \(\varepsilon>0\).
(b) \(D\subset E\), since for \(y\notin E\) with \(N:=f^{-1}(B(y,\varepsilon))\) null the modification \(\tilde f:=f\) off \(N\), \(\tilde f:=y+\varepsilon\) on \(N\), has \(|\tilde f-y|\ge\varepsilon\) everywhere, so the closed set \(\{|z-y|\ge\varepsilon\}\) contains \(\overline{\tilde f(X)}\) and omits \(y\).
Let \(\mu\) be a nonnegative measure and let \(f\) be a \(\mu\)-measurable function that has a bounded modification. Such functions are called essentially bounded. The essential supremum \(\operatorname{esssup} f\) and essential infimum \(\operatorname{essinf} f\) of the function \(f\) are defined as follows:
\begin{equation*} \operatorname{esssup} f:=\inf\bigl\{M:\ f(x)\le M\ \mu\text{-a.e.}\bigr\},\qquad \operatorname{essinf} f:=\sup\bigl\{m:\ f(x)\ge m\ \mu\text{-a.e.}\bigr\}. \end{equation*}
A bounded measurable function \(f\) on \([a,b]\) is called reduced if, for every interval \((\alpha,\beta)\subset[a,b]\), one has
\begin{equation*} \inf_{(\alpha,\beta)}f=\operatorname{essinf}_{[\alpha,\beta]}f,\qquad \sup_{(\alpha,\beta)}f=\operatorname{esssup}_{[\alpha,\beta]}f . \end{equation*}
Prove that each bounded measurable function \(f\) on \([a,b]\) with Lebesgue measure has a reduced modification.
Take \(\tilde f := m\) on \((a,b)\cap N\) and \(\tilde f := f\) elsewhere, where \(N\) is the null set below and, the sup and inf running over open intervals \(J\) with \(x \in J \subset [a,b]\),
\begin{equation*} m(x) := \sup_{J \ni x} \operatorname{essinf}_{J} f , \qquad M(x) := \inf_{J \ni x} \operatorname{esssup}_{J} f \end{equation*}
(both finite, since \(f\) is bounded). Reducedness of a modification \(\tilde f\) amounts to
\begin{equation*} \operatorname{essinf}_{J} f \ \le\ \tilde f(x)\ \le\ \operatorname{esssup}_{J} f \qquad \text{for every open } J \subset [a,b] \text{ and } x \in J, \tag{R} \end{equation*}
because \(\operatorname{esssup}_J\tilde f=\operatorname{esssup}_Jf\le\sup_J\tilde f\) always (the equality as \(\tilde f=f\) a.e., the inequality as \(\tilde f\le\sup_J\tilde f\) on \(J\)) and symmetrically for the infima, while \(\{\alpha,\beta\}\) is null, so the open and closed versions of the definition agree.
Two standard facts are used: \(f\le\operatorname{esssup}_Sf\) a.e. on any \(S\) with \(\lambda(S)>0\) (take admissible \(M_j\downarrow\operatorname{esssup}_Sf\) and discard countably many null sets), symmetrically for the infimum; and both essential bounds are monotone in \(S\), an inequality holding a.e. on \(T\) holding a.e. on \(S\subset T\). Hence \(m\le M\) on \((a,b)\): for \(J_1,J_2\ni x\) the intersection has positive length, so
\begin{equation*} \operatorname{essinf}_{J_1}f \le \operatorname{essinf}_{J_1\cap J_2}f \le \operatorname{esssup}_{J_1\cap J_2}f \le \operatorname{esssup}_{J_2}f , \tag{1} \end{equation*}
and one takes the supremum over \(J_1\) and the infimum over \(J_2\). Only rational endpoints matter: every \(J\ni x\) contains some \(J^{\prime}\ni x\) from the countable family \(\mathcal{J}\) of intervals with rational endpoints, so monotonicity gives
\begin{equation*} m(x)=\sup_{J\in\mathcal{J},\,J\ni x}\operatorname{essinf}_{J}f, \qquad M(x)=\inf_{J\in\mathcal{J},\,J\ni x}\operatorname{esssup}_{J}f. \tag{2} \end{equation*}
Hence the countable union
\begin{equation*} N := \bigcup_{J\in\mathcal{J}}\bigl(\{x\in J: f(x)>\operatorname{esssup}_{J}f\} \cup \{x\in J: f(x)<\operatorname{essinf}_{J}f\}\bigr) \end{equation*}
is null, and \(m\le f\le M\) on \((a,b)\setminus N\) by (2).
So \(\tilde f = f\) off \(N\), a modification, measurable by completeness of \(\lambda\) and bounded since \(m\) is; and \(m(x)\le\tilde f(x)\le M(x)\) for every \(x\in(a,b)\), by (1) where \(\tilde f=m\) and by the previous sentence elsewhere. Applying the definitions of \(m\) and \(M\) to the particular \(J\) gives (R).
(\(\circ\)) Let \(\mu\) be a probability measure, \(\varepsilon_n>0\), \(\sum_{n=1}^{\infty}\varepsilon_n<\infty\), and let \(f_n\) be \(\mu\)-measurable functions such that
\begin{equation*} \sum_{n=1}^{\infty}\mu\bigl(x:\ |f_n(x)|>\varepsilon_n\bigr)<\infty . \end{equation*}
Prove that
\begin{equation*} \sum_{n=1}^{\infty}|f_n(x)|<\infty\qquad\text{a.e.} \end{equation*}
This is the easy half of Borel–Cantelli. With \(A_n:=\{|f_n|>\varepsilon_n\}\) and \(E:=\limsup_n A_n=\bigcap_n\bigcup_{m\ge n}A_m\), countable subadditivity gives
\begin{equation*} \mu(E) \ \le\ \mu\Bigl(\bigcup_{m=n}^{\infty}A_m\Bigr) \ \le\ \sum_{m=n}^{\infty}\mu(A_m) \xrightarrow[\ n\to\infty\ ]{} 0 , \end{equation*}
the tails of the convergent series of the hypothesis, so \(\mu(E)=0\). For \(x\notin E\) there is an \(n\) with \(|f_m(x)|\le\varepsilon_m\) for all \(m\ge n\), whence
\begin{equation*} \sum_{m=1}^{\infty}|f_m(x)| \le \sum_{m=1}^{n-1}|f_m(x)| + \sum_{m=n}^{\infty}\varepsilon_m < \infty , \end{equation*}
the head being a finite sum of finite terms, once the null set on which some \(f_n\) is infinite has been discarded as well.
(\(\circ\)) Let \(f,g\colon[0,1]\to[0,1]\), where \(f\) is continuous and \(g\) is Riemann integrable. Show that the composition \(g\circ f\) may fail to be Riemann integrable.
Take \(f:=\operatorname{dist}(\cdot,K)\) and \(g:=I_{\{0\}}\), where, with \(\{r_n\}\) enumerating \(\mathbb{Q}\cap[0,1]\),
\begin{equation*} K := [0,1]\setminus\bigcup_{n=1}^{\infty}\bigl(r_n - 4^{-n-1},\ r_n + 4^{-n-1}\bigr) \end{equation*}
is compact with empty interior (its complement contains every rational, hence is dense) and \(\lambda(K) \ge 1 - 2\sum_n 4^{-n-1} = 5/6\), as in Example 1.7.6. Both map \([0,1]\) into \([0,1]\); \(f\) is continuous, distance functions being \(1\)-Lipschitz, and \(g\) is Riemann integrable, being bounded and differing from \(0\) at one point only (all lower Darboux sums vanish, and the upper ones are at most the total length of the at most two intervals containing \(0\)). As \(K\) is closed, \(\operatorname{dist}(x,K)=0\) exactly on \(K\), so \(g\circ f = I_K\).
But \(I_K\) is not Riemann integrable: for any partition \(0=x_0<\cdots<x_N=1\) with \(\Delta_i:=[x_{i-1},x_i]\), every \(\Delta_i\) contains points off \(K\) (\(K\) has empty interior), so the lower Darboux sum is \(0\), whereas the \(\Delta_i\) meeting \(K\) cover \(K\) and carry \(\sup_{\Delta_i} I_K=1\), so
\begin{equation*} U = \sum_{i:\,\Delta_i\cap K\ne\emptyset}(x_i-x_{i-1}) \ \ge\ \lambda(K) \ \ge\ \frac{5}{6} . \end{equation*}
(\(\circ\)) Let \(\mu\) be a nonnegative measure, let \(f\in L^2(\mu)\cap L^4(\mu)\), and let
\begin{equation*} \int f^2\,d\mu=\int f^3\,d\mu=\int f^4\,d\mu . \end{equation*}
Prove that \(f(x)\in\{0,1\}\) a.e.
The nonnegative function \((f^2-f)^2\) has zero integral. Indeed \(f^3\in L^1(\mu)\), since \(|f|,f^2\in L^2(\mu)\) (the latter because \(\int f^4\,d\mu<\infty\)) and Cauchy–Bunyakowsky (Corollary 2.11.3) gives
\begin{equation*} \int|f|^3\,d\mu=\int |f|\cdot f^2\,d\mu \ \le\ \Bigl(\int f^2\,d\mu\Bigr)^{1/2}\Bigl(\int f^4\,d\mu\Bigr)^{1/2}<\infty . \end{equation*}
Let \(A\) denote the common finite value of \(\int f^2\,d\mu\), \(\int f^3\,d\mu\), \(\int f^4\,d\mu\). Then \((f^2-f)^2=f^4-2f^3+f^2\) is a linear combination of integrable functions, so by linearity
\begin{equation*} \int (f^2-f)^2\,d\mu=A-2A+A=0 . \end{equation*}
By Corollary 2.5.4 the integrand vanishes a.e., i.e. \(f(f-1)=0\) a.e., that is \(f(x)\in\{0,1\}\) a.e.
(\(\circ\)) Let \(1<p<\infty\), \(p^{-1}+q^{-1}=1\). Prove that for all nonnegative \(a\) and \(b\) one has the inequality
\begin{equation*} ab\le\frac{a^p}{p}+\frac{b^q}{q}, \end{equation*}
where the equality is only possible if \(b=a^{p-1}\).
Fix \(b\ge 0\) and minimise \(\varphi(t):=t^p/p+b^q/q-tb\) on \([0,+\infty)\). Here \(p^{-1}+q^{-1}=1\) gives \(q-1=1/(p-1)\) and \((p-1)(q-1)=1\). Since \(\varphi^{\prime}(t)=t^{p-1}-b\) and \(t\mapsto t^{p-1}\) is strictly increasing on \([0,+\infty)\), the function \(\varphi\) strictly decreases on \([0,t_0]\) and strictly increases on \([t_0,+\infty)\), where \(t_0:=b^{1/(p-1)}=b^{q-1}\) is the unique root of \(t^{p-1}=b\); so \(t_0\) is the unique minimum point. As \(t_0^p=b^{q}=t_0b\),
\begin{equation*} \varphi(t_0)=\frac{b^{q}}{p}+\frac{b^{q}}{q}-b^{q} =b^{q}\Bigl(\frac1p+\frac1q-1\Bigr)=0 . \end{equation*}
Hence \(\varphi(a)\ge 0\) for every \(a\ge 0\), which is the asserted inequality, with equality exactly for \(a=t_0=b^{q-1}\); applying the strictly increasing map \(t\mapsto t^{p-1}\) turns this into
\begin{equation*} a^{p-1}=b^{(q-1)(p-1)}=b . \end{equation*}
Thus equality is only possible when \(b=a^{p-1}\).
Method (2): since \(x\mapsto x^{p-1}\) has inverse \(y\mapsto y^{q-1}\), the rectangle \(R=[0,a]\times[0,b]\) is covered by
\begin{equation*} A=\{0\le x\le a,\ 0\le y\le x^{p-1}\},\qquad B=\{0\le y\le b,\ 0\le x\le y^{q-1}\} \end{equation*}
(a point of \(R\) with \(y>x^{p-1}\) has \(x<y^{q-1}\)), and \(\lambda_2(A)=a^p/p\), \(\lambda_2(B)=b^q/q\) (Check!), whence \(ab=\lambda_2( R)\le a^p/p+b^q/q\).
Exercises 2.12.88–2.12.94
(\(\circ\)) Justify the relation (2.12.8).
For \(\lambda(x_0)\) one may take any number between the one-sided derivatives at an interior point \(x_0\); everything follows from monotonicity of the difference quotient. Convexity applied to the convex combination \(v=\frac{w-v}{w-u}u+\frac{v-u}{w-u}w\) for \(u<v<w\) in \(\mathrm{Dom}(\Psi)\) gives \((w-u)\Psi(v)\le(w-v)\Psi(u)+(v-u)\Psi(w)\); subtracting \((w-u)\Psi(u)\), respectively \((w-u)\Psi(w)\), and dividing by the positive numbers \((w-u)(v-u)\), respectively \((w-u)(w-v)\), yields the chord chain
\begin{equation*} \frac{\Psi(v)-\Psi(u)}{v-u} \le \frac{\Psi(w)-\Psi(u)}{w-u} \le \frac{\Psi(w)-\Psi(v)}{w-v}. \end{equation*}
Hence \(Q(h):=h^{-1}\bigl(\Psi(x_0+h)-\Psi(x_0)\bigr)\) is nondecreasing in \(h\): for admissible \(h_1<h_2\) apply the chain to (i) \(x_0<x_0+h_1<x_0+h_2\) when \(0<h_1\), reading off its first inequality; (ii) \(x_0+h_1<x_0+h_2<x_0\) when \(h_2<0\), reading off its second; (iii) \(x_0+h_1<x_0<x_0+h_2\) when \(h_1<0<h_2\), reading off its outer terms. Consequently the one-sided limits exist and, \(Q\) being monotone,
\begin{equation*} \Psi^{\prime}_-(x_0)=\liminf_{h\to0}Q(h)=\sup_{h<0}Q(h)=:D^- , \end{equation*}
\begin{equation*} \Psi^{\prime}_+(x_0)=\limsup_{h\to0}Q(h)=\inf_{h>0}Q(h)=:D^+ , \end{equation*}
with \(D^-\le D^+\), both finite for interior \(x_0\) because \(Q(h_1)\le D^-\le D^+\le Q(h_2)\) for admissible \(h_1<0<h_2\). Fix \(\lambda\in[D^-,D^+]\) and put \(h=x-x_0\): if \(h>0\) then \(Q(h)\ge D^+\ge\lambda\), while if \(h<0\) then \(Q(h)\le D^-\le\lambda\) and multiplication by \(h<0\) reverses the inequality, so in both cases
\begin{equation*} \Psi(x)-\Psi(x_0)=h\,Q(h)\ge\lambda h=\lambda(x-x_0), \end{equation*}
trivially also at \(x=x_0\). This is (2.12.8).
At an endpoint adjoined to \(\mathrm{Dom}(\Psi)\) no real \(\lambda(x_0)\) need exist (for \(\Psi(x)=-\sqrt{x}\) on \([0,1]\) at \(x_0=0\) one has \(D^+=\inf_{h>0}(-h^{-1/2})=-\infty\)), and this costs nothing: in the only application, Theorem 2.12.19, the case \(x_0=\int f\,d\mu=a\) makes \(f-a\ge0\) have zero integral, so \(f=a\) a.e. and (2.12.9) is an equality, and likewise at \(b\).
(\(\circ\)) Let \(1 < p < \infty\), \(p^{-1} + q^{-1} = 1\), \(f \in \mathcal{L}^p(\mu)\), \(g \in \mathcal{L}^q(\mu)\), and let
\begin{equation*} \int f g \,d\mu = \|f\|_p \|g\|_q > 0. \end{equation*}
Prove that \(g = \mathrm{sign}\, f \cdot |f|^{p-1}\) a.e.
Put \(A=\|f\|_p\), \(B=\|g\|_q\), both finite and positive since \(AB>0\), and normalize: \(F=|f|/A\), \(G=|g|/B\), so that \(\int F^p\,d\mu=\int G^q\,d\mu=1\). By Young’s inequality with its equality case (Exercise 2.12.87) one has pointwise \(FG\le F^p/p+G^q/q\), with equality exactly when \(G=F^{p-1}\); the majorant is integrable with integral \(1/p+1/q=1\), so \(\int FG\,d\mu\le1\) and, using \(fg\le|f||g|\) pointwise,
\begin{equation*} AB=\int fg\,d\mu\ \le\ \int|f||g|\,d\mu=AB\int FG\,d\mu\ \le\ AB . \end{equation*}
Hence both inequalities are equalities, and each of two nonnegative integrands has zero integral:
(i) \(\int(|fg|-fg)\,d\mu=0\), so \(f(x)g(x)\ge0\) a.e.;
(ii) \(\int\bigl(F^p/p+G^q/q-FG\bigr)\,d\mu=1-1=0\), so by the equality case applied at a.e. \(x\) (with \(a=F(x)\), \(b=G(x)\)) we get \(G=F^{p-1}\) a.e., i.e. \(|g|=c|f|^{p-1}\) a.e. with \(c:=BA^{1-p}>0\).
On \(\{f\ne0\}\) the sign of \(g\) agrees with that of \(f\) by (i), and on \(\{f=0\}\) (ii) forces \(g=0\); therefore
\begin{equation*} g=c\,\mathrm{sign}\, f\cdot|f|^{p-1}\quad\text{a.e.}, \qquad c=\frac{\|g\|_q}{\|f\|_p^{\,p-1}} . \end{equation*}
The printed statement omits a normalization: the hypothesis is invariant under \(g\mapsto tg\) with \(t>0\), and since \(\|g\|_q=c\|f\|_p^{p-1}\) the constant is \(c=1\) precisely when \(\|g\|_q=\|f\|_p^{p-1}\), in particular when \(\|f\|_p=\|g\|_q=1\).
(\(\circ\)) Let \(\mu\) be a probability measure and let \(f\) be a nonnegative \(\mu\)-integrable function such that \(\ln f \in \mathcal{L}^1(\mu)\). Prove that
\begin{equation*} \lim_{p \to 0+} \int \frac{f^p - 1}{p} \, d\mu = \int \ln f \, d\mu . \end{equation*}
Dominated convergence with the majorant \(\Phi:=|f-1|+|\ln f|\). Integrability of \(f\) and of \(\ln f\) gives \(0<f<\infty\) a.e., and discarding a null set we assume this everywhere. The key bound is
\begin{equation*} \frac{|t^p-1|}{p}\ \le\ |t-1|+|\ln t|\qquad (t>0,\ p\in(0,1)) , \end{equation*}
proved by cases:
(i) \(t\ge1\): with \(u=\ln t\ge0\), convexity of \(p\mapsto e^{pu}\) on \([0,1]\) bounds it by the chord, \(e^{pu}\le(1-p)+pe^{u}\), i.e. \(t^p-1\le p(t-1)\), and \(t^p\ge1\), so the left side is \(\le t-1\);
(ii) \(0<t\le1\): then \(t^p\le1\) and, with \(s=p|\ln t|\ge0\), the inequality \(e^{-s}\ge1-s\) gives \((1-e^{-s})/p\le s/p=|\ln t|\).
Hence \(h_p:=(f^p-1)/p\) satisfies \(|h_p|\le\Phi\) for all \(p\in(0,1)\), and \(\Phi\in\mathcal{L}^1(\mu)\) since \(f,\ln f\in\mathcal{L}^1(\mu)\) and \(\mu(X)=1\). Pointwise, with \(a=\ln f(x)\in\mathbb{R}\),
\begin{equation*} h_p(x)=\frac{e^{pa}-1}{p}\ \longrightarrow\ a=\ln f(x)\qquad (p\to0+), \end{equation*}
the derivative at \(0\) of \(p\mapsto e^{pa}\). Applying the dominated convergence theorem (Theorem 2.8.1) along an arbitrary sequence \(p_n\to0+\) and noting that the limit is the same for all such sequences,
\begin{equation*} \lim_{p\to0+}\int\frac{f^p-1}{p}\,d\mu=\int\ln f\,d\mu . \end{equation*}
(\(\circ\)) Let \(\mu\) be a probability measure and let \(f\) be a nonnegative \(\mu\)-integrable function such that \(\ln f \in \mathcal{L}^1(\mu)\). Prove that
\begin{equation*} \lim_{p \to 0+} \left( \int f^p \, d\mu \right)^{1/p} = \exp \int \ln f \, d\mu . \end{equation*}
Take logarithms and quote Exercise 2.12.90. As there we may assume \(0<f<\infty\) everywhere, and \(f^p\le1+f\) is integrable for \(p\in(0,1)\). Put \(I_p=\int f^p\,d\mu\) and \(L=\int\ln f\,d\mu\in\mathbb{R}\); here \(I_p>0\), since \(I_p=0\) would force \(f^p=0\) a.e. Because \(\mu(X)=1\),
\begin{equation*} \frac{I_p-1}{p}=\int\frac{f^p-1}{p}\,d\mu\ \xrightarrow[p\to0+]{}\ L \end{equation*}
by Exercise 2.12.90; in particular \(u_p:=I_p-1=p\bigl(L+o(1)\bigr)\to0\). Since \(\ln(1+u)=u+o(u)\), the function \(\psi(u)=u^{-1}\ln(1+u)\), \(\psi(0)=1\), is continuous at \(0\), whence
\begin{equation*} \frac{\ln I_p}{p}=\psi(u_p)\cdot\frac{I_p-1}{p}\ \longrightarrow\ 1\cdot L=L \end{equation*}
(both sides vanish when \(u_p=0\)). By continuity of the exponential,
\begin{equation*} \lim_{p\to0+}\Bigl(\int f^p\,d\mu\Bigr)^{1/p} =\lim_{p\to0+}\exp\frac{\ln I_p}{p}=\exp\int\ln f\,d\mu . \end{equation*}
(\(\circ\)) Let \(\mu\) be a probability measure and let \(f \in \mathcal{L}^1(\mu)\). Prove that
\begin{equation*} 1 + \left( \int |f| \, d\mu \right)^{2} \le \left( \int \sqrt{1 + |f|^2} \, d\mu \right)^{2} \le \left( 1 + \int |f| \, d\mu \right)^{2} . \end{equation*}
The right-hand bound is \(\sqrt{1+s^2}\le1+s\), the left-hand one is Jensen’s inequality for \(\varphi(t)=\sqrt{1+t^2}\). Put \(m=\int|f|\,d\mu<\infty\). Since \(1+s^2\le(1+s)^2\) for \(s\ge0\), we have \(1\le\varphi(|f|)\le1+|f|\), so \(\varphi(|f|)\) is integrable (\(\mu\) is finite, \(|f|\) integrable) and
\begin{equation*} \int\sqrt{1+|f|^2}\,d\mu\ \le\ 1+m , \end{equation*}
which squares to the second inequality. For the first, \(\varphi\) is convex because \(\varphi^{\prime\prime}(t)=(1+t^2)^{-3/2}>0\), and the hypotheses of Theorem 2.12.19 hold (\(\mu\) a probability measure, \(|f|\) integrable with values in \(\mathrm{Dom}(\varphi)=\mathbb{R}\), \(\varphi(|f|)\) integrable as just shown), so Jensen’s inequality gives
\begin{equation*} \sqrt{1+m^2}=\varphi\Bigl(\int|f|\,d\mu\Bigr)\ \le\ \int\sqrt{1+|f|^2}\,d\mu . \end{equation*}
Squaring the nonnegative sides of the two displays yields the asserted chain.
(\(\circ\)) Let \(f, g \ge 0\) be integrable functions on a space with a probability measure \(\mu\) and let \(fg \ge 1\). Show that
\begin{equation*} \int f \, d\mu \int g \, d\mu \ge 1 . \end{equation*}
Apply Cauchy–Bunyakowsky to \(u=\sqrt{f}\), \(v=\sqrt{g}\). These lie in \(\mathcal{L}^2(\mu)\), since \(\int u^2\,d\mu=\int f\,d\mu<\infty\) and \(\int v^2\,d\mu=\int g\,d\mu<\infty\), and \(uv=\sqrt{fg}\ge1\) a.e., so \(\int uv\,d\mu\ge\mu(X)=1\). Hence by Corollary 2.11.3 (Hölder with \(p=q=2\)),
\begin{equation*} 1\ \le\ \int uv\,d\mu \ \le\ \Bigl(\int f\,d\mu\Bigr)^{1/2}\Bigl(\int g\,d\mu\Bigr)^{1/2}, \end{equation*}
and squaring the nonnegative sides gives \(\int f\,d\mu\int g\,d\mu\ge1\).
(\(\circ\)) Let \(\mu\) be a countably additive measure with values in \([0, +\infty]\) and let \(f \in \mathcal{L}^1(\mu)\) be such that \(f - 1 \in \mathcal{L}^p(\mu)\) for some \(p \in [1, \infty)\). Prove that the measure \(\mu\) is finite.
Split \(X=A\cup B\) with \(A=\{f\ge1/2\}\), \(B=\{f<1/2\}\) and apply Chebyshev’s inequality twice. That inequality (Theorem 2.5.3, extended to measures with values in \([0,+\infty]\) by Proposition 2.6.4) reads \(\mu(|h|\ge R)\le R^{-1}\int|h|\,d\mu\) for integrable \(h\) and \(R>0\); its right-hand side being finite is what forces the left-hand side to be finite. On \(A\) one has \(|f|\ge1/2\), and on \(B\) one has \(f-1<-1/2\), hence \(|f-1|^p>2^{-p}\); so with \(h=f\), \(R=1/2\) and with \(h=|f-1|^p\), \(R=2^{-p}\),
\begin{equation*} \mu(A)\le2\int_X|f|\,d\mu=2\|f\|_1,\qquad \mu(B)\le2^{p}\int_X|f-1|^{p}\,d\mu=2^{p}\|f-1\|_p^{p}, \end{equation*}
both finite because \(f\in\mathcal{L}^1(\mu)\) and \(f-1\in\mathcal{L}^p(\mu)\). By additivity \(\mu(X)=\mu(A)+\mu(B)<\infty\).
Exercises 2.12.95–2.12.101
Let \(\mu\) be a probability measure, let \(\{f_n\}\subset L^1(\mu)\), and let \(I_n\) be the integral of \(f_n\). Suppose that there exists \(c>0\) such that
\begin{equation*} \|f_n - I_n\|_p^p \le c\,\|f_n\|_1, \qquad \forall\, n\in\mathbb{N}. \end{equation*}
Prove that either
\begin{equation*} \limsup_{n\to\infty}\|f_n\|_1<\infty \quad\text{and}\quad \liminf_{n\to\infty}|f_n(x)|<\infty \ \text{ a.e.}, \end{equation*}
or
\begin{equation*} \limsup_{n\to\infty}\|f_n\|_1=\infty \quad\text{and}\quad \limsup_{n\to\infty}|f_n(x)|=\infty \ \text{ a.e.} \end{equation*}
The printed statement omits the exponent range: \(p>1\) is meant, since for \(p=1\) one always has \(\|f_n-I_n\|_1\le2\|f_n\|_1\) and the hypothesis is vacuous. Here \(I_n\) is read as the constant function \(I_n\), so \(f_n\in L^p(\mu)\) for every \(n\). The first halves of the two alternatives being exhaustive, two implications are needed.
(i) Let \(\limsup_n\|f_n\|_1=:M<\infty\) and pick a subsequence with \(\int|f_{n_k}|\,d\mu\to M\). The \(|f_{n_k}|\) are nonnegative with uniformly bounded integrals, so Fatou’s theorem (Theorem 2.8.3, in the form of Corollary 2.8.4) gives
\begin{equation*} \int_X\liminf_{k\to\infty}|f_{n_k}|\,d\mu \ \le\ \liminf_{k\to\infty}\int_X|f_{n_k}|\,d\mu=M<\infty , \end{equation*}
whence \(\liminf_n|f_n|\le\liminf_k|f_{n_k}|<\infty\) a.e. (this half does not use the hypothesis).
(ii) Let \(\limsup_n\|f_n\|_1=\infty\). Pass to a subsequence with \(\|f_n\|_1\to\infty\) and \(\|f_n\|_1>0\): the hypothesis is inherited by subsequences, and an upper limit along a subsequence does not exceed the one along the whole sequence, so it suffices to conclude there. Put \(g_n:=f_n\|f_n\|_1^{-1/p}\) and \(C_n:=I_n\|f_n\|_1^{-1/p}\); dividing the hypothesis by \(\|f_n\|_1\),
\begin{equation*} \|g_n-C_n\|_p^{p}=\frac{\|f_n-I_n\|_p^{p}}{\|f_n\|_1}\le c . \tag{1} \end{equation*}
The sequence \(\{C_n\}\) is unbounded. Indeed, if \(|C_n|\le K\) for all \(n\), then the constant \(C_n\) has \(L^p\)-norm \(|C_n|\) (as \(\mu(X)=1\)), so (1) and the triangle inequality give \(\|g_n\|_p\le c^{1/p}+K=:A^{1/p}\), that is \(\|f_n\|_p^{p}\le A\|f_n\|_1\le A\|f_n\|_p\), the last step by Hölder’s inequality on a probability space; hence \(\|f_n\|_p\le A^{1/(p-1)}\) (here \(p>1\) is used; the bound is trivial when \(\|f_n\|_p=0\)) and \(\|f_n\|_1\le\|f_n\|_p\) is bounded, contradicting \(\|f_n\|_1\to\infty\).
Passing to a further subsequence we may assume \(C_n\to+\infty\) or \(C_n\to-\infty\), and in the second case we replace \(f_n\) by \(-f_n\), which changes neither the hypothesis nor the conclusion; so \(C_n\to+\infty\). By (1) and Fatou’s theorem,
\begin{equation*} \int_X\liminf_{n\to\infty}|g_n-C_n|^{p}\,d\mu \ \le\ \liminf_{n\to\infty}\int_X|g_n-C_n|^{p}\,d\mu\ \le\ c<\infty , \end{equation*}
so \(M(x):=\liminf_n|g_n(x)-C_n|<\infty\) for all \(x\) outside a null set \(N\). Fix such an \(x\) and indices \(n_1<n_2<\cdots\) with \(|g_{n_j}(x)-C_{n_j}|\le M(x)+1\); then \(g_{n_j}(x)\ge C_{n_j}-M(x)-1\), so once \(C_{n_j}>M(x)+1\),
\begin{equation*} |f_{n_j}(x)|\ \ge\ \|f_{n_j}\|_1^{1/p}\bigl(C_{n_j}-M(x)-1\bigr) \ \longrightarrow\ +\infty . \end{equation*}
Hence \(\limsup_n|f_n(x)|=\infty\) for every \(x\notin N\), along this subsequence and therefore along the original sequence.
(\(\circ\)) Let \(f\in L^1[a,b]\) and let
\begin{equation*} \int_a^b t^k f(t)\,dt=0 \end{equation*}
for all nonnegative integer \(k\). Show that \(f=0\) a.e.
Approximate \(g:=\operatorname{sign}f\) by uniformly bounded polynomials and pass to the limit. By linearity the hypothesis says
\begin{equation*} \int_a^b P(t)f(t)\,dt=0\qquad\text{for every polynomial }P. \tag{1} \end{equation*}
Here \(|g|\le1\). By Lusin’s theorem (Theorem 2.2.10) choose for each \(j\) a continuous \(u_j\) and a compact \(K_j\subset[a,b]\) with \(\lambda\bigl([a,b]\setminus K_j\bigr)<2^{-j}\) and \(u_j=g\) on \(K_j\); the truncation \(h_j:=\max\bigl(-1,\min(1,u_j)\bigr)\) is continuous, satisfies \(|h_j|\le1\), and still equals \(g\) on \(K_j\). By the Weierstrass approximation theorem pick a polynomial \(P_j\) with \(\sup_{[a,b]}|P_j-h_j|\le2^{-j}\); then \(|P_j|\le2\) on \([a,b]\) for every \(j\). For
\begin{equation*} N:=\bigcap_{m=1}^{\infty}\ \bigcup_{j\ge m}\bigl([a,b]\setminus K_j\bigr) \end{equation*}
we get \(\lambda(N)\le\sum_{j\ge m}2^{-j}=2^{1-m}\) for every \(m\), so \(\lambda(N)=0\); and if \(t\notin N\), then \(t\in K_j\) for all \(j\ge m\) with some \(m\), whence \(|P_j(t)-g(t)|\le2^{-j}\to0\). Thus \(P_j\to g\) a.e., \(|P_jf|\le2|f|\in L^1[a,b]\), and the dominated convergence theorem (Theorem 2.8.1) together with (1) gives
\begin{equation*} 0=\int_a^b P_j(t)f(t)\,dt\ \longrightarrow\ \int_a^b g(t)f(t)\,dt=\int_a^b|f(t)|\,dt . \end{equation*}
Hence \(\int_a^b|f|\,dt=0\) and \(f=0\) a.e.
(G. Hardy) Let \(f\) be a nonnegative measurable function on \([0,+\infty)\) and let \(1\le q<\infty\), \(0<r<\infty\). Show that
\begin{equation*} \int_0^{\infty}\left(\int_0^{t}f(s)\,ds\right)^{q}t^{-r-1}\,dt\ \le\ \left(\frac{q}{r}\right)^{q}\int_0^{\infty}s^{q-r-1}f(s)^{q}\,ds . \end{equation*}
Tonelli plus one Hölder split. All integrands are nonnegative and measurable (the function \((s,t)\mapsto f(s)\mathbf{1}_{\{s<t\}}\) is product-measurable, \(\{s<t\}\) being open, and \(F(t)=\int_0^tf\,ds\) is nondecreasing), so every integral is an element of \([0,+\infty]\), Tonelli’s theorem (Theorem 3.4.5) applies on \((0,\infty)^2\), and there is nothing to prove when the right-hand side is infinite.
(i) \(q=1\). By Tonelli,
\begin{equation*} \int_0^{\infty}\Bigl(\int_0^{t}f(s)\,ds\Bigr)t^{-r-1}\,dt =\int_0^{\infty}f(s)\Bigl(\int_s^{\infty}t^{-r-1}\,dt\Bigr)ds =\frac1r\int_0^{\infty}f(s)s^{-r}\,ds , \end{equation*}
which is the right-hand side, with equality.
(ii) \(q>1\). Let \(p=q/(q-1)\) and \(\alpha:=p^{-1}(1-r/q)\), so that \(\alpha p=1-r/q<1\). Hölder’s inequality on \((0,t)\) applied to \(f(s)=\bigl(f(s)s^{\alpha}\bigr)s^{-\alpha}\), together with
\begin{equation*} \int_0^{t}s^{-\alpha p}\,ds=\frac{t^{\,1-\alpha p}}{1-\alpha p}=\frac{q}{r}\,t^{\,r/q}, \end{equation*}
gives on raising to the power \(q\), for every \(t>0\) and in \([0,+\infty]\),
\begin{equation*} \Bigl(\int_0^{t}f(s)\,ds\Bigr)^{q} \ \le\ \Bigl(\frac{q}{r}\Bigr)^{q/p}t^{\,r/p}\int_0^{t}f(s)^{q}s^{\alpha q}\,ds . \tag{1} \end{equation*}
Multiply (1) by \(t^{-r-1}\), note \(r/p-r-1=-r/q-1\), and integrate; by Tonelli the resulting double integral equals
\begin{equation*} \int_0^{\infty}f(s)^{q}s^{\alpha q}\Bigl(\int_s^{\infty}t^{-r/q-1}\,dt\Bigr)ds =\frac{q}{r}\int_0^{\infty}f(s)^{q}s^{\alpha q-r/q}\,ds . \end{equation*}
Since \(\alpha q=(q-r)/p\) and \(1/p=1-1/q\), the exponent is \(\alpha q-r/q=(q-r)(1-1/q)-r/q=q-r-1\), while the constant is \((q/r)^{q/p+1}=(q/r)^{q}\). This is the asserted inequality.
(P.Yu. Glazyrina) Let \(f\ge 0\) be a \(\mu\)-measurable function. Prove the inequality
\begin{equation*} \int f^{p}\,d\mu\int f^{s-p}\,d\mu\ \le\ \int f^{q}\,d\mu\int f^{s-q}\,d\mu \end{equation*}
assuming that \(p,q,s\) are real numbers such that \(|p-s/2|<|q-s/2|\) and the above integrals exist.
Split the exponents by Hölder twice, with the conjugate pair \(r,t\) below. As the statement implies, \(0<f<+\infty\) \(\mu\)-a.e., so \(f^{\lambda+\nu}=f^{\lambda}f^{\nu}\) a.e. and the four integrals are elements of \([0,+\infty]\). Write \(u:=p-s/2\), \(v:=q-s/2\), so the hypothesis reads \(|u|<|v|\) and \(v\ne0\); since \(s-2q=-2v\), \(p-q=u-v\), \(s-p-q=-u-v\),
\begin{equation*} r:=\frac{s-2q}{p-q}=\frac{2v}{v-u},\qquad t:=\frac{s-2q}{s-p-q}=\frac{2v}{v+u}, \end{equation*}
with nonzero denominators, both of the sign of \(v\). They are conjugate,
\begin{equation*} \frac1r+\frac1t=\frac{(p-q)+(s-p-q)}{s-2q}=\frac{s-2q}{s-2q}=1 , \end{equation*}
and both exceed \(1\): for \(v>0\) the inequalities \(r>1\), \(t>1\) read \(u+v>0\) and \(v>u\), true since \(|u|<v\); for \(v<0\) division by negative denominators reverses them into \(u+v<0\) and \(v<u\), true since \(|u|<-v\). (Check!)
Put \(\alpha:=q/t\). Then \(\alpha t=q\) and
\begin{equation*} (p-\alpha)r=\Bigl(p-q+\frac{q}{r}\Bigr)r=(p-q)r+q=(s-2q)+q=s-q , \end{equation*}
so Hölder’s inequality with exponents \(t,r\) applied to \(f^{p}=f^{\alpha}f^{p-\alpha}\) gives
\begin{equation*} \int f^{p}\,d\mu \ \le\ \Bigl(\int f^{q}\,d\mu\Bigr)^{1/t}\Bigl(\int f^{s-q}\,d\mu\Bigr)^{1/r}. \tag{1} \end{equation*}
Put \(\gamma:=q/r\), \(\delta:=(s-q)/t\). Then \(\gamma r=q\), \(\delta t=s-q\) and
\begin{equation*} \gamma+\delta=\frac{q(p-q)+(s-q)(s-p-q)}{s-2q}=\frac{(s-p)(s-2q)}{s-2q}=s-p \end{equation*}
(both numerators equal \(2pq+s^{2}-2sq-sp\)), so \(f^{s-p}=f^{\gamma}f^{\delta}\) and Hölder with exponents \(r,t\) gives
\begin{equation*} \int f^{s-p}\,d\mu \ \le\ \Bigl(\int f^{q}\,d\mu\Bigr)^{1/r}\Bigl(\int f^{s-q}\,d\mu\Bigr)^{1/t}. \tag{2} \end{equation*}
Multiplying (1) by (2) and using \(r^{-1}+t^{-1}=1\) yields the assertion.
(Fukuda, Vakhania, Kvaratskhelia) Let \(\mu\) be a probability measure and let \(f\in L^{p}(\mu)\) be such that \(\|f\|_{L^{p}(\mu)}\le C\|f\|_{L^{q}(\mu)}\) for some \(q\in[1,p)\) and \(C\ge 1\). Show that
\begin{equation*} \|f\|_{L^{r}(\mu)}\le C^{\kappa}\|f\|_{L^{s}(\mu)} \end{equation*}
whenever \(1\le s<r\le p\), where \(\kappa=1\) if \(q\le s<r\le p\), \(\kappa=q(p-s)\bigl(s(p-q)\bigr)^{-1}\) if \(s<q<r\le p\), and \(\kappa=p(q-s)\bigl(s(p-q)\bigr)^{-1}\) if \(s<r\le q\).
Everything follows from the monotonicity of \(a\mapsto\|f\|_{L^{a}(\mu)}\) on a probability space plus one interpolation lemma. By Hölder’s inequality,
\begin{equation*} \|f\|_{L^{a}(\mu)}\le\|f\|_{L^{b}(\mu)}\qquad (1\le a\le b), \tag{M} \end{equation*}
and all norms below are finite as \(f\in L^{p}(\mu)\); if \(\|f\|_{L^{p}(\mu)}=0\) then \(f=0\) a.e. and all is trivial, so we assume \(\|f\|_{L^{a}(\mu)}\in(0,\infty)\) for \(1\le a\le p\).
Lemma. If \(1\le s<q\), then \(\|f\|_{L^{q}(\mu)}\le C^{\,p(q-s)/(s(p-q))}\|f\|_{L^{s}(\mu)}\).
Take \(t:=(p-s)/(q-s)>1\) with conjugate \(t^{\prime}=(p-s)/(p-q)\), and split \(|f|^{q}=|f|^{\alpha q}|f|^{\beta q}\) where \(\alpha q:=p/t=p(q-s)/(p-s)\) and \(\beta q:=s/t^{\prime}=s(p-q)/(p-s)\); these are positive and sum to \(q\), since \(p(q-s)+s(p-q)=q(p-s)\). Writing \(a:=\alpha q\), so that \(q-a=\beta q=s(p-q)/(p-s)>0\), Hölder’s inequality with the exponents \(t,t^{\prime}\) and then the hypothesis give
\begin{equation*} \|f\|_{L^{q}(\mu)}^{q}\ \le\ \|f\|_{L^{p}(\mu)}^{\,a}\|f\|_{L^{s}(\mu)}^{\,q-a} \ \le\ C^{\,a}\|f\|_{L^{q}(\mu)}^{\,a}\|f\|_{L^{s}(\mu)}^{\,q-a} . \end{equation*}
Dividing by the finite positive number \(\|f\|_{L^{q}(\mu)}^{a}\) and raising to the power \((q-a)^{-1}=(p-s)/\bigl(s(p-q)\bigr)>0\) proves the Lemma.
The three cases exhaust \(1\le s<r\le p\): if \(q\le s\) we are in (i); if \(s<q\), then (iii) applies when \(r\le q\) and (ii) when \(q<r\).
(i) \(q\le s<r\le p\). By (M), the hypothesis and (M) again, \(\|f\|_{L^{r}(\mu)}\le\|f\|_{L^{p}(\mu)}\le C\|f\|_{L^{q}(\mu)}\le C\|f\|_{L^{s}(\mu)}\), i.e. \(\kappa=1\).
(ii) \(s<q<r\le p\). By (M), the hypothesis and the Lemma, \(\|f\|_{L^{r}(\mu)}\le C^{\,1+p(q-s)/(s(p-q))}\|f\|_{L^{s}(\mu)}\), and
\begin{equation*} 1+\frac{p(q-s)}{s(p-q)}=\frac{s(p-q)+p(q-s)}{s(p-q)}=\frac{q(p-s)}{s(p-q)}=\kappa . \end{equation*}
(iii) \(s<r\le q\). By (M) and the Lemma, \(\|f\|_{L^{r}(\mu)}\le\|f\|_{L^{q}(\mu)}\le C^{\,p(q-s)/(s(p-q))}\|f\|_{L^{s}(\mu)}\), the asserted estimate with \(\kappa=p(q-s)\bigl(s(p-q)\bigr)^{-1}\).
(i) Let \(E\) be a partially ordered real vector space such that if \(x\le y\), then \(tx\le ty\) for all \(t\ge 0\) and \(x+z\le y+z\) for all \(z\in E\). Suppose that \(E_{0}\) is a linear subspace in \(E\) such that, for each \(x\in E\), there exists an element \(x_{0}\in E_{0}\) with \(x\le x_{0}\). Let \(L_{0}\) be a linear function on \(E_{0}\) such that \(L_{0}(v)\ge 0\) whenever \(v\in E_{0}\) and \(v\ge 0\). Prove that \(L_{0}\) can be extended to a linear function \(L\) on \(E\) such that \(L(x)\ge 0\) for all \(x\ge 0\).
(ii) Deduce from (i) the existence of a nonnegative finitely additive function on the class of all subsets of \([0,1]\) extending Lebesgue measure.
(iii) Deduce from (i) the existence of a generalized limit on the space \(m\) of all bounded sequences, i.e., a linear function \(\Lambda\) on \(m\) such that \(\Lambda(x)\ge 0\) for all \(x=(x_{n})\) with \(x_{n}\ge 0\) and \(\Lambda(x)=\lim_{n\to\infty}x_{n}\) for all convergent sequences \(x=(x_{n})\).
(i) Apply the Hahn–Banach theorem to the majorant
\begin{equation*} p(x):=\inf\{L_{0}(v):\ v\in E_{0},\ x\le v\},\qquad x\in E . \end{equation*}
It is real-valued: the set is nonempty by the hypothesis on \(E_{0}\), so \(p(x)<+\infty\); and choosing \(u_{0}\in E_{0}\) with \(-x\le u_{0}\) we get \(0\le x+u_{0}\le v+u_{0}\in E_{0}\) for every admissible \(v\), so \(L_{0}(v)\ge-L_{0}(u_{0})\) by positivity of \(L_{0}\), whence \(p(x)>-\infty\). It is sublinear: for \(\alpha>0\) one has \(x\le v\) iff \(\alpha x\le\alpha v\) (multiply by \(\alpha\), resp. \(\alpha^{-1}\)), so \(p(\alpha x)=\alpha p(x)\), while \(p(0)=0\) because \(0\le0\) gives \(p(0)\le L_{0}(0)=0\) and positivity gives \(p(0)\ge0\); and if \(x\le v\), \(y\le w\) with \(v,w\in E_{0}\), then \(x+y\le v+y\le v+w\in E_{0}\), so \(p(x+y)\le L_{0}(v)+L_{0}(w)\) and, taking the two infima separately, \(p(x+y)\le p(x)+p(y)\). Finally \(p=L_{0}\) on \(E_{0}\): here \(x\le x\) gives \(p(x)\le L_{0}(x)\), while \(x\le v\) gives \(0\le v-x\in E_{0}\), hence \(L_{0}(v)\ge L_{0}(x)\).
By the Hahn–Banach theorem (Theorem 1.12.26) there is a linear \(L\) on \(E\) with \(L|_{E_{0}}=L_{0}\) and \(L\le p\). If \(x\ge0\), then \(-x\le 0\in E_{0}\), so \(L(-x)\le p(-x)\le L_{0}(0)=0\) and \(L(x)=-L(-x)\ge0\).
(ii) Let \(E\) be the bounded real functions on \([0,1]\) with the pointwise order (which satisfies the two compatibility conditions), \(E_{0}\) the bounded Lebesgue measurable ones, and \(L_{0}(v)=\int_{0}^{1}v\,d\lambda\), linear and nonnegative on nonnegative functions; the hypothesis of (i) holds since \(x\le c\cdot\mathbf{1}\in E_{0}\) with \(c=\sup_{s}|x(s)|<\infty\). With \(L\) as in (i) put \(\nu(A):=L(\mathbf{1}_{A})\) for \(A\subset[0,1]\). Then \(\nu\ge0\); for disjoint \(A,B\) we have \(\mathbf{1}_{A\cup B}=\mathbf{1}_{A}+\mathbf{1}_{B}\), so \(\nu\) is finitely additive on all subsets; and \(\nu(A)=L_{0}(\mathbf{1}_{A})=\lambda(A)\) for Lebesgue measurable \(A\).
(iii) Let \(E=m\) with the coordinatewise order, \(E_{0}\) the convergent sequences and \(L_{0}(x)=\lim_{n}x_{n}\), linear and nonnegative on nonnegative convergent sequences; again \(x\le c\cdot(1,1,\dots)\in E_{0}\) with \(c=\sup_{n}|x_{n}|<\infty\). The functional \(\Lambda:=L\) furnished by (i) is linear on \(m\), satisfies \(\Lambda(x)\ge0\) whenever \(x_{n}\ge0\) for all \(n\), and \(\Lambda(x)=\lim_{n}x_{n}\) on convergent sequences; that is, \(\Lambda\) is a generalized limit.
(S. Banach) (i) Prove that on the space \(L\) of all bounded functions on \([0,1)\) there exists a linear function \(\Lambda\) with the following properties:
(a) if \(f\in L\) is Lebesgue integrable, then \(\Lambda(f)\) coincides with the Lebesgue integral of \(f\) over \([0,1)\),
(b) if \(f\in L\) and \(f\ge 0\), then \(\Lambda(f)\ge 0\),
(c) \(\Lambda\bigl(f(\,\cdot\,+s)\bigr)=\Lambda(f)\) for all \(f\in L\) and \(s\in[0,1]\), where \(f(t+s)=f\bigl(\mathrm{fr}(t+s)\bigr)\), \(\mathrm{fr}(s)\) is the fractional part of \(s\).
(ii) Construct a linear function on \(L\) that coincides with the integral on the set of all Riemann integrable functions, but differs from the Lebesgue integral at some Lebesgue integrable function.
(i) Hahn–Banach with the translation-invariant sublinear majorant of Example 1.12.27(ii). Pass to periodic functions: \(f\mapsto\widetilde f\), \(\widetilde f(t)=f(\mathrm{fr}(t))\), is an order-preserving linear bijection of \(L\) onto the space \(\widetilde L\) of bounded \(1\)-periodic functions on \(\mathbb{R}\), carrying the Lebesgue integrable functions onto those \(f\in\widetilde L\) whose restriction to \([0,1)\) is measurable and the operation \(f\mapsto f(\,\cdot\,+s)\) of the statement into genuine translation; we write \(f\) for \(\widetilde f\). Put
\begin{equation*} p(f):=\inf_{n,\ a_{1},\dots,a_{n}}S(f,a_{1},\dots,a_{n}),\qquad S(f,a_{1},\dots,a_{n}):=\sup_{t\in\mathbb{R}}\frac1n\sum_{i=1}^{n}f(t+a_{i}), \end{equation*}
sublinear by Example 1.12.27(ii) and real-valued on \(\widetilde L\) because \(\inf f\le S\le\sup f\). On the subspace \(L_{0}\subset\widetilde L\) of functions integrable on \([0,1)\) let \(\Lambda_{0}(f)=\int_{0}^{1}f(x)\,dx\). Translation invariance of Lebesgue measure and periodicity give \(\int_{0}^{1}f(t+a)\,dt=\int_{0}^{1}f(t)\,dt\) for every \(a\in\mathbb{R}\), so
\begin{equation*} \Lambda_{0}(f)=\int_{0}^{1}\frac1n\sum_{i=1}^{n}f(t+a_{i})\,dt \ \le\ S(f,a_{1},\dots,a_{n}), \end{equation*}
whence \(\Lambda_{0}\le p\) on \(L_{0}\). By the Hahn–Banach theorem (Theorem 1.12.26) there is a linear \(\Lambda\) on \(\widetilde L\) with \(\Lambda|_{L_{0}}=\Lambda_{0}\) and \(\Lambda\le p\).
(a) holds by construction. (b): if \(f\le0\), every average is \(\le0\), so \(p(f)\le0\) and \(\Lambda(f)\le0\); applying this to \(-f\) gives \(\Lambda(f)\ge0\) for \(f\ge0\). (c): fix \(h\in\mathbb{R}\), put \(g:=f(\,\cdot\,+h)-f\) and take the shifts \(a_{i}=ih\), \(i=1,\dots,n\); the sum telescopes,
\begin{equation*} \frac1n\sum_{i=1}^{n}\bigl[f(t+ih+h)-f(t+ih)\bigr] =\frac1n\bigl[f\bigl(t+(n+1)h\bigr)-f(t+h)\bigr], \end{equation*}
so \(p(g)\le S(g,h,2h,\dots,nh)\le 2\sup|f|/n\) for every \(n\), i.e. \(p(g)\le0\); the same computation applied to \(-f\) gives \(p(-g)\le0\). Hence \(\Lambda(g)\le p(g)\le0\) and \(-\Lambda(g)=\Lambda(-g)\le p(-g)\le0\), so \(\Lambda(g)=0\). Transporting \(\Lambda\) back by \(f\mapsto\widetilde f\) yields the required functional on \(L\).
(ii) Use the upper Darboux integral \(\overline I\) as majorant, with \(\underline I(f)=-\overline I(-f)\). It is sublinear on \(L\) (refine to a common partition and use \(\sup(f+g)\le\sup f+\sup g\) on each subinterval), and a bounded \(f\) is Riemann integrable exactly when \(\overline I(f)=\underline I(f)\), the common value being the Lebesgue integral. Let \(R\subset L\) be the subspace of Riemann integrable functions and \(f_{0}:=\mathbf{1}_{\mathbb{Q}\cap[0,1)}\); as every subinterval contains rational and irrational points, \(\overline I(f_{0})=1\) and \(\underline I(f_{0})=0\), so \(f_{0}\notin R\) although \(f_{0}\) is Lebesgue integrable with integral \(0\). Hence every element of \(R+\mathbb{R}f_{0}\) is uniquely \(g+cf_{0}\) with \(g\in R\), and
\begin{equation*} \Lambda_{1}(g+cf_{0}):=\int_{0}^{1}g(x)\,dx+c \end{equation*}
is a well-defined linear function there. For \(g\in R\) and bounded \(h\), subadditivity applied to \(g+h\) and to \(h=(g+h)+(-g)\) gives \(\overline I(g+h)=\int_{0}^{1}g\,dx+\overline I(h)\); and \(\overline I(cf_{0})=c\overline I(f_{0})=c\) for \(c\ge0\), while \(\overline I(cf_{0})=-|c|\underline I(f_{0})=0>c\) for \(c<0\). So \(\Lambda_{1}\le\overline I\) on \(R+\mathbb{R}f_{0}\), and by Theorem 1.12.26 the function \(\Lambda_{1}\) extends to a linear \(\Lambda\) on \(L\) with \(\Lambda\le\overline I\). Then \(\Lambda(g)=\int_{0}^{1}g\,dx\) for every \(g\in R\), whereas \(\Lambda(f_{0})=1\ne0\), the Lebesgue integral of \(f_{0}\).
Exercises 2.12.102–2.12.108
(S. Banach) Prove that Lebesgue measure on \([0,1]\) can be extended to an additive but not countably additive nonnegative set function \(\nu\) that is defined on the class of all subsets of \([0,1]\) and has the following invariance property: \(\nu(E+h)=\nu(E)\) for all \(E\subset(0,1]\) and \(h\in(0,1]\), where in the formation of the sum \(E+h\) the numbers \(e+h>1\) are replaced by \(e+h-1\) (in this and the previous example one can deal with the circle and rotations in place of \((0,1]\) and translations).
Put \(\nu(E):=\Lambda(\widetilde I_{E})\), where \(\Lambda\) is the invariant extension of the integral constructed in Exercise 2.12.101 and \(\widetilde I_{E}\) is the \(1\)-periodic extension of \(I_{E}\). Identify \((0,1]\) with the circle group \(\mathbb{R}/\mathbb{Z}\), so that the addition of the statement is addition modulo \(1\); Exercise 2.12.101 gives a linear \(\Lambda\) on the space of bounded \(1\)-periodic functions with (a) \(\Lambda(f)=\int_{0}^{1}f\,dx\) whenever \(f\) is Lebesgue integrable on \((0,1]\), (b) \(\Lambda(f)\ge0\) for \(f\ge0\), (c) \(\Lambda\bigl(f(\,\cdot+h)\bigr)=\Lambda(f)\) for all \(h\in\mathbb{R}\). For arbitrary \(E\subset[0,1]\) set \(\nu(E):=\Lambda(\widetilde I_{E\cap(0,1]})\); this changes nothing on measurable sets, \(\{0\}\) being null. Then \(\nu\ge0\) by (b); \(\nu\) is additive, since \(\widetilde I_{E_{1}\cup E_{2}}=\widetilde I_{E_{1}}+\widetilde I_{E_{2}}\) for disjoint \(E_{1},E_{2}\) and \(\Lambda\) is linear; \(\nu=\lambda\) on Lebesgue measurable sets by (a), in particular \(\nu\bigl((0,1]\bigr)=1\); and \(\nu\) is invariant, because \(\widetilde I_{E+h}(t)=\widetilde I_{E}(t-h)\) for all \(t\) (as \(t\) lies in \(E+h\) modulo \(1\) exactly when \(t-h\) lies in \(E\) modulo \(1\)), so \(\nu(E+h)=\nu(E)\) by (c).
Countable additivity fails. Following Example 1.7.7, let \(V\subset(0,1]\) contain, by the axiom of choice, exactly one point of each class of the relation \(x-y\in\mathbb{Q}\). Then
\begin{equation*} (0,1]=\bigcup_{q\in\mathbb{Q}\cap(0,1]}(V+q) \end{equation*}
disjointly, with \(+\) modulo \(1\): given \(x\in(0,1]\) there is \(v\in V\) with \(x-v\in\mathbb{Q}\), so \(x=v+q\) modulo \(1\) for a unique \(q\in\mathbb{Q}\cap(0,1]\); and \(v+q\equiv v^{\prime}+q^{\prime}\) forces \(v-v^{\prime}\in\mathbb{Q}\), hence \(v=v^{\prime}\) and \(q=q^{\prime}\). Countable additivity together with the invariance just proved would give
\begin{equation*} 1=\nu\bigl((0,1]\bigr)=\sum_{q\in\mathbb{Q}\cap(0,1]}\nu(V+q) =\sum_{q\in\mathbb{Q}\cap(0,1]}\nu(V), \end{equation*}
which is impossible: the right-hand side is \(0\) if \(\nu(V)=0\) and \(+\infty\) if \(\nu(V)>0\).
Let \(f\in\mathcal{L}^1(\mathbb{R}^1)\) and \(a>0\).
(i) Show that the series \(\sum_{n=-\infty}^{+\infty}f(n+a^{-1}x)\) converges absolutely for a.e. \(x\).
(ii) Let \(g(x)=\sum_{n=-\infty}^{+\infty}f(n+a^{-1}x)\) if the series converges and \(g(x)=0\) otherwise. Show that
\begin{equation*} \int_0^a g(x)\,dx=a\int_{-\infty}^{+\infty}f(x)\,dx . \end{equation*}
(iii) Show that for a.e. \(x\) for each \(a>0\) one has \(\lim_{n\to\infty}n^{-a}f(nx)=0\).
(i) The \(a\)-periodic majorant \(G(x):=\sum_{n=-\infty}^{+\infty}|f(n+a^{-1}x)|\) is integrable on \([0,a]\). Indeed the substitution \(u=n+a^{-1}x\) gives \(\int_{0}^{a}|f(n+a^{-1}x)|\,dx=a\int_{n}^{n+1}|f(u)|\,du\), so the partial sums \(G_{M}=\sum_{|n|\le M}|f(n+a^{-1}\,\cdot\,)|\) satisfy \(\int_{0}^{a}G_{M}\,dx=a\int_{-M}^{M+1}|f|\,du\le a\|f\|_{1}\); they are nonnegative and increase to \(G\), so by the monotone convergence theorem (Theorem 2.8.2)
\begin{equation*} \int_{0}^{a}G(x)\,dx=a\|f\|_{1}<\infty . \end{equation*}
Hence \(G<\infty\) a.e. on \([0,a]\), i.e. the series converges absolutely a.e. there. Replacing \(x\) by \(x+a\) turns \(n+a^{-1}x\) into \((n+1)+a^{-1}x\) and so only shifts the summation index; thus \(G\) is \(a\)-periodic and \(\{G=+\infty\}\) is the countable union of the translates by \(ka\), \(k\in\mathbb{Z}\), of a null set, hence null. So the series converges absolutely for a.e. \(x\in\mathbb{R}^1\).
(ii) On the full-measure set \(D=\{G<\infty\}\) the symmetric partial sums \(S_{M}=\sum_{|n|\le M}f(n+a^{-1}\,\cdot\,)\) converge to \(g\), with \(|S_{M}|\le G\). Since \(G\) is integrable on \([0,a]\) by (i), the dominated convergence theorem (Theorem 2.8.1) and the substitution above give
\begin{equation*} \int_{0}^{a}g(x)\,dx=\lim_{M\to\infty}a\int_{-M}^{M+1}f(u)\,du =a\int_{-\infty}^{+\infty}f(u)\,du , \end{equation*}
the last step because \(f\) is integrable.
(iii) Fix \(b>0\). The substitution \(y=nx\) gives \(\int_{-\infty}^{+\infty}|f(nx)|\,dx=n^{-1}\|f\|_{1}\), whence
\begin{equation*} \sum_{n=1}^{\infty}n^{-b}\int_{-\infty}^{+\infty}|f(nx)|\,dx =\|f\|_{1}\sum_{n=1}^{\infty}n^{-b-1}<\infty , \end{equation*}
since \(b+1>1\). By the monotone convergence theorem (Theorem 2.8.2), applied to the partial sums on bounded intervals exhausting \(\mathbb{R}^1\), the function \(\Sigma_{b}(x):=\sum_{n}n^{-b}|f(nx)|\) is integrable, hence finite off a null set \(Z_{b}\); the terms of a convergent series tend to \(0\), so \(n^{-b}|f(nx)|\to0\) for \(x\notin Z_{b}\). To interchange the quantifiers put \(Z:=\bigcup_{k}Z_{1/k}\), a null set: for \(x\notin Z\) and arbitrary \(a>0\) choose \(k\) with \(1/k\le a\), so that
\begin{equation*} n^{-a}|f(nx)|\le n^{-1/k}|f(nx)|\xrightarrow[n\to\infty]{}0 . \end{equation*}
Thus for a.e. \(x\) the limit is \(0\) simultaneously for all \(a>0\).
Let \(f\in\mathcal{L}^1(\mathbb{R}^1)\). Prove the equality
\begin{equation*} \left|\int_{-\infty}^{+\infty}f(x)\,dx\right|=\inf\left\{\int_{-\infty}^{+\infty}\Bigl|\sum_{i=1}^{n}\alpha_i f(x+x_i)\Bigr|\,dx\right\}, \end{equation*}
where inf is taken over all numbers \(x_i\in\mathbb{R}^1\), \(n\in\mathbb{N}\) and \(\alpha_i\ge0\) with \(\alpha_1+\cdots+\alpha_n=1\).
Write \(Tu=\sum_{i=1}^{n}\alpha_iu(\,\cdot\,+x_i)\) for admissible data (\(\alpha_i\ge0\), \(\sum_i\alpha_i=1\), \(x_i\in\mathbb{R}^1\)), \(L(u)=\bigl|\int u\,dx\bigr|\) and \(R(u)=\inf_T\int|Tu|\,dx\); we prove \(L(f)=R(f)\).
\(L(f)\le R(f)\): by translation invariance \(\int Tf\,dx=\sum_i\alpha_i\int f\,dx=\int f\,dx\), so \(L(f)=\bigl|\int Tf\bigr|\le\int|Tf|\) for every admissible \(T\).
\(R\le L\) for a continuous \(\varphi\) vanishing off \([-R_{0},R_{0}]\): for \(T>0\) and \(n\in\mathbb{N}\) take \(\alpha_i=1/n\), \(x_i=iT/n\), and compare the Riemann sum \(S_{n,T}(x)=\frac1n\sum_{i}\varphi(x+iT/n)\) with \(\Psi_{T}(x)=\frac1T\int_x^{x+T}\varphi(u)\,du\). Since
\begin{equation*} S_{n,T}(x)-\Psi_{T}(x)=\frac1T\sum_{i=1}^{n}\int_{(i-1)T/n}^{iT/n} \Bigl[\varphi\Bigl(x+\frac{iT}{n}\Bigr)-\varphi(x+t)\Bigr]dt , \end{equation*}
uniform continuity gives \(|S_{n,T}-\Psi_{T}|\le\omega(T/n)\), where \(\omega\) is the modulus of continuity of \(\varphi\) and \(\omega(\delta)\to0\). Both functions vanish off \(J:=[-R_{0}-T,R_{0}]\), of measure \(2R_{0}+T\), so
\begin{equation*} \int_{-\infty}^{+\infty}|S_{n,T}|\,dx\le\int_J|\Psi_{T}|\,dx+(2R_{0}+T)\,\omega(T/n). \end{equation*}
Writing \(\Psi_{T}(x)=\frac1T\int\varphi-\frac1T\int_{\mathbb{R}\setminus[x,x+T]}\varphi\) and applying Tonelli’s theorem (Theorem 3.4.5) to the nonnegative remainder,
\begin{equation*} \frac1T\int_J\int_{\mathbb{R}\setminus[x,x+T]}|\varphi(u)|\,du\,dx =\frac1T\int|\varphi(u)|\,\lambda\{x\in J:u\notin[x,x+T]\}\,du \le\frac{2R_{0}}{T}\|\varphi\|_1 , \end{equation*}
because for \(|u|\le R_{0}\) the set \(\{x:u\in[x,x+T]\}=[u-T,u]\subset J\) has measure \(T\), leaving \(2R_{0}\), while for \(|u|>R_{0}\) the integrand vanishes. Hence
\begin{equation*} \int_{-\infty}^{+\infty}|S_{n,T}|\,dx\le\Bigl|\int\varphi\Bigr| +\frac{2R_{0}}{T}\Bigl(\Bigl|\int\varphi\Bigr|+\|\varphi\|_1\Bigr) +(2R_{0}+T)\,\omega(T/n). \end{equation*}
Given \(\varepsilon>0\), choose \(T\) so large that the middle term is \(<\varepsilon/2\), then \(n\) so large that the last is \(<\varepsilon/2\); thus \(R(\varphi)\le L(\varphi)+\varepsilon\) for all \(\varepsilon>0\).
General \(f\): given \(\varepsilon>0\), Corollary 4.2.2 gives a continuous \(\varphi\) with bounded support and \(\|f-\varphi\|_1<\varepsilon\). For admissible \(T\), translation invariance gives \(\int|Tf|\le\int|T\varphi|+\sum_i\alpha_i\int|f-\varphi|\,dx=\int|T\varphi|+\|f-\varphi\|_1\), so \(R(f)\le R(\varphi)+\varepsilon\); also \(|L(f)-L(\varphi)|\le\|f-\varphi\|_1<\varepsilon\). Hence \(R(f)\le L(\varphi)+\varepsilon\le L(f)+2\varepsilon\) for every \(\varepsilon>0\), i.e. \(R(f)\le L(f)\), which with the first step gives the asserted equality.
(Fréchet, Slutsky) Let \(\mu\) be a probability measure on a space \(X\) and let \(f\) be a \(\mu\)-measurable function. We call a number \(m\) a median of \(f\) if \(\mu(f<c)\le1/2\) for all \(c<m\) and \(\mu(f<c)\ge1/2\) for all \(c>m\).
(i) Prove that a median of \(f\) exists, but may not be unique.
(ii) Prove that a median is unique if \(f\) has a continuous strictly increasing distribution function \(\Phi_f\) and then \(m=\Phi_f^{-1}(1/2)\).
(iii) Suppose that measurable functions \(f_n\) converge to \(f\) in measure \(\mu\). Prove that the set of medians of the functions \(f_n\) is bounded and that if \(m_n\) is a median of \(f_n\) and \(m\) is a limit point of \(\{m_n\}\), then \(m\) is a median of \(f\).
(i) The set of medians is the nonempty interval \([m_-,m_+]\), where
\begin{equation*} m_-:=\sup\{c:\ \Phi_f( c)<1/2\},\qquad m_+:=\sup\{c:\ \Phi_f( c)\le1/2\} \end{equation*}
and \(\Phi_f( c)=\mu(f<c)\) is the distribution function (3.6.2): it is nondecreasing, left continuous (\(\{f<c\}=\bigcup_k\{f<c-1/k\}\) increases), with limits \(0\) at \(-\infty\) and \(1\) at \(+\infty\) by countable additivity, \(f\) being finite. Hence both sets are nonempty and bounded above, so \(m_-\le m_+\) are finite. If \(m\in[m_-,m_+]\) and \(c<m\), then \(c<m_+\) gives some \(c^{\prime}>c\) with \(\Phi_f(c^{\prime})\le1/2\), so \(\Phi_f( c)\le1/2\); if \(c>m\), then \(c>m_-\) puts \(c\) outside \(\{\Phi_f<1/2\}\subset(-\infty,m_-]\), so \(\Phi_f( c)\ge1/2\): thus \(m\) is a median. Conversely a median \(m\) satisfies \(m\ge m_-\) (else some \(c>m\) has \(\Phi_f( c)<1/2\)) and \(m\le m_+\) (else any \(c\in(m_+,m)\) has \(\Phi_f( c)\le1/2\), i.e. \(c\le m_+\)). Medians need not be unique: for \(f=I_{[1/2,1]}\) on \([0,1]\) with Lebesgue measure, \(\Phi_f=1/2\) on \((0,1]\), so \(m_-=0\), \(m_+=1\) and every point of \([0,1]\) is a median. We record for later
\begin{equation*} \Phi_f( c)<1/2\ \ (c<m_-),\qquad \Phi_f( c)>1/2\ \ (c>m_+), \end{equation*}
both immediate from the definitions of \(m_-\) and \(m_+\).
(ii) Unique, and \(m=\Phi_f^{-1}(1/2)\): by continuity and the limits \(0,1\) the intermediate value theorem gives \(m\) with \(\Phi_f(m)=1/2\), unique by strict monotonicity, which also yields \(\{\Phi_f<1/2\}=(-\infty,m)\) and \(\{\Phi_f\le1/2\}=(-\infty,m]\); so \(m_-=m_+=m\) and by (i) the median set is \(\{m\}\).
(iii) Convergence in measure means \(\mu(|f_n-f|\ge\varepsilon)\to0\) for every \(\varepsilon>0\). From \(\{f_n<c^{\prime}\}\subset\{f<c\}\cup\{|f_n-f|\ge c-c^{\prime}\}\) and its symmetric counterpart, for \(c^{\prime}<c\),
\begin{equation*} \Phi_{f_n}(c^{\prime})\le\Phi_f( c)+\mu(|f_n-f|\ge c-c^{\prime}),\qquad \Phi_{f_n}( c)\ge\Phi_f(c^{\prime})-\mu(|f_n-f|\ge c-c^{\prime}), \tag{\(\ast\)} \end{equation*}
while the median property of \(m_n\) reformulates as
\begin{equation*} \Phi_{f_n}( c)<1/2\ \Rightarrow\ m_n\ge c,\qquad \Phi_{f_n}( c)>1/2\ \Rightarrow\ m_n\le c \tag{\(\ast\ast\)} \end{equation*}
(else \(c>m_n\) would force \(\Phi_{f_n}( c)\ge1/2\), resp. \(c<m_n\) would force \(\Phi_{f_n}( c)\le1/2\)).
Boundedness. Let \([m_-,m_+]\) be the median interval of \(f\). By the consequences recorded in (i), \(\Phi_f(m_–1/2)<1/2\) and \(\Phi_f(m_++1/2)>1/2\), so \((\ast)\) with \((c^{\prime},c)=(m_–1,m_–1/2)\), respectively \((m_++1/2,m_++1)\), and \(\mu(|f_n-f|\ge1/2)\to0\) give
\begin{equation*} \limsup_n\Phi_{f_n}(m_–1)<\tfrac12,\qquad \liminf_n\Phi_{f_n}(m_++1)>\tfrac12 . \end{equation*}
By \((\ast\ast)\), \(m_–1\le m_n\le m_++1\) for all large \(n\); the finitely many remaining \(f_n\) have bounded median sets by (i).
Limit points. Let \(m_{n_k}\to m\). If some \(c<m\) had \(\Phi_f( c)>1/2\), left continuity would give \(c_1<c\) with \(\Phi_f(c_1)>1/2\), so by the second inequality in \((\ast)\) we would get \(\Phi_{f_n}( c)>1/2\), hence \(m_n\le c\), for all large \(n\), contradicting \(m_{n_k}\to m>c\). If some \(c>m\) had \(\Phi_f( c)<1/2\), pick \(m<c_2<c\); the first inequality in \((\ast)\) gives \(\Phi_{f_n}(c_2)<1/2\), hence \(m_n\ge c_2\), for all large \(n\), contradicting \(m_{n_k}\to m<c_2\). So \(\Phi_f( c)\le1/2\) for \(c<m\) and \(\Phi_f( c)\ge1/2\) for \(c>m\), i.e. \(m\) is a median of \(f\).
Let \(f\) be a nonnegative continuous function on \([0,+\infty)\) with the infinite integral over \([0,+\infty)\). Show that there exists \(a>0\) with \(\sum_{n=1}^{\infty}f(na)=\infty\).
Suppose \(\sum_{n=1}^{\infty}f(na)<\infty\) for every \(a>0\); the Baire category theorem plus one integration gives a contradiction. The sets
\begin{equation*} E_N:=\Bigl\{a\in[1,2]:\ \sum_{n=1}^{\infty}f(na)\le N\Bigr\} =\bigcap_{M=1}^{\infty}\{a\in[1,2]:\ g_M(a)\le N\}, \quad g_M:=\sum_{n\le M}f(n\,\cdot\,), \end{equation*}
are closed, each \(g_M\) being continuous and the terms nonnegative, and they cover the complete metric space \([1,2]\); so some \(E_N\) contains an interval \([\alpha,\beta]\subset[1,2]\) with \(\alpha<\beta\). The \(g_M\) are nonnegative, continuous, nondecreasing in \(M\), with \(\int_\alpha^\beta g_M\,da\le N(\beta-\alpha)\), so by the monotone convergence theorem (Theorem 2.8.2) and the substitution \(t=na\),
\begin{equation*} \sum_{n=1}^{\infty}\frac1n\int_{n\alpha}^{n\beta}f(t)\,dt =\int_\alpha^\beta\sum_{n=1}^{\infty}f(na)\,da\ \le\ N(\beta-\alpha) . \end{equation*}
A second application of Theorem 2.8.2, to the partial sums of \(t\mapsto n^{-1}I_{[n\alpha,n\beta]}(t)f(t)\), turns the left-hand side into \(\int_0^{+\infty}f(t)w(t)\,dt\) with \(w(t)=\sum_{n\in[t/\beta,\,t/\alpha]}n^{-1}\), every integer in that interval being \(\ge1\) since \(t/\beta>0\). Now \([t/\beta,t/\alpha]\) contains at least \(t(\beta-\alpha)/(\alpha\beta)-1\) integers, each \(\le t/\alpha\), i.e. with \(1/n\ge\alpha/t\); since \(w\ge0\) in any case,
\begin{equation*} w(t)\ \ge\ \Bigl(t\,\frac{\beta-\alpha}{\alpha\beta}-1\Bigr)\frac{\alpha}{t} =\frac{\beta-\alpha}{\beta}-\frac{\alpha}{t}\ \ge\ c:=\frac{\beta-\alpha}{2\beta}>0 \end{equation*}
for all \(t\ge T_0:=2\alpha\beta/(\beta-\alpha)\). Hence \(N(\beta-\alpha)\ge c\int_{T_0}^{+\infty}f(t)\,dt\). But \(f\) is continuous, so \(\int_0^{T_0}f<\infty\) and therefore \(\int_{T_0}^{+\infty}f=+\infty\) by hypothesis, a contradiction. Thus \(\sum_{n=1}^{\infty}f(na)=\infty\) for some \(a>0\).
(Buczolich, Mauldin) Prove that there exist an open set \(E\subset(0,+\infty)\) and intervals \(I_1\) and \(I_2\) in \([1/2,1)\) such that \(\sum_{n=1}^{\infty}I_E(nx)=\infty\) for all \(x\in I_1\) and \(\sum_{n=1}^{\infty}I_E(nx)<\infty\) for all \(x\in I_2\).
In the printed “for all \(x\)” form the assertion holds for one-point intervals and is false for nondegenerate ones; the intended statement is the almost everywhere theorem of Buczolich and Mauldin. Here \(\sum_{n}I_E(nx)\) counts the \(n\in\mathbb{N}\) with \(nx\in E\).
(i) Degenerate intervals. Take \(x_1=1/2\) and \(x_2=1/\sqrt2\) in \([1/2,1)\). The sets \(P=\{n/2:n\in\mathbb{N}\}\) and \(Q=\{m/\sqrt2:m\in\mathbb{N}\}\) are disjoint (an equality \(n/2=m/\sqrt2\) would make \(\sqrt2\) rational) and \(Q\) is closed and discrete, so \(\delta_n:=\mathrm{dist}(n/2,Q)>0\). Put
\begin{equation*} E:=\bigcup_{n=1}^{\infty}\Bigl(\frac n2-\varepsilon_n,\ \frac n2+\varepsilon_n\Bigr), \qquad \varepsilon_n:=\tfrac12\min(\delta_n,\ n/2). \end{equation*}
Then \(E\subset(0,+\infty)\) is open, \(nx_1\in E\) for every \(n\) while \(mx_2\notin E\) for every \(m\), so the two sums are \(\infty\) and \(0\).
(ii) Nondegenerate intervals: no open \(E\) works in the everywhere sense. Put \(H_n:=\{x>0:nx\in E\}=n^{-1}E\), open, and
\begin{equation*} D:=\Bigl\{x>0:\ \sum_{n=1}^{\infty}I_E(nx)=\infty\Bigr\}=\bigcap_{N=1}^{\infty}V_N, \qquad V_N:=\bigcup_{n\ge N}H_n , \end{equation*}
since \(x\in D\) means precisely \(x\in H_n\) for infinitely many \(n\); thus \(D\) is a \(G_\delta\) set. If \(E\subset(0,T]\) is bounded, then \(nx>T\) for all large \(n\), so \(D=\emptyset\) and no nondegenerate \(I_1\) exists. If \(E\) is unbounded, every \(V_N\) is dense in \((0,+\infty)\): for \(0<\alpha<\beta\) the intervals \((n\alpha,n\beta)\) and \(\bigl((n+1)\alpha,(n+1)\beta\bigr)\) overlap as soon as \(n(\beta-\alpha)>\alpha\), so \(\bigcup_{n\ge N}n(\alpha,\beta)\) contains a ray \((T_1,+\infty)\); choosing \(y\in E\) with \(y>T_1\) we get \(y=nx\) with \(n\ge N\) and \(x\in(\alpha,\beta)\), i.e. \(x\in H_n\subset V_N\). Hence \(D\) is a dense \(G_\delta\), its complement \(C=\{x>0:\sum_nI_E(nx)<\infty\}\) is of the first category, and by the Baire category theorem \(C\) contains no nondegenerate interval; so no nondegenerate \(I_2\) exists.
The theorem of Buczolich and Mauldin is the almost everywhere version: there exist an open \(E\subset(0,+\infty)\) and nondegenerate intervals \(I_1,I_2\subset[1/2,1)\) with \(\sum_nI_E(nx)=\infty\) for a.e. \(x\in I_1\) and \(\sum_nI_E(nx)<\infty\) for a.e. \(x\in I_2\), which answers the Haight–von Weizsäcker zero-one problem in the negative and is consistent with (ii), measure and category being independent notions of smallness. Its convergence half reduces to the Borel–Cantelli lemma, since \(\sum_n\lambda(n^{-1}E\cap I_2)=\sum_nn^{-1}\lambda(E\cap nI_2)<\infty\) forces \(\sum_nI_E(nx)<\infty\) for a.e. \(x\in I_2\); the difficulty is that for large \(n\) the dilates \(nI_1\) and \(nI_2\) overlap heavily, so the components of \(E\) must be placed along integers of carefully controlled multiplicative structure. (A fully rigorous proof of the remaining step is beyond the scope of this page; see the reference given in the book.)
Suppose we are given two measurable sets \(A\) and \(B\) in the circle of length \(1\) having linear Lebesgue measures \(\alpha\) and \(\beta\), respectively. Let \(B_\varphi\) be the image of the set \(B\) under the rotation in the angle \(\varphi\) counter-clockwise. Show that for some \(\varphi\) the set \(A\cap B_\varphi\) has measure at least \(\alpha\beta\).
Average over the rotation: the mean of \(s\mapsto\lambda(A\cap B_s)\) equals \(\alpha\beta\), so the value \(\alpha\beta\) is attained or exceeded. Identify the circle of length \(1\) with \([0,1)\) under addition modulo \(1\), so that the rotation by \(\varphi\) is the translation by the arc length \(s\) and \(\lambda\) is an invariant probability measure; since \(t\in B_s\) exactly when \(t-s\in B\) modulo \(1\),
\begin{equation*} \lambda(A\cap B_s)=\int_0^1 I_A(t)\,I_B(t-s)\,dt . \end{equation*}
The integrand \(F(t,s)=I_A(t)I_B(t-s)\) is Lebesgue measurable on the square: writing \(B=B_0\cup Z\) with \(B_0\) Borel and \(Z\) inside a Borel null set \(Z_1\), the set \(\{(t,s):t-s\in B_0\}\) is Borel, while \(\{(t,s):t-s\in Z_1\}\) is Borel of measure \(\int_0^1\lambda(Z_1+s)\,ds=0\) by Tonelli and invariance, so all of its subsets are measurable by completeness. Hence Tonelli’s theorem (Theorem 3.4.5) applies to \(F\ge0\), and the inner integral \(\int_0^1I_B(t-s)\,ds=\lambda(t-B)=\beta\) is constant in \(t\):
\begin{equation*} \int_0^1\lambda(A\cap B_s)\,ds =\int_0^1 I_A(t)\Bigl(\int_0^1 I_B(t-s)\,ds\Bigr)dt=\alpha\beta . \end{equation*}
Were \(\lambda(A\cap B_s)<\alpha\beta\) for every \(s\), the measurable function \(h(s):=\alpha\beta-\lambda(A\cap B_s)\) would be strictly positive with \(\int_0^1h\,ds=0\), forcing \(h=0\) a.e. Hence \(\lambda(A\cap B_\varphi)\ge\alpha\beta\) for some \(\varphi\).
Exercises 2.12.109–2.12.115
Let \(f\) be an integrable complex-valued function on a space \(X\) with a probability measure \(\mu\). Prove that
\begin{equation*} \int_X f\,d\mu = 0 \end{equation*}
precisely when
\begin{equation*} \int_X |1 + zf(x)|\,\mu(dx) \ge 1 \end{equation*}
for all complex numbers \(z\).
Necessity is the triangle inequality, sufficiency is dominated convergence as \(r\to0^+\).
Necessity. If \(\int_Xf\,d\mu=0\), then \(\int_X(1+zf)\,d\mu=\mu(X)+z\cdot0=1\) for every \(z\in\mathbb{C}\), so
\begin{equation*} \int_X|1+zf(x)|\,\mu(dx)\ \ge\ \Bigl|\int_X(1+zf)\,d\mu\Bigr|=1 . \end{equation*}
Sufficiency. Fix \(\theta\in\mathbb{R}^1\) and put \(z=re^{i\theta}\), \(r>0\). Since \(\mu(X)=1\), the hypothesis reads \(\int_Xg_r\,d\mu\ge0\) for \(g_r:=r^{-1}\bigl(|1+re^{i\theta}f|-1\bigr)\). With \(w=e^{i\theta}f(x)\),
\begin{equation*} |1+rw|=\bigl(1+2r\,\mathrm{Re}\,w+r^{2}|w|^{2}\bigr)^{1/2}, \end{equation*}
a function of \(r\) differentiable at \(r=0\) with derivative \(\mathrm{Re}\,w\), so \(g_r(x)\to\mathrm{Re}\bigl(e^{i\theta}f(x)\bigr)\) for every \(x\); and \(\bigl||1+rw|-1\bigr|\le r|w|\) gives \(|g_r|\le|f|\in\mathcal{L}^1(\mu)\). Hence the dominated convergence theorem (Theorem 2.8.1), applied along an arbitrary sequence \(r_j\downarrow0\), yields
\begin{equation*} \mathrm{Re}\Bigl(e^{i\theta}\int_Xf\,d\mu\Bigr) =\int_X\mathrm{Re}\bigl(e^{i\theta}f\bigr)\,d\mu =\lim_{j\to\infty}\int_Xg_{r_j}\,d\mu\ \ge\ 0 \end{equation*}
for every \(\theta\). If \(c:=\int_Xf\,d\mu\) were nonzero, write \(c=|c|e^{i\psi}\) and take \(\theta=\pi-\psi\): then \(\mathrm{Re}(e^{i\theta}c)=-|c|<0\), a contradiction. Hence \(\int_Xf\,d\mu=0\).
Let \(\{f_n\}\) be a sequence of integrable complex-valued functions on \([0,1]\) such that
\begin{equation*} \lim_{n\to\infty}\int_0^1 |\mathrm{Re}\,f_n(x)|\,dx = 1, \qquad \lim_{n\to\infty}\int_0^1 \bigl|1 - |f_n(x)|\bigr|\,dx = 0 . \end{equation*}
Show that
\begin{equation*} \lim_{n\to\infty}\int_0^1 |\mathrm{Im}\,f_n(x)|\,dx = 0 . \end{equation*}
Split the integral at \(|f_n|=2\). Write \(a_n=\mathrm{Re}\,f_n\), \(b_n=\mathrm{Im}\,f_n\), so that \(|a_n|,|b_n|\le|f_n|=(a_n^2+b_n^2)^{1/2}\) and all functions below are integrable, and put \(g_n:=|f_n|-|a_n|\ge0\). The second hypothesis gives \(\bigl|\int_0^1|f_n|\,dx-1\bigr|\le\int_0^1\bigl|1-|f_n|\bigr|\,dx\to0\), so \(\int_0^1|f_n|\,dx\to1\), and with the first hypothesis
\begin{equation*} \int_0^1 g_n\,dx=\int_0^1|f_n|\,dx-\int_0^1|a_n|\,dx\ \longrightarrow\ 0 , \end{equation*}
while pointwise \(b_n^2=\bigl(|f_n|-|a_n|\bigr)\bigl(|f_n|+|a_n|\bigr)\le 2g_n|f_n|\).
(i) On \(A_n:=\{|f_n|\le2\}\) this gives \(b_n^2\le4g_n\), so by Cauchy–Bunyakowsky on the probability space \(([0,1],\lambda)\),
\begin{equation*} \int_{A_n}|b_n|\,dx\ \le\ \Bigl(\int_{A_n}b_n^2\,dx\Bigr)^{1/2} \ \le\ \Bigl(4\int_0^1 g_n\,dx\Bigr)^{1/2}\ \longrightarrow\ 0 . \end{equation*}
(ii) On \(B_n:=\{|f_n|>2\}\) one has \(|f_n|-1>1\), hence \(|f_n|\le2(|f_n|-1)\le2\bigl|1-|f_n|\bigr|\), and \(|b_n|\le|f_n|\) gives
\begin{equation*} \int_{B_n}|b_n|\,dx\ \le\ 2\int_0^1\bigl|1-|f_n|\bigr|\,dx\ \longrightarrow\ 0 . \end{equation*}
Adding (i) and (ii), \(\int_0^1|\mathrm{Im}\,f_n(x)|\,dx\to0\).
(Kakutani) Let \(f\) and \(g\) be two nonnegative measurable functions on \([0,1]\) having the following property: if the integral of \(f\) over some measurable set \(E\) is finite, then the integral of \(g\) over \(E\) is finite as well. Prove that there exist a constant \(K\) and a nonnegative integrable function \(h\) such that \(g(x) \le Kf(x) + h(x)\).
Take \(K = n\) and \(h = g\,I_{A_n}\), where \(A_n := \{g > nf\}\) and \(n\) is any index with \(\int_{A_n} g\,d\lambda < \infty\); such an \(n\) exists, as the main step below shows.
First, \(\lambda(N) = 0\) for \(N := \{g = +\infty\}\cap\{f<+\infty\}\): if \(E \subset \{g=+\infty\}\cap\{f\le m\}\), then \(\int_E f \le m<\infty\), so \(\int_E g<\infty\) by hypothesis and \(\lambda(E)=0\); let \(m \to \infty\). On \(\{f=+\infty\}\) the desired inequality is trivial, so replacing \(f, g\) by \(f I_P\), \(g I_P\) with \(P := \{f<+\infty\}\setminus N\) (the hypothesis is inherited; Check!) we may assume \(f\) and \(g\) finite everywhere.
Selection lemma: if \(\int_E g = \infty\) and \(c>0\), there is a measurable \(E^{\prime}\subset E\) with \(\int_{E^{\prime}}g = c\). Indeed \(E_m := E\cap\{g\le m\}\uparrow E\) since \(g\) is finite, so \(\int_{E_m}g\uparrow\infty\) by monotone convergence; fix \(m\) with \(\int_{E_m}g>c\), note \(g I_{E_m}\) is integrable (its integral is at most \(m\)), so that \(\varphi(t) := \int_{E_m\cap[0,t]}g\,d\lambda\) is continuous by absolute continuity of the Lebesgue integral (Theorem 2.5.7) with \(\varphi(0)=0\) and \(\varphi(1)>c\); take \(E^{\prime} = E_m\cap[0,t_0]\) with \(\varphi(t_0)=c\) from the intermediate value theorem.
Suppose now \(\int_{A_n}g = \infty\) for every \(n\). Construct inductively disjoint \(E_k \subset A_{2^k}\) with \(\int_{E_k}g = 1\): with \(F := E_1\cup\dots\cup E_{k-1}\) one has \(\int_F g = k-1<\infty\), hence \(\int_{A_{2^k}\setminus F}g = \infty\), and the selection lemma applies to it. Since \(f<2^{-k}g\) on \(A_{2^k}\), countable additivity of the integral over the disjoint union \(E := \bigcup_k E_k\) gives
\begin{equation*} \int_E f\,d\lambda \le \sum_{k=1}^{\infty}2^{-k}\int_{E_k}g\,d\lambda = 1 < \infty , \qquad \int_E g\,d\lambda = \sum_{k=1}^{\infty} 1 = \infty , \end{equation*}
contradicting the hypothesis. So fix \(n\) with \(\int_{A_n}g<\infty\); then \(h = g I_{A_n}\ge 0\) is integrable, and \(g \le Kf+h\) holds on \(A_n\) (there \(h=g\)) and off \(A_n\) (there \(g\le nf\)), hence everywhere on \(P\). On the original \([0,1]\) the same \(K, h\) work outside the null set \(N\), and everywhere if one admits the value \(+\infty\) for \(h\) on \(N\).
(\(\circ\)) Suppose that increasing functions \(f_n\) converge in measure on the interval \([a,b]\) with Lebesgue measure. Show that they converge almost everywhere.
The whole sequence converges at every continuity point of the nondecreasing function \(F(x) := \sup\{f(y): y \in D,\ y<x\}\) defined below, and those points form a set of full measure.
By the Riesz theorem (Theorem 2.2.5(i)) some subsequence \(f_{n_k}\to f\) almost everywhere; let \(D\) be the set of \(x \in (a,b)\) with \(|f(x)|<\infty\) and \(f_{n_k}(x)\to f(x)\). Then \(\lambda([a,b]\setminus D)=0\), so \(D\) meets every nondegenerate subinterval, and \(f\) is nondecreasing on \(D\) (pass to the limit in \(f_{n_k}(x)\le f_{n_k}(y)\)). Hence the supremum defining \(F(x)\) is over a nonempty set and is bounded by \(f(z)\) for any \(z \in D\cap(x,b)\), so \(F\colon (a,b)\to\mathbb{R}^1\) is well defined, nondecreasing, and satisfies
\begin{equation*} \text{(i) } y \in D,\ y<x \Longrightarrow f(y)\le F(x); \qquad \text{(ii) } y \in D \Longrightarrow F(y)\le f(y) , \end{equation*}
the second because \(f\) is nondecreasing on \(D\). Being monotone, \(F\) has at most countably many discontinuities, so its continuity set \(C\subset(a,b)\) has full measure.
Fix \(x \in C\) and \(\varepsilon>0\), and choose \(a<u<x<v<b\) with \(F(v)-F(u)<\varepsilon\), so that \(F(x)-F(u)<\varepsilon\) and \(F(v)-F(x)<\varepsilon\). Put \(\alpha := \min(x-u, v-x)>0\) and take \(N\) with \(\lambda(\{|f_n-f|\ge\varepsilon\})<\alpha\) for \(n\ge N\). For such \(n\), each of \((u,x)\cap D\) and \((x,v)\cap D\) has measure \(\ge\alpha\), so each meets \(\{|f_n-f|<\varepsilon\}\); pick \(s\) and \(t\) in the respective intersections, so \(u<s<x<t<v\). Since \(f_n\) is nondecreasing, \(f_n(s)\le f_n(x)\le f_n(t)\), whence
\begin{equation*} f_n(x) \ge f_n(s) > f(s)-\varepsilon \ge F(s)-\varepsilon \ge F(u)-\varepsilon > F(x)-2\varepsilon \end{equation*}
by (ii) and monotonicity of \(F\), and symmetrically
\begin{equation*} f_n(x) \le f_n(t) < f(t)+\varepsilon \le F(v)+\varepsilon < F(x)+2\varepsilon \end{equation*}
by (i). Thus \(|f_n(x)-F(x)|\le 2\varepsilon\) for \(n \ge N\), i.e. \(f_n(x)\to F(x)\) for every \(x \in C\), and \(\lambda([a,b]\setminus C)=0\).
(Lovasz, Simonovits) Suppose we are given lower semicontinuous integrable functions \(u_1\) and \(u_2\) on \(\mathbb{R}^n\). Prove that there exist \(a,b \in \mathbb{R}^n\) and an affine function \(L \colon (0,1)\to(0,+\infty)\) such that
\begin{equation*} \int_0^1 u_i\bigl((1-t)a+tb\bigr)L(t)^{n-1}\,dt > 0, \qquad i=1,2 . \end{equation*}
The printed statement omits the hypothesis \(\int_{\mathbb{R}^n}u_i\,dx>0\), without which \(u_1=u_2=-\exp(-|x|^2)\) is a counterexample; we prove the localization lemma under it.
For \(n=1\) one has \(L^{n-1}\equiv1\), and after the substitution \(s=(1-t)a+tb\) the assertion reads \(\int_a^b u_i\,ds>0\), which holds for \(a=-R\), \(b=R\) with \(R\) large by dominated convergence. Let \(n\ge2\).
Reduction. Dominated convergence (dominant \(|u_i|\)) gives \(R\) with \(\int_B u_i\,dx>0\) for \(B := \{|x|\le R\}\). A lower semicontinuous function on the compact \(B\) is bounded below and is the increasing limit of continuous \(v_i^{(k)}\); applying monotone convergence (Theorem 2.8.2) to \(v_i^{(k)}-v_i^{(1)}\) and adding back the bounded \(v_i^{(1)}\) gives continuous \(v_i\le u_i\) on \(B\) with \(\int_B v_i\,dx>0\). Since \(u_i\ge v_i\) and all segments used below lie in \(B\), it suffices to treat \(v_1,v_2\). Choose \(\eta>0\) with \(\int_B(v_i-\eta)\,dx>0\) and put \(w_i := v_i-\eta\); it now suffices to find \(a\ne b\) in \(B\) and affine \(L>0\) on \((0,1)\) with \(\int_0^1 w_i((1-t)a+tb)L(t)^{n-1}dt\ge0\), since then the corresponding integral of \(v_i\) is at least \(\eta\int_0^1L^{n-1}dt>0\).
Why the weight is a power of an affine function. Let \(a\ne b\), let \(K\) be compact convex with \(0\in K\) and nonempty interior in the hyperplane through \(0\) orthogonal to \(b-a\), and let \(\ell(t)=(1-t)r_0+tr_1\) with \(r_0,r_1>0\). The cone \(T_\varepsilon := \{(1-t)a+tb+\varepsilon\ell(t)y:\ t\in(0,1),\ y\in K\}\) is convex, because \(\lambda_1K+\lambda_2K=(\lambda_1+\lambda_2)K\) for \(\lambda_i\ge0\) and \(\ell\) is affine. For continuous \(v\) on \(B\) and \(T_\varepsilon\subset B\), Fubini’s theorem and the scaling \(y=\varepsilon\ell(t)w\) of \((n-1)\)-dimensional Lebesgue measure give, the inner integrals converging uniformly in \(t\) by uniform continuity of \(v\),
\begin{equation*} \lim_{\varepsilon\to0^{+}}\varepsilon^{-(n-1)}\int_{T_\varepsilon}v\,dx = |b-a|\,\lambda_{n-1}(K)\int_0^1 v\bigl((1-t)a+tb\bigr)\ell(t)^{n-1}dt . \end{equation*}
Bisection. Call a compact convex \(Z\) good if \(\int_Z w_i\,dx>0\) for \(i=1,2\); then \(B\) is good, and a good set has interior. For good \(Z\) and any \((n-2)\)-dimensional affine subspace \(F\), let \(H^{+}_\theta\), \(\theta\in[0,\pi]\), be the pencil of closed half-spaces bounded by hyperplanes through \(F\), chosen continuously with \(H^{+}_\pi=H^{-}_0\). Then \(\psi(\theta):=\int_{Z\cap H^{+}_\theta}w_1\) is continuous by absolute continuity of the Lebesgue integral (Theorem 2.5.7) applied to \(|w_1|I_Z\), and \(\psi(0)+\psi(\pi)=\int_Zw_1>0\) since hyperplanes are null; so some \(\theta_0\) gives \(\psi(\theta_0)=\tfrac12\int_Zw_1>0\). As \(\int_{Z\cap H^{+}}w_2+\int_{Z\cap H^{-}}w_2>0\), one of the two halves is good.
Degeneration to a point or a segment. Let \(\mathcal{F}\) be the countable family of affine hulls of affinely independent \((n-1)\)-tuples of rational points. If \(S\) is convex with \(\dim\operatorname{aff}S\ge2\), pick \(p\) in its relative interior and a 2-plane \(P\) with \(p\in P\subset\operatorname{aff}S\), so \(p\) is interior to \(D:=S\cap P\) in \(P\); perturbing a basis of a complement \(W\) of \(\operatorname{dir}P\) to rational points gives \(F\in\mathcal{F}\) whose direction still complements \(\operatorname{dir}P\) and whose (unique, Cramer) intersection point with \(P\) stays interior to \(D\). Every hyperplane \(H\supset F\) then meets \(P\) in a line through that point, so a defining affine function of \(H\) changes sign on \(D\): \(S\) has points strictly on both sides of \(H\). Enumerate \(\mathcal{F}=\{F_j\}\), put \(Z_0:=B\) and let \(Z_j\subset Z_{j-1}\) be the good half given by the bisection step for \(F_j\), with boundary hyperplane \(H_j\). Then \(Z_\infty:=\bigcap_jZ_j\) is a nonempty compact convex set, and \(\dim\operatorname{aff}Z_\infty\le1\), since otherwise some \(H_j\) would have points of \(Z_\infty\subset Z_j\subset H^{+}_j\) strictly on both sides. Moreover \(\varepsilon_j:=\sup_{x\in Z_j}\operatorname{dist}(x,Z_\infty)\to0\) (Check!, using compactness of \(B\) and closedness of each \(Z_m\)).
(i) \(Z_\infty=\{p\}\). Since \(\int_{Z_j}w_1>0\), each \(Z_j\) carries \(y_j\) with \(w_1(y_j)>0\), and \(y_j\to p\), so \(v_1(p)\ge\eta\) and likewise \(v_2(p)\ge\eta\); by continuity \(v_i>\eta/2\) on \(U(p,r)\cap B\), so \(a:=p\), any \(b\in B\cap U(p,r)\) with \(b\ne p\), and \(L\equiv1\) settle the exercise.
(ii) \(Z_\infty=[a,b]\), \(a\ne b\). Put \(e:=(b-a)/|b-a|\), \(\alpha:=\langle a,e\rangle<\beta:=\langle b,e\rangle\), and let \(\varphi_j(s):=\lambda_{n-1}(K_j(s))\) be the section volume of \(Z_j\) at level \(s\), supported on a compact interval \([\alpha_j,\beta_j]\supset[\alpha,\beta]\) with \(\alpha_j\to\alpha\), \(\beta_j\to\beta\) (by \(\varepsilon_j\to0\)) and \(\int\varphi_j=\lambda_n(Z_j)>0\). Convexity gives \(\theta K_j(s_1)+(1-\theta)K_j(s_2)\subset K_j(\theta s_1+(1-\theta)s_2)\), so Brunn–Minkowski (Theorem 3.10.24; for \(n=2\) elementary) and homogeneity make \(\varphi_j^{1/(n-1)}\) concave. Comparing it at a point with the tent vanishing at \(\alpha_j,\beta_j\) bounds the probability densities \(g_j:=\varphi_j/\lambda_n(Z_j)\) by \(M:=n/(\beta-\alpha)\), so \(\rho_j:=g_j^{1/(n-1)}\) are concave, bounded by \(M^{1/(n-1)}\), and – by monotonicity of difference quotients of concave functions – uniformly Lipschitz on each \([\alpha^{\prime},\beta^{\prime}]\subset(\alpha,\beta)\); Arzela–Ascoli and a diagonal argument give \(\rho_j\to\rho\) locally uniformly on \((\alpha,\beta)\) with \(\rho\ge0\) concave. Let \(\pi(s)\) be the point of \([a,b]\) at level \(\operatorname{med}(\alpha,s,\beta)\) and \(\widetilde w_i:=w_i\circ\pi\). Every \(x\in Z_j\) satisfies \(|x-\pi(\langle x,e\rangle)|\le2\varepsilon_j\), so with a common modulus of continuity \(\omega\) of \(w_1,w_2\) on \(B\), Fubini’s theorem and \(\int_{Z_j}w_i>0\) give \(\int_{\alpha_j}^{\beta_j}\widetilde w_ig_j\,ds>-\omega(2\varepsilon_j)\to0\). Splitting off the edge pieces \([\alpha_j,\alpha+1/k]\) and \([\beta-1/k,\beta_j]\) (bounded by \(4CM/k\), \(C:=\max_B(|w_1|+|w_2|)\)) and letting \(j\to\infty\), then \(k\to\infty\), yields \(\int_\alpha^\beta\widetilde w_i\rho^{n-1}ds\ge0\) and, applied to \(\int g_j=1\), also \(\int_\alpha^\beta\rho^{n-1}ds=1\), so \(\rho\not\equiv0\). Reparametrizing by \(s=\alpha+t(\beta-\alpha)\), with \(x(t):=(1-t)a+tb\), \(\varrho(t):=\rho(\alpha+t(\beta-\alpha))\) and \(h_i(t):=v_i(x(t))-\eta/2\),
\begin{equation*} \int_0^1 h_i(t)\,\varrho(t)^{n-1}\,dt \ \ge\ \frac{\eta}{2}\int_0^1\varrho(t)^{n-1}\,dt \ >\ 0 . \end{equation*}
What remains. It suffices to prove Lemma C: for \(h_1,h_2\) continuous on \([0,1]\) and \(\varrho\ge0\) concave, \(\varrho\not\equiv0\), with \(\int_0^1h_i\varrho^{n-1}dt>0\), there are \(0\le\alpha^{\prime}<\beta^{\prime}\le1\) and an affine \(\ell>0\) on \((\alpha^{\prime},\beta^{\prime})\) with \(\int_{\alpha^{\prime}}^{\beta^{\prime}}h_i\ell^{n-1}dt\ge0\); for then \(a^{\prime}:=x(\alpha^{\prime})\), \(b^{\prime}:=x(\beta^{\prime})\) and \(L(t):=\ell(\alpha^{\prime}+t(\beta^{\prime}-\alpha^{\prime}))\) answer the exercise, since \(x\) is affine and the substitution turns the displayed inequality into \(\int_0^1v_i((1-t)a^{\prime}+tb^{\prime})L^{n-1}dt\ge(\eta/2)\int_0^1L^{n-1}dt>0\). One may assume \(\varrho\) piecewise affine: its interpolant \(\varrho_N\) at the nodes \(k/N\) is concave, \(\le\varrho\), and converges locally uniformly, so dominated convergence keeps both integrals positive for large \(N\). Removing the finitely many kinks is the concluding argument of Lovasz–Simonovits [623, Section 2] and Kannan–Lovasz–Simonovits [489, Section 2]: one maximizes \(\mu\mapsto\int h_2\,d\mu-\varepsilon\iint|x-y|\,d\mu\,d\mu\) over the (weakly sequentially compact, by the estimates above) set of admissible \(\mu\) with \(\int h_1\,d\mu\ge0\), the second term being strictly convex along mixtures, and shows that any admissible \(\mu\) with nonaffine weight is a nontrivial mixture of two admissible measures. I have not been able to reproduce that step in full rigor. (A fully rigorous proof of the remaining step is beyond the scope of this page; see the reference given in the book.)
Suppose that a sequence of convex functions \(f_n\) on a ball \(U \subset \mathbb{R}^d\) is uniformly bounded. Prove that it contains a subsequence convergent in \(L^p(U)\) for all \(p \in [1,\infty)\).
Extract by Arzela–Ascoli a subsequence converging pointwise on the open ball; dominated convergence then gives convergence in \(L^p(U)\) for all \(p<\infty\) at once.
Let \(U=\{|x|<R\}\) (the sphere is Lebesgue null, so this changes no \(L^p(U)\)) and \(|f_n|\le M\) on \(U\). A convex \(f\) with \(|f|\le M\) is Lipschitz with constant \(2M/\delta\) on \(V_\delta:=\{|x|<R-\delta\}\): for \(x\ne y\) in \(V_\delta\), the point \(z:=y+\delta(y-x)/|y-x|\) satisfies \(|z|\le|y|+\delta<R\), and \(y=(1-\lambda)x+\lambda z\) with \(\lambda=|y-x|/(|y-x|+\delta)\le|y-x|/\delta\), so convexity gives
\begin{equation*} f(y)-f(x)\ \le\ \lambda\bigl(f(z)-f(x)\bigr) \ \le\ 2M\lambda\ \le\ \frac{2M}{\delta}|y-x| , \end{equation*}
and symmetrically with \(x,y\) interchanged. Hence on each compact \(W_k:=\{|x|\le R-2/k\}\subset V_{1/k}\) the family \(\{f_n\}\) is uniformly bounded by \(M\) and equicontinuous, so the Arzela–Ascoli theorem and a diagonal extraction over \(k\) produce a subsequence \(f_{n_j}\) converging at every point of \(\bigcup_kW_k=\{|x|<R\}\), say to \(f\), with \(|f|\le M\) (set \(f:=0\) on the sphere).
Fix \(p\in[1,\infty)\). Then \(|f_{n_j}-f|^p\to0\) almost everywhere on \(U\) and \(|f_{n_j}-f|^p\le(2M)^p\), which is integrable since \(U\) has finite measure, so the dominated convergence theorem (Theorem 2.8.1) gives
\begin{equation*} \lim_{j\to\infty}\int_U|f_{n_j}(x)-f(x)|^p\,dx = 0 . \end{equation*}
The subsequence was chosen independently of \(p\), so it works for all \(p\in[1,\infty)\) simultaneously.
Let \(\mu\) be a probability measure on a measurable space \((X,\mathcal{A})\), let \(1<p<\infty\), and let \(f_n \in \mathcal{L}^p(\mu)\) be nonnegative functions such that
\begin{equation*} \|f_n\|_{L^p(\mu)} \le C\|f_n\|_{L^1(\mu)} \end{equation*}
with some constant \(C\) (or, more generally, \(\bigl\|\sum_{n=1}^{N}f_n\bigr\|_{L^p(\mu)} \le C\sum_{n=1}^{N}\|f_n\|_{L^1(\mu)}\)). Prove that the series \(\sum_{n=1}^{\infty}f_n\) converges \(\mu\)-a.e. if and only if
\begin{equation*} \sum_{n=1}^{\infty}\int_X f_n\,d\mu<\infty . \end{equation*}
Monotone convergence gives one implication and the Paley–Zygmund bound of Proposition 2.11.7 the other, applied to \(S_N:=\sum_{n\le N}f_n\) and \(A_N:=\sum_{n\le N}\int_Xf_n\,d\mu=\|S_N\|_{L^1(\mu)}\) (the last equality since \(f_n\ge0\)); write \(q:=p/(p-1)\).
The first hypothesis implies the second by the triangle inequality in \(L^p(\mu)\), so we work under estimate (A): \(\|S_N\|_{L^p(\mu)}\le C A_N\) for all \(N\). Since \(f_n\ge0\), the partial sums increase, so \(\sum_nf_n(x)\) converges exactly when \(\sup_NS_N(x)<\infty\).
(i) If \(\sup_NA_N<\infty\), then \(S_N\uparrow\sum_nf_n\) with \(\sup_N\int_XS_N\,d\mu<\infty\), so the monotone convergence theorem (Theorem 2.8.2) makes \(\sum_nf_n\) integrable, hence finite \(\mu\)-a.e.: the series converges \(\mu\)-a.e. (Neither (A) nor \(p>1\) is used here.)
(ii) If \(A_N\uparrow\infty\), discard the finitely many \(N\) with \(A_N=0\) and apply Proposition 2.11.7 to \(S_N\in\mathcal{L}^p(\mu)\) with \(\lambda=1/2\), then (A):
\begin{equation*} \mu(E_N)\ \ge\ \bigl(1-\tfrac12\bigr)^{q} \frac{\|S_N\|^q_{L^1(\mu)}}{\|S_N\|^q_{L^p(\mu)}} \ \ge\ 2^{-q}\frac{A_N^q}{(C A_N)^q} = 2^{-q}C^{-q} =: c>0 , \end{equation*}
where \(E_N:=\{S_N\ge\tfrac12A_N\}\). The sets \(G_M:=\bigcup_{N\ge M}E_N\) decrease to \(E:=\limsup_NE_N\) and satisfy \(\mu(G_M)\ge\mu(E_M)\ge c\), so continuity from above of the finite measure \(\mu\) gives \(\mu(E)\ge c>0\). For \(x\in E\) one has \(S_N(x)\ge\frac12A_N\) for infinitely many \(N\) while \(A_N\to\infty\), so \(\sum_nf_n(x)=\lim_NS_N(x)=+\infty\): the series diverges on a set of positive measure.
Exercises 2.12.116–2.12.117
(Kadec, Pelczynski [476]) Let \(\mu\) be a probability measure on a measurable space \((X,\mathcal{A})\) and let \(p\ge 1\), \(\varepsilon>0\). Set
\begin{equation*} M^p_\varepsilon := \Bigl\{ f\in\mathcal{L}^p(\mu):\ \mu\bigl(\{x:\ |f(x)|\ge \varepsilon\|f\|_{L^p(\mu)}\}\bigr)\ge\varepsilon \Bigr\}. \end{equation*}
(i) Show that \(\mathcal{L}^p(\mu)=\bigcup_{\varepsilon>0}M^p_\varepsilon\).
(ii) Suppose that \(f\in\mathcal{L}^p(\mu)\), where \(p>1\), and that \(\|f\|_{L^p(\mu)}\le C\|f\|_{L^r(\mu)}\), where \(r\in(1,p)\). Show that \(f\in M^p_\varepsilon\) with \(\varepsilon=C^{rp/(p-1)}2^{p/(1-p)}\).
(i) Every \(f\in\mathcal{L}^p(\mu)\) lies in some \(M^p_\varepsilon\); (ii) the printed exponent \(p-1\) should be \(p-r\), the correct conclusion being \(f\in M^p_{\varepsilon_0}\) with \(\varepsilon_0=(2C^{r})^{-p/(p-r)}\) for \(f\) not \(\mu\)-a.e. zero (for \(f=0\) a.e. one has \(f\in M^p_\varepsilon\) for every \(\varepsilon\in(0,1]\), while \(M^p_\varepsilon=\varnothing\) for \(\varepsilon>1\)).
Throughout \(a:=\|f\|_{L^p(\mu)}\) and \(E_\varepsilon:=\{|f|\ge\varepsilon a\}\), so that \(f\in M^p_\varepsilon\) means \(\mu(E_\varepsilon)\ge\varepsilon\).
(i) The inclusion \(\bigcup_\varepsilon M^p_\varepsilon\subset\mathcal{L}^p(\mu)\) is the definition. Conversely, if \(a=0\) then \(E_1=X\) and \(f\in M^p_1\). If \(a>0\) and \(f\notin M^p_\varepsilon\) for every \(\varepsilon>0\), then the sets \(B_n:=\{|f|<a/n\}\) satisfy \(\mu(B_n)=1-\mu(E_{1/n})>1-1/n\) and decrease to \(\{f=0\}\), so continuity from above of the probability measure \(\mu\) gives \(\mu(f=0)=1\), i.e. \(a=0\) – a contradiction.
(ii) Let \(f\) be not \(\mu\)-a.e. zero, so \(a>0\). Holder’s inequality on the probability space with exponents \(p/r\) and \(p/(p-r)\), applied to \(|f|^r\cdot1\), gives \(\|f\|_{L^r(\mu)}\le a\); with the hypothesis \(a\le C\|f\|_{L^r(\mu)}\) this forces \(C\ge1\) and hence \(\varepsilon_0\le2^{-p/(p-r)}<1\). Assume \(\mu(E)<\varepsilon_0\) for \(E:=E_{\varepsilon_0}\). The same Holder inequality on \(E\), the pointwise bound \(|f|<\varepsilon_0a\) off \(E\), and \(\varepsilon_0^{r}\le\varepsilon_0^{(p-r)/p}\) (since \(0<\varepsilon_0<1\) and \(r>1>(p-r)/p\)) yield
\begin{equation*} \|f\|^r_{L^r(\mu)} \le a^r\mu(E)^{(p-r)/p} + (\varepsilon_0a)^r < 2a^r\varepsilon_0^{(p-r)/p} . \end{equation*}
Combining with \(a^r\le C^r\|f\|^r_{L^r(\mu)}\) and dividing by \(a^r>0\) gives \(1<2C^r\varepsilon_0^{(p-r)/p}\), i.e. \(\varepsilon_0>(2C^r)^{-p/(p-r)}=\varepsilon_0\), a contradiction. Hence \(\mu(E_{\varepsilon_0})\ge\varepsilon_0\), i.e. \(f\in M^p_{\varepsilon_0}\).
That the printed value fails is seen on \(X=[0,1]\) with \(p=2\), \(r=3/2\), \(f=8\cdot\mathbf{1}_{[0,1/64]}\): here \(\|f\|_{L^{3/2}}=\tfrac12\|f\|_{L^2}\), the printed value is \(C^{rp/(p-1)}2^{p/(1-p)}=1/32\) in the hint’s normalization, yet \(\mu(\{|f|\ge\tfrac1{32}\|f\|_{L^2}\})=1/64<1/32\), whereas the corrected \(\varepsilon_0=1/1024\le1/64\) is admissible.
(Sarason [845]) Let \((X,\mathcal{A},\mu)\) be a probability space and let \(f>0\) be a \(\mu\)-measurable function such that
\begin{equation*} \int_X f\,d\mu\int_X f^{-1}\,d\mu\le 1+c^3 \end{equation*}
for some \(c\in(0,1/2)\). Let \(J\) be the integral of \(f\) and let \(I\) be the integral of \(\ln f\). Show that
\begin{equation*} \int_X|\ln f-\ln J|\,d\mu\le 8c,\qquad \int_X|\ln f-I|\,d\mu\le 16c . \end{equation*}
Everything follows from the fact that \(f\) is within \(c\) of \(1\) off a set of measure at most \(2c\): with \(A:=\{(1+c)^{-1}<f<1+c\}\) one has \(\mu(X\setminus A)\le2c\).
Both \(\int_Xf\,d\mu\) and \(\int_Xf^{-1}\,d\mu\) are positive, and finite since otherwise their product would be \(+\infty\); so \(0<J<\infty\) and we may replace \(f\) by \(f/J\), which leaves both asserted integrals unchanged (\(\ln(f/J)=\ln f-\ln J\), and \(I\) becomes \(I-\ln J\)) and gives
\begin{equation*} \int_Xf\,d\mu=1,\qquad \int_Xf^{-1}\,d\mu\le 1+c^3 . \end{equation*}
The function \(\varphi(t)=t+t^{-1}\) decreases on \((0,1]\) and increases on \([1,\infty)\), so \(\varphi\ge2\) everywhere and \(\varphi(t)\ge1+c+(1+c)^{-1}\) outside \(((1+c)^{-1},1+c)\). Integrating \(\varphi(f)=f+f^{-1}\) (integrable by the above) and using \(1+c+(1+c)^{-1}-2=c^2/(1+c)\),
\begin{equation*} 2+c^3\ \ge\ 2+\frac{c^2}{1+c}\,\mu(X\setminus A), \qquad\text{i.e.}\qquad \mu(X\setminus A)\le c(1+c)\le 2c , \end{equation*}
the last step by \(c<1/2\). On \(A\) both \(f\) and \(f^{-1}\) exceed \((1+c)^{-1}\), so \(\int_Af\,d\mu\ge(1-2c)/(1+c)\) and likewise for \(f^{-1}\), whence
\begin{equation*} \int_{X\setminus A}f\,d\mu\le 1-\frac{1-2c}{1+c}\le 3c , \qquad \int_{X\setminus A}f^{-1}\,d\mu\le c^3+\frac{3c}{1+c}\le 4c . \end{equation*}
Since \(|\ln t|\le t+t^{-1}\) for every \(t>0\) (Check!), \(\ln f\) is \(\mu\)-integrable, so \(I\) is finite; and \(|\ln f|<\ln(1+c)\le c\) on \(A\). Splitting the integral,
\begin{equation*} \int_X|\ln f|\,d\mu \le c\,\mu(A)+\int_{X\setminus A}(f+f^{-1})\,d\mu \le c+3c+4c=8c , \end{equation*}
which, undoing the normalization, is \(\int_X|\ln f-\ln J|\,d\mu\le8c\). Consequently \(|I|\le\int_X|\ln f|\,d\mu\le8c\) and therefore
\begin{equation*} \int_X|\ln f-I|\,d\mu \le \int_X|\ln f|\,d\mu+|I| \le 8c+8c = 16c . \end{equation*}
Operations on Measures and Functions
Exercises 3.10.29–3.10.35
(\(\circ\)) Let \(\mu\) be a signed Borel measure on \(\mathbb{R}^n\) that is bounded on bounded sets. Prove that if every continuous function with bounded support has the zero integral with respect to the measure \(\mu\), then \(\mu = 0\).
Bounded open sets get measure \(0\), and Dynkin’s theorem spreads this to all Borel sets.
The hypothesis is read as: \(\mu\) is countably additive and its restriction \(\mu_R\) to each ball \(B_R=\{|x|<R\}\) is a bounded signed measure (by Corollary 3.1.3 boundedness on a \(\sigma\)-algebra is automatic), so every continuous \(f\) supported in \(B_R\) is \(\mu\)-integrable, with \(|\int f\,d\mu|\le\sup_x|f(x)|\cdot|\mu_R|(\mathbb{R}^n)<\infty\).
Let \(U\ne\emptyset\) be bounded and open; then \(\partial U\ne\emptyset\), since \(\mathbb{R}^n\) is connected and \(U\ne\mathbb{R}^n\). The sets \(K_j:=\{x\in U:\operatorname{dist}(x,\partial U)\ge j^{-1}\}\) are compact and increase to \(U\) (Check!), and
\begin{equation*} f_j(x) := \frac{\operatorname{dist}(x,\mathbb{R}^n\setminus U)} {\operatorname{dist}(x,\mathbb{R}^n\setminus U)+\operatorname{dist}(x,K_j)} \end{equation*}
(with \(f_j:=0\) when \(K_j=\emptyset\)) has nonvanishing denominator, the two closed sets \(\mathbb{R}^n\setminus U\) and \(K_j\) being disjoint; so \(f_j\) is continuous, \(0\le f_j\le1\), \(f_j=1\) on \(K_j\) and \(f_j=0\) off \(U\), whence \(f_j\) has bounded support and \(\int f_j\,d\mu=0\). Since \(f_j\to I_U\) pointwise with \(|f_j-I_U|\le2I_{\overline U}\), choosing \(R\) with \(\overline U\subset B_R\) and applying the dominated convergence theorem (Theorem 2.8.1) to the finite measure \(|\mu_R|\) gives
\begin{equation*} |\mu(U)| = \Bigl|\int(f_j-I_U)\,d\mu\Bigr| \le \int|f_j-I_U|\,d|\mu_R| \longrightarrow 0 , \end{equation*}
i.e. \(\mu(U)=0\).
Fix \(R\). The open subsets of \(B_R\) form a \(\pi\)-system containing \(B_R\) and generating \(\mathcal{B}(B_R)\), and \(\mu\) vanishes on it; the class \(\{A\in\mathcal{B}(B_R):\mu(A)=0\}\) is a \(\lambda\)-system, because \(\mu_R\) is finite (so differences subtract) and countably additive (so increasing limits pass). Dynkin’s theorem gives \(\mu=0\) on \(\mathcal{B}(B_R)\), hence on every bounded Borel set. Finally an arbitrary Borel \(A\) is the disjoint union of \(A\cap B_1\) and the sets \(A\cap(B_j\setminus B_{j-1})\), \(j\ge2\), so \(\mu(A)=0\) by countable additivity.
Let \(\mathcal{A}\) be the algebra of all finite subsets of \(\mathbb{R}\) and their complements. If \(A\) is finite, then we set
\begin{equation*} \mu(A) := \operatorname{Card}\big(A \cap (-\infty, 0]\big) - \operatorname{Card}\big(A \cap (0, +\infty)\big), \end{equation*}
where \(\operatorname{Card}(M)\) is the cardinality of \(M\), and if the complement of \(A\) is finite, then we set \(\mu(A) := -\mu(\mathbb{R}^1 \setminus A)\). Show that \(\mu\) is a countably additive signed measure on the algebra \(\mathcal{A}\), but \(\mu\) has no countably additive extensions to the \(\sigma\)-algebra \(\sigma(\mathcal{A})\) (even if we admit measures with values in \([-\infty, +\infty)\) or \((-\infty, +\infty]\)).
Countable additivity is trivial because only finitely many pairwise disjoint members of \(\mathcal{A}\) can be nonempty, while no extension exists because it would have to take both the value \(-\infty\) on \(\{1,2,\dots\}\) and the value \(+\infty\) on \(\{-1,-2,\dots\}\).
Write \(c(A) := \operatorname{Card}(A\cap(-\infty,0])-\operatorname{Card}(A\cap(0,+\infty))\) for finite \(A\). The family \(\mathcal{A}\) is an algebra and \(\mu\) is unambiguous, since no subset of \(\mathbb{R}\) is at once finite and cofinite; \(\mu\) is real-valued with \(\mu(\emptyset)=\mu(\mathbb{R})=0\).
Finite additivity. Two cofinite sets always meet, the complement of their intersection being finite. So for disjoint \(A,B\in\mathcal{A}\) either both are finite, and \(\mu(A\cup B)=c(A)+c(B)\) by additivity of cardinality, or (say) \(A\) is finite and \(B\) cofinite; then \(A\subset F:=\mathbb{R}\setminus B\), the set \(F\) is finite, \(\mathbb{R}\setminus(A\cup B)=F\setminus A\), and \(c(F)=c(A)+c(F\setminus A)\), so
\begin{equation*} \mu(A\cup B) = -c(F\setminus A) = -c(F)+c(A) = \mu(B)+\mu(A) . \end{equation*}
Countable additivity. Let \(A_n\in\mathcal{A}\) be pairwise disjoint with \(A:=\bigcup_nA_n\in\mathcal{A}\); at most one \(A_n\) is cofinite. (i) All \(A_n\) finite: then \(A\) is at most countable, hence not cofinite (cofinite subsets of \(\mathbb{R}\) are uncountable), hence finite, so only finitely many \(A_n\) are nonempty. (ii) Exactly one \(A_{n_0}\) cofinite: the remaining \(A_n\) are disjoint subsets of the finite set \(\mathbb{R}\setminus A_{n_0}\), so again only finitely many are nonempty. Either way \(\sum_n\mu(A_n)\) has finitely many nonzero terms and equals \(\mu(A)\) by finite additivity.
No extension. \(\sigma(\mathcal{A})\) is the \(\sigma\)-algebra of all at most countable sets and their complements: this family is a \(\sigma\)-algebra containing \(\mathcal{A}\), and every at most countable set is a countable union of singletons. Hence \(C:=\{1,2,\dots\}\) and \(D:=\{-1,-2,\dots\}\) lie in \(\sigma(\mathcal{A})\). A countably additive extension \(\lambda\) satisfies \(\lambda(\{n\})=-1\) and \(\lambda(\{-n\})=1\) for \(n\ge1\), so applying countable additivity to \(C=\bigcup_n\{n\}\) and \(D=\bigcup_n\{-n\}\) forces \(\lambda( C)=-\infty\) and \(\lambda(D)=+\infty\) – impossible for values in \(\mathbb{R}\), in \([-\infty,+\infty)\), or in \((-\infty,+\infty]\).
(i) Let \(\mu\) be a finite nonnegative measure on a \(\sigma\)-algebra \(\mathcal{A}\) in a space \(X\) and let \(\nu\) be a countably additive measure on \(\mathcal{A}\) with values in \([0, +\infty]\) such that \(\nu \ll \mu\). Show that there exists a set \(S \in \mathcal{A}\) such that the measure \(\nu|_S\) assumes only the values \(0\) and \(+\infty\) and the measure \(\nu|_{X \setminus S}\) is \(\sigma\)-finite.
(ii) Deduce from (i) that, given \(\sigma\)-finite measures \(\mu \ge 0\) and \(\nu \ge 0\) with \(\nu \ll \mu\) on a \(\sigma\)-algebra \(\mathcal{A}\), for every sub-\(\sigma\)-algebra \(\mathcal{B} \subset \mathcal{A}\), there is a \(\mathcal{B}\)-measurable function \(\xi\) such that \(\nu|_B = \xi \cdot \mu|_B\) for every \(B \in \mathcal{B}\) with \(\mu(B) + \nu(B) < \infty\). Show that this is not true for all \(B \in \mathcal{B}\) in the case where \(\mu\) is Lebesgue measure on \(\mathbb{R}^1\), \(\nu = \varrho \cdot \mu\) is a probability measure, and \(\mathcal{B}\) is generated by all singletons.
(i) Take for \(S\) a set of maximal \(\mu\)-measure in
\begin{equation*} \mathcal{S} := \{A\in\mathcal{A}:\ \nu(B)\in\{0,+\infty\} \ \text{for every } B\in\mathcal{A},\ B\subset A\} . \end{equation*}
This class is hereditary, and closed under countable unions: if \(A_n\in\mathcal{S}\), \(B\subset\bigcup_nA_n\) and \(\nu(B)<\infty\), then \(\nu(B\cap A_n)\le\nu(B)<\infty\) forces \(\nu(B\cap A_n)=0\), so \(\nu(B)\le\sum_n\nu(B\cap A_n)=0\). Put \(c:=\sup\{\mu(A):A\in\mathcal{S}\}\le\mu(X)<\infty\), choose \(A_n\in\mathcal{S}\) with \(\mu(A_n)\to c\), and set \(S:=\bigcup_nA_n\in\mathcal{S}\), so that \(\mu(S)=c\). Since \(A\cap S\subset S\) for every \(A\in\mathcal{A}\), the measure \(\nu|_S\) takes only the values \(0\) and \(+\infty\). Two consequences of \(\nu\ll\mu\): a set \(A\in\mathcal{S}\) with \(\nu(A)=+\infty\) has \(\mu(A)>0\); hence every \(A\in\mathcal{S}\) with \(A\subset Y:=X\setminus S\) has \(\nu(A)=0\), since otherwise \(S\cup A\in\mathcal{S}\) would have \(\mu(S\cup A)=c+\mu(A)>c\).
For the \(\sigma\)-finiteness of \(\nu|_Y\) let \(\mathcal{G}:=\{B\in\mathcal{A}: B\subset Y,\ \nu|_B\ \sigma\text{-finite}\}\), again closed under countable unions, put \(t:=\sup\{\mu(B):B\in\mathcal{G}\}<\infty\) and choose as above \(T\in\mathcal{G}\) with \(\mu(T)=t\). If some \(C\in\mathcal{A}\), \(C\subset Z:=Y\setminus T\), had \(0<\nu( C)<\infty\), then \(\mu( C)>0\) and \(T\cup C\in\mathcal{G}\) would give \(\mu(T\cup C)=t+\mu( C)>t\); so \(Z\in\mathcal{S}\) and therefore \(\nu(Z)=0\). Writing \(T=\bigcup_kT_k\) with \(\nu(T_k)<\infty\), we get \(Y=Z\cup\bigcup_kT_k\), so \(\nu|_{X\setminus S}\) is \(\sigma\)-finite.
(ii) Both \(\mu\) and \(\nu\) are \([0,+\infty]\)-valued measures on \((X,\mathcal{B})\), not necessarily \(\sigma\)-finite there. Decomposing \(X=\bigcup_nX_n\) into disjoint \(X_n\in\mathcal{A}\) with \((\mu+\nu)(X_n)<\infty\) and setting \(w:=\sum_n2^{-n}(1+(\mu+\nu)(X_n))^{-1}I_{X_n}\in(0,1]\), the measure \(\lambda:=w\cdot(\mu+\nu)\) is finite with the same null sets as \(\mu+\nu\), so \(\mu\ll\lambda\) and \(\nu\ll\lambda\) on \(\mathcal{B}\). Applying (i) on \((X,\mathcal{B})\) to \((\lambda,\mu)\) and to \((\lambda,\nu)\) gives \(S_\mu,S_\nu\in\mathcal{B}\) off which \(\mu\), resp. \(\nu\), is \(\sigma\)-finite along \(\mathcal{B}\), while on them each takes only the values \(0\) and \(+\infty\). On \(T:=X\setminus(S_\mu\cup S_\nu)\) both are \(\sigma\)-finite, so intersecting the two exhaustions and disjointifying, \(T=\bigsqcup_mC_m\) with \(C_m\in\mathcal{B}\) and \(\mu(C_m)+\nu(C_m)<\infty\). On each trace \(\sigma\)-algebra \(\mathcal{B}_m=\{E\in\mathcal{B}:E\subset C_m\}\) the two measures are finite with \(\nu\ll\mu\), so the Radon–Nikodym theorem (Theorem 3.2.2) supplies \(\mathcal{B}_m\)-measurable \(\xi_m\ge0\) with \(\nu=\xi_m\cdot\mu\) on \(\mathcal{B}_m\). Put \(\xi:=\sum_m\xi_mI_{C_m}\) on \(T\) and \(\xi:=0\) elsewhere; countable additivity and monotone convergence give
\begin{equation*} \nu(E\cap T) = \sum_m\int_{E\cap C_m}\xi\,d\mu = \int_E\xi\,d\mu \qquad (E\in\mathcal{B}) . \end{equation*}
If now \(B\in\mathcal{B}\) has \(\mu(B)+\nu(B)<\infty\) and \(E\in\mathcal{B}\), \(E\subset B\), then \(\nu(B\cap S_\nu)\in\{0,+\infty\}\) is \(0\) and \(\mu(B\cap S_\mu)=0\), whence \(\nu(B\cap S_\mu)=0\) by \(\nu\ll\mu\); since \(E\setminus T\subset(B\cap S_\mu)\cup(B\cap S_\nu)\), we get \(\nu(E)=\nu(E\cap T)=\int_E\xi\,d\mu\), i.e. \(\nu|_B=\xi\cdot\mu|_B\).
Counterexample for all \(B\in\mathcal{B}\). Here \(\mathcal{B}\) consists of the at most countable sets and their complements (Exercise 3.10.30). Every \(\mathcal{B}\)-measurable \(\xi\) is constant off an at most countable set: each \(\{\xi\le s\}\) is countable or co-countable and the latter property is monotone in \(s\), so with \(c:=\inf\{s:\{\xi\le s\}\ \text{co-countable}\}\) one checks the three cases \(c=+\infty\) (then \(\{\xi<+\infty\}=\bigcup_n\{\xi\le n\}\) is countable), \(c=-\infty\) (then \(\{\xi=-\infty\}=\bigcap_n\{\xi\le-n\}\) is co-countable), and \(c\in\mathbb{R}\) (then \(\{\xi\le c\}\) is co-countable and \(\{\xi<c\}\) countable). With \(N\) countable and \(\xi\equiv c\) off \(N\) we have \(\mu(N)=0\) and \(\mu(\mathbb{R}\setminus N)=+\infty\), so \(\int_{\mathbb{R}}\xi\,d\mu\) equals \(0\), \(+\infty\) or \(-\infty\) according to the sign of \(c\) – never \(1=\nu(\mathbb{R})\), although \(\mathbb{R}\in\mathcal{B}\).
(\(\circ\)) Suppose we are given three bounded measures \(\mu_1\), \(\mu_2\), and \(\mu_3\) on a \(\sigma\)-algebra \(\mathcal{A}\) such that \(\mu_1 \ll \mu_2\) and \(\mu_2 \ll \mu_3\). Show that one has \(\mu_1 \ll \mu_3\) and \(d\mu_1/d\mu_3 = (d\mu_1/d\mu_2)(d\mu_2/d\mu_3)\).
If \(|\mu_3|(A)=0\) then \(|\mu_2|(A)=0\), hence \(|\mu_1|(A)=0\): thus \(\mu_1\ll\mu_3\) (the measures may be signed, and \(\ll\) is meant in the sense of Definition 3.2.1).
By the Radon–Nikodym theorem (Theorem 3.2.2), applicable since all three measures are bounded, \(\mu_1=f\cdot\mu_2\) and \(\mu_2=g\cdot\mu_3\) with \(f\in L^1(\mu_2)\), \(g\in L^1(\mu_3)\). The identity is then a change-of-variables lemma: for a bounded measure \(\lambda\), \(g\in L^1(\lambda)\) and \(\nu:=g\cdot\lambda\), one has \(|\nu|=|g|\cdot|\lambda|\), and every \(h\in L^1(\nu)\) satisfies \(hg\in L^1(\lambda)\) and \(\int h\,d\nu=\int hg\,d\lambda\).
Proof of the lemma. Writing \(\lambda=\xi\cdot|\lambda|\) with \(|\xi|=1\) and replacing \(g\) by \(g\xi\), we may assume \(\lambda\ge0\). Then \((\{g\ge0\},\{g<0\})\) is a Hahn decomposition for \(\nu\), so \(\nu^{\pm}=g^{\pm}\cdot\lambda\) by Corollary 3.1.2 and \(|\nu|=|g|\cdot\lambda\). Hence \(\int h\,d|\nu|=\int h|g|\,d\lambda\) holds for indicators, so for nonnegative simple \(h\), so for all measurable \(h\ge0\) by monotone convergence; in particular \(h\in L^1(\nu)\) iff \(hg\in L^1(\lambda)\). Likewise \(\int h\,d\nu^{\pm}=\int hg^{\pm}\,d\lambda\) for simple \(h\ge0\); for \(h\ge0\) in \(L^1(\nu)\) take simple \(0\le h_k\uparrow h\) and pass to the limit in \(\int h_k\,d\nu^{+}-\int h_k\,d\nu^{-}=\int h_kg^{+}\,d\lambda-\int h_kg^{-}\,d\lambda\), all four limits being finite because \(\int h\,d|\nu|=\int h|g|\,d\lambda<\infty\). For general \(h\) subtract the parts \(h^{\pm}\).
Applying the lemma with \(\lambda=\mu_3\), \(\nu=\mu_2\) and \(h=fI_A\in L^1(\mu_2)\),
\begin{equation*} \mu_1(A) = \int_A f\,d\mu_2 = \int_A fg\,d\mu_3 \qquad\text{for all } A\in\mathcal{A} , \end{equation*}
and \(A=X\) shows \(fg\in L^1(\mu_3)\). So \(\mu_1=(fg)\cdot\mu_3\), that is,
\begin{equation*} \frac{d\mu_1}{d\mu_3} = \frac{d\mu_1}{d\mu_2}\cdot\frac{d\mu_2}{d\mu_3} \qquad |\mu_3|\text{-a.e.} \end{equation*}
The product is unambiguous although \(f\) is determined only \(|\mu_2|\)-a.e.: since \(|\mu_2|=|g|\cdot|\mu_3|\), a \(|\mu_2|\)-null set carries \(g=0\) \(|\mu_3|\)-a.e.
(\(\circ\)) Let \(\mu\) and \(\nu\) be two probability measures on a \(\sigma\)-algebra \(\mathcal{A}\) such that for some \(\alpha \in (0, 1)\), one has \(\| \alpha \mu - (1 - \alpha) \nu \| = 1\). Prove that \(\mu \perp \nu\).
The separating set is \(A:=\{f>0\}\), where \(f=d\mu/d\sigma\) and \(\sigma:=(\mu+\nu)/2\).
Both \(\mu\ll\sigma\) and \(\nu\ll\sigma\), so the Radon–Nikodym theorem (Theorem 3.2.2) gives \(\mu=f\cdot\sigma\), \(\nu=g\cdot\sigma\) with \(f,g\in L^1(\sigma)\); the densities may be taken nonnegative everywhere, since \(\int_{\{f<0\}}f\,d\sigma=\mu(\{f<0\})\ge0\) forces \(\sigma(f<0)=0\), and likewise for \(g\). Here \(\int f\,d\sigma=\int g\,d\sigma=1\). For a nonnegative \(\sigma\) and \(h\in L^1(\sigma)\) one has \(|h\cdot\sigma|=|h|\cdot\sigma\) (the first part of the lemma in Exercise 3.10.32), so with \(h:=\alpha f-(1-\alpha)g\),
\begin{equation*} 1 = \|\alpha\mu-(1-\alpha)\nu\| = \int_X|\alpha f-(1-\alpha)g|\,d\sigma , \end{equation*}
while \(\int_X(\alpha f+(1-\alpha)g)\,d\sigma=\alpha+(1-\alpha)=1\). Subtracting and using \(a+b-|a-b|=2\min(a,b)\) for \(a,b\ge0\),
\begin{equation*} \int_X 2\min\bigl(\alpha f,(1-\alpha)g\bigr)\,d\sigma = 0 , \end{equation*}
so the nonnegative integrand vanishes \(\sigma\)-a.e.; as \(\alpha\in(0,1)\), this means \(fg=0\) \(\sigma\)-a.e. Hence \(g=0\) \(\sigma\)-a.e. on \(A\) and \(f=0\) on \(X\setminus A\), giving
\begin{equation*} \nu(A)=\int_Ag\,d\sigma=0 , \qquad \mu(X\setminus A)=\int_{X\setminus A}f\,d\sigma=0 , \end{equation*}
i.e. \(\mu\perp\nu\) with \(\Omega=A\) in Definition 3.2.1(ii).
(\(\circ\)) Let \(\mu\) and \(\nu\) be two probability measures such that \(\nu \ll \mu\). Show that if a sequence of \(\mu\)-measurable functions \(f_n\) converges in measure \(\mu\) to a function \(f\), then it converges to \(f\) in measure \(\nu\) as well.
Yes: \(\nu\ll\mu\) with \(\nu\) finite is equivalent to the \(\varepsilon\)-\(\delta\) form, and the sets \(A_n(\varepsilon):=\{|f_n-f|\ge\varepsilon\}\) have \(\mu(A_n(\varepsilon))\to0\).
Since \(\nu\ll\mu\), the measure \(\nu\) extends to the \(\mu\)-completion by \(\nu(A\triangle Z):=\nu(A)\) for \(Z\) inside a \(\mu\)-null set (well defined and countably additive, Check!), so the \(\mu\)-measurable sets \(A_n(\varepsilon)\) are measurable for both measures.
Claim: if \(\nu\ll\mu\) and \(\nu\) is finite, then for every \(\eta>0\) there is \(\delta>0\) with \(\mu(A)<\delta\Rightarrow\nu(A)\le\eta\). Otherwise there are \(B_n\) with \(\mu(B_n)<2^{-n}\) and \(\nu(B_n)>\eta\); then \(B:=\bigcap_N\bigcup_{n\ge N}B_n\) has \(\mu(B)\le\sum_{n\ge N}2^{-n}=2^{-N+1}\) for every \(N\), so \(\mu(B)=0\), while the sets \(\bigcup_{n\ge N}B_n\) decrease with \(\nu\)-measure \(>\eta\), so \(\nu(B)\ge\eta>0\) by continuity from above of the finite \(\nu\) – contradicting \(\nu\ll\mu\). (Alternatively: \(\nu=\varrho\cdot\mu\) by Theorem 3.2.2, and the claim is the absolute continuity of the Lebesgue integral, Theorem 2.5.7.)
Given \(\varepsilon>0\) and \(\eta>0\), take \(\delta\) from the claim; since \(\mu(A_n(\varepsilon))\to0\) there is \(n_0\) with \(\mu(A_n(\varepsilon))<\delta\), hence \(\nu(A_n(\varepsilon))\le\eta\), for \(n\ge n_0\). As \(\eta\) and \(\varepsilon\) were arbitrary, \(f_n\to f\) in measure \(\nu\).
(\(\circ\)) Let \(\mu\) and \(\nu\) be two probability measures and let \(f_n\), \(n \in \mathbb{N}\), and \(f\) be \(\mu \otimes \nu\)-measurable functions such that for \(\mu\)-a.e. fixed \(x\) the functions \(f_n(\,\cdot\,, x)\) converge to \(f(\,\cdot\,, x)\) in measure \(\nu\). Show that the functions \(f_n\) converge to \(f\) in measure \(\mu \otimes \nu\).
Test convergence in measure by the functional \(d_\lambda(h):=\int_Z\frac{|h|}{1+|h|}\,d\lambda\) and integrate it in \(x\).
Lemma: for a finite measure \(\lambda\) and \(\lambda\)-measurable \(h_n\), one has \(h_n\to0\) in measure \(\lambda\) iff \(d_\lambda(h_n)\to0\). Indeed \(t\mapsto t/(1+t)\) increases and is bounded by \(1\), so
\begin{equation*} d_\lambda(h_n) \le \varepsilon\lambda(Z)+\lambda(\{|h_n|>\varepsilon\}) , \qquad \lambda(\{|h_n|\ge\varepsilon\}) \le \frac{1+\varepsilon}{\varepsilon}\,d_\lambda(h_n) , \end{equation*}
the second because \(|h_n|/(1+|h_n|)\ge\varepsilon/(1+\varepsilon)\) on \(\{|h_n|\ge\varepsilon\}\); let \(n\to\infty\), then \(\varepsilon\to0\).
Put \(g_n:=|f_n-f|/(1+|f_n-f|)\), a \(\mu\otimes\nu\)-measurable function with \(0\le g_n\le1\), hence integrable for the probability measure \(\mu\otimes\nu\). Fubini’s theorem (Theorem 3.4.4, with measurability of the sections from Theorem 3.4.1) makes \(\varphi_n(x):=\int_Yg_n(y,x)\,\nu(dy)\) defined for \(\mu\)-a.e. \(x\), \(\mu\)-measurable, with \(0\le\varphi_n\le1\) and \(\int g_n\,d(\mu\otimes\nu)=\int_X\varphi_n\,d\mu\). For \(\mu\)-a.e. \(x\) the lemma applied to \(\lambda=\nu\) and \(h_n=f_n(\cdot,x)-f(\cdot,x)\) gives \(\varphi_n(x)=d_\nu(h_n)\to0\), so dominated convergence gives \(\int g_n\,d(\mu\otimes\nu)\to0\). By the lemma again, now for \(\lambda=\mu\otimes\nu\),
\begin{equation*} (\mu\otimes\nu)\bigl(\{|f_n-f|\ge\varepsilon\}\bigr) \le \frac{1+\varepsilon}{\varepsilon}\int g_n\,d(\mu\otimes\nu) \longrightarrow 0 \end{equation*}
for every \(\varepsilon>0\), i.e. \(f_n\to f\) in measure \(\mu\otimes\nu\).
Exercises 3.10.36–3.10.42
Suppose that a sequence of measures \(\mu_n\) on a measurable space \((X,\mathcal{A})\) converges in variation to a measure \(\mu\) and a sequence of measures \(\nu_n\) converges in variation to a measure \(\nu\). Let \(\nu_n = \nu_n^{ac} + \nu_n^{s}\), \(\nu = \nu^{ac} + \nu^{s}\), where \(\nu_n^{ac} \ll \mu_n\), \(\nu_n^{s} \perp \mu_n\), \(\nu^{ac} \ll \mu\), \(\nu^{s} \perp \mu\). Prove that \(\mathcal{A}\)-measurable versions of the Radon-Nikodym densities \(d\nu_n^{ac}/d\mu_n\) converge to \(d\nu^{ac}/d\mu\) in measure \(|\mu|\). In particular, if \(\mu_n \ll \mu\) and \(\nu_n \ll \mu_n\), then \(d\nu_n/d\mu_n \to d\nu/d\mu\) in measure \(|\mu|\).
The densities are \(h_n=I_{\{f_n\ne0\}}g_n/f_n\) and \(h=I_{\{f\ne0\}}g/f\), where \(f_n,g_n,f,g\) are densities of \(\mu_n,\nu_n,\mu,\nu\) with respect to one dominating measure, and they converge because \(f_n\to f\), \(g_n\to g\) in measure there.
Dominating measure. Put \(\sigma:=|\mu|+|\nu|+\sum_n2^{-n}(|\mu_n|+|\nu_n|)/(\|\mu_n\|+\|\nu_n\|)\), omitting the summands with \(\|\mu_n\|+\|\nu_n\|=0\); each summand is nonnegative of mass \(2^{-n}\), so \(\sigma\) is a finite nonnegative measure dominating \(|\mu|,|\nu|,|\mu_n|,|\nu_n|\). By the Radon–Nikodym theorem (Theorem 3.2.2) there are \(f,g,f_n,g_n\in\mathcal{L}^1(\sigma)\) with \(\mu=f\cdot\sigma\), \(\nu=g\cdot\sigma\), \(\mu_n=f_n\cdot\sigma\), \(\nu_n=g_n\cdot\sigma\).
If \(\lambda=\varphi\cdot\sigma\), then \((\{\varphi\ge0\},\{\varphi<0\})\) is a Hahn decomposition for \(\lambda\), so \(|\lambda|=|\varphi|\cdot\sigma\) and \(\|\lambda\|=\int|\varphi|\,d\sigma\); applied to \(\mu_n-\mu\) and \(\nu_n-\nu\) this gives
\begin{equation*} \int_X|f_n-f|\,d\sigma = \|\mu_n-\mu\|\to0 , \qquad \int_X|g_n-g|\,d\sigma = \|\nu_n-\nu\|\to0 , \end{equation*}
hence \(f_n\to f\) and \(g_n\to g\) in measure \(\sigma\) (Chebyshev). Also \(|\mu|=|f|\cdot\sigma\) and \(|\mu_n|=|f_n|\cdot\sigma\).
The versions. Split \(\nu_n=\alpha_n+\beta_n\) with \(\alpha_n:=I_{\{f_n\ne0\}}g_n\cdot\sigma\) and \(\beta_n:=I_{\{f_n=0\}}g_n\cdot\sigma\). With \(\Omega_n:=\{f_n=0\}\) one has \(|\mu_n|(\Omega_n)=0\) and \(|\beta_n|(X\setminus\Omega_n)=0\), so \(\beta_n\perp\mu_n\); and \(|\mu_n|(A)=0\) gives \(\sigma(A\cap\{f_n\ne0\})=0\), hence \(|\alpha_n|(A)=0\), so \(\alpha_n\ll\mu_n\). The Lebesgue decomposition (Theorem 3.2.3) is unique: if \(\nu=\alpha^{(1)}+\beta^{(1)}=\alpha^{(2)}+\beta^{(2)}\), then \(\lambda:=\alpha^{(1)}-\alpha^{(2)}=\beta^{(2)}-\beta^{(1)}\) is concentrated on the \(|\mu|\)-null set \(\Omega^{(1)}\cup\Omega^{(2)}\) and is \(\ll\mu\), so \(\lambda=0\). Hence \(\nu_n^{ac}=\alpha_n\), and since \(h_n\cdot\mu_n=(h_nf_n)\cdot\sigma=\alpha_n\) with \(\int|h_n|\,d|\mu_n|=\int_{\{f_n\ne0\}}|g_n|\,d\sigma\le\|\nu_n\|<\infty\), the function \(h_n\) is a version of \(d\nu_n^{ac}/d\mu_n\); the same computation without \(n\) applies to \(h\).
Convergence. Note \(|\mu|\) is finite, \(|\mu|\ll\sigma\), and \(|\mu|\) vanishes off \(S:=\{f\ne0\}\). If \(h_n\not\to h\) in measure \(|\mu|\), there are \(\varepsilon,\delta>0\) and a subsequence with \(|\mu|(\{|h_{n_k}-h|\ge\varepsilon\})\ge\delta\). By the Riesz theorem a further subsequence has \(f_{n_{k_j}}\to f\) and \(g_{n_{k_j}}\to g\) \(\sigma\)-a.e.; at such a point of \(S\) we have \(f_{n_{k_j}}\ne0\) for large \(j\) and therefore \(h_{n_{k_j}}\to g/f=h\). So \(h_{n_{k_j}}\to h\) \(\sigma\)-a.e. on \(S\), hence \(|\mu|\)-a.e., hence in measure \(|\mu|\) as \(|\mu|\) is finite – a contradiction.
Arbitrary versions. If \(\tilde h_n\cdot\mu_n=\nu_n^{ac}\), then \(\int_A(\tilde h_n-h_n)f_n\,d\sigma=0\) for all \(A\), so \((\tilde h_n-h_n)f_n=0\) \(\sigma\)-a.e. and \(|\mu_n|(N_n)=0\) for \(N_n:=\{\tilde h_n\ne h_n\}\); consequently
\begin{equation*} |\mu|(N_n) \le \int_{N_n}|f_n|\,d\sigma+\int_X|f-f_n|\,d\sigma = \|\mu-\mu_n\|\to0 , \end{equation*}
and \(\{|\tilde h_n-h|\ge\varepsilon\}\subset\{|h_n-h|\ge\varepsilon\}\cup N_n\) gives \(\tilde h_n\to h\) in measure \(|\mu|\); any version of \(d\nu^{ac}/d\mu\) equals \(h\) outside a \(|\mu|\)-null set.
Particular case. If \(\mu_n\ll\mu\) and \(\nu_n\ll\mu_n\), then \(\nu_n\ll\mu\); and \(\nu\ll\mu\), since \(|\mu|(A)=0\) gives \(\nu_n(B)=0\) for every measurable \(B\subset A\), so \(|\nu(B)|=|\nu(B)-\nu_n(B)|\le\|\nu-\nu_n\|\to0\) and \(|\nu|(A)=0\) by formula (3.1.3). Hence \(\nu_n^{ac}=\nu_n\), \(\nu^{ac}=\nu\), and the statement reads \(d\nu_n/d\mu_n\to d\nu/d\mu\) in measure \(|\mu|\).
(Nikodym) Let \(\mu\) be a bounded nonnegative measure on a \(\sigma\)-algebra \(\mathcal{A}\) in a space \(X\), let \(G\) be a nonmeasurable set. Let \(\sigma(\mathcal{A}\cup G)\) be the \(\sigma\)-algebra generated by \(\mathcal{A}\) and \(G\), and let \(\underline{G}\) and \(\widetilde{G}\) be a measurable kernel and a measurable envelope of \(G\). Denote by \(\gamma_1\) and \(\gamma_2\) the Radon-Nikodym densities of the measures \(A \mapsto \mu(A \cap \underline{G})\) and \(A \mapsto \mu(A \cap \widetilde{G})\) with respect to \(\mu\). Let \(\gamma\) be a \(\mu\)-measurable function such that \(\gamma_1 \le \gamma \le \gamma_2\). Show that the formula
\begin{equation*} \nu(E) = \int_A \gamma(x)\,\mu(dx) + \int_B \bigl(1 - \gamma(x)\bigr)\,\mu(dx), \end{equation*}
where \(E = (A \cap G) \cup \bigl(B \cap (X\setminus G)\bigr)\), \(A, B \in \mathcal{A}\), defines a countably additive extension of \(\mu\) to \(\sigma(\mathcal{A}\cup G)\) and that every countably additive extension of \(\mu\) to \(\sigma(\mathcal{A}\cup G)\) has such a form.
Both halves rest on the two facts \(\mu(C\setminus\underline G)=0\) for measurable \(C\subset G\) and \(\mu(C\cap\widetilde G)=0\) for measurable \(C\subset G^{c}:=X\setminus G\), which say that \(\gamma\) may be prescribed to be \(1\) on \(\underline G\) and \(0\) off \(\widetilde G\).
Those two facts follow from 1.12(iv): \(\underline G\cup C\in\mathcal{A}\) lies in \(G\), so \(\mu(\underline G)+\mu(C\setminus\underline G)\le\mu_{*}(G)=\mu(\underline G)\); and \(\widetilde G\setminus C\in\mathcal{A}\) contains \(G\), so \(\mu(\widetilde G)-\mu(\widetilde G\cap C)\ge\mu^{*}(G)=\mu(\widetilde G)\). Since \(I_{\underline G}\) and \(I_{\widetilde G}\) are versions of \(\gamma_1,\gamma_2\), the hypothesis \(\gamma_1\le\gamma\le\gamma_2\) says
\begin{equation*} 0\le\gamma\le1,\qquad \gamma=1\ \text{ on }\underline G, \qquad \gamma=0\ \text{ on } X\setminus\widetilde G , \end{equation*}
\(\mu\)-a.e.; modifying \(\gamma\) on a \(\mu\)-null set (which changes no integral below) we assume these hold everywhere with \(\gamma\) being \(\mathcal{A}\)-measurable, so that \(\gamma\cdot\mu\) and \((1-\gamma)\cdot\mu\) are finite nonnegative measures with sum \(\mu\).
Structure of \(\sigma(\mathcal{A}\cup G)\). The class \(\mathcal{E}:=\{(A\cap G)\cup(B\cap G^{c}):A,B\in\mathcal{A}\}\) contains \(\mathcal{A}\) (take \(A=B\)) and \(G\) (take \(A=X\), \(B=\emptyset\)), and is a \(\sigma\)-algebra because complements and countable unions act coordinatewise:
\begin{equation*} X\setminus\bigl[(A\cap G)\cup(B\cap G^{c})\bigr] = \bigl((X\setminus A)\cap G\bigr)\cup\bigl((X\setminus B)\cap G^{c}\bigr) . \end{equation*}
Hence \(\mathcal{E}=\sigma(\mathcal{A}\cup G)\) (Example 1.2.7, as in the proof of Theorem 1.12.14), so every \(E\) has the stated form.
\(\nu\) is well defined. If \((A_1\cap G)\cup(B_1\cap G^{c})=(A_2\cap G)\cup(B_2\cap G^{c})\), intersecting with \(G\) gives \(A_1\triangle A_2\subset G^{c}\), so \(\mu((A_1\triangle A_2)\cap\widetilde G)=0\) and, as \(\gamma=0\) off \(\widetilde G\), \(\int_{A_1}\gamma\,d\mu=\int_{A_2}\gamma\,d\mu\); symmetrically \(B_1\triangle B_2\subset G\), so \(\mu((B_1\triangle B_2)\setminus\underline G)=0\) and, as \(1-\gamma=0\) on \(\underline G\), \(\int_{B_1}(1-\gamma)\,d\mu=\int_{B_2}(1-\gamma)\,d\mu\).
\(\nu\) is a countably additive extension. \(\nu\ge0\), and \(A=(A\cap G)\cup(A\cap G^{c})\) gives \(\nu(A)=\int_A\gamma\,d\mu+\int_A(1-\gamma)\,d\mu=\mu(A)\) for \(A\in\mathcal{A}\). If the \(E_n=(A_n\cap G)\cup(B_n\cap G^{c})\) are disjoint, replace \(A_n,B_n\) by \(A_n^{\prime}:=A_n\setminus\bigcup_{m<n}A_m\) and \(B_n^{\prime}:=B_n\setminus\bigcup_{m<n}B_m\): disjointness gives \(A_n^{\prime}\cap G=A_n\cap G\) and \(B_n^{\prime}\cap G^{c}=B_n\cap G^{c}\), so \(E_n\) is unchanged while the \(A_n^{\prime}\), resp. \(B_n^{\prime}\), are disjoint with union representing \(E=\bigcup_nE_n\). Countable additivity of the finite measures \(\gamma\cdot\mu\) and \((1-\gamma)\cdot\mu\) now gives \(\nu(E)=\sum_n\nu(E_n)\).
Every extension has this form. Let \(\nu\ge0\) be countably additive on \(\mathcal{E}\) with \(\nu|_{\mathcal{A}}=\mu\), and put \(\theta(A):=\nu(A\cap G)\); as \(A\mapsto A\cap G\) preserves disjointness and countable unions, \(\theta\) is a nonnegative measure on \(\mathcal{A}\) with \(0\le\theta\le\mu\), hence \(\theta\ll\mu\) and \(\theta=\gamma\cdot\mu\) with \(0\le\gamma\le1\) \(\mu\)-a.e. by Theorem 3.2.2. For measurable \(A\subset\underline G\) one has \(\theta(A)=\nu(A)=\mu(A)\), so \(\gamma=1\) \(\mu\)-a.e. on \(\underline G\); for measurable \(A\subset X\setminus\widetilde G\) one has \(A\cap G=\emptyset\) and \(\theta(A)=0\), so \(\gamma=0\) \(\mu\)-a.e. there. Thus \(\gamma_1\le\gamma\le\gamma_2\) \(\mu\)-a.e., and since \(A\cap G\) and \(B\cap G^{c}\) are disjoint,
\begin{equation*} \nu(E) = \theta(A)+\bigl[\nu(B)-\nu(B\cap G)\bigr] = \int_A\gamma\,d\mu+\int_B(1-\gamma)\,d\mu . \end{equation*}
(\(\circ\)) Let \((X,\mathcal{A})\) and \((Y,\mathcal{B})\) be two measurable spaces. Show that every set in \(\mathcal{A}\otimes\mathcal{B}\) is contained in the \(\sigma\)-algebra generated by sets \(A_n\times B_n\) for some at most countable collections \(\{A_n\}\subset\mathcal{A}\) and \(\{B_n\}\subset\mathcal{B}\).
The collection \(\mathcal{M}\) of sets lying in \(\sigma(\{A_n\times B_n\})\) for some at most countable \(\{A_n\}\subset\mathcal{A}\), \(\{B_n\}\subset\mathcal{B}\) is itself a \(\sigma\)-algebra containing all measurable rectangles, hence contains \(\mathcal{A}\otimes\mathcal{B}\).
Every rectangle \(A\times B\) lies in \(\mathcal{M}\) (take the one-element families), so \(X\times Y\in\mathcal{M}\). If \(E\in\sigma(\mathcal{R}_0)\) for an at most countable family \(\mathcal{R}_0\) of rectangles, then \((X\times Y)\setminus E\in\sigma(\mathcal{R}_0)\) too. And if \(E_k\in\sigma(\mathcal{R}_k)\) with each \(\mathcal{R}_k\) at most countable, then \(\mathcal{R}_\infty:=\bigcup_k\mathcal{R}_k\) is an at most countable family of rectangles with \(\sigma(\mathcal{R}_k)\subset\sigma(\mathcal{R}_\infty)\) for every \(k\), so \(\bigcup_kE_k\in\sigma(\mathcal{R}_\infty)\subset\mathcal{M}\) after re-indexing by a single index. Since \(\mathcal{A}\otimes\mathcal{B}=\sigma(\mathcal{R})\) with \(\mathcal{R}\) the family of all measurable rectangles, the inclusion \(\mathcal{A}\otimes\mathcal{B}\subset\mathcal{M}\) is exactly the assertion. (This is the instance \(\mathcal{F}=\mathcal{R}\) of Problem 1.12.54.)
(\(\circ\)) Let \((X,\mathcal{A})\) and \((Y,\mathcal{B})\) be two measurable spaces and let a mapping \(f\colon X \to Y\) be \((\mathcal{A},\mathcal{B})\)-measurable. Show that the mapping \(\varphi\colon x \mapsto \bigl(x, f(x)\bigr)\) from \(X\) to \(X\times Y\) is \((\mathcal{A}, \mathcal{A}\otimes\mathcal{B})\)-measurable. Deduce from this that, given a measurable space \((Z,\mathcal{E})\) and a mapping \(g\colon X\times Y \to Z\) measurable with respect to the pair \((\mathcal{A}\otimes\mathcal{B}, \mathcal{E})\), the mapping \(x \mapsto g\bigl(x, f(x)\bigr)\) from \(X\) to \(Z\) is \((\mathcal{A},\mathcal{E})\)-measurable.
The class \(\mathcal{S}:=\{S\subset X\times Y:\varphi^{-1}(S)\in\mathcal{A}\}\) is a \(\sigma\)-algebra, because preimages commute with complements and countable unions, and it contains every measurable rectangle:
\begin{equation*} \varphi^{-1}(A\times B) = \{x: x\in A,\ f(x)\in B\} = A\cap f^{-1}(B) \in \mathcal{A} , \end{equation*}
using the \((\mathcal{A},\mathcal{B})\)-measurability of \(f\). Since \(\mathcal{A}\otimes\mathcal{B}\) is generated by such rectangles, \(\mathcal{A}\otimes\mathcal{B}\subset\mathcal{S}\), i.e. \(\varphi\) is \((\mathcal{A},\mathcal{A}\otimes\mathcal{B})\)-measurable.
For the second claim, \(x\mapsto g(x,f(x))\) is the composition \(g\circ\varphi\), and for \(C\in\mathcal{E}\) one has \(g^{-1}( C)\in\mathcal{A}\otimes\mathcal{B}\), hence \((g\circ\varphi)^{-1}( C)=\varphi^{-1}(g^{-1}( C))\in\mathcal{A}\).
Let \(T = \{(x,y)\in[0,1]^2 : x - y \in \mathbb{Q}\}\). Show that \(T\) has measure zero, but meets every set of the form \(A\times B\), where \(A\) and \(B\) are sets of positive measure in \([0,1]\). See also Exercise 3.10.63.
All sections of \(T\) are countable, which kills its measure; and \(A-B\) contains an interval, which supplies a rational difference.
\(T=\bigcup_{q\in\mathbb{Q}}L_q\) with \(L_q:=\{(x,y)\in[0,1]^2:x-y=q\}\) closed, so \(T\) is Borel, and its sections \(T^{y}=(y+\mathbb{Q})\cap[0,1]\) are at most countable, whence by Theorem 3.4.1 (formula (3.4.1))
\begin{equation*} \lambda_2(T) = \int_0^1\lambda(T^{y})\,dy = 0 . \end{equation*}
Let \(\lambda(A)>0\) and \(\lambda(B)>0\). Then \(-B\) is measurable with \(\lambda(-B)=\lambda(B)>0\), so by Example 3.9.7 the algebraic sum \(A-B=A+(-B)\) contains an interval \((\alpha,\beta)\). Pick a rational \(q\in(\alpha,\beta)\) and write \(q=a-b\) with \(a\in A\), \(b\in B\); then \((a,b)\in(A\times B)\cap T\).
(\(\circ\)) Suppose that a function \(f\) on \([0,1]^2\) is Lebesgue measurable and that, for a.e. \(x\) and a.e. \(y\), the functions \(z \mapsto f(x,z)\) and \(z \mapsto f(z,y)\) are constant. Show that \(f = c\) a.e. for some constant \(c\).
If some level set split the square nontrivially, a full vertical segment and a full horizontal segment would have to be disjoint, which they are not.
Let \(X_0, Y_0\subset[0,1]\) be the full-measure sets on which \(z\mapsto f(x,z)\), resp. \(z\mapsto f(z,y)\), is constant, and put \(F( r):=\lambda_2(\{f<r\})\). If \(F( r)\in\{0,1\}\) for every \(r\), then \(c:=\sup\{r:F( r)=0\}\) is finite (\(F\) is nondecreasing with limits \(0\) and \(1\) at \(\mp\infty\), \(f\) being finite), and letting \(n\to\infty\) in \(\lambda_2(\{f<c-1/n\})=0\) and \(\lambda_2(\{f\ge c+1/n\})=0\) gives \(f=c\) a.e.
So suppose \(0<F( r)<1\) for some \(r\), i.e. the disjoint sets \(P:=\{f<r\}\) and \(Q:=\{f\ge r\}\) with \(P\cup Q=[0,1]^2\) both have positive measure. By Theorem 3.4.1 the sections \(P_x\) are measurable for a.e. \(x\) and \(0<\lambda_2(P)=\int_0^1\lambda(P_x)\,dx\), so \(\{x:\lambda(P_x)>0\}\) has positive measure and meets the full-measure set \(X_0\): pick \(x_0\in X_0\) with \(\lambda(P_{x_0})>0\). Then \(f(x_0,y_1)<r\) for some \(y_1\), and constancy of \(y\mapsto f(x_0,y)\) gives
\begin{equation*} \{x_0\}\times[0,1]\subset P . \end{equation*}
Symmetrically, from \(0<\lambda_2(Q)=\int_0^1\lambda(Q^{y})\,dy\) pick \(y_0\in Y_0\) with \(\lambda(Q^{y_0})>0\); constancy of \(x\mapsto f(x,y_0)\) gives \([0,1]\times\{y_0\}\subset Q\). But then \((x_0,y_0)\in P\cap Q=\emptyset\).
Let \(\mu\) and \(\nu\) be finite nonnegative measures on measurable spaces \((X,\mathcal{A})\) and \((Y,\mathcal{B})\), \(A \subset X\), \(B \subset Y\). Prove the equality \((\mu\otimes\nu)^{*}(A\times B) = \mu^{*}(A)\,\nu^{*}(B)\).
Measurable envelopes give \(\le\), and the section formula gives \(\ge\).
Let \(\widetilde A\in\mathcal{A}\), \(\widetilde B\in\mathcal{B}\) be measurable envelopes of \(A\) and \(B\), which exist by (1.12.4): \(A\subset\widetilde A\), \(\mu(\widetilde A)=\mu^{*}(A)\), and similarly for \(B\). Then \(A\times B\subset\widetilde A\times\widetilde B\in\mathcal{A}\otimes\mathcal{B}\), so by the definition of the product measure
\begin{equation*} (\mu\otimes\nu)^{*}(A\times B) \le (\mu\otimes\nu)(\widetilde A\times\widetilde B) = \mu^{*}(A)\,\nu^{*}(B) . \end{equation*}
Conversely let \(E\in\mathcal{A}\otimes\mathcal{B}\) contain \(A\times B\). Its sections \(E^{y}=\{x:(x,y)\in E\}\) lie in \(\mathcal{A}\) and \(y\mapsto\mu(E^{y})\) is \(\mathcal{B}\)-measurable by Proposition 3.3.2. For \(y\in B\) we have \(A\times\{y\}\subset E\), i.e. \(A\subset E^{y}\), so \(\mu(E^{y})\ge\mu^{*}(A)\); hence \(B\subset B_0:=\{y:\mu(E^{y})\ge\mu^{*}(A)\}\in\mathcal{B}\) and \(\nu(B_0)\ge\nu^{*}(B)\). Theorem 3.4.1 (formula (3.4.1)) then gives
\begin{equation*} (\mu\otimes\nu)(E) = \int_Y\mu(E^{y})\,\nu(dy) \ \ge\ \mu^{*}(A)\,\nu(B_0) \ \ge\ \mu^{*}(A)\,\nu^{*}(B) , \end{equation*}
and taking the infimum over such \(E\) yields \((\mu\otimes\nu)^{*}(A\times B)\ge\mu^{*}(A)\nu^{*}(B)\).
Exercises 3.10.43–3.10.49
Let \((X,\mathcal{A})\) and \((Y,\mathcal{B})\) be measurable spaces. Show that, for every \(E \in \mathcal{A}\otimes\mathcal{B}\), the family of sections \(E_x=\{y\in Y:\ (x,y)\in E\}\) contains at most continuum of distinct sets.
The section \(E_x\) depends on \(x\) only through a sequence of zeros and ones, so there are at most \(\mathfrak{c}:=\operatorname{card}\{0,1\}^{\mathbb{N}}\) of them.
By Exercise 3.10.38 there are at most countable families \(\{A_n\}\subset\mathcal{A}\) and \(\{B_n\}\subset\mathcal{B}\) (indexed by all of \(\mathbb{N}\), repeating sets if necessary) with \(E\in\mathcal{E}:=\sigma(\{A_n\times B_n: n\in\mathbb{N}\})\). Put \(\Phi(x):=(I_{A_n}(x))_{n\in\mathbb{N}}\in\{0,1\}^{\mathbb{N}}\) and let
\begin{equation*} \mathcal{D} := \{S\subset X\times Y:\ \Phi(x_1)=\Phi(x_2) \ \Longrightarrow\ S_{x_1}=S_{x_2}\} . \end{equation*}
Since sectioning commutes with complements and countable unions, \(\mathcal{D}\) is a \(\sigma\)-algebra; and it contains every generator, because \((A_n\times B_n)_x\) equals \(B_n\) if \(x\in A_n\) and \(\emptyset\) otherwise, so the section depends on \(x\) only through \(I_{A_n}(x)\). Hence \(\mathcal{E}\subset\mathcal{D}\) and in particular \(E\in\mathcal{D}\).
So \(x\mapsto E_x\) factors through \(\Phi\), say \(E_x=\Psi(\Phi(x))\), and therefore
\begin{equation*} \operatorname{card}\{E_x:\ x\in X\} \le \operatorname{card}\Phi(X) \le \operatorname{card}\{0,1\}^{\mathbb{N}} = \mathfrak{c} . \end{equation*}
Let \((X,\mathcal{A})\) be a measurable space of cardinality greater than that of the continuum. Show that the diagonal \(D=\{(x,x),\ x\in X\}\) does not belong to the \(\sigma\)-algebra \(\mathcal{A}\otimes\mathcal{A}\).
The sections of \(D\) are the singletons \(D_x=\{x\}\), so \(x\mapsto D_x\) is injective and \(D\) has \(\operatorname{card}X>\mathfrak{c}\) distinct sections, where \(\mathfrak{c}\) is the cardinality of the continuum. By Exercise 3.10.43 every \(E\in\mathcal{A}\otimes\mathcal{A}\) has at most \(\mathfrak{c}\) distinct sections; hence \(D\notin\mathcal{A}\otimes\mathcal{A}\).
(\(\circ\)) Construct examples showing that (a) the existence and equality of the repeated integrals in (3.4.3) do not guarantee the \(\mu\otimes\nu\)-integrability of a measurable function \(f\); (b) it may occur that both repeated integrals exist for some measurable function \(f\), but are not equal; (c) there exists a measurable function \(f\) such that one of the repeated integrals exists, but the other one does not.
(a) \(f(x,y)=xy(x^2+y^2)^{-2}\) on \([-1,1]^2\); (b) \(f(x,y)=(x^2-y^2)(x^2+y^2)^{-2}\) on \([0,1]^2\); (c) \(f(n,m)=1\) if \(m=2n-1\), \(=-1\) if \(m=2n\), \(=0\) otherwise, on \(\mathbb{N}\times\mathbb{N}\) with counting measures. In each case \(f(0,0):=0\), and recall that in Bogachev’s terminology a repeated integral exists when the inner integral is defined for a.e. value of the outer variable and the resulting function is integrable.
(a) For \(0<|y|\le1\) the denominator is \(\ge y^4>0\), so \(x\mapsto f(x,y)\) is continuous and bounded, hence integrable, and it is odd: \(\int_{-1}^1f(x,y)\,dx=0\). By the symmetry \(f(x,y)=f(y,x)\) both repeated integrals exist and vanish. But for \(0<|y|\le1\), using \(\frac{d}{dx}\bigl(-\tfrac12(x^2+y^2)^{-1}\bigr)=x(x^2+y^2)^{-2}\),
\begin{equation*} \int_{-1}^{1}|f(x,y)|\,dx \ \ge\ \int_0^{|y|}\frac{x|y|}{(x^2+y^2)^2}\,dx = |y|\Bigl(\frac{1}{2y^2}-\frac{1}{4y^2}\Bigr) = \frac{1}{4|y|} , \end{equation*}
which is not integrable in \(y\); were \(f\) product integrable, Fubini’s theorem (Theorem 3.4.4) applied to \(|f|\) would make \(y\mapsto\int_{-1}^1|f(x,y)|\,dx\) integrable. So \(f\notin L^1(\lambda\otimes\lambda)\).
(b) Since \(\partial_y\bigl(y(x^2+y^2)^{-1}\bigr)=(x^2-y^2)(x^2+y^2)^{-2}\), for fixed \(x\in(0,1]\) the bounded continuous function \(y\mapsto f(x,y)\) integrates to \((1+x^2)^{-1}\), whence
\begin{equation*} \int_0^1\Bigl(\int_0^1f(x,y)\,dy\Bigr)dx = \int_0^1\frac{dx}{1+x^2} = \frac{\pi}{4} , \end{equation*}
while \(f(x,y)=-f(y,x)\) gives \(\int_0^1(\int_0^1f\,dx)dy=-\pi/4\). Both repeated integrals exist and differ.
(c) Every function on \(\mathbb{N}\times\mathbb{N}\) is measurable. For fixed \(n\) the function \(m\mapsto f(n,m)\) is nonzero only at \(m=2n-1\) and \(m=2n\), so it is integrable with \(\int_Yf\,\nu(dm)=1-1=0\), and the zero function is integrable: that repeated integral exists and is \(0\). For fixed \(m\) there is exactly one \(n\) with \(f(n,m)\ne0\), namely \(n=(m+1)/2\) for odd \(m\) and \(n=m/2\) for even \(m\), so \(g(m):=\int_Xf\,\mu(dn)=(-1)^{m+1}\); as \(\int_Y|g|\,d\nu=\sum_m1=+\infty\), the other repeated integral does not exist.
(Minkowski’s inequality for integrals) Let \((X,\mathcal{A},\mu)\) and \((Y,\mathcal{B},\nu)\) be spaces with nonnegative \(\sigma\)-finite measures and let \(f\) be an \(\mathcal{A}\otimes\mathcal{B}\)-measurable function. Prove that whenever \(1\le p<q<\infty\) one has \[ \int_Y\Bigl(\int_X|f(x,y)|^p\,\mu(dx)\Bigr)^{q/p}\nu(dy)\ \le\ \Bigl(\int_X\Bigl(\int_Y|f(x,y)|^q\,\nu(dy)\Bigr)^{p/q}\mu(dx)\Bigr)^{q/p}. \]
The whole inequality reduces to the case \(p=1\), where it is Holder’s inequality applied to \(F^{r-1}\cdot g\) with \(F(y)=\int_Xg(x,y)\,\mu(dx)\).
Replacing \(f\) by \(|f|\) we assume \(f\ge0\); all integrands below are nonnegative and \(\mathcal{A}\otimes\mathcal{B}\)-measurable, so all integrals live in \([0,+\infty]\). Throughout we use Tonelli’s theorem in the form: for \(h\ge0\) product measurable and \(\mu,\nu\) \(\sigma\)-finite, the inner integrals \(\int_Xh\,d\mu\) and \(\int_Yh\,d\nu\) are measurable functions of the free variable and the two repeated integrals both equal \(\int h\,d(\mu\otimes\nu)\) in \([0,+\infty]\). (Sections are measurable by Proposition 3.3.2(i), measurability of \(y\mapsto\mu(E^{y})\) is Proposition 3.3.2(ii) extended to \(\sigma\)-finite \(\mu\) by monotone convergence, and the identity follows from Theorem 3.4.4 applied to the integrable \(h_n:=\min(h,n)I_{X_n\times Y_n}\) followed by monotone convergence; it complements Theorem 3.4.5.)
Reduction. Suppose that for every \(r>1\) and nonnegative product-measurable \(g\),
\begin{equation*} \int_Y\Bigl(\int_Xg\,\mu(dx)\Bigr)^{r}\nu(dy) \le \Bigl(\int_X\Bigl(\int_Yg^{r}\,\nu(dy)\Bigr)^{1/r}\mu(dx)\Bigr)^{r} . \tag{*} \end{equation*}
Given \(1\le p<q<\infty\), the choice \(r=q/p>1\) and \(g=f^{p}\) turns the two sides of \((\ast)\) into the two sides of the assertion, since \(g^{r}=f^{q}\) and \(1/r=p/q\).
Proof of \((\ast)\). Fix \(r>1\) and \(g\ge0\), and put \(F(y):=\int_Xg(x,y)\,\mu(dx)\), \(J:=\int_YF^{r}\,d\nu\) and \(M:=\int_X(\int_Yg(z,y)^{r}\nu(dy))^{1/r}\mu(dz)\). We may assume \(M<\infty\), and first also \(J<\infty\). By Tonelli applied to the nonnegative product-measurable \((z,y)\mapsto F(y)^{r-1}g(z,y)\) (measurable since \(F\) is),
\begin{equation*} J = \int_YF^{r-1}\Bigl(\int_Xg(z,y)\,\mu(dz)\Bigr)\nu(dy) = \int_X\Bigl(\int_YF(y)^{r-1}g(z,y)\,\nu(dy)\Bigr)\mu(dz) . \end{equation*}
Holder’s inequality on \((Y,\mathcal{B},\nu)\) with the conjugate exponents \(r/(r-1)\) and \(r\), legitimate because \(\int_Y(F^{r-1})^{r/(r-1)}d\nu=J<\infty\), bounds the inner integral by \(J^{(r-1)/r}(\int_Yg(z,y)^{r}\nu(dy))^{1/r}\); integrating in \(z\) gives \(J\le J^{(r-1)/r}M\). If \(J=0\) there is nothing to prove, and if \(0<J<\infty\) division by \(J^{(r-1)/r}\) gives \(J\le M^{r}\).
For arbitrary \(J\), choose \(X_n\uparrow X\), \(Y_n\uparrow Y\) of finite measure and set \(g_n:=\min(g,n)I_{X_n\times Y_n}\uparrow g\). Then \(F_n(y)\le n\mu(X_n)\) and \(F_n=0\) off \(Y_n\), so \(J_n\le(n\mu(X_n))^{r}\nu(Y_n)<\infty\) and the previous paragraph applies to \(g_n\), giving \(J_n\le M^{r}\) since \(g_n\le g\). Monotone convergence gives \(F_n\uparrow F\), hence \(F_n^{r}\uparrow F^{r}\) and \(J_n\uparrow J\), so \(J\le M^{r}\).
(\(\circ\)) Prove the equalities \[ \frac{1}{\sqrt{2\pi}}\int_{-\infty}^{\infty}\exp\Bigl(-\frac{1}{2}t^2\Bigr)\,dt=1,\qquad \frac{1}{\sqrt{2\pi}}\int_{-\infty}^{\infty}t^2\exp\Bigl(-\frac{1}{2}t^2\Bigr)\,dt=1. \]
Both follow from \(I:=\int_{\mathbb{R}}e^{-x^2}\,dx=\sqrt{\pi}\). Since \(x^2\ge2|x|-1\) gives \(e^{-x^2}\le e\,e^{-2|x|}\), we have \(0<I<\infty\), so the nonnegative continuous \(e^{-x^2-y^2}\) has finite repeated integral \(I^2\) and is therefore \(\lambda_2\)-integrable by Tonelli (Theorem 3.4.5), with \[ \int_{\mathbb{R}^2}e^{-x^2-y^2}\,dx\,dy=I^{2} \] by Fubini (Theorem 3.4.4). The polar map \(F(r,\theta)=(r\cos\theta,r\sin\theta)\) is injective and \(C^1\) on \(U=(0,\infty)\times(0,2\pi)\) with Jacobian \(r>0\), and \(F(U)=\mathbb{R}^2\setminus L\) with \(L\) a half-line, \(\lambda_2(L)=0\); so Theorem 3.7.1 and Theorem 3.4.4 give \[ I^{2}=\int_{U}e^{-r^2}r\,dr\,d\theta=2\pi\cdot\tfrac12=\pi , \] whence \(I=\sqrt{\pi}\). The substitution \(x=t/\sqrt2\) (Corollary 3.6.4) now yields \[ \int_{-\infty}^{\infty}e^{-t^2/2}\,dt=\sqrt2\,I=\sqrt{2\pi}, \] the first equality.
For the second, \(\frac{d}{dt}(-e^{-t^2/2})=t\,e^{-t^2/2}\), so integration by parts on \([-R,R]\) gives \[ \int_{-R}^{R}t^2e^{-t^2/2}\,dt=-2R\,e^{-R^2/2}+\int_{-R}^{R}e^{-t^2/2}\,dt . \] As \(R\to\infty\) the first term vanishes and the second tends to \(\sqrt{2\pi}\); the left side tends to \(\int_{\mathbb{R}}t^2e^{-t^2/2}\,dt\) by monotone convergence. Divide by \(\sqrt{2\pi}\).
(\(\circ\)) Let \(e_1,\dots,e_n\) be a basis in \(\mathbb{R}^n\). Prove that a Lebesgue measurable set \(A\subset\mathbb{R}^n\) has measure zero precisely when it can be written in the following form: \(A=A_1\cup\cdots\cup A_n\), where the sets \(A_j\) are measurable and, for every index \(j\) and every \(x\in\mathbb{R}^n\), the set \(\{t\in\mathbb{R}:\ x+te_j\in A_j\}\) has measure zero on the real line (in other words, the sections of \(A_j\) by the straight lines parallel to \(e_j\) have zero linear measures).
Sufficiency. Fix \(j\), pick a linear isomorphism \(T_j\) with \(T_j\varepsilon_n=e_j\), and put \(C=T_j^{-1}(A_j)\); writing points as \((u,t)\), \(u\in\mathbb{R}^{n-1}\), the identity \(T_j(u,t)=T_j(u,0)+te_j\) gives \[ C^{u}=\{t:\ T_j(u,0)+t\,e_j\in A_j\}, \] which is \(\lambda_1\)-null by hypothesis (with \(x=T_j(u,0)\)). Each \(C_N=C\cap[-N,N]^n\) is measurable of finite measure with null sections, so Corollary 3.4.2 (legitimate since \(\lambda_n=\lambda_{n-1}\otimes\lambda_1\) with both factors \(\sigma\)-finite) gives \(\lambda_n(C_N)=\int\lambda_1((C_N)^{u})\,du=0\); letting \(N\to\infty\), \(\lambda_n( C)=0\), whence \(\lambda_n(A_j)=|\det T_j|\,\lambda_n( C)=0\) by Corollary 3.6.4. A finite union of null sets is null.
Necessity. It suffices to treat the standard basis: with \(S\varepsilon_j=e_j\), Corollary 3.6.4 makes \(S^{\pm1}\) preserve measurability and nullity, and \(S^{-1}(x+te_j)=S^{-1}x+t\varepsilon_j\) transports a decomposition of \(S^{-1}(A)\) into one of \(A\). Induct on \(n\).
(i) \(n=1\): take \(A_1=A\), as \(\{t:x+t\in A\}=A-x\) is null by translation invariance.
(ii) \(n\ge2\): write points as \((y,t)\), \(y\in\mathbb{R}^{n-1}\), and \(A^{y}=\{t:(y,t)\in A\}\). Since \(\lambda_n(A)=0\), Corollary 3.4.2 gives \(\int_{\mathbb{R}^{n-1}}\lambda_1(A^{y})\,dy=0\), so \[ B=\{y:\ A^{y}\ \text{nonmeasurable, or}\ \lambda_1(A^{y})>0\} \] lies in a \(\lambda_{n-1}\)-null set, hence is measurable and null by completeness. Put \(A_n=A\cap((\mathbb{R}^{n-1}\setminus B)\times\mathbb{R})\), measurable because \(B\times\mathbb{R}=\bigcup_N B\times[-N,N]\) is \(\lambda_n\)-null; for \(x=(y,s)\),
\begin{equation*} \{t:\ x+t\varepsilon_n\in A_n\}=\emptyset\ \ (y\in B),\qquad =A^{y}-s\ \ (y\notin B), \end{equation*}
null in both cases. By induction \(B=B_1\cup\cdots\cup B_{n-1}\) with all \(\varepsilon_j\)-sections null; put \(A_j=A\cap(B_j\times\mathbb{R})\), measurable since \(B_j\times\mathbb{R}\) is \(\lambda_n\)-null. As \(x+t\varepsilon_j=(y+t\varepsilon_j,s)\) for \(j\le n-1\), \[ \{t:\ x+t\varepsilon_j\in A_j\}\subset\{t:\ y+t\varepsilon_j\in B_j\}, \] null, hence measurable and null by completeness of \(\lambda_1\). Finally \(\bigcup_{j\le n-1}A_j=A\cap(B\times\mathbb{R})\), so \(A_1\cup\cdots\cup A_n=A\).
(Sierpinski [872]) (i) Show that in the plane (or in the unit square) there exists a Lebesgue nonmeasurable set that meets every straight line parallel to one of the coordinate axes in at most one point.
(ii) Show that in the plane there is a nonmeasurable set whose intersection with every straight line has at most two points.
(i) Take \(A=\{(x_\alpha,y_\alpha):\alpha<\omega(\mathfrak{c})\}\), where \(\{K_\alpha\}_{\alpha<\omega(\mathfrak{c})}\) enumerates the compact subsets of \(Q=[0,1]^2\) of positive plane measure (\(\omega(\mathfrak{c})\) is the least ordinal of cardinality \(\mathfrak{c}\)) and the points \((x_\alpha,y_\alpha)\in K_\alpha\) are chosen by transfinite recursion with all \(x_\alpha\) distinct and all \(y_\alpha\) distinct. Two cardinality facts run the recursion: there are exactly \(\mathfrak{c}\) such compact sets (open sets are unions of rational balls, so at most \(\mathfrak{c}\) closed sets exist; closed balls give \(\mathfrak{c}\)), and every measurable set of positive measure has cardinality \(\mathfrak{c}\) (inner regularity, Cantor-Bendixson, and card \(=\mathfrak{c}\) for a nonempty perfect set in a Polish space). At stage \(\alpha\), with \(P=\{x_\beta\}_{\beta<\alpha}\) and \(R=\{y_\beta\}_{\beta<\alpha}\) of cardinality \(<\mathfrak{c}\), Theorem 3.4.1 gives \[ 0<\lambda_2(K_\alpha)=\int_0^1\lambda_1\bigl((K_\alpha)_x\bigr)\,dx , \] so \(G=\{x:\lambda_1((K_\alpha)_x)>0\}\) has positive measure, hence cardinality \(\mathfrak{c}\); pick \(x_\alpha\in G\setminus P\), and then \(y_\alpha\in(K_\alpha)_{x_\alpha}\setminus R\), legitimate because that section has positive measure, hence cardinality \(\mathfrak{c}\).
Distinctness of the coordinates makes every horizontal and every vertical line meet \(A\) at most once. Were \(A\) measurable, all vertical sections being at most singletons, Corollary 3.4.2 would give \(\lambda_2(A)=0\); then \(\lambda_2(Q\setminus A)=1\), inner regularity would supply a compact \(K=K_\alpha\subset Q\setminus A\) of positive measure, and \((x_\alpha,y_\alpha)\in K_\alpha\cap A\) contradicts \(K_\alpha\subset Q\setminus A\).
(ii) The same recursion over an enumeration \(\{K_\alpha\}\) of all compact subsets of \(\mathbb{R}^2\) of positive measure, with “no two chosen points share a coordinate” replaced by “no three chosen points are collinear”. At stage \(\alpha\) let \(\mathcal{L}\) be the family of lines through two of the earlier points; \(\operatorname{card}\mathcal{L}\le\max(\kappa^2,1)<\mathfrak{c}\), where \(\kappa=\operatorname{card}\{\beta<\alpha\}\) and \(\kappa^2=\kappa\) for infinite \(\kappa\). Choose \(x\in G\) (as in (i), \(\operatorname{card}G=\mathfrak{c}\)) with \(\{x\}\times\mathbb{R}\notin\mathcal{L}\); every \(\ell\in\mathcal{L}\) then meets \(\{x\}\times\mathbb{R}\) in at most one point, so \(\bigcup\mathcal{L}\) hits that line in fewer than \(\mathfrak{c}\) points, whereas \(\operatorname{card}(K_\alpha)_x=\mathfrak{c}\). Pick \(y\in(K_\alpha)_x\) lying on no \(\ell\in\mathcal{L}\) and set \(p_\alpha=(x,y)\in K_\alpha\).
No line carries three points of \(A=\{p_\alpha\}\): if \(p_{\alpha_1},p_{\alpha_2},p_{\alpha_3}\in\ell\) with \(\alpha_1<\alpha_2<\alpha_3\) (smallest representing indices, distinct since the points are), then \(\ell\in\mathcal{L}\) at stage \(\alpha_3\), contradicting \(p_{\alpha_3}\notin\ell\). Nonmeasurability as in (i): vertical sections carry at most two points, so Corollary 3.4.2 gives \(\lambda_2(A\cap[-N,N]^2)=0\) for every \(N\), hence \(\lambda_2(A)=0\), and a compact \(K_\alpha\subset\mathbb{R}^2\setminus A\) of positive measure would contain \(p_\alpha\in A\). Enumerating only the compact subsets of \([0,1]^2\) places \(A\) in the square.
Exercises 3.10.50–3.10.56
Show that there exists a bounded nonnegative function \(f\) on the square \([0,1]\times[0,1]\) such that it is not Lebesgue measurable, but the repeated integrals
\begin{equation*} \int_0^1\!\!\int_0^1 f(x,y)\,dx\,dy \qquad\text{and}\qquad \int_0^1\!\!\int_0^1 f(x,y)\,dy\,dx \end{equation*}
exist and vanish.
Take \(f=I_A\), where \(A\subset[0,1]^2\) is the Lebesgue nonmeasurable set of Exercise 3.10.49(i), which meets every line parallel to a coordinate axis in at most one point. Then \(0\le f\le1\), and \(f\) is not Lebesgue measurable, since \(\{f>1/2\}=A\).
For each \(x\) the section \(A_x\) is at most a singleton, so \(y\mapsto f(x,y)\) is measurable with
\begin{equation*} \int_0^1 f(x,y)\,dy=\lambda_1(A_x)=0 , \end{equation*}
and this inner integral, the zero function of \(x\), integrates to \(0\). Symmetrically every \(A^y\) is at most a singleton, so \(\int_0^1f(x,y)\,dx=0\) for every \(y\) and the other repeated integral vanishes as well.
(Sierpinski) (i) Assuming the continuum hypothesis construct a set \(S\subset[0,1]^2\) such that all its vertical sections are at most countable and all its horizontal sections have at most countable complements. Observe that the repeated integrals of \(I_S\) exist and are different.
(ii) Without use of the continuum hypothesis construct a measurable space \(X\) with a probability measure \(\mu\) and a set \(S\subset X^2\) such that the repeated integrals
\begin{equation*} \int_X\!\int_X I_S(x,y)\,\mu(dx)\,\mu(dy) \qquad\text{and}\qquad \int_X\!\int_X I_S(x,y)\,\mu(dy)\,\mu(dx) \end{equation*}
exist and are not equal.
(iii) Under the continuum hypothesis construct a set \(E\subset[0,1]^2\) such that its indicator function \(I_E\) is measurable in every variable separately, the function
\begin{equation*} x\mapsto\int_0^1 I_E(x,y)\,dy \end{equation*}
is measurable, but the function
\begin{equation*} y\mapsto\int_0^1 I_E(x,y)\,dx \end{equation*}
is not.
(i) Take \(S=\{(x,y)\in[0,1]^2:\ y\prec x\}\), where \(\prec\) is transported from \(\omega_1\) by a bijection \(\iota:[0,1]\to\omega_1\) (available under the continuum hypothesis) via \(u\prec v:\Longleftrightarrow\iota(u)<\iota(v)\); every \(v\) has \(\{u:u\prec v\}\) in bijection with the ordinal \(\iota(v)<\omega_1\), hence at most countably many predecessors. Writing \(S_x=\{y:(x,y)\in S\}\) and \(S^y=\{x:(x,y)\in S\}\), the section \(S_x=\{y:y\prec x\}\) is at most countable and \(S^y=\{x:y\prec x\}\) has the at most countable complement \(\{x:x\prec y\}\cup\{y\}\). Every section is therefore Borel, and null or of full measure, so
\begin{equation*} \int_0^1 I_S(x,y)\,dy=0\ \ (\forall x),\qquad \int_0^1 I_S(x,y)\,dx=1\ \ (\forall y), \end{equation*}
both constant in the outer variable; the repeated integrals are thus \(0\) and \(1\). By Theorem 3.4.1 this forces \(S\) to be nonmeasurable for \(\lambda\otimes\lambda\).
(ii) Take \(X=\omega_1\), \(\mathcal{A}=\{A\subset X:\ A\ \text{or}\ X\setminus A\ \text{is at most countable}\}\), \(\mu(A)=0\) in the first case and \(1\) in the second, and \(S=\{(x,y):x\le y\}\). Here \(\mathcal{A}\) is a \(\sigma\)-algebra: it is complementation-symmetric by definition, and a countable union of its members is at most countable unless some term is co-countable, in which case so is the union. And \(\mu\) is a probability measure: \(\omega_1\) is uncountable, and two disjoint members cannot both be co-countable (else \(A_1\subset X\setminus A_2\) would be countable with countable complement), so for a disjoint sequence either all terms are at most countable, and both sides of additivity vanish, or exactly one is co-countable, and both sides equal \(1\). Now \(S^y=\{x:x\le y\}\) is at most countable and \(S_x=\{y:y\ge x\}\) is co-countable, so all inner integrals exist and are constants, giving
\begin{equation*} \int_X\!\int_X I_S\,\mu(dx)\,\mu(dy)=0,\qquad \int_X\!\int_X I_S\,\mu(dy)\,\mu(dx)=1 . \end{equation*}
(iii) Take \(E=S\cap([0,1]\times D)\), with \(S\) from (i) and \(D\subset[0,1]\) nonmeasurable (a Vitali set). For fixed \(x\), \(E_x=S_x\cap D\) is at most countable, so \(y\mapsto I_E(x,y)\) is measurable and \(\int_0^1I_E(x,y)\,dy=0\), an identically zero, hence measurable, function of \(x\). For fixed \(y\), \(E^y=S^y\) if \(y\in D\) and \(E^y=\varnothing\) otherwise, both measurable, and \(\lambda(S^y)=1\) gives
\begin{equation*} \int_0^1 I_E(x,y)\,dx=I_D(y), \end{equation*}
a nonmeasurable function of \(y\).
(\(\circ\)) Prove that the graph of a measurable real function on a measure space \((X,\mathcal{A},\mu)\) with a finite measure \(\mu\) has measure zero with respect to \(\mu\otimes\lambda\), where \(\lambda\) is Lebesgue measure.
The graph \(\Gamma_f=\{(x,t):t=f(x)\}\) is a \(\mu\otimes\lambda\)-null set. It is product measurable, being \(\{u=v\}\) for the \(\mathcal{A}\otimes\mathcal{B}(\mathbb{R}^1)\)-measurable functions \(u(x,t)=f(x)\) and \(v(x,t)=t\). For \(n\in\mathbb{N}\) let \(\lambda_n\) be the restriction of \(\lambda\) to \([-n,n]\), a finite measure; each section of \(\Gamma_f\cap(X\times[-n,n])\) is the single point \(f(x)\) or empty, so Theorem 3.4.1 (applicable, both \(\mu\) and \(\lambda_n\) being finite) gives
\begin{equation*} (\mu\otimes\lambda_n)\bigl(\Gamma_f\cap(X\times[-n,n])\bigr) =\int_X\lambda_n\bigl((\Gamma_f)_x\bigr)\,\mu(dx)=0 , \end{equation*}
and countable additivity along the increasing union yields \((\mu\otimes\lambda)(\Gamma_f)=0\).
Method (2): a covering argument, without Fubini. Since \(\Gamma_f=\bigcup_k\Gamma_f\cap(X_k\times\mathbb{R}^1)\) with \(X_k=\{|f|\le k\}\uparrow X\), we may assume \(|f|\le M\). For \(n\in\mathbb{N}\) let \(\Delta_i=[i/n,(i+1)/n)\) over the finitely many \(i\) with \(|i|\le Mn+1\); these are disjoint and cover \([-M,M]\), and \(t=f(x)\) puts \((x,t)\) in \(f^{-1}(\Delta_i)\times\Delta_i\) for exactly one \(i\). The sets \(f^{-1}(\Delta_i)\in\mathcal{A}\) are disjoint with union \(X\) and \(\lambda(\Delta_i)=n^{-1}\), so
\begin{equation*} (\mu\otimes\lambda)\Bigl(\bigcup_i f^{-1}(\Delta_i)\times\Delta_i\Bigr) \le\frac1n\sum_i\mu\bigl(f^{-1}(\Delta_i)\bigr)=\frac{\mu(X)}{n}, \end{equation*}
and \(\Gamma_f\) lies inside this product measurable set for every \(n\).
(\(\circ\)) Let \((X,\mathcal{A}_X)\) and \((Y,\mathcal{A}_Y)\) be measurable spaces and let \(f:X\to Y\) be a mapping. Construct examples showing that:
(i) even if \(f\) is \((\mathcal{A}_X,\mathcal{A}_Y)\)-measurable, its graph may not belong to \(\mathcal{A}_X\otimes\mathcal{A}_Y\);
(ii) the graph of \(f\) may belong to \(\mathcal{A}_X\otimes\mathcal{A}_Y\) without \(f\) being measurable.
Prove that if the set \(\{(y,y),\ y\in Y\}\) belongs to \(\mathcal{A}_Y\otimes\mathcal{A}_Y\), then the graph of any \((\mathcal{A}_X,\mathcal{A}_Y)\)-measurable mapping belongs to \(\mathcal{A}_X\otimes\mathcal{A}_Y\). (See also Corollary 6.10.10 in Chapter 6.)
(i) Take \(X=Y=[0,1]\) with \(\mathcal{A}_X=\mathcal{A}_Y=\mathcal{E}\), the \(\sigma\)-algebra of at most countable and co-countable sets (Exercise 3.10.51(ii)), and \(f=\mathrm{id}\), which is \((\mathcal{E},\mathcal{E})\)-measurable since \(f^{-1}(A)=A\). Its graph is the diagonal \(\Delta\), and \(\Delta\notin\mathcal{E}\otimes\mathcal{E}\). Indeed, every member of a product \(\sigma\)-algebra lies in \(\sigma(\mathcal{R}_0)\) for some countable family \(\mathcal{R}_0\) of measurable rectangles: the sets with that property form a \(\sigma\)-algebra (closure under countable unions by uniting the countable families) containing all rectangles. So suppose \(\Delta\in\sigma(\{A_n\times B_n\})\) with \(A_n,B_n\in\mathcal{E}\). Let \(N\) be the union of those \(A_n,B_n\) that are at most countable together with the complements of the co-countable ones; \(N\) is at most countable, so we may pick distinct \(u,v\in[0,1]\setminus N\). Each \(A_n\) and each \(B_n\) either contains \([0,1]\setminus N\) or is disjoint from it, so
\begin{equation*} u\in A_n\iff v\in A_n,\qquad u\in B_n\iff v\in B_n\qquad(n\in\mathbb{N}). \end{equation*}
Hence every \(A_n\times B_n\) lies in the \(\sigma\)-algebra \(\mathcal{D}=\{E:\ (u,u)\in E\iff(u,v)\in E\}\), so \(\Delta\in\mathcal{D}\); but \((u,u)\in\Delta\) and \((u,v)\notin\Delta\).
(ii) Take \(X=[0,1]\) with \(\mathcal{A}_X=\mathcal{B}([0,1])\), \(Y=[0,1]\) with \(\mathcal{A}_Y=\mathcal{L}\) the Lebesgue \(\sigma\)-algebra, and \(f=\mathrm{id}\). It is not \((\mathcal{B},\mathcal{L})\)-measurable: all \(2^{\mathfrak{c}}\) subsets of the Cantor set are Lebesgue measurable, while \(|\mathcal{B}([0,1])|=\mathfrak{c}\), so some \(L\in\mathcal{L}\setminus\mathcal{B}\) has \(f^{-1}(L)=L\notin\mathcal{B}\). Yet its graph \(\Delta\) is product measurable: with \(I_{n,1},\dots,I_{n,2^n}\) the dyadic intervals of length \(2^{-n}\),
\begin{equation*} \Delta=\bigcap_{n=1}^{\infty}\ \bigcup_{k=1}^{2^n} I_{n,k}\times I_{n,k}, \end{equation*}
since lying in a common interval of length \(2^{-n}\) for every \(n\) forces \(x=y\).
Final assertion. Assume \(\Delta_Y=\{(y,y)\}\in\mathcal{A}_Y\otimes\mathcal{A}_Y\) and let \(f\) be measurable. The map \(F(x,y)=(f(x),y)\) satisfies \(F^{-1}(B\times C)=f^{-1}(B)\times C\in\mathcal{A}_X\otimes\mathcal{A}_Y\), and \(\{E:F^{-1}(E)\in\mathcal{A}_X\otimes\mathcal{A}_Y\}\) is a \(\sigma\)-algebra containing all rectangles, hence containing \(\mathcal{A}_Y\otimes\mathcal{A}_Y\). Therefore
\begin{equation*} \Gamma_f=\{(x,y):f(x)=y\}=F^{-1}(\Delta_Y)\in\mathcal{A}_X\otimes\mathcal{A}_Y . \end{equation*}
Show that under the continuum hypothesis the plane can be covered by countably many graphs of functions \(y=y(x)\) and \(x=x(y)\). In particular, there exists a nonmeasurable graph among them.
Under the continuum hypothesis \(|\mathbb{R}|=|\omega_1|\), so as in Exercise 3.10.51(i) a bijection \(\iota:\mathbb{R}\to\omega_1\) transports the ordinal order to a linear order \(\prec\) on \(\mathbb{R}\) in which every point has at most countably many predecessors. Put
\begin{equation*} P(x)=\{y:\ y\preceq x\},\qquad Q(y)=\{x:\ x\prec y\}, \end{equation*}
both at most countable, with \(P(x)\ne\varnothing\) since \(x\in P(x)\). Choose surjections \(n\mapsto f_n(x)\) of \(\mathbb{N}\) onto \(P(x)\), and \(n\mapsto g_n(y)\) of \(\mathbb{N}\) onto \(Q(y)\) when \(Q(y)\ne\varnothing\) (setting \(g_n(y)=y\) otherwise), and let \(\Gamma_n\), \(\Gamma_n^{*}\) be the graphs of \(y=f_n(x)\) and of \(x=g_n(y)\).
These countably many graphs cover the plane: for \((x,y)\) linearity gives either \(x\prec y\), so \(x\in Q(y)\) and \((x,y)\in\Gamma_n^{*}\) for some \(n\), or \(y\preceq x\), so \(y\in P(x)\) and \((x,y)\in\Gamma_n\) for some \(n\).
One of them is nonmeasurable. A Lebesgue measurable graph \(\Gamma\) of a function \(y=h(x)\) is null: every vertical section of \(\Gamma\cap[-m,m]^2\) is at most a singleton, so Theorem 3.4.1, applied to the product of the finite measures \(\lambda|_{[-m,m]}\) with itself, gives
\begin{equation*} \lambda_2\bigl(\Gamma\cap[-m,m]^2\bigr) =\int_{-m}^{m}\lambda_1\bigl(\Gamma_x\cap[-m,m]\bigr)\,dx=0 , \end{equation*}
and \(m\to\infty\); horizontal sections settle a graph \(x=h(y)\) the same way. Were all these graphs measurable, \(\mathbb{R}^2\) would be a countable union of null sets.
(Fichtenholz) There exists a measurable function \(f\) on \([0,1]^2\) such that \(f\) is not integrable, but for all measurable sets \(A,B\subset[0,1]\), the repeated integrals
\begin{equation*} \int_A\!\int_B f(x,y)\,dx\,dy\qquad\text{and}\qquad \int_B\!\int_A f(x,y)\,dy\,dx \end{equation*}
exist, are finite and equal.
Take \(f=\gamma_p\,h^{(2p)}_{ij}\) on \(J_{p,i}\times J_{p,j}\) and \(f=0\) off the blocks \(Q_p:=J_p\times J_p\), where \(J_p=(2^{-p},2^{-p+1}]\), \(\ell_p:=\lambda(J_p)=2^{-p}\), \(N_p:=\gamma_p:=4^{p}\), the \(J_{p,1},\dots,J_{p,N_p}\) are the consecutive subintervals of \(J_p\) of length \(\ell_p/N_p\), and \(H^{(p)}=(h^{(p)}_{ij})\) is the Hadamard matrix defined by \(H^{(0)}=(1)\) and
\begin{equation*} H^{(p)}:=\begin{pmatrix} H^{(p-1)} & H^{(p-1)}\\ H^{(p-1)} & -H^{(p-1)}\end{pmatrix}. \end{equation*}
Induction gives \(H^{(p)}\) symmetric with entries \(\pm1\) and \((H^{(p)})^{\top}H^{(p)}=2^{p}I\) (the block product has diagonal blocks \(2\cdot2^{p-1}I\) and vanishing off-diagonal blocks), hence
\begin{equation*} \|H^{(p)}\beta\|_2=\sqrt{N}\,\|\beta\|_2,\qquad N=2^{p}. \tag{1} \end{equation*}
Thus \(f\) is measurable, being constant on each member of a countable disjoint family of rectangles and zero elsewhere, and \(f(x,y)=f(y,x)\) by symmetry of \(H^{(2p)}\).
\(f\) is not integrable: \(|f|=\gamma_p\) on the disjoint blocks \(Q_p\), of measure \(\ell_p^{2}\), so \(\int_{[0,1]^2}|f|=\sum_p\gamma_p\ell_p^{2}=\sum_p1=+\infty\). All sections are integrable: for \(x\in J_p\) the section \(f(x,\cdot)\) vanishes off \(J_p\) and is bounded by \(\gamma_p\), so \(\int_0^1|f(x,y)|\,dy=\gamma_p\ell_p=2^{p}\), and symmetrically in \(y\).
Key estimate. For measurable \(A\subset[0,1]\) put \(\psi_A(x)=\int_Af(x,y)\,dy\) and \(\beta_j=\lambda(A\cap J_{p,j})\), so that \(0\le\beta_j\le\ell_p/N_p\) and \(\sum_j\beta_j\le\ell_p\), whence
\begin{equation*} \|\beta\|_2^{2}\le\bigl(\max_j\beta_j\bigr)\sum_j\beta_j\le\frac{\ell_p^{2}}{N_p}, \qquad\text{i.e.}\qquad \|\beta\|_2\le\frac{\ell_p}{\sqrt{N_p}} . \tag{2} \end{equation*}
For \(x\in J_{p,i}\) we have \(\psi_A(x)=\gamma_p(H^{(2p)}\beta)_i\), so Cauchy-Schwarz with (1) and (2) gives
\begin{equation*} \int_{J_p}|\psi_A|\,dx =\gamma_p\frac{\ell_p}{N_p}\sum_{i=1}^{N_p}\bigl|(H^{(2p)}\beta)_i\bigr| \le\gamma_p\ell_p\|\beta\|_2\le\frac{\gamma_p\ell_p^{2}}{\sqrt{N_p}}=2^{-p}, \end{equation*}
so \(\int_0^1|\psi_A|\le1\), and \(\psi_A\in L^1[0,1]\). (It is Borel, being constant on each \(J_{p,i}\) and zero at \(0\).) By symmetry the same holds for \(\theta_B(y)=\int_Bf(x,y)\,dx\).
Equality of the repeated integrals. Fix measurable \(A,B\subset[0,1]\). On the block \(Q_p\) the function \(f\) is bounded and \(\lambda_2(Q_p)<\infty\), so \(fI_{Q_p}\) is integrable and Fubini (Theorem 3.4.4) gives
\begin{equation*} \int_{B\cap J_p}\!\!\int_{A\cap J_p}\!\!f\,dy\,dx =\int_{(B\times A)\cap Q_p}\!\!f\,d\lambda_2 =\int_{A\cap J_p}\!\!\int_{B\cap J_p}\!\!f\,dx\,dy=:c_p . \end{equation*}
For \(x\in J_p\) we have \(\psi_A(x)=\int_{A\cap J_p}f(x,y)\,dy\), since \(f(x,\cdot)\) vanishes off \(J_p\); and \(\sum_p|c_p|\le\sum_p\int_{J_p}|\psi_A|\le1\). So countable additivity of the integral of \(\psi_A\in L^1\) over \([0,1]=\{0\}\cup\bigcup_pJ_p\) yields
\begin{equation*} \int_B\!\int_A f\,dy\,dx=\sum_{p\ge1}c_p=\int_A\!\int_B f\,dx\,dy , \end{equation*}
both repeated integrals existing and finite.
Let \(f\) be a Riemann integrable function on \([0,1]^2\).
(i) Prove that for almost every \(x\in[0,1]\), the function \(y\mapsto f(x,y)\) is Riemann integrable and the function \(\varphi:x\mapsto\varphi(x)\), where \(\varphi(x)\) equals the Riemann integral
\begin{equation*} \int_0^1 f(x,y)\,dy \end{equation*}
if it exists and the lower Riemann integral otherwise, is Riemann integrable.
(ii) Prove that if at all points \(x\) where the Riemann integral in \(y\) does not exist, we redefine \(\varphi\) to be zero, then the obtained function may not be Riemann integrable (although it remains Lebesgue integrable and its Lebesgue integral is unchanged).
(i) Let \(|f|\le M\), and let \(\underline{\varphi}(x)\), \(\overline{\varphi}(x)\) be the lower and upper Riemann integrals of the bounded section \(y\mapsto f(x,y)\); then \(\varphi=\underline{\varphi}\), and the section is Riemann integrable exactly where \(\underline{\varphi}(x)=\overline{\varphi}(x)\). Take a grid partition \(P\) of the square with cells \(R_{ij}=[x_{i-1},x_i]\times[y_{j-1},y_j]\), put \(m_{ij}=\inf_{R_{ij}}f\), \(M_{ij}=\sup_{R_{ij}}f\), and write \(L,U\) for the Darboux sums. For \(x\in[x_{i-1},x_i]\), comparison with the Darboux sums of the section along \(\{y_j\}\) gives \(\underline{\varphi}(x)\ge\sum_jm_{ij}\Delta y_j\) and \(\overline{\varphi}(x)\le\sum_jM_{ij}\Delta y_j\); hence, for \(g\) either of \(\underline{\varphi}\), \(\overline{\varphi}\) and \(P_1=\{x_i\}\), using \(\underline{\varphi}\le\overline{\varphi}\),
\begin{equation*} L(f,P)\ \le\ L(g,P_1)\ \le\ U(g,P_1)\ \le\ U(f,P). \tag{4} \end{equation*}
Given \(\varepsilon>0\), the Darboux criterion supplies \(P\) with \(U(f,P)-L(f,P)<\varepsilon\), so \(U(g,P_1)-L(g,P_1)<\varepsilon\): both \(\underline{\varphi}\) and \(\overline{\varphi}\) are Riemann integrable, and by (4) their integrals and \(\int_{[0,1]^2}f\) all lie in \([L(f,P),U(f,P)]\), whence
\begin{equation*} \int_0^1\underline{\varphi}(x)\,dx=\int_0^1\overline{\varphi}(x)\,dx =\int_{[0,1]^2} f . \end{equation*}
Thus \(h:=\overline{\varphi}-\underline{\varphi}\ge0\) is Riemann integrable with \(\int_0^1h=0\), so by Theorem 2.10.1 its Lebesgue integral vanishes as well and \(h=0\) a.e. Hence almost every section is Riemann integrable, and \(\varphi=\underline{\varphi}\) is Riemann integrable.
(ii) Take \(f=1+g\), where \(g(x,y)=1/q\) if \(x=p/q\in\mathbb{Q}\) in lowest terms and \(y\in\mathbb{Q}\), and \(g(x,y)=0\) otherwise. Then \(1\le f\le2\), and \(f\) is continuous at every \((x_0,y_0)\) with \(x_0\) irrational: given \(\varepsilon>0\), pick \(n>1/\varepsilon\) and then \(\delta>0\) so that no rational of denominator at most \(n\) lies within \(\delta\) of \(x_0\); for \(|x-x_0|<\delta\) this gives \(0\le g(x,y)\le1/n<\varepsilon\) for all \(y\), while \(g(x_0,y_0)=0\). So the discontinuity set lies in \((\mathbb{Q}\cap[0,1])\times[0,1]\), of plane measure zero, and \(f\) is Riemann integrable by the Lebesgue-Vitali criterion (Exercise 2.12.38(i)), with \(\int_{[0,1]^2}f=1\) since \(f=1\) a.e. (Theorem 2.10.1).
For irrational \(x\) the section is identically \(1\); for \(x=p/q\) in lowest terms it is \(1+q^{-1}I_{\mathbb{Q}}\), with infimum \(1\) and supremum \(1+1/q\) on every nondegenerate interval, so its lower integral is \(1\), its upper integral \(1+1/q\), and the Riemann integral does not exist. Hence \(\varphi\equiv1\). Redefining \(\varphi\) to be \(0\) at the rational \(x\) leaves
\begin{equation*} \widetilde{\varphi}=I_{[0,1]\setminus\mathbb{Q}}, \end{equation*}
whose upper and lower Darboux sums are \(1\) and \(0\) for every partition (\(\mathbb{Q}\) and its complement being dense), so it is not Riemann integrable; but it is bounded and measurable, and equals \(\varphi\) off the null set \(\mathbb{Q}\cap[0,1]\), so \(\int_0^1\widetilde{\varphi}\,d\lambda=1\) is unchanged.
Exercises 3.10.57–3.10.63
(Fichtenholz, Lichtenstein) Let \(f\) be a bounded function on the square \([0,1]\times[0,1]\) such that, for every fixed \(y\), the function \(x\mapsto f(x,y)\) is Riemann integrable, and, for every fixed \(x\), the function \(y\mapsto f(x,y)\) is Lebesgue integrable.
(i) Prove that the function
\begin{equation*} F_1(x)=\int_0^1 f(x,y)\,dy \end{equation*}
is Riemann integrable, the function
\begin{equation*} F_2(y)=\int_0^1 f(x,y)\,dx \end{equation*}
is Lebesgue integrable, and their respective integrals are equal.
(ii) Prove that if the function \(y\mapsto f(x,y)\) also is Riemann integrable for every \(x\), then the repeated Riemann integrals of \(f\) exist and are equal. Note, however, that in this situation \(f\) may not be Lebesgue integrable over the square.
(i) With \(M:=\sup|f|\) and \(J:=\int_0^1F_2(y)\,dy\), the function \(F_2\) is Lebesgue integrable, \(F_1\) is Riemann integrable, and \(\int_0^1F_1=J\). Indeed, \(S_n(y)=n^{-1}\sum_{k\le n}f(k/n,y)\) is Lebesgue measurable, being a finite combination of the measurable sections \(f(k/n,\cdot)\), and it is the Riemann sum of the Riemann integrable \(x\mapsto f(x,y)\) for the uniform partition of mesh \(1/n\to0\) with right-endpoint tags, so \(S_n(y)\to F_2(y)\) for every \(y\). Hence \(F_2\) is measurable, and \(|F_2|\le M\) makes it integrable.
For a tagged partition \(0=a_1<\cdots<a_{m+1}=1\) with tags \(x_i\) and increments \(\Delta_i\), the function \(T(y)=\sum_if(x_i,y)\Delta_i\) is Lebesgue integrable with \(|T|\le M\), and linearity of the integral gives the key identity
\begin{equation*} \int_0^1 T(y)\,dy=\sum_{i}\Delta_i\int_0^1f(x_i,y)\,dy=\sum_{i}F_1(x_i)\,\Delta_i , \end{equation*}
the Riemann sum of \(F_1\) for that partition. Take any sequence of tagged partitions of mesh tending to \(0\), with associated functions \(T_n\). For each fixed \(y\) the number \(T_n(y)\) is a Riemann sum of the Riemann integrable \(x\mapsto f(x,y)\), so \(T_n(y)\to F_2(y)\), while \(|T_n|\le M\), which is integrable on \([0,1]\). Dominated convergence gives \(\int_0^1T_n\to J\), i.e. the Riemann sums of the bounded function \(F_1\) converge to \(J\) along every such sequence; by the sequential criterion for Riemann integrability,
\begin{equation*} \int_0^1F_1(x)\,dx=J=\int_0^1F_2(y)\,dy . \end{equation*}
(ii) Apply (i) to \(g(u,v):=f(v,u)\): for fixed \(v\) the map \(u\mapsto f(v,u)\) is Riemann integrable by the extra assumption, and for fixed \(u\) the map \(v\mapsto f(v,u)\) is Riemann integrable by the original assumption, hence bounded and Lebesgue integrable with the same integral (Theorem 2.10.1). So \(u\mapsto\int_0^1g(u,v)\,dv=F_2(u)\) is Riemann integrable too, and Theorem 2.10.1 identifies \(F_1(x)\) and \(F_2(y)\) with the Riemann integrals of the respective sections. Both repeated Riemann integrals therefore exist and, by (i), equal \(J\).
Lebesgue integrability may nevertheless fail. Take \(f=I_A\), where \(A\subset[0,1]^2\) is the nonmeasurable set of Exercise 3.10.49(i), meeting every line parallel to a coordinate axis at most once. Each section vanishes off at most one point, hence is Riemann integrable with integral \(0\); so \(f\) satisfies all hypotheses of (ii) and both repeated Riemann integrals are \(0\), while \(\{f=1\}=A\) is not Lebesgue measurable.
Let \(X=Y=[0,1]\), let \(\lambda^*\) be Lebesgue outer measure, and let \(\nu^*(A)\) be the cardinality of a set \(A\). Show that the diagonal \(D\) of the square \([0,1]^2\) is measurable with respect to \(\lambda^*\times\nu^*\) in the sense of Theorem 3.10.1, but the repeated integrals of \(I_D\) against \(d\nu^*\,d\lambda^*\) and \(d\lambda^*\,d\nu^*\) equal, respectively, \(1\) and \(0\).
The diagonal is measurable and the two repeated integrals are \(1\) and \(0\). Both \(\lambda^*\) and \(\nu^*\) are Caratheodory outer measures in the sense of Definition 1.11.1: cardinality vanishes on \(\varnothing\) and is monotone and countably subadditive. Moreover \(M_{\nu^*}=2^{[0,1]}\), since cardinality is additive on the disjoint decomposition \(E=(E\cap A)\cup(E\setminus A)\) for arbitrary \(A,E\subset[0,1]\); so \(\nu\) is counting measure on all subsets, while \(M_{\lambda^*}\) is the Lebesgue \(\sigma\)-algebra and \(\lambda\) is Lebesgue measure.
Measurability. By the last assertion of Theorem 3.10.1, \(A\times B\in M_{\lambda^*\times\nu^*}\) whenever \(A\in M_{\lambda^*}\) and \(B\in M_{\nu^*}=2^{[0,1]}\); in particular every open rectangle is measurable. The rectangles with rational-endpoint sides form a countable base of \([0,1]^2\), so every open set lies in the \(\sigma\)-algebra \(M_{\lambda^*\times\nu^*}\) (Theorem 1.11.4), which therefore contains the Borel \(\sigma\)-algebra. The diagonal \(D\) is closed, hence \(\lambda^*\times\nu^*\)-measurable.
The repeated integrals. The sections are singletons, \(D_x=\{x\}\) and \(D^y=\{y\}\), so
\begin{equation*} \int_Y I_D(x,y)\,d\nu(y)=\nu(\{x\})=1,\qquad \int_X I_D(x,y)\,d\lambda(x) =\lambda(\{y\})=0 . \end{equation*}
The first is the constant \(1\), Lebesgue measurable in \(x\); the second is the constant \(0\), \(\nu\)-measurable in \(y\) (every function is, since \(M_{\nu^*}=2^{[0,1]}\)). Integrating, the repeated integrals are \(\lambda([0,1])=1\) and \(0\).
(i) (Davies) Let \(E\subset\mathbb{R}^2\) be a Lebesgue measurable set of finite measure. Then, there exists a family \(L\) of straight lines in \(\mathbb{R}^2\) such that the union of all these lines is measurable and has the same measure as \(E\) and every point of \(E\) belongs to at least one line from \(L\). A multidimensional analog is obtained in Falconer.
(ii) (Csornyei) Prove that the assertion analogous to (i) is true for every \(\sigma\)-finite Borel measure on the plane.
Both parts reduce, by completeness of the measure, to choosing a line \(\ell_p\) through each \(p\in E\) so that \(\bigcup_{p\in E}\ell_p\setminus E\) is null: if the excess is null, then \(\bigcup L=E\cup(\bigcup L\setminus E)\) is measurable with \(\lambda_2(\bigcup L)=\lambda_2(E)\); conversely a measurable \(\bigcup L\supset E\) of the same finite measure has null excess. (Measurability of \(\bigcup L\) is not automatic for an arbitrary family of lines; it is forced here by the smallness of the excess.)
A cone over a null set of directions is null: for \(\Theta\subset[0,\pi)\) with \(\lambda_1(\Theta)=0\) and \(p=0\), the set \(C\setminus\{0\}\) is the image under \(\Phi(r,\theta)=(r\cos\theta,r\sin\theta)\) of \((0,\infty)\times\Theta^{\prime}\), where \(\Theta^{\prime}=\Theta\cup(\Theta+\pi)\) is null, so that domain has plane measure zero, and \(\Phi\) has locally bounded Jacobian \(r\), hence carries null sets to null sets (Theorem 3.7.1). So (i) is immediate for any \(E\) inside such a cone; and since a single line is null, \(L\) must be uncountable as soon as \(\lambda_2(E)>0\).
(i) follows from two facts of the cited papers:
(N) (Nikodym, Davies) there is a \(\lambda_2\)-null set \(M\subset\mathbb{R}^2\) such that every \(p\notin M\) lies on a line \(\ell_p\) with \(\ell_p\setminus\{p\}\subset M\);
(D) (Davies) every plane set of measure zero is contained in the union of a family of lines whose union is again of measure zero.
Granting these, split \(E=E_1\cup E_2\) with \(E_1=E\setminus M\) and \(E_2=E\cap M\). The lines of (N) through the points of \(E_1\) satisfy \(\bigcup_{p\in E_1}\ell_p\subset E_1\cup M\), and (D) gives \(L^{\prime}\) with \(E_2\subset\bigcup L^{\prime}\) and \(\lambda_2(\bigcup L^{\prime})=0\). For \(L=\{\ell_p:p\in E_1\}\cup L^{\prime}\) the excess \(\bigcup L\setminus E\subset M\cup\bigcup L^{\prime}\) is null, which by the reformulation is (i).
(ii) The reformulation used only completeness of the completed measure, so it applies verbatim to any \(\sigma\)-finite Borel \(\mu\). One case is elementary: if \(\mu(\mathbb{R}^2\setminus\ell_0)=0\), take through each \(p\in E\cap\ell_0\) any line other than \(\ell_0\), so that \(\ell_p\cap\ell_0=\{p\}\), and through each \(q\in E\setminus\ell_0\) the parallel to \(\ell_0\), so that \(\ell_q\cap\ell_0=\varnothing\); then \((\bigcup L)\cap\ell_0=E\cap\ell_0\subset E\), so the excess lies in the \(\mu\)-null Borel set \(\mathbb{R}^2\setminus\ell_0\). In general one needs the analogues
(N\(_\mu\)) a \(\mu\)-null set \(M_\mu\) such that every \(p\notin M_\mu\) lies on a line \(\ell_p\) with \(\ell_p\setminus\{p\}\subset M_\mu\);
(D\(_\mu\)) every \(\mu\)-null set is covered by a family of lines whose union is \(\mu\)-null;
and then the deduction given for (i) goes through word for word with \(\lambda_2\) replaced by \(\mu\).
The deep ingredients are quoted, not proved: Example 1.12.25 of the book yields only the bounded prototype of (N), the line \(\ell_p\) being uncontrolled outside the unit square, and (N), (D) and their \(\mu\)-analogues are the Besicovitch-type constructions of Davies [206], [208] and Csornyei [195]. (A fully rigorous proof of the remaining step is beyond the scope of this page; see the reference given in the book.)
(Falconer) Let \(A\) be a set of Lebesgue measure zero in \(\mathbb{R}^n\) and let \(1<k<n\). Denote by \(G_{n,k}\) the space of all \(k\)-dimensional linear subspaces in \(\mathbb{R}^n\) equipped with its natural measure (see Federer; for the purposes of this exercise it suffices to embed \(G_{n,k}\) into \(\mathbb{R}^{kn}\) and consider the corresponding measure). Prove that, for almost all \(\Pi\in G_{n,k}\), all sections of \(A\) by the planes parallel to \(\Pi\) have \(k\)-dimensional measure zero.
In coordinates adapted to \(\mathbb{R}^n=\Pi\oplus\Pi^{\perp}\) one has \(\lambda_n=\lambda_k\otimes\lambda_{n-k}\), so Theorem 3.4.1 applied to a measurable hull of \(A\) gives, for every \(\Pi\in G_{n,k}\),
\begin{equation*} \lambda_k\bigl(A\cap(\Pi+a)\bigr)=0 \qquad\text{for }\lambda_{n-k}\text{-a.e. }a\in\Pi^{\perp}. \tag{1} \end{equation*}
Here \(\Pi+a\), \(a\in\Pi^{\perp}\), parametrizes the \(k\)-planes parallel to \(\Pi\), and \(\gamma\) is the rotation-invariant probability on \(G_{n,k}\). The content of the exercise is the upgrade of “a.e. \(a\)” to “every \(a\)”, at the price of a \(\gamma\)-null set of directions.
Reductions. As \(\lambda_k(A\cap(\Pi+a))\le\sum_j\lambda_k(A_j\cap(\Pi+a))\) for \(A_j=A\cap B(0,j)\), we may assume \(A\subset B(0,R)\). By outer regularity choose bounded open \(U_1\supset U_2\supset\cdots\supset A\) inside \(B(0,R+1)\) with \(\lambda_n(U_m)\le2^{-m}\) and put
\begin{equation*} \Phi_m(\Pi):=\sup_{a\in\Pi^{\perp}}\lambda_k\bigl(U_m\cap(\Pi+a)\bigr) . \end{equation*}
Since \(\lambda_k(A\cap(\Pi+a))\le\Phi_m(\Pi)\) for all \(a\), it suffices to prove
\begin{equation*} \lim_{m\to\infty}\Phi_m(\Pi)=0\qquad\text{for }\gamma\text{-a.e. }\Pi. \tag{2} \end{equation*}
Each \(\Phi_m\) is Borel: for open \(U\) the map \((\Pi,a)\mapsto\lambda_k(U\cap(\Pi+a))\) is lower semicontinuous (transport by rotations \(T_j\to I\) with \(T_j\Pi=\Pi_j\) and apply Fatou to the indicators on a fixed copy of \(\mathbb{R}^k\)), and a supremum of lower semicontinuous functions is lower semicontinuous.
The hypothesis \(k>1\) is essential: a Nikodym set of \(\lambda_n\)-measure zero contains a whole line in every direction, so for \(k=1\) statement (2) fails for every \(\Pi\in G_{n,1}\).
Case \(2k>n\), proved in full. This covers hyperplane sections \(k=n-1\) for every \(n\ge3\), hence the whole exercise for \(n=3\) (Marstrand’s theorem).
(a) Slice Fourier identity. For \(f\in L^1(\mathbb{R}^n)\) put \(g_\Pi(a)=\int_\Pi f(a+p)\,d\lambda_k(p)\). The isometry \((p,a)\mapsto p+a\) of \(\Pi\times\Pi^{\perp}\) onto \(\mathbb{R}^n\) carries \(\lambda_k\otimes\lambda_{n-k}\) to \(\lambda_n\), so for \(f\ge0\) Tonelli gives \(g_\Pi\in L^1(\Pi^{\perp})\) with \(\|g_\Pi\|_{L^1}=\|f\|_{L^1}\), and Fubini together with \(\langle\xi,a+p\rangle=\langle\xi,a\rangle\) for \(\xi\in\Pi^{\perp}\), \(p\in\Pi\), gives
\begin{equation*} \widehat{g_\Pi}(\xi)=\hat f(\xi),\qquad \xi\in\Pi^{\perp}, \end{equation*}
the transform on the left taken in \(\Pi^{\perp}\).
(b) Supremum bound. For bounded open \(U\), \(f=I_U\) and \(I(\Pi):=\int_{\Pi^{\perp}}|\hat f|\,d\lambda_{n-k}\),
\begin{equation*} \sup_{a\in\Pi^{\perp}}\lambda_k\bigl(U\cap(\Pi+a)\bigr) \le(2\pi)^{k-n}\,I(\Pi). \tag{3} \end{equation*}
Only \(I(\Pi)<\infty\) needs proof: then \(g_\Pi\in L^1(\Pi^{\perp})\) has integrable Fourier transform, so by the inversion theorem (Section 3.8, in the present normalization) it agrees \(\lambda_{n-k}\)-a.e. with a continuous \(h\) satisfying \(\|h\|_\infty\le(2\pi)^{k-n}I(\Pi)\); the set \(\{g_\Pi=h\}\) has full measure, hence is dense, so taking \(a_j\to a\) inside it and using the lower semicontinuity of \(g_\Pi\) gives \(g_\Pi(a)\le\liminf_jg_\Pi(a_j)=h(a)\).
(c) Integral-geometric identity. For every nonnegative Borel \(\varphi\) on \(\mathbb{R}^n\),
\begin{equation*} \int_{G_{n,k}}\!\int_{\Pi^{\perp}}\varphi\,d\lambda_{n-k}\,d\gamma(\Pi) =c_{n,k}\int_{\mathbb{R}^n}\varphi(\xi)\,|\xi|^{-k}\,d\xi,\qquad c_{n,k} =\frac{s_{n-k}}{s_n}, \tag{4} \end{equation*}
where \(s_d\) is the surface measure of the unit sphere of \(\mathbb{R}^d\). Realize \(\gamma\) as the image of the Haar probability \(\theta\) of \(SO(n)\) under \(\rho\mapsto\rho\Pi_0\); then the left side is the integral of \(\varphi\) against
\begin{equation*} m(B)=\int_{SO(n)}\lambda_{n-k}\bigl(V_0\cap\rho^{-1}B\bigr)\,d\theta(\rho),\qquad V_0 =\Pi_0^{\perp} \end{equation*}
(Tonelli for the measurability in \(\rho\), monotone convergence for countable additivity). Polar coordinates in \(\rho V_0\) evaluate \(m\) on a polar rectangle \(E=\{r\omega:r\in[r_1,r_2],\ \omega\in\Omega\}\) as \(\tau(\Omega)\int_{r_1}^{r_2}r^{n-k-1}dr\), with \(\tau\) a rotation-invariant Borel measure on \(S^{n-1}\) of total mass \(s_{n-k}\); hence \(\tau=(s_{n-k}/s_n)\sigma\), because a compact group acting continuously and transitively on a compact metric space admits exactly one invariant Borel probability (averaging a continuous function over the group is constant, and any invariant probability integrates the function to that constant). The right side gives \(E\) the value \(\sigma(\Omega)\int_{r_1}^{r_2}r^{n-k-1}dr\). Polar rectangles form a \(\pi\)-system generating \(\mathcal{B}(\mathbb{R}^n\setminus\{0\})\), both measures are finite on annuli and kill \(\{0\}\), so they coincide; monotone convergence extends (4) from indicators to all nonnegative Borel \(\varphi\).
(d) Conclusion. Apply (a)-(c) to \(f_m=I_{U_m}\) with \(\varphi=|\hat f_m|\), writing \(I_m(\Pi)\) for the corresponding integral. Splitting at \(|\xi|=1\), bounding \(|\hat f_m|\le\|f_m\|_{L^1}=\lambda_n(U_m)\) inside and using Cauchy-Schwarz with Parseval, \(\|\hat f_m\|_{L^2}^2=(2\pi)^n\lambda_n(U_m)\) (Section 3.8), outside,
\begin{equation*} c_{n,k}^{-1}\int_{G_{n,k}}I_m\,d\gamma \le\lambda_n(U_m)\int_{|\xi|\le1}|\xi|^{-k}\,d\xi +(2\pi)^{n/2}\lambda_n(U_m)^{1/2}\Bigl(\int_{|\xi|>1}|\xi|^{-2k}\,d\xi\Bigr)^{1/2}. \end{equation*}
The first integral is finite because \(k<n\), the second precisely because \(2k>n\), and \(\lambda_n(U_m)\le2^{-m}\), so the right side tends to \(0\). Fatou then gives \(\liminf_mI_m(\Pi)=0\) for \(\gamma\)-a.e. \(\Pi\), and for such \(\Pi\), the sequence \(\Phi_m(\Pi)\) being nonincreasing, (3) yields \(\lim_m\Phi_m(\Pi)=0\), which is (2).
The range \(2\le k\le n/2\). Here (2) is quoted from Falconer’s continuity theorem for the \(k\)-plane transform: for \(2\le k\le n-1\) and nonnegative Borel \(f\in L^p(\mathbb{R}^n)\) with \(p>n/k\), for \(\gamma\)-a.e. \(\Pi\) the function \(a\mapsto\int_{\Pi+a}f\,d\lambda_k\) is finite for all \(a\in\Pi^{\perp}\) and continuous there. Taking \(f=I_G\) for a \(G_\delta\) hull \(G\supset A\) with \(\lambda_n(G)=0\), that function is continuous and, by (1), vanishes a.e., hence identically, so \(\lambda_k(A\cap(\Pi+a))=0\) for all \(a\), for \(\gamma\)-a.e. \(\Pi\). The elementary \(L^1\)-\(L^2\) argument above cannot reach it, since \(\int_{|\xi|>1}|\xi|^{-2k}\,d\xi\) diverges when \(2k\le n\). (A fully rigorous proof of the remaining step is beyond the scope of this page; see the reference given in the book.)
(Talagrand) Let \((X,\mathcal{A},\mu)\) and \((Y,\mathcal{B},\nu)\) be probability spaces and let \(E\in\mathcal{A}\otimes\mathcal{B}\), \(\mu\otimes\nu(E)=\varepsilon>0\). Show that there exists a set \(A\in\mathcal{A}\) with the following property: \(\mu(A)>0\) and for every \(k\in\mathbb{N}\) there exists \(\varepsilon_k>0\) such that
\begin{equation*} \nu\Bigl(\bigcap_{i=1}^{k}E_{x_i}\Bigr)\ge\varepsilon_k \end{equation*}
for all \(x_1,\dots,x_k\in A\), where \(E_x:=\{y:\ (x,y)\in E\}\).
Take \(A=B\setminus\bigcup_{m\ge1}\{x\in B:\ \nu(C_m\setminus E_x)>\nu(C_m)/(4m)\}\) with \(B\), \(C_m\) built below, and \(\varepsilon_k=\tfrac34\nu(C_k)\). Sections are measurable by Proposition 3.3.2, and Theorem 3.4.1, applied to \((A\times C)\cap E\), gives for \(A\in\mathcal{A}\), \(C\in\mathcal{B}\)
\begin{equation*} \int_A\nu(C\cap E_x)\,d\mu(x)=\int_C\mu(A\cap E^y)\,d\nu(y), \qquad \int_X\nu(E_x)\,d\mu=\varepsilon . \tag{1} \end{equation*}
Separability. The sets lying in a \(\sigma\)-algebra generated by countably many measurable rectangles form a \(\sigma\)-algebra containing all rectangles, so \(E\in\mathcal{A}_0\otimes\mathcal{B}_0\) with \(\mathcal{A}_0=\sigma(\{A_n\})\), \(\mathcal{B}_0=\sigma(\{B_n\})\) countably generated; in particular \(E^y\in\mathcal{A}_0\) for every \(y\). Then \(L^1(X,\mathcal{A}_0,\mu)\) is separable: the countable algebra \(\mathcal{R}\) generated by the \(A_n\) approximates every \(S\in\mathcal{A}_0\) in measure, because \(\{S:\inf_{R\in\mathcal{R}}\mu(S\triangle R)=0\}\) contains \(\mathcal{R}\) and is a \(\sigma\)-algebra (it is complementation-stable as \(\mu(S^c\triangle R^c)=\mu(S\triangle R)\), and stable under countable unions since \(\mu(\bigcup_{j\le m}S_j\triangle\bigcup_{j\le m}R_j)\le\sum_{j\le m}\mu(S_j\triangle R_j)\)), whence rational combinations of indicators from \(\mathcal{R}\) are dense.
Hence \(\Phi:Y\to L^1(X,\mathcal{A}_0,\mu)\), \(\Phi(y)=I_{E^y}\), is \(\mathcal{B}\)-measurable: by separability it suffices that \(y\mapsto\|\Phi(y)-h\|_{L^1}\) be measurable for simple \(h=\sum_{j\le N}a_jI_{D_j}\) with \(\{D_j\}\) a measurable partition, and
\begin{equation*} \|I_{E^y}-h\|_{L^1}=\sum_{j\le N}\bigl(|1-a_j|\,\mu(D_j\cap E^y) +|a_j|\,\mu(D_j\setminus E^y)\bigr), \end{equation*}
which is measurable in \(y\) by (1).
For \(\mu(A)>0\) put \(q(A)=\operatorname*{ess\,sup}_y\mu(A\cap E^y)/\mu(A)\in[0,1]\), the essential supremum taken for \(\nu\); by its definition
\begin{equation*} \nu\bigl(\{y:\ \mu(A\cap E^y)/\mu(A)>q(A)-\tau\}\bigr)>0\qquad(\tau>0), \tag{2} \end{equation*}
and \(q(X)\ge\int_Y\mu(E^y)\,d\nu=\varepsilon>0\).
Improvement Lemma. If \(\mu(A)>0\) and \(c:=q(A)>0\), then for all \(\tau\in(0,c)\) and \(\epsilon>0\) there is \(A^{\prime}\subset A\) in \(\mathcal{A}\) with \(\mu(A^{\prime})\ge(c-\tau)\mu(A)\) and \(q(A^{\prime})\ge1-\epsilon\). Indeed \(C=\{y:\mu(A\cap E^y)>(c-\tau)\mu(A)\}\) has \(\nu( C)>0\) by (2); put \(r=\epsilon(c-\tau)\mu(A)>0\) and cover the separable space \(L^1(\mathcal{A}_0)\) by countably many balls of radius \(r/2\), so that \(C^{\prime}=C\cap\Phi^{-1}(B_i)\) has \(\nu(C^{\prime})>0\) for some \(i\), and \(\mu(E^{y}\triangle E^{y^{\prime}})<r\) for \(y,y^{\prime}\in C^{\prime}\). Fix \(y_0\in C^{\prime}\) and set \(A^{\prime}=A\cap E^{y_0}\), so \(\mu(A^{\prime})>(c-\tau)\mu(A)>0\); for \(y\in C^{\prime}\), \(\mu(A^{\prime}\setminus E^{y})\le\mu(E^{y_0}\triangle E^{y})<r\), whence
\begin{equation*} \frac{\mu(A^{\prime}\cap E^{y})}{\mu(A^{\prime})}\ge1-\frac{r}{(c-\tau)\mu(A)} =1-\epsilon , \end{equation*}
and \(\nu(C^{\prime})>0\) gives \(q(A^{\prime})\ge1-\epsilon\).
(i) A set with \(q=1\). With \(\epsilon_m=2^{-m-4}\) and \(c_0=q(X)\ge\varepsilon\), apply the lemma to \(X\) with \(\tau=c_0/2\), \(\epsilon=\epsilon_1\), obtaining \(A_1\) with \(\mu(A_1)\ge c_0/2\) and \(q(A_1)\ge1-\epsilon_1\), then inductively to \(A_{m-1}\) with \(\tau=\epsilon_{m-1}\), \(\epsilon=\epsilon_m\), obtaining \(A_m\subset A_{m-1}\) with \(\mu(A_m)\ge(1-2\epsilon_{m-1})\mu(A_{m-1})\) and \(q(A_m)\ge1-\epsilon_m\). Since \(\prod_m(1-2\epsilon_m)\ge1-2\sum_m\epsilon_m=\tfrac78\), we get \(\mu(A_m)\ge c_0/4\) for all \(m\), so \(B:=\bigcap_mA_m\) has \(\mu(B)=\lim_m\mu(A_m)\ge c_0/4>0\). For \(y\) in \(C^{\prime}_m=\{y:\mu(A_m\cap E^y)>(1-2\epsilon_m)\mu(A_m)\}\), of positive \(\nu\)-measure by (2), we get \(\mu(B\setminus E^y)\le\mu(A_m\setminus E^y)<2\epsilon_m\), hence \(q(B)\ge1-8\epsilon_m/c_0\) for every \(m\), i.e. \(q(B)=1\).
(ii) Removing the bad points. Put \(\lambda_m=1/(4m)\), \(\eta_m=\lambda_m2^{-m-1}\) and \(C_m=\{y:\mu(B\setminus E^y)\le\eta_m\mu(B)\}\), which has \(\nu(C_m)>0\) by (2) since \(q(B)=1\). By (1),
\begin{equation*} \int_B\nu(C_m\setminus E_x)\,d\mu(x)=\int_{C_m}\mu(B\setminus E^y)\,d\nu(y) \le\eta_m\,\mu(B)\,\nu(C_m), \end{equation*}
so Markov’s inequality bounds \(\mu(\{x\in B:\nu(C_m\setminus E_x)>\lambda_m\nu(C_m)\})\) by \(2^{-m-1}\mu(B)\). Deleting these sets from \(B\) leaves \(A\in\mathcal{A}\) with \(\mu(A)\ge\mu(B)(1-\sum_m2^{-m-1})=\mu(B)/2>0\) and
\begin{equation*} \nu(C_m\setminus E_x)\le\lambda_m\nu(C_m)=\frac{\nu(C_m)}{4m} \qquad(m\in\mathbb{N},\ x\in A). \tag{3} \end{equation*}
(iii) Conclusion. Fix \(k\) and let \(x_1,\dots,x_k\in A\). Then (3) with \(m=k\) gives
\begin{equation*} \nu\Bigl(\bigcap_{i=1}^{k}E_{x_i}\Bigr) \ge\nu(C_k)-\sum_{i=1}^{k}\nu\bigl(C_k\setminus E_{x_i}\bigr) \ge\nu(C_k)-k\cdot\frac{\nu(C_k)}{4k}=\tfrac34\nu(C_k)=\varepsilon_k>0 . \end{equation*}
(Erdos, Oxtoby) Let \((X_1,\mathcal{A}_1,\mu_1)\) and \((X_2,\mathcal{A}_2,\mu_2)\) be probability spaces with atomless measures. Show that there exists a set \(A\in\mathcal{A}_1\otimes\mathcal{A}_2\) such that \(\mu_1\otimes\mu_2(A)>0\) and if \(A_i\in\mathcal{A}_i\) and \(\mu_1(A_1)\mu_2(A_2)>0\), then \(\mu_1\otimes\mu_2\bigl((A_1\times A_2)\setminus A\bigr)>0\).
Take \(A=(T_1\times T_2)^{-1}(B)\), where \(T_i:X_i\to[0,1]\) push \(\mu_i\) forward to Lebesgue measure and \(B\subset[0,1)^2\) is the Borel set constructed below, with
\begin{equation*} \lambda_2(B)\ge\tfrac34\quad\text{and} \quad\lambda_2\bigl((S\times T)\setminus B\bigr)>0\ \text{ if }\lambda(S)\lambda(T)>0. \tag{\(\star\)} \end{equation*}
The maps \(T_i\). Every atomless probability space carries a measurable \(T\) with \(\mu\circ T^{-1}=\lambda\): by Sierpinski’s theorem (Corollary 1.12.10) split \(X=X_0\sqcup X_1\) into halves and then each \(X_s\) of length \(|s|=n\) and measure \(2^{-n}\) into \(X_{s0}\sqcup X_{s1}\), and set \(T(x)=\sum_ns_n(x)2^{-n}\) along the branch through \(x\). Each digit \(s_n\) is measurable, and \(T^{-1}\) of a dyadic interval of rank \(n\) is the corresponding \(X_s\) up to the null set where the expansion is ambiguous, so \(\mu\circ T^{-1}=\lambda\) on dyadic intervals and hence on \(\mathcal{B}([0,1])\).
Since \((T_1\times T_2)\) carries \(\mu_1\otimes\mu_2\) to \(\lambda\otimes\lambda\), we get \(\mu_1\otimes\mu_2(A)=\lambda_2(B)>0\). Suppose \(\mu_1(A_1)\mu_2(A_2)>0\) and \(\mu_1\otimes\mu_2((A_1\times A_2)\setminus A)=0\). The images \(\rho_i\) of \(\mu_i|_{A_i}\) under \(T_i\) satisfy \(\rho_i\le\lambda\), so \(\rho_i=g_i\cdot\lambda\) with Borel \(0\le g_i\le1\) and \(\int_0^1g_i=\mu_i(A_i)>0\); the image of \((\mu_1|_{A_1})\otimes(\mu_2|_{A_2})\) being \(\rho_1\otimes\rho_2\),
\begin{equation*} 0=(\rho_1\otimes\rho_2)\bigl([0,1]^2\setminus B\bigr) =\int_{[0,1]^2\setminus B}g_1(s)g_2(t)\,d\lambda_2(s,t). \end{equation*}
With \(S=\{g_1>0\}\) and \(T=\{g_2>0\}\), both of positive measure, \(g_1g_2>0\) on \(S\times T\) forces \(\lambda_2((S\times T)\setminus B)=0\), contradicting (\(\star\)).
A combinatorial lemma. For \(\epsilon,\delta\in(0,1]\) there are arbitrarily large \(n\) and a set \(\mathcal{Z}\subset\mathcal{D}_n\times\mathcal{D}_n\), where \(\mathcal{D}_n\) is the family of the \(N=2^n\) dyadic intervals of rank \(n\), with \(|\mathcal{Z}|\le2\epsilon N^2\) and \((\mathcal{I}\times\mathcal{J})\cap\mathcal{Z}\ne\varnothing\) whenever \(|\mathcal{I}|,|\mathcal{J}|\ge\delta N\). Choose the pairs independently with probability \(\epsilon\). For \(|\mathcal{I}|=|\mathcal{J}|=m:=\lceil\delta N\rceil\) the probability of missing \(\mathcal{I}\times\mathcal{J}\) is \((1-\epsilon)^{m^2}\le e^{-\epsilon\delta^2N^2}\), and there are at most \(\binom{N}{m}^2\le2^{2N}\) such pairs, so the probability that some pair is missed is at most \(2^{2N}e^{-\epsilon\delta^2N^2}<1/2\) for every large \(N=2^n\) (the exponent grows quadratically, the binomial bound only linearly); larger subfamilies contain ones of size \(m\). The expectation of \(|\mathcal{Z}|\) is \(\epsilon N^2\), so Markov’s inequality gives \(P(|\mathcal{Z}|\ge2\epsilon N^2)\le1/2\). The two probabilities sum to less than \(1\), so some realization has both properties.
Construction of \(B\). With \(\epsilon_k=2^{-k-3}\) and \(\delta_k=1/k\), pick inductively \(n_1<n_2<\cdots\) and sets \(\mathcal{Z}_k\) as in the lemma, and put \(Z_k=\bigcup_{(I,J)\in\mathcal{Z}_k}I\times J\) and \(B=[0,1)^2\setminus\bigcup_kZ_k\), a Borel set with
\begin{equation*} \lambda_2(Z_k)=|\mathcal{Z}_k|\,4^{-n_k}\le2\epsilon_k=2^{-k-2},\qquad \lambda_2(B) \ge1-\tfrac14=\tfrac34 . \end{equation*}
Verification of (\(\star\)). Let \(\lambda(S),\lambda(T)>0\) and suppose \(\lambda_2((S\times T)\setminus B)=0\). By the Lebesgue density theorem (Section 5.8(ii)) almost every \(s\in S\) is a density point, and for such \(s\) also \(\lambda(S\cap I_n(s))2^{n}\to1\), where \(I_n(s)\in\mathcal{D}_n\) contains \(s\): indeed \(I_n(s)\subset\Delta_n=[s-2^{-n},s+2^{-n}]\) with \(\lambda(\Delta_n\setminus I_n(s))=2^{-n}\), so \(\lambda(S\cap I_n(s))\ge2^{-n}(1-2\eta_n)\), \(\eta_n\to0\). Hence the sets \(S^{(N)}=\{s\in S:\lambda(S\cap I_n(s))\ge0.9\cdot2^{-n}\ \text{for all}\ n\ge N\}\) increase to a subset of \(S\) of full measure, so \(\lambda(S^{(N_1)})\ge\lambda(S)/2\) for some \(N_1\), and likewise \(\lambda(T^{(N_2)})\ge\lambda(T)/2\). For \(n\ge\max(N_1,N_2)\) the families \(\mathcal{I}_n=\{I\in\mathcal{D}_n:\lambda(S\cap I)\ge0.9\lambda(I)\}\) and \(\mathcal{J}_n=\{J\in\mathcal{D}_n:\lambda(T\cap J)\ge0.9\lambda(J)\}\) cover \(S^{(N_1)}\) and \(T^{(N_2)}\), so \(|\mathcal{I}_n|\ge(\lambda(S)/2)2^{n}\) and \(|\mathcal{J}_n|\ge(\lambda(T)/2)2^{n}\). Choose \(k\) with \(\delta_k\le\min\{\lambda(S),\lambda(T)\}/2\) and \(n_k\ge\max(N_1,N_2)\); then some \((I,J)\in\mathcal{Z}_k\) has \(I\in\mathcal{I}_{n_k}\), \(J\in\mathcal{J}_{n_k}\), so \(I\times J\subset Z_k\) is disjoint from \(B\), while
\begin{equation*} \lambda_2\bigl((S\times T)\cap(I\times J)\bigr)=\lambda(S\cap I)\,\lambda(T\cap J) \ge0.81\,\lambda(I)\lambda(J)>0 . \end{equation*}
As \(B\cap(I\times J)=\varnothing\), this positive-measure set lies inside \((S\times T)\setminus B\), a contradiction.
(i) (Brodskii, Eggleston) Let a set \(E\subset[0,1]\times[0,1]\) have Lebesgue measure \(1\). Prove that there exist a nonempty perfect set \(P\subset[0,1]\) and a compact set \(K\subset[0,1]\) of positive measure such that \(P\times K\subset E\).
(ii) (Davies) Suppose that every union of less than \(\mathfrak{c}\) Lebesgue measure zero sets has measure zero (which holds, e.g., under the continuum hypothesis or Martin’s axiom). Prove that every measurable set \(E\subset[0,1]^2\) of Lebesgue measure \(1\) contains a product-set \(X\times Y\) such that \(X\) and \(Y\) in \([0,1]\) have outer measure \(1\).
(ii) Put \(Z=[0,1]^2\setminus E\), so \(\lambda_2(Z)=0\), with sections \(Z_x,Z^y\). By Theorem 3.4.1 the sets \(G_1=\{x:\lambda(Z_x)=0\}\) and \(G_2=\{y:\lambda(Z^y)=0\}\) have full measure. Note that \(\lambda^*(S)=1\) exactly when \(S\) meets every closed set of positive measure: if \(\lambda^*(S)<1\), a measurable hull \(M\supset S\) with \(\lambda(M)<1\) has, by inner regularity, a closed \(F\subset[0,1]\setminus M\) with \(\lambda(F)>0\); conversely a closed \(F\) of positive measure disjoint from \(S\) gives \(\lambda^*(S)\le1-\lambda(F)<1\). The closed sets of positive measure are \(\mathfrak{c}\) in number (a closed set is the complement of a union of rational intervals); enumerate them as \((F_\alpha)_{\alpha<\mathfrak{c}}\), identifying \(\mathfrak{c}\) with the least ordinal of that cardinality.
Choose points by recursion on \(\alpha<\mathfrak{c}\) so that
\begin{equation*} x_\alpha\in F_\alpha\cap G_1\setminus\bigcup_{\beta<\alpha}Z^{y_\beta},\qquad y_\alpha\in F_\alpha\cap G_2\setminus\bigcup_{\beta\le\alpha}Z_{x_\beta}. \tag{\(\ast\)} \end{equation*}
This is possible: each \(Z^{y_\beta}\) is null because \(y_\beta\in G_2\), and fewer than \(\mathfrak{c}\) of them are united, so the union is null by hypothesis, while \(F_\alpha\cap G_1\) has measure \(\lambda(F_\alpha)>0\); symmetrically for \(y_\alpha\), all \(x_\beta\) with \(\beta\le\alpha\) lying in \(G_1\).
Put \(X=\{x_\alpha\}_{\alpha<\mathfrak{c}}\) and \(Y=\{y_\alpha\}_{\alpha<\mathfrak{c}}\). Then \(X\times Y\subset E\): for \(\beta<\alpha\) the first condition of (\(\ast\)) gives \((x_\alpha,y_\beta)\notin Z\), and for \(\beta\ge\alpha\) the second condition at stage \(\beta\) gives \((x_\alpha,y_\beta)\notin Z\). Finally \(x_\alpha\in F_\alpha\cap X\) and \(y_\alpha\in F_\alpha\cap Y\) for every \(\alpha\), so \(\lambda^*(X)=\lambda^*(Y)=1\).
(i) No extra set-theoretic hypothesis is needed here. By inner regularity (Theorem 1.4.8) fix a compact \(C\subset E\) with \(\lambda_2( C)>3/4\); since \(P\times K\subset C\) implies \(P\times K\subset E\), it suffices to produce \(P,K\) for \(C\). All sections \(C_x\) are compact. Two facts:
- \(A_K=\{x:K\subset C_x\}\) is closed for compact \(C\) and arbitrary \(K\): if \(x_j\to x\) in \(A_K\) and \(y\in K\), then \((x_j,y)\in C\) and \((x_j,y)\to(x,y)\in C\).
- \(x\mapsto\lambda(C_x)\) is upper semicontinuous: for open \(U\supset C_x\) and \(x_j\to x\) we have \(C_{x_j}\subset U\) for large \(j\) (otherwise points \(y_m\in C_{x_{j_m}}\setminus U\) have a limit \(y\in C_x\setminus U\)), so \(\limsup_j\lambda(C_{x_j})\le\lambda(U)\), and \(\lambda(U)\downarrow\lambda(C_x)\) by outer regularity.
Hence \(B=\{x:\lambda(C_x)\ge1/2\}\) is closed, and since \(\lambda(C_x)\le1\) everywhere and \(<1/2\) off \(B\),
\begin{equation*} \tfrac34<\int_0^1\lambda(C_x)\,dx\le\lambda(B)+\tfrac12\bigl(1-\lambda(B)\bigr), \qquad\text{so}\qquad\lambda(B)>\tfrac12 . \end{equation*}
\(L^1\)-continuity of the section map. Put \(\Psi(x)=I_{C_x}\in L^1[0,1]\), so \(\|\Psi(x)-\Psi(x^{\prime})\|_{L^1}=\lambda(C_x\triangle C_{x^{\prime}})\). For simple \(h=\sum_{j\le N}a_jI_{D_j}\) with \(\{D_j\}\) a Borel partition,
\begin{equation*} \|\Psi(x)-h\|_{L^1}=\sum_{j\le N}\bigl(|1-a_j|\,\lambda(D_j\cap C_x) +|a_j|\,\lambda(D_j\setminus C_x)\bigr), \end{equation*}
measurable in \(x\) by Theorem 3.4.1 applied to \(C\cap([0,1]\times D_j)\); a general \(h\in L^1\) follows because \(h\mapsto\|\Psi(x)-h\|_{L^1}\) is \(1\)-Lipschitz and simple functions are dense. Take \(\{h_j\}\) dense in \(L^1[0,1]\), put \(f_j(x)=\|\Psi(x)-h_j\|_{L^1}\), and by Lusin’s theorem (Theorem 2.2.10) choose compact \(H_j\) with \(\lambda([0,1]\setminus H_j)<2^{-j-3}\) and \(f_j|_{H_j}\) continuous. Then \(H=B\cap\bigcap_jH_j\) is compact with \(\lambda(H)>\tfrac12-\tfrac18=\tfrac38\), and \(\Psi|_H\) is continuous: for \(x_k\to x\) in \(H\) and \(\eta>0\) pick \(j\) with \(f_j(x)<\eta\); then \(f_j(x_k)<\eta\) for large \(k\), so \(\lambda(C_{x_k}\triangle C_x)\le f_j(x_k)+f_j(x)<2\eta\). Compactness makes \(\Psi|_H\) uniformly continuous: for every \(\theta>0\) there is \(\rho>0\) with
\begin{equation*} x,x^{\prime}\in H,\quad|x-x^{\prime}|\le\rho\ \Longrightarrow\ \lambda(C_x\triangle C_{x^{\prime}})\le\theta. \tag{1} \end{equation*}
The countable scheme. Let \(H_d\) be the set of density points of \(H\) lying in \(H\); then \(\lambda(H_d)=\lambda(H)>3/8\), and every \(x\in H_d\) is an accumulation point of \(H_d\), since \(H\cap(x-\rho,x+\rho)\) has positive measure for every \(\rho>0\) (Lebesgue density theorem, Section 5.8). Put \(\theta_n=2^{-2n-2}\), so that \(\sum_n2^{n-1}\theta_n=1/8\), and use (1) to pick \(\rho_n\in(0,2^{-n}]\) forcing \(\lambda(C_x\triangle C_{x^{\prime}})\le\theta_n\) for \(x,x^{\prime}\in H\) with \(|x-x^{\prime}|\le\rho_n\). Fix \(x_0\in H_d\), set \(D_0=\{x_0\}\), and given the finite \(D_{n-1}\subset H_d\) choose \(\sigma_n(x)\in H_d\) with \(0<|\sigma_n(x)-x|\le\rho_n\) for each \(x\in D_{n-1}\), putting \(D_n=D_{n-1}\cup\sigma_n(D_{n-1})\). Then \(|D_{n-1}|\le2^{n-1}\) and \(\lambda(C_x\setminus C_{\sigma_n(x)})\le\theta_n\). Let \(D=\bigcup_nD_n\), a countable subset of \(H_d\).
The sets \(K\) and \(P\). Put \(K=\bigcap_{x\in D}C_x\), a countable intersection of compacta. Then
\begin{equation*} C_{x_0}\setminus K\subset \bigcup_{n\ge1}\ \bigcup_{x\in D_{n-1}}\bigl(C_x\setminus C_{\sigma_n(x)}\bigr), \end{equation*}
because a point \(z\in C_{x_0}\) outside the right-hand side lies in \(C_y\) for every \(y\in D_n\), by induction on \(n\): for \(y=\sigma_n(x)\) with \(x\in D_{n-1}\) use \(z\in C_x\) and \(z\notin C_x\setminus C_{\sigma_n(x)}\). Hence \(\lambda(C_{x_0}\setminus K)\le\sum_n2^{n-1}\theta_n=1/8\), and \(x_0\in H_d\subset B\) gives \(\lambda(K)\ge\tfrac12-\tfrac18=\tfrac38>0\). Finally \(P=\overline{D}\) is a nonempty perfect set: each \(x\in D_m\) lies in \(D_{n-1}\) for every \(n>m\), so \(\sigma_n(x)\in D\) with \(0<|\sigma_n(x)-x|\le2^{-n}\to0\), and a point of \(P\setminus D\) is by definition a limit of points of \(D\). Since \(D\subset A_K\) and \(A_K\) is closed, \(P\subset A_K\), i.e. \(P\times K\subset C\subset E\).
Exercises 3.10.64–3.10.70
Let \((X,\mathcal{A},\mu)\) and \((Y,\mathcal{B},\nu)\) be measure spaces, where \(\mu\) and \(\nu\) take values in \([0,+\infty]\). Denote by \(\lambda_{max}\) the measure corresponding to the Carathéodory outer measure generated by the set function \(\tau(A\times B)=\mu(A)\nu(B)\) on the class of all sets \(A\times B\), where \(A\in\mathcal{A}\), \(B\in\mathcal{B}\). Let \(\Lambda\) be the domain of definition of \(\lambda_{max}\) according to the Carathéodory construction. Let \(\lambda_{min}\) denote the set function on \(\Lambda\) with values in \([0,+\infty]\) defined by the formula
\begin{equation*} \lambda_{min}(L)=\sup\bigl\{\lambda_{max}\bigl(L\cap(A\times B)\bigr):\ A\in\mathcal{A},\ \mu(A)<\infty,\ B\in\mathcal{B},\ \nu(B)<\infty\bigr\}. \end{equation*}
(i) Show that \(\mathcal{A}\otimes\mathcal{B}\subset\Lambda\) and \(\lambda_{max}(A\times B)=\mu(A)\nu(B)\) for all \(A\in\mathcal{A}\), \(B\in\mathcal{B}\).
(ii) Show that \(\lambda_{min}(A\times B)=\mu(A)\nu(B)\) if \(A\in\mathcal{A}\), \(B\in\mathcal{B}\) and \(\mu(A)\nu(B)<\infty\).
(iii) Show that \(\lambda_{min}(E)=\lambda_{max}(E)\) if \(\lambda_{max}(E)<\infty\).
(iv) Let \(\lambda\) be a measure on \(\mathcal{A}\otimes\mathcal{B}\) with values in \([0,+\infty]\) such that \(\lambda(A\times B)=\mu(A)\nu(B)\) for all \(A\in\mathcal{A}\), \(B\in\mathcal{B}\). Show that \(\lambda_{min}(E)\le\lambda(E)\le\lambda_{max}(E)\) for all \(E\in\mathcal{A}\otimes\mathcal{B}\).
(v) Show that the measures \(\lambda_{min}\) and \(\lambda_{max}\) possess equal collections of integrable functions and the corresponding integrals coincide.
Throughout, products in \([0,+\infty]\) use \(0\cdot\infty=0\); with this convention \(\tau\) is a well-defined function of the rectangle (a nonempty product determines its factors, and every representation of \(\varnothing\) gives \(\tau=0\)), so Example 1.11.5 applies and the Method I outer measure is
\begin{equation*} \theta(S) =\inf\Bigl\{\sum_{n}\mu(A_n)\nu(B_n): \ A_n\in\mathcal{A},\ B_n\in\mathcal{B},\ S\subset\bigcup_nA_n\times B_n\Bigr\}. \end{equation*}
By Theorem 1.11.4(iii), \(\Lambda=\mathcal{M}_{\theta}\) is a \(\sigma\)-algebra and \(\lambda_{max}=\theta|_{\Lambda}\) a complete measure; \(\mathcal{D}\) denotes the rectangles \(A^{\prime}\times B^{\prime}\) with \(\mu(A^{\prime}),\nu(B^{\prime})<\infty\), so that \(\lambda_{min}(L)=\sup_{R\in\mathcal{D}}\lambda_{max}(L\cap R)\), and \(\mathcal{D}\) is directed upwards by inclusion. Two facts recur: distributivity \((a_1+a_2)(b_1+b_2)=\sum_{i,j\le2}a_ib_j\) in \([0,+\infty]\) (Check! the cases where a sum vanishes or a factor is infinite), and \(\lambda_{max}(L_m)\uparrow\lambda_{max}(L)\) for \(L_m\uparrow L\) in \(\Lambda\). The book’s part (i) means \(\mathcal{A}\otimes\mathcal{B}\subset\Lambda\).
(i) First,
\begin{equation*} \theta(A\times B)=\mu(A)\nu(B)\qquad(A\in\mathcal{A},\ B\in\mathcal{B}). \tag{1} \end{equation*}
Only \(\ge\) needs proof. Let \(A\times B\subset\bigcup_nA_n\times B_n\) with \(S:=\sum_n\mu(A_n)\nu(B_n)<\infty\); replacing \(A_n,B_n\) by \(A_n\cap A\), \(B_n\cap B\) only decreases \(S\), and we may assume \(\mu(A)>0\), \(\nu(B)>0\). Fix a finite \(\gamma<\nu(B)\). For \(x\in A\) the \(B_n\) with \(x\in A_n\) cover \(B\), so \(\sum_{n:\,x\in A_n}\nu(B_n)\ge\nu(B)>\gamma\), and the sets \(C_N=\{x\in A:\sum_{n\le N,\,x\in A_n}\nu(B_n)>\gamma\}\in\mathcal{A}\) increase to \(A\). Decomposing \(C_N\) into the atoms \(\mathcal{P}_N\) of the finite algebra generated by \(A_1,\dots,A_N\) and using distributivity,
\begin{equation*} \sum_{n\le N}\nu(B_n)\mu(A_n)\ \ge\ \sum_{P\in\mathcal{P}_N}\mu(P)\!\!\!\sum_{n\le N:\,P\subset A_n}\!\!\!\nu(B_n)\ \ge\ \gamma\,\mu(C_N), \end{equation*}
so \(\gamma\mu(A)\le S\); let \(\gamma\uparrow\nu(B)\) and take the infimum over covers.
Every rectangle \(R=A\times B\) is \(\theta\)-measurable: for \(S\) with \(\theta(S)<\infty\) take a cover with \(\sum_n\mu(A_n)\nu(B_n)\le\theta(S)+\varepsilon\) and split each \(A_n\times B_n\) into the four disjoint rectangles cut out by \(A\) and \(B\); distributivity makes their four \(\tau\)-values sum to \(\mu(A_n)\nu(B_n)\), the first pieces cover \(S\cap R\) and the other three cover \(S\setminus R\), whence \(\theta(S\cap R)+\theta(S\setminus R)\le\theta(S)+\varepsilon\). Since \(\Lambda\) is a \(\sigma\)-algebra containing all rectangles, \(\mathcal{A}\otimes\mathcal{B}\subset\Lambda\); with (1) this is (i).
(ii) By (i) and monotonicity \(\lambda_{min}(A\times B)=\sup\{\mu(A\cap A^{\prime})\nu(B\cap B^{\prime})\}\le\mu(A)\nu(B)\), the supremum over \(\mu(A^{\prime}),\nu(B^{\prime})<\infty\). Assume \(\mu(A)\nu(B)<\infty\). (a) If both factors are finite, \(A^{\prime}=A\), \(B^{\prime}=B\) attains the value. (b) If \(\mu(A)=\infty\), then \(\nu(B)=0\) and the product is \(0\), while every term vanishes because \(\nu(B\cap B^{\prime})=0\) and \(\mu(A\cap A^{\prime})\le\mu(A^{\prime})<\infty\); \(\nu(B)=\infty\) is symmetric.
(iii) \(\lambda_{min}\le\lambda_{max}\) by monotonicity. Let \(\lambda_{max}(E)=\theta(E)<\infty\) and cover \(E\subset\bigcup_nA_n\times B_n\) with \(\sum_n\mu(A_n)\nu(B_n)<\infty\). Put \(J=\{n:\mu(A_n)\nu(B_n)>0\}\) and let \(N\) be the union of the remaining rectangles, each \(\theta\)-null by (1), so \(N\in\Lambda\) and \(\lambda_{max}(N)=0\); for \(n\in J\) both \(\mu(A_n)\) and \(\nu(B_n)\) are finite. The rectangles \(R_m=(\bigcup_{n\in J,\,n\le m}A_n)\times(\bigcup_{n\in J,\,n\le m}B_n)\) lie in \(\mathcal{D}\) and increase with \(\bigcup_mR_m\supset E\setminus N\), so \(E\cap R_m\setminus N\uparrow E\setminus N\) and
\begin{equation*} \lambda_{max}(E)=\lambda_{max}(E\setminus N) =\lim_m\lambda_{max}(E\cap R_m\setminus N)\le\lambda_{min}(E). \end{equation*}
\(\lambda_{min}\) is countably additive on \(\Lambda\) (needed in (v)). For disjoint \(L_n\) with union \(L\) and any \(R\in\mathcal{D}\) we have \(\lambda_{max}(L\cap R)=\sum_n\lambda_{max}(L_n\cap R)\le\sum_n\lambda_{min}(L_n)\), so \(\lambda_{min}(L)\le\sum_n\lambda_{min}(L_n)\). Conversely, given \(N\) and \(c_n<\lambda_{min}(L_n)\) for \(n\le N\), choose \(R_n\in\mathcal{D}\) with \(\lambda_{max}(L_n\cap R_n)>c_n\) and, by directedness, \(R\in\mathcal{D}\) containing \(R_1\cup\cdots\cup R_N\); the \(L_n\cap R\) being disjoint members of \(\Lambda\),
\begin{equation*} \lambda_{min}(L)\ge\lambda_{max}(L\cap R)\ge\sum_{n\le N}\lambda_{max}(L_n\cap R_n) \ge\sum_{n\le N}c_n . \end{equation*}
(iv) Upper bound: countable subadditivity of \(\lambda\) over a cover gives \(\lambda(E)\le\sum_n\mu(A_n)\nu(B_n)\), hence \(\lambda(E)\le\theta(E)=\lambda_{max}(E)\). Lower bound: for \(R=A^{\prime}\times B^{\prime}\in\mathcal{D}\) the two measures \(F\mapsto\lambda(F\cap R)\) and \(F\mapsto\lambda_{max}(F\cap R)\) on \(\mathcal{A}\otimes\mathcal{B}\) are finite of total mass \(\mu(A^{\prime})\nu(B^{\prime})\) and agree on rectangles by (i); rectangles are closed under finite intersections and generate \(\mathcal{A}\otimes\mathcal{B}\), so after normalizing (if the mass is \(0\) both vanish) Lemma 1.9.4 gives equality. Hence \(\lambda_{max}(E\cap R)=\lambda(E\cap R)\le\lambda(E)\), and the supremum over \(R\in\mathcal{D}\) gives \(\lambda_{min}(E)\le\lambda(E)\).
(v) The identity map induces a bijection of \(L^1(\lambda_{max})\) onto \(L^1(\lambda_{min})\) preserving the integrals. (a) Since \(\lambda_{min}\le\lambda_{max}\), every \(\lambda_{max}\)-a.e. statement holds \(\lambda_{min}\)-a.e.; and a \(\lambda_{max}\)-simple \(g=\sum_{i\le k}c_iI_{E_i}\) has \(\lambda_{max}(E_i)<\infty\) whenever \(c_i\ne0\) (Definition 2.6.1), so \(\lambda_{min}(E_i)=\lambda_{max}(E_i)\) by (iii) and \(g\) is \(\lambda_{min}\)-simple with the same integral. Applied to \(|g_n-g_m|\), this shows a sequence of simple functions is fundamental in the mean for one measure iff for the other, so a defining sequence in the sense of Definition 2.4.1 serves both and the integrals coincide.
(b) If \(\lambda_{min}(E)<\infty\), take \(R_m\in\mathcal{D}\) increasing (directedness) with \(\lambda_{max}(E\cap R_m)\to\lambda_{min}(E)\) and put \(T=\bigcup_m(E\cap R_m)\subset E\); then \(\lambda_{max}(T)=\lambda_{min}(E)<\infty\), so \(\lambda_{min}(T)=\lambda_{max}(T)=\lambda_{min}(E)\) by (iii), and \(\lambda_{min}(E\setminus T)=0\) by additivity.
(c) Let \(f\) be \(\lambda_{min}\)-integrable. By Proposition 2.6.2(i), \(\{f\ne0\}\subset\bigcup_nS_n\) with \(\lambda_{min}(S_n)<\infty\); take \(T_n\subset S_n\) as in (b), \(T=\bigcup_nT_n\), \(T_n^{\prime}=T_1\cup\cdots\cup T_n\), so \(g:=fI_T\) equals \(f\) \(\lambda_{min}\)-a.e. On measurable subsets of \(T\) the two measures agree, by (iii) on each \(T_n^{\prime}\) and then by monotone limits; hence their completions there agree and \(g\) is \(\Lambda\)-measurable. Since
\begin{equation*} \sup_n\int_{T_n^{\prime}}|g|\,d\lambda_{max} =\sup_n\int_{T_n^{\prime}}|g|\,d\lambda_{min}\le\int|f|\,d\lambda_{min}<\infty \end{equation*}
and \(T_n^{\prime}\uparrow T\) with \(\lambda_{max}(T_n^{\prime})<\infty\), Proposition 2.6.2(ii) makes \(g\) \(\lambda_{max}\)-integrable with \(\int g\,d\lambda_{max}=\int f\,d\lambda_{min}\).
(d) Injectivity: if \(g\) is \(\lambda_{max}\)-integrable and \(g=0\) \(\lambda_{min}\)-a.e., then each \(\{|g|>1/n\}\) has finite \(\lambda_{max}\)-measure, so by (iii) its two measures agree and both vanish; hence \(g=0\) \(\lambda_{max}\)-a.e.
Let \(\mu\), \(\nu\), \(\lambda_{min}\), and \(\lambda_{max}\) be the same as in Exercise 3.10.64. Show that the following conditions are equivalent: (i) \(\lambda_{min}=\lambda_{max}\), (ii) \(\lambda_{max}\) is semifinite, (iii) \(\lambda_{max}\) is locally determined.
We keep the notation of Exercise 3.10.64: \(\theta\), \(\Lambda=\mathcal{M}_{\theta}\), \(\lambda_{max}=\theta|_{\Lambda}\), and \(\mathcal{D}\) the rectangles of finite \(\mu\)- and \(\nu\)-measure. Semifinite, locally measurable, saturated and locally determined are as in Exercises 1.12.132, 1.12.133 and 1.12.135.
Lemma A. If \(\theta(S)<\infty\), then \(S\subset(\bigcup_nR_n)\cup Z\) with \(R_n\in\mathcal{D}\) and \(\theta(Z)=0\). Indeed, take a cover \(S\subset\bigcup_nA_n\times B_n\) with \(\sum_n\mu(A_n)\nu(B_n)<\infty\) and let \(Z\) be the union of those \(A_n\times B_n\) with \(\mu(A_n)\nu(B_n)=0\), each \(\theta\)-null by Exercise 3.10.64(i); for the remaining indices \(0<\mu(A_n)\nu(B_n)<\infty\) forces both factors finite.
(i) implies (ii). If \(\lambda_{max}(E)=\infty\), then \(\lambda_{min}(E)=\infty\), so some \(R=A^{\prime}\times B^{\prime}\in\mathcal{D}\) has \(\lambda_{max}(E\cap R)>0\), while \(\lambda_{max}(E\cap R)\le\mu(A^{\prime})\nu(B^{\prime})<\infty\) by Exercise 3.10.64(i). Thus \(E\) contains a measurable subset of finite positive measure.
(ii) implies (i). Equality holds on sets of finite \(\lambda_{max}\)-measure by Exercise 3.10.64(iii), so let \(\lambda_{max}(E)=\infty\). A semifinite measure has measurable subsets of arbitrarily large finite measure inside any set of infinite measure: were \(s=\sup\{\lambda_{max}(F):F\subset E,\ F\in\Lambda,\ \lambda_{max}(F)<\infty\}\) finite, pick \(F_n\subset E\) with \(\lambda_{max}(F_n)\to s\) and put \(F_{\infty}=\bigcup_nF_n\); then \(\lambda_{max}(F_1\cup\cdots\cup F_n)\in[\lambda_{max}(F_n),s]\), so \(\lambda_{max}(F_{\infty})=s<\infty\), hence \(\lambda_{max}(E\setminus F_{\infty})=\infty\) and semifiniteness supplies \(G\subset E\setminus F_{\infty}\) with \(0<\lambda_{max}(G)<\infty\), making \(\lambda_{max}(F_{\infty}\cup G)=s+\lambda_{max}(G)>s\), a contradiction. So for any \(c<\infty\) choose \(F\subset E\) with \(c<\lambda_{max}(F)<\infty\); then \(\lambda_{min}(E)\ge\lambda_{min}(F)=\lambda_{max}(F)>c\) by Exercise 3.10.64(iii), and \(\lambda_{min}(E)=\infty\).
(iii) implies (ii). Immediate: locally determined means semifinite and saturated.
(ii) implies (iii). It suffices that \(\lambda_{max}\) is always saturated. Let \(L\) be locally measurable and \(\theta(S)<\infty\) (the case \(\theta(S)=\infty\) is trivial). By Lemma A, \(S\subset M\cup Z\) with \(M=\bigcup_nR_n\), \(R_n\in\mathcal{D}\), \(\theta(Z)=0\). Each \(R_n\in\Lambda\) has \(\lambda_{max}(R_n)<\infty\) by Exercise 3.10.64(i), so \(L\cap R_n\in\Lambda\) and \(L\cap M=\bigcup_n(L\cap R_n)\in\Lambda\); testing its \(\theta\)-measurability on \(S\cap M\), and using \(\theta(Z)=0\) with subadditivity,
\begin{equation*} \theta(S\cap L)+\theta(S\setminus L)\le\theta(S\cap M\cap L) +\theta\bigl((S\cap M)\setminus L\bigr)=\theta(S\cap M)\le\theta(S). \end{equation*}
Hence \(L\in\Lambda\), and with (ii) the measure \(\lambda_{max}\) is locally determined.
Let \(\mu\), \(\nu\), \(\lambda_{min}\), and \(\lambda_{max}\) be the same as in Exercise 3.10.64.
(i) Let \(\mu\) and \(\nu\) be decomposable measures. Prove that the measure \(\lambda_{min}\) is decomposable.
(ii) Show that there exist a Maharam measure \(\mu\) and a probability measure \(\nu\) such that the measure \(\lambda_{min}\) is not Maharam.
We keep the notation \(\theta,\Lambda,\mathcal{D}\) and Lemma A of Exercises 3.10.64 and 3.10.65; decomposable, Maharam and semifinite are as in Exercises 1.12.131 and 1.12.134.
(i) The cells \(G_{\alpha\beta}=X_{\alpha}\times Y_{\beta}\) of decompositions \(\{X_{\alpha}\}_{\alpha\in\Lambda_1}\), \(\{Y_{\beta}\}_{\beta\in\Lambda_2}\) of \(\mu\) and \(\nu\) decompose \((X\times Y,\Lambda,\lambda_{min})\): they are disjoint, cover \(X\times Y\), lie in \(\Lambda\) by Exercise 3.10.64(i), and \(\lambda_{min}(G_{\alpha\beta})=\mu(X_{\alpha})\nu(Y_{\beta})<\infty\) by Exercise 3.10.64(ii).
For \(A\in\mathcal{A}\) with \(\mu(A)<\infty\), property (b) of decomposability makes \(\mu(A)=\sum_{\alpha}\mu(A\cap X_{\alpha})\) finite, so only countably many terms are nonzero; with \(A^{\ast}\) the union of the corresponding \(A\cap X_{\alpha}\), the set \(A\setminus A^{\ast}\) meets every \(X_{\alpha}\) in a null set, whence \(\mu(A\setminus A^{\ast})=0\) by (b). So for \(R=A\times B\in\mathcal{D}\) there is a countable \(C\subset\Lambda_1\times\Lambda_2\) with
\begin{equation*} R\subset\Bigl(\bigcup_{(\alpha,\beta)\in C}G_{\alpha\beta}\Bigr) \cup\bigl[(A\setminus A^{\ast})\times Y\bigr] \cup\bigl[X\times(B\setminus B^{\ast})\bigr], \tag{2} \end{equation*}
the last two rectangles being \(\theta\)-null by Exercise 3.10.64(i) and \(0\cdot\infty=0\).
Property (a): let \(W\cap G_{\alpha\beta}\in\Lambda\) for all \(\alpha,\beta\), and \(\theta(S)<\infty\). Lemma A together with (2) applied to each \(R_n\) gives \(S\subset M\cup Z^{\prime}\) with \(M=\bigcup_{(\alpha,\beta)\in C}G_{\alpha\beta}\), \(C\) countable, and \(\theta(Z^{\prime})=0\). Then \(W\cap M\in\Lambda\) as a countable union, and exactly as in the saturation proof of Exercise 3.10.65,
\begin{equation*} \theta(S\cap W)+\theta(S\setminus W)\le\theta(S\cap M\cap W) +\theta\bigl((S\cap M)\setminus W\bigr)=\theta(S\cap M)\le\theta(S), \end{equation*}
so \(W\in\Lambda\).
Property (b): for finite \(F\subset\Lambda_1\times\Lambda_2\), additivity of \(\lambda_{min}\) gives \(\sum_{F}\lambda_{min}(W\cap G_{\alpha\beta})\le\lambda_{min}(W)\), hence \(\sum_{\alpha,\beta}\lambda_{min}(W\cap G_{\alpha\beta})\le\lambda_{min}(W)\). Conversely, for \(R\in\mathcal{D}\) take \(C\) as in (2), so that \(\theta(R\setminus\bigcup_CG_{\alpha\beta})=0\); since \(\lambda_{max}(W\cap G_{\alpha\beta})\le\mu(X_{\alpha})\nu(Y_{\beta})<\infty\), Exercise 3.10.64(iii) makes the two measures agree on subsets of cells, and countable additivity gives
\begin{equation*} \lambda_{max}(W\cap R)=\sum_{C}\lambda_{max}(W\cap R\cap G_{\alpha\beta}) \le\sum_{C}\lambda_{min}(W\cap G_{\alpha\beta}). \end{equation*}
Taking the supremum over \(R\in\mathcal{D}\) gives the reverse inequality, so \(\lambda_{min}\) is decomposable.
(ii) Take \(C\) with \(\operatorname{card}C>\mathfrak{c}\), \(\mathcal{I}=\mathcal{P}( C)\), \(X=\{0,1\}^{\mathcal{I}}\), and \(x_{\gamma}\in X\) defined by \(x_{\gamma}(\Gamma)=1\) iff \(\gamma\in\Gamma\). With \(\mathcal{K}\) the at most countable subsets of \(\mathcal{I}\) and \(F_{\gamma K}=\{x:x|_K=x_{\gamma}|_K\}\), put
\begin{equation*} \Sigma_{\gamma} =\{E\subset X:\ \exists K\in\mathcal{K}\ \ F_{\gamma K}\subset E \ \text{or}\ F_{\gamma K}\subset X\setminus E\}, \qquad \Sigma=\bigcap_{\gamma\in C}\Sigma_{\gamma}, \end{equation*}
and \(\mu(E)=\operatorname{card}\{\gamma:x_{\gamma}\in E\}\), \(=+\infty\) when infinite. Each \(\Sigma_{\gamma}\) is a \(\sigma\)-algebra: the definition is symmetric in \(E\) and \(X\setminus E\), and for \(E=\bigcup_nE_n\) either some \(E_n\) contains an \(F_{\gamma K}\), hence so does \(E\), or \(F_{\gamma K_n}\subset X\setminus E_n\) for every \(n\), and then \(F_{\gamma K}=\bigcap_nF_{\gamma K_n}\subset X\setminus E\) with \(K=\bigcup_nK_n\in\mathcal{K}\). And \(\mu\) is countably additive, cardinality being additive over disjoint unions in \([0,+\infty]\).
For \(D\subset C\) the set \(G_D=\{x:x(D)=1\}\) lies in \(\Sigma\) (as \(F_{\gamma,\{D\}}\) equals \(G_D\) if \(\gamma\in D\) and \(X\setminus G_D\) otherwise) and \(\{\gamma:x_{\gamma}\in G_D\}=D\). In particular \(E_{\gamma}:=G_{\{\gamma\}}\in\Sigma\) satisfies \(x_{\gamma}\in E_{\gamma}\), \(\mu(E_{\gamma})=1\) and \(\mu(E_{\gamma}\cap E_{\delta})=0\) for \(\gamma\ne\delta\). Since \(\mu(F)=0\) exactly when no \(x_{\gamma}\) lies in \(F\), the measure \(\mu\) is semifinite, and it is Maharam: for \(\mathcal{M}\subset\Sigma\) and \(D=\{\delta:x_{\delta}\in M\ \text{for some}\ M\in\mathcal{M}\}\),
\begin{equation*} \mu(M\setminus H^{\prime})=0\ \ \forall M\in\mathcal{M} \iff x_{\delta}\in H^{\prime}\ \ \forall\delta\in D\iff \mu(G_D\setminus H^{\prime}) =0, \end{equation*}
so \(G_D\) is an essential supremum of \(\mathcal{M}\). (It is also complete: a subset of a \(\mu\)-null \(E\) misses every \(x_{\gamma}\), so some \(F_{\gamma K}\subset X\setminus E\) works for each \(\gamma\).)
Let \(Y=\{0,1\}^{C}\) carry the product \(\nu\) of the uniform measures on \(\{0,1\}\), a probability measure on \(\mathcal{T}=\bigotimes_{\gamma}\mathcal{P}(\{0,1\})\) by Theorem 3.5.1 and Lemma 3.5.2. For \(F_{\gamma}=\{y:y(\gamma)=1\}\) we have \(\nu(F_{\gamma})=1/2\) and \(\nu(F_{\gamma}\triangle F_{\delta})=1/2\) for \(\gamma\ne\delta\).
Lemma D (Fremlin 2A1P). If \(\operatorname{card}A>\mathfrak{c}\), \(\langle f_{\alpha}\rangle_{\alpha\in A}\subset\{0,1\}^{\mathcal{I}}\) and the \(K_{\alpha}\subset\mathcal{I}\) are at most countable, then \(f_{\alpha}=f_{\beta}\) on \(K_{\alpha}\cap K_{\beta}\) for some distinct \(\alpha,\beta\). Proof: recursively choose disjoint \(M_{\xi}\), \(\xi<\omega_1\), letting \(M_{\xi}\) be any set of cardinality at most \(\mathfrak{c}\), disjoint from \(\bigcup_{\eta<\xi}M_{\eta}\), with \(\operatorname{card}\{\alpha:K_{\alpha}\cap M_{\xi}=\varnothing\}\le\mathfrak{c}\), and \(M_{\xi}=\varnothing\) if none exists; then \(M=\bigcup_{\xi}M_{\xi}\) has cardinality at most \(\mathfrak{c}\). By Zorn’s lemma the family of sets \(P\subset A\) with \(K_{\alpha}\cap K_{\beta}\subset M\) for distinct \(\alpha,\beta\in P\) (a condition on pairs, so chains are admissible) has a maximal element \(B\). If \(\operatorname{card}B\le\mathfrak{c}\), then \(N_0=(\bigcup_{\alpha\in B}K_{\alpha})\setminus M\) has cardinality at most \(\mathfrak{c}\), is disjoint from \(M\), and by maximality \(\{\gamma:K_{\gamma}\cap N_0=\varnothing\}\subset B\); so the recursion never idled, each \(C_{\xi}=\{\alpha:K_{\alpha}\cap M_{\xi}=\varnothing\}\) has cardinality at most \(\mathfrak{c}\), and \(\operatorname{card}\bigcup_{\xi<\omega_1}C_{\xi}\le\mathfrak{c}<\operatorname{card}A\) yields an \(\alpha\) with \(K_{\alpha}\cap M_{\xi}\ne\varnothing\) for all \(\xi<\omega_1\), making the countable set \(K_{\alpha}\cap M\) surject onto \(\omega_1\), absurd. So \(\operatorname{card}B>\mathfrak{c}\). The functions from at most countable subsets of \(M\) to \(\{0,1\}\) number at most \(\mathfrak{c}^{\aleph_0}\cdot\mathfrak{c}=\mathfrak{c}\), so \(\alpha\mapsto f_{\alpha}|_{K_{\alpha}\cap M}\) is not injective on \(B\); distinct \(\alpha,\beta\in B\) with equal restrictions have \(K_{\alpha}\cap K_{\beta}\subset K_{\alpha}\cap M\), where \(f_{\alpha}=f_{\beta}\).
Lemma C. If \(\theta(W)=0\), there is \(N\in\mathcal{A}\) with \(\mu(N)=0\) and \(\nu^{\ast}(W_x)=0\) for all \(x\notin N\). Proof: for each \(m\) cover \(W\subset\bigcup_nA_n^m\times B_n^m\) with \(\sum_n\mu(A_n^m)\nu(B_n^m)<4^{-m}\) and put \(D^m=\{x:\sum_{n:x\in A_n^m}\nu(B_n^m)>2^{-m}\}\); exactly as in the proof of (1) in Exercise 3.10.64, with \(2^{-m}\) in place of \(\gamma\), \(D^m\in\mathcal{A}\) with \(2^{-m}\mu(D^m)\le4^{-m}\). Then \(N=\bigcap_M\bigcup_{m\ge M}D^m\) has \(\mu(N)\le2^{-M+1}\) for every \(M\), so \(\mu(N)=0\); and for \(x\notin N\), \(W_x\subset\bigcup\{B_n^m:x\in A_n^m\}\) gives \(\nu^{\ast}(W_x)\le2^{-m}\) for all large \(m\).
\(\lambda_{min}\) is not Maharam. It is semifinite: if \(\lambda_{min}(E)=\infty\), some \(R=A^{\prime}\times B^{\prime}\in\mathcal{D}\) has \(0<\lambda_{max}(E\cap R)\le\mu(A^{\prime})\nu(B^{\prime})<\infty\), so \(E\cap R\) has finite positive \(\lambda_{min}\)-measure by Exercise 3.10.64(iii). Suppose \(V\in\Lambda\) were an essential supremum of \(\mathcal{W}=\{E_{\gamma}\times F_{\gamma}:\gamma\in C\}\). Any \(L\in\Lambda\) inside \(E_{\gamma}\times Y\) has \(\lambda_{max}(L)\le\mu(E_{\gamma})\nu(Y)=1\), so \(\lambda_{min}(L)=\theta(L)\) by Exercise 3.10.64(iii); in particular \(\lambda_{min}(L)=0\) forces \(\theta(L)=0\). Hence \(\theta((E_{\gamma}\times F_{\gamma})\setminus V)=0\), and Lemma C gives \(N_{\gamma}^{\prime}\in\Sigma\) of measure \(0\) with \(\nu^{\ast}(F_{\gamma}\setminus V_x)=0\) for \(x\in E_{\gamma}\setminus N_{\gamma}^{\prime}\). Next, \(W_{\gamma}=(X\times Y)\setminus(E_{\gamma}\times(Y\setminus F_{\gamma}))\) is an essential upper bound of \(\mathcal{W}\), since \((E_{\delta}\times F_{\delta})\setminus W_{\gamma}\subset(E_{\gamma}\cap E_{\delta})\times Y\) is \(\theta\)-null for \(\delta\ne\gamma\) and empty for \(\delta=\gamma\); so minimality gives \(\lambda_{min}(V\cap(E_{\gamma}\times(Y\setminus F_{\gamma})))=0\), hence \(\theta=0\), and Lemma C gives \(N_{\gamma}^{\prime\prime}\in\Sigma\) of measure \(0\) with \(\nu^{\ast}(V_x\setminus F_{\gamma})=0\) for \(x\in E_{\gamma}\setminus N_{\gamma}^{\prime\prime}\). Put \(T_{\gamma}=E_{\gamma}\setminus(N_{\gamma}^{\prime}\cup N_{\gamma}^{\prime\prime})\in\Sigma\), so \(\mu(T_{\gamma})=1\), \(x_{\gamma}\in T_{\gamma}\) (its only \(x_{\delta}\) can be \(x_{\gamma}\)), and \(\nu^{\ast}(V_x\triangle F_{\gamma})=0\) on \(T_{\gamma}\). The \(T_{\gamma}\) are pairwise disjoint: \(x\in T_{\gamma}\cap T_{\delta}\) with \(\gamma\ne\delta\) would give
\begin{equation*} \tfrac12=\nu(F_{\gamma}\triangle F_{\delta})\le\nu^{\ast}(V_x\triangle F_{\gamma}) +\nu^{\ast}(V_x\triangle F_{\delta})=0 . \end{equation*}
Since \(T_{\gamma}\in\Sigma_{\gamma}\) and \(x_{\gamma}\in F_{\gamma K}\cap T_{\gamma}\) for every \(K\in\mathcal{K}\), the alternative \(F_{\gamma K}\subset X\setminus T_{\gamma}\) is impossible, so \(F_{\gamma,K_{\gamma}}\subset T_{\gamma}\) for some \(K_{\gamma}\in\mathcal{K}\). Lemma D applied to \(\langle x_{\gamma}\rangle_{\gamma\in C}\) and \(\langle K_{\gamma}\rangle\) yields distinct \(\gamma,\delta\) with \(x_{\gamma}=x_{\delta}\) on \(K_{\gamma}\cap K_{\delta}\); the point \(x\) equal to \(x_{\gamma}\) on \(K_{\gamma}\), to \(x_{\delta}\) on \(K_{\delta}\) and to \(0\) elsewhere is then well defined and lies in \(F_{\gamma,K_{\gamma}}\cap F_{\delta,K_{\delta}}\subset T_{\gamma}\cap T_{\delta}=\varnothing\).
(Luther) Let \(X=Y=[0,1]\), let \(\mathcal{A}=\mathcal{B}([0,1])\), and let the measure \(\mu=\nu\) with values in \([0,+\infty]\) be defined as follows: we fix a non-Borel set \(E\); then every point \(x\) is assigned the measure \(2\) or \(1\) depending on whether \(x\) belongs to \(E\) or not, finally, the measure extends naturally to all Borel sets (in particular, all infinite sets obtain infinite measures). Let \(\pi\) be the Carathéodory extension of the measure \(\mu\otimes\nu\). Prove that the measure \(\pi\) is semifinite, \(\mu=\nu\) is semifinite and complete, but for the diagonal \(D\) in \([0,1]\times[0,1]\) the function \(\nu(D_x)=I_E(x)+1\) is not measurable with respect to \(\mu\).
Write \(w(x)=I_E(x)+1\), so that \(\mu(A)=\nu(A)=\sum_{x\in A}w(x)\in[0,+\infty]\), the sum being the supremum of its finite partial sums. Such a weighted counting measure is countably additive on any \(\sigma\)-algebra (for disjoint \(A_n\) the finite partial sums over \(\bigcup_nA_n\) are dominated by \(\sum_n\mu(A_n)\), and conversely), and \(\mu(A)\ge\operatorname{card}A\), so infinite sets get infinite measure.
\(\mu=\nu\) is semifinite and complete: \(\mu(\{x\})=w(x)\in\{1,2\}\), so every nonempty Borel set contains a Borel subset of finite positive measure, while \(\mu(A)=0\) only for \(A=\varnothing\), whose only subset is \(\varnothing\). Hence the \(\mu\)-measurable sets are exactly the Borel sets.
\(\pi=c\), where \(c(S)=\sum_{(x,y)\in S}w(x)w(y)\) for \(S\subset[0,1]^2\). Indeed \(c\) is a measure on all subsets of the square, and
\begin{equation*} c(A\times B)=\Bigl(\sum_{x\in A}w(x)\Bigr)\Bigl(\sum_{y\in B}w(y)\Bigr)=\mu(A)\nu(B) \end{equation*}
in \([0,+\infty]\) (a vanishing factor means an empty set and both sides vanish; otherwise both sides are infinite together or given by finite arithmetic). So for the Caratheodory outer measure \(\theta\) of Exercise 3.10.64, countable subadditivity of \(c\) over any cover gives \(c\le\theta\); conversely an at most countable \(S\) is covered by the singleton rectangles \(\{x\}\times\{y\}\), giving \(\theta(S)\le c(S)\), while an uncountable \(S\) has \(c(S)=+\infty\), all weights being at least \(1\). Thus \(\theta=c\), and \(c\) being countably additive on all subsets makes every set Caratheodory measurable, so \(\Lambda=2^{[0,1]^2}\) and \(\pi=c\).
\(\pi\) is semifinite: any singleton \(\{(x,y)\}\subset S\) is \(\pi\)-measurable with \(\pi(\{(x,y)\})=w(x)w(y)\in\{1,2,4\}\).
The diagonal. \(D\in\mathcal{A}\otimes\mathcal{B}\), since with \(I_{n,i}=[(i-1)2^{-n},i2^{-n})\) for \(i<2^n\) and \(I_{n,2^n}=[1-2^{-n},1]\),
\begin{equation*} D=\bigcap_{n=1}^{\infty}\ \bigcup_{i=1}^{2^n}\,(I_{n,i}\times I_{n,i}), \end{equation*}
a point of the right-hand side satisfying \(|x-y|\le2^{-n}\) for every \(n\). Its sections are \(D_x=\{x\}\), so
\begin{equation*} \nu(D_x)=w(x)=I_E(x)+1 , \end{equation*}
and \(\{x:\nu(D_x)=2\}=E\) is not Borel, hence not \(\mu\)-measurable.
(\(\circ\)) Construct a signed bounded measure \(\mu\) on \(\mathbb{N}\), a mapping \(f\colon\mathbb{N}\to\mathbb{N}\) and a function \(g\) on \(\mathbb{N}\) such that \(\mu\circ f^{-1}=0\), but the function \(g\circ f\) is not integrable with respect to \(\mu\) (although \(g\) is integrable against the measure \(\mu\circ f^{-1}\)).
Take \(\mathbb{N}\) with all subsets measurable and
\begin{equation*} \mu(\{2n\})=\frac{1}{n^{2}},\qquad \mu(\{2n-1\})=-\frac{1}{n^{2}}, \end{equation*}
\begin{equation*} \mu(A)=\sum_{k\in A}\mu(\{k\}),\qquad f(2n)=f(2n-1)=n,\qquad g(n)=n . \end{equation*}
Since \(\sum_k|\mu(\{k\})|=\sum_n2n^{-2}=\pi^2/3<\infty\), the series defining \(\mu(A)\) converges absolutely, so \(\mu\) is a bounded signed measure (countable additivity is the rearrangement theorem for absolutely convergent series), with \(|\mu|(A)=\sum_{k\in A}|\mu(\{k\})|\) and \(\|\mu\|=\pi^2/3\).
\(\mu\circ f^{-1}=0\): for \(A\subset\mathbb{N}\) the preimage \(f^{-1}(A)\) is the disjoint union of the pairs \(\{2n-1,2n\}\), \(n\in A\), so
\begin{equation*} \mu\circ f^{-1}(A)=\sum_{n\in A}\Bigl(-\frac{1}{n^{2}}+\frac{1}{n^{2}}\Bigr)=0 . \end{equation*}
Hence every function on \(\mathbb{N}\), \(g\) included, is integrable against \(\mu\circ f^{-1}\), with integral \(0\).
But \(g\circ f\) is not \(\mu\)-integrable: \((g\circ f)(2n)=(g\circ f)(2n-1)=n\), so
\begin{equation*} \int_{\mathbb{N}}|g\circ f|\,d|\mu|=\sum_{n=1}^{\infty}\Bigl(n\cdot\frac{1}{n^{2}} +n\cdot\frac{1}{n^{2}}\Bigr)=+\infty , \end{equation*}
and integrability against a signed measure means integrability against \(|\mu|\).
Let \(f\in\mathcal{L}^1(\mathbb{R}^1)\). Prove that the function \(f(x-x^{-1})\) is integrable and one has
\begin{equation*} \int_{-\infty}^{+\infty}f(x-x^{-1})\,dx=\int_{-\infty}^{+\infty}f(x)\,dx . \end{equation*}
Put \(\varphi(x)=x-x^{-1}\) and \(\sigma(x)=-x^{-1}\), both defined off the null set \(\{0\}\). The identity
\begin{equation*} \varphi(\sigma(x))=-\frac1x+x=\varphi(x),\qquad x\ne0, \tag{3} \end{equation*}
carries the proof. On each of \((0,+\infty)\) and \((-\infty,0)\) the map \(\varphi\) is \(C^1\) with \(\varphi^{\prime}(x)=1+x^{-2}>0\) and increases from \(-\infty\) to \(+\infty\), hence is a diffeomorphism onto \(\mathbb{R}\); and \(\sigma\) is a diffeomorphism interchanging the two half-lines, with \(\sigma^{\prime}(x)=x^{-2}\).
Let \(u\ge0\) be Borel. Substituting \(x=\sigma(t)\), \(t\in(0,+\infty)\), and using (3) with Theorem 3.7.1 (which extends from \(\mathcal{L}^1\) to nonnegative Borel integrands by monotone convergence along \(h_n=\min(h,n)I_{[-n,-1/n]}\); here \(h=u\circ\varphi\) on \((-\infty,0)\), Borel since \(\varphi\) is continuous there),
\begin{equation*} \int_{-\infty}^{0}u(\varphi(x))\,dx=\int_{0}^{+\infty}u(\varphi(t))\,t^{-2}\,dt , \end{equation*}
and symmetrically with the half-lines exchanged. Adding the two,
\begin{equation*} 2\int_{\mathbb{R}}u(\varphi(x))\,dx=\int_{\mathbb{R}}u(\varphi(x))\bigl(1 +x^{-2}\bigr)\,dx . \end{equation*}
Applying Theorem 3.7.1 to the two diffeomorphisms \(\varphi:(0,+\infty)\to\mathbb{R}\) and \(\varphi:(-\infty,0)\to\mathbb{R}\), whose Jacobians are \(1+x^{-2}\), the right-hand side equals \(2\int_{\mathbb{R}}u(z)\,dz\), so
\begin{equation*} \int_{\mathbb{R}}u\bigl(x-x^{-1}\bigr)\,dx=\int_{\mathbb{R}}u(z)\,dz \tag{4} \end{equation*}
for every nonnegative Borel \(u\), both sides lying in \([0,+\infty]\).
Now let \(f\in\mathcal{L}^1(\mathbb{R}^1)\) and choose a Borel \(\tilde f\) equal to \(f\) off a null set \(N\). On each half-line \(\varphi^{-1}\) is \(C^1\), hence locally Lipschitz, so \(\varphi^{-1}(N)\) is null and \(f\circ\varphi=\tilde f\circ\varphi\) a.e.; in particular \(x\mapsto f(x-x^{-1})\) is Lebesgue measurable and neither side of the assertion changes when \(f\) is replaced by \(\tilde f\). So assume \(f\) Borel. Then (4) with \(u=|f|\) gives \(\int_{\mathbb{R}}|f(x-x^{-1})|\,dx=\|f\|_{L^1}<\infty\), and (4) with \(u=f^{+}\) and \(u=f^{-}\), subtracted (note \(f^{\pm}\circ\varphi=(f\circ\varphi)^{\pm}\)), gives the asserted identity.
Prove that there exists a continuous function \(f\) on \([0,1]\) that is constant on no interval, but \(f(x)\) is a rational number for a.e. \(x\).
Take \(\mu=\sum_{n\ge1}c_n\delta_{q_n}\), with \(\{q_n\}\) an enumeration of \(\mathbb{Q}\cap[0,1]\) and \(c_n>0\), \(\sum_nc_n=1\), and put
\begin{equation*} F=\{f\in C[0,1]:\ f([0,1])\subset[0,1],\ \lambda\circ f^{-1}=\mu\}. \end{equation*}
Every \(f\in F\) has \(\lambda(\{x:f(x)\in\mathbb{Q}\})=\mu(\mathbb{Q})=1\), so it suffices to find an \(f\in F\) constant on no nondegenerate interval.
\(F\ne\varnothing\). The function \(G(t)=\mu([0,t])\) is nondecreasing, right continuous, \(G(1)=1\), and strictly increasing, since \((s,t]\) contains some \(q_n\) and then \(G(t)-G(s)\ge c_n>0\). Put \(f_1(x)=\inf\{t\in[0,1]:G(t)\ge x\}\) for \(0<x\le1\) and \(f_1(0)=0\); as usual
\begin{equation*} f_1(x)\le t\iff x\le G(t)\qquad (0<x\le1,\ t\in[0,1]), \tag{5} \end{equation*}
the nontrivial implication using right continuity of \(G\). Then \(f_1\) is continuous: a jump at \(x_0\in(0,1]\), with \(f_1\le a\) to the left and \(f_1\ge b>a\) to the right, would give through (5) that \(G(t)=x_0\) for every \(t\in(a,b)\), contradicting strict monotonicity; and \(G(0)=\mu(\{0\})>0\) makes \(f_1\equiv0\) on \((0,G(0)]\), giving continuity at \(0\). Finally (5) yields \(\lambda(\{f_1\le t\})=G(t)=\mu([0,t])\) for every \(t\), and the sets \([0,t]\) form a class closed under finite intersections generating \(\mathcal{B}([0,1])\), so \(\lambda\circ f_1^{-1}=\mu\) by Lemma 1.9.4.
\(F\) is a complete metric space under \(d(\phi,\psi)=\sup_t|\phi(t)-\psi(t)|\), being closed in \(C[0,1]\): if \(f_k\in F\) converge uniformly to \(f\), then \(\int\phi\,d\mu=\int_0^1\phi(f_k(x))\,dx\to\int_0^1\phi(f(x))\,dx\) for every \(\phi\in C[0,1]\); approximating \(I_C\) for closed \(C\) by \(\phi_m(t)=\max(0,1-m\operatorname{dist}(t,C))\downarrow I_C\) and using dominated convergence, \(\mu\) and \(\lambda\circ f^{-1}\) agree on closed sets, hence everywhere by Lemma 1.9.4.
The Baire argument. Suppose every \(f\in F\) is constant on some nondegenerate interval. If \(f\equiv c\) there, then \(\mu(\{c\})=\lambda(f^{-1}( c))>0\), so \(c\) is an atom of \(\mu\), i.e. \(c\in\mathbb{Q}\), and the interval contains one with rational endpoints. Hence \(F=\bigcup F_{p,q,r}\) over the countably many rational triples with \(0\le p<q\le1\) and \(r\in\mathbb{Q}\cap[0,1]\), where \(F_{p,q,r}=\{f\in F:f\equiv r\ \text{on}\ [p,q]\}\) is closed, uniform convergence preserving that identity. By Baire’s theorem (Exercise 1.12.83) some \(F_{p,q,r}\) contains a ball \(U=\{f\in F:d(f,f_0)<d_0\}\) with \(f_0\in F_{p,q,r}\), \(d_0>0\).
Let \([a,b]\supset[p,q]\) be the largest interval on which \(f_0\equiv r\) (closed, by continuity). If \([a,b]=[0,1]\), then \(\mu=\lambda\circ f_0^{-1}=\delta_r\), false since \(\mu\) has infinitely many atoms; so, after reflecting \([0,1]\) if necessary, \(a>0\). Fix \(e\in(p,q)\) and, by continuity of \(f_0\) at \(a\), choose \(\varepsilon\in(0,a)\) with \(|f_0-r|<d_0/2\) on \([c,a]\), where \(c:=a-\varepsilon\); since \(e<q\le b\) gives \(f_0\equiv r\) on \([a,e]\), we get \(\sup_K|f_0-r|<d_0/2\) for \(K=[c,e]\). By maximality of \([a,b]\) there is \(s_0\in[c,a)\) with \(f_0(s_0)\ne r\).
Put \(L=e-c\), take an odd \(N=2k+1\) with \(L/N<(e-p)/2\), and let \(\sigma:K\to K\) be the zigzag map: divide \(K\) into the \(N\) consecutive intervals \(J_i\) of length \(L/N\) and send each affinely onto \(K\), increasingly for odd \(i\) and decreasingly for even \(i\). Then \(\sigma\) is continuous, fixes \(c\) and \(e\) (the last branch is increasing, \(N\) being odd), and preserves \(\lambda|_K\), each \(\sigma^{-1}(A)\cap J_i\) being an affine copy of \(A\) scaled by \(1/N\). Put \(\psi=f_0\circ\sigma\) on \(K\) and \(\psi=f_0\) off \(K\); it is continuous with values in \([0,1]\), and
\begin{equation*} \lambda\circ\psi^{-1}=\lambda|_{[0,1]\setminus K}\circ f_0^{-1} +(\lambda|_K\circ\sigma^{-1})\circ f_0^{-1}=\lambda\circ f_0^{-1}=\mu , \end{equation*}
so \(\psi\in F\). Also \(d(\psi,f_0)<d_0\): the difference vanishes off \(K\), and on \(K\) it is at most \(2\sup_K|f_0-r|<d_0\), that supremum being attained on the compact \(K\). Finally the last zigzag interval \(J_{i_0}=[e-L/N,e]\) has left endpoint \(e-L/N>(e+p)/2>p\), so \(J_{i_0}\subset[p,q]\), and \(\sigma\) maps it onto all of \(K\), which contains \(s_0\); hence \(\psi(t_0)=f_0(s_0)\ne r\) for some \(t_0\in[p,q]\), so \(\psi\in U\setminus F_{p,q,r}\), contradicting \(U\subset F_{p,q,r}\).
Exercises 3.10.71–3.10.77
(\(\circ\)) Let \(E\) be a set of finite measure on the real line and let \(\alpha_n\to+\infty\). Prove that
\begin{equation*} \lim_{n\to\infty}\int_E (\sin\alpha_n t)^2\,dt=\lambda(E)/2 . \end{equation*}
From \(2(\sin\alpha t)^2=1-\cos(2\alpha t)\) and \(\lambda(E)<\infty\) with a bounded integrand,
\begin{equation*} \int_E(\sin\alpha_n t)^2\,dt =\frac{\lambda(E)}{2}-\frac12\int_E\cos(2\alpha_n t)\,dt , \end{equation*}
so everything reduces to \(\int_E\cos(\beta t)\,dt\to0\) as \(|\beta|\to\infty\), the Riemann-Lebesgue lemma for \(I_E\in L^1(\mathbb{R}^1)\).
For a bounded interval \(J=(a,b)\) and \(\beta\ne0\),
\begin{equation*} \Bigl|\int_J\cos(\beta t)\,dt\Bigr|=\Bigl|\frac{\sin\beta b-\sin\beta a}{\beta}\Bigr| \le\frac{2}{|\beta|}\longrightarrow0 , \end{equation*}
hence the same for any finite union \(V\) of bounded intervals, by additivity. Given \(\varepsilon>0\), the definition of Lebesgue outer measure supplies bounded open intervals \(I_k\) with \(E\subset U:=\bigcup_kI_k\) and \(\sum_k\lambda(I_k)\le\lambda(E)+\varepsilon\), so \(\lambda(U\setminus E)\le\varepsilon\); choosing \(N\) with \(\sum_{k>N}\lambda(I_k)<\varepsilon\) and putting \(V=\bigcup_{k\le N}I_k\), the inclusions \(E\setminus V\subset\bigcup_{k>N}I_k\) and \(V\setminus E\subset U\setminus E\) give
\begin{equation*} \|I_E-I_V\|_{L^1}=\lambda(E\bigtriangleup V)\le2\varepsilon . \end{equation*}
Since \(|\int_E\cos(\beta t)\,dt-\int_V\cos(\beta t)\,dt|\le\|I_E-I_V\|_{L^1}\) for every \(\beta\), letting \(|\beta|\to\infty\) gives \(\limsup_{|\beta|\to\infty}|\int_E\cos(\beta t)\,dt|\le2\varepsilon\), and \(\varepsilon>0\) was arbitrary. Taking \(\beta=2\alpha_n\to+\infty\),
\begin{equation*} \lim_{n\to\infty}\int_E(\sin\alpha_n t)^2\,dt=\frac{\lambda(E)}{2}. \end{equation*}
(\(\circ\)) Let a sequence of real numbers \(\alpha_n\) be such that \(f(x):=\lim_{n\to\infty}\sin(\alpha_n x)\) exists on a set \(E\) of positive measure. Prove that \(\{\alpha_n\}\) has a finite limit.
The sequence converges to a finite \(\alpha\); we show it is bounded and has one limit point. Shrinking \(E\) to \(E\cap[-N,N]\) we may assume \(0<\lambda(E)<\infty\), so that \(f\) is measurable with \(|f|\le1\) and hence \(f\in L^1(E)\cap L^2(E)\).
(i) Boundedness. Suppose \(|\alpha_{n_k}|\to\infty\) along a subsequence. The Riemann-Lebesgue lemma of Exercise 3.10.71, proved there for indicators of bounded intervals and extended to all of \(L^1(\mathbb{R}^1)\) by density of step functions and the bound \(|\int(g-h)\sin(\beta x)dx|\le\|g-h\|_{L^1}\), applied to \(fI_E\in L^1\) gives
\begin{equation*} \lim_{k\to\infty}\int_E f(x)\sin(\alpha_{n_k}x)\,dx=0 , \end{equation*}
whereas dominated convergence (\(|f\sin(\alpha_{n_k}x)|\le1\), \(\lambda(E)<\infty\)) evaluates the same limit as \(\int_Ef^2dx\); hence \(f=0\) a.e. on \(E\). But Exercise 3.10.71 with \(|\alpha_{n_k}|\to+\infty\) and \((\sin\alpha x)^2=(\sin|\alpha|x)^2\) gives \(\int_E(\sin\alpha_{n_k}x)^2dx\to\lambda(E)/2>0\), while dominated convergence gives \(\int_Ef^2dx=0\) for the same limit – a contradiction.
(ii) A single limit point. If \(\alpha_{n_k}\to\alpha\) and \(\alpha_{m_k}\to\beta\) with \(\alpha\ne\beta\), continuity of the sine forces \(\sin\alpha x=\sin\beta x\) on \(E\). By
\begin{equation*} \sin\alpha x-\sin\beta x =2\sin\tfrac{(\alpha-\beta)x}{2}\,\cos\tfrac{(\alpha+\beta)x}{2}, \end{equation*}
the solution set lies in \(\frac{2\pi}{\alpha-\beta}\mathbb{Z}\cup\{x:\tfrac{(\alpha+\beta)x}{2}\in\tfrac{\pi}{2}+\pi\mathbb{Z}\}\), which is countable (the second set is countable if \(\alpha+\beta\ne0\) and empty if \(\alpha+\beta=0\)), while \(\lambda(E)>0\) makes \(E\) uncountable.
A bounded sequence with exactly one limit point converges to it, so \(\alpha_n\to\alpha\in\mathbb{R}^1\).
(\(\circ\)) Prove that there exists a Lebesgue measurable one-to-one mapping \(f\) of the real line onto itself such that the inverse mapping is not Lebesgue measurable.
Take \(f\) mapping \(\mathbb{R}^1\setminus C\) Borel-bijectively onto \([0,+\infty)\) and the Cantor set \(C\) bijectively onto \((-\infty,0)\) in such a way that a compact piece of \(C\) goes onto a Vitali set. (Measurable means that preimages of Borel sets are Lebesgue measurable.)
Off \(C\): the open set \(\mathbb{R}^1\setminus C\) is a countable disjoint union of open intervals \(I_k\). Fix increasing homeomorphisms \(h_k\colon I_k\to(0,1)\) and put
\begin{equation*} g(x)=x\ \ (x\notin\{1/m:m\ge2\}),\qquad g(1/2)=0,\qquad g(1/(m+1))=1/m\ \ (m\ge2), \end{equation*}
a Borel bijection of \((0,1)\) onto \([0,1)\) (it is the identity off a countable set). Then \(\varphi_k:=g\circ h_k+k-1\) maps \(I_k\) injectively onto \([k-1,k)\), and \(f:=\varphi_k\) on \(I_k\) is an injective Borel map of \(\mathbb{R}^1\setminus C\) onto \([0,+\infty)\), since \(f^{-1}(B)=\bigcup_k\varphi_k^{-1}(B)\) for Borel \(B\).
On \(C\): let \(V\subset(-1,0)\) contain one point of \((-1,0)\) from each coset of \(\mathbb{Q}\); it is not measurable, because the disjoint translates \(V+q\), \(q\in\mathbb{Q}\cap(-1,1)\), lie in \((-2,1)\) and cover \((-1,0)\). Put \(K:=C\cap[0,1/3]\). All four of \(K\), \(C\setminus K\supset C\cap[2/3,1]\), \(V\) and \((-\infty,0)\setminus V\supset(-\infty,-1]\) have cardinality \(\mathfrak{c}\), so choose bijections \(K\to V\) and \(C\setminus K\to(-\infty,0)\setminus V\) and let \(f\) be their union on \(C\).
Thus \(f\) is a bijection of \(\mathbb{R}^1\) onto \([0,+\infty)\cup(-\infty,0)=\mathbb{R}^1\). For Borel \(B\), \(f^{-1}(B)\) is the union of a Borel set and a subset of the null set \(C\), hence Lebesgue measurable by completeness. But \(K\) is compact, and
\begin{equation*} (f^{-1})^{-1}(K)=f(K)=V \end{equation*}
is not Lebesgue measurable, so \(f^{-1}\) is not Lebesgue measurable.
Prove that there exists a Borel one-to-one function \(f\colon[0,1]\to[0,1]\) such that \(f(x)=x\) for all \(x\), with the exception of points of a countable set, but the inverse function is discontinuous at all points of \((0,1]\).
Take \(Z:=\{1/m:m\ge2\}\), \(T:=(\mathbb{Q}\cap(0,1])\setminus Z\), any bijection \(\theta\colon T\to Z\), and
\begin{equation*} f(x):=\theta(x)\ (x\in T),\qquad f(x):=\theta^{-1}(x)\ (x\in Z), \qquad f(x):=x\ \text{else}. \end{equation*}
Both \(T\) and \(Z\) are countably infinite (\(T\) contains all \(2/m\) with odd \(m\ge3\)), so \(\theta\) exists; and \(T\) is dense in \([0,1]\), since any interval \((a,b)\subset(0,1]\) contains infinitely many rationals but only the finitely many points \(1/m\) with \(m<1/a\).
Since \(f\) interchanges \(T\) and \(Z\) and fixes the rest, \(f\circ f=\mathrm{id}\): thus \(f\) is a bijection of \([0,1]\) onto itself with \(f^{-1}=f\), equal to \(x\) off the countable set \(T\cup Z\). It is Borel, because for Borel \(B\)
\begin{equation*} f^{-1}(B)=\big(B\setminus(T\cup Z)\big)\cup\{x\in T\cup Z:\ f(x)\in B\}, \end{equation*}
a Borel set together with a countable one.
Now fix \(y\in(0,1]\). Then \(f(y)>0\), since either \(f(y)=y\) or \(f(y)\in T\cup Z\subset(0,1]\). By density choose pairwise distinct \(t_k\in T\) with \(t_k\to y\); the values \(f(t_k)=\theta(t_k)=1/m_k\) are pairwise distinct, so \(m_k\to\infty\) and
\begin{equation*} f(t_k)\to0\ne f(y). \end{equation*}
Hence \(f^{-1}=f\) is discontinuous at every point of \((0,1]\).
(Aleksandrov, Ivanov) Let \(K\) be a compact set in \(\mathbb{R}^n\) such that the intersection of \(K\) with every straight line is a finite union of intervals (possibly degenerate). Prove the Jordan measurability of \(K\), i.e., the equality \(\lambda_n(\partial K)=0\), where \(\lambda_n\) is Lebesgue measure.
Only a partial proof is given: the structural consequences of the hypothesis are established in full, and the theorem is reduced to a threading construction that is not carried out here. Since \(K\) is closed, \(\lambda_n(\partial K)=\lambda_n(K)-\lambda_n(\mathrm{int}\,K)\), so the assertion says that almost every point of \(K\) is interior; for \(n=1\) it is trivial, \(K\) being a finite union of intervals.
Lemma 1 (dichotomy along rays). For \(x\in K\) and a unit vector \(e\), exactly one of (a) \(x+te\in K\) for all \(t\in[0,\delta]\), (b) \(x+te\notin K\) for all \(t\in(0,\delta]\) holds for some \(\delta>0\). Indeed \(A:=\{t:x+te\in K\}\) is compact and a finite union of intervals, so \(A=\bigcup_{i\le m}[a_i,b_i]\) with \(b_i<a_{i+1}\) and \(0\in[a_i,b_i]\) for some \(i\); if \(b_i>0\) then (a) holds with \(\delta=b_i\), and otherwise \(A\cap(0,a_{i+1})=\varnothing\) (or \(A\cap(0,\infty)=\varnothing\) when \(i=m\)) gives (b). Finiteness of the number of intervals is exactly what is used here.
Put \(D^m_x:=\{e\in S^{n-1}:x+te\in K\ \forall t\in[0,1/m]\}\), closed because \(K\) is, and \(D_x:=\bigcup_mD^m_x=\{e:\text{(a) holds}\}\), a Borel subset of the sphere.
Lemma 2 (the density exists at every point of \(K\)). With \(\sigma\) the surface measure on \(S^{n-1}\),
\begin{equation*} \theta(x):=\lim_{r\to0}\frac{\lambda_n(K\cap B(x,r))}{\lambda_n(B(x,r))} =\frac{\sigma(D_x)}{\sigma(S^{n-1})},\qquad x\in K . \end{equation*}
Polar coordinates centred at \(x\) (Section 3.7) write the ratio as \(\sigma(S^{n-1})^{-1}\int_{S^{n-1}}\phi_r\,d\sigma\) with \(\phi_r(e)=nr^{-n}\int_0^rI_K(x+se)s^{n-1}ds\in[0,1]\), and by Lemma 1 \(\phi_r\to I_{D_x}\) pointwise, so dominated convergence applies.
Corollary 3. By the Lebesgue density theorem \(\theta=1\) a.e. on \(K\), so \(\sigma(S^{n-1}\setminus D_x)=0\) for almost every \(x\in K\), hence for almost every point of \(\partial K\) as well whenever \(\lambda_n(\partial K)>0\).
Lemma 4 (a directional version, without the density theorem). For a fixed unit vector \(e\), the set \(A_e:=\{x:\exists\delta>0,\ x+te\in K\ \forall|t|\le\delta\}\) is the union of the closed sets \(F_m:=\{x:x+te\in K\ \forall|t|\le1/m\}\), hence Borel, and \(\lambda_n(\partial K\setminus A_e)=0\). Indeed, writing \(x=y+te\) with \(y\in e^{\perp}\), the section \(A(y):=\{t:y+te\in K\}\) is a compact finite union of intervals, so all but finitely many of its points are interior to it and therefore lie in \(A_e\); every section of \(\partial K\setminus A_e\) is thus finite, and Fubini’s theorem (Theorem 3.4.4) applies.
Remark 5 (no pointwise criterion at boundary points). The condition (G) \(\sigma(S^{n-1}\setminus D_x)>0\) for every \(x\in\partial K\), which with Corollary 3 would finish the proof, is false. For \(n=2\) let \(B\) be the closed unit disc and
\begin{equation*} U:=\{(u,v):\ 0<u<1,\ v^2<u^4\},\qquad K:=B\setminus U . \end{equation*}
Parametrizing a line by \(t\mapsto(u(t),v(t))\) with affine \(u,v\) gives \(L\cap U=\{0<u<1\}\cap\{p>0\}\) with \(p:=u^4-v^2\) a polynomial, so \(L\cap U\), and hence \(L\cap K\), is a finite union of intervals; \((t,0)\in U\) for \(0<t<1\) gives \(0\in\partial K\); and for \(e=(\cos\varphi,\sin\varphi)\) with \(\sin\varphi\ne0\) the inclusion \(te\in U\) would force \(t^2\sin^2\varphi<t^4\cos^4\varphi\), which fails for \(0<t\le|\sin\varphi|/\cos^2\varphi\). So \(S^1\setminus D_0=\{(1,0)\}\) is \(\sigma\)-null although \(0\in\partial K\); multiplying by a cube gives the same in \(\mathbb{R}^n\).
Lemma 6 (grid structure at density points). Suppose \(\lambda_n(\partial K)>0\). By Lemma 4 applied to \(e_1,\dots,e_n\), almost every point of \(\partial K\) lies in \(\bigcap_iA_{e_i}\), an increasing union of closed sets; by inner regularity there are a compact \(N^{\prime}\subset\partial K\) with \(\lambda_n(N^{\prime})>0\) and \(\delta>0\) with \(x+[-\delta,\delta]e_i\subset K\) for all \(x\in N^{\prime}\) and all \(i\). Let \(x_0\) be a density point of \(N^{\prime}\), let \(0<r\le\delta/2\), let \(Q\) be the closed cube of half-edge \(r\) centred at \(x_0\), and \(P_i:=\pi_i(N^{\prime}\cap Q)\). Then:
(i) for \(y\in P_i\) the whole chord \(Q\cap\pi_i^{-1}(y)\) lies in \(K\), the segment through the corresponding \(x\in N^{\prime}\cap Q\) having half-length \(\delta\ge2r\);
(ii) \(Q\setminus K\subset Q\cap\bigcap_i\pi_i^{-1}(P_i^c)\) and \(\lambda_{n-1}(\pi_i(Q)\setminus P_i)\le\varepsilon( r)(2r)^{n-1}\) with \(\varepsilon( r):=1-\lambda_n(N^{\prime}\cap Q)(2r)^{-n}\to0\), since \(Q\cap\pi_i^{-1}(P_i^c)\) misses \(N^{\prime}\cap Q\) (divide its measure by the chord length \(2r\));
(iii) for \(x\in N^{\prime}\) with \(|x-x_0|_\infty\le r/2\), the projection \(\pi_i(x)\) lies in the boundary of \(P_i\) in \(e_i^{\perp}\): \(x\) is a limit of points \(z_k\notin K\) eventually in \(Q\), and \(\pi_i(z_k)\notin P_i\) by (ii) while \(\pi_i(x)\in P_i\).
So positive boundary measure would force this fat-Cantor grid at almost every boundary point simultaneously; a proof must exclude it using the finiteness of the number of components of \(K\cap L\) along the lines of the slots.
Lemma 7 (moat lemma). Let \(x\in K\), \(r>0\), and suppose the connected component \(C\) of \(x\) in \(K_r:=K\cap\overline{B}(x,r)\) misses \(\partial B(x,r)\). Then every unit vector \(e\) admits \(t\in(0,r)\) with \(x+te\notin K\). In a compact metric space the component of a point is its quasicomponent, so, \(F:=K_r\cap\partial B(x,r)\) being compact and disjoint from \(C\), finitely many clopen subsets of \(K_r\) containing \(x\) intersect in a clopen \(A\ni x\) with \(A\cap\partial B(x,r)=\varnothing\), i.e. \(A\subset B(x,r)\). Then \(A\) is compact and open in \(K\), so either \(K=A\subset B(x,r)\) and the claim is clear, or \(d:=\mathrm{dist}(A,K\setminus A)>0\). With \(t^{*}:=\max\{t\in[0,r]:x+te\in A\}<r\), each \(t\in(t^{*},\min(t^{*}+d,r))\) gives \(x+te\notin A\) by maximality and \(\mathrm{dist}(x+te,A)<d\), hence \(x+te\notin K\).
Corollary 8. Let \(S\) be the set of \(x\in K\) whose component in \(K\cap\overline{B}(x,2^{-k})\) misses \(\partial B(x,2^{-k})\) for infinitely many \(k\). Then \(\lambda_n(S)=0\), and for almost every \(x\in K\) there is \(m(x)\) such that for every \(r\le2^{-m(x)}\) the component of \(x\) in \(K\cap\overline{B}(x,r)\) contains points at distance \(>r/2\) from \(x\).
(i) On a line \(L\), Lemma 7 applied to \(\pm e\) at radii \(2^{-k_i}\downarrow0\) puts points of \(L\setminus K\) on both sides of each \(x\in S\cap L\) arbitrarily near, so \(\{x\}\) is a component of the finite union of intervals \(K\cap L\); hence \(S\cap L\) has at most \(N(L)<\infty\) points. (ii) \(S=\bigcap_m\bigcup_{k\ge m}(K\setminus D_k)\), where \(D_k\) is the set of \(x\in K\) whose component in \(K\cap\overline{B}(x,2^{-k})\) meets the sphere; \(D_k\) is closed, for if \(x_j\to x\) in \(D_k\) then Blaschke selection gives a Hausdorff limit \(C\) of the corresponding continua, and \(C\) is a continuum inside \(K\cap\overline{B}(x,2^{-k})\) containing \(x\) and a point of the sphere (a splitting of \(C\) into pieces at mutual distance \(3\varepsilon\) would split \(C_j\) for large \(j\)). So \(S\) is Borel. (iii) Every line meets \(S\) in a finite set, so Fubini’s theorem (Theorem 3.4.4) gives \(\lambda_n(S)=0\). For the last claim take the largest \(k\) with \(2^{-k}\le r\); then \(k\ge m(x)\) and \(2^{-k}>r/2\).
What remains. Fix \(R\) with \(K\subset B(0,R)\) and let \(\mathcal{L}\) be the compact metric space of lines \(L_{e,y}(t)=y+te\), \(e\in S^{n-1}\), \(y\in e^{\perp}\), \(|y|\le R\).
Proposition 9 (reduction to a single bad line). Suppose that for every \(j\) there is a nonempty compact \(T_j\subset\mathcal{L}\) with \(T_{j+1}\subset T_j\) such that every \(L\in T_j\) admits parameters
\begin{equation*} t_1<u_1<t_2<\dots<u_{j-1}<t_j\quad\text{with}\quad L(t_i)\notin K,\ \ L(u_i)\in K . \end{equation*}
Then \(K\) violates the hypothesis: by the finite intersection property there is \(L\in\bigcap_jT_j\), and for each \(j\) the points \(L(u_1),\dots,L(u_{j-1})\) lie in pairwise distinct components of \(K\cap L\), consecutive ones being separated by \(L(t_{i+1})\notin K\).
Assuming \(\lambda_n(\partial K)>0\), all the raw material is available: balls disjoint from \(K\) arbitrarily near every boundary point, giving the stable conditions \(L(t_i)\notin K\); the uniform grid of Lemma 6; and the continua of Corollary 8 as blockers for the unstable conditions \(L(u_i)\in K\), a connected set with points strictly on both sides of a planar line being forced to meet it. What I cannot do is the recursive threading of balls and blockers along a shrinking compact family of lines: the continuum supplied near a prescribed point has unknown shape and may be a segment nearly parallel to the family, while the points where blockers are guaranteed are located only up to sets of small measure. For \(n\ge3\) there is a further obstruction of principle, since a connected set with points on both sides of a line need not meet the line, and nothing above supplies \((n-1)\)-dimensional sheets inside \(K\). This is the content of the papers of Aleksandrov [14] and Ivanov [451] cited in the statement. (A fully rigorous proof of the remaining step is beyond the scope of this page; see the reference given in the book.)
(\(\circ\)) Let \(f\in L^2(\mathbb{R}^n)\), where we consider the space of complex-valued functions. Let \(f_j(x)=f(x)\) if \(|x_i|\le j\), \(i=1,\dots,n\), and \(f_j(x)=0\) at all other points.
(i) (Plancherel’s theorem) Show that the sequence of functions \(\widehat{f_j}\) converges in \(L^2(\mathbb{R}^n)\) to some function, called the Fourier transform of \(f\) in \(L^2(\mathbb{R}^n)\) and denoted by \(\widehat{f}\).
(ii) Show that the mapping \(f\mapsto\widehat{f}\) is a bijection of \(L^2(\mathbb{R}^n)\) and
\begin{equation*} \int_{\mathbb{R}^n}f(x)\overline{g(x)}\,dx=\int_{\mathbb{R}^n}\widehat{f}(x)\overline{\widehat{g}(x)}\,dx\quad\text{for all }f,g\in L^2(\mathbb{R}^n). \end{equation*}
(iii) Show that the Fourier transform defined in (i) is uniquely determined by the property that on \(L^2(\mathbb{R}^n)\cap L^1(\mathbb{R}^n)\) it coincides with the previously defined Fourier transform and satisfies the equality in (ii).
(iv) Show that there exists a sequence \(j_k\to\infty\) such that \(\widehat{f_{j_k}}(x)\to\widehat{f}(x)\) a.e.
Everything follows from the Plancherel identity \(\|\widehat u\|_2=\|u\|_2\) on \(L^1\cap L^2\), established first. Throughout \(\widehat u\) is as in Definition 3.8.1, \(Q_j:=\{x:|x_i|\le j\}\) and \(f_j=fI_{Q_j}\). By the Cauchy-Bunyakowsky inequality \(\int_{Q_j}|f|dx\le\|f\|_2\lambda_n(Q_j)^{1/2}<\infty\), so \(f_j\in L^1\cap L^2\), while \(\|f-f_j\|_2^2=\int_{\mathbb{R}^n\setminus Q_j}|f|^2dx\to0\); thus \(f_j\to f\) in \(L^2\) and \(L^1\cap L^2\) is dense in \(L^2\). Also \(C_0^\infty(\mathbb{R}^n)\) is dense in \(L^2\), by Corollary 4.2.2 and mollification (Section 3.9).
For \(\psi\in C_0^\infty\) the Parseval equality (3.8.3) of Corollary 3.8.11 gives \(\|\widehat\psi\|_2=\|\psi\|_2\), all integrals being finite since \(\widehat\psi\) decays faster than any power. For \(u\in L^1\cap L^2\) choose \(\psi_k\in C_0^\infty\) with \(\psi_k\to u\) in \(L^1\) and in \(L^2\) at once: given \(j\) and \(\varepsilon>0\), take \(\psi\in C_0^\infty\) with \(\|\psi-u_j\|_2<\varepsilon\), where \(u_j:=uI_{Q_j}\), and \(\chi\in C_0^\infty\) with \(0\le\chi\le1\), \(\chi=1\) on \(Q_j\), \(\chi=0\) off \(Q_{j+1}\); then \(\chi u_j=u_j\) gives
\begin{equation*} \|\chi\psi-u_j\|_2\le\varepsilon,\qquad \|\chi\psi-u_j\|_1\le\lambda_n(Q_{j+1})^{1/2}\varepsilon , \end{equation*}
and \(u_j\to u\) in both norms, so a diagonal choice works. Now \(\|\widehat{\psi_k}-\widehat u\|_\infty\le(2\pi)^{-n/2}\|\psi_k-u\|_1\to0\) by the bound of Proposition 3.8.4, while \(\|\widehat{\psi_k}-\widehat{\psi_l}\|_2=\|\psi_k-\psi_l\|_2\to0\), so by completeness of \(L^2\) (Theorem 4.1.3) \(\widehat{\psi_k}\to w\) in \(L^2\) and a.e. along a subsequence; comparing limits, \(w=\widehat u\) and
\begin{equation*} \|\widehat u\|_2=\lim_k\|\widehat{\psi_k}\|_2=\lim_k\|\psi_k\|_2=\|u\|_2, \qquad u\in L^1\cap L^2. \tag{1} \end{equation*}
(i) Since \(f_j-f_k\in L^1\cap L^2\) and the classical transform is linear on \(L^1\), (1) gives \(\|\widehat{f_j}-\widehat{f_k}\|_2=\|f_j-f_k\|_2\to0\), so \(\widehat{f_j}\) converges in \(L^2\) to a function \(\widehat f\). The limit is independent of the approximating sequence: \(\|\widehat{h_k}-\widehat{f_j}\|_2=\|h_k-f_j\|_2\) for any \(h_k\in L^1\cap L^2\) with \(h_k\to f\) in \(L^2\); taking \(h_k\equiv f\) shows that for \(f\in L^1\cap L^2\) the new \(\widehat f\) is the classical one.
(ii) The map \(F:f\mapsto\widehat f\) is linear (\((af+bg)I_{Q_j}=af_j+bg_j\) and \(L^2\)-limits) and isometric, \(\|Ff\|_2=\lim_j\|f_j\|_2=\|f\|_2\) by (1); hence injective and continuous. Polarization \((u,v)=\frac14\sum_{k=0}^3i^k\|u+i^kv\|^2\) then gives \(\int\widehat f\,\overline{\widehat g}\,dx=(Ff,Fg)=(f,g)=\int f\overline g\,dx\). The range is closed (a Cauchy image sequence has a Cauchy preimage sequence) and contains \(C_0^\infty\): for \(\psi\in C_0^\infty\) put \(v(y):=\widehat\psi(-y)\in L^1\cap L^2\); the change of variable \(z=-y\) and the inversion formula (3.8.2) of Corollary 3.8.9 give \(\widehat v=\psi\). By density of \(C_0^\infty\), \(F(L^2)=L^2\), so \(F\) is a bijection.
(iii) If \(T\) coincides with the classical transform on \(L^1\cap L^2\) and preserves inner products, then
\begin{equation*} \|Tu-Tv\|_2^2=\|u\|_2^2-2\operatorname{Re}(u,v)+\|v\|_2^2=\|u-v\|_2^2 , \end{equation*}
so \(T\) is continuous and agrees with \(F\) on the dense set \(L^1\cap L^2\); hence \(T=F\).
(iv) Choose \(j_k\to\infty\) with \(\|\widehat{f_{j_k}}-\widehat f\|_2\le2^{-k}\). By monotone convergence,
\begin{equation*} \int_{\mathbb{R}^n}\sum_{k\ge1}\big|\widehat{f_{j_k}}(x)-\widehat f(x)\big|^2dx =\sum_{k\ge1}\|\widehat{f_{j_k}}-\widehat f\|_2^2\le\sum_{k\ge1}4^{-k}<\infty , \end{equation*}
so the series converges a.e. and its terms tend to \(0\) a.e.
(\(\circ\)) The Laplace transform of a complex-valued function \(f\in L^2[0,+\infty)\) is defined by
\begin{equation*} Lf(s)=\int_0^{\infty}e^{-st}f(t)\,dt,\qquad s>0 . \end{equation*}
Show that \(Lf\in L^2[0,+\infty)\) and that \(\|Lf\|_2\le\sqrt{\pi}\,\|f\|_2\).
The bound comes from the Cauchy-Bunyakowsky splitting \(e^{-st}|f(t)|=(e^{-st/2}|f(t)|t^{1/4})(e^{-st/2}t^{-1/4})\) followed by Tonelli’s theorem. That splitting gives, for \(s>0\),
\begin{equation*} |Lf(s)|^2\le\int_0^{\infty}e^{-st}|f(t)|^2t^{1/2}dt \cdot\int_0^{\infty}e^{-st}t^{-1/2}dt, \end{equation*}
and the substitution \(u=st\) evaluates the second factor as \(\Gamma(1/2)s^{-1/2}=\sqrt{\pi}\,s^{-1/2}\), whence
\begin{equation*} |Lf(s)|^2\le\sqrt{\pi}\,s^{-1/2}\int_0^{\infty}e^{-st}|f(t)|^2t^{1/2}\,dt . \tag{1} \end{equation*}
Here \(Lf(s)\) is defined for every \(s>0\), since \(\int_0^{\infty}e^{-st}|f|\,dt\le\|f\|_2(2s)^{-1/2}<\infty\), and \(Lf\) is continuous on \((0,+\infty)\) by dominated convergence with the majorant \(e^{-s_0t}|f(t)|\), hence measurable.
Integrating (1) in \(s\) and interchanging the order by Tonelli’s theorem (Theorem 3.4.5; the integrand is nonnegative and jointly measurable),
\begin{equation*} \int_0^{\infty}|Lf(s)|^2ds\le\sqrt{\pi}\int_0^{\infty}|f(t)|^2t^{1/2} \Big(\int_0^{\infty}s^{-1/2}e^{-st}ds\Big)dt, \end{equation*}
and the inner integral is \(\Gamma(1/2)t^{-1/2}=\sqrt{\pi}\,t^{-1/2}\). Therefore
\begin{equation*} \int_0^{\infty}|Lf(s)|^2\,ds\le\pi\int_0^{\infty}|f(t)|^2\,dt=\pi\|f\|_2^2 , \end{equation*}
so \(Lf\in L^2[0,+\infty)\) and \(\|Lf\|_2\le\sqrt{\pi}\,\|f\|_2\).
Exercises 3.10.78–3.10.84
(\(\circ\)) Give an example of a function \(f \in L^1(\mathbb{R}^1)\) such that its Fourier transform is neither in \(L^1(\mathbb{R}^1)\) nor in \(L^2(\mathbb{R}^1)\), and an example of a function \(g\) in \(L^2(\mathbb{R}^1)\) such that its Fourier transform does not belong to \(L^1(\mathbb{R}^1)\).
Take \(f(x)=x^{-1/2}e^{-x}I_{(0,\infty)}(x)\) and \(g=I_{[-1,1]}\), with the normalization of Definition 3.8.1.
Here \(f\ge0\) is Borel with \(\int f\,dx=\Gamma(1/2)=\sqrt{\pi}\), so \(f\in L^1(\mathbb{R}^1)\). For \(\operatorname{Re}a>0\) the integral \(F(a):=\int_0^{\infty}x^{-1/2}e^{-ax}dx\) converges absolutely and is holomorphic there, differentiation under the integral sign being dominated by \(x^{1/2}e^{-\delta x}\in L^1(0,\infty)\) on \(\{\operatorname{Re}a\ge\delta\}\); the substitution \(x=t/a\) gives \(F(a)=\sqrt{\pi}\,a^{-1/2}\) for real \(a>0\), and two holomorphic functions on the connected right half-plane agreeing on the positive axis agree throughout (principal branch). With \(a=1+iy\),
\begin{equation*} \widehat f(y)=(2\pi)^{-1/2}\sqrt{\pi}\,(1+iy)^{-1/2},\qquad |\widehat f(y)|=\tfrac{1}{\sqrt2}(1+y^2)^{-1/4}. \end{equation*}
Thus \(|\widehat f(y)|\) is of order \(|y|^{-1/2}\) and \(|\widehat f(y)|^2\) of order \(|y|^{-1}\) at infinity, so \(\widehat f\notin L^1(\mathbb{R}^1)\) and \(\widehat f\notin L^2(\mathbb{R}^1)\).
Next \(g\in L^1\cap L^2\), so its \(L^2\)-transform of Exercise 3.10.76 is the classical one,
\begin{equation*} \widehat g(y)=(2\pi)^{-1/2}\frac{2\sin y}{y}\ \ (y\ne0), \qquad \widehat g(0)=2(2\pi)^{-1/2}, \end{equation*}
and it is not integrable: \(\int_{k\pi}^{(k+1)\pi}|\sin y|y^{-1}dy\ge2/((k+1)\pi)\) for every \(k\ge1\), and \(\sum_k1/(k+1)\) diverges.
Find a uniformly continuous function \(f\) on \(\mathbb{R}^1\) that satisfies the condition \(\lim_{|x|\to\infty} f(x) = 0\), but is not the Fourier transform of a function from \(L^1(\mathbb{R}^1)\).
Take the odd function \(f\) with \(f(x)=1/\ln x\) for \(x\ge2\) and \(f(x)=x/(2\ln2)\) for \(0\le x\le2\).
The two formulas agree at \(x=2\) and \(f(0)=0\), so \(f\) is continuous and odd, bounded by \(1/\ln2\), with \(f(x)\to0\) as \(|x|\to\infty\); and a continuous function with finite limits at \(\pm\infty\) is uniformly continuous (given \(\varepsilon\), take \(M\) with \(|f|<\varepsilon/2\) off \([-M,M]\) and use uniform continuity on \([-M-1,M+1]\)). So \(f\) satisfies all the necessary conditions of Proposition 3.8.4.
Claim: if \(h\in L^1(\mathbb{R}^1)\) and \(\widehat h\) is odd, then \(\sup_{R>1}|\int_1^{R}\widehat h(y)y^{-1}dy|<\infty\). Replacing \(h\) by its odd part \(h^{-}(x):=\frac12(h(x)-h(-x))\in L^1\), which has \(\widehat{h^{-}}=\widehat h\) because \(\widehat{h(-\cdot)}(y)=\widehat h(-y)\), we may assume \(h\) odd, so that \(\widehat h(y)=-i(2\pi)^{-1/2}\int h(x)\sin(xy)dx\). Since \(\int_1^R\int|h(x)\sin(xy)|y^{-1}dxdy\le\|h\|_{L^1}\ln R<\infty\), Fubini’s theorem (Theorem 3.4.5) gives
\begin{equation*} \int_1^{R}\frac{\widehat h(y)}{y}dy =-i(2\pi)^{-1/2}\int_{\mathbb{R}^1}h(x)\Big[\int_1^{R}\frac{\sin(xy)}{y}dy\Big]dx , \end{equation*}
and for \(x>0\) the substitution \(u=xy\) turns the inner integral into \(\mathrm{Si}(Rx)-\mathrm{Si}(x)\), of modulus at most \(2C_0\) with \(C_0:=\sup_{T\ge0}|\mathrm{Si}(T)|<\infty\) (\(\mathrm{Si}\) is continuous with limit \(\pi/2\)); it is odd in \(x\) and vanishes at \(x=0\). Hence the left-hand side is bounded by \(2C_0(2\pi)^{-1/2}\|h\|_{L^1}\) for all \(R>1\).
Our \(f\) is odd, and the substitution \(u=\ln y\) gives
\begin{equation*} \int_1^{R}\frac{f(y)}{y}dy=\frac{1}{2\ln2}+\ln\ln R-\ln\ln2\longrightarrow+\infty , \end{equation*}
so no \(h\in L^1(\mathbb{R}^1)\) satisfies \(\widehat h=f\).
(\(\circ\)) For \(f\) in the complex space \(L^2(\mathbb{R}^1)\) we set
\begin{equation*} \mathcal{H}_{\varepsilon}f(x) = \frac{1}{\pi}\int_{-\infty}^{+\infty}\frac{y}{y^2+\varepsilon^2}\,f(x-y)\,dy . \end{equation*}
Show that there exists the limit \(\mathcal{H}_0 f := \lim_{\varepsilon\to 0}\mathcal{H}_{\varepsilon}f\) in \(L^2(\mathbb{R}^1)\) as \(\varepsilon\to 0\); then \(\mathcal{H}_0 f\) is called the Hilbert transform of \(f\). In addition, one has \(\mathcal{H}_0 = \mathcal{F}^{-1}\mathcal{M}\mathcal{F}\), where \(\mathcal{F}\) is the Fourier transform in \(L^2(\mathbb{R}^1)\) and \(\mathcal{M}g(x) = i(2\pi)^{-1/2}(\operatorname{sign}x)g(x)\).
The multiplier is \(\mathcal{M}g(x)=-i(\operatorname{sign}x)g(x)\); the printed formula omits the factor \((2\pi)^{1/2}\) of Corollary 3.9.3 and carries the opposite sign, and we prove the corrected statement. Throughout \(\mathcal{F}\) is the unitary \(L^2\)-transform of Exercise 3.10.76, agreeing with the classical one on \(L^1\cap L^2\).
Put \(g_\varepsilon(y):=\pi^{-1}y/(y^2+\varepsilon^2)\), which lies in \(L^2(\mathbb{R}^1)\) but not in \(L^1(\mathbb{R}^1)\), so that \(\mathcal{H}_\varepsilon f=g_\varepsilon*f\) converges absolutely at every \(x\) by the Cauchy-Bunyakowsky inequality.
Lemma. For \(u,v\in L^2(\mathbb{R}^1)\) and every \(x\), \((u*v)(x)=\int\widehat u(\xi)\widehat v(\xi)e^{ix\xi}d\xi\), absolutely since \(\widehat u\widehat v\in L^1\). Indeed, with \(b(y):=\overline{u(x-y)}\) one has \((u*v)(x)=\int v\overline b\,dy\), and \(w(y):=u(x-y)\) has \(\widehat w(\xi)=e^{-ix\xi}\widehat u(-\xi)\) (a change of variable for \(u\in L^1\cap L^2\), then \(L^2\)-continuity of both sides), so \(\widehat b(\xi)=\overline{\widehat w(-\xi)}=e^{-ix\xi}\overline{\widehat u(\xi)}\) and Parseval gives the claim.
Next \(m_\varepsilon(x):=-i(\operatorname{sign}x)e^{-\varepsilon|x|}\in L^1\cap L^2\), and since \(\int_0^{\infty}e^{-\varepsilon x}\sin(xy)dx=y/(\varepsilon^2+y^2)\) and \(2(2\pi)^{-1/2}=(2\pi)^{1/2}\pi^{-1}\),
\begin{equation*} \check{m_\varepsilon}(y)=(2\pi)^{-1/2}(-i)\int_0^{\infty}e^{-\varepsilon x} \big(e^{ixy}-e^{-ixy}\big)dx=(2\pi)^{1/2}g_\varepsilon(y). \end{equation*}
As \(m_\varepsilon\in L^1\cap L^2\), this classical inverse transform is also the \(L^2\) one (Exercise 3.10.76(iii) for the inverse transform), so applying the unitary \(\mathcal{F}\) gives \((2\pi)^{1/2}\widehat{g_\varepsilon}=m_\varepsilon\). Hence the Lemma with \(u=g_\varepsilon\), \(v=f\) reads
\begin{equation*} \mathcal{H}_\varepsilon f(x) =\int_{\mathbb{R}^1}\widehat{g_\varepsilon}(\xi)\widehat f(\xi)e^{ix\xi}d\xi =\big(m_\varepsilon\widehat f\big)^{\vee}(x), \end{equation*}
and \(m_\varepsilon\widehat f\) lies in \(L^2\) (as \(|m_\varepsilon|\le1\)) and in \(L^1\) (a product of two \(L^2\) functions up to a constant), so its classical and \(L^2\) inverse transforms agree and \(\mathcal{H}_\varepsilon f\in L^2\) with \(\mathcal{F}\mathcal{H}_\varepsilon f=m_\varepsilon\widehat f\).
Finally \(m_\varepsilon\to m_0:=-i\operatorname{sign}\) pointwise off the origin with \(|m_\varepsilon\widehat f-m_0\widehat f|\le2|\widehat f|\in L^2\), so dominated convergence and the isometry of \(\mathcal{F}^{-1}\) give
\begin{equation*} \|\mathcal{H}_\varepsilon f-\mathcal{F}^{-1}(m_0\widehat f)\|_{L^2} =\|m_\varepsilon\widehat f-m_0\widehat f\|_{L^2}\longrightarrow0 . \end{equation*}
Thus \(\mathcal{H}_0f:=\lim_{\varepsilon\to0}\mathcal{H}_\varepsilon f\) exists in \(L^2(\mathbb{R}^1)\) and equals \(\mathcal{F}^{-1}\mathcal{M}\mathcal{F}f\).
Suppose that \(f \in L^1(\mathbb{R}^1)\), \(\varphi \in L^{\infty}(\mathbb{R}^1)\) and that, for some \(\beta>0\) and all \(x\), we have \(\varphi(x+\beta) = -\varphi(x)\) (e.g., \(\varphi(x) = \sin x\), \(\beta = \pi\)). Show that
\begin{equation*} \lim_{n\to\infty}\int_{-\infty}^{+\infty} f(x)\varphi(nx)\,dx = 0 . \end{equation*}
The primitive \(\Phi(t):=\int_0^t\varphi(y)dy\) is bounded, which settles the case of an interval, and density does the rest. Fix a representative with \(\varphi(x+\beta)=-\varphi(x)\) at every point and put \(K:=\|\varphi\|_{L^{\infty}}\); then \(\varphi\) is \(2\beta\)-periodic, and each \(x\mapsto\varphi(nx)\) is measurable with \(|\varphi(nx)|\le K\) a.e., dilations preserving measurable and null sets (Corollary 3.6.4).
For every \(T\) the change of variable \(y=x+\beta\) gives \(\int_T^{T+\beta}\varphi=-\int_{T+\beta}^{T+2\beta}\varphi\), hence \(\int_T^{T+2\beta}\varphi=0\); so \(\Phi\) is \(2\beta\)-periodic and continuous, and \(C:=\sup|\Phi|\le2\beta K<\infty\). For \(f=I_{[a,b]}\), substituting \(y=nx\),
\begin{equation*} \int_{\mathbb{R}^1}I_{[a,b]}(x)\varphi(nx)dx=\frac1n\int_{na}^{nb}\varphi(y)dy =\frac{\Phi(nb)-\Phi(na)}{n}, \end{equation*}
of modulus at most \(2C/n\to0\); by linearity \(T_n(h):=\int h(x)\varphi(nx)dx\to0\) for every step function \(h\).
These functionals satisfy \(|T_n(u)|\le K\|u\|_{L^1}\) for all \(n\), and step functions are dense in \(L^1(\mathbb{R}^1)\): simple integrable functions are dense by the construction of the integral, and any set of finite measure is approximated in measure by a finite union of intervals, Lebesgue measure being the Caratheodory extension from that algebra (Theorem 1.5.6). Given \(\delta>0\), choose a step function \(h\) with \(\|f-h\|_{L^1}<\delta\); then
\begin{equation*} \limsup_{n\to\infty}|T_n(f)|\le \limsup_{n\to\infty}\big(K\delta+|T_n(h)|\big)=K\delta , \end{equation*}
and \(\delta\) was arbitrary, so \(\int f(x)\varphi(nx)dx\to0\).
Let us define the standard surface measure \(\sigma_{n-1}\) on the unit sphere \(S^{n-1}\) in \(\mathbb{R}^n\) by the equality
\begin{equation*} \sigma_{n-1}(B) := n\lambda_n\big(\{x:\ 0<|x|\le 1,\ x/|x|\in B\}\big),\qquad B\in\mathcal{B}(S^{n-1}). \end{equation*}
Show that \(\sigma_{n-1}\) is a unique Borel measure on \(S^{n-1}\) that satisfies the equality
\begin{equation*} r^{n-1}dr\otimes\sigma_{n-1} = \lambda_n\circ\Phi^{-1}, \end{equation*}
where \(\Phi\colon \mathbb{R}^n\setminus\{0\}\to(0,\infty)\times S^{n-1}\), \(\Phi(x) = \big(|x|,\ x/|x|\big)\). In particular, if \(f\) is integrable over \(\mathbb{R}^n\), then one has
\begin{equation*} \int_{\mathbb{R}^n} f(x)\,dx = \int_0^{\infty}\int_{S^{n-1}} r^{n-1}f(ry)\,\sigma_{n-1}(dy)\,dr . \end{equation*}
The measure is \(\sigma_{n-1}\) itself, and the whole content is the sector formula \(\lambda_n(S(a,b;B))=(\int_a^br^{n-1}dr)\sigma_{n-1}(B)\).
Write \(S(a,b;B):=\{x:a<|x|\le b,\ x/|x|\in B\}\), a Borel set because \(x\mapsto x/|x|\) is continuous, and \(C_t(B):=S(0,t;B)\). Countable additivity of \(\lambda_n\) on the disjoint sets \(C_1(B_j)\) makes \(\sigma_{n-1}\) a Borel measure, finite since \(\sigma_{n-1}(S^{n-1})=n\lambda_n(\{|x|\le1\})\). Since \(C_t(B)=t\,C_1(B)\), Corollary 3.6.4 with \(L=tI\) gives \(\lambda_n(C_t(B))=t^n\sigma_{n-1}(B)/n\), whence for \(0<a<b<\infty\)
\begin{equation*} \lambda_n\big(S(a,b;B)\big)=\frac{b^n-a^n}{n}\sigma_{n-1}(B) =\Big(\int_a^br^{n-1}dr\Big)\sigma_{n-1}(B). \tag{1} \end{equation*}
Existence. \(\Phi\) is a homeomorphism onto \((0,\infty)\times S^{n-1}\) with inverse \((r,y)\mapsto ry\), so \(\lambda_n\circ\Phi^{-1}\) is a Borel measure there, \(\sigma\)-finite because it gives \((0,N]\times S^{n-1}\) the value \(\lambda_n(\{0<|x|\le N\})\). Both factors are separable metric spaces with countable bases, so the Borel \(\sigma\)-algebra of the product is the product \(\sigma\)-algebra. As \(\Phi^{-1}((a,b]\times B)=S(a,b;B)\), by (1) the two measures agree on the class \(\mathcal{E}\) of rectangles \((a,b]\times B\), which is closed under finite intersections; both are finite and equal on \(X_k:=(1/k,k]\times S^{n-1}\in\mathcal{E}\), so after normalization Lemma 1.9.4 gives equality on \(\mathcal{B}(X_k)\), and \(X_k\uparrow(0,\infty)\times S^{n-1}\).
Uniqueness. If \(r^{n-1}dr\otimes\tau=\lambda_n\circ\Phi^{-1}\), testing on \((1/2,1]\times B\) and using (1) gives \(\tau(B)=\sigma_{n-1}(B)\), since \(\int_{1/2}^1r^{n-1}dr=(1-2^{-n})/n\in(0,\infty)\).
Integration formula. For integrable \(f\) we may integrate over \(\mathbb{R}^n\setminus\{0\}\), and Theorem 3.6.1 on image measures applied to \(g:=f\circ\Phi^{-1}\), i.e. \(g(r,y)=f(ry)\), followed by Fubini’s theorem (Theorems 3.4.4 and 3.4.5, both measures being \(\sigma\)-finite), gives
\begin{equation*} \int_{\mathbb{R}^n}f(x)\,dx =\int_0^{\infty}\int_{S^{n-1}}r^{n-1}f(ry)\,\sigma_{n-1}(dy)\,dr . \end{equation*}
(\(\circ\)) (i) Show that \(\sigma_{n-1}(S^{n-1}) = 2\pi^{n/2}/\Gamma(n/2)\).
(ii) Let \(c_k\) be the volume of a ball of radius \(1\) in \(\mathbb{R}^k\). Show that
\begin{equation*} c_n = \pi^{n/2}/\Gamma(1+n/2),\qquad c_{2k} = \pi^k/k!,\qquad c_{2k+1} = 2^{2k+1}k!\,\pi^k/(2k+1)! . \end{equation*}
(i) Compare two evaluations of the Gaussian integral. By Example 3.8.2 and Fubini’s theorem \(\int_{\mathbb{R}^n}e^{-|x|^2}dx=\pi^{n/2}\), while the polar formula of Exercise 3.10.82 and the substitution \(s=r^2\) give
\begin{equation*} \int_{\mathbb{R}^n}e^{-|x|^2}dx=\sigma_{n-1}(S^{n-1})\int_0^{\infty}r^{n-1}e^{-r^2}dr =\sigma_{n-1}(S^{n-1})\,\tfrac12\Gamma(n/2), \end{equation*}
so \(\sigma_{n-1}(S^{n-1})=2\pi^{n/2}/\Gamma(n/2)\).
(ii) Taking \(B=S^{n-1}\) in the definition, \(\sigma_{n-1}(S^{n-1})=n\lambda_n(\{|x|\le1\})=nc_n\), so by (i) and \(\Gamma(1+n/2)=(n/2)\Gamma(n/2)\),
\begin{equation*} c_n=\frac1n\cdot\frac{2\pi^{n/2}}{\Gamma(n/2)}=\frac{\pi^{n/2}}{\Gamma(1+n/2)} . \end{equation*}
For \(n=2k\) this is \(\pi^k/\Gamma(1+k)=\pi^k/k!\). For \(n=2k+1\), induction from \(\Gamma(1/2)=\sqrt{\pi}\) and \(\Gamma(s+1)=s\Gamma(s)\) gives \(\Gamma(k+3/2)=(2k+1)!!\,2^{-k-1}\sqrt{\pi}\) (Check!), and \((2k+1)!!=(2k+1)!/(2^kk!)\), so
\begin{equation*} c_{2k+1}=\frac{\pi^{k+1/2}2^{k+1}}{(2k+1)!!\,\sqrt{\pi}} =\frac{2^{k+1}\pi^k2^kk!}{(2k+1)!}=\frac{2^{2k+1}k!\,\pi^k}{(2k+1)!} . \end{equation*}
(Schechtman, Schlumprecht, Zinn [850]) Let \(\sigma\) be a probability measure on the unit sphere \(S\) in \(\mathbb{R}^n\) that is proportional to the standard surface measure and let \(\nu\) be a probability measure on \((0,+\infty)\). Let us consider the measure \(\mu = \nu\otimes\sigma\) on \(\mathbb{R}^n\) (more precisely, \(\mu\) is the image of \(\nu\otimes\sigma\) under the mapping \((t,y)\mapsto ty\)). Let \(\mathcal{U}_n\) be the group of all orthogonal matrices \(n\times n\) with its natural Borel \(\sigma\)-algebra and a Borel probability measure \(m\) with the following property: for each Borel set \(B\subset \mathcal{U}_n\) and each \(U\in\mathcal{U}_n\), letting \(L_U\) and \(R_U\) be the left and right multiplications in \(\mathcal{U}_n\) by \(U\), we have \(m\big(L_U(B)\big) = m\big(R_U(B)\big) = m(B)\) (the existence of such a measure – Haar’s measure – is proved in Chapter 9). Prove that, for all centrally symmetric convex Borel sets \(A\) and \(B\) in \(\mathbb{R}^n\), one has the inequality
\begin{equation*} \int_{\mathcal{U}_n}\mu\big(A\cap U(B)\big)\,m(dU) \ge \mu(A)\mu(B). \end{equation*}
In particular, if \(B\) is spherically symmetric, then \(\mu(A\cap B)\ge \mu(A)\mu(B)\). These inequalities are true for any probability measure \(\mu\) with a spherically symmetric density.
The point is that radial sections of centrally symmetric convex sets are nested, so that \(\nu(A_\varphi\cap B_\psi)=\min(\nu(A_\varphi),\nu(B_\psi))\ge\nu(A_\varphi)\nu(B_\psi)\). Write \(T(t,y)=ty\), so \(\mu=(\nu\otimes\sigma)\circ T^{-1}\); all spaces here are separable metric with countable bases, so Borel \(\sigma\)-algebras of products are product \(\sigma\)-algebras and Fubini’s theorem (Theorem 3.4.4) applies to the bounded Borel integrands below.
(1) For Borel \(C\) and \(\varphi\in S\) let \(C_\varphi:=\{r>0:r\varphi\in C\}\), which is the section of the Borel set \(T^{-1}( C)\); then \(\varphi\mapsto\nu(C_\varphi)\) is Borel and \(\mu( C)=\int_S\nu(C_\varphi)\sigma(d\varphi)\).
(2) If \(A=-A\) is convex, \(r\in A_\varphi\) and \(0<s<r\), then with \(\lambda:=(1+s/r)/2\in(0,1)\) we have \(s\varphi=\lambda(r\varphi)+(1-\lambda)(-r\varphi)\in A\), so each \(A_\varphi\) is downward closed in \((0,\infty)\). Two downward closed sets are nested (points \(r\in I\setminus J\) and \(s\in J\setminus I\) contradict downward closedness whichever is the larger, and \(r=s\) is impossible), so \(A_\varphi\cap B_\psi\) equals one of the two sets and
\begin{equation*} \nu(A_\varphi\cap B_\psi)=\min\big(\nu(A_\varphi),\nu(B_\psi)\big) \ge\nu(A_\varphi)\nu(B_\psi), \end{equation*}
the last step because \(\min(a,b)\ge\min(a,b)\max(a,b)=ab\) for \(a,b\in[0,1]\).
(3) The image of \(m\) under \(\Psi_\varphi\colon U\mapsto U^{-1}\varphi\) is \(\sigma\). First \(\sigma\) is \(\mathcal{U}_n\)-invariant: \(|Vx|=|x|\) and \((Vx)/|Vx|=V(x/|x|)\) turn the set defining \(\sigma_{n-1}(V(E))\) into the \(V\)-image of the set defining \(\sigma_{n-1}(E)\), and \(\lambda_n\) is orthogonally invariant (Corollary 3.6.4). Second, such a measure is unique: right invariance of \(m\) makes \(c:=\int_{\mathcal{U}_n}h(U\varphi)m(dU)\) independent of \(\varphi\) (pick \(W\) with \(W\varphi=\varphi^{\prime}\)), so any invariant Borel probability \(\tau\) satisfies \(\int_Sh\,d\tau=\int\int h(U\varphi)\tau(d\varphi)m(dU)=c\) by Fubini, for every bounded Borel \(h\). Third, \(\Psi_\varphi^{-1}(V(E))=R_V^{-1}(\Psi_\varphi^{-1}(E))\) and right invariance give the invariance of \(\rho:=m\circ\Psi_\varphi^{-1}\); hence \(\rho=\sigma\).
Since \((U(B))_\varphi=\{r:rU^{-1}\varphi\in B\}=B_{U^{-1}\varphi}\), we get \((A\cap U(B))_\varphi=A_\varphi\cap B_{U^{-1}\varphi}\), and (1), Fubini on \(\mathcal{U}_n\times S\) for the Borel integrand \((U,\varphi)\mapsto\nu(A_\varphi\cap B_{U^{-1}\varphi})\), and (3) with Theorem 3.6.1 give
\begin{equation*} \int_{\mathcal{U}_n}\mu\big(A\cap U(B)\big)m(dU) =\int_S\int_S\nu\big(A_\varphi\cap B_\psi\big)\sigma(d\psi)\sigma(d\varphi), \end{equation*}
which by (2) and two applications of (1) is at least \(\mu(A)\mu(B)\).
If \(U(B)=B\) for every \(U\), the left-hand side is \(\mu(A\cap B)\), \(m\) being a probability measure. Finally a probability measure \(\mu(dx)=q(|x|)dx\) is of the above form: with \(s_n:=\sigma_{n-1}(S^{n-1})\), \(\sigma:=\sigma_{n-1}/s_n\) and \(\nu\) the measure with density \(s_nr^{n-1}q( r)\), Exercise 3.10.82 gives \(\mu( C)=\int_0^{\infty}\int_SI_C(ry)\sigma(dy)\nu(dr)\), i.e. \(\mu=(\nu\otimes\sigma)\circ T^{-1}\), and \(C=\mathbb{R}^n\) shows that \(\nu\) is a probability measure.
Exercises 3.10.85–3.10.91
(Sard’s theorem) Let \(U \subset \mathbb{R}^n\) be open and let \(F : U \to \mathbb{R}^n\) be continuously differentiable. Prove that the image of the set of all points where the derivative of \(F\) is not invertible has measure zero.
Immediate from the Sard inequality: Proposition 3.7.3 (formula (3.7.7)) gives \(\lambda_n(F(A))\le\int_A|J_F(x)|dx\) for measurable \(A\subset U\), and the integrand vanishes identically on the critical set \(Z:=\{x\in U:J_F(x)=0\}\), so \(\lambda_n(F(Z))=0\). (Here \(F(Z)\) is measurable, \(F\) being locally Lipschitz and hence having Lusin’s property (N).)
Method (2): a proof not presupposing Proposition 3.7.3. Writing \(U\) as a countable union of closed cubes \(Q\subset U\), it suffices to prove \(\lambda_n(F(Z\cap Q))=0\) for one such cube, of side \(\ell\); note \(Z\cap Q\) is compact, so its image is compact. Put \(M:=1+\sup_Q\|F^{\prime}\|<\infty\) and fix \(\varepsilon\in(0,1]\). Uniform continuity of \(F^{\prime}\) on \(Q\) gives \(\delta>0\) with \(\|F^{\prime}(x)-F^{\prime}(z)\|\le\varepsilon\) for \(x,z\in Q\), \(|x-z|\le\delta\), and convexity of \(Q\) together with
\begin{equation*} F(z)-F(x)-F^{\prime}(x)(z-x) =\int_0^1\big[F^{\prime}(x+t(z-x))-F^{\prime}(x)\big](z-x)dt \end{equation*}
yields \(|F(z)-F(x)-F^{\prime}(x)(z-x)|\le\varepsilon|z-x|\). Split \(Q\) into \(N^n\) subcubes of side \(\ell/N\), with \(r:=\sqrt n\,\ell/N\le\delta\). If a subcube \(Q^{\prime}\) contains a critical point \(x\), the range of \(F^{\prime}(x)\) lies in a hyperplane \(H\), so \(F(Q^{\prime})\) lies in \(F(x)+\{h+v:h\in H,|h|\le Mr,\ v\perp H,|v|\le\varepsilon r\}\), a box with \(n-1\) edges of length \(2Mr\) and one of length \(2\varepsilon r\); by invariance of \(\lambda_n\) under orthogonal maps and translations (Corollary 3.6.4),
\begin{equation*} \lambda_n^{*}\big(F(Q^{\prime})\big)\le(2Mr)^{n-1}2\varepsilon r =2^nM^{n-1}(\sqrt n\,\ell)^n\,\varepsilon\,N^{-n}. \end{equation*}
The at most \(N^n\) such subcubes cover \(Z\cap Q\), so subadditivity gives \(\lambda_n^{*}(F(Z\cap Q))\le C\varepsilon\) with \(C:=2^nM^{n-1}(\sqrt n\,\ell)^n\) independent of \(\varepsilon\) and \(N\); let \(\varepsilon\to0\).
Let \(f\) be a continuously differentiable function on \(\mathbb{R}^n\) that vanishes outside a cube \(Q\) and let
\begin{equation*} \int_Q f(x)\, dx = 0 . \end{equation*}
Show that there exist continuously differentiable functions \(f_1, \ldots, f_n\) on \(\mathbb{R}^n\) such that \(f_i = 0\) outside \(Q\) and \(f = \sum_{i=1}^n \partial_{x_i} f_i\).
Induct on \(n\), the step splitting off the last variable by means of a fixed bump \(\zeta\).
Reduction to \(Q=[0,1]^n\): with \(T(x)=a+rx\) and \(Q=T([0,1]^n)\), the function \(\tilde f:=f\circ T\) is \(C^1\), vanishes outside \([0,1]^n\) and has integral \(r^{-n}\int_Qf=0\) (Theorem 3.7.1); given \(\tilde f=\sum_i\partial_{x_i}\tilde f_i\) as required, the functions \(f_i:=r(\tilde f_i\circ T^{-1})\) are \(C^1\), vanish outside \(Q\), and \(\partial_{y_i}f_i(y)=(\partial_{x_i}\tilde f_i)(T^{-1}y)\), so \(\sum_i\partial_{y_i}f_i=f\).
For \(n=1\), \(f_1(t):=\int_{-\infty}^tf(s)ds\) is \(C^1\) with \(f_1^{\prime}=f\), vanishes for \(t<0\), and vanishes for \(t>1\) because \(\int_0^1f=0\).
Step \(n\to n+1\). Write \(x=(y,t)\) and \(g(y):=\int_{\mathbb{R}}f(y,t)dt\). Since \(f\) and \(\partial_{y_i}f\) are continuous and vanish off the compact set \([0,1]^{n+1}\), hence are bounded, differentiation under the integral sign is legitimate (difference quotients are dominated via the mean value theorem), so \(g\in C^1(\mathbb{R}^n)\) with \(\partial_{y_i}g=\int\partial_{y_i}f\,dt\); moreover \(g\) vanishes outside \([0,1]^n\) and \(\int_{[0,1]^n}g=\int_{[0,1]^{n+1}}f=0\) by Fubini’s theorem. The inductive hypothesis gives \(g=\sum_{i\le n}\partial_{y_i}g_i\) with \(g_i\in C^1(\mathbb{R}^n)\) vanishing outside \([0,1]^n\). Fix \(\zeta\in C^\infty(\mathbb{R})\) with \(\zeta\ge0\), \(\zeta=0\) outside \([0,1]\) and \(\int\zeta=1\), and put
\begin{equation*} f_i(y,t):=g_i(y)\zeta(t)\ \ (i\le n),\qquad f_{n+1}(y,t):=\int_{-\infty}^{t}\big[f(y,s)-\zeta(s)g(y)\big]ds . \end{equation*}
These are \(C^1\) (for \(f_{n+1}\) by the same dominated-convergence argument, its integrand being continuous, bounded and supported in a compact set), with \(\partial_tf_{n+1}=f-\zeta g\) and \(\partial_{y_i}f_{n+1}(y,t)=\int_{-\infty}^{t}[\partial_{y_i}f(y,s)-\zeta(s)\partial_{y_i}g(y)]ds\). They vanish outside \([0,1]^{n+1}\): for \(t<0\) the integrand vanishes on \((-\infty,t)\); for \(t>1\) the full integral is \(g(y)-g(y)\int\zeta=0\); and for \(y\notin[0,1]^n\) both \(f(y,\cdot)\) and \(g(y)\) vanish. Finally
\begin{equation*} \sum_{i\le n}\partial_{y_i}f_i+\partial_tf_{n+1} =\zeta(t)g(y)+f(y,t)-\zeta(t)g(y)=f(y,t). \end{equation*}
Let \(U\) be a closed ball in \(\mathbb{R}^n\) and let \(F : U \to \mathbb{R}^n\) be a mapping that is infinitely differentiable in a neighborhood of \(U\). Suppose that \(y \notin F(\partial U)\), where \(\partial U\) is the boundary of \(U\). Let \(W\) be a cube containing \(y\) in its interior and not meeting \(F(\partial U)\), and let \(\varrho\) be a nonnegative smooth function vanishing outside \(W\) and having the integral \(1\). Show that the quantity defined by the following formula and called the degree of the mapping \(F\) on \(U\) at the point \(y\) is independent of our choice of a function \(\varrho\) with the stated properties:
\begin{equation*} d(F,U;y) := \int_U \varrho\bigl(F(x)\bigr) J_F(x)\, dx, \qquad J_F = \det F^{\prime} . \end{equation*}
Two admissible densities give the same integral because their difference has zero integral over a common cube, and any such function is a sum \(\sum_i\partial_{x_i}g_i\) each term of which integrates to zero against \(J_F\).
Write \(A:=F^{\prime}\) and \((\operatorname{cof}A)_{ij}:=(-1)^{i+j}\det A^{(i|j)}\), so that the Laplace expansion gives
\begin{equation*} \sum_{j=1}^nA_{kj}(\operatorname{cof}A)_{ij}=\delta_{ki}\det A \tag{1} \end{equation*}
(for \(k\ne i\) the left side expands a determinant with two equal rows).
Piola identity: if \(F\) is twice continuously differentiable, then \(\sum_j\partial_{x_j}(\operatorname{cof}A)_{ij}=0\). Indeed \(A_{kl}=\partial_{x_l}F_k\), so the product rule gives
\begin{equation*} \sum_j\partial_{x_j}(\operatorname{cof}A)_{ij} =\sum_{\sigma\in S_n}\operatorname{sgn}(\sigma) \sum_{m\ne i}\big(\partial_{x_{\sigma(i)}}\partial_{x_{\sigma(m)}}F_m\big) \prod_{k\ne i,m}\partial_{x_{\sigma(k)}}F_k , \end{equation*}
where \(j=\sigma(i)\) ranges over all values; for fixed \(m\) the involution \(\sigma\mapsto\sigma\tau\) with \(\tau=(i\,m)\) reverses the sign and fixes the summand, mixed second partials commuting.
Hence for \(u\in C^1(\mathbb{R}^n)\) the field \(V_j:=(u\circ F)(\operatorname{cof}A)_{ij}\) is \(C^1\) and satisfies, by the product rule, the chain rule \(\partial_{x_j}(u\circ F)=\sum_k(\partial_ku)(F)A_{kj}\) and (1),
\begin{equation*} \operatorname{div}V=(\partial_{x_i}u)(F)\,J_F . \tag{2} \end{equation*}
Lemma. Let \(B\subset\mathbb{R}^n\) be compact, \(F\) twice continuously differentiable on an open \(O\supset B\), \(Q\) a closed cube with \(Q\cap F(\partial B)=\emptyset\), and \(u\in C^1(\mathbb{R}^n)\) vanishing outside \(Q\). Then \(\int_B(\partial_{x_i}u)(F)J_F\,dx=0\). Indeed \(K:=\{x\in B:F(x)\in Q\}\) is compact and misses \(\partial B\), hence is a compact subset of \(\operatorname{int}B\), while \(V\) vanishes on the open set \(O^{\prime}:=\{x\in O:F(x)\notin Q\}\supset B\setminus K\). Take \(\chi\in C_c^\infty(\operatorname{int}B)\) with \(0\le\chi\le1\) and \(\chi\equiv1\) on an open \(G\supset K\) (take \(\chi\equiv0\) if \(K=\emptyset\)). Then \(\chi V\) is \(C^1\) with compact support in \(\operatorname{int}B\), so Fubini’s theorem and the fundamental theorem of calculus give \(\int_{\mathbb{R}^n}\operatorname{div}(\chi V)dx=0\), hence \(\int_B\operatorname{div}(\chi V)dx=0\). Since \(B\subset G\cup O^{\prime}\), with \(\nabla\chi=0\) on \(G\) and \(V\equiv0\) on the open set \(O^{\prime}\), we have \(\operatorname{div}(\chi V)=\operatorname{div}V\) on \(B\), and (2) finishes.
Consequently, if \(g\in C^1\) vanishes outside such a cube \(Q\) with \(\int g\,dx=0\), then Exercise 3.10.86 gives \(g=\sum_i\partial_{x_i}g_i\) with \(g_i\in C^1\) vanishing outside \(Q\), and the Lemma applied to each \(u=g_i\) yields
\begin{equation*} \int_Bg\big(F(x)\big)J_F(x)\,dx=0 . \tag{3} \end{equation*}
Now let \(\varrho_1,\varrho_2\) be admissible for \(y\notin F(\partial U)\), with cubes \(W_1,W_2\). If \(W_1=W_2=W\), then \(g:=\varrho_1-\varrho_2\) vanishes outside \(W\) with \(\int g=0\), so (3) with \(B=U\), \(Q=W\) gives equality of the two integrals. In general choose a closed cube \(W_0\) with \(y\in\operatorname{int}W_0\) and \(W_0\subset\operatorname{int}W_1\cap\operatorname{int}W_2\), and \(\varrho_0\) admissible for \(W_0\); since \(\varrho_0\) vanishes outside \(W_0\subset W_m\) it is admissible for \(W_m\) as well, so the previous case gives \(\int_U\varrho_m(F)J_F\,dx=\int_U\varrho_0(F)J_F\,dx\) for \(m=1,2\). The integral itself exists, the integrand being continuous on the compact set \(U\).
Show that if the point \(y\) in the previous exercise is such that \(F^{-1}(y) = \{x_1,\ldots,x_k\}\), where \(J_F(x_i) \ne 0\), then \(d(F,U;y) = \sum_{i=1}^k \operatorname{sign} J_F(x_i)\).
Choose a cube \(W\ni y\) so small that \(\{x\in U:F(x)\in W\}\subset\bigcup_iV_i\), where the \(V_i\) are disjoint neighbourhoods of the \(x_i\) on which \(F\) is a diffeomorphism; then each \(V_i\) contributes exactly \(\varepsilon_i:=\operatorname{sign}J_F(x_i)\).
Since \(y\notin F(\partial U)\), all the \(x_i\) lie in \(\operatorname{int}U\), and as \(J_F(x_i)\ne0\) the inverse function theorem gives connected pairwise disjoint open sets \(V_i\ni x_i\) inside \(\operatorname{int}U\) with \(F\) mapping \(V_i\) diffeomorphically onto the open set \(F(V_i)\ni y\) and \(J_F\ne0\) on \(V_i\); being continuous and nonvanishing on a connected set, \(J_F\) has there the constant sign \(\varepsilon_i\).
All sufficiently small closed cubes \(W_m\ni y\) of side \(1/m\) lie in \(\bigcap_iF(V_i)\). If none satisfied the second requirement, there would be \(z_m\in U\setminus\bigcup_iV_i\) with \(F(z_m)\in W_m\), so \(F(z_m)\to y\); the set \(U\setminus\bigcup_iV_i\) being compact, a subsequence would converge to a point \(z\) there with \(F(z)=y\), contradicting \(F^{-1}(y)\subset\bigcup_iV_i\). Fix such a cube \(W\); it misses \(F(\partial U)\), since a preimage in \(\partial U\) would lie in \(\bigcup_iV_i\subset\operatorname{int}U\). So \(W\) is admissible, and for an admissible \(\varrho\) the integrand \(\varrho(F)J_F\) vanishes off \(\bigcup_iV_i\), whence
\begin{equation*} d(F,U;y)=\sum_{i=1}^k\int_{V_i}\varrho\big(F(x)\big)J_F(x)\,dx . \end{equation*}
On \(V_i\) we have \(J_F=\varepsilon_i|J_F|\) and \(F\) injective, so the change of variables formula (Theorem 3.7.1, with the bounded Borel \(\varrho\in L^1(\mathbb{R}^n)\)) gives
\begin{equation*} \int_{V_i}\varrho\big(F(x)\big)|J_F(x)|\,dx=\int_{F(V_i)}\varrho(u)\,du =\int_{\mathbb{R}^n}\varrho(u)\,du=1 , \end{equation*}
the middle step because \(\varrho\) vanishes outside \(W\subset F(V_i)\). Hence \(d(F,U;y)=\sum_{i=1}^k\varepsilon_i\).
(i) Show that in Exercise 3.10.87 the number \(d(F,U;y)\) is an integer for all \(y \notin F(\partial U)\) and that this number is locally constant as a function of \(y\). Deduce that the degree of the mapping at \(y\) is unchanged if one replaces \(F\) with \(F_1\) with \(\|F(x) - F_1(x)\| + |J_F(x) - J_{F_1}(x)| \le \varepsilon\), where \(\varepsilon > 0\) is sufficiently small. (ii) Let \(F : U \to U\) be continuous. Prove that there exists \(x \in U\) with \(F(x) = x\).
(i) The degree is locally constant because one admissible pair \((W,\varrho)\) serves every \(y^{\prime}\in\operatorname{int}W\), the only requirements being \(y^{\prime}\in\operatorname{int}W\) and \(W\cap F(\partial U)=\emptyset\); so by Exercise 3.10.87
\begin{equation*} d(F,U;y^{\prime})=\int_U\varrho\big(F(x)\big)J_F(x)\,dx=d(F,U;y), \qquad y^{\prime}\in\operatorname{int}W , \end{equation*}
and \(d\) is constant on each component of the open set \(\mathbb{R}^n\setminus F(\partial U)\).
Integrality: by Sard’s theorem (Exercise 3.10.85) the set \(F(Z)\) is null, where \(Z:=\{x\in U:J_F(x)=0\}\), so some \(y^{\prime}\in\operatorname{int}W\setminus F(Z)\) exists. Every \(x\in U\) with \(F(x)=y^{\prime}\) then lies in \(\operatorname{int}U\) and has \(J_F(x)\ne0\), hence is isolated in the compact set \(F^{-1}(y^{\prime})\cap U\) by the inverse function theorem; a compact set of isolated points is finite. Thus Exercise 3.10.88 applies and
\begin{equation*} d(F,U;y)=d(F,U;y^{\prime})=\sum_{i=1}^k\operatorname{sign}J_F(x_i)\in\mathbb{Z}. \end{equation*}
Stability: put \(\delta:=\operatorname{dist}(y,F(\partial U))>0\), fix a closed cube \(W\) centred at \(y\) of diameter \(<\delta/2\) and an admissible \(\varrho\), and set \(L:=\sup\|\nabla\varrho\|\), \(M:=\sup_U|J_F|\). If \(\sup_U(\|F-F_1\|+|J_F-J_{F_1}|)\le\varepsilon<\delta/2\), then \(\|F_1(x)-y\|\ge\delta/2\) for \(x\in\partial U\) while \(W\) lies within \(\delta/4\) of \(y\), so \(W\cap F_1(\partial U)=\emptyset\) and \((W,\varrho)\) is admissible for \(F_1\) too. Pointwise on \(U\), by the mean value theorem for \(\varrho\),
\begin{equation*} \big|\varrho(F)J_F-\varrho(F_1)J_{F_1}\big| \le L\varepsilon M+\|\varrho\|_\infty\varepsilon , \end{equation*}
so \(|d(F,U;y)-d(F_1,U;y)|\le C\varepsilon\) with \(C:=\lambda_n(U)(LM+\|\varrho\|_\infty)\) independent of \(F_1\); both degrees being integers, \(\varepsilon<\min\{\delta/2,1/C\}\) forces them to coincide.
(ii) Brouwer’s theorem. After an affine change of variables let \(U=\{|x|\le1\}\).
(a) \(F\) infinitely differentiable near \(U\) with \(F(U)\subset U\). If \(F\) had no fixed point, put \(G_t(x):=x-tF(x)\), \(t\in[0,1]\). Then \(0\notin G_t(\partial U)\): for \(t<1\) and \(|x|=1\), \(|G_t(x)|\ge1-t>0\), and \(G_1=\mathrm{id}-F\) never vanishes. Also \(\|G_t-G_s\|\le|t-s|\) on \(U\), and \(J_{G_t}(x)=\det(I-tF^{\prime}(x))\) is a polynomial in \(t\) with coefficients continuous on the compact \(U\), so \(\sup_U|J_{G_t}-J_{G_s}|\to0\) as \(|t-s|\to0\); by the stability just proved, \(t\mapsto d(G_t,U;0)\) is locally constant on \([0,1]\), hence constant. But \(d(G_0,U;0)=1\) by Exercise 3.10.88 (\(G_0=\mathrm{id}\), \(J_{G_0}\equiv1\)), whereas \(d(G_1,U;0)=0\), since \(0\notin G_1(U)\) permits a cube \(W\ni0\) disjoint from the compact set \(G_1(U)\), making \(\varrho(G_1)\equiv0\) on \(U\). Contradiction.
(b) \(F\) merely continuous. Extend it by the radial retraction \(\widetilde F(x):=F(x/\max(1,|x|))\) and mollify: \(F_k:=\widetilde F*\varphi_k\) with \(\varphi_k(z)=k^n\varphi(kz)\), where \(\varphi\in C_c^\infty\) is nonnegative with \(\int\varphi=1\) and support in the unit ball. Each \(F_k\) is infinitely differentiable, and \(F_k(x)\in U\) because \(F_k(x)\) is a \(\varphi_k\)-average of points of the closed convex set \(U\) (a separating hyperplane would contradict the averaging). Moreover \(F_k\to\widetilde F\) uniformly on \(U\), by uniform continuity of \(\widetilde F\) on \(\{|x|\le2\}\) and
\begin{equation*} \|F_k(x)-\widetilde F(x)\| \le\int_{|z|\le1/k}\|\widetilde F(x-z)-\widetilde F(x)\|\varphi_k(z)\,dz . \end{equation*}
By (a) each \(F_k\) has a fixed point \(x_k\in U\); passing to \(x_{k_l}\to x\in U\),
\begin{equation*} \|F(x)-x\|\le\|F(x)-F(x_{k_l})\|+\|F(x_{k_l})-F_{k_l}(x_{k_l})\|+\|x_{k_l}-x\|\to0 . \end{equation*}
(Faber, Mycielski) (i) Let \(P \subset \mathbb{R}^n\) be a compact set that is a finite union of compact \(n\)-dimensional simplexes and let \(f : P \to \mathbb{R}\) be a smooth function in a neighborhood of \(P\) such that \(f\) vanishes outside \(P\). Show that
\begin{equation*} \int_P \det\Bigl(\frac{\partial^2 f}{\partial x_i \partial x_j}\Bigr)_{i,j \le n} dx = 0 . \end{equation*}
Construct an example showing that an analogous assertion for a ball \(P\) may fail. (ii) Let \(B \subset \mathbb{R}^n\) be a compact set and let \(F : B \to \mathbb{R}^n\) be a smooth mapping in a neighborhood of \(B\) such that \(F(\partial B)\) has measure zero and the connected complement. Show that
\begin{equation*} \int_B \det\bigl(F^{\prime}(x)\bigr)\, dx = 0 . \end{equation*}
Both parts follow from one proposition: if \(B\subset\mathbb{R}^n\) is compact, \(F\) is twice continuously differentiable on an open \(O\supset B\), \(N:=F(\partial B)\) is Lebesgue null and every connected component of \(\Omega:=\mathbb{R}^n\setminus N\) is unbounded, then \(\int_BJ_F\,dx=0\), where \(J_F:=\det F^{\prime}\).
Proof. Here \(N\) is compact and \(\Omega\) open. For a closed cube \(Q\) disjoint from \(N\), the Lemma of Exercise 3.10.87, stated there for an arbitrary compact \(B\), together with Exercise 3.10.86 gives \(\int_Bu(F)J_F\,dx=0\) for every \(u\in C^1\) vanishing outside \(Q\) with \(\int u\,dx=0\); splitting \(u=(\int u)\varrho_Q+(u-(\int u)\varrho_Q)\) for one admissible \(\varrho_Q\),
\begin{equation*} \int_Bu\big(F(x)\big)J_F(x)dx=\Big(\int_{\mathbb{R}^n}u\,dx\Big)d_Q,\qquad d_Q:=\int_B\varrho_Q(F)J_F\,dx , \tag{1} \end{equation*}
so \(d_Q\) does not depend on \(\varrho_Q\), and \(Q\subset Q^{\prime}\) gives \(d_Q=d_{Q^{\prime}}\) (any \(\varrho_Q\) is admissible for \(Q^{\prime}\)). Hence \(d(y):=d_Q\), for any closed cube \(Q\subset\Omega\) with \(y\in\operatorname{int}Q\), is well defined (two such cubes contain a common one about \(y\)) and locally constant on \(\Omega\), hence constant on components. Taking \(R>\sup_B\|F\|\) and \(|y|>R+1\) with a small cube \(Q\ni y\) disjoint from \(\{|u|\le R\}\supset F(B)\) gives \(d(y)=0\); every component being unbounded, \(d\equiv0\) on \(\Omega\).
So \(\int_Bu(F)J_F\,dx=0\) for every \(u\in C^1\) vanishing outside a closed cube contained in \(\Omega\), hence, via a smooth partition of unity subordinate to a finite cover of \(\operatorname{supp}\psi\) by open cubes with closures in \(\Omega\), for every \(\psi\in C_c^\infty(\Omega)\). Choosing \(\psi_m\in C_c^\infty(\Omega)\) with \(0\le\psi_m\le1\) and \(\psi_m=1\) on \(K_m:=\{|y|\le m,\ \operatorname{dist}(y,N)\ge1/m\}\), which increase to \(\Omega\), dominated convergence (\(|\psi_m(F)J_F|\le\sup_B|J_F|\)) gives
\begin{equation*} \int_{B\cap F^{-1}(\Omega)}J_F(x)\,dx=0 . \tag{2} \end{equation*}
On the rest of \(B\) the integrand vanishes where \(J_F=0\); the open set \(S:=\{x\in O:J_F(x)\ne0\}\) is, being Lindelof, a countable union of balls \(B_k\) on which \(F\) is injective (inverse function theorem), and the change of variables formula (Theorem 3.7.1) with the bounded Borel \(g=\mathbf1_N\in L^1\) gives
\begin{equation*} \int_{B_k\cap F^{-1}(N)}|J_F(x)|\,dx=\int_{F(B_k)}\mathbf1_N(y)\,dy \le\lambda_n(N)=0 , \end{equation*}
so \(\int_{B\cap F^{-1}(N)}|J_F|dx=0\); adding this to (2) over the disjoint decomposition \(B=(B\cap F^{-1}(\Omega))\cup(B\cap F^{-1}(N))\) proves the Proposition.
(ii) A connected complement is its own component and contains the complement of a large ball, \(F(\partial B)\) being compact; hence it is unbounded.
(i) Here \(F:=\nabla f\), so \(F^{\prime}\) is the Hessian. Read literally the hypothesis would give \(f\in C_c^\infty\) with support in \(P\), hence \(F(\partial P)=\{0\}\) and the conclusion for every compact \(P\), ball included, so that no counterexample could exist; the intended hypothesis is that \(f\) vanishes on \(\partial P\), and we assume it, taking \(n\ge2\) (for \(n=1\) it fails: \(P=[0,1]\), \(f(x)=x(1-x)\) gives \(\int_0^1f^{\prime\prime}=-2\)).
Let \(H_1,\dots,H_r\) be the distinct hyperplanes spanned by the \((n-1)\)-faces of the simplexes \(S_k\) whose union is \(P\), with unit normals \(\nu_j\), and put \(\Lambda:=\bigcup_j\mathbb{R}\nu_j\), a closed null set since \(n\ge2\), and \(L:=\bigcup_{i<j}(H_i\cap H_j)\), a finite union of affine subspaces of dimension at most \(n-2\).
(A) \(\partial P\subset\bigcup_jH_j\): a point of \(\partial P\) lies in some \(S_k\) but in no \(\operatorname{int}S_k\), hence on an \((n-1)\)-face.
(B) If \(x\in\partial P\setminus L\), then \(\nabla f(x)\in\Lambda\). Such an \(x\) lies on exactly one \(H_j\), so a small ball \(B_0\subset O\) about \(x\) meets no other \(H_i\), and by (A) each of the two components \(B^{\pm}\) of \(B_0\setminus H_j\) misses \(\partial P\) and hence lies in \(\operatorname{int}P\) or in \(\mathbb{R}^n\setminus P\). Not both lie in \(\operatorname{int}P\), else \(B_0\subset\overline{\operatorname{int}P}\subset P\) would be open and contradict \(x\in\partial P\); and not both lie outside \(P\), since \(S_k=\overline{\operatorname{int}S_k}\) makes \(B_0\) meet the open set \(\operatorname{int}P\) off \(H_j\). So every \(z\in B_0\cap H_j\) is a limit both of points of \(P\) and of points of its complement, i.e. \(z\in\partial P\); thus \(f\equiv0\) on the relatively open set \(B_0\cap H_j\), all derivatives of \(f\) tangent to \(H_j\) vanish at \(x\), and \(\nabla f(x)\in\mathbb{R}\nu_j\).
(C) \(\partial P\setminus L\) is dense in \(\partial P\): given \(x\in\partial P\) and \(\varepsilon>0\), the ball \(B(x,\varepsilon)\) contains a point \(q\notin P\) and a nonempty open piece of \(\operatorname{int}P\); the points \(p\) for which \([p,q]\) meets \(L\) lie in finitely many cones \(q+\mathbb{R}(M-q)\) of dimension at most \(n-1\), a null set, so some \(p\in B(x,\varepsilon)\cap\operatorname{int}P\) has \([p,q]\cap L=\emptyset\); that segment meets \(\partial P\), else \(\operatorname{int}P\) and \(\mathbb{R}^n\setminus P\) would disconnect it, and any such point is as required.
By (B), (C), continuity of \(\nabla f\) and closedness of \(\Lambda\) we get \(F(\partial P)\subset\Lambda\), a null set. Each component \(C\) of \(\mathbb{R}^n\setminus F(\partial P)\) is open, hence of positive measure, so it contains some \(z\notin\Lambda\); the component \(D\ni z\) of \(\mathbb{R}^n\setminus\Lambda\) is a connected subset of \(\mathbb{R}^n\setminus F(\partial P)\) meeting \(C\), so \(D\subset C\), and \(D\) contains the ray \(\{tz:t>0\}\), \(\Lambda\) being a union of lines through the origin. So \(C\) is unbounded, and the Proposition gives \(\int_P\det(\partial^2f/\partial x_i\partial x_j)dx=0\).
For the ball, \(P=\{|x|\le1\}\) and \(f(x)=|x|^2-1\) is smooth and vanishes on \(\partial P\), with Hessian \(2I\), so
\begin{equation*} \int_P\det\Big(\frac{\partial^2f}{\partial x_i\partial x_j}\Big)dx =2^n\lambda_n(P)>0 . \end{equation*}
Prove Proposition 3.10.16.
Both assertions come from splitting an arbitrary disjoint family, first along \(A\) and \(B\), then along the sign of \(m\). Proposition 3.10.16 says that for an additive \(m\colon\mathcal{R}\to(-\infty,+\infty]\) on a ring \(\mathcal{R}\) the variation
\begin{equation*} v(m)(A)=\sup\Big\{\sum_{j\le k}|m(A_j)|: \ A_j\in\mathcal{R}\ \text{disjoint},\ A_j\subset A\Big\} \end{equation*}
is additive and \(m^+=(v(m)+m)/2\), where \(m^+(E)=\sup_{F\in\mathcal{R},F\subset E}m(F)\). Note \(m(\emptyset)=0\), no expression \(\infty-\infty\) occurs, and if \(m(A)<\infty\) then \(m\) is real-valued on every \(F\in\mathcal{R}\) with \(F\subset A\), since \(m(A)=m(F)+m(A\setminus F)\) with both summands \(>-\infty\).
Additivity. For disjoint \(A,B\in\mathcal{R}\), concatenating a family inside \(A\) with one inside \(B\) gives \(v(m)(A)+v(m)(B)\le v(m)(A\cup B)\). Conversely, for disjoint \(C_i\in\mathcal{R}\) with \(C_i\subset A\cup B\) put \(C_i^{\prime}:=C_i\cap A\in\mathcal{R}\) and \(C_i^{\prime\prime}:=C_i\setminus C_i^{\prime}\in\mathcal{R}\); then \(m(C_i)=m(C_i^{\prime})+m(C_i^{\prime\prime})\) gives \(|m(C_i)|\le|m(C_i^{\prime})|+|m(C_i^{\prime\prime})|\), and the two subfamilies lie in \(A\), resp. \(B\), so \(\sum_i|m(C_i)|\le v(m)(A)+v(m)(B)\). With \(v(m)(\emptyset)=0\), induction gives additivity for finitely many disjoint sets.
The identity. Fix \(A\in\mathcal{R}\). If \(m(A)=+\infty\), then \(F=A\) gives \(m^+(A)=+\infty\) and the family \(\{A\}\) gives \(v(m)(A)=+\infty\), so both sides are \(+\infty\). If \(m(A)\in\mathbb{R}\), then \(m\) is real-valued below \(A\) and we show \(v(m)(A)=2m^+(A)-m(A)\). For \(F\in\mathcal{R}\), \(F\subset A\), the family \(\{F,A\setminus F\}\) gives
\begin{equation*} v(m)(A)\ge|m(F)|+|m(A\setminus F)|\ge m(F)-\big(m(A)-m(F)\big)=2m(F)-m(A); \end{equation*}
take the supremum over \(F\). Conversely, adjoining \(A_0:=A\setminus\bigcup_{j\le k}A_j\) to a disjoint family only increases \(\sum_j|m(A_j)|\), so it suffices to bound the sums coming from finite partitions of \(A\); for such a partition, with \(P:=\bigcup_{m(A_j)\ge0}A_j\in\mathcal{R}\) and \(N:=A\setminus P\), additivity of the finite values gives
\begin{equation*} \sum_{j}|m(A_j)|=m(P)-m(N)=2m(P)-m(A)\le2m^+(A)-m(A). \end{equation*}
Exercises 3.10.92–3.10.98
Prove that if a function \(\psi\) is positive definite, then
\begin{equation*} |\psi(y) - \psi(z)|^2 \le 2\psi(0)\bigl[\psi(0) - \operatorname{Re}\psi(y-z)\bigr]. \end{equation*}
Test Definition 3.8.3 with \(k=3\), \(y_1=0\), \(y_2=y\), \(y_3=z\), \(c_1=c\), \(c_2=1\), \(c_3=-1\), and optimize in \(c\).
First \(\psi(0)\ge0\) (take \(k=1\), \(c_1=1\)), and \(\psi(-u)=\overline{\psi(u)}\) with \(|\psi(u)|\le\psi(0)\) as in Proposition 3.10.18(ii): with \(k=2\), \(y_1=0\), \(y_2=u\),
\begin{equation*} Q(c_1,c_2)=\big(|c_1|^2+|c_2|^2\big)\psi(0)+c_1\overline{c_2}\psi(-u) +c_2\overline{c_1}\psi(u)\ \ge\ 0 , \end{equation*}
where \(c_1=c_2=1\) forces \(\operatorname{Im}\psi(u)+\operatorname{Im}\psi(-u)=0\) and \(c_1=1\), \(c_2=i\) forces \(\operatorname{Re}\psi(u)-\operatorname{Re}\psi(-u)=0\), the displayed quantity being real; then for unimodular \(c_1,c_2\) the form equals \(2\psi(0)+2\operatorname{Re}[c_2\overline{c_1}\psi(u)]\), and the choice \(c_2\overline{c_1}=-\overline{\psi(u)}/|\psi(u)|\) gives \(|\psi(u)|\le\psi(0)\).
If \(\psi(0)=0\), then \(\psi\equiv0\) and both sides vanish; so let \(\psi(0)>0\). With the choice above, the nine terms of \(Q\) group into the diagonal \((|c|^2+2)\psi(0)\), the pairs \((1,2),(2,1)\) giving \(2\operatorname{Re}[\overline c\psi(y)]\), the pairs \((1,3),(3,1)\) giving \(-2\operatorname{Re}[\overline c\psi(z)]\), and the pairs \((2,3),(3,2)\) giving \(-2\operatorname{Re}\psi(y-z)\), all by \(\psi(-u)=\overline{\psi(u)}\); so with \(w:=\psi(y)-\psi(z)\),
\begin{equation*} \big(|c|^2+2\big)\psi(0)+2\operatorname{Re}\big[\overline c\,w\big] -2\operatorname{Re}\psi(y-z)\ \ge\ 0,\qquad c\in\mathbb{C}. \end{equation*}
For \(w=0\) there is nothing to prove, since \(\operatorname{Re}\psi(y-z)\le|\psi(y-z)|\le\psi(0)\). For \(w\ne0\) take \(c=-tw/|w|\) with \(t\ge0\), so that \(|c|=t\) and \(\overline cw=-t|w|\):
\begin{equation*} \psi(0)t^2-2|w|t+2\big[\psi(0)-\operatorname{Re}\psi(y-z)\big]\ \ge\ 0 \qquad\text{for all }t\ge0 , \end{equation*}
and the admissible minimizer \(t=|w|/\psi(0)\) gives \(|w|^2\le2\psi(0)[\psi(0)-\operatorname{Re}\psi(y-z)]\).
Prove that if a function \(\psi\) on \(\mathbb{R}^n\) is positive definite and continuous at the origin, then it is continuous everywhere.
Exercise 3.10.92 gives uniform continuity at once. If \(\psi(0)=0\), then \(|\psi(u)|\le\psi(0)=0\) by Proposition 3.10.18(ii) and \(\psi\equiv0\); so let \(\psi(0)>0\). Given \(\varepsilon>0\), continuity at the origin provides \(\delta>0\) with \(|\psi(u)-\psi(0)|<\varepsilon^2/(2\psi(0))\) for \(|u|<\delta\), whence, \(\psi(0)\) being real (Proposition 3.10.18(i)),
\begin{equation*} \psi(0)-\operatorname{Re}\psi(u)=\operatorname{Re}\big[\psi(0)-\psi(u)\big] \le|\psi(0)-\psi(u)|<\frac{\varepsilon^2}{2\psi(0)}\qquad(|u|<\delta). \end{equation*}
Exercise 3.10.92 then gives, for \(|y-z|<\delta\),
\begin{equation*} |\psi(y)-\psi(z)|^2\le2\psi(0)\cdot\frac{\varepsilon^2}{2\psi(0)}=\varepsilon^2 . \end{equation*}
Prove that a complex function \(\varphi\) equals the characteristic functional of a nonnegative absolutely continuous measure precisely when there exists a complex function \(\psi\in L^2(\mathbb{R}^n)\) such that
\begin{equation*} \varphi(x)=\int_{\mathbb{R}^n}\psi(x+y)\overline{\psi(y)}\,dy . \end{equation*}
The measure is the one with density \(|\widehat\psi|^2\); everything rests on
\begin{equation*} \int_{\mathbb{R}^n}\psi(x+y)\overline{\psi(y)}\,dy =\int_{\mathbb{R}^n}e^{i(u,x)}\big|\widehat\psi(u)\big|^2du, \qquad\psi\in L^2(\mathbb{R}^n). \tag{1} \end{equation*}
Indeed, with \(\psi_x(y):=\psi(x+y)\) the change of variable \(y\mapsto y-x\) gives \(\widehat{\psi_x}(u)=e^{i(u,x)}\widehat\psi(u)\) for \(\psi\in L^1\cap L^2\), hence for every \(\psi\in L^2\), translation, multiplication by a unimodular function and the Fourier transform all being isometries of \(L^2\) (Exercise 3.10.76). The left side of (1) converges absolutely for every \(x\) by the Cauchy-Bunyakowsky inequality, and the Parseval identity of Exercise 3.10.76(ii) applied to \(\psi_x,\psi\) turns it into the right side.
Sufficiency. Given \(\psi\in L^2\), put \(f:=|\widehat\psi|^2\ge0\); then \(\|f\|_{L^1}=\|\widehat\psi\|_{L^2}^2=\|\psi\|_{L^2}^2<\infty\), so the measure \(\mu:=f\,dx\) is nonnegative, bounded and absolutely continuous, and (1) says precisely that \(\varphi=\widetilde\mu\).
Necessity. Let \(\varphi=\widetilde\mu\) with \(\mu\ge0\) bounded and absolutely continuous, take a density \(f\ge0\) of \(\mu\) and put \(h:=\sqrt f\). Then \(h\ge0\) is measurable with \(\int h^2dx=\mu(\mathbb{R}^n)<\infty\), so \(h\in L^2(\mathbb{R}^n)\), and by the surjectivity in Exercise 3.10.76(ii) there is \(\psi\in L^2\) with \(\widehat\psi=h\). Then \(|\widehat\psi|^2=f\) a.e., and (1) gives \(\int\psi(x+y)\overline{\psi(y)}dy=\widetilde\mu(x)=\varphi(x)\) for every \(x\).
Let \(\mu\) be a probability measure on the real line with the characteristic functional \(\widetilde{\mu}\) and let \(F_{\mu}(t):=\mu\bigl((-\infty,t)\bigr)\).
(i) Prove that, for every \(t\), the limit
\begin{equation*} \lim_{T\to\infty}\frac{1}{2T}\int_{-T}^{T}\exp(-its)\,\widetilde{\mu}(s)\,ds \end{equation*}
exists and equals the jump of the function \(F_{\mu}\) at the point \(t\).
(ii) Let \(\{t_j\}\) be all points of discontinuity of \(F_{\mu}\) and let \(d_j\) be the size of the jump at \(t_j\). Prove the equality
\begin{equation*} \lim_{T\to\infty}\frac{1}{2T}\int_{-T}^{T}\bigl|\widetilde{\mu}(s)\bigr|^{2}\,ds=\sum_{j=1}^{\infty}d_j^{2}. \end{equation*}
Deduce that a necessary and sufficient condition for the continuity of \(F_{\mu}\) is that the limit on the left be zero.
(i) The limit is \(\mu(\{t\})\), which is the jump; (ii) the limit is \((\mu\otimes\mu)(\Delta)=\sum_jd_j^2\). Throughout \(\widetilde\mu(s)=\int e^{isx}\mu(dx)\) is bounded continuous with \(|\widetilde\mu|\le1\) (Proposition 3.8.4(ii)). Since \(\int_{-T}^Te^{isa}ds=2\sin(Ta)/a\),
\begin{equation*} K_T(a):=\frac{1}{2T}\int_{-T}^{T}e^{isa}ds=\frac{\sin(Ta)}{Ta}\ \ (a\ne0), \qquad K_T(0)=1 , \end{equation*}
with \(|K_T|\le1\) (as \(|\sin u|\le|u|\)) and \(K_T(a)\to I_{\{0\}}(a)\) as \(T\to\infty\). Moreover \(F_\mu\) is nondecreasing and left continuous, with \(F_\mu(t+)=\mu((-\infty,t])\) by countable additivity, so its jump at \(t\) is \(\mu(\{t\})\).
(i) Fubini’s theorem for the finite measure \(ds\otimes\mu\) on \([-T,T]\times\mathbb{R}\) (the integrand has modulus \(1\)) and then dominated convergence with the \(\mu\)-integrable majorant \(1\) give
\begin{equation*} \frac{1}{2T}\int_{-T}^{T}e^{-its}\widetilde\mu(s)\,ds =\int_{\mathbb{R}}K_T(x-t)\,\mu(dx)\longrightarrow\mu(\{t\}). \end{equation*}
(ii) Likewise \(|\widetilde\mu(s)|^2=\int\int e^{is(x-y)}\mu(dx)\mu(dy)\), and Fubini on \([-T,T]\times\mathbb{R}^2\) followed by dominated convergence give
\begin{equation*} \frac{1}{2T}\int_{-T}^{T}\big|\widetilde\mu(s)\big|^2ds =\int\int K_T(x-y)\,\mu(dx)\mu(dy)\longrightarrow(\mu\otimes\mu)(\Delta), \end{equation*}
\(\Delta\) being the closed diagonal. By Fubini, \((\mu\otimes\mu)(\Delta)=\int\mu(\{y\})\mu(dy)\). The atoms of \(\mu\) are at most countably many (at most \(m\) of mass \(>1/m\)) and are exactly the discontinuity points \(t_j\), with \(\mu(\{t_j\})=d_j\), so \(\phi(y):=\mu(\{y\})=\sum_jd_jI_{\{t_j\}}\) is Borel and vanishes off \(\{t_j\}\), whence by countable additivity of the integral
\begin{equation*} \int_{\mathbb{R}}\phi\,d\mu=\sum_j\phi(t_j)\mu(\{t_j\})=\sum_jd_j^2 . \end{equation*}
Finally \(F_\mu\) is continuous exactly when \(\mu\) has no atoms, i.e. when all \(d_j\) vanish, i.e. when this limit is \(0\).
(\(\circ\)) Let \(f\) be a Lebesgue integrable function on \(\mathbb{R}^n\) such that, for every orthogonal linear operator \(U\) on \(\mathbb{R}^n\), the functions \(f\) and \(f\circ U\) coincide almost everywhere. Prove that there exists a function \(g\) on \([0,\infty)\) such that \(f(x)=g(|x|)\) for almost all \(x\).
Mollify with a radial kernel and pass to an almost everywhere convergent subsequence. Put \(\psi:=c^{-1}\psi_0\) with \(\psi_0(t):=\max(0,1-|t|)\) and \(c:=\int_{\mathbb{R}^n}\psi_0(|y|)dy\in(0,\infty)\), so that \(\varrho(y):=\psi(|y|)\) is continuous with bounded support and \(\int\varrho\,dy=1\); set \(\varrho_\varepsilon(y):=\varepsilon^{-n}\varrho(y/\varepsilon)\), which has integral \(1\) and satisfies \(\varrho_\varepsilon(Uv)=\varrho_\varepsilon(v)\) for every orthogonal \(U\).
Since \(f\in L^1\) and \(\varrho_\varepsilon\in L^{\infty}\), Corollary 3.9.6 (with \(p=1\), \(q=\infty\)) makes \(u_\varepsilon:=f*\varrho_\varepsilon\) defined at every point, bounded and continuous. It is spherically invariant: the substitution \(y=Uz\) preserves Lebesgue measure (Corollary 3.6.4), so by radiality of \(\varrho_\varepsilon\) and \(f\circ U=f\) a.e.,
\begin{equation*} u_\varepsilon(Ux)=\int f(Uz)\varrho_\varepsilon(x-z)\,dz =\int f(z)\varrho_\varepsilon(x-z)\,dz =u_\varepsilon(x). \end{equation*}
As the orthogonal group is transitive on each sphere, \(u_\varepsilon(x)=g_\varepsilon(|x|)\) with \(g_\varepsilon( r):=u_\varepsilon(re_1)\) continuous on \([0,\infty)\).
By Theorem 4.2.4 with \(p=1\), \(\|f*\varrho_\varepsilon-f\|_{L^1}\to0\), so \(u_{1/j}\to f\) in measure on every ball \(B_m\) (Chebyshev); by the Riesz theorem (Theorem 2.2.5(i)) applied on \(B_1\), then on \(B_2\), and a diagonal choice, there are \(\varepsilon_k\to0\) with \(u_{\varepsilon_k}\to f\) outside a null set \(N\). Put
\begin{equation*} E:=\bigcap_{m}\bigcup_{K}\bigcap_{k,l\ge K} \big\{r\ge0:\ |g_{\varepsilon_k}( r)-g_{\varepsilon_l}( r)|\le1/m\big\}, \end{equation*}
the set where \((g_{\varepsilon_k}( r))_k\) is Cauchy; it is Borel, the innermost sets being closed since the \(g_{\varepsilon_k}\) are continuous. Let \(g:=\lim_kg_{\varepsilon_k}\) on \(E\) and \(g:=0\) off \(E\); then \(g\) is Borel, being the pointwise limit of the \(g_{\varepsilon_k}I_E\). For \(x\notin N\) we have \(g_{\varepsilon_k}(|x|)=u_{\varepsilon_k}(x)\to f(x)\), so \(|x|\in E\) and \(g(|x|)=f(x)\).
Prove that a bounded Borel measure on \(\mathbb{R}^n\) is spherically invariant precisely when its characteristic functional is a function of \(|x|\).
Everything follows from the identity, valid for orthogonal \(U\) and \(\mu_U:=\mu\circ U^{-1}\),
\begin{equation*} \widetilde{\mu_U}(x)=\int e^{i(x,Uy)}\,\mu(dy)=\int e^{i(U^{*}x,y)}\,\mu(dy) =\widetilde{\mu}\bigl(U^{-1}x\bigr), \tag{1} \end{equation*}
by the change of variables formula for image measures (Theorem 3.6.1, applied to the Jordan components of the bounded measure \(\mu\), the integrand being bounded) together with \(U^{*}=U^{-1}\).
(i) If \(\mu\circ U^{-1}=\mu\) for all orthogonal \(U\), then (1) gives \(\widetilde{\mu}(U^{-1}x)=\widetilde{\mu}(x)\) for all such \(U\), so \(\widetilde{\mu}\) is constant on each sphere \(\{|x|=r\}\), the orthogonal group acting transitively on spheres. Hence \(\widetilde{\mu}(x)=g(|x|)\) with \(g( r):=\widetilde{\mu}(re_1)\).
(ii) If \(\widetilde{\mu}(x)=g(|x|)\), then by (1) and \(|U^{-1}x|=|x|\),
\begin{equation*} \widetilde{\mu_U}(x)=g\bigl(|U^{-1}x|\bigr)=g(|x|)=\widetilde{\mu}(x), \end{equation*}
so the bounded Borel measures \(\mu_U\) and \(\mu\) have equal Fourier transforms and therefore coincide by Proposition 3.8.6.
Let \(A\) and \(B\) be two sets of positive measure in \(\mathbb{R}^n\) and let \(C\) be a set in \(\mathbb{R}^{2n}\) that coincides with the set \(A\times B\) up to a measure zero set. Show that the set \(D:=\{x+y:\ x,y\in\mathbb{R}^n,\ (x,y)\in C\}\) coincides up to a measure zero set with a set that contains an open ball.
The required set is \(D^{\prime}:=D\cup U\), where \(U:=\{x:\ I_A*I_B(x)>0\}\); we show \(U\setminus Z\subset D\) for some null \(Z\), whence \(D^{\prime}\triangle D\subset Z\) is null (Lebesgue measure being complete) and \(D^{\prime}\supset U\) contains an open ball. Replacing \(A,B\) by their intersections \(A_0,B_0\) with large balls and \(C\) by \(C\cap(A_0\times B_0)\) only shrinks \(D\) and preserves the hypothesis, the new symmetric difference lying in \((A\times B)\setminus C\); so assume \(\lambda_n(A),\lambda_n(B)\in(0,\infty)\).
Since \(I_A\in L^1\) and \(I_B\in L^{\infty}\), Corollary 3.9.6 makes \(h:=I_A*I_B\) everywhere defined, bounded and continuous, while Fubini’s theorem gives \(\int h\,dx=\lambda_n(A)\lambda_n(B)>0\); so \(U=\{h>0\}\) is open and nonempty.
Now \(N:=C\triangle(A\times B)\) is null by hypothesis, and the shear \(S(x,y)=(x-y,y)\) has \(|\det S|=1\), so by Corollary 3.6.4 the set \(N^{\prime}:=S^{-1}(N)\) is null in \(\mathbb{R}^{2n}\) and \(S^{-1}( C)\) is measurable. By Fubini’s theorem there is a null set \(Z\subset\mathbb{R}^n\) off which the section \(N^{\prime}_x\) is null and \(y\mapsto I_C(x-y,y)=I_{S^{-1}( C)}(x,y)\) is measurable; for \(x\notin Z\),
\begin{equation*} \int_{\mathbb{R}^n}I_C(x-y,y)\,dy=\int_{\mathbb{R}^n}I_A(x-y)I_B(y)\,dy=h(x). \end{equation*}
Hence for \(x\in U\setminus Z\) this integral is positive, so \((x-y,y)\in C\) for some \(y\), i.e. \(x=(x-y)+y\in D\).
Exercises 3.10.99–3.10.105
Prove Proposition 3.9.9.
Everything reduces to the majorant
\begin{equation*} g(x) := \int_{\mathbb{R}^n} |f(x-y)|\,\nu(dy)\in[0,+\infty], \qquad \nu:=|\mu|,\quad M:=\|\mu\|=\nu(\mathbb{R}^n), \end{equation*}
which is Borel by Tonelli’s theorem (Theorem 3.4.5; \((x,y)\mapsto f(x-y)\) is Borel and \(\lambda,\nu\) are \(\sigma\)-finite): where \(g(x)<\infty\) the integral \(f*\mu(x)\) converges absolutely with \(|f*\mu(x)|\le g(x)\), so it suffices to prove \(\|g\|_p\le M\|f\|_p\). Translation invariance gives \(\int|f(x-y)|^p\,dx=\|f\|_p^p\) for each \(y\).
(i) \(p=1\). Tonelli:
\begin{equation*} \int g\,dx=\int\Bigl(\int|f(x-y)|\,dx\Bigr)\nu(dy)=M\|f\|_1<\infty . \end{equation*}
(ii) \(1<p<\infty\), \(p^{\prime}=p/(p-1)\). Hoelder against the finite measure \(\nu\) gives \(g(x)^p\le M^{p/p^{\prime}}\int|f(x-y)|^p\,\nu(dy)\), so by Tonelli
\begin{equation*} \int g^p\,dx\le M^{p/p^{\prime}}\int\Bigl(\int|f(x-y)|^p\,dx\Bigr)\nu(dy)=M^{p}\|f\|_p^p , \end{equation*}
since \(p/p^{\prime}+1=p\).
(iii) \(p=\infty\). With \(N:=\{|f|>\|f\|_{\infty}\}\), a Borel null set, Tonelli applied to \(I_N(x-y)\) yields \(\int\nu(\{y: x-y\in N\})\,dx=\int\lambda(N+y)\,\nu(dy)=0\), so for a.e. \(x\) one has \(|f(x-y)|\le\|f\|_{\infty}\) for \(\nu\)-a.e. \(y\), whence \(g(x)\le M\|f\|_{\infty}\).
In each case \(g<\infty\) a.e., so \(f*\mu\) is defined a.e.; on each ball \(B_R\) the bound \(\int_{B_R}g\,d\lambda<\infty\) (Hoelder for \(p>1\)) makes \(f(x-y)\) integrable for \(\lambda|_{B_R}\otimes\mu^{\pm}\), so Fubini’s theorem (Theorem 3.4.4) gives measurability of \(f*\mu\) on \(B_R\), hence on \(\mathbb{R}^n\). Therefore \(\|f*\mu\|_p\le\|g\|_p\le\|f\|_p\|\mu\|\).
Let \(f\in L^1(\mathbb{R}^1)\). Prove the equalities
\begin{equation*} \left|\int_{-\infty}^{+\infty} f(x)\,dx\right| = \lim_{T\to+\infty}\int_{-\infty}^{+\infty}\left|(2T)^{-1}\int_{-T}^{T} f(x+t)\,dt\right|dx, \end{equation*}
\begin{equation*} \int_0^1\left|\sum_{n=-\infty}^{\infty} f(x+n)\right|dx = \lim_{N\to\infty}\int_{-\infty}^{+\infty}\left|(2N+1)^{-1}\sum_{n=-N}^{N} f(x+n)\right|dx . \end{equation*}
Both equalities say that a \(1\)-Lipschitz functional on \(L^1\) converges to another, so in each case it suffices to check the limit on the dense set of \(f\) vanishing off \([-k,k]\).
(i) With \(\psi_T:=(2T)^{-1}I_{[-T,T]}\) (even, \(\|\psi_T\|_1=1\)) the inner average is \(f*\psi_T\), so the claim is \(\Lambda_T(f):=\|f*\psi_T\|_1\to\Lambda(f):=|\int f\,dx|\). Young’s inequality (Theorem 3.9.4, \(p=q=r=1\)) gives \(\|h*\psi_T\|_1\le\|h\|_1\), so \(|\Lambda_T(f)-\Lambda_T(h)|\le\|f-h\|_1\), and \(\Lambda\) is \(1\)-Lipschitz too. Let now \(g\) vanish off \([-k,k]\), \(S:=\int g\), and \(T>k\). From \(g*\psi_T(x)=(2T)^{-1}\int_{x-T}^{x+T}g\), we get \(g*\psi_T(x)=S/(2T)\) for \(|x|\le T-k\), \(g*\psi_T(x)=0\) for \(|x|>T+k\), and \(|g*\psi_T(x)|\le(2T)^{-1}\|g\|_1\) on the remaining set of measure \(4k\). Hence
\begin{equation*} \Bigl|\Lambda_T(g)-\tfrac{T-k}{T}|S|\Bigr|\le\frac{2k}{T}\|g\|_1\xrightarrow[T\to\infty]{}0 , \end{equation*}
so \(\Lambda_T(g)\to|S|=\Lambda(g)\). Since \(\|f-fI_{[-k,k]}\|_1\to0\) as \(k\to\infty\), the Lipschitz bounds give \(\Lambda_T(f)\to\Lambda(f)\).
(ii) Put \(A_Nf(x):=(2N+1)^{-1}\sum_{|n|\le N}f(x+n)\) and \(Pf(y):=\sum_{n\in\mathbb{Z}}f(y+n)\). Tonelli’s theorem (Theorem 3.4.5) gives \(\int_0^1\sum_n|f(y+n)|\,dy=\|f\|_1<\infty\), so \(Pf\) converges absolutely a.e. and \(\Theta(f):=\int_0^1|Pf|\,dy\le\|f\|_1\); translation invariance gives \(\|A_Nf\|_1\le\|f\|_1\), so \(\Theta_N(f):=\|A_Nf\|_1\) and \(\Theta\) are again \(1\)-Lipschitz. Let \(f\) vanish off \([-k,k]\) with \(k\in\mathbb{N}\), let \(N>k+1\), and write \(x=m+y\), \(m\in\mathbb{Z}\), \(y\in[0,1)\); then \(\sum_{|n|\le N}f(x+n)=\sum_{j=m-N}^{m+N}f(y+j)\), whose terms vanish unless \(-k-1\le j\le k\). So \(A_Nf(m+y)=(2N+1)^{-1}Pf(y)\) when \(|m|\le N-k-1\), and \(A_Nf(m+y)=0\) when \(m>N+k\) or \(m<-N-k-1\); for every \(m\), \(\int_m^{m+1}|A_Nf|\,dx\le(2N+1)^{-1}\|f\|_1\). At most \(4k+3\) integers \(m\) fall in neither class, so
\begin{equation*} \Bigl|\Theta_N(f)-\tfrac{2N-2k-1}{2N+1}\Theta(f)\Bigr| \le\frac{(4k+3)\|f\|_1}{2N+1}\xrightarrow[N\to\infty]{}0 , \end{equation*}
and the density argument of (i) finishes the proof.
Let \((X,\mathcal{A},\mu)\) be a probability space and let \(\nu\) be a bounded nonnegative measure on \(\mathcal{A}\). Prove that, for every \(\varepsilon>0\), the family \(\mathcal{A}_\varepsilon:=\{A\in\mathcal{A}:\ \mu(A)\le\varepsilon\}\) contains a set \(A_\varepsilon\) such that \(\nu(A_\varepsilon)\) is maximal in the following sense: if \(B\in\mathcal{A}_\varepsilon\) and \(\mu(B)\le\mu(A_\varepsilon)\), then \(\nu(A_\varepsilon)\ge\nu(B)\).
Take \(A_\varepsilon:=A_k=\{h>k\,g\}\) for \(k\in\mathbb{N}\) large, where \(g,h\ge0\) are the Radon-Nikodym densities of \(\mu,\nu\) with respect to the finite measure \(\sigma:=\mu+\nu\) (Theorem 3.2.2; \(\mu\ll\sigma\) and \(\nu\ll\sigma\) trivially), so that
\begin{equation*} \mu(A)=\int_A g\,d\sigma,\qquad \nu(A)=\int_A h\,d\sigma\qquad (A\in\mathcal{A}). \end{equation*}
(i) Every \(A_c:=\{h>cg\}\), \(c\ge0\), is maximal in the stated sense. Let \(\mu(B)\le\mu(A_c)\); splitting both sets along \(A_c\cap B\),
\begin{equation*} \nu(A_c)-\nu(B) = \int_{A_c\setminus B}h\,d\sigma - \int_{B\setminus A_c}h\,d\sigma . \end{equation*}
On \(A_c\) we have \(h\ge c g\), and on \(X\setminus A_c\) we have \(h\le c g\). Hence
\begin{equation*} \int_{A_c\setminus B}h\,d\sigma \ge c\int_{A_c\setminus B}g\,d\sigma = c\,\mu(A_c\setminus B),\qquad \int_{B\setminus A_c}h\,d\sigma \le c\int_{B\setminus A_c}g\,d\sigma = c\,\mu(B\setminus A_c), \end{equation*}
and therefore
\begin{equation*} \nu(A_c)-\nu(B) \ge c\big(\mu(A_c\setminus B) - \mu(B\setminus A_c)\big) = c\big(\mu(A_c)-\mu(B)\big)\ge 0 , \end{equation*}
using \(\mu(A_c\setminus B)=\mu(A_c)-\mu(A_c\cap B)\), \(\mu(B\setminus A_c)=\mu(B)-\mu(A_c\cap B)\) (all finite) and \(c\ge0\).
(ii) Some \(A_k\), \(k\in\mathbb{N}\), lies in \(\mathcal{A}_\varepsilon\). The \(A_k\) decrease with
\begin{equation*} \bigcap_{k\ge1}A_k=\{g=0,\ h>0\}=:A_\infty,\qquad \mu(A_\infty)=\int_{A_\infty}g\,d\sigma=0 \end{equation*}
(if \(g(x)>0\) then \(h(x)>kg(x)\) fails for large \(k\)), so \(\mu(A_k)\to0\) by continuity from above of the finite measure \(\mu\); fix \(k\) with \(\mu(A_k)\le\varepsilon\).
Let \((X,\mu)\) be a space with a nonnegative measure \(\mu\) and let \(f\) be a \(\mu\)-measurable function. The nonincreasing rearrangement of the function \(f\) is the function \(f^*\) on \([0,+\infty)\) with values in \([0,+\infty]\) defined by the equality
\begin{equation*} f^*(t) = \inf\big\{s\ge 0:\ \mu\big(x:\ |f(x)|>s\big)\le t\big\},\quad\text{where }\inf\varnothing = +\infty . \end{equation*}
(i) Show that if \(f\) assumes finitely many values \(0<c_1<\cdots<c_n\) on measurable sets \(A_0,A_1,\ldots,A_n\) and \(0<\mu(A_i)<\infty\) if \(1\le i\le n\), then
\begin{equation*} f^*(t) = \sum_{j=1}^{n}c_j I_{[\mu(B_{n-j}),\,\mu(B_{n+1-j}))}(t) = \sum_{j=1}^{n}b_j I_{[0,\,\mu(B_j))}(t), \end{equation*}
where \(B_j = A_{n+1-j}\cup\cdots\cup A_n\), \(B_0=\varnothing\), \(b_j = c_{n+1-j}-c_{n-j}\), \(c_0=0\).
(ii) Show that \(f^*(t) = \sup\big\{s\ge 0:\ \mu\big(x:\ |f(x)|>s\big)>t\big\}\).
(iii) Show that if measurable functions \(f_n\) monotonically increase to \(|f|\), then the functions \(f_n^*\) monotonically increase to \(f^*\).
(iv) Show that the functions \(f\) and \(f^*\) are equimeasurable, i.e., one has
\begin{equation*} \mu\big(x:\ |f(x)|>s\big) = \lambda\big(t:\ f^*(t)>s\big), \end{equation*}
where \(\lambda\) is Lebesgue measure.
(v) Prove the following Hardy and Littlewood inequality:
\begin{equation*} \int_X |fg|\,d\mu \le \int_0^{\infty} f^*(t)g^*(t)\,dt, \end{equation*}
where \(f\) and \(g\) are measurable functions.
Write \(m_f(s):=\mu(|f|>s)\), so \(f^*(t)=\inf\{s\ge0:\ m_f(s)\le t\}\); \(m_f\) is nonincreasing and right-continuous (if \(s_j\downarrow s\) then \(\{|f|>s_j\}\uparrow\{|f|>s\}\)), and \(f^*\) is nonincreasing.
(i) Since \(\mu(A_i)>0\), we have \(\mu(B_0)=0<\mu(B_1)<\cdots<\mu(B_n)<\infty\), and
\begin{equation*} m_f(s)=\mu(B_{n+1-j})\ \ (c_{j-1}\le s<c_j,\ 1\le j\le n),\qquad m_f(s)=0\ \ (s\ge c_n), \end{equation*}
because \(\{|f|>s\}=A_j\cup\cdots\cup A_n=B_{n+1-j}\) on the \(j\)-th range. Fix \(1\le k\le n\) and \(\mu(B_{k-1})\le t<\mu(B_k)\): for \(s\in[c_{j-1},c_j)\) the condition \(m_f(s)\le t\) reads \(\mu(B_{n+1-j})\le t\), i.e. \(n+1-j\le k-1\) by strict monotonicity of \(j\mapsto\mu(B_j)\), so
\begin{equation*} \{s\ge0:\ m_f(s)\le t\}=[c_{n+1-k},+\infty),\qquad f^*(t)=c_{n+1-k}, \end{equation*}
while \(f^*(t)=0\) for \(t\ge\mu(B_n)=m_f(0)\). Re-indexing by \(j=n+1-k\) gives the first formula. For the second, the indices with \(\mu(B_j)>t\) are exactly \(j\ge k\), and \(\sum_{j\ge k}(c_{n+1-j}-c_{n-j})=c_{n+1-k}\) telescopes (and the sum is \(0\) for \(t\ge\mu(B_n)\)).
(ii) Put \(E:=\{s\ge0:\ m_f(s)\le t\}\) and \(F:=\{s\ge0:\ m_f(s)>t\}\), which partition \([0,+\infty)\). As \(m_f\) is nonincreasing, \(F\) is an initial segment, so every \(s>\sup F\) lies in \(E\) and no \(s<\sup F\) does, whence \(\inf E=\sup F\); if \(F=\varnothing\) both sides are \(0\), and if \(E=\varnothing\) both are \(+\infty\).
(iii) Here \(0\le f_n\uparrow|f|\) pointwise. If \(0\le u\le v\), then \(m_u\le m_v\) and hence \(u^*\le v^*\), so \(f_n^*\le f_{n+1}^*\le f^*\) and \(\varphi:=\lim_n f_n^*\le f^*\). If \(\varphi(t)<s<f^*(t)\) for some \(t\), then \(m_f(s)>t\) by definition of \(f^*\); but \(f_n^*(t)<s\) gives some \(s^{\prime}<s\) with \(m_{f_n}(s^{\prime})\le t\), so \(m_{f_n}(s)\le t\) for all \(n\), and \(\{f_n>s\}\uparrow\{|f|>s\}\) forces \(m_f(s)=\lim_n m_{f_n}(s)\le t\) – a contradiction.
(iv) For every \(s\ge0\),
\begin{equation*} \{t\ge0:\ f^*(t)>s\}=[0,\,m_f(s)), \end{equation*}
which gives \(\lambda(f^*>s)=m_f(s)\). Indeed, \(t\ge m_f(s)\) puts \(s\) in \(\{\sigma: m_f(\sigma)\le t\}\), so \(f^*(t)\le s\); and if \(t<m_f(s)\), right-continuity yields \(s^{\prime\prime}>s\) with \(m_f(s^{\prime\prime})>t\), so \(m_f(\sigma)>t\) for all \(\sigma\le s^{\prime\prime}\) and \(f^*(t)\ge s^{\prime\prime}>s\).
(v) By the layer-cake formula \(\int h\,d\theta=\int_0^\infty\theta(h>u)\,du\) (monotone convergence from simple \(h\), the integrand being nonincreasing in \(u\)), applied first to \(|g|\) and the measure \(\theta(A):=\int_A|f|\,d\mu\) and then, for fixed \(v\), to \(|f|I_{\{|g|>v\}}\) and \(\mu\),
\begin{equation*} \int_X|fg|\,d\mu=\int_0^{\infty}\!\!\int_0^{\infty} \mu\bigl(\{|f|>u\}\cap\{|g|>v\}\bigr)\,du\,dv , \end{equation*}
and the same computation on \(([0,\infty),\lambda)\) for \(f^*,g^*\) gives, using (iv) twice,
\begin{equation*} \int_0^{\infty}f^*g^*\,dt=\int_0^{\infty}\!\!\int_0^{\infty} \min\bigl(m_f(u),m_g(v)\bigr)\,du\,dv . \end{equation*}
Since \(\mu(\{|f|>u\}\cap\{|g|>v\})\le\min(m_f(u),m_g(v))\) pointwise, integrating in \((u,v)\) gives \(\int_X|fg|\,d\mu\le\int_0^\infty f^*g^*\,dt\).
Let us consider the measures \(H^s_\delta\) and \(H^s\) from 3.10(iii). Verify that if \(s<t\) and \(H^s(A)<\infty\), then \(H^t(A)=0\), and if \(H^s_\delta(A)=0\) for some \(\delta>0\), then \(H^s(A)=0\).
Both assertions follow from comparing one admissible cover with another; write \(\kappa_s:=\alpha(s)2^{-s}>0\) for the constant of Definition 3.10.8, so \(H^s_\delta\) is the infimum of \(\sum_j\kappa_s(\operatorname{diam}C_j)^s\) over covers with \(\operatorname{diam}C_j\le\delta\), and \(H^s=\sup_{\delta>0}H^s_\delta\).
(i) Let \(s<t\), \(M:=H^s(A)<\infty\) and \(\delta>0\). Since \(H^s_\delta(A)\le M\), pick a cover with \(\operatorname{diam}C_j\le\delta\) and \(\sum_j\kappa_s(\operatorname{diam}C_j)^s\le M+1\); it is admissible for \(H^t_\delta\), and \(\kappa_t(\operatorname{diam}C_j)^t\le(\kappa_t/\kappa_s)\delta^{t-s}\kappa_s(\operatorname{diam}C_j)^s\), so
\begin{equation*} H^t_\delta(A)\le\frac{\kappa_t}{\kappa_s}\,\delta^{t-s}(M+1)\xrightarrow[\delta\to0]{}0 . \end{equation*}
(ii) Let \(H^s_\delta(A)=0\) and \(0<\delta^{\prime}<\delta\); we show \(H^s_{\delta^{\prime}}(A)=0\), whence \(H^s(A)=0\). For \(s=0\) the hypothesis forces \(A=\varnothing\), so let \(s>0\) and \(0<\varepsilon<\kappa_s(\delta^{\prime})^s\). Some cover with \(\operatorname{diam}C_j\le\delta\) has \(\sum_j\kappa_s(\operatorname{diam}C_j)^s<\varepsilon\), so each \(\kappa_s(\operatorname{diam}C_j)^s<\kappa_s(\delta^{\prime})^s\), i.e. \(\operatorname{diam}C_j<\delta^{\prime}\); the cover is thus admissible for \(H^s_{\delta^{\prime}}\) and \(H^s_{\delta^{\prime}}(A)<\varepsilon\) for every such \(\varepsilon\).
(i) Show that, for every \(\alpha\in(0,1)\), there exists a set \(B_\alpha\subset[0,1]\) with the Hausdorff measure of order \(\alpha\) equal to \(1\).
(ii) Show that for the Cantor set \(C\) and \(\alpha=\ln 2/\ln 3\) we have \(0<H^\alpha( C)<\infty\).
Both parts come from the Cantor-type set \(C_r:=\bigcap_m E_m\), where \(r\in(0,1/2)\), \(E_0=[0,1]\), and \(E_{m+1}\) replaces each of the \(2^m\) intervals \(J\) of \(E_m\) (of length \(r^m\)) by the two end subintervals of length \(r|J|\); for the exponent \(s:=\ln2/\ln(1/r)\in(0,1)\) one has \(r^s=1/2\), and \(r\mapsto s\) maps \((0,1/2)\) onto \((0,1)\). We write \(\kappa_s:=\alpha(s)2^{-s}>0\) as in Definition 3.10.8. We prove \(\kappa_s(1-2r)^s\le H^s(C_r)\le\kappa_s\).
Carry the natural measure \(\nu:=P\circ\Phi^{-1}\), where \(P\) is the uniform product measure on \(\Omega=\{0,1\}^{\mathbb{N}}\) (see 3.5) and \(\Phi(\omega):=\sum_{k\ge1}\omega_k(1-r)r^{k-1}\) is continuous with \(\Phi(\Omega)=C_r\); since \(r<1/2\) the intervals of \(E_m\) are disjoint and \(\Phi^{-1}(J_w)=[w]\) for each word \(w\) of length \(m\), so \(\nu\) is a Borel probability measure on \(C_r\) with \(\nu(J_w)=2^{-m}\).
Upper bound: the \(2^m\) intervals of \(E_m\) have diameter \(r^m\) and cover \(C_r\), so for \(\delta\ge r^m\)
\begin{equation*} H^s_\delta(C_r)\le 2^m\kappa_s (r^m)^s = \kappa_s\,2^m (r^s)^m = \kappa_s\,2^m2^{-m}=\kappa_s . \end{equation*}
Letting \(m\to\infty\) gives \(H^s(C_r)\le\kappa_s<\infty\).
Lower bound, by the mass distribution principle: every interval \(I\) satisfies \(\nu(I)\le(1-2r)^{-s}|I|^{s}\). This is clear when \(|I|\ge1-2r\); for \(0<|I|<1-2r\) let \(m\ge1\) be largest with \(r^{m-1}(1-2r)>|I|\). Distinct intervals of \(E_m\) are separated by gaps of length at least \(r^{m-1}(1-2r)\) (the smallest gaps being those cut at step \(m\)), so \(I\) meets at most one interval \(J\) of \(E_m\) and, by maximality \(r^m\le|I|/(1-2r)\),
\begin{equation*} \nu(I)\le\nu(J)=2^{-m}=(r^{m})^{s}\le(1-2r)^{-s}|I|^{s} \end{equation*}
(and \(\nu\) has no atoms, so \(|I|=0\) is trivial). Replacing each set \(U_j\) of a cover of \(C_r\) by \([\inf U_j,\sup U_j]\), of the same diameter, subadditivity of \(\nu\) gives
\begin{equation*} 1=\nu(C_r)\le(1-2r)^{-s}\sum_j(\operatorname{diam}U_j)^s , \end{equation*}
so \(H^s_\delta(C_r)\ge\kappa_s(1-2r)^s\) for every \(\delta>0\), whence \(H^s(C_r)\ge\kappa_s(1-2r)^s>0\).
(ii) Take \(r=1/3\): then \(C_r=C\), \(s=\ln2/\ln3\), and the bounds read \(0<\kappa_s3^{-s}\le H^s( C)\le\kappa_s<\infty\).
(i) Let \(s=\alpha\in(0,1)\), \(r=2^{-1/s}\in(0,1/2)\) and \(m_0:=H^s(C_r)\in(0,\infty)\) by the above. By 3.10(iii), \(H^s\) is translation invariant with \(H^s(\lambda A)=\lambda^sH^s(A)\), and its restriction to Borel sets is a countably additive measure (Proposition 3.10.9 with Theorem 1.11.4). Put
\begin{equation*} A_N:=\bigcup_{k=0}^{N-1}\Big(\frac{k}{N}+\frac{1}{2N}C_r\Big)\subset[0,1], \end{equation*}
whose \(N\) summands are compact, pairwise disjoint (the \(k\)-th lies in \([k/N,k/N+1/(2N)]\)) and each of measure \((2N)^{-s}m_0\), so
\begin{equation*} H^s(A_N)=N(2N)^{-s}m_0=2^{-s}N^{1-s}m_0\xrightarrow[N\to\infty]{}+\infty \end{equation*}
as \(1-s>0\); fix \(N\) with \(H^s(A_N)\ge1\) and set \(A:=A_N\), so \(1\le H^s(A)<\infty\). Then \(\varphi(u):=H^s(A\cap[0,u])\) is nondecreasing with \(\varphi(0)=0\) (points are \(H^s\)-null for \(s>0\)) and \(\varphi(1)\ge1\), and it is continuous, the finite Borel measure \(B\mapsto H^s(A\cap B)\) being continuous from below and above and having no atoms. By the intermediate value theorem \(\varphi(u_0)=1\) for some \(u_0\), and \(B_\alpha:=A\cap[0,u_0]\subset[0,1]\) has \(H^\alpha(B_\alpha)=1\).
Let \(H^s\) be the Hausdorff measure on \(\mathbb{R}^n\). Prove that the \(H^s\)-measure of every Borel set \(B\subset\mathbb{R}^n\) equals the supremum of the \(H^s\)-measures of compact subsets of \(B\).
Only \(H^s(B)\le\iota(B):=\sup\{H^s(K):K\subset B\ \text{compact}\}\) needs proof, monotonicity giving the reverse. By Proposition 3.10.9 all Borel sets are \(H^s\)-measurable, so \(H^s|_{\mathcal{B}(\mathbb{R}^n)}\) is a countably additive measure (Theorem 1.11.4); we use this throughout, and write \(\kappa_s=\alpha(s)2^{-s}\), \(H^s=\sup_{\delta>0}H^s_\delta\).
(i) \(H^s(B)<\infty\). Then \(\mu(A):=H^s(A\cap B)\) is a finite Borel measure, so by Theorem 1.4.8 there are compact \(K_\varepsilon\subset B\) with \(H^s(B\setminus K_\varepsilon)<\varepsilon\), whence \(H^s(K_\varepsilon)\ge H^s(B)-\varepsilon\) by subadditivity and \(\iota(B)\ge H^s(B)\).
(ii) \(H^s(B)=\infty\). Fix \(C>0\); we seek compact \(K\subset B\) with \(H^s(K)>C\). Continuity from below of the Borel measure \(H^s\) applied to \(B\cap\{|x|\le R\}\uparrow B\) supplies a bounded Borel \(B^{\prime}\subset B\) with \(H^s(B^{\prime})>C\), and then \(\delta>0\) with \(H^s_\delta(B^{\prime})>C\) (also \(H^s_\delta(B^{\prime})<\infty\), a bounded set being coverable by finitely many sets of diameter \(\le\delta\)). Since \(H^s\ge H^s_\delta\), it now suffices to find compact \(K\subset B^{\prime}\) with \(H^s_\delta(K)>C\).
This last step is not the routine repetition of (i) that the book’s hint suggests: Borel sets need not be \(H^s_\delta\)-measurable, since \(H^s_\delta\) is not a metric outer measure. For \(n=1\), \(s=0\), \(H^0_\delta(A)\) is the least number of sets of diameter \(\le\delta\) covering \(A\), and \(H^0_\delta(\{0,\delta\})=1<2=H^0_\delta(\{0\})+H^0_\delta(\{\delta\})\), so \(\{0\}\) is not \(H^0_\delta\)-measurable. The content of the step is that Borel sets are capacitable for the Hausdorff pre-measure:
\begin{equation*} H^s_\delta(B^{\prime})=\sup\{H^s_\delta(K):\ K\subset B^{\prime},\ K\ \text{compact}\} \end{equation*}
for Borel \(B^{\prime}\). The proof is that of Theorem 1.10.5, whose argument for \(\mu^*(A)=\sup\{\mu(E):E\subset A,\ E\in\mathcal{E}\}\) on \(\mathcal{E}\)-Souslin sets uses only two properties of the set function, applied to the class \(\mathcal{E}\) of compact sets, which is stable under finite unions and countable intersections and generates all Borel sets by the \(A\)-operation (Theorem 1.10.4(ii)). It is convenient to use the strict-diameter variant
\begin{equation*} G^s_\delta(A):=\inf\Big\{\sum_j\kappa_s(\operatorname{diam}C_j)^s:\ A\subset\bigcup_jC_j,\ \operatorname{diam}C_j<\delta\Big\}. \end{equation*}
Here \(H^s_{\delta}\le G^s_\delta\le H^s_{\delta^{\prime}}\) for \(\delta^{\prime}<\delta\), so \(\sup_{\delta>0}G^s_\delta=H^s\), and a compact \(K\) with \(G^s_\delta(K)>C\) has \(H^s(K)>C\); thus it suffices to prove capacitability for \(G^s_\delta\), for which Theorem 1.10.5 needs exactly
(A) \(G^s_\delta(A_k)\to G^s_\delta(A)\) for arbitrary \(A_k\uparrow A\);
(B) \(G^s_\delta(K_k)\to G^s_\delta(K)\) for compact \(K_k\downarrow K\).
(B) holds: given \(\varepsilon>0\), cover \(K\) by sets \(C_j\) of diameter \(<\delta\) with \(\sum_j\kappa_s(\operatorname{diam}C_j)^s<G^s_\delta(K)+\varepsilon\) and fatten each to an open \(V_j:=\{\operatorname{dist}(\cdot,C_j)<\eta_j\}\) with \(\operatorname{diam}V_j<\delta\) and \(\kappa_s(\operatorname{diam}V_j)^s\le\kappa_s(\operatorname{diam}C_j)^s+\varepsilon2^{-j}\); compactness gives \(K_k\subset\bigcup_jV_j\) for large \(k\), so \(G^s_\delta(K_k)\le G^s_\delta(K)+2\varepsilon\), and monotonicity gives the reverse. Granting (A), the proof of Theorem 1.10.5 runs verbatim with \(\mu^*\) replaced by \(G^s_\delta\): from a monotone Souslin scheme of compact sets for the Borel set \(B^{\prime}\) one selects \(m_1,m_2,\dots\) with \(G^s_\delta(M_{m_1,\ldots,m_k})>G^s_\delta(B^{\prime})-\varepsilon\) (this is where (A) enters, as Proposition 1.5.12 does in the book), and \(K:=\bigcap_kD_{m_1,\ldots,m_k}\subset B^{\prime}\) is compact with \(G^s_\delta(K)\ge G^s_\delta(B^{\prime})-\varepsilon>C\) by (B).
Property (A) is the increasing sets lemma for the Hausdorff content, a theorem of R. O. Davies (Proc. London Math. Soc. (3) 20 (1970), 222–236), and is the one ingredient taken on trust here; it cannot be replaced by the corresponding property of \(H^s\), which satisfies (A) but fails (B) (in \(\mathbb{R}^2\) with \(s=1\), balls of radius \(1/k\) about a point have \(H^1=\infty\) while the point is \(H^1\)-null).
Case (ii) is complete without (A) whenever \(H^s\) is \(\sigma\)-finite on \(B\): writing \(B=\bigcup_kB_k\) with \(B_k\uparrow\) Borel of finite measure, \(H^s(B_k)\to\infty\), and case (i) applies to some \(B_k\) with \(H^s(B_k)>C\). This covers \(s\ge n\) (for \(s=n\), \(H^n\) is Lebesgue measure by Proposition 3.10.11; for \(s>n\), \(H^s\equiv0\)) and \(s=0\) (counting measure: an infinite \(B\) has finite subsets of arbitrarily large cardinality). Only a Borel set on which \(H^s\) is not \(\sigma\)-finite, \(0<s<n\), needs (A). (A fully rigorous proof of the remaining step is beyond the scope of this page; see the reference given in the book.)
Exercises 3.10.106–3.10.112
Let \(H^s\) be the Hausdorff measure on \(\mathbb{R}^n\) and let \(K \subset \mathbb{R}^n\) be a compact set with \(H^s(K) = \infty\). Prove that there exists a compact set \(C \subset K\) with \(0 < H^s( C) < \infty\).
It suffices to produce a Borel \(F \subset K\) with \(0 < H^s(F) < \infty\): Exercise 3.10.105 then gives a compact \(C \subset F\) with \(H^s( C)>0\), and \(H^s( C)\le H^s(F)<\infty\). For \(s=0\) the counting measure \(H^0\) makes \(K\) infinite and any point works, so let \(s>0\); by translation invariance and \(H^s(rA)=r^sH^s(A)\) we may assume \(K \subset [0,1)^n\).
The tool is the net measures built on the families \(\mathcal{D}_k\) of the \(2^{kn}\) half-open dyadic cubes \[ \prod_{i=1}^n [j_i 2^{-k}, (j_i+1)2^{-k}), \qquad j_i \in \{0,\dots,2^k-1\}, \] contained in \([0,1)^n\). Every cube of \(\mathcal{D}_k\) has diameter \(d_k := \sqrt{n}\,2^{-k}\), the cubes of \(\mathcal{D}_k\) partition \([0,1)^n\), and every cube of \(\mathcal{D}_l\) with \(l \ge k\) is contained in exactly one cube of \(\mathcal{D}_k\). For \(A \subset [0,1)^n\) and \(k \in \mathbb{N}\) put \[ \mathcal{M}_k(A) = \inf\Big\{ \sum_i d_{l_i}^s : A \subset \bigcup_i Q_i,\ Q_i \in \mathcal{D}_{l_i},\ l_i \ge k \Big\}. \] Their relevant properties, with \(\mathcal{M}:=\lim_k\mathcal{M}_k\), are:
(P1) each \(\mathcal{M}_k\) is monotone and countably subadditive, and \(\mathcal{M}_k \le \mathcal{M}_{k+1}\) (a finer net admits fewer covers);
(P2) \(\mathcal{M}_k(A) = \sum_{Q \in \mathcal{D}_k} \mathcal{M}_k(A \cap Q)\), since every admissible cube lies in exactly one \(Q \in \mathcal{D}_k\), so an admissible cover splits into admissible covers of the \(A \cap Q\);
(P3) \(\mathcal{M}_k(A) \le d_k^s\) if \(A \subset Q\in\mathcal{D}_k\) (cover \(A\) by \(Q\));
(P4) \(H^s_{d_k} \le \mathcal{M}_k\), the admissible cubes having diameter at most \(d_k\); hence \(H^s(A) \le \liminf_k \mathcal{M}_k(A)\);
(P5) \(\mathcal{M} \le 2^n (2\sqrt{n})^s H^s\): given a cover \(\{U_i\}\) with \(\operatorname{diam}U_i\le\delta\), discard singletons (coverable by dyadic cubes of arbitrarily small cost as \(s>0\)) and take \(l_i\) with \(2^{-l_i-1} \le \operatorname{diam} U_i < 2^{-l_i}\); each \(U_i\) meets at most \(2^n\) cubes of \(\mathcal{D}_{l_i}\), each costing \(d_{l_i}^s \le (2\sqrt{n}\operatorname{diam}U_i)^s\).
So \(\mathcal{M}_k(K)\uparrow\infty\) by (P4), and it suffices to find Borel \(F\subset K\) with \(0<\mathcal{M}(F)<\infty\), positivity of \(H^s(F)\) following from (P5) and finiteness from (P4).
Finiteness comes from a greedy nested selection: we construct integers \(k_1 < k_2 < \cdots\) and finite families \(\mathcal{C}_m \subset \mathcal{D}_{k_m}\) such that, putting \(K_m := K \cap \bigcup_{Q \in \mathcal{C}_m} Q\), \[ 1 < \mathcal{M}_{k_m}(K_m) \le 1 + 2^{-m}, \qquad K_1 \supset K_2 \supset \cdots. \] Choose \(k_1\) with \(\mathcal{M}_{k_1}(K) > 1\) and \(d_{k_1}^s \le 2^{-1}\). By (P2) the sum \(\sum_{Q \in \mathcal{D}_{k_1}} \mathcal{M}_{k_1}(K \cap Q)=\mathcal{M}_{k_1}(K)\) is finite with terms at most \(d_{k_1}^s\) by (P3); let \(\mathcal{C}_1\) be the shortest initial segment of the cubes with nonzero terms whose sum exceeds \(1\), so that (P2) applied to \(K_1\) gives \(1 < \mathcal{M}_{k_1}(K_1) \le 1 + 2^{-1}\), the discarded last term being at most \(2^{-1}\). Given \(\mathcal{C}_m\), we have \(\mathcal{M}_{k}(K_m)>1\) for all \(k \ge k_m\) by (P1); choose \(k_{m+1} > k_m\) with \(d_{k_{m+1}}^s \le 2^{-m-1}\) and repeat the selection for \(K_m\) at level \(k_{m+1}\). Each \(Q \in \mathcal{C}_{m+1}\) meets \(K_m\), hence lies in a cube of \(\mathcal{C}_m\), so \(K_{m+1} \subset K_m\) and \(1 < \mathcal{M}_{k_{m+1}}(K_{m+1}) \le 1 + 2^{-m-1}\). Then \(F := \bigcap_m K_m\) is Borel with, by (P4), \[ H^s_{d_{k_m}}(F) \le \mathcal{M}_{k_m}(K_m) \le 1 + 2^{-m}, \quad\text{so}\quad H^s(F)\le1 . \]
For the lower bound the device is the mass distribution principle: a nonzero finite Borel measure \(\mu\) concentrated on \(F\) with \(\mu(B(x,r)) \le C r^s\) for \(r<r_0\) forces \(H^s(F) \ge \mu(F)/(C 2^s)>0\), since each \(U_i\) of a cover meeting \(F\) lies in a ball \(B(x_i,\operatorname{diam}U_i)\), whence \(\mu(F) \le C2^s\sum_i (\operatorname{diam} U_i)^s\). Such a \(\mu\) is manufactured on a Cantor scheme with consecutive scales \(k_m = k_1+m-1\) by carrying weights \(\mu_m(Q) \ge 0\), \(Q \in \mathcal{C}_m\), with the invariants \[ \mu_m(Q) \le \mathcal{M}_{k_m}(K_m \cap Q), \qquad \sum_{Q \in \mathcal{C}_m} \mu_m(Q) = \beta, \qquad \sum_{\substack{Q \in \mathcal{C}_{m+1} \\ Q \subset Q^{\prime}}} \mu_{m+1}(Q) = \mu_m(Q^{\prime}) . \] The first invariant propagates: for \(Q^{\prime} \in \mathcal{C}_m\), (P1) and (P2) give \(\sum_{Q \subset Q^{\prime}} \mathcal{M}_{k_{m+1}}(K_m \cap Q) = \mathcal{M}_{k_{m+1}}(K_m \cap Q^{\prime}) \ge \mu_m(Q^{\prime})\), so the mass of \(Q^{\prime}\) can be split among its children within their caps; by (P3) the cap forces \(\mu_m(Q)\le d_{k_m}^s\). A weak limit point \(\mu\) of \(\sum_{Q \in \mathcal{C}_m} \mu_m(Q)\delta_{x_Q}\), \(x_Q \in K_m\cap Q\), then has mass \(\beta\), sits on \(\bigcap_m(K \cap \bigcup_{Q \in \mathcal{C}_m} \overline{Q})\), and satisfies \(\mu(B(x,r)) \le 4^n(2\sqrt{n}\,r)^s\) for small \(r\), a ball of radius \(r\) lying in an open cube of side \(3\cdot2^{-k}\) (with \(2^{-k}\) the largest dyadic scale exceeding \(2r\)) that meets at most \(4^n\) cubes of \(\mathcal{D}_k\), each of \(\mu\)-mass at most \(d_k^s\).
What is missing is running the two constructions at once: the greedy selection must be arranged so that the mass scheme survives with \(\beta\) bounded away from \(0\). The greedy step guarantees \(\mathcal{M}_{k_m}(K_m) > 1\) only at the current scale, whereas a lower bound for \(H^s(\bigcap_m K_m)\) needs \(\mathcal{M}_{k_j}(K_m) \ge c\) for every coarse scale \(k_j\), \(j \le m\), uniformly in \(m\) – and \(\mathcal{M}_{k_j} \le \mathcal{M}_{k_m}\) makes the retained inequality the wrong one; nor does a compactness passage to the limit repair it, since covering \(\bigcap_mK_m\) by cubes of levels \(\ge k\) is priced by \(\mathcal{M}_k(K_m)\), which is not bounded below. This is Besicovitch’s original construction, for which the book’s hint refers to Federer [282, Theorem 2.10.47]; granting it, \(F\) is a Borel subset of \(K\) with \(0<H^s(F)<\infty\) and the first reduction yields \(C\). (A fully rigorous proof of the remaining step is beyond the scope of this page; see the reference given in the book.)
(Erdos, Taylor [272]) Let \(A_n\) be Lebesgue measurable sets in \([0,1]\) with \(\lambda(A_n) \ge \varepsilon > 0\) for all \(n \in \mathbb{N}\). Show that, for every continuous monotonically increasing function \(\varphi\) with \(\varphi(0) = 0\) and \(\lim_{t \to 0+} \varphi(t)/t = +\infty\), there exists a subsequence \(n_k\) such that the set \(\bigcap_{k=1}^\infty A_{n_k}\) has infinite measure with respect to the Hausdorff measure generated by the function \(\varphi\).
The witness is a Cantor scheme inside a weak-star limit density, carrying a measure \(\mu\) with \(\mu(J)\le2^{-k}\varphi(|J|)\) for \(|J|<\ell_k\). Write \(H^\varphi_\delta(A)=\inf\{\sum_i\varphi(\operatorname{diam}U_i):\operatorname{diam}U_i\le\delta\}\), \(H^\varphi=\lim_{\delta\to0}H^\varphi_\delta\), and \[ \Psi(\delta) := \sup_{0<t\le\delta} t/\varphi(t) \xrightarrow[\delta\to0+]{} 0 , \] the hypothesis \(\varphi(t)/t\to+\infty\) being exactly this.
Since \(\{I_{A_n}\}\) is bounded in \(L^\infty[0,1]=(L^1[0,1])^{\prime}\) with \(L^1[0,1]\) separable, a weak-star convergent subsequence exists; passing to it and relabelling (every later subsequence is one of the original sequence), there is \(g \in L^\infty[0,1]\), \(0 \le g \le 1\), with \[ \lim_{n \to \infty} \lambda(A_n \cap B) = \int_B g \, d\lambda \quad \text{for every measurable } B \subset [0,1]. \] With \(B=[0,1]\) this gives \(\int_0^1g\ge\varepsilon\); put \(\eta := \varepsilon/2\) and \(G := \{g \ge \eta\}\), so \(\varepsilon \le \int_0^1 g \le \eta + \lambda(G)\) yields \(\lambda(G) \ge \varepsilon/2>0\) and \[ \lim_{n \to \infty} \lambda(A_n \cap B) = \int_B g \, d\lambda \ge \eta \lambda(B) \quad \text{whenever } B \subset G. \tag{1} \]
The density device: for compact \(F \subset [0,1]\) and \(\ell > 0\) put \[ F^\sharp(\ell) := \Big\{ x \in F : \lambda\big(F \cap [x-r,x+r]\big) \ge \tfrac{3}{2} r \ \text{ for all } 0 < r \le \ell \Big\}. \] Then \(F^\sharp(\ell)\) is compact (the function \((x,r) \mapsto \lambda(F \cap [x-r,x+r])\) being continuous), increases as \(\ell\) decreases, and by the Lebesgue density theorem covers almost all of \(F\) in the limit, so \(\lambda(F^\sharp(\ell) \cap B) \to \lambda(F \cap B)\) as \(\ell \to 0+\) for every measurable \(B\).
The scheme: we construct indices \(n_1 < n_2 < \cdots\), numbers \(\ell_1 > \ell_2 > \cdots \to 0\), compact sets \(F_1 \supset F_2 \supset \cdots\) and finite families \(\mathcal{I}_k\) of non-overlapping closed intervals of length \(\ell_k\) (non-overlapping meaning that their interiors are pairwise disjoint, so that two of them can meet only in a common endpoint, a set of measure zero), such that for all \(k\):
(a) \(F_k \subset A_{n_1} \cap \cdots \cap A_{n_k} \cap G\) and \(F_k \subset \bigcup_{I \in \mathcal{I}_k} I\);
(b) \(\lambda(F_k \cap I) \ge \tfrac{1}{2}\ell_k\) for every \(I \in \mathcal{I}_k\);
(c) every \(I \in \mathcal{I}_{k+1}\) is contained in exactly one \(I^{\prime} \in \mathcal{I}_k\), and every \(I^{\prime} \in \mathcal{I}_k\) contains at least \[ M_{k+1} := \frac{\eta\, \ell_k}{64\, \ell_{k+1}} \] intervals of \(\mathcal{I}_{k+1}\).
To start, (1) with \(B = G\) gives \(n_1\) with \(\lambda(A_{n_1} \cap G) \ge \eta\lambda(G)/2 > 0\), and inner regularity a compact \(F \subset A_{n_1} \cap G\) with \(\lambda(F) > 0\); take \(\ell_1\) so small that \(\lambda(F^\sharp(\ell_1)) \ge \lambda(F)/2\), let \(x_1,\dots,x_N\) be a maximal \(\ell_1\)-separated subset of \(F^\sharp(\ell_1)\), and let \(\mathcal{I}_1\) consist of the non-overlapping intervals \([x_i \pm \ell_1/2]\), each with \(\lambda(F \cap I) \ge \tfrac{3}{4}\ell_1\) by the definition of \(F^\sharp\) at \(r=\ell_1/2\). Then \(F_1 := F \cap \bigcup_{I \in \mathcal{I}_1} I\) satisfies (a), (b), and maximality gives \(\lambda(F^\sharp(\ell_1)) \le 2N\ell_1\), i.e. \(N \ge \lambda(F)/(4\ell_1)\).
Given \(\mathcal{I}_k\), \(F_k\): for \(I \in \mathcal{I}_k\) the set \(F_k \cap I\) lies in \(G\) and has measure at least \(\ell_k/2\), so by (1) \[ \lim_{n \to \infty}\lambda(A_n \cap F_k \cap I) \ge \eta\,\lambda(F_k \cap I) \ge \tfrac{\eta}{2}\ell_k . \] As \(\mathcal{I}_k\) is finite, there is \(n_{k+1} > n_k\) with \(\lambda(A_{n_{k+1}} \cap F_k \cap I) \ge \tfrac{\eta}{4}\ell_k\) for all \(I \in \mathcal{I}_k\). By inner regularity choose a compact \(F \subset F_k \cap A_{n_{k+1}}\) with \(\lambda(F \cap I) \ge \tfrac{\eta}{8}\ell_k\) for all \(I \in \mathcal{I}_k\). By Step 2 choose \(\ell_{k+1} > 0\) so small that \[ \lambda\big(F^\sharp(\ell_{k+1}) \cap I\big) \ge \tfrac{\eta}{16}\,\ell_k \quad \text{for all } I \in \mathcal{I}_k, \tag{2} \] and additionally so small that \(\ell_{k+1} \le \eta \ell_k / 100\) and that the smallness condition (4) below is satisfied. In each \(I \in \mathcal{I}_k\) take a maximal \(\ell_{k+1}\)-separated subset of the part of \(F^\sharp(\ell_{k+1}) \cap I\) lying at distance more than \(\ell_{k+1}\) from the endpoints of \(I\) (discarding the endpoint strips costs at most \(2\ell_{k+1} \le \eta\ell_k/50\) of measure, so by (2) the remaining set still has measure at least \(\eta \ell_k/32\)), and let \(\mathcal{I}_{k+1}\) consist of the intervals of length \(\ell_{k+1}\) centred at all the chosen points. These intervals are non-overlapping, each lies in its parent, and by maximality of the separated set the union of the intervals of length \(2\ell_{k+1}\) centred at the chosen points covers a set of measure at least \(\eta\ell_k/32\), so the number of chosen points inside a given \(I \in \mathcal{I}_k\) is at least \[ M_{k+1} := \frac{\eta\,\ell_k}{64\,\ell_{k+1}} . \] Finally put \(F_{k+1} := F \cap \bigcup_{I \in \mathcal{I}_{k+1}} I\). Property (b) at level \(k+1\) holds because each \(J \in \mathcal{I}_{k+1}\) is centred at a point of \(F^\sharp(\ell_{k+1})\), so \(\lambda(F_{k+1} \cap J) = \lambda(F \cap J) \ge \tfrac{3}{4}\ell_{k+1}\).
Put \(E := \bigcap_k F_k\) and \(E^{\prime} := \bigcap_k \bigcup_{I \in \mathcal{I}_k} I\); both are compact and \(E \subset E^{\prime}\). Conversely, if \(z \in E^{\prime}\), then for each \(k\) there is a unique \(I_k \in \mathcal{I}_k\) containing \(z\), and these are nested by (c); the centre \(x_{I_j}\) lies in \(F_j \subset F_k\) for \(j \ge k\) and \(|z - x_{I_j}| \le \ell_j/2 \to 0\), so \(z \in F_k\) for every \(k\), i.e. \(z \in E\). Thus \(E = E^{\prime}\), and by (a) \[ E \subset \bigcap_{k=1}^\infty A_{n_k}. \]
Give each \(I \in \mathcal{I}_1\) mass \(1/\#\mathcal{I}_1\) and split each interval’s mass equally among its children; pushing the resulting product measure on branches forward by the coding map gives a Borel probability measure \(\mu\) on \(E\) with \(\mu(I)\) the assigned mass (unambiguous, as distinct intervals of a generation share at most an endpoint and points are \(\mu\)-null by \(p_m\to0\) below). Let \[ p_k := \max_{I \in \mathcal{I}_k} \mu(I), \qquad \text{so} \qquad p_{k+1} \le \frac{p_k}{M_{k+1}} \le \frac{64\, p_k\, \ell_{k+1}}{\eta\, \ell_k}. \tag{3} \] Since \(\#\mathcal{I}_1 \ge \lambda(F)/(4\ell_1)\) we have \(p_1 \le 4\ell_1/\lambda(F)\). Moreover \(\ell_{k+1} \le \eta\ell_k/100\) gives \(M_{k+1} \ge 100/64 > 1\), so the sequence \(p_k\) decreases geometrically to \(0\); in particular \(\mu\) has no atoms.
Growth estimate: let \(J\) be an interval with \(\ell_{k+1} \le |J| < \ell_k\). Since the intervals of \(\mathcal{I}_k\) are non-overlapping of length \(\ell_k > |J|\), \(J\) meets at most two of them, so \(\mu(J) \le 2p_k\). Since the intervals of \(\mathcal{I}_{k+1}\) are non-overlapping of length \(\ell_{k+1} \le |J|\), \(J\) meets at most \(|J|/\ell_{k+1} + 2 \le 3|J|/\ell_{k+1}\) of them, so by (3) \[ \mu(J) \le \frac{3|J|}{\ell_{k+1}}\,p_{k+1} \le \frac{192\,p_k}{\eta\,\ell_k}\,|J| . \] Consequently, for every interval \(J\) with \(\ell_{k+1} \le |J| < \ell_k\), \[ \frac{\mu(J)}{\varphi(|J|)} \le \max\Big\{ \frac{192\,p_k}{\eta\,\ell_k}\cdot\frac{|J|}{\varphi(|J|)},\ \frac{2p_k}{\varphi(|J|)} \Big\} , \] and we use the first bound when \(|J| \le \eta\ell_k/200\) and the second one when \(|J| > \eta \ell_k /200\); in the latter case \(\varphi(|J|) \ge \varphi(\eta\ell_k/200)\) because \(\varphi\) is increasing. Therefore, for all intervals \(J\) in the window \(\ell_{k+1} \le |J| < \ell_k\), \[ \frac{\mu(J)}{\varphi(|J|)} \le \frac{192\,p_k}{\eta\,\ell_k}\,\Psi(\ell_k) + \frac{2p_k}{\varphi(\eta\ell_k/200)} =: \varkappa_k . \] We now impose on the choice of \(\ell_{k+1}\) in Step 3 the additional smallness requirement \[ \frac{192\,p_{k+1}}{\eta\,\ell_{k+1}}\,\Psi(\ell_{k+1}) + \frac{2p_{k+1}}{\varphi(\eta\ell_{k+1}/200)} \le 2^{-k-1}, \tag{4} \] which by (3) is implied by \[ \frac{192 \cdot 64\, p_k}{\eta^2 \ell_k}\,\Psi(\ell_{k+1}) + \frac{128\, p_k \ell_{k+1}}{\eta\,\ell_k\,\varphi(\eta \ell_{k+1}/200)} \le 2^{-k-1} . \] Both terms on the left tend to \(0\) as \(\ell_{k+1} \to 0+\) with \(k\) fixed: the first because \(\Psi(\delta) \to 0\), the second because \(\ell_{k+1}/\varphi(\eta\ell_{k+1}/200) \le (200/\eta)\Psi(\eta\ell_{k+1}/200) \to 0\). Hence (4) can indeed be met, and likewise the analogous requirement \(\varkappa_1 \le 1/2\) can be met at the first step by choosing \(\ell_1\) small (recall \(p_1 \le 4\ell_1/\lambda(F)\), so \(\varkappa_1 \le \frac{768}{\eta \lambda(F)}\Psi(\ell_1) + \frac{8\ell_1}{\lambda(F)\varphi(\eta\ell_1/200)}\), and both terms tend to \(0\) with \(\ell_1\)).
Thus \(\varkappa_k \le 2^{-k}\) for all \(k\), and consequently \[ \mu(J) \le \varkappa_k\,\varphi(|J|) \le 2^{-k}\varphi(|J|) \quad \text{for every interval } J \text{ with } |J| < \ell_k . \tag{5} \] Indeed, given such a \(J\) with \(|J| > 0\), let \(j\) be the unique index with \(\ell_{j+1} \le |J| < \ell_j\) (it exists since \(\ell_m \downarrow 0\)); then \(j \ge k\), since \(j < k\) would force \(|J| \ge \ell_{j+1} \ge \ell_k\), and therefore \(\mu(J) \le \varkappa_j \varphi(|J|) \le 2^{-j}\varphi(|J|) \le 2^{-k}\varphi(|J|)\). For a degenerate \(J\) both sides vanish, because \(\mu(\{x\}) \le p_m \to 0\) for every \(x\).
Now let \(\delta < \ell_1\) and let \(\{U_i\}\) be any cover of \(E\) with \(\operatorname{diam} U_i \le \delta\). Let \(k(\delta)\) be the largest \(k\) with \(\ell_k > \delta\), so that every \(U_i\) satisfies \(\operatorname{diam} U_i \le \delta < \ell_{k(\delta)}\). Replacing \(U_i\) by the closed interval \(J_i\) of the same diameter containing it, we obtain from (5) \[ 1 = \mu(E) \le \sum_i \mu(J_i) \le 2^{-k(\delta)} \sum_i \varphi(\operatorname{diam} U_i). \] Hence \(H^\varphi_\delta(E) \ge 2^{k(\delta)}\), and \(k(\delta) \to \infty\) as \(\delta \to 0\) because \(\ell_k \to 0\). Therefore \(H^\varphi(E) = \infty\), and since \(E \subset \bigcap_{k} A_{n_k}\), also \[ H^\varphi\Big( \bigcap_{k=1}^\infty A_{n_k} \Big) = \infty . \]
(Darst [204]) Prove that there exist an infinitely differentiable function \(f\) on the real line and a set \(Z\) of Lebesgue measure zero such that the set \(f^{-1}(Z)\) is not Lebesgue measurable.
Take \(f(x):=\int_0^xg\,dt\), where \(g\in C^\infty\) is nonnegative and vanishes exactly on a fat Cantor set \(P\), and \(Z:=f(N)\) for a nonmeasurable \(N\subset P\): then \(f\) is strictly increasing, \(f(P)\) is null, and \(f^{-1}(Z)=N\).
Let \(P \subset (0,1)\) be compact nowhere dense with \(\lambda(P) > 0\) (remove from \([1/4,3/4]\) middle intervals of total length \(<1/4\)), \(\alpha := \min P\), \(\beta := \max P\), and let \((a_j,b_j)\subset(\alpha,\beta)\) be the bounded components of \(\mathbb{R}\setminus P\), the unbounded ones being \((-\infty,\alpha)\) and \((\beta,+\infty)\). For a bounded interval \((a,b)\) set \[ \psi_{a,b}(x) = \exp\Big( -\frac{1}{(x-a)(b-x)} \Big) \ \text{ for } x \in (a,b), \qquad \psi_{a,b}(x) = 0 \ \text{ for } x \notin (a,b). \] Then \(\psi_{a,b} \in C^\infty(\mathbb{R})\), all derivatives of \(\exp(-1/t)\) tending to \(0\) as \(t \to 0+\), with \(0 \le \psi_{a,b} \le 1\) and \(\psi_{a,b} > 0\) exactly on \((a,b)\). Similarly put \(\psi_-(x) = \exp\big(-1/(\alpha-x)\big)\) for \(x < \alpha\) and \(\psi_-(x)=0\) for \(x \ge \alpha\), and \(\psi_+(x) = \exp\big(-1/(x-\beta)\big)\) for \(x > \beta\), \(\psi_+(x)=0\) for \(x \le \beta\); these are again \(C^\infty\), bounded together with all their derivatives, positive exactly on the respective unbounded components.
Choose numbers \(c_j > 0\) such that \[ c_j \sup_{x \in \mathbb{R}} \big| \psi_{a_j,b_j}^{(k)}(x) \big| \le 2^{-j} \quad \text{for all } k = 0,1,\dots,j, \] which is possible because each \(\psi_{a_j,b_j}\) has bounded derivatives of every order. Put \[ g := \psi_- + \psi_+ + \sum_{j=1}^\infty c_j\, \psi_{a_j,b_j} . \] For each \(k\) the series of \(k\)-th derivatives converges uniformly (its terms with \(j \ge k\) are bounded by \(2^{-j}\)), so \(g \in C^\infty(\mathbb{R})\), \(g \ge 0\), and \(g>0\) exactly off \(P\), the components being disjoint. Hence \(f:=\int_0^{\cdot}g\,dt\) is \(C^\infty\) with \(f^{\prime}=g\), and it is strictly increasing: for \(x<y\) the interval \((x,y)\) is not inside the nowhere dense \(P\), so \(\int_x^yg>0\).
Now \(f(P)\) is null. Since \(f\) is a strictly increasing continuous bijection of \([\alpha,\beta]\) onto \([f(\alpha),f(\beta)]\), the images \(f\big((a_j,b_j)\big) = \big(f(a_j),f(b_j)\big)\) of the bounded components are pairwise disjoint open subintervals of \(\big(f(\alpha),f(\beta)\big)\), disjoint also from \(f(P)\), and \[ [f(\alpha),f(\beta)] = f(P) \cup \bigcup_j \big(f(a_j),f(b_j)\big) . \] Because \(g\) vanishes on \(P\) and \(\lambda(P \cap [\alpha,\beta]) = \lambda(P)\), we get \[ \sum_j \big( f(b_j) - f(a_j) \big) = \sum_j \int_{a_j}^{b_j} g\,dt = \int_{[\alpha,\beta] \setminus P} g \,dt = \int_\alpha^\beta g\,dt = f(\beta) - f(\alpha). \] Hence the disjoint open intervals \(\big(f(a_j),f(b_j)\big)\) fill \(\big(f(\alpha),f(\beta)\big)\) up to a set of measure zero, and since \(f(P)\) is disjoint from all of them, \[ \lambda\big( f(P) \big) = \big(f(\beta)-f(\alpha)\big) - \sum_j \big( f(b_j)-f(a_j) \big) = 0 . \]
Next, \(P\) contains a nonmeasurable set: by the axiom of choice pick \(N \subset P\) meeting each class of the relation \(x-y\in\mathbb{Q}\) exactly once. The sets \(N + q\), \(q \in \mathbb{Q}\), are pairwise disjoint and \(P \subset \bigcup_{q \in \mathbb{Q}} (N+q)\). If \(N\) were measurable, then \(\lambda(N)=0\) would give \(\lambda(P)=0\), a contradiction, while \(\lambda(N) > 0\) would give \[ \lambda\Big( \bigcup_{q \in \mathbb{Q} \cap [0,1]} (N+q) \Big) = \sum_{q \in \mathbb{Q}\cap[0,1]} \lambda(N) = \infty , \] although this union lies in \([0,2]\). Finally \(Z := f(N) \subset f(P)\) is null, hence Lebesgue measurable by completeness, while injectivity of \(f\) gives \(f^{-1}(Z) = N\), which is not measurable.
(Kaufman, Rickert [497]) (i) Let \(\mu\) be a complex measure with \(\|\mu\| = 1\) (see the definition before Proposition 3.10.16). Prove that there exists a measurable set \(E\) such that \(|\mu(E)| \ge 1/\pi\).
(ii) Prove that in (i) one can pick a set \(E\) with \(|\mu(E)| > 1/\pi\) precisely when the Radon-Nikodym density \(f\) of the measure \(\mu\) with respect to \(|\mu|\) satisfies the equality \[ \int f(t)^k \, |\mu|(dt) = 0 \] for all \(k \in \{-1, 1, -2n, 2n\}\), \(n \in \mathbb{N}\).
(iii) Let \(\mu\) be a measure with values in \(\mathbb{R}^n\) such that \(\|\mu\| = 1\). Prove that there exists a measurable set \(E\) such that \[ |\mu(E)| \ge \Gamma(n/2)\Big[ 2\sqrt{\pi}\,\Gamma\big((n+1)/2\big) \Big]^{-1} . \]
All three parts average the halfspace sets \(E_y=\{(f,y)>0\}\) over directions, where \(\mu(A)=\int_Af\,d|\mu|\) is the polar decomposition: each component of \(\mu\) is a bounded signed measure with \(\ll|\mu|\), so the Radon-Nikodym theorem supplies \(f\in L^1(|\mu|)\) with values in \(\mathbb{C}\) (resp. \(\mathbb{R}^n\)), and \(|f|=1\) \(|\mu|\)-a.e. because summing \(|\mu(A_j)|\le\int_{A_j}|f|\,d|\mu|\) over disjoint \(A_j\subset A\) gives \(|\mu|(A)\le\int_A|f|\,d|\mu|\), while partitioning \(A\) into sets where \(f\) is nearly constant gives \(\int_A|f|\,d|\mu|\le(1+\varepsilon)|\mu|(A)\).
(i) Write \(f = e^{i\varphi}\), \(\varphi:=\arg f\in[0,2\pi)\), and for \(\theta\in\mathbb{R}\) put \[ E_\theta := \{ t : \cos(\theta - \varphi(t)) > 0 \} , \] \[ h(\theta) := \int_X \big( \cos(\theta - \varphi(t)) \big)^+ \, |\mu|(dt) = \operatorname{Re}\Big( e^{-i\theta}\mu(E_\theta) \Big) , \] the second equality because \(\cos(\theta-\varphi)\) is its own positive part on \(E_\theta\) and \(\le0\) off it. As \(|\mu|\) is finite and the integrand bounded, Fubini’s theorem gives \[ \frac{1}{2\pi}\int_0^{2\pi} h(\theta)\,d\theta = \int_X \Big( \frac{1}{2\pi}\int_0^{2\pi} \big(\cos(\theta-\varphi(t))\big)^+ d\theta \Big) |\mu|(dt) = \frac{1}{\pi}\,|\mu|(X) = \frac{1}{\pi}, \] because \(\frac{1}{2\pi}\int_0^{2\pi}(\cos\theta)^+ d\theta = \frac{1}{2\pi}\int_{-\pi/2}^{\pi/2}\cos\theta\,d\theta = \frac{1}{\pi}\) and the inner integral does not depend on \(\varphi(t)\) by periodicity. Hence \(h(\theta_0) \ge 1/\pi\) for some \(\theta_0\), and then \[ |\mu(E_{\theta_0})| \ge \operatorname{Re}\big( e^{-i\theta_0}\mu(E_{\theta_0}) \big) = h(\theta_0) \ge \frac{1}{\pi} . \]
(ii) The printed statement asserts the opposite implication, a slip; we prove that a set with \(|\mu(E)|>1/\pi\) exists precisely when at least one of those integrals is nonzero. First, \[ \sup_{E} |\mu(E)| = \max_{\theta} h(\theta) , \tag{1} \] since for any \(E\) and \(\theta\), \[ \operatorname{Re}\big( e^{-i\theta}\mu(E) \big) = \int_E \cos(\theta - \varphi)\,d|\mu| \le \int_X \big( \cos(\theta-\varphi) \big)^+ d|\mu| = h(\theta), \] and \(\theta\) with \(e^{-i\theta}\mu(E) = |\mu(E)|\) gives \(|\mu(E)| \le \sup_\theta h\), while \(h(\theta) \le |\mu(E_\theta)|\) as in (i); \(h\) is Lipschitz and \(2\pi\)-periodic, so the supremum is attained. Since \(\frac{1}{2\pi}\int_0^{2\pi}h = 1/\pi\) by (i), \[ \text{there exists } E \text{ with } |\mu(E)| > 1/\pi \iff h \not\equiv 1/\pi . \tag{2} \] For the Fourier coefficients of \(h\) put \(u(\theta) := (\cos\theta)^+\), so that \(h(\theta) = \int_X u(\theta - \varphi(t))\,|\mu|(dt)\). By Fubini’s theorem, \[ \widehat h(k) = \frac{1}{2\pi}\int_0^{2\pi} h(\theta)e^{-ik\theta}\,d\theta = \int_X \Big( \frac{1}{2\pi}\int_0^{2\pi} u(\theta-\varphi(t))e^{-ik\theta}d\theta \Big)|\mu|(dt) = \widehat u(k) \int_X e^{-ik\varphi(t)}\,|\mu|(dt), \] that is, since \(e^{-ik\varphi} = f^{-k}\), \[ \widehat h(k) = \widehat u(k) \int_X f(t)^{-k}\,|\mu|(dt) . \tag{3} \] For \(k \ne \pm 1\), \[ \widehat u(k) = \frac{1}{2\pi}\int_{-\pi/2}^{\pi/2}\cos\theta\, \cos(k\theta)\,d\theta = \frac{1}{4\pi}\Big[ \frac{2\sin\big((k+1)\pi/2\big)}{k+1} + \frac{2\sin\big((k-1)\pi/2\big)}{k-1} \Big], \] whence \(\widehat u(0) = 1/\pi\), \(\widehat u(k) = 0\) for odd \(k\) with \(|k| \ge 3\) (both sines vanish), and for \(k = 2m\) with \(m \ne 0\) \[ \widehat u(2m) = \frac{(-1)^m}{2\pi}\Big( \frac{1}{2m+1} - \frac{1}{2m-1} \Big) = \frac{(-1)^{m+1}}{\pi(4m^2-1)} \ne 0 ; \] finally \(\widehat u(\pm 1) = \frac{1}{2\pi}\int_{-\pi/2}^{\pi/2}\cos^2\theta\,d\theta = 1/4 \ne 0\). Therefore \[ \{ k \ne 0 : \widehat u(k) \ne 0 \} = \{-1,\,1\} \cup \{ -2n,\, 2n : n \in \mathbb{N} \} , \] exactly the index set of the formulation. A continuous function on the circle is constant iff its coefficients with \(k \ne 0\) vanish, so by (3), the index set being symmetric under \(k \mapsto -k\), \[ h \equiv \tfrac{1}{\pi} \iff \int_X f(t)^k\,|\mu|(dt) = 0 \ \text{ for all } k \in \{-1,1,-2n,2n\},\ n \in \mathbb{N}, \tag{4} \] which with (2) is the assertion.
(iii) Replace the circle by the sphere: let \(\sigma\) be the normalized surface measure on \(S^{n-1}\) and \(E_y := \{ t : (f(t),y) > 0 \}\) for \(y \in S^{n-1}\), so that \[ \big( \mu(E_y), y \big) = \int_{E_y} (f,y)\,d|\mu| = \int_X (f,y)^+ \,d|\mu| . \] Averaging over \(y\), Fubini’s theorem gives \[ \int_{S^{n-1}} \Big( \int_X (f(t),y)^+ |\mu|(dt) \Big)\sigma(dy) = \int_X \Big( \int_{S^{n-1}} (f(t),y)^+\sigma(dy) \Big)|\mu|(dt) = c_n\,|\mu|(X) = c_n , \] where, by the rotation invariance of \(\sigma\), the inner integral equals the constant \[ c_n = \int_{S^{n-1}} (y_1)^+\,\sigma(dy) = \frac{1}{2}\int_{S^{n-1}}|y_1|\,\sigma(dy) \] for every unit vector \(f(t)\). Hence there is \(y_0 \in S^{n-1}\) with \(\int_X (f,y_0)^+ d|\mu| \ge c_n\), and then \[ |\mu(E_{y_0})| \ge \big( \mu(E_{y_0}), y_0 \big) \ge c_n . \]
To compute \(c_n\), let \(s_{n-1} = 2\pi^{n/2}/\Gamma(n/2)\) be the total surface measure of \(S^{n-1}\) and \(\omega_{n-1} = \pi^{(n-1)/2}/\Gamma((n+1)/2)\) the volume of the unit ball of \(\mathbb{R}^{n-1}\); the projection \(y \mapsto (y_2,\dots,y_n)\) maps \(\{y_1>0\}\) bijectively onto that ball with surface-measure factor \(|y_1|^{-1}\), so \[ \int_{\{y_1>0\}} |y_1| \, dS(y) = \omega_{n-1}, \qquad \int_{S^{n-1}}|y_1|\,dS(y) = 2\omega_{n-1} , \] where \(dS\) is the unnormalized surface measure. Therefore \[ c_n = \frac{1}{2}\cdot\frac{2\omega_{n-1}}{s_{n-1}} = \frac{\pi^{(n-1)/2}/\Gamma\big((n+1)/2\big)}{2\pi^{n/2}/\Gamma(n/2)} = \frac{\Gamma(n/2)}{2\sqrt{\pi}\,\Gamma\big((n+1)/2\big)} , \] which is the asserted bound.
(i) Suppose that the values of two Borel probability measures \(\mu\) and \(\nu\) on \(\mathbb{R}^n\) coincide on every half-space of the form \(\{x : (x,y) \le c\}\), \(y \in \mathbb{R}^n\), \(c \in \mathbb{R}^1\). Prove that \(\mu = \nu\). Prove the same for open half-spaces.
(ii) (Ptak, Tkadlec [771]) Suppose that the values of two Borel probability measures \(\mu\) and \(\nu\) on \(\mathbb{R}^n\) coincide on every open ball with the origin at the boundary. Prove that \(\mu = \nu\).
(iii) Prove the analog of (ii) for closed balls.
(i) Fix \(y\) and let \(\pi_y(x) := (x,y)\), a continuous map \(\mathbb{R}^n\to\mathbb{R}^1\). For every \(c\), \[ \mu\circ\pi_y^{-1}\big( (-\infty,c] \big) = \mu\big( \{x : (x,y) \le c\} \big) = \nu\big( \{x : (x,y) \le c\} \big) = \nu\circ\pi_y^{-1}\big( (-\infty,c] \big) . \] The rays \((-\infty,c]\) are closed under finite intersections and generate \(\mathcal{B}(\mathbb{R}^1)\), so \(\mu\circ\pi_y^{-1} = \nu\circ\pi_y^{-1}\) by uniqueness of extension (Theorem 1.5.6). Hence, by the change of variables formula (Theorem 3.6.1) for the bounded continuous \(t \mapsto \exp(it)\), \[ \widehat{\mu}(y) = \int_{\mathbb{R}^n} \exp\big( i(x,y) \big)\,\mu(dx) = \int_{\mathbb{R}^1} \exp(it)\, \mu\circ\pi_y^{-1}(dt) = \int_{\mathbb{R}^1} \exp(it)\, \nu\circ\pi_y^{-1}(dt) = \widehat{\nu}(y) \] for every \(y\), so \(\mu = \nu\) by Proposition 3.8.6. For open half-spaces, \(\{(x,y) \le c\} = \bigcap_k \{(x,y) < c + 1/k\}\) reduces the hypothesis to the closed case.
(ii) Let \[ F(x) = \frac{x}{|x|^2} \ \text{ for } x \ne 0, \qquad F(0) = 0 , \] a Borel involution of \(\mathbb{R}^n\), so \(F^{-1}=F\). An open ball with the origin on the boundary is \(U = \{x : |x - a| < |a|\}\) with \(a \ne 0\), and \[ x \in U \iff |x|^2 - 2(x,a) < 0 \iff |x|^2 < 2(x,a). \] Every \(x\in U\) is nonzero, so dividing by \(|x|^2\), \[ x \in U \iff (F(x), a) > \tfrac{1}{2} . \] Thus \(F^{-1}(H_a) = U\) with \(H_a := \{ z : (z,a) > 1/2 \}\), using also \(F(0)=0\notin H_a\), so with \(\mu^{\prime} := \mu \circ F^{-1}\), \(\nu^{\prime} := \nu \circ F^{-1}\), \[ \mu^{\prime}(H_a) = \mu(U) = \nu(U) = \nu^{\prime}(H_a) \qquad \text{for every } a \ne 0 . \tag{1} \] As \(a\) runs over \(\mathbb{R}^n\setminus\{0\}\), the \(H_a\) are exactly the open half-spaces \(\{(z,b)>c\}\) with \(b\ne0\), \(c>0\) (take \(a=b/(2c)\)). The other open half-spaces follow: for \(d>0\), \(\{(z,b)\ge d\}=\bigcap_{k\ge2}\{(z,b)>d-d/k\}\) is a decreasing intersection of such sets, so \(\mu^{\prime},\nu^{\prime}\) agree on it and on its complement \(\{(z,-b)>-d\}\), i.e. on every open half-space with negative constant; and \(\{(z,b)>0\}=\bigcup_k\{(z,b)>1/k\}\). By (i), \(\mu^{\prime}=\nu^{\prime}\), and applying \(F\) once more (an involution) gives \(\mu=\nu\).
(iii) First \(\mu(\{0\}) = \nu(\{0\})\): fix a unit vector \(e\) and consider the closed balls \[ \overline{U}_t := \{ x : |x - te| \le t \} = \{ x : |x|^2 \le 2t(x,e) \}, \qquad t > 0 , \] each with the origin on its boundary. For \(0<s<t\) and \(x\in\overline U_s\) one has \((x,e)\ge|x|^2/(2s)\ge0\), so \(|x|^2\le2t(x,e)\) and \(x\in\overline U_t\); thus \(\overline U_t\) decreases as \(t\downarrow0\), and \(\bigcap_{t>0}\overline U_t=\{0\}\) since \(|x|^2\le2t|x|\) for all \(t>0\) forces \(x=0\). Continuity from above gives \[ \mu(\{0\}) = \lim_{k \to \infty}\mu\big( \overline U_{1/k} \big) = \lim_{k \to \infty}\nu\big( \overline U_{1/k} \big) = \nu(\{0\}) . \]
Next, for \(\overline U = \{ x : |x-a| \le |a| \}\), \(a \ne 0\), exactly as in (ii), for \(x \ne 0\), \[ x \in \overline U \iff |x|^2 \le 2(x,a) \iff \big( F(x), a \big) \ge \tfrac{1}{2}, \] while \(0 \in \overline U\) and \(F(0) = 0 \notin \overline H_a := \{z : (z,a) \ge 1/2\}\). Hence \(F^{-1}(\overline H_a) = \overline U \setminus \{0\}\) and \[ \mu^{\prime}(\overline H_a) = \mu(\overline U) - \mu(\{0\}) = \nu(\overline U) - \nu(\{0\}) = \nu^{\prime}(\overline H_a) \] for every \(a \ne 0\), i.e. on all \(\{(z,b) \ge d\}\) with \(b \ne 0\), \(d > 0\). Complements give \(\{(z,b)<d\}\), hence all open half-spaces with negative constant, while \(\{(z,b)>c\}=\bigcup_k\{(z,b)\ge c+1/k\}\) for \(c\ge0\). So (i) yields \(\mu^{\prime} = \nu^{\prime}\) and \(\mu = \nu\) as in (ii).
(\(\circ\)) Let a function \(\Phi\) be strictly increasing and continuous on \([0,1]\). Prove that for every bounded Borel function \(f\) one has \[ \int_0^1 f(x)\,d\Phi(x) = \int_{\Phi(0)}^{\Phi(1)} f\big( \Phi^{-1}(y) \big)\,dy \] with the Lebesgue-Stieltjes integral on the left and the Lebesgue integral on the right.
The formula is the change of variables for the image measure identity \[ \mu_\Phi = \lambda_J \circ (\Phi^{-1})^{-1} , \tag{1} \] where \(\mu_\Phi\) is the Lebesgue-Stieltjes measure of \(\Phi\), determined by \(\mu_\Phi((a,b])=\Phi(b)-\Phi(a)\), and \(\lambda_J\) is Lebesgue measure on \(J := [\Phi(0),\Phi(1)]\); \(\Phi\) being a homeomorphism of \([0,1]\) onto \(J\), the map \(\Phi^{-1}\) is continuous and increasing, so \(f\circ\Phi^{-1}\) is a bounded Borel function. For (1): \(\Phi^{-1}(y)\in(a,b]\) iff \(y\in(\Phi(a),\Phi(b)]\), so both sides give \((a,b]\) the value \(\Phi(b)-\Phi(a)\), and both give \(\{0\}\) the value \(0\) (continuity of \(\Phi\) makes \(\mu_\Phi(\{0\})=\lim_{b\to0+}(\Phi(b)-\Phi(0))=0\)); these sets generate the algebra of their finite disjoint unions, so Theorem 1.5.6(iii) applies. Now Theorem 3.6.1 applied to \(\Phi^{-1}\), \(\lambda_J\) and the bounded (hence \(\mu_\Phi\)-integrable) \(f\) gives \[ \int_{[0,1]} f\,d\mu_\Phi = \int_J f\big( \Phi^{-1}(y) \big)\,dy . \]
Let \(\mu\) be a Borel (possibly signed) measure on \([0,1]\) with the following property: if continuous functions \(f_n\) are uniformly bounded and converge to zero almost everywhere with respect to Lebesgue measure \(\lambda\), then \[ \int f_n \, d\mu \to 0 . \] Prove that \(\mu \ll \lambda\).
The test functions are \(f_n(x) := \max\{0, 1 - n\,\mathrm{dist}(x,K)\}\) for compact \(K\subset[0,1]\) with \(\lambda(K)=0\): they are continuous, \(0\le f_n\le1\), equal \(1\) on \(K\), and vanish eventually off the closed set \(K\), so \(f_n\to I_K\) pointwise, hence \(f_n\to0\) \(\lambda\)-a.e. The hypothesis gives \(\int f_n\,d\mu\to0\), while dominated convergence against the finite measure \(|\mu|\) gives \(\int f_n\,d\mu\to\mu(K)\). Thus \[ \mu(K) = 0 \quad \text{for every compact } K \subset [0,1] \text{ with } \lambda(K) = 0. \tag{1} \]
Now let \(B\) be Borel with \(\lambda(B)=0\); we show \(|\mu|(B)=0\). Take a Hahn decomposition \([0,1]=P\cup N\), so \(\mu^+(A)=\mu(A\cap P)\) and \(\mu^-(A)=-\mu(A\cap N)\). By inner regularity of the finite measure \(|\mu|\) (Theorem 1.4.8) there are compact \(K_j \subset B \cap P\) with \[ |\mu|\big( (B\cap P) \setminus K_j \big) \to 0 . \] Each \(K_j\) satisfies \(\lambda(K_j) \le \lambda(B) = 0\), so \(\mu(K_j) = 0\) by (1). Consequently \[ \big| \mu^+(B) \big| = \big| \mu(B\cap P) \big| = \big| \mu(B \cap P) - \mu(K_j) \big| = \Big| \int_{(B\cap P)\setminus K_j} d\mu \Big| \le |\mu|\big( (B\cap P)\setminus K_j \big) \to 0, \] whence \(\mu^+(B) = 0\), and likewise \(\mu^-(B)=0\) from \(B\cap N\). So \(|\mu|(B)=0\), i.e. \(|\mu|\ll\lambda\) and a fortiori \(\mu\ll\lambda\).
Exercises 3.10.113–3.10.119
(i) Let \((X,\mathcal{A},\mu)\) and \((Y,\mathcal{B},\nu)\) be complete probability spaces, let \(A\subset X\) be a set that is not measurable with respect to \(\mu\), and let \(B\subset Y\) be a set such that \(A\times B\) is measurable with respect to \(\mu\otimes\nu\). Prove that \(\nu(B)=0\).
(ii) Let \((X_n,\mathcal{A}_n,\mu_n)\), where \(n\in\mathbb{N}\), be complete probability spaces and let sets \(A_n\subset X_n\) be such that \(\prod_{n=1}^{\infty}A_n\) is measurable with respect to \(\bigotimes_{n=1}^{\infty}\mu_n\). Prove that either every \(A_n\) is measurable with respect to \(\mu_n\), or \(\mu\bigl(\prod_{n=1}^{\infty}A_n\bigr)=0\) and then
\begin{equation*} \lim_{n\to\infty}\prod_{i=1}^{n}\mu_i^{*}(A_i)=0 . \end{equation*}
(i) The sections are \((A\times B)^{y}=A\) for \(y\in B\) and \(=\emptyset\) otherwise, so \(B\) is exactly the set of \(y\) with non-measurable section, \(A\) being non-measurable and \(\emptyset\) measurable. Theorem 3.4.1 applied to \(A\times B\in(\mathcal{A}\otimes\mathcal{B})_{\mu\otimes\nu}\) puts that set inside some \(N\in\mathcal{B}\) with \(\nu(N)=0\), so \(\nu^{*}(B)=0\) and \(B\) is \(\nu\)-measurable with \(\nu(B)=0\), \(\nu\) being complete.
(ii) The device, also used in Exercise 3.10.114, is:
Lemma. Let \(N\) be a nonempty at most countable index set, let \((Y_n,\mathcal{B}_n,\rho_n)\), \(n\in N\), be complete probability spaces, let \(B_n\subset Y_n\), and put
\begin{equation*} \rho:=\bigotimes_{n\in N}\rho_n,\qquad Q:=\prod_{n\in N}B_n . \end{equation*}
If \(\rho^{*}(Q)=0\), then \(\prod_{n\in N}\rho_n^{*}(B_n)=0\), i.e. the partial products of the numbers \(\rho_n^{*}(B_n)\) (taken along any enumeration of \(N\)) tend to zero.
Proof. Put \(\mathcal{B}^{\prime}_n:=\sigma(\mathcal{B}_n\cup\{B_n\})\) and \(\sigma_n:=\rho_n\) when \(B_n\in\mathcal{B}_n\). Otherwise \(\rho_{n*}(B_n)<\rho_n^{*}(B_n)\) (equality would force measurability, \(\rho_n\) being complete and finite), so Theorem 1.12.14 with \(\gamma:=\rho_n^{*}(B_n)\) gives a probability measure \(\sigma_n\) on \(\mathcal{B}^{\prime}_n\) with
\begin{equation*} \sigma_n|_{\mathcal{B}_n}=\rho_n,\qquad \sigma_n(B_n)=\rho_n^{*}(B_n). \end{equation*}
Let \(\sigma:=\bigotimes_{n\in N}\sigma_n\). Fix an enumeration \(N=\{n_1,n_2,\dots\}\) (in the finite case the lists terminate and limits are attained) and put
\begin{equation*} Q_k:=\bigcap_{j\le k}\mathrm{pr}_{n_j}^{-1}(B_{n_j}), \end{equation*}
a cylinder with base in \(\mathcal{B}^{\prime}_{n_1}\otimes\dots\otimes\mathcal{B}^{\prime}_{n_k}\), so that
\begin{equation*} \sigma(Q_k)=\prod_{j\le k}\sigma_{n_j}(B_{n_j})=\prod_{j\le k}\rho_{n_j}^{*}(B_{n_j}). \end{equation*}
The \(Q_k\) decrease to \(Q\), so \(Q\in\bigotimes_{n}\mathcal{B}^{\prime}_n\) and
\begin{equation*} \sigma(Q)=\lim_{k\to\infty}\prod_{j\le k}\rho_{n_j}^{*}(B_{n_j}). \end{equation*}
Since the factors agree on the \(\mathcal{B}_n\), the measures \(\sigma\) and \(\rho\) agree on the algebra of cylinders with bases in \(\mathcal{B}_{n_1}\otimes\dots\otimes\mathcal{B}_{n_k}\), hence on \(\bigotimes_{n}\mathcal{B}_n\) by Theorem 1.5.6(iii). By \(\rho^{*}(Q)=0\) and Corollary 1.5.8 there is \(F\in\bigotimes_n\mathcal{B}_n\) with \(Q\subset F\), \(\rho(F)=0\), so \(\sigma(Q)\le\sigma(F)=\rho(F)=0\).
Throughout, products of probability measures are the Lebesgue completions of the measures on the product \(\sigma\)-algebra (Section 3.5), so they are complete and split as \(\bigotimes_{\alpha}\mu_\alpha=(\bigotimes_{\Lambda_1}\mu_\alpha)\otimes(\bigotimes_{\Lambda_2}\mu_\alpha)\); in particular (i) applies to any such splitting. Now put \(\mu:=\bigotimes_{n=1}^{\infty}\mu_n\) and \(P:=\prod_{n=1}^{\infty}A_n\), and suppose that not every \(A_n\) is \(\mu_n\)-measurable, say \(A_{n_0}\) is not. Split the index set into \(\{n_0\}\) and its complement: writing
\begin{equation*} \widetilde X:=\prod_{n\neq n_0}X_n,\qquad \widetilde\mu:=\bigotimes_{n\neq n_0}\mu_n,\qquad \widetilde P:=\prod_{n\neq n_0}A_n, \end{equation*}
we have, by the splitting property of infinite products noted after Lemma 3.5.2, \(\mu=\mu_{n_0}\otimes\widetilde\mu\), and under the corresponding identification \(P=A_{n_0}\times\widetilde P\). Both \(\mu_{n_0}\) and \(\widetilde\mu\) are complete probability measures, the former by hypothesis and the latter because products of measures are defined as Lebesgue completions. Since \(P\) is \(\mu\)-measurable and \(A_{n_0}\) is not \(\mu_{n_0}\)-measurable, part (i) applies and yields that \(\widetilde P\) is \(\widetilde\mu\)-measurable with \(\widetilde\mu(\widetilde P)=0\).
Hence \(X_{n_0}\times\widetilde P\) is \(\mu\)-measurable with
\begin{equation*} \mu\bigl(X_{n_0}\times\widetilde P\bigr)=\mu_{n_0}(X_{n_0})\,\widetilde\mu(\widetilde P)=0 , \end{equation*}
and \(P\subset X_{n_0}\times\widetilde P\), so \(\mu^{*}(P)=0\); since \(P\) is measurable, \(\mu(P)=0\).
The Lemma with \(N=\mathbb{N}\), \(\rho_n=\mu_n\), \(B_n=A_n\) then turns \(\mu^{*}(P)=0\) into
\begin{equation*} \lim_{n\to\infty}\prod_{i=1}^{n}\mu_i^{*}(A_i)=0 . \end{equation*}
Let \((X_\alpha,\mathcal{A}_\alpha,\mu_\alpha)\), where \(\alpha\in\Lambda\) and \(\Lambda\neq\emptyset\), be measurable spaces with complete probability measures and let \(E_\alpha\subset X_\alpha\) be such that \(E=\prod_{\alpha\in\Lambda}E_\alpha\) is measurable with respect to \(\bigotimes_{\alpha}\mu_\alpha\), but does not belong to \(\bigotimes_{\alpha}\mathcal{A}_\alpha\). Prove that \(\prod_{\alpha\in\Lambda}\mu^{*}_\alpha(E_\alpha)=0\), i.e., there exists an at most countable family of indices \(\alpha_n\) such that the product of numbers \(\mu^{*}_{\alpha_n}(E_{\alpha_n})\) diverges to zero.
Write \(p_\alpha:=\mu_\alpha^{*}(E_\alpha)\), \(\mu:=\bigotimes_{\alpha\in\Lambda}\mu_\alpha\). We use two facts from Section 3.5 – every set of \(\bigotimes_{\alpha}\mathcal{A}_\alpha\) is a cylinder over an at most countable set of coordinates (Lemma 3.5.2), and products split as \(\bigotimes_{\Lambda}\mu_\alpha=(\bigotimes_{\Lambda^{\prime}}\mu_\alpha)\otimes(\bigotimes_{\Lambda^{\prime\prime}}\mu_\alpha)\) into complete factors – together with
Lemma L (from Exercise 3.10.113(ii)). If \(N\) is a nonempty at most countable set, \((Y_n,\mathcal{B}_n,\rho_n)\) are complete probability spaces, \(B_n\subset Y_n\), \(\rho=\bigotimes_{n\in N}\rho_n\), \(Q=\prod_{n\in N}B_n\) and \(\rho^{*}(Q)=0\), then \(\prod_{n\in N}\rho_n^{*}(B_n)=0\).
First we may assume \(E_\alpha\ne X_\alpha\) for all \(\alpha\). Let \(\Lambda_0:=\{\alpha:\ E_\alpha=X_\alpha\}\) and \(\Lambda^{\prime}:=\Lambda\setminus\Lambda_0\). If \(\Lambda^{\prime}=\emptyset\), then \(E=\prod_\alpha X_\alpha\) belongs to \(\bigotimes_\alpha\mathcal{A}_\alpha\), contrary to the hypothesis; so \(\Lambda^{\prime}\neq\emptyset\). Assume \(\Lambda_0\neq\emptyset\) (otherwise there is nothing to do) and write
\begin{equation*} \pi^{\prime}:=\bigotimes_{\alpha\in\Lambda^{\prime}}\mu_\alpha \ \text{ on } X^{\prime}:=\prod_{\alpha\in\Lambda^{\prime}}X_\alpha,\qquad \pi_0:=\bigotimes_{\alpha\in\Lambda_0}\mu_\alpha \ \text{ on } X_0:=\prod_{\alpha\in\Lambda_0}X_\alpha , \end{equation*}
so that \(\mu=\pi^{\prime}\otimes\pi_0\) and \(E=E^{\prime}\times X_0\) with \(E^{\prime}:=\prod_{\alpha\in\Lambda^{\prime}}E_\alpha\). Every section \(E^{y}\), \(y\in X_0\), equals \(E^{\prime}\), and Theorem 3.4.1 makes \(E^y\) \(\pi^{\prime}\)-measurable for \(\pi_0\)-a.e. \(y\), so \(E^{\prime}\) is \(\pi^{\prime}\)-measurable; also \(E^{\prime}\notin\bigotimes_{\Lambda^{\prime}}\mathcal{A}_\alpha\), else \(E\) would be a cylinder with measurable base. As \(p_\alpha=1\) on \(\Lambda_0\), replacing \(\Lambda\) by \(\Lambda^{\prime}\) changes nothing.
(i) \(\Lambda\) at most countable. A singleton \(\Lambda\) is vacuous (completeness would put \(E\) in \(\mathcal{A}_{\alpha_0}\)), so let \(2\le\operatorname{card}\Lambda\le\aleph_0\). If every \(E_\alpha\) belonged to \(\mathcal{A}_\alpha\), then \(E=\bigcap_{\alpha\in\Lambda}\mathrm{pr}_\alpha^{-1}(E_\alpha)\) would be an at most countable intersection of cylinders with measurable bases, hence an element of \(\bigotimes_\alpha\mathcal{A}_\alpha\), contrary to the hypothesis. So there is \(\alpha_0\) with \(E_{\alpha_0}\notin\mathcal{A}_{\alpha_0}\), and by the completeness of \(\mu_{\alpha_0}\) the set \(E_{\alpha_0}\) is not \(\mu_{\alpha_0}\)-measurable. Splitting off the index \(\alpha_0\) and applying Exercise 3.10.113(i) exactly as in the proof of its part (ii), we get that \(\widetilde E:=\prod_{\alpha\neq\alpha_0}E_\alpha\) satisfies \(\widetilde\pi(\widetilde E)=0\), where \(\widetilde\pi:=\bigotimes_{\alpha\neq\alpha_0}\mu_\alpha\). Lemma L now gives \(\prod_{\alpha\neq\alpha_0}p_\alpha=0\), whence \(\prod_{\alpha\in\Lambda}p_\alpha=0\).
From now on \(\Lambda\) is uncountable. Put
\begin{equation*} \Lambda_1:=\{\alpha\in\Lambda:\ p_\alpha=1\},\qquad \Lambda_2:=\Lambda\setminus\Lambda_1 , \end{equation*}
\begin{equation*} \Pi_i:=\prod_{\alpha\in\Lambda_i}E_\alpha,\qquad \pi_i:=\bigotimes_{\alpha\in\Lambda_i}\mu_\alpha\quad (i=1,2), \end{equation*}
so that \(E=\Pi_1\times\Pi_2\) and \(\mu=\pi_1\otimes\pi_2\) (with the obvious conventions if one index set is empty).
(ii) \(\Lambda_2\) uncountable. Since \(\Lambda_2=\bigcup_{k\ge1}\{p_\alpha\le1-1/k\}\), one of these sets is infinite, giving \(q<1\) and distinct \(\alpha_n\) with \(p_{\alpha_n}\le q\), so \(\prod_{n\le N}p_{\alpha_n}\le q^N\to0\).
(iii) \(\Lambda_2\) at most countable, so \(\Lambda_1\) is uncountable with \(p_\alpha=1\), in particular \(E_\alpha\ne\emptyset\), on \(\Lambda_1\). First, \(\pi_1^{*}(\Pi_1)>0\): were it \(0\), Corollary 1.5.8 would give \(F\in\bigotimes_{\alpha\in\Lambda_1}\mathcal{A}_\alpha\) with \(\Pi_1\subset F\) and \(\pi_1(F)=0\). By Lemma 3.5.2 there is an at most countable set \(S\subset\Lambda_1\) (which we may enlarge so as to be nonempty) with
\begin{equation*} F=F_0\times\prod_{\alpha\in\Lambda_1\setminus S}X_\alpha,\qquad F_0\in\bigotimes_{\alpha\in S}\mathcal{A}_\alpha , \end{equation*}
and then \(\rho(F_0)=\pi_1(F)=0\) for \(\rho:=\bigotimes_{\alpha\in S}\mu_\alpha\) by splitting. As \(E_\alpha\ne\emptyset\) on \(\Lambda_1\), every point of \(\prod_{\alpha\in S}E_\alpha\) extends to a point of \(\Pi_1\subset F\), so \(\prod_{\alpha\in S}E_\alpha\subset F_0\) and \(\rho^{*}(\prod_{\alpha\in S}E_\alpha)=0\); Lemma L then forces \(\prod_{\alpha\in S}p_\alpha=0\), whereas that product is \(1\).
Next, \(\Pi_1\) contains no nonempty \(G\in\bigotimes_{\alpha\in\Lambda_1}\mathcal{A}_\alpha\): such a \(G\) depends only on coordinates in some at most countable \(S\subset\Lambda_1\) (Lemma 3.5.2), and \(\Lambda_1\) being uncountable there is \(\beta\in\Lambda_1\setminus S\) with some \(z\in X_\beta\setminus E_\beta\); changing the \(\beta\)-th coordinate of a \(g\in G\) to \(z\) keeps the point in \(G\subset\Pi_1\) yet puts it outside \(\Pi_1\).
Hence \(\Pi_1\) is not \(\pi_1\)-measurable: otherwise \(\pi_1(\Pi_1)>0\) and Corollary 1.5.8 would supply a nonempty \(G\in\bigotimes_{\alpha\in\Lambda_1}\mathcal{A}_\alpha\) inside \(\Pi_1\). In particular \(\Lambda_2\ne\emptyset\), since \(\Lambda_2=\emptyset\) would make \(E=\Pi_1\) non-\(\mu\)-measurable. So Exercise 3.10.113(i) applies to \(E=\Pi_1\times\Pi_2\) and \(\mu=\pi_1\otimes\pi_2\) and gives
\begin{equation*} \pi_2(\Pi_2)=0,\qquad\text{in particular}\quad \pi_2^{*}(\Pi_2)=0 . \end{equation*}
Applying Lemma L to the at most countable family \(\{(X_\alpha,\mathcal{A}_\alpha,\mu_\alpha)\}_{\alpha\in\Lambda_2}\) and the sets \(E_\alpha\), we conclude that \(\prod_{\alpha\in\Lambda_2}p_\alpha=0\). Since \(p_\alpha=1\) for \(\alpha\in\Lambda_1\), this means exactly that
\begin{equation*} \prod_{\alpha\in\Lambda}\mu_\alpha^{*}(E_\alpha)=0 , \end{equation*}
witnessed by an at most countable set of indices in \(\Lambda_2\).
(\(\circ\)) Let \(\mu\) be a Borel probability measure with a density \(\varrho\) on \(\mathbb{R}^2\).
(i) Show that the distribution of \(f(x,y)=x+y\) on \((\mathbb{R}^2,\mu)\) has the density
\begin{equation*} \varrho_1(t)=\int_{-\infty}^{+\infty}\varrho(t-s,s)\,ds . \end{equation*}
(ii) Show that the distribution of \(g(x,y)=x/y\) on \((\mathbb{R}^2,\mu)\) has the density
\begin{equation*} \varrho_2(t)=\int_{-\infty}^{+\infty}|s|\varrho(ts,s)\,ds . \end{equation*}
Both densities come from Theorem 3.7.1, \(\int_U h(F(z))|J_F(z)|\,dz=\int_{F(U)}h(u)\,du\) for injective \(C^1\) maps \(F\) on open \(U\) and nonnegative Borel \(h\) (monotone convergence on \(\min(h,n)\mathbb{1}_{[-n,n]^2}\) extends it from integrable \(h\)). Take \(\varrho\) Borel, which alters neither \(\mu\) nor the integrals below.
(i) Let \(T(t,s):=(t-s,s)\), a linear bijection of \(\mathbb{R}^2\) with \(|J_T|\equiv1\), and fix Borel \(B\subset\mathbb{R}\). For \(h(x,y)=\mathbb{1}_B(x+y)\varrho(x,y)\),
\begin{equation*} h(T(t,s))=\mathbb{1}_B\bigl((t-s)+s\bigr)\varrho(t-s,s)=\mathbb{1}_B(t)\varrho(t-s,s), \end{equation*}
so that
\begin{equation*} \mu\circ f^{-1}(B)=\int_{\mathbb{R}^2}\mathbb{1}_B(x+y)\varrho(x,y)\,dx\,dy=\int_{\mathbb{R}^2}\mathbb{1}_B(t)\varrho(t-s,s)\,dt\,ds , \end{equation*}
and Fubini’s theorem for nonnegative functions (Theorems 3.4.4, 3.4.5) makes \(\varrho_1(t)=\int\varrho(t-s,s)\,ds\) Borel with \(\mu\circ f^{-1}(B)=\int_B\varrho_1\,dt\); \(B=\mathbb{R}\) gives \(\int\varrho_1=1\).
(ii) Here \(g\) is defined off the null line \(\{y=0\}\), so \(\mu\)-a.e. Put \(V:=\{y\ne0\}\), \(U:=\{s\ne0\}\) and \(\Phi(t,s):=(ts,s)\), a \(C^1\) injection of \(U\) onto \(V\) with inverse \((x,y)\mapsto(x/y,y)\) and
\begin{equation*} J_\Phi(t,s)=\det\begin{pmatrix}s & t\\ 0 & 1\end{pmatrix}=s,\qquad |J_\Phi(t,s)|=|s| . \end{equation*}
For \(h(x,y)=\mathbb{1}_B(x/y)\varrho(x,y)\) on \(V\),
\begin{equation*} h(\Phi(t,s))=\mathbb{1}_B(ts/s)\varrho(ts,s)=\mathbb{1}_B(t)\varrho(ts,s)\qquad (s\neq 0), \end{equation*}
so that
\begin{equation*} \mu\circ g^{-1}(B)=\int_V \mathbb{1}_B(x/y)\varrho(x,y)\,dx\,dy=\int_U \mathbb{1}_B(t)\,|s|\varrho(ts,s)\,dt\,ds , \end{equation*}
and Fubini again makes \(\varrho_2(t)=\int|s|\varrho(ts,s)\,ds\) Borel with \(\mu\circ g^{-1}(B)=\int_B\varrho_2\,dt\) and \(\int\varrho_2=\mu(V)=1\).
Let \(\varphi(x)=\exp\bigl(il(x)\bigr)\), where \(l\) is a nonmeasurable additive function on the real line (such a function is easily constructed by using a Hamel basis). Show that \(\varphi\) is positive definite and \(\varphi(0)=1\).
Additivity of the real-valued \(l\) turns the quadratic form of Definition 3.8.3 into a square: \(\varphi(x_j-x_m)=a_j\overline{a_m}/(c_j\overline{c_m})\) with \(a_j:=c_j\exp(il(x_j))\), as shown below.
A nonmeasurable additive \(l\) exists: let \(H\) be a Hamel basis of \(\mathbb{R}\) over \(\mathbb{Q}\) (Lemma 1.12.20, Example 1.12.21), so each \(x\) has a unique finite representation \(x=\sum_{h}q_h(x)h\) with rational \(q_h(x)\), fix \(h_0\in H\) and put \(l:=q_{h_0}\). Uniqueness makes \(l\) \(\mathbb{Q}\)-linear, in particular
\begin{equation*} l(x+y)=l(x)+l(y)\qquad\text{for all }x,y\in\mathbb{R}, \end{equation*}
so \(l\) is additive. Were it measurable, then, \(l\) taking only rational values, \(\mathbb{R}=\bigcup_{q\in\mathbb{Q}}l^{-1}(q)\) would force \(\lambda(l^{-1}(q))>0\) for some \(q\), whence by Steinhaus’ theorem (Exercise 1.12.62) \(l^{-1}(q)-l^{-1}(q)\supset(-\delta,\delta)\); additivity gives \(l=0\) there, and \(l(x)=n\,l(x/n)=0\) for large \(n\) then gives \(l\equiv0\), contradicting \(l(h_0)=1\).
Now \(\varphi:=\exp(il)\) satisfies \(|\varphi|\equiv1\), and \(l(0)=2l(0)\) gives \(\varphi(0)=1\). For positive definiteness, additivity gives
\begin{equation*} l(x-y)=l(x)-l(y). \end{equation*}
so \(\varphi(x_j-x_m)=\exp(il(x_j))\overline{\exp(il(x_m))}\), \(l\) being real. With \(a_j:=c_j\exp(il(x_j))\),
\begin{equation*} \sum_{j,m=1}^{k}c_j\overline{c_m}\,\varphi(x_j-x_m) =\sum_{j,m=1}^{k}\Bigl(c_j\exp\bigl(il(x_j)\bigr)\Bigr)\overline{\Bigl(c_m\exp\bigl(il(x_m)\bigr)\Bigr)} =\sum_{j,m=1}^{k}a_j\overline{a_m}=\Bigl|\sum_{j=1}^{k}a_j\Bigr|^{2}\ge 0 . \end{equation*}
(i) Let \(\mu\) be a probability measure on \(\mathbb{R}^n\). Prove that
\begin{equation*} 0\le 1-\operatorname{Re}\widehat{\mu}(2y)\le 4\bigl(1-\operatorname{Re}\widehat{\mu}(y)\bigr),\qquad y\in\mathbb{R}^n . \end{equation*}
(ii) Show that if \(\widehat{\mu}(y)=1\) in some neighborhood of the origin, then \(\mu\) is Dirac’s measure at the origin.
(i) The pointwise inequality
\begin{equation*} 0\le 1-\cos 2t=2(1-\cos t)(1+\cos t)\le 4(1-\cos t) \end{equation*}
at \(t=(y,x)\), integrated against the probability measure \(\mu\) (so \(\int1\,d\mu=1\)), is the assertion, since \(\operatorname{Re}\widehat{\mu}(y)=\int\cos(y,x)\,\mu(dx)\).
(ii) Say \(\widehat{\mu}=1\) on \(\{|y|<\varepsilon\}\). Since \(|\widehat\mu|\le1\), \(\operatorname{Re}\widehat\mu(y)=1\) forces \(\operatorname{Im}\widehat\mu(y)=0\), i.e. \(\widehat\mu(y)=1\). We induct on \(k\ge0\): \(\widehat{\mu}=1\) on \(\{|y|<2^{k}\varepsilon\}\), the case \(k=0\) being the hypothesis. For \(|z|<2^{k+1}\varepsilon\) write \(z=2y\) with \(|y|<2^{k}\varepsilon\); by (i) and the inductive hypothesis,
\begin{equation*} 0\le 1-\operatorname{Re}\widehat{\mu}(z)\le 4\bigl(1-\operatorname{Re}\widehat{\mu}(y)\bigr)=0 , \end{equation*}
i.e. \(\widehat{\mu}(z)=1\). As \(2^{k}\varepsilon\to\infty\), \(\widehat\mu\equiv1=\widehat{\delta_0}\) on \(\mathbb{R}^n\), so \(\mu=\delta_0\) by Proposition 3.8.6.
(Gneiting) Let \(E\subset\mathbb{R}\) be a closed set symmetric about the origin and let \(0\in E\). Show that there exist probability measures \(\mu\) and \(\nu\) on \(\mathbb{R}\) such that \(\widehat{\mu}(t)=\widehat{\nu}(t)\) for all \(t\in E\) and \(\widehat{\mu}(t)\neq\widehat{\nu}(t)\) for all \(t\notin E\).
Take \(\mu,\nu\) with densities \(p\pm q\), \(p\) the Cauchy density and \(q\) real, even, with \(|q|\le p/2\) and \(\int q(x)e^{itx}\,dx\) vanishing exactly on \(E\); then \(\widehat\mu-\widehat\nu=2\int e^{itx}q\,dx\) has zero set \(E\).
Put \(U:=\mathbb{R}\setminus E\), open, symmetric, and missing \(0\). (If \(U=\emptyset\) the second requirement is vacuous and \(\mu=\nu\) serves.) No component of \(U\) contains \(0\), so each lies in \((0,\infty)\) or in \((-\infty,0)\), and \(x\mapsto-x\) permutes them; listing the components in \((0,\infty)\) as \(J_n=(a_n,b_n)\), \(0\le a_n<b_n\le\infty\),
\begin{equation*} U=\bigcup_n\bigl(J_n\cup(-J_n)\bigr), \end{equation*}
a disjoint union. Bumps sitting exactly on these components: with \(\theta(u):=e^{-1/u}\) for \(u>0\), \(\theta(u):=0\) otherwise – smooth, \(0\le\theta<1\), positive exactly on \(u>0\), with every \(\theta^{(k)}=P_k(1/u)e^{-1/u}\) bounded – put, for bounded \(J_n\),
\begin{equation*} \psi_n(t):=\theta\bigl(t^2-a_n^2\bigr)\,\theta\bigl(b_n^2-t^2\bigr). \end{equation*}
and, for the at most one unbounded component,
\begin{equation*} \psi_n(t):=\theta\bigl(t^2-a_n^2\bigr)e^{-t^2}. \end{equation*}
Each \(\psi_n\) is smooth, real, even, with \(0\le\psi_n\le1\) and \(\{\psi_n>0\}=J_n\cup(-J_n)\); in the first case it has compact support, in the second it and all its derivatives are \(O(P(t)e^{-t^2})\). So in both cases
\begin{equation*} \psi_n,\ \psi_n^{\prime},\ \psi_n^{\prime\prime}\in L^1(\mathbb{R}),\qquad \lim_{|t|\to\infty}\psi_n(t)=\lim_{|t|\to\infty}\psi_n^{\prime}(t)=0 . \end{equation*}
Now define
\begin{equation*} q_n(x):=\frac{1}{2\pi}\int_{-\infty}^{+\infty}\psi_n(t)e^{-itx}\,dt=\frac{1}{2\pi}\int_{-\infty}^{+\infty}\psi_n(t)\cos(tx)\,dt, \end{equation*}
\(\psi_n\) being real and even, so \(q_n\) is real, even, continuous with \(|q_n|\le\|\psi_n\|_{L^1}/(2\pi)\). Two integrations by parts, legitimate by the decay above, give
\begin{equation*} \int\psi_n^{\prime\prime}(t)e^{-itx}\,dt=(ix)^2\int\psi_n(t)e^{-itx}\,dt=-x^2\int\psi_n(t)e^{-itx}\,dt , \end{equation*}
so that \(|x^2q_n(x)|\le\frac{1}{2\pi}\|\psi_n^{\prime\prime}\|_{L^1}\). Hence
\begin{equation*} M_n:=\sup_{x\in\mathbb{R}}\bigl(1+x^2\bigr)|q_n(x)|\le\frac{1}{2\pi}\Bigl(\|\psi_n\|_{L^1}+\|\psi_n^{\prime\prime}\|_{L^1}\Bigr)<\infty . \end{equation*}
In particular \(q_n\in L^1\). Since \(q_n=(2\pi)^{-1/2}\widehat{\psi_n}\) in the book’s normalization, with \(\psi_n\) and \(\widehat{\psi_n}\) both integrable, the inversion of Corollary 3.8.12 gives a continuous modification of \(\psi_n\) equal to \(\int q_n(x)e^{itx}\,dx\); \(\psi_n\) being continuous, it equals that everywhere:
\begin{equation*} \int_{-\infty}^{+\infty}q_n(x)e^{itx}\,dx=\psi_n(t)\qquad\text{for every }t\in\mathbb{R}. \end{equation*}
With \(p(x):=1/(\pi(1+x^2))\) the Cauchy density, choose \(c_n>0\) with
\begin{equation*} c_nM_n\le\frac{2^{-n}}{2\pi}\qquad (n=1,2,\dots) \end{equation*}
and set
\begin{equation*} q(x):=\sum_n c_nq_n(x),\qquad h(t):=\sum_n c_n\psi_n(t). \end{equation*}
The first series converges absolutely and uniformly:
\begin{equation*} \sum_n c_n|q_n(x)|\le\frac{1}{2\pi(1+x^2)}\sum_n 2^{-n}\le\frac{1}{2\pi(1+x^2)}=\frac{p(x)}{2}, \end{equation*}
so \(q\) is real, even, continuous with \(|q|\le p/2\). In the second series at most one summand is nonzero at each \(t\), the sets \(J_n\cup(-J_n)\) being disjoint, so \(h>0\) on \(U\) and \(h=0\) on \(E\): the zero set of \(h\) is exactly \(E\). By dominated convergence (partial sums dominated by \(p/2\in L^1\)),
\begin{equation*} \int_{-\infty}^{+\infty}q(x)e^{itx}\,dx=\sum_n c_n\int_{-\infty}^{+\infty}q_n(x)e^{itx}\,dx=\sum_n c_n\psi_n(t)=h(t) \end{equation*}
for every \(t\); in particular \(\int q\,dx=h(0)=0\), as \(0\in E\). Finally set \(p_\mu:=p+q\) and \(p_\nu:=p-q\), both continuous, integrable and \(\ge p/2>0\), with
\begin{equation*} \int p_\mu\,dx=\int p\,dx+\int q\,dx=1+0=1,\qquad \int p_\nu\,dx=\int p\,dx-\int q\,dx=1 . \end{equation*}
Let \(\mu\) and \(\nu\) be the Borel probability measures with densities \(p_\mu\) and \(p_\nu\). Then for every \(t\),
\begin{equation*} \widehat{\mu}(t)-\widehat{\nu}(t)=\int_{-\infty}^{+\infty}e^{itx}\bigl(p_\mu(x)-p_\nu(x)\bigr)dx=2\int_{-\infty}^{+\infty}e^{itx}q(x)\,dx=2h(t). \end{equation*}
Since the zero set of \(h\) is exactly \(E\), we conclude
\begin{equation*} \widehat{\mu}(t)=\widehat{\nu}(t)\ \text{ for }t\in E,\qquad \widehat{\mu}(t)\neq\widehat{\nu}(t)\ \text{ for }t\notin E, \end{equation*}
which is the required statement.
Let \(\mu\) and \(\nu\) be two Borel probability measures on the real line. Prove that
\begin{equation*} \int\!\!\int (x+y)^2\,\mu(dx)\,\nu(dy)<\infty\quad\text{precisely when}\quad \int x^2\,\mu(dx)+\int y^2\,\nu(dy)<\infty . \end{equation*}
Both directions are the inequality \((u+v)^2\le2u^2+2v^2\); the integrand being nonnegative Borel, Theorems 3.4.4 and 3.4.5 let us read the double integral as either repeated integral.
(i) If \(\int x^2\,d\mu\) and \(\int y^2\,d\nu\) are finite, then, \(\mu\) and \(\nu\) being probability measures,
\begin{equation*} \int\!\!\int (x+y)^2\,\mu(dx)\,\nu(dy)\le 2\int\!\!\int x^2\,\mu(dx)\,\nu(dy)+2\int\!\!\int y^2\,\mu(dx)\,\nu(dy) =2\int x^2\,\mu(dx)+2\int y^2\,\nu(dy)<\infty . \end{equation*}
(ii) Conversely, let \(I:=\int\!\!\int(x+y)^2\,\mu(dx)\,\nu(dy)<\infty\) and write \(I=\int F\,d\nu\) with \(F(y):=\int(x+y)^2\,\mu(dx)\). Then \(\{F=+\infty\}\) is \(\nu\)-null, hence not all of \(\mathbb{R}\), so \(F(y_0)<\infty\) for some \(y_0\); from \(x^2\le2(x+y_0)^2+2y_0^2\),
\begin{equation*} \int x^2\,\mu(dx)\le 2\int (x+y_0)^2\,\mu(dx)+2y_0^2<\infty . \end{equation*}
The other repeated integral gives, in the same way, an \(x_0\) with \(\int(x_0+y)^2\,d\nu<\infty\), whence \(\int y^2\,d\nu\le2\int(x_0+y)^2\,d\nu+2x_0^2<\infty\).
Exercises 3.10.120–3.10.125
(Gromov) Suppose that in \(\mathbb{R}^n\) we are given \(k\le n+1\) balls \(B(x_i,r_i)\) with the centers \(x_i\) and radii \(r_i\) and \(k\) balls \(B(y_i,r_i)\) with the centers \(y_i\) and radii \(r_i\) such that \(|x_i-x_j|\ge|y_i-y_j|\) for all \(i,j\). Then the following inequality holds:
\begin{equation*} \lambda_n\Big(\bigcap_{i=1}^k B(x_i,r_i)\Big)\le\lambda_n\Big(\bigcap_{i=1}^k B(y_i,r_i)\Big), \end{equation*}
where \(\lambda_n\) is Lebesgue measure.
As far as I know, the following question raised in the 1950s by several authors (M. Kneser, E.T. Poulsen, and H. Hadwiger; see Meyer, Reisner, Schmuckenschläger) remains open: suppose that in \(\mathbb{R}^n\) we are given \(k\) balls \(B(x_i,r)\) of radius \(r\) centered at the points \(x_1,\dots,x_k\) and \(k\) balls \(B(y_i,r)\) of radius \(r\) centered at the points \(y_1,\dots,y_k\) such that \(|x_i-x_j|\le|y_i-y_j|\) for all \(i,j\); is it true that
\begin{equation*} \lambda_n\Big(\bigcup_{i=1}^k B(x_i,r)\Big)\le\lambda_n\Big(\bigcup_{i=1}^k B(y_i,r)\Big)? \end{equation*}
Throughout the radii \(r_1,\dots,r_k>0\) are fixed, balls are taken closed (this changes no volume), and for a configuration \(q=(q_1,\dots,q_k)\in(\mathbb{R}^n)^k\) we write
\begin{equation*} K(q)=\bigcap_{i=1}^k B(q_i,r_i),\qquad V(q)=\lambda_n\big(K(q)\big),\qquad S_i(q)=\partial B(q_i,r_i). \end{equation*}
Each \(K(q)\) is compact convex, and the proof is the Csikos variation formula \(\frac{d}{dt}V(p(t))=-\sum_{i<j}\mathcal{H}^{n-1}(W_{ij}(t))\dot d_{ij}(t)\) along a contraction \(p\) joining \(x\) to \(y\) inside \(\mathbb{R}^n\), which exists because \(k\le n+1\).
(a) \(V\) is continuous: if \(\delta=\max_i|q_i^{\prime}-q_i|\), then \(K(q^{\prime})\subseteq\bigcap_i B(q_i,r_i+\delta)\) and \(\bigcap_i B(q_i,r_i-\delta)\subseteq K(q^{\prime})\). As \(\delta\downarrow0\) the outer sets decrease to \(K(q)\) and the inner ones increase to the interior of \(K(q)\), of the same measure \(V(q)\) (a convex set with empty interior is null).
(b) For \(n=1\) we have \(k\le2\), and
\begin{equation*} \lambda_1\big(B(q_1,r_1)\cap B(q_2,r_2)\big)=\max\big(0,\,r_1+r_2-|q_1-q_2|\big), \end{equation*}
is nonincreasing in \(|q_1-q_2|\); so let \(n\ge2\).
(c) The \(x_i\) may be taken pairwise distinct: if \(x_i=x_j\) then \(y_i=y_j\) and both intersections are unchanged by deleting the index \(j\) and replacing \(r_i\) by \(\min(r_i,r_j)\), with \(k-1\le n+1\) still.
(d) They may be taken affinely independent as well. Let \(\rho=\min_{i\ne j}|x_i-x_j|>0\) and let \(\theta>1\). The configuration \(\theta x=(\theta x_1,\dots,\theta x_k)\) has \(|\theta x_i-\theta x_j|\ge|x_i-x_j|+(\theta-1)\rho\). Since \(k\le n+1\), affinely independent \(k\)-tuples are dense in \((\mathbb{R}^n)^k\), so there is an affinely independent \(x^{\prime}\) with \(\max_i|x_i^{\prime}-\theta x_i|\) so small that all pairwise distances differ from those of \(\theta x\) by less than \((\theta-1)\rho/2\); then \(|x_i^{\prime}-x_j^{\prime}|\ge|x_i-x_j|\ge|y_i-y_j|\) for all \(i,j\), and \(x^{\prime}\to x\) as \(\theta\downarrow1\). Granting the theorem for such \(x^{\prime}\), letting \(\theta\downarrow1\) and using (a) gives \(V(x)\le V(y)\).
The contraction: for a configuration \(q\) put \(d_{ij}(q)=|q_i-q_j|\) and let \(G(q)\) be the \((k-1)\times(k-1)\) matrix
\begin{equation*} G(q)_{ij}=\langle q_i-q_k,\;q_j-q_k\rangle=\tfrac12\big(d_{ik}^2+d_{jk}^2-d_{ij}^2\big),\qquad 1\le i,j\le k-1 . \end{equation*}
This Gram matrix is positive semidefinite, and positive definite exactly when \(q_1,\dots,q_k\) are affinely independent. Conversely, for symmetric positive semidefinite \(G\) the columns \(v_i\) of \(G^{1/2}\) have \(\langle v_i,v_j\rangle=G_{ij}\), so the points \(p_i=v_i\) \((i<k)\), \(p_k=0\) of \(\mathbb{R}^{k-1}\) satisfy
\begin{equation*} |p_i-p_j|^2=G_{ii}+G_{jj}-2G_{ij},\qquad |p_i-p_k|^2=G_{ii}. \end{equation*}
So squared distance matrices of \(k\)-point configurations correspond affinely and bijectively to positive semidefinite \((k-1)\times(k-1)\) matrices. Set \(G(t):=(1-t)G(x)+tG(y)\), again positive semidefinite, with squared distances
\begin{equation*} d_{ij}(t)^2=(1-t)|x_i-x_j|^2+t|y_i-y_j|^2 . \end{equation*}
Let \(p(t)\) be the configuration built from \(G(t)^{1/2}\); as \(k-1\le n\) it sits in \(\mathbb{R}^n\), and:
(i) \(|p_i(t)-p_j(t)|=d_{ij}(t)\) with \(d_{ij}^2\) affine in \(t\) and \(d_{ij}(1)\le d_{ij}(0)\), so each \(d_{ij}\) is nonincreasing;
(ii) \(p(0)\) and \(p(1)\) have the distance matrices of \(x\) and \(y\), so \(V(p(0))=V(x)\), \(V(p(1))=V(y)\), configurations being determined by distances up to isometry and \(V\) being isometry invariant;
(iii) with \(G(x)\) positive definite by (d), \(G(t)\succeq(1-t)G(x)\succ0\) for \(t<1\), and \(A\mapsto A^{1/2}\) is real analytic on positive definite and continuous on positive semidefinite matrices, so \(p\) is analytic on \([0,1)\) and continuous on \([0,1]\);
(iv) the centers \(p_i(t)\), \(t<1\), are distinct, since \(d_{ij}(t)^2\ge(1-t)|x_i-x_j|^2>0\).
By (i), \(d_{ij}\) is differentiable where \(d_{ij}>0\), with
\begin{equation*} \dot d_{ij}(t)=\frac{|y_i-y_j|^2-|x_i-x_j|^2}{2\,d_{ij}(t)}\le0 . \end{equation*}
so it suffices that \(t\mapsto V(p(t))\) be nondecreasing on \([0,1)\), continuity of \(V\) and \(p\) then giving \(V(x)\le V(y)\).
First variation of \(V\): let \(q\) have distinct centers, so that each \(S_i\cap S_j\), \(i\ne j\), lies in a hyperplane and is \(\mathcal{H}^{n-1}\)-null (\(n\ge2\)). Put
\begin{equation*} A_i=\bigcap_{l\ne i}B(q_l,r_l),\qquad \Sigma_i=S_i\cap A_i,\qquad \nu_i(z)=\frac{z-q_i}{r_i}. \end{equation*}
so that \(\Sigma_i\) is the part of \(\partial K(q)\) on \(S_i\). We claim
\begin{equation*} \nabla_{q_i}V(q)=\int_{\Sigma_i}\nu_i\,d\mathcal{H}^{n-1}. \end{equation*}
Only the \(i\)-th ball moves, so with \(B_i=B(q_i,r_i)\),
\begin{equation*} V(q_1,\dots,q_i+h,\dots,q_k)-V(q)=\lambda_n\big(A_i\cap(B_i+h)\big)-\lambda_n\big(A_i\cap B_i\big) =\lambda_n\big(A_i\cap((B_i+h)\setminus B_i)\big)-\lambda_n\big(A_i\cap(B_i\setminus(B_i+h))\big). \end{equation*}
In polar coordinates \(z=q_i+s\omega\) around \(q_i\), the part of a ray inside \((B_i+h)\setminus B_i\) has length \(\langle h,\omega\rangle^{+}+o(|h|)\) uniformly in \(\omega\) and lies within \(|h|\) of \(q_i+r_i\omega\), so
\begin{equation*} \frac1{|h|}\lambda_n\big(A_i\cap((B_i+h)\setminus B_i)\big) =\int_{S^{n-1}}r_i^{\,n-1}\,\big\langle h/|h|,\omega\big\rangle^{+}\,\chi(\omega,h)\,d\omega+o(1), \end{equation*}
where \(\chi(\omega,h)\in[0,1]\) and \(\chi(\omega,h)\to1\) for every \(\omega\) with \(q_i+r_i\omega\) in the interior of \(A_i\) and \(\chi(\omega,h)\to0\) for every \(\omega\) with \(q_i+r_i\omega\notin A_i\). The exceptional \(\omega\) correspond to \(S_i\cap\partial A_i\subseteq\bigcup_{l\ne i}(S_i\cap S_l)\), which is null, so dominated convergence gives the limit \(\int_{\Sigma_i}\langle h/|h|,\nu_i\rangle^{+}d\mathcal{H}^{n-1}\); likewise \(|h|^{-1}\lambda_n(A_i\cap(B_i\setminus(B_i+h)))\to\int_{\Sigma_i}\langle h/|h|,\nu_i\rangle^{-}d\mathcal{H}^{n-1}\), and subtracting proves the claim.
For the variation formula, use the power functions \(\pi_i(z)=|z-q_i|^2-r_i^2\), which are affine up to the common term \(|z|^2\), so that each \(\pi_i-\pi_j\) is affine. Put
\begin{equation*} D_i=\{z:\pi_i(z)\ge\pi_l(z)\ \ \forall\,l\},\qquad F_i=D_i\cap B(q_i,r_i),\qquad H_{ij}=\{\pi_i=\pi_j\},\qquad W_{ij}=F_i\cap H_{ij}. \end{equation*}
Each \(D_i\) is a closed convex polyhedron, so \(F_i\) is compact convex, and \(H_{ij}\perp(q_j-q_i)\). Then (Check!): \(F_i\subseteq K(q)\) (as \(\pi_l\le\pi_i\le0\) on \(F_i\)); \(\bigcup_iF_i=K(q)\) (maximize \(\pi_i(z)\)); \(\partial F_i\cap S_i=\Sigma_i\); \(W_{ij}=F_i\cap F_j\) and \(\partial F_i\subseteq\Sigma_i\cup\bigcup_{j\ne i}W_{ij}\), since \(\partial D_i\subseteq\bigcup_{j\ne i}H_{ij}\); and the outer unit normal of \(D_i\) along \(H_{ij}\) is \(-e_{ij}\), \(e_{ij}:=(q_j-q_i)/|q_j-q_i|\), because \(\nabla(\pi_i-\pi_j)=2(q_j-q_i)\) and \(D_i\subseteq\{\pi_i\ge\pi_j\}\). Gauss-Green for constant fields on \(F_i\) gives \(\int_{\partial F_i}\nu\,d\mathcal{H}^{n-1}=0\) (all surface terms vanish when \(F_i\) has empty interior), i.e.
\begin{equation*} \int_{\Sigma_i}\nu_i\,d\mathcal{H}^{n-1}=\sum_{j\ne i}\mathcal{H}^{n-1}(W_{ij})\,e_{ij}. \end{equation*}
Combining with the first variation,
\begin{equation*} \nabla_{q_i}V(q)=\sum_{j\ne i}\mathcal{H}^{n-1}(W_{ij})\,e_{ij},\qquad i=1,\dots,k . \end{equation*}
These are continuous on the open set \(U\) of configurations with distinct centers, as is seen from the first-variation formula rather than the right-hand side: with \(z=q_i+r_i\omega\),
\begin{equation*} \nabla_{q_i}V(q)=r_i^{\,n-1}\int_{S^{n-1}}\omega\,\mathbf 1_{A_i(q)}(q_i+r_i\omega)\,d\omega . \end{equation*}
For \(q^{(m)}\to q\) in \(U\), each strict inequality \(|q_i+r_i\omega-q_l|<r_l\) or \(>r_l\) persists for large \(m\), so the integrand is eventually \(\omega\) or \(0\); the remaining \(\omega\) are those with \(q_i+r_i\omega\in S_i\cap S_l\), a null set as the centers are distinct. Dominated convergence gives \(\nabla_{q_i}V(q^{(m)})\to\nabla_{q_i}V(q)\), so \(V\in C^1(U)\).
By (iv), \(p(t)\in U\) and \(p\) is differentiable for \(t<1\), so the chain rule gives
\begin{equation*} \frac{d}{dt}V(p(t))=\sum_{i=1}^k\big\langle\nabla_{q_i}V(p(t)),\dot p_i(t)\big\rangle =\sum_{i}\sum_{j\ne i}\mathcal{H}^{n-1}(W_{ij})\big\langle\dot p_i,e_{ij}\big\rangle =\sum_{i<j}\mathcal{H}^{n-1}(W_{ij})\big\langle\dot p_i-\dot p_j,e_{ij}\big\rangle . \end{equation*}
Since \(\langle \dot p_j-\dot p_i,e_{ij}\rangle=\dot d_{ij}\), this is the variation formula
\begin{equation*} \frac{d}{dt}\lambda_n\Big(\bigcap_{i=1}^kB(p_i(t),r_i)\Big) =-\sum_{i<j}\mathcal{H}^{n-1}(W_{ij}(t))\,\dot d_{ij}(t) , \end{equation*}
which is \(\ge0\), all \(\dot d_{ij}\le0\) by (i). Hence \(V(p(\cdot))\) is nondecreasing on \([0,1)\), and by continuity \(V(x)=V(p(0))\le V(p(1))=V(y)\).
As for the question raised in the statement: \(k\le n+1\) entered only in realizing the interpolants \(G(t)\) inside \(\mathbb{R}^n\) rather than \(\mathbb{R}^{k-1}\), and for \(k\le n+1\) the same computation on the nearest-point power diagram, \(C_i=\{\pi_i\le\pi_l\ \forall l\}\) with \(\bigcup_lB(q_l,r_l)\cap C_i=C_i\cap B(q_i,r_i)\), gives
\begin{equation*} \frac{d}{dt}\lambda_n\Big(\bigcup_{i}B(p_i(t),r_i)\Big)=\sum_{i<j}\mathcal{H}^{n-1}(\widetilde W_{ij}(t))\,\dot d_{ij}(t)\le 0 \end{equation*}
along a contraction; for \(k>n+1\) a contraction need not be realizable by a continuous contraction inside \(\mathbb{R}^n\).
(i) Let \((X_i,\mathcal{A}_i,\mu_i)\), \(i=1,\dots,n\), be measurable spaces with nonnegative \(\sigma\)-finite measures and let \(f_i\) be nonnegative \(\bigotimes_{i=1}^n\mu_i\)-measurable functions on \(\prod_{i=1}^nX_i\) such that \(f_i\) is independent of the \(i\)-th variable. Prove the inequality
\begin{equation*} \Big(\int f_1\cdots f_n\,d\mu_1\cdots d\mu_n\Big)^{n-1}\le\prod_{i=1}^n\int f_i^{\,n-1}\prod_{j\ne i}d\mu_j . \end{equation*}
(ii) Let \(E\) be a Borel set in \(\mathbb{R}^3\) and let \(E_i\) be its orthogonal projection to the coordinate plane \(x_i=0\). Prove the inequality \(\lambda_3(E)^2\le\lambda_2(E_1)\lambda_2(E_2)\lambda_2(E_3)\).
(i) Induct on \(n\), writing
\begin{equation*} I=\int f_1\cdots f_n\,d\mu_1\cdots d\mu_n,\qquad I_i=\int f_i^{\,n-1}\prod_{j\ne i}d\mu_j \end{equation*}
for the two sides of \(I^{\,n-1}\le\prod_iI_i\); here \(f_i\) is a section of a product-measurable function, hence \(\bigotimes_{j\ne i}\mathcal{A}_j\)-measurable, so \(I_i\) is defined, and Tonelli’s theorem applies throughout, the measures being \(\sigma\)-finite and the functions nonnegative. We may assume \(0<I_i<\infty\): \(I_i=\infty\) is trivial, and \(I_i=0\) forces \(f_i=0\) a.e., hence \(I=0\).
For \(n=2\), Tonelli gives
\begin{equation*} I=\int f_1(x_2)f_2(x_1)\,\mu_1(dx_1)\mu_2(dx_2)=\Big(\int f_2\,d\mu_1\Big)\Big(\int f_1\,d\mu_2\Big)=I_2I_1 , \end{equation*}
with equality. For \(n\ge3\), assuming the case \(n-1\): since \(f_1\) is free of \(x_1\), Tonelli gives
\begin{equation*} I=\int f_1\Big(\int f_2\cdots f_n\,d\mu_1\Big)\,d\mu_2\cdots d\mu_n . \end{equation*}
The generalized Hoelder inequality on \((X_1,\mu_1)\) with the \(n-1\) equal exponents \(n-1\) (whose reciprocals sum to \(1\)) applied to \(f_2,\dots,f_n\) gives
\begin{equation*} \int f_2\cdots f_n\,d\mu_1\le\prod_{i=2}^n\Big(\int f_i^{\,n-1}d\mu_1\Big)^{1/(n-1)}=\prod_{i=2}^ng_i^{1/(n-1)}, \qquad g_i:=\int f_i^{\,n-1}d\mu_1 . \end{equation*}
By Tonelli each \(g_i\) is measurable on \(\prod_{j\ge2}X_j\) and free of the \(i\)-th variable, so
\begin{equation*} I\le\int f_1\,g_2^{1/(n-1)}\cdots g_n^{1/(n-1)}\,d\mu_2\cdots d\mu_n . \end{equation*}
Hoelder on \(\prod_{j\ge2}X_j\) with \(p=n-1\), \(q=(n-1)/(n-2)\) now gives
\begin{equation*} I\le\Big(\int f_1^{\,n-1}d\mu_2\cdots d\mu_n\Big)^{1/(n-1)} \Big(\int\big(g_2\cdots g_n\big)^{1/(n-2)}\,d\mu_2\cdots d\mu_n\Big)^{(n-2)/(n-1)} =I_1^{1/(n-1)}\,J^{(n-2)/(n-1)}, \end{equation*}
where \(J=\int h_2\cdots h_n\,d\mu_2\cdots d\mu_n\) with \(h_i:=g_i^{1/(n-2)}\).
Each \(h_i\) is free of the \(i\)-th variable, so the inductive hypothesis applies to \(h_2,\dots,h_n\):
\begin{equation*} J^{\,n-2}\le\prod_{i=2}^n\int h_i^{\,n-2}\prod_{2\le j\le n,\ j\ne i}d\mu_j =\prod_{i=2}^n\int g_i\prod_{j\ge2,\ j\ne i}d\mu_j . \end{equation*}
By Tonelli,
\begin{equation*} \int g_i\prod_{j\ge2,\ j\ne i}d\mu_j=\int f_i^{\,n-1}\,d\mu_1\prod_{j\ge2,\ j\ne i}d\mu_j =\int f_i^{\,n-1}\prod_{j\ne i}d\mu_j=I_i , \end{equation*}
so \(J\le(\prod_{i\ge2}I_i)^{1/(n-2)}\), and substituting,
\begin{equation*} I\le I_1^{1/(n-1)}\Big(\prod_{i=2}^nI_i\Big)^{\frac1{n-2}\cdot\frac{n-2}{n-1}} =\Big(\prod_{i=1}^nI_i\Big)^{1/(n-1)} , \end{equation*}
i.e. \(I^{\,n-1}\le\prod_{i=1}^nI_i\).
(ii) This is Loomis-Whitney in \(\mathbb{R}^3\); the projections \(E_i\) are Souslin, hence Lebesgue measurable (Chapter 6), and the argument below in fact bounds \(\lambda_3(E)^2\) by the outer measures. Let \(K\subseteq E\) be compact with compact projections \(K_1,K_2,K_3\), take \(X_i=\mathbb{R}^1\), \(\mu_i=\lambda_1\), \(n=3\), and put
\begin{equation*} f_1(x_2,x_3)=\mathbf 1_{K_1}(x_2,x_3),\quad f_2(x_1,x_3)=\mathbf 1_{K_2}(x_1,x_3),\quad f_3(x_1,x_2)=\mathbf 1_{K_3}(x_1,x_2), \end{equation*}
so that \(f_i\) is free of the \(i\)-th variable and, the projections of a point of \(K\) lying in the \(K_i\),
\begin{equation*} \mathbf 1_K(x_1,x_2,x_3)\le f_1f_2f_3 . \end{equation*}
Since \(f_i^2=f_i\) and \(\int f_i\,d\lambda_2=\lambda_2(K_i)\), part (i) with \(n=3\) gives
\begin{equation*} \lambda_3(K)^2\le\Big(\int f_1f_2f_3\,dx_1dx_2dx_3\Big)^2\le\prod_{i=1}^3\int f_i^{\,2}\prod_{j\ne i}dx_j =\lambda_2(K_1)\lambda_2(K_2)\lambda_2(K_3). \end{equation*}
Since \(K_i\subseteq E_i\), and since inner regularity gives \(\lambda_3(E)=\sup_K\lambda_3(K)\) (also when \(\lambda_3(E)=\infty\)),
\begin{equation*} \lambda_3(E)^2\le\lambda_2^*(E_1)\lambda_2^*(E_2)\lambda_2^*(E_3)=\lambda_2(E_1)\lambda_2(E_2)\lambda_2(E_3). \end{equation*}
(i) (T. Carleman) Suppose we are given a sequence of numbers \(\sigma_n\) with \(\sum_{n=1}^{\infty}\sigma_{2n}^{-1/(2n)}=\infty\). Prove that two probability measures \(\mu\) and \(\nu\) on the real line coincide if they have equal moments
\begin{equation*} \int_{-\infty}^{+\infty}t^n\,\mu(dt)=\int_{-\infty}^{+\infty}t^n\,\nu(dt)=\sigma_n,\qquad\forall\,n\in\mathbb{N}. \end{equation*}
(ii) Prove that for all \(n\) one has
\begin{equation*} \int_0^{\infty}x^n\exp\big(-x^{1/4}\big)\sin\big(x^{1/4}\big)\,dx=0 . \end{equation*}
Deduce the existence of two different probability measures on the real line with equal moments for all \(n\).
(iii) (M.G. Krein) Show that a probability density \(\varrho\) on the real line is not uniquely determined by its moments in the class of all probability measures precisely when the function \((1+x^2)^{-1}\min\big(\ln\varrho(x),0\big)\) has a finite integral over \(\mathbb{R}^1\).
(i) Equal moments together with Carleman’s condition force \(\mu=\nu\). If \(\sigma_{2m}=0\) for some \(m\ge1\), then \(\mu=\nu=\delta_0\), so assume \(\sigma_{2m}>0\) for all \(m\). All absolute moments are finite: \(\int|t|^{2m}\mu(dt)=\sigma_{2m}\), and \(\int|t|^{2m+1}\mu(dt)\le\sigma_{2m}^{1/2}\sigma_{2m+2}^{1/2}<\infty\) by the Cauchy–Schwarz inequality, and likewise for \(\nu\). Put \(\varrho=\tfrac12(\mu+\nu)\), \(M_k=2\int|t|^k\varrho(dt)\) and
\begin{equation*} \phi(s)=\int_{\mathbb{R}}e^{ist}\,\mu(dt)-\int_{\mathbb{R}}e^{ist}\,\nu(dt). \end{equation*}
Differentiation under the integral sign (dominated by the integrable \(|t|^k\)) gives \(\phi\in C^{\infty}(\mathbb{R})\) with
\begin{equation*} \phi^{(k)}(s)=\int(it)^ke^{ist}\mu(dt)-\int(it)^ke^{ist}\nu(dt), \end{equation*}
hence \(\phi^{(k)}(0)=i^k(\sigma_k-\sigma_k)=0\) and \(\sup_{\mathbb{R}}|\phi^{(k)}|\le M_k\) for every \(k\ge0\). The sequence \((M_k)\) is logarithmically convex, \(M_k^2\le M_{k-1}M_{k+1}\), by the Cauchy–Schwarz inequality applied to \(|t|^{(k-1)/2}\cdot|t|^{(k+1)/2}\), and, since \(M_{2n}=2\sigma_{2n}\),
\begin{equation*} \sum_{k\ge1}M_k^{-1/k}\ \ge\ \sum_{n\ge1}M_{2n}^{-1/(2n)} \ \ge\ \tfrac12\sum_{n\ge1}\sigma_{2n}^{-1/(2n)}=\infty . \end{equation*}
Thus \(\phi\) belongs to the Denjoy–Carleman class attached to \((M_k)\), which is quasi-analytic precisely because \(\sum_kM_k^{-1/k}=\infty\); as \(\phi\) vanishes to infinite order at \(0\), \(\phi\equiv0\). Equal Fourier transforms give \(\mu=\nu\) by Proposition 3.8.6.
(ii) The substitution \(x=u^4\) turns the integral into a Gamma integral whose imaginary part vanishes:
\begin{equation*} \int_0^{\infty}x^ne^{-x^{1/4}}\sin\big(x^{1/4}\big)\,dx =4\int_0^{\infty}u^{4n+3}e^{-u}\sin u\,du , \end{equation*}
while for every integer \(m\ge0\), by absolute convergence (\(\operatorname{Re}(1-i)=1>0\)) and \(1-i=\sqrt2\,e^{-i\pi/4}\),
\begin{equation*} \int_0^{\infty}u^me^{-u}\sin u\,du=\operatorname{Im}\frac{\Gamma(m+1)}{(1-i)^{m+1}} =m!\,2^{-(m+1)/2}\sin\frac{(m+1)\pi}{4}. \end{equation*}
For \(m=4n+3\) the sine is \(\sin(n+1)\pi=0\), so the integral vanishes for every \(n\ge0\). Now take
\begin{equation*} \varrho_1(x)=\tfrac1{24}e^{-x^{1/4}}\mathbf 1_{(0,\infty)}(x),\qquad \varrho_2=\varrho_1\big(1+\sin(x^{1/4})\big), \end{equation*}
both nonnegative, with \(\int_0^{\infty}e^{-x^{1/4}}dx=4\cdot3!=24\) and all moments finite, \(\int_0^{\infty}x^ne^{-x^{1/4}}dx=4\,\Gamma(4n+4)\). The vanishing just proved gives
\begin{equation*} \int_{\mathbb{R}}x^n(\varrho_2-\varrho_1)\,dx =\frac1{24}\int_0^{\infty}x^ne^{-x^{1/4}}\sin\big(x^{1/4}\big)dx=0 \end{equation*}
for all \(n\ge0\): \(n=0\) makes \(\varrho_2\) a probability density, \(n\ge1\) equates all moments, and \(\varrho_2-\varrho_1=\varrho_1\sin(x^{1/4})\) is nonzero on a set of positive measure.
(iii) The stated condition is equivalent to Krein’s condition \(\int(1+x^2)^{-1}|\ln\varrho|\,dx<\infty\), because \(\ln^{+}\varrho\le\varrho\in L^1\) makes the positive part always integrable; Krein’s condition implies indeterminacy. It forces \(\varrho>0\) a.e. and makes
\begin{equation*} F(z)=\exp\Big(\frac1{\pi i}\int_{\mathbb{R}} \Big(\frac1{t-z}-\frac{t}{1+t^2}\Big)\ln\varrho(t)\,dt\Big) \end{equation*}
analytic on \(\mathbb{C}^{+}\) (the integral converges absolutely), the outer function with boundary modulus \(|F(x)|=\varrho(x)\) a.e.; write \(F(x)=\varrho(x)e^{i\theta(x)}\) and \(F_\lambda(z)=e^{i\lambda z}F(z)\) for \(\lambda>0\) (inner times outer, same boundary modulus). All moments of \(\varrho\) being finite, \(|x+i|^n\varrho(x)\in L^1(\mathbb{R})\), so \((z+i)^nF_\lambda\) lies in the Nevanlinna class with integrable boundary modulus, hence in \(H^1(\mathbb{C}^{+})\), where \(\int_{\mathbb{R}}G(x)\,dx=0\). Taking linear combinations over \(n\) and then imaginary parts,
\begin{equation*} \int_{\mathbb{R}}x^n\varrho(x)\sin\big(\theta(x)+\lambda x\big)\,dx=0\qquad(n\ge0). \end{equation*}
For two values \(\lambda_1\ne\lambda_2\) the phases cannot both vanish a.e. (that would force \((\lambda_1-\lambda_2)x\in\pi\mathbb{Z}\) a.e.), so fix \(\lambda\) with \(h:=\sin(\theta+\lambda x)\) not a.e. zero; then \(\varrho_{\pm}=\varrho(1\pm h)\) are two distinct probability densities with the moments of \(\varrho\).
The converse half is the delicate one and is not proved here: divergence of the logarithmic integral yields determinacy only under additional regularity of \(\varrho\) (Krein, Lin, Pedersen), so the “precisely when” is to be read within such a class of regular densities. (A fully rigorous proof of the remaining step is beyond the scope of this page; see the reference given in the book.)
Let \(f\) and \(g\) be nonnegative Lebesgue measurable functions on \(\mathbb{R}^n\) and let the mapping \(f*g\) with values in \([0,+\infty]\) be defined as follows: \(f*g(x)\) is the integral of the function \(y\mapsto f(x-y)g(y)\) if it is integrable and \(f*g(x)=+\infty\) otherwise. Show that \(f*g\) is Borel measurable.
\(f*g\) is an everywhere defined pointwise limit of continuous functions, hence Borel. With \(f_k=\min(f,k)\) and \(g_k=\min(g,k)\mathbf 1_{[-k,k]^n}\) we have \(f_k\in L^{\infty}(\mathbb{R}^n)\) and \(g_k\in L^1(\mathbb{R}^n)\), so each \(f_k*g_k\) is bounded and continuous by Corollary 3.9.6 (the exponents \(\infty\) and \(1\) are conjugate). Since \(f_k\uparrow f\) and \(g_k\uparrow g\) pointwise, for each fixed \(x\) the nonnegative functions \(y\mapsto f_k(x-y)g_k(y)\) increase to \(y\mapsto f(x-y)g(y)\), whence by the monotone convergence theorem
\begin{equation*} \lim_{k\to\infty}f_k*g_k(x)=\int_{\mathbb{R}^n}f(x-y)g(y)\,dy=f*g(x)\in[0,+\infty] \end{equation*}
at every point \(x\) (the integrand is Lebesgue measurable, \(y\mapsto x-y\) being an affine bijection, and the prescribed value \(+\infty\) is exactly the value of the integral when it is not finite). Therefore, for every \(c\in\mathbb{R}\),
\begin{equation*} \{x:f*g(x)>c\}=\bigcup_{m\ge1}\bigcup_{N\ge1}\bigcap_{k\ge N}\{x:f_k*g_k(x)>c+\tfrac1m\} \end{equation*}
is a Borel set, so \(f*g\) is Borel measurable although \(f\) and \(g\) are only Lebesgue measurable.
Let \(B\) be an open ball in \(\mathbb{R}^n\) and let \(f\colon B\to\mathbb{R}\) be a measurable function such that
\begin{equation*} \int_B\int_B\frac{|f(x)-f(y)|}{|x-y|^{n+1}}\,dx\,dy<\infty . \end{equation*}
Prove that \(f=c\) a.e., where \(c\) is a constant.
The constant is the \(L^1_{\mathrm{loc}}\)-limit of the mollifications: each \(f_\varepsilon=g_\varepsilon*(f\mathbf 1_B)\) is smooth with finite double integral, hence constant. Write \(D\) for the double integral and \(R\) for the diameter of \(B\).
First \(f\in L^1(B)\): since \(|x-y|\le R\) on \(B\times B\), \(\int_B\int_B|f(x)-f(y)|dx\,dy\le R^{n+1}D<\infty\), so by Tonelli’s theorem there is \(y_0\in B\) with \(|f(y_0)|<\infty\) and \(\int_B|f-f(y_0)|dx<\infty\).
Fix \(g\in C_0^{\infty}(\mathbb{R}^n)\), \(g\ge0\), \(\operatorname{supp}g\subseteq\{|z|\le1\}\), \(\int g=1\), put \(g_\varepsilon(z)=\varepsilon^{-n}g(z/\varepsilon)\), \(B_\varepsilon=\{x\in B:\operatorname{dist}(x,\partial B)>\varepsilon\}\) and \(f_\varepsilon=g_\varepsilon*(f\mathbf 1_B)\), which on \(B_\varepsilon\) integrates only values of \(f\) inside \(B\). Then \(f_\varepsilon\in C^{\infty}\) by Corollary 3.9.5 (\(g_\varepsilon\) is bounded with bounded continuous derivatives of all orders, \(f\mathbf 1_B\in L^1\)), and from \(|f_\varepsilon(x)-f_\varepsilon(y)|\le\int g_\varepsilon(z)|f(x-z)-f(y-z)|dz\), Tonelli’s theorem and the substitution \(x^{\prime}=x-z\), \(y^{\prime}=y-z\) (which preserves \(|x-y|\) and maps \(B_\varepsilon\times B_\varepsilon\) into \(B\times B\)),
\begin{equation*} \int_{B_\varepsilon}\int_{B_\varepsilon} \frac{|f_\varepsilon(x)-f_\varepsilon(y)|}{|x-y|^{n+1}}\,dx\,dy\le D<\infty . \end{equation*}
A smooth \(u\) on an open ball \(U\) with finite double integral is constant. On a concentric ball \(U^{\prime}\) with \(\overline{U^{\prime}}\subset U\), Taylor’s formula gives \(|u(x)-u(y)-\langle\nabla u(y),x-y\rangle|\le C|x-y|^2\), and \(\int_{U^{\prime}}\int_{U^{\prime}}|x-y|^{1-n}dx\,dy<\infty\) by polar coordinates, so the triangle inequality yields
\begin{equation*} \int_{U^{\prime}}\int_{U^{\prime}} \frac{|\langle\nabla u(y),x-y\rangle|}{|x-y|^{n+1}}\,dx\,dy<\infty . \end{equation*}
If \(\nabla u(y_0)\ne0\), continuity gives \(\delta>0\), \(\eta>0\) and a ball \(U^{\prime\prime}\) around \(y_0\) with \(\overline{U^{\prime\prime}}\subset U^{\prime}\), \(|\nabla u|\ge\delta\) on \(U^{\prime\prime}\) and \(B(y,\eta)\subseteq U^{\prime}\) for \(y\in U^{\prime\prime}\); then for \(v=\nabla u(y)\), \(x=y+r\omega\),
\begin{equation*} \int_{U^{\prime}}\frac{|\langle v,x-y\rangle|}{|x-y|^{n+1}}\,dx \ge\Big(\int_{S^{n-1}}|\langle v,\omega\rangle|\,d\omega\Big)\int_0^{\eta}\frac{dr}{r}=\infty , \end{equation*}
since \(\int_{S^{n-1}}|\langle v,\omega\rangle|d\omega=\kappa_n|v|\ge\kappa_n\delta>0\). As \(\lambda_n(U^{\prime\prime})>0\), this contradicts the previous display, so \(\nabla u\equiv0\) and \(u\) is constant on the connected set \(U\).
Hence \(f_\varepsilon\equiv c_\varepsilon\) on the ball \(B_\varepsilon\). For a ball \(B^{\prime\prime}\) with \(\overline{B^{\prime\prime}}\subset B\) we have \(B^{\prime\prime}\subseteq B_\varepsilon\) for small \(\varepsilon\), and \(\|g_\varepsilon*(f\mathbf 1_B)-f\mathbf 1_B\|_{L^1(\mathbb{R}^n)}\to0\) by Theorem 4.2.4, so \(c_\varepsilon\to f\) in \(L^1(B^{\prime\prime})\); since \(\|c_\varepsilon-c_{\varepsilon^{\prime}}\|_{L^1(B^{\prime\prime})}=|c_\varepsilon-c_{\varepsilon^{\prime}}|\lambda_n(B^{\prime\prime})\), the numbers \(c_\varepsilon\) converge to some \(c\) and \(f=c\) a.e. on \(B^{\prime\prime}\). Exhausting \(B\) by such balls, any two of which overlap in a set of positive measure, gives \(f=c\) a.e. on \(B\).
(Kolmogorov) Let \(E\) be a Lebesgue measurable set on the real line. Let \(L(E)\) be the supremum of lengths of the intervals onto which \(E\) can be mapped by means of a nonexpanding (i.e., Lipschitzian with the constant 1) mapping. Show that \(L(E)\) coincides with Lebesgue measure of \(E\).
\(L(E)=\lambda(E)\), the extremal maps being \(f_K(x)=\lambda\big(K\cap(-\infty,x)\big)\) for compact \(K\subseteq E\).
\(L(E)\le\lambda(E)\): a nonexpanding \(\varphi\) does not increase outer measure. Cover \(E\) by intervals \(I_k\) with \(\sum_k|I_k|\le\lambda(E)+\varepsilon\) (nothing to prove if \(\lambda(E)=\infty\)); each \(\varphi(E\cap I_k)\) has diameter at most \(|I_k|\), so countable subadditivity gives \(\lambda^*(\varphi(E))\le\lambda(E)+\varepsilon\). If \(\varphi(E)=I\) is an interval, then \(|I|=\lambda^*(\varphi(E))\le\lambda(E)\); take the supremum.
\(L(E)\ge\lambda(E)\): let \(K\subseteq E\) be compact and nonempty. Then \(f_K\) is nonexpanding, since \(0\le f_K(x)-f_K(y)=\lambda(K\cap[y,x))\le x-y\) for \(y<x\), nondecreasing, with range in \([0,\lambda(K)]\), and \(f_K(K)=[0,\lambda(K)]\): with \(m=\min K\), \(M=\max K\) one has \(f_K(m)=0\) and \(f_K(M)=\lambda(K)\) (as \(\{M\}\) is null), while for \(t\in(0,\lambda(K))\) the least \(a\) with \(f_K(a)=t\) (the level set is nonempty and closed, \(f_K\) being continuous with limits \(0\) and \(\lambda(K)\) at \(\mp\infty\)) satisfies
\begin{equation*} \lambda\big(K\cap(a-\eta,a)\big)=f_K(a)-f_K(a-\eta)>0\qquad\text{for every }\eta>0 , \end{equation*}
so \(a\) is a limit point of the closed set \(K\), i.e. \(a\in K\). Consequently
\begin{equation*} [0,\lambda(K)]=f_K(K)\subseteq f_K(E)\subseteq[0,\lambda(K)] , \end{equation*}
so \(f_K\) maps \(E\) onto an interval of length \(\lambda(K)\) and \(L(E)\ge\lambda(K)\). By inner regularity \(\lambda(E)=\sup\{\lambda(K):K\subseteq E\text{ compact}\}\), which gives \(L(E)\ge\lambda(E)\), also when \(\lambda(E)=\infty\).
The Spaces \(L^p\) and Spaces of Measures
Exercises 4.7.42–4.7.48
(\(\circ\)) Let \(f \in L^p(\mathbb{R}^1)\) and \(f \in L^q(\mathbb{R}^1)\), where \(p \le q\). Prove that \(f \in L^r(\mathbb{R}^1)\) for all \(r \in [p, q]\).
\(f\in L^r(\mathbb{R}^1)\) for every \(r\in[p,q]\), by splitting at the level \(1\). On \(E_0=\{|f|\le1\}\) we have \(|f|^r=|f|^p|f|^{r-p}\le|f|^p\) since \(r\ge p\); on \(E_1=\{|f|>1\}\) we have \(|f|^r=|f|^q|f|^{r-q}\le|f|^q\) since \(r\le q\). Hence, for \(q<\infty\),
\begin{equation*} \begin{aligned} \int_{\mathbb{R}^1}|f|^r\,d\lambda &\le\int_{E_0}|f|^p\,d\lambda+\int_{E_1}|f|^q\,d\lambda\\ &\le\|f\|_{L^p}^p+\|f\|_{L^q}^q<\infty . \end{aligned} \end{equation*}
If \(q=\infty\) and \(r<\infty\), the bound on \(E_0\) is unchanged, while Chebyshev’s inequality (Theorem 2.5.3) applied to \(|f|^p\) gives \(\lambda(E_1)\le\int|f|^pd\lambda<\infty\), whence \(\int_{E_1}|f|^rd\lambda\le\|f\|_{L^{\infty}}^r\lambda(E_1)<\infty\).
Method (2): for \(p<r<q<\infty\) define \(\theta\in(0,1)\) by \(1/r=\theta/p+(1-\theta)/q\), which is possible as \(1/r\) lies between \(1/q\) and \(1/p\). Then \(\alpha=p/(\theta r)\) and \(\beta=q/((1-\theta)r)\) are conjugate, \(1/\alpha+1/\beta=r(\theta/p+(1-\theta)/q)=1\), and Hoelder’s inequality applied to \(|f|^r=|f|^{\theta r}|f|^{(1-\theta)r}\) gives
\begin{equation*} \|f\|_{L^r}\le\|f\|_{L^p}^{\theta}\|f\|_{L^q}^{1-\theta}<\infty . \end{equation*}
(\(\circ\)) Let \(f\) be a bounded measurable function on a space with a nonnegative measure \(\mu\). Prove that
\begin{equation*} \|f\|_{L^\infty(\mu)} = \inf \big\{ a \ge 0 : \ \mu\big( x : |f(x)| > a \big) = 0 \big\} . \end{equation*}
Both numbers equal the least \(a\) with \(|f|\le a\) almost everywhere. Write \(N=\|f\|_{L^\infty(\mu)}\), which by the definition in Section 4.1 is the infimum of \(\sup_\Omega|f|\) over the sets \(\Omega\) with \(\mu(X\setminus\Omega)=0\), and put
\begin{equation*} S=\{a\ge0:\mu(\{|f|>a\})=0\},\qquad M=\inf S ; \end{equation*}
\(S\ne\emptyset\) because \(f\) is bounded, so \(0\le M<\infty\).
(i) \(N\le M\): for \(a\in S\) the set \(\Omega=\{|f|\le a\}\) has null complement and \(\sup_\Omega|f|\le a\), so \(N\le a\).
(ii) \(M\le N\): if \(\mu(X\setminus\Omega)=0\) and \(s=\sup_\Omega|f|\), then \(\{|f|>a\}\subseteq X\setminus\Omega\) for every \(a>s\), so \(\mu(\{|f|>a\})=0\) by monotonicity, i.e. \(a\in S\) and \(M\le a\); let \(a\downarrow s\) and take the infimum over \(\Omega\).
Hence \(N=M\), and this infimum is attained, since \(\{|f|>M\}=\bigcup_{n\ge1}\{|f|>M+n^{-1}\}\) is a countable union of null sets; thus \(|f|\le\|f\|_{L^\infty(\mu)}\) a.e.
(\(\circ\)) Show that \(\|f\|_{L^\infty(\mu)} = \lim\limits_{p \to \infty} \|f\|_{L^p(\mu)}\) if the measure \(\mu\) is bounded and \(f \in L^\infty(\mu)\).
\(\lim_{p\to\infty}\|f\|_{L^p(\mu)}=N:=\|f\|_{L^\infty(\mu)}\), by squeezing between \(N\mu(X)^{1/p}\) and \(t\mu(E_t)^{1/p}\). By Exercise 4.7.43 we have \(|f|\le N\) a.e. and \(\mu(\{|f|>t\})>0\) for every \(t<N\). If \(\mu(X)=0\) both sides vanish, so assume \(\mu(X)>0\).
Upper bound: \(|f|^p\le N^p\) a.e. gives \(f\in L^p(\mu)\) and
\begin{equation*} \|f\|_{L^p(\mu)}\le N\mu(X)^{1/p}\longrightarrow N\qquad(p\to\infty), \end{equation*}
since \(0<\mu(X)<\infty\); hence \(\limsup_{p\to\infty}\|f\|_{L^p(\mu)}\le N\).
Lower bound: for \(0<t<N\) put \(E_t=\{|f|>t\}\) and \(\delta=\mu(E_t)\in(0,\infty)\). Then
\begin{equation*} \|f\|_{L^p(\mu)}^p\ge\int_{E_t}|f|^p\,d\mu\ge t^p\delta , \end{equation*}
so \(\|f\|_{L^p(\mu)}\ge t\delta^{1/p}\to t\) and \(\liminf_{p\to\infty}\|f\|_{L^p(\mu)}\ge t\); let \(t\uparrow N\).
(\(\circ\)) Let \(\mu\) be a probability measure and let \(f\) be a measurable function such that \(\sup_{p \ge 1} \|f\|_{L^p(\mu)} < \infty\). Prove that \(f \in L^\infty(\mu)\).
\(|f|\le C:=\sup_{p\ge1}\|f\|_{L^p(\mu)}\) almost everywhere, so \(f\in L^\infty(\mu)\). Indeed, if \(a>C\) and \(\delta:=\mu(\{|f|>a\})>0\), then Chebyshev’s inequality (Theorem 2.5.3) applied to the integrable \(|f|^p\) gives, for every \(p\ge1\),
\begin{equation*} \delta\le\mu\big(\{|f|^p\ge a^p\}\big)\le\frac1{a^p}\int_X|f|^p\,d\mu \le\Big(\frac{C}{a}\Big)^p\longrightarrow0 \end{equation*}
as \(p\to\infty\), since \(0\le C/a<1\), a contradiction. Hence \(\mu(\{|f|>a\})=0\) for every \(a>C\), and \(\{|f|>C\}=\bigcup_{n\ge1}\{|f|>C+n^{-1}\}\) is null. The function \(fI_{\{|f|\le C\}}\) is everywhere defined, measurable and bounded by \(C\), and coincides with \(f\) outside a null set, so \(f\) determines an element of \(L^\infty(\mu)\) with \(\|f\|_{L^\infty(\mu)}\le C\).
(\(\circ\)) Let \(A \subset \mathbb{R}^1\) be a set of positive Lebesgue measure. Prove that the spaces \(L^p\) on the set \(A\) equipped with Lebesgue measure are infinite-dimensional.
The indicators of a disjoint sequence of subsets of \(A\) of positive measure form an infinite linearly independent family in \(L^p(A)\). Such a sequence exists. Since \(\lambda(A\cap[-N,N])\to\lambda(A)>0\), there is \(N\) with \(m:=\lambda(A\cap J_0)>0\), \(J_0=[-N,N]\). If a bounded interval \(J\) with endpoints \(c<d\) has \(\lambda(A\cap J)=m^{\prime}>0\), then \(t\mapsto\lambda(A\cap J\cap[c,t])\) is \(1\)-Lipschitz and runs from \(0\) to \(m^{\prime}\), so by the intermediate value theorem some \(t_0\in(c,d)\) splits \(J\) into disjoint intervals \(J\cap[c,t_0]\) and \(J\cap(t_0,d]\) each meeting \(A\) in measure \(m^{\prime}/2\). Iterating from \(K_0=J_0\) produces \(K_{n-1}=I_n\cup K_n\) with
\begin{equation*} \lambda(A\cap I_n)=\lambda(A\cap K_n)=2^{-n}m>0,\qquad I_n\cap K_n=\emptyset , \end{equation*}
and the \(I_n\) are pairwise disjoint, since \(I_k\subseteq K_{k-1}\subseteq K_n\) for \(k>n\).
Put \(B_n=A\cap I_n\) and \(f_n=I_{B_n}\). Each \(f_n\) is a nonzero element of \(L^p(A)\), as \(\|f_n\|_{L^p(A)}^p=\lambda(B_n)\in(0,\infty)\) for \(p<\infty\) and \(\|f_n\|_{L^\infty(A)}=1\). If \(\sum_{j=1}^{k}c_jf_{n_j}=0\) a.e. on \(A\), then, the \(B_{n_j}\) being pairwise disjoint, the left-hand side equals \(c_i\) on \(B_{n_i}\), whose measure is positive, so \(c_i=0\) for every \(i\). Hence \(L^p(A)\) contains linearly independent sets of every finite cardinality and is infinite-dimensional, for every \(p\in[1,\infty]\).
(\(\circ\)) Prove the formula for the Legendre polynomials in Example 4.3.7.
The orthogonalization of \(1,x,x^2,\dots\) in \(L^2[-1,1]\) produces exactly the multiples \(c_nP_n\) of \(P_n:=\frac{d^n}{dx^n}(x^2-1)^n\), with \(c_n=\pm\|P_n\|^{-1}\) in the orthonormal case.
Write \(u_n(x)=(x^2-1)^n=(x-1)^n(x+1)^n\), so \(P_n=u_n^{(n)}\) is a polynomial of degree \(n\) with leading coefficient \(2n(2n-1)\cdots(n+1)=(2n)!/n!\ne0\); thus \(P_n\in V_n\setminus V_{n-1}\), where \(V_n=\operatorname{span}(1,x,\dots,x^n)\) has dimension \(n+1\) (a nonzero polynomial has finitely many roots, hence is not \(0\) a.e.).
By the Leibniz rule, every summand of \(u_n^{(j)}\) with \(j\le n-1\) keeps a factor \((x-1)^{n-i}\) with \(n-i\ge1\) and a factor \((x+1)^{n-j+i}\) with \(n-j+i\ge1\), so
\begin{equation*} u_n^{(j)}(1)=u_n^{(j)}(-1)=0,\qquad j=0,1,\dots,n-1 . \end{equation*}
Hence, integrating by parts \(k+1\) times for \(0\le k\le n-1\) (each boundary term involves some \(u_n^{(n-i)}\) with \(0\le n-i\le n-1\) and so vanishes),
\begin{equation*} \int_{-1}^{1}x^kP_n\,dx=(-1)^kk!\int_{-1}^{1}u_n^{(n-k)}dx =(-1)^kk!\Big[u_n^{(n-k-1)}\Big]_{-1}^{1}=0 , \end{equation*}
i.e. \(P_n\perp V_{n-1}\); in particular \((P_n,P_m)=0\) for \(m<n\), since \(P_m\in V_m\subseteq V_{n-1}\).
Gram–Schmidt applied to \(1,x,x^2,\dots\) (proof of Corollary 4.3.4) yields the unit vector \(e_n\in V_n\cap V_{n-1}^{\perp}\). By Proposition 4.3.1 inside the finite-dimensional \(V_n\) one has \(V_n=V_{n-1}\oplus W_n\) with \(W_n=V_{n-1}^{\perp}\cap V_n\) of dimension \((n+1)-n=1\), so \(W_n\) contains exactly the two unit vectors \(\pm P_n/\|P_n\|\) and \(e_n=c_nP_n\), the sign being fixed by the positivity of the leading coefficient that Gram–Schmidt delivers.
The constants: integrating by parts \(n\) times (all boundary terms vanish as above), using \(u_n^{(2n)}=(2n)!\) and the Beta integral \(\int_{-1}^{1}(1-x)^n(1+x)^ndx=2^{2n+1}(n!)^2/(2n+1)!\),
\begin{equation*} \|P_n\|_{L^2[-1,1]}^2=(-1)^n(2n)!\int_{-1}^{1}(x^2-1)^ndx=\frac{2^{2n+1}(n!)^2}{2n+1} , \end{equation*}
so the orthonormalized system is \(e_n(x)=\frac1{2^nn!}\sqrt{\frac{2n+1}{2}}\frac{d^n}{dx^n}(x^2-1)^n\). With Rodrigues’ normalization \(c_n=1/(2^nn!)\) the Leibniz rule at \(x=1\) leaves only \(u_n^{(n)}(1)=n!\,2^n\), i.e. \(L_n(1)=1\) and \(L_0=1\), as in Example 4.3.7.
(\(\circ\)) Prove that the functions \(\sqrt{2/\pi} \sin nt\), \(n \in \mathbb{N}\), form an orthonormal basis in \(L^2[0, \pi]\). Prove the same for the functions \(\sqrt{1/\pi}\), \(\sqrt{2/\pi} \cos nt\), \(n \in \mathbb{N}\).
Both systems are orthonormal bases, completeness following from the trigonometric basis of \(L^2[-\pi,\pi]\) by odd, resp. even, extension. All spaces are real, with \((u,v)=\int uv\).
The system \((2\pi)^{-1/2}\), \(\pi^{-1/2}\cos nx\), \(\pi^{-1/2}\sin nx\) is an orthonormal basis of \(L^2[0,2\pi]\) by Example 4.3.7(i), hence also of \(L^2[-\pi,\pi]\): the map \(\Phi h=h\) on \([0,\pi]\), \(\Phi h=h(\cdot-2\pi)\) on \((\pi,2\pi]\) is unitary by translation invariance of Lebesgue measure and, the system being \(2\pi\)-periodic, carries it onto itself. So an \(h\in L^2[-\pi,\pi]\) orthogonal to \(1\) and to all \(\cos nt\), \(\sin nt\) is zero a.e.
Orthonormality on \([0,\pi]\): since \(\int_0^{\pi}\cos kt\,dt=\sin(k\pi)/k=0\) for every nonzero integer \(k\),
\begin{equation*} \int_0^{\pi}\sin nt\sin mt\,dt =\tfrac12\int_0^{\pi}\big[\cos(n-m)t-\cos(n+m)t\big]dt=\tfrac{\pi}2\delta_{nm}, \end{equation*}
and likewise \(\int_0^{\pi}\cos nt\cos mt\,dt=\frac{\pi}2\delta_{nm}\) for \(n,m\ge1\), while \(\int_0^\pi\psi_0^2=1\) and \(\psi_0\perp\psi_n\) for \(\psi_0=\pi^{-1/2}\). Hence \(\varphi_n=\sqrt{2/\pi}\sin nt\) and \(\psi_0,\ \psi_n=\sqrt{2/\pi}\cos nt\) are orthonormal systems.
An orthonormal system with trivial orthogonal complement is a basis: \(L^2[0,\pi]=E\oplus E^{\perp}\) for the closed span \(E\) by Proposition 4.3.1, and if \(E^{\perp}=\{0\}\), then for \(f\in L^2[0,\pi]\) Bessel’s inequality (Theorem 4.3.6) and the Riesz–Fischer theorem 4.3.5 make \(s=\sum_n(f,\varphi_n)\varphi_n\) converge, while \(f-s\perp\varphi_m\) for every \(m\) forces \(f=s\). So it suffices to rule out a nonzero orthogonal \(g\).
(i) Sine system: let \(\int_0^{\pi}g\sin nt\,dt=0\) for all \(n\ge1\) and let \(h\) be the odd extension of \(g\) to \([-\pi,\pi]\), which is measurable with \(\int_{-\pi}^{\pi}|h|^2=2\int_0^{\pi}|g|^2<\infty\). Substituting \(s=-t\) on \([-\pi,0]\),
\begin{equation*} \int_{-\pi}^{\pi}h\sin nt\,dt=2\int_0^{\pi}g\sin nt\,dt=0,\qquad \int_{-\pi}^{\pi}h\cos nt\,dt=0\quad(n\ge0), \end{equation*}
so \(h=0\) a.e. and \(g=0\) a.e.
(ii) Cosine system: let \(\int_0^{\pi}g\,dt=0\) and \(\int_0^{\pi}g\cos nt\,dt=0\) for all \(n\ge1\), and let \(h\) be the even extension. The same substitution gives \(\int_{-\pi}^{\pi}h\sin nt\,dt=0\) (odd integrand) and \(\int_{-\pi}^{\pi}h\cos nt\,dt=2\int_0^{\pi}g\cos nt\,dt=0\) for \(n\ge0\), so again \(h=0\) a.e. and \(g=0\) a.e.
Exercises 4.7.49–4.7.55
(\(\circ\)) Let \(\mu\) be the measure on \((0,+\infty)\) with density \(e^{-x}\) with respect to Lebesgue measure. Prove that the Laguerre polynomials obtained by the orthogonalization of the functions \(1, x, x^2, \ldots\), form an orthonormal basis in \(L^2(\mu)\).
The Laguerre system is complete because a \(g\in L^2(\mu)\) orthogonal to all powers makes the Fourier transform of \(ge^{-x}\) vanish. Each \(x^n\) lies in \(L^2(\mu)\), since \(\int_0^{\infty}x^{2n}e^{-x}dx=(2n)!\), and the powers are linearly independent there (a nonzero polynomial has finitely many roots), so Gram–Schmidt produces an orthonormal \(\{\ell_n\}\) with \(\operatorname{span}\{\ell_0,\dots,\ell_n\}=\operatorname{span}\{1,\dots,x^n\}\); it is a basis precisely when the only \(g\in L^2(\mu)\) with \(\int_0^{\infty}g(x)x^ne^{-x}dx=0\) for all \(n\ge0\) is \(g=0\).
Let \(g\) be such a function. For \(c>1/2\), the Cauchy–Bunyakowsky inequality applied to \(|g|e^{-x/2}\cdot e^{-(c-1/2)x}\) gives
\begin{equation*} \int_0^{\infty}|g(x)|e^{-cx}\,dx\le\frac{\|g\|_{L^2(\mu)}}{\sqrt{2c-1}}<\infty . \end{equation*}
Hence
\begin{equation*} G(z):=\int_0^{\infty}g(x)e^{-x}e^{izx}\,dx,\qquad z\in S:=\{|\operatorname{Im}z|<1/2\}, \end{equation*}
converges absolutely, because \(|g(x)e^{-x}e^{izx}|=|g(x)|e^{-(1+\operatorname{Im}z)x}\) with \(1+\operatorname{Im}z>1/2\). It is continuous on \(S\) by dominated convergence (near \(z_0\) the integrands are dominated by \(|g|e^{-(1-\eta)x}\) with \(|\operatorname{Im}z_0|<\eta<1/2\)) and holomorphic: for a closed triangle \(T\subset S\) the double integral is absolutely convergent, so Fubini’s theorem gives
\begin{equation*} \oint_{\partial T}G(z)\,dz =\int_0^{\infty}g(x)e^{-x}\Big(\oint_{\partial T}e^{izx}dz\Big)dx=0 , \end{equation*}
and Morera’s theorem applies. For \(|z|<1/2\) the displayed bound with \(c=1-|z|\) justifies term-by-term integration of \(e^{izx}=\sum_n(iz)^nx^n/n!\), whence
\begin{equation*} G(z)=\sum_{n=0}^{\infty}\frac{(iz)^n}{n!}\int_0^{\infty}g(x)x^ne^{-x}\,dx=0 , \end{equation*}
and by the uniqueness theorem \(G\equiv0\) on the connected set \(S\), in particular on \(\mathbb{R}\).
Now \(h:=ge^{-x}\mathbf 1_{(0,\infty)}\) lies in \(L^1(\mathbb{R})\) (the bound with \(c=1\)) and has \(\widehat h(y)=G(-y)\equiv0\), so \(h=0\) a.e. by Proposition 3.8.6, i.e. \(g=0\) in \(L^2(\mu)\).
(\(\circ\)) (i) Let \((X,\mathcal{A},\mu)\) and \((Y,\mathcal{B},\nu)\) be two probability spaces. Suppose that for some \(p \in [1,+\infty)\) sets \(F \subset L^p(\mu)\) and \(G \subset L^p(\nu)\) are everywhere dense. Show that the set of linear combinations of products \(fg\), where \(f \in F\), \(g \in G\), is everywhere dense in \(L^p(\mu\otimes\nu)\). Prove that if \(\{f_n\}\) and \(\{g_n\}\) are orthonormal bases in \(L^2(\mu)\) and \(L^2(\nu)\), respectively, then \(\{f_n g_k\}\) is an orthonormal basis in \(L^2(\mu\otimes\nu)\).
(ii) Let \((X_\alpha,\mathcal{A}_\alpha,\mu_\alpha)\) be a family of probability spaces. Suppose that for some \(p \in [1,+\infty)\) and every \(\alpha\), we are given an everywhere dense set \(F_\alpha \subset L^p(\mu_\alpha)\). Show that the set of linear combinations of products \(f_{\alpha_1}\cdots f_{\alpha_n}\), where \(f_{\alpha_i} \in F_{\alpha_i}\), is everywhere dense in \(L^p\bigl(\bigotimes_\alpha \mu_\alpha\bigr)\). Deduce that if, for every \(\alpha\), we have an orthonormal basis \(\{f_{\alpha,\beta}\}\) in \(L^2(\mu_\alpha)\), then the elements \(f_{\alpha_1,\beta_1}\cdots f_{\alpha_n,\beta_n}\), where the indices \(\alpha_i\) are distinct, form an orthonormal basis in \(L^2\bigl(\bigotimes_\alpha \mu_\alpha\bigr)\).
(i) Dense sets multiply to a dense set. Write \((f\otimes g)(x,y)=f(x)g(y)\); Tonelli’s theorem (Theorem 3.4.5) applied to \(|f\otimes g|^p=|f|^p\otimes|g|^p\) gives
\begin{equation*} \|f\otimes g\|_{L^p(\mu\otimes\nu)}=\|f\|_{L^p(\mu)}\|g\|_{L^p(\nu)} , \end{equation*}
and the closure \(\overline{\mathcal{L}}\) of the span \(\mathcal{L}\) of the products \(f\otimes g\), \(f\in F\), \(g\in G\), is a closed subspace, so it suffices to place a dense family inside it.
Indicators of rectangles have dense span: simple functions are dense in \(L^p(\lambda)\), \(\lambda:=\mu\otimes\nu\), by Lemma 4.2.1, and \(\lambda\) is the unique countably additive extension of its restriction to the algebra \(\mathcal{R}\) of finite disjoint unions of measurable rectangles (construction of Section 3.3 and Theorem 1.5.6), so by Definition 1.5.1 and Theorem 1.5.6 every \(E\) in the (possibly completed) domain admits \(R\in\mathcal{R}\) with \(\lambda(E\bigtriangleup R)<\varepsilon\), i.e. \(\|I_E-I_R\|_{L^p}^p<\varepsilon\), and \(I_R\) is a finite sum of functions \(I_A\otimes I_B\).
Each \(I_A\otimes I_B\) lies in \(\overline{\mathcal{L}}\): choosing \(f_n\to I_A\) in \(L^p(\mu)\) and \(g_n\to I_B\) in \(L^p(\nu)\) and splitting
\begin{equation*} f_n\otimes g_n-I_A\otimes I_B=(f_n-I_A)\otimes g_n+I_A\otimes(g_n-I_B) , \end{equation*}
the norm identity gives
\begin{equation*} \|f_n\otimes g_n-I_A\otimes I_B\|_{L^p(\lambda)} \le\|f_n-I_A\|_{L^p(\mu)}\|g_n\|_{L^p(\nu)}+\mu(A)^{1/p}\|g_n-I_B\|_{L^p(\nu)} , \end{equation*}
which tends to \(0\) since \(\|g_n\|_{L^p(\nu)}\) is bounded. Hence \(\overline{\mathcal{L}}=L^p(\mu\otimes\nu)\).
For orthonormal bases \(\{f_n\}\), \(\{g_k\}\), Tonelli and Fubini (Theorems 3.4.4, 3.4.5) give \(\langle f_n\otimes g_k,f_m\otimes g_l\rangle=\delta_{nm}\delta_{kl}\), and the first part with \(p=2\), \(F=\operatorname{span}\{f_n\}\), \(G=\operatorname{span}\{g_k\}\) makes \(\operatorname{span}\{f_n\otimes g_k\}\) dense; an orthonormal system with dense span is a basis.
(ii) The same two steps, with cylinders in place of rectangles. Splitting \(\mu=\bigotimes_\alpha\mu_\alpha\) as \((\mu_{\alpha_1}\otimes\cdots\otimes\mu_{\alpha_n})\otimes\mu^{\prime}\) with \(\mu^{\prime}\) a probability measure (Section 3.5) and applying Tonelli to \(\prod_i|h_i|^p\),
\begin{equation*} \Big\|\prod_{i=1}^{n}h_i(x_{\alpha_i})\Big\|_{L^p(\mu)} =\prod_{i=1}^{n}\|h_i\|_{L^p(\mu_{\alpha_i})} \qquad(\alpha_i\text{ distinct}). \end{equation*}
The cylindrical sets form the algebra \(\mathcal{R}\) with \(\sigma(\mathcal{R})=\bigotimes_\alpha\mathcal{A}_\alpha\), and \(\mu\) is the unique countably additive extension from \(\mathcal{R}\) (Theorem 3.5.1), completed as in Definition 1.5.1; so, exactly as in (i), every \(E\) in the domain satisfies \(\mu(E\bigtriangleup R)<\varepsilon\) for some \(R\in\mathcal{R}\), and with Lemma 4.2.1 the span of the indicators \(I_C=\prod_{i}I_{A_i}(x_{\alpha_i})\) of cylinders is dense in \(L^p(\mu)\). Approximating \(v_i=I_{A_i}\) by \(u_i^{(m)}\in F_{\alpha_i}\) and telescoping,
\begin{equation*} \prod_{i=1}^{n}u_i-\prod_{i=1}^{n}v_i =\sum_{i=1}^{n}v_1\cdots v_{i-1}(u_i-v_i)u_{i+1}\cdots u_n , \end{equation*}
the norm factorization and \(\|v_j\|_{L^p}\le1\) give \(\|\prod_iu_i^{(m)}-I_C\|_{L^p(\mu)}\to0\) (the norms \(\|u_j^{(m)}\|\) being bounded in \(m\)), so \(I_C\in\overline{\mathcal{L}}\) and \(\mathcal{L}\) is dense in \(L^p(\mu)\).
The basis statement needs the normalization, implicit in the printed formulation, that each basis \(\{f_{\alpha,\beta}\}_\beta\) contains the constant \(\mathbf 1\) and that only factors \(\ne\mathbf 1\) occur in the products, \(\mathbf 1\) itself (the empty product) belonging to the system: otherwise \(f_{\alpha_1,\beta_1}\) and \(f_{\alpha_1,\beta_1}f_{\alpha_2,\beta_2}\) would have inner product \(\int f_{\alpha_2,\beta_2}d\mu_{\alpha_2}\ne0\). Then every factor has zero mean, and for \(P=\prod_{\alpha\in S}f_{\alpha,\beta_\alpha}\), \(Q=\prod_{\alpha\in T}f_{\alpha,\gamma_\alpha}\) with \(S,T\) finite, splitting the product measure over \(S\cup T\) and its complement and using Tonelli and Fubini,
\begin{equation*} \begin{aligned} \langle P,Q\rangle&=\prod_{\alpha\in S\cap T} \langle f_{\alpha,\beta_\alpha},f_{\alpha,\gamma_\alpha}\rangle \prod_{\alpha\in S\setminus T}\Big(\int f_{\alpha,\beta_\alpha}d\mu_\alpha\Big) \prod_{\alpha\in T\setminus S}\Big(\int f_{\alpha,\gamma_\alpha}d\mu_\alpha\Big)\\ &=\begin{cases}\prod_{\alpha\in S}\delta_{\beta_\alpha\gamma_\alpha},&S=T,\\ 0,&S\ne T,\end{cases} \end{aligned} \end{equation*}
so the system is orthonormal. Completeness: apply the first part of (ii) with \(p=2\) and \(F_\alpha=\operatorname{span}\{f_{\alpha,\beta}\}_\beta\); by multilinearity each product \(f_{\alpha_1}\cdots f_{\alpha_n}\) with \(f_{\alpha_i}\in F_{\alpha_i}\) is a finite combination of products of basis elements over distinct indices, and deleting the factors \(\mathbf 1\) turns each into a member of the system, so the span of the system contains \(\mathcal{L}\) and is dense.
(\(\circ\)) Prove that if a series is Cesaro summable to a number \(s\), then it is summable to \(s\) in the sense of Abel (see Section 4.3).
Abel summation twice turns \(S( r)=\sum_{n\ge1}\alpha_nr^n\) into \((1-r)^2\sum_{n\ge1}n\sigma_nr^n\), where the Cesaro limit passes to the limit. Put \(s_n=\sum_{k\le n}\alpha_k\), \(\sigma_n=(s_1+\cdots+s_n)/n\), \(S_n=s_1+\cdots+s_n=n\sigma_n\) and \(s_0=S_0=0\), and assume \(\sigma_n\to s\).
Convergence: \(|\sigma_n|\le C\) for all \(n\) gives \(|S_n|\le Cn\), hence \(|s_n|=|S_n-S_{n-1}|\le2Cn\) and \(|\alpha_n|\le4Cn\), so \(\sum_n|\alpha_n|r^n\), \(\sum_n|s_n|r^n\) and \(\sum_n|S_n|r^n\) all converge for \(0<r<1\) and the rearrangements below are legitimate. Using \(\alpha_n=s_n-s_{n-1}\) and then \(s_n=S_n-S_{n-1}\),
\begin{equation*} S( r)=(1-r)\sum_{n\ge1}s_nr^n=(1-r)^2\sum_{n\ge1}S_nr^n=(1-r)^2\sum_{n\ge1}n\sigma_nr^n . \end{equation*}
Since \((1-r)^2\sum_{n\ge1}nr^n=r\), subtracting \(rs\) gives \(S( r)-rs=(1-r)^2\sum_{n\ge1}n(\sigma_n-s)r^n\). Given \(\varepsilon>0\), choose \(N\) with \(|\sigma_n-s|<\varepsilon\) for \(n>N\) and put \(M_N=\sum_{n\le N}n|\sigma_n-s|<\infty\); then
\begin{equation*} |S( r)-rs|\le(1-r)^2M_N+\varepsilon(1-r)^2\frac{r}{(1-r)^2}\le(1-r)^2M_N+\varepsilon , \end{equation*}
so \(|S( r)-s|\le|S( r)-rs|+(1-r)|s|\) has \(\limsup_{r\to1^-}|S( r)-s|\le\varepsilon\). As \(\varepsilon\) was arbitrary, \(S( r)\to s\).
Let \(\{\varphi_n\}\) be an orthonormal basis in \(L^2[0,1]\).
(i) Prove that there exist numbers \(c_n\), \(n \ge 2\), such that the sums \(\sum_{n=2}^{N} c_n\varphi_n(x)\) converge to \(\varphi_1\) in measure.
(ii) Prove that, for every \(\varepsilon > 0\), there exists a set \(E_\varepsilon\) with measure greater than \(1-\varepsilon\) such that the linear span of the functions \(\varphi_n\), \(n \ge 2\), is everywhere dense in \(L^2(E_\varepsilon)\), where \(E_\varepsilon\) is equipped with Lebesgue measure.
(iii) Prove that there exists a positive bounded measurable function \(\theta\) such that the linear span of the functions \(\theta\varphi_n\), \(n \ge 2\), is everywhere dense in the space \(L^2[0,1]\).
(i) The coefficients are built block by block from the density in \(L^0[0,1]\) of \(\mathcal{E}_k:=\operatorname{span}\{\varphi_n:n\ge k\}\); this gives a subsequence of partial sums converging to \(\varphi_1\) in measure, the full assertion being Talalyan’s theorem. Here \(L^0=L^0[0,1]\) carries the complete translation invariant metric
\begin{equation*} d(f,g)=\int_0^1\frac{|f-g|}{1+|f-g|}\,dx \end{equation*}
inducing convergence in measure (Exercise 4.7.60); \(L^0\) is a topological vector space, since \(t_jf_j-tf=t_j(f_j-f)+(t_j-t)f\to0\) in measure for bounded \(t_j\to t\) and \(f\) finite a.e., and \(d(h,0)\le\|h\|_{L^2}\), so \(L^2\)-convergence implies convergence in measure.
Claim: \(\mathcal{E}_k\) is dense in \(L^0\) for every \(k\ge1\). First, \(\operatorname{span}\{\varphi_n:n\ge1\}\) is dense in \(L^0\): the truncations \(hI_{\{|h|\le m\}}\) converge to \(h\) in measure and lie in \(L^2[0,1]\), where \(\{\varphi_n\}\) is a basis. Let \(F=\overline{\mathcal{E}_k}\), a closed linear subspace.
Lemma A. If \(F\subseteq L^0\) is a closed subspace and \(v_1,\dots,v_m\) are linearly independent modulo \(F\), then \(Y=F+\operatorname{span}\{v_1,\dots,v_m\}\) is closed and the coordinate functionals \(L_i(f+\sum_jt_jv_j)=t_i\) are continuous on \(Y\). Indeed, let \(y_j=f_j+\sum_it_{i,j}v_i\to y\) and \(T_j=\max_i|t_{i,j}|\). If \(T_j\to\infty\) along a subsequence, then \(y_j/T_j\to0\) and, after passing to a further subsequence with \(s_{i,j}:=t_{i,j}/T_j\to s_i\), \(\max_i|s_i|=1\),
\begin{equation*} F\ni\frac{f_j}{T_j}=\frac{y_j}{T_j}-\sum_is_{i,j}v_i\longrightarrow-\sum_is_iv_i\in F , \end{equation*}
contradicting independence modulo \(F\). So \((T_j)\) is bounded; every subsequence has a further one along which \(t_{i,j}\to t_i\) and \(f_j\to y-\sum_it_iv_i\in F\), whence \(y\in Y\), and by uniqueness of the representation \(t_{i,j}\to L_i(y)\) for the full sequence.
Take a maximal subfamily \(v_1,\dots,v_m\) of \(\{\varphi_1,\dots,\varphi_{k-1}\}\) independent modulo \(F\) and put \(Y=F+\operatorname{span}\{v_i\}\). Every \(\varphi_j\) with \(j<k\) lies in \(Y\): maximality gives a nontrivial relation \(t_0\varphi_j+\sum_it_iv_i\in F\) in which \(t_0\ne0\). Since also \(\varphi_n\in F\) for \(n\ge k\), the closed set \(Y\) contains a dense set, i.e. \(Y=L^0\). If \(m\ge1\), then \(L_1\) would be a nonzero continuous linear function on \(L^0\), impossible by Exercise 4.7.61 (Nikodym). Hence \(m=0\) and \(F=L^0\).
Blocks: put \(N_0=1\), \(S_N=\sum_{n=2}^{N}c_n\varphi_n\) and, given \(c_2,\dots,c_{N_{j-1}}\), use the density of \(\mathcal{E}_{N_{j-1}+1}\) to choose \(N_j>N_{j-1}\) and coefficients \(c_n\) for \(N_{j-1}<n\le N_j\) (the unused ones set to \(0\)) with
\begin{equation*} d\Big(\sum_{n=N_{j-1}+1}^{N_j}c_n\varphi_n,\ \varphi_1-S_{N_{j-1}}\Big)<2^{-j} , \end{equation*}
i.e. \(d(S_{N_j},\varphi_1)<2^{-j}\) by translation invariance. Thus \(S_{N_j}\to\varphi_1\) in measure.
The intermediate partial sums are not controlled by this construction, and no \(L^2\) bound can control them: if \(S_N\to\varphi_1\) in measure and \(\sup_i\|S_{M_i}\|_{L^2}=C<\infty\) along a subsequence, then \(S_{M_i}\to\varphi_1\) a.e. after a further passage (Riesz), and in fact weakly in \(L^2\): for \(g\in L^2\) and \(E\) from Egorov with \(\lambda([0,1]\setminus E)<\delta\) one has \(\int_E(S_{M_i}-\varphi_1)g\to0\) by uniform convergence, while
\begin{equation*} \Big|\int_{[0,1]\setminus E}(S_{M_i}-\varphi_1)g\,dx\Big| \le(C+1)\big\|gI_{[0,1]\setminus E}\big\|_{L^2} \end{equation*}
is small by absolute continuity of \(\int g^2\), so \(\langle S_{M_i},\varphi_1\rangle\to1\), contradicting \(\langle S_N,\varphi_1\rangle=0\) for every \(N\). Hence \(\|S_N\|_{L^2}\to\infty\) for any admissible coefficients, and the full statement (A. A. Talalyan’s theorem on the representation of measurable functions by series convergent in measure in a complete orthonormal system with finitely many elements deleted) needs a genuinely finer construction. Only the subsequence statement is used in (ii), and (iii) is independent of (i). (A fully rigorous proof of the remaining step is beyond the scope of this page; see the reference given in the book.)
(ii) Take \(E_\varepsilon\) from Egorov’s theorem. By (i) there are \(u_j\in\operatorname{span}\{\varphi_n:n\ge2\}\) with \(u_j\to\varphi_1\) in measure, hence a.e. along a subsequence by the Riesz theorem (Theorem 2.2.5(i)), again denoted \(u_j\); Egorov’s theorem (Theorem 2.2.1) gives a measurable \(E_\varepsilon\) with \(\lambda(E_\varepsilon)>1-\varepsilon\) and \(u_j\to\varphi_1\) uniformly on \(E_\varepsilon\). If \(g\in L^2(E_\varepsilon)\) is orthogonal to every \(\varphi_n|_{E_\varepsilon}\), \(n\ge2\), extend it by \(0\) to \([0,1]\); then \(\int_{E_\varepsilon}gu_j\,dx=0\) for every \(j\) by linearity, and \(g\in L^1(E_\varepsilon)\), so
\begin{equation*} \Big|\int_{E_\varepsilon}gu_j\,dx-\int_{E_\varepsilon}g\varphi_1\,dx\Big| \le\sup_{E_\varepsilon}|u_j-\varphi_1|\int_{E_\varepsilon}|g|\,dx\longrightarrow0 , \end{equation*}
whence \(\int_0^1g\varphi_1\,dx=0\) as well. Thus \(g\) is orthogonal to the whole basis \(\{\varphi_n\}_{n\ge1}\) and \(g=0\) a.e.; the span of the restrictions has trivial orthogonal complement in \(L^2(E_\varepsilon)\), i.e. it is dense.
(iii) Take \(\theta\) positive and bounded with \(\varphi_1/\theta\notin L^2[0,1]\). Since \(\|\varphi_1\|_{L^2}=1\), the set \(B=\{|\varphi_1|>1/m\}\) has \(\lambda(B)>0\) for some \(m\in\mathbb{N}\); Lebesgue measure being atomless, partition \(B\) into measurable \(B_1,B_2,\dots\) with \(\lambda(B_k)>0\) and put \(\theta=\min(1,m^{-1}\lambda(B_k)^{1/2})\) on \(B_k\) and \(\theta=1\) off \(B\). Then \(0<\theta\le1\), and on \(B_k\) one has \(\theta^{-2}\ge m^2/\lambda(B_k)\) and \(\varphi_1^2>m^{-2}\), so
\begin{equation*} \int_{B_k}\frac{\varphi_1^2}{\theta^2}\,dx \ge\frac1{m^2}\cdot\frac{m^2}{\lambda(B_k)}\cdot\lambda(B_k)=1 , \end{equation*}
and summing over \(k\) gives \(\int_0^1\varphi_1^2/\theta^2dx=\infty\). Now let \(g\in L^2[0,1]\) be orthogonal to every \(\theta\varphi_n\), \(n\ge2\). Then \(g\theta\in L^2[0,1]\) satisfies \(\langle g\theta,\varphi_n\rangle=0\) for \(n\ge2\), so its expansion in the basis \(\{\varphi_n\}\) is \(g\theta=c\varphi_1\); \(c\ne0\) would give \(\varphi_1/\theta=g/c\in L^2[0,1]\), so \(c=0\), \(g\theta=0\) a.e. and, \(\theta\) being everywhere positive, \(g=0\) a.e. Hence \(\operatorname{span}\{\theta\varphi_n:n\ge2\}\) is dense in \(L^2[0,1]\).
(\(\circ\)) Let \(\sum_{n=1}^{\infty}\alpha_n^2 = \infty\). Prove that there exist numbers \(\beta_n\) such that \(\sum_{n=1}^{\infty}\beta_n^2 < \infty\) and \(\sum_{n=1}^{\infty}\alpha_n\beta_n = \infty\).
Take \(\beta_n=\alpha_n/A_n\) for \(n\ge n_0\) and \(\beta_n=0\) for \(n<n_0\), where \(a_n=\alpha_n^2\), \(A_n=a_1+\cdots+a_n\) and \(n_0\) is the least index with \(\alpha_{n_0}\ne0\); then \(A_n\) is nondecreasing, \(A_n\to\infty\) and \(A_n\ge A_{n_0}>0\) for \(n\ge n_0\). For \(n>n_0\) one has \(0<A_{n-1}\le A_n\), hence
\begin{equation*} \frac{a_n}{A_n^2}\le\frac{a_n}{A_{n-1}A_n} =\frac{A_n-A_{n-1}}{A_{n-1}A_n}=\frac1{A_{n-1}}-\frac1{A_n} , \end{equation*}
so telescoping and adding the term \(a_{n_0}/A_{n_0}^2=1/A_{n_0}\),
\begin{equation*} \sum_{n=1}^{\infty}\beta_n^2=\sum_{n\ge n_0}\frac{a_n}{A_n^2}\le\frac2{A_{n_0}}<\infty . \end{equation*}
On the other hand, given \(m\ge n_0\) choose \(k>m\) with \(A_k\ge2A_m\) (possible as \(A_k\to\infty\)); since \(A_n\le A_k\) for \(n\le k\),
\begin{equation*} \sum_{n=m+1}^{k}\alpha_n\beta_n=\sum_{n=m+1}^{k}\frac{a_n}{A_n} \ge\frac{A_k-A_m}{A_k}\ge\frac12 , \end{equation*}
so the partial sums are not fundamental and, all terms being nonnegative, \(\sum_n\alpha_n\beta_n=\infty\).
(\(\circ\)) Let \(\alpha_n \ge 0\) and \(\sum_{n=1}^{\infty}\alpha_n = \infty\). Prove that there exist numbers \(c_n \ge 0\) such that \(\sum_{n=1}^{\infty}\alpha_n c_n = \infty\) and \(\sum_{n=1}^{\infty}\alpha_n c_n^2 < \infty\).
Take \(c_n=1/A_n\) for \(n\ge n_0\) and \(c_n=0\) for \(n<n_0\), where \(A_n=\alpha_1+\cdots+\alpha_n\) and \(n_0\) is the least index with \(\alpha_{n_0}>0\); then \(c_n\ge0\), \(A_n\) is nondecreasing, \(A_n\to\infty\) and \(A_n\ge A_{n_0}>0\) for \(n\ge n_0\). For \(n>n_0\),
\begin{equation*} \frac{\alpha_n}{A_n^2}\le\frac{\alpha_n}{A_{n-1}A_n} =\frac{A_n-A_{n-1}}{A_{n-1}A_n}=\frac1{A_{n-1}}-\frac1{A_n} , \end{equation*}
so telescoping and adding the term \(\alpha_{n_0}/A_{n_0}^2=1/A_{n_0}\),
\begin{equation*} \sum_{n=1}^{\infty}\alpha_nc_n^2 =\sum_{n\ge n_0}\frac{\alpha_n}{A_n^2}\le\frac2{A_{n_0}}<\infty . \end{equation*}
On the other hand, given \(m\ge n_0\) choose \(k>m\) with \(A_k\ge2A_m\); since \(A_n\le A_k\) for \(n\le k\),
\begin{equation*} \sum_{n=m+1}^{k}\frac{\alpha_n}{A_n} \ge\frac1{A_k}\sum_{n=m+1}^{k}\alpha_n=\frac{A_k-A_m}{A_k}\ge\frac12 , \end{equation*}
so the partial sums are not fundamental and \(\sum_n\alpha_nc_n=\infty\).
(\(\circ\)) Let \(A \subset \mathbb{R}^1\) be a set of infinite Lebesgue measure. Prove that there exists a function \(f \in L^2(\mathbb{R}^1)\) that is not integrable on \(A\).
Take \(f=c_n\) on \(A\cap[n,n+1)\) and \(f=0\) off \(A\), the \(c_n\ge0\) coming from Exercise 4.7.54 applied to \(\alpha_n=\lambda(A\cap[n,n+1))\), \(n\in\mathbb{Z}\). The intervals \([n,n+1)\) partition \(\mathbb{R}\), so countable additivity gives \(\sum_n\alpha_n=\lambda(A)=\infty\), and all terms being nonnegative this is independent of the enumeration of \(\mathbb{Z}\) by \(\mathbb{N}\); Exercise 4.7.54 then provides \(c_n\ge0\) with \(\sum_n\alpha_nc_n=\infty\) and \(\sum_n\alpha_nc_n^2<\infty\). The function \(f\ge0\) is measurable, being the pointwise sum of the functions \(c_nI_{A\cap[n,n+1)}\) with pairwise disjoint supports, and term-by-term integration of series of nonnegative measurable functions gives
\begin{equation*} \int_{\mathbb{R}}f^2\,dx=\sum_{n\in\mathbb{Z}}\alpha_nc_n^2<\infty,\qquad \int_A|f|\,dx=\sum_{n\in\mathbb{Z}}\alpha_nc_n=\infty , \end{equation*}
i.e. \(f\in L^2(\mathbb{R}^1)\) is not integrable on \(A\).
Exercises 4.7.56–4.7.62
(\(\circ\)) Let \(f \in L^1(\mathbb{R})\), \(f > 0\). Prove that \(1/f \notin L^1(\mathbb{R})\).
No: if \(1/f\) were integrable, then \(f^{1/2}\) and \(f^{-1/2}\) would both lie in \(L^2(\mathbb{R})\), and the Cauchy–Bunyakowsky inequality applied to \(1=f^{1/2}f^{-1/2}\) (valid a.e., since an integrable \(f\) is finite a.e. and \(f>0\)) would give, for every measurable \(E\) of finite measure,
\begin{equation*} \lambda(E)=\int_Ef^{1/2}f^{-1/2}dx \le\Big(\int_{\mathbb{R}}f\,dx\Big)^{1/2}\Big(\int_{\mathbb{R}}\frac{dx}{f}\Big)^{1/2}=:C<\infty , \end{equation*}
both integrands being nonnegative. Taking \(E=[-n,n]\) yields \(2n\le C\) for every \(n\in\mathbb{N}\), which is impossible. Hence \(1/f\notin L^1(\mathbb{R})\).
(\(\circ\)) Prove that the set of nonnegative functions is closed and nowhere dense in the space \(L^1[0,1]\).
The set \(P=\{f\in L^1[0,1]:f\ge0\ \text{a.e.}\}\) is closed with empty interior, hence nowhere dense.
Closed: if \(f_n\in P\) and \(\|f_n-f\|_{L^1}\to0\), then \(|\int_Af-\int_Af_n|\le\|f-f_n\|_{L^1}\to0\) for every measurable \(A\subseteq[0,1]\), so \(\int_Af\,dx\ge0\) for every such \(A\); taking \(A=\{f<0\}\), where the integrand is strictly negative, forces \(\lambda(A)=0\), i.e. \(f\in P\).
Empty interior: let \(f\in P\) and \(\varepsilon>0\). By absolute continuity of the integral of \(f+1\in L^1[0,1]\) there is \(\delta>0\) with \(\int_B(f+1)dx<\varepsilon\) whenever \(\lambda(B)<\delta\). Put \(\beta=\min(\delta,1)/2\), \(B=[0,\beta]\) and \(g=f-(f+1)I_B\), i.e. \(g=-1\) on \(B\) and \(g=f\) off \(B\). Then \(g\in L^1[0,1]\),
\begin{equation*} \|f-g\|_{L^1}=\int_B(f+1)\,dx<\varepsilon , \end{equation*}
while \(g=-1<0\) on the set \(B\) of positive measure, so \(g\notin P\). Hence no ball centred at a point of \(P\) is contained in \(P\).
(Müntz’s theorem) Suppose we are given a sequence of real numbers \(p_i > -1/2\) with \(\lim_{i \to \infty} p_i = +\infty\). Prove that \(\sum_{i \, : \, p_i \ne 0} 1/p_i = \infty\) precisely when the linear span of the functions \(x^{p_i}\) is everywhere dense in \(L^2[0,1]\).
The span is dense exactly when \(\sum_{i:p_i\ne0}1/p_i=\infty\), and everything follows from the distance formula
\begin{equation*} \operatorname{dist}\big(x^{\lambda_0},\operatorname{span}\{x^{\lambda_1},\dots,x^{\lambda_n}\}\big) =\frac1{\sqrt{2\lambda_0+1}}\prod_{k=1}^{n}\frac{|\lambda_0-\lambda_k|}{\lambda_0+\lambda_k+1} \end{equation*}
for pairwise distinct \(\lambda_0,\dots,\lambda_n>-1/2\). The \(p_i\) are assumed pairwise distinct: a repetition leaves the span unchanged but alters the series, so the series is read over the distinct exponents. Note \(x^p\in L^2[0,1]\) iff \(p>-1/2\), and \(p_i\to+\infty\) leaves only finitely many \(p_i\) in any bounded set, so the series converges or diverges with its tail over \(p_i>1\) and the hypothesis is unambiguous.
Proof of the formula. For linearly independent \(v_1,\dots,v_n\) in a real Hilbert space and \(v_0=u+w\) with \(u\in\operatorname{span}(v_i)\), \(w\perp\operatorname{span}(v_i)\), subtracting from the zeroth row of the Gram matrix \(G(v_0,\dots,v_n)=(\langle v_i,v_j\rangle)\) the combination of the other rows with the coefficients of \(u\) turns that row into \((\|w\|^2,0,\dots,0)\), so expanding along it,
\begin{equation*} \operatorname{dist}\big(v_0,\operatorname{span}(v_i)\big)^2 =\frac{\det G(v_0,v_1,\dots,v_n)}{\det G(v_1,\dots,v_n)} . \end{equation*}
For \(v_i=x^{\lambda_i}\) one has \(\langle x^{\lambda_i},x^{\lambda_j}\rangle=(\lambda_i+\lambda_j+1)^{-1}\), a Cauchy matrix, and Cauchy’s determinant identity
\begin{equation*} \det\Big(\frac1{a_i+b_j}\Big)_{i,j=1}^{m} =\frac{\prod_{i<j}(a_j-a_i)(b_j-b_i)}{\prod_{i,j}(a_i+b_j)} \end{equation*}
with \(a_i=\lambda_i\), \(b_j=\lambda_j+1\) gives
\begin{equation*} \det G(x^{\lambda_0},\dots,x^{\lambda_n}) =\frac{\prod_{0\le i<j\le n}(\lambda_j-\lambda_i)^2}{\prod_{0\le i,j\le n}(\lambda_i+\lambda_j+1)} , \end{equation*}
and likewise over the indices \(1,\dots,n\); dividing and using
\begin{equation*} \frac{\prod_{0\le i,j\le n}(\lambda_i+\lambda_j+1)}{\prod_{1\le i,j\le n}(\lambda_i+\lambda_j+1)} =(2\lambda_0+1)\prod_{k=1}^{n}(\lambda_0+\lambda_k+1)^2 \end{equation*}
yields the formula. (Positivity of these Gram determinants also shows that distinct powers are linearly independent, as assumed.) Each factor lies in \([0,1)\), since for \(\lambda_0\ge\lambda_k\) the inequality \(\lambda_0-\lambda_k<\lambda_0+\lambda_k+1\) is \(2\lambda_k+1>0\).
(i) Divergence implies density. Discard the finitely many indices with \(p_i\le1\), so that \(p_k>1\) for all \(k\) and still \(\sum_k1/p_k=\infty\). Fix an integer \(m\ge0\); if \(m=p_k\) for some \(k\), then \(x^m\) is in the span, and otherwise, by the formula,
\begin{equation*} \operatorname{dist}\big(x^m,\operatorname{span}\{x^{p_1},\dots,x^{p_n}\}\big) =\frac1{\sqrt{2m+1}}\prod_{k=1}^{n}\Big|1-\frac{2m+1}{m+p_k+1}\Big| . \end{equation*}
For \(p_k>2m+1\) the \(k\)th factor equals \(1-\frac{2m+1}{m+p_k+1}\le\exp\big(-\frac{2m+1}{m+p_k+1}\big)\) and the remaining finitely many factors are at most \(1\), while \(m+p_k+1\le3p_k\) for \(p_k\ge m+1\) makes \(\sum_k(m+p_k+1)^{-1}=\infty\). Hence the product tends to \(0\), every \(x^m\) lies in the closed span, and so does every polynomial. Polynomials are dense in \(L^2[0,1]\): extend \(f\) by zero, approximate it in \(L^2(\mathbb{R})\) by \(\varphi\in C_0^{\infty}(\mathbb{R})\) (Corollary 4.2.2) and \(\varphi\) uniformly on \([0,1]\) by a polynomial (Weierstrass).
(ii) Convergence implies non-density. Only finitely many \(p_i\) lie in \((0,1)\), so choose \(\lambda_0\in(0,1)\) distinct from every \(p_i\). All factors of
\begin{equation*} d_n:=\operatorname{dist}\big(x^{\lambda_0},\operatorname{span}\{x^{p_1},\dots,x^{p_n}\}\big) =\frac1{\sqrt{2\lambda_0+1}}\prod_{k=1}^{n}\frac{|\lambda_0-p_k|}{\lambda_0+p_k+1} \end{equation*}
are strictly positive, and \(c_k:=1-|\lambda_0-p_k|/(\lambda_0+p_k+1)\in[0,1)\) satisfies \(c_k=(2\lambda_0+1)/(\lambda_0+p_k+1)\le(2\lambda_0+1)/p_k\) for \(p_k>\lambda_0\), so \(\sum_kc_k<\infty\) by the assumed convergence. Then \(\prod_k(1-c_k)>0\), because \(c_k\to0\) gives \(c_k\le1/2\) for \(k\ge k_0\) and \(\ln(1-c)\ge-2c\) on \([0,1/2]\), while the finitely many earlier factors are positive constants. Hence \(d_n\downarrow c>0\); as every element of the span of all the \(x^{p_i}\) lies in the span of finitely many of them, the distance from \(x^{\lambda_0}\) to that span is \(\inf_nd_n=c>0\), and the span is not dense.
Prove that the Haar functions \(h_n\) form a Schauder basis in \(L^p[0,1]\) for all \(p \in [1,+\infty)\). The Haar functions \(h_n\) are defined as follows: for all \(n \ge 1\) and \(1 \le i \le 2^n\) we set \(h_{2^n+i}(t) = I_{[(2i-2)/2^{n+1},\, (2i-1)/2^{n+1}]}(t) - I_{((2i-1)/2^{n+1},\, 2i/2^{n+1}]}(t)\).
Yes, with coefficient functionals \(f\mapsto\|h_k\|_2^{-2}\langle f,h_k\rangle\); the partial sum operators are the conditional expectations on dyadic partitions, which are \(L^p\)-contractions. The system starts with \(h_1\equiv1\) (without it the closed span would contain only functions with zero integral), the displayed formula being used for all \(n\ge0\); how endpoints are shared between halves is irrelevant in \(L^p\). For \(\Delta_{n,i}=[(i-1)2^{-n},i2^{-n}]\) with halves \(\Delta^{-}_{n,i}\), \(\Delta^{+}_{n,i}\),
\begin{equation*} h_{2^n+i}=I_{\Delta^{-}_{n,i}}-I_{\Delta^{+}_{n,i}},\qquad \int_0^1h_{2^n+i}\,dt=0,\qquad\|h_{2^n+i}\|_2^2=2^{-n} . \end{equation*}
Orthogonality and spans: two Haar functions of the same generation have disjoint supports, and for generations \(n<m\) the later function has zero integral and is supported in one half of some \(\Delta_{n,i}\), where the earlier one is constant; also \(\langle h_1,h_k\rangle=\int_0^1h_k=0\) for \(k\ge2\). For \(2^n<N=2^n+m\le2^{n+1}\) let \(\mathcal{P}_N\) consist of the two halves of \(\Delta_{n,i}\) for \(i\le m\) and of \(\Delta_{n,i}\) for \(i>m\) (for \(N=2^n\), of the \(\Delta_{n,i}\)), and let \(W_N\) be the functions constant on its elements. Then \(\dim W_N=2m+(2^n-m)=N\), and \(h_1,\dots,h_N\in W_N\) are orthogonal and nonzero, hence
\begin{equation*} \operatorname{span}\{h_1,\dots,h_N\}=W_N , \end{equation*}
the mesh of \(\mathcal{P}_N\) being at most \(2^{-n}\to0\).
Partial sums are averages: with
\begin{equation*} S_Nf=\sum_{k=1}^{N}\frac{\langle f,h_k\rangle}{\|h_k\|_2^2}h_k ,\qquad E_Nf=\sum_{J\in\mathcal{P}_N}\Big(\frac1{\lambda(J)}\int_Jf\,dx\Big)I_J \end{equation*}
(all \(\langle f,h_k\rangle\) are finite, as \(L^p[0,1]\subseteq L^1[0,1]\) and the \(h_k\) are bounded), one has \(S_N=E_N\): \(E_Nf\in W_N\) and \(\int_J(f-E_Nf)dx=0\) for every \(J\in\mathcal{P}_N\) give \(f-E_Nf\perp W_N\), so \(\langle f,h_k\rangle=\langle E_Nf,h_k\rangle\) for \(k\le N\) and the orthogonal expansion of \(E_Nf\in W_N\) is \(S_Nf\). Hoelder’s inequality on each \(J\) gives \(|\lambda(J)^{-1}\int_Jf|^p\le\lambda(J)^{-1}\int_J|f|^p\), hence
\begin{equation*} \|S_Nf\|_p=\|E_Nf\|_p\le\|f\|_p\qquad(1\le p<\infty,\ N\ge1) . \end{equation*}
Convergence: for continuous \(g\) with modulus of continuity \(\omega_g\), the value of \(E_Ng\) on \(J\in\mathcal{P}_N\) is an average of values of \(g\) on \(J\), so \(\|S_Ng-g\|_p\le\|S_Ng-g\|_{\infty}\le\omega_g(2^{-n})\to0\). Given \(f\in L^p[0,1]\) and \(\varepsilon>0\), choose a continuous \(g\) with \(\|f-g\|_p<\varepsilon/3\) (Corollary 4.2.2, restriction to \([0,1]\) not increasing the norm); then
\begin{equation*} \|S_Nf-f\|_p\le\|S_N(f-g)\|_p+\|S_Ng-g\|_p+\|g-f\|_p<\varepsilon \end{equation*}
for all large \(N\), i.e. \(f=\sum_{k\ge1}\|h_k\|_2^{-2}\langle f,h_k\rangle h_k\) in \(L^p[0,1]\).
Uniqueness: if \(f=\sum_{k\ge1}a_kh_k\) in \(L^p[0,1]\), the functional \(u\mapsto\int_0^1uh_j\,dx\) is continuous on \(L^p[0,1]\) (Hoelder, \(h_j\) bounded), and applying it to the partial sums with \(N\ge j\) and using orthogonality gives \(\langle f,h_j\rangle=a_j\|h_j\|_2^2\).
(\(\circ\)) Let \(\mu\) be a finite nonnegative measure on a space \(X\). For \(f, g \in L^0(\mu)\), we set \[ d_0(f,g) := \int_X \frac{|f-g|}{1+|f-g|}\, d\mu, \qquad d_1(f,g) := \int_X \min(|f-g|, 1)\, d\mu . \] Prove that \(d_0\) and \(d_1\) are metrics, with respect to which \(L^0(\mu)\) is complete, and that a sequence converges in one of these metrics precisely when it converges in measure (similarly for fundamental sequences).
Both are metrics inducing convergence in measure, and \(L^0(\mu)\) is complete in each, because \(\varphi(t)=t/(1+t)\) and \(\psi(t)=\min(t,1)\) are nondecreasing, bounded by \(1\), vanish only at \(0\), are subadditive on \([0,\infty)\) and satisfy \(\tfrac12\psi\le\varphi\le\psi\). Subadditivity:
\begin{equation*} \varphi(a+b)=\frac{a}{1+a+b}+\frac{b}{1+a+b} \le\frac{a}{1+a}+\frac{b}{1+b}=\varphi(a)+\varphi(b) , \end{equation*}
while for \(\psi\), either \(\max(a,b)\ge1\) and \(\psi(a)+\psi(b)\ge1\ge\psi(a+b)\), or \(a,b<1\) and \(\psi(a)+\psi(b)=a+b\ge\psi(a+b)\); the comparison reads \(t/2\le t/(1+t)\le t\) for \(t\le1\) and \(1/2\le t/(1+t)\le1\) for \(t\ge1\).
Metrics: \(d_0(f,g)\le\mu(X)<\infty\) and \(d_1(f,g)\le\mu(X)\), symmetry is clear, \(d_0(f,g)=0\) forces \(|f-g|=0\) a.e., i.e. \(f=g\) in \(L^0(\mu)\), and monotonicity with subadditivity give pointwise \(\varphi(|f-g|)\le\varphi(|f-h|)+\varphi(|h-g|)\), which integrates to the triangle inequality; the same for \(\psi\). Moreover \(\tfrac12d_1\le d_0\le d_1\), so the two metrics have the same convergent and the same fundamental sequences, and it suffices to treat \(d_0\).
Convergence in measure: \(\varphi\) being strictly increasing, \(\{|f-g|\ge\varepsilon\}=\{\varphi(|f-g|)\ge\varphi(\varepsilon)\}\), so Chebyshev’s inequality gives
\begin{equation*} \mu\big(|f-g|\ge\varepsilon\big)\le\frac{1+\varepsilon}{\varepsilon}\,d_0(f,g) ,\tag{a} \end{equation*}
while splitting the integral at the level \(\varepsilon\) and bounding the integrand by \(\varphi(\varepsilon)\le\varepsilon\), resp. by \(1\),
\begin{equation*} d_0(f,g)\le\varepsilon\mu(X)+\mu\big(|f-g|\ge\varepsilon\big) .\tag{b} \end{equation*}
Thus \(d_0(f_n,f)\to0\) gives convergence in measure by (a); conversely, given \(\delta>0\) choose \(\varepsilon\) with \(\varepsilon\mu(X)<\delta/2\) and apply (b). Applied to the pairs \((f_n,f_m)\), (a) and (b) identify \(d_0\)-fundamental sequences with those fundamental in measure.
Completeness: a \(d_0\)-fundamental sequence is fundamental in measure, hence converges in measure to some \(\mu\)-measurable \(f\) by Theorem 2.2.5(ii) (the measure being finite), and then \(d_0(f_n,f)\to0\) by the equivalence just proved. The same holds for \(d_1\).
(Nikodym) Prove that on the space \(L^0[0,1]\) of all Lebesgue measurable functions equipped with the metric \[ d(f,g) = \int_0^1 \frac{|f-g|}{1+|f-g|}\, dx \] corresponding to convergence in measure, there exists no continuous linear function except for the identically zero one. Extend this assertion to the case of an arbitrary atomless probability measure.
Only \(L\equiv0\), because every \(f\) is the barycentre of \(n\) elements of an arbitrarily small ball. Since \(L(0)=0\), continuity at the origin gives \(r>0\) with \(|L(g)|<1\) whenever \(d(g,0)<r\). Fix \(f\in L^0[0,1]\) and an integer \(n>1/r\), and put \(f_k=nfI_{[(k-1)/n,k/n)}\), \(k=1,\dots,n\), so that \(f=\frac1n(f_1+\cdots+f_n)\) and, the integrand being bounded by \(1\) and vanishing off an interval of length \(1/n\),
\begin{equation*} d(f_k,0)=\int_{(k-1)/n}^{k/n}\frac{n|f|}{1+n|f|}\,dx\le\frac1n<r . \end{equation*}
Hence \(|L(f)|\le\frac1n\sum_{k=1}^{n}|L(f_k)|<1\) for every \(f\in L^0[0,1]\); applying this to \(tf\) gives \(t|L(f)|<1\) for all \(t>0\), so \(L(f)=0\).
For an arbitrary atomless probability measure \(\mu\) the only ingredient used, a partition of \(X\) into \(n\) sets of measure \(1/n\), is again available: by Corollary 1.12.10 an atomless measure takes every value in \([0,\mu(X)]\), so choose \(X_1\) with \(\mu(X_1)=1/n\), then \(X_2\subseteq X\setminus X_1\) with \(\mu(X_2)=1/n\) (the restriction to \(X\setminus X_1\) is atomless of mass \((n-1)/n\)), and after \(n-1\) steps put \(X_n=X\setminus(X_1\cup\cdots\cup X_{n-1})\). With \(f_k=nfI_{X_k}\) one has \(f=\frac1n(f_1+\cdots+f_n)\) and
\begin{equation*} d(f_k,0)=\int_{X_k}\frac{n|f|}{1+n|f|}\,d\mu\le\mu(X_k)=\frac1n<r , \end{equation*}
so the argument above again gives \(L\equiv0\).
Let \(\mu\) be a nonnegative measure, \(0<p<1\), and let \(L^p(\mu)\) be the set of all equivalence classes of \(\mu\)-measurable functions \(f\) such that \(|f|^p \in L^1(\mu)\).
(i) Prove that the function \[ d_p(f,g) := \int |f-g|^p \, d\mu \] is a complete metric on the space \(L^p(\mu)\).
(ii) Prove that \(L^p(\mu)\) is a linear space such that the operations of addition and multiplication by real numbers are continuous on \(L^p(\mu)\) with the metric \(d_p\) (i.e., \(L^p(\mu)\) is a complete metrizable topological vector space).
(iii) Prove that in the case where \(\mu\) is Lebesgue measure on \([a,b]\), there is no nonzero linear function on the space \(L^p(\mu)\) continuous with respect to the metric \(d_p\). In particular, convergence in the metric \(d_p\) cannot be described by any norm.
All three parts rest on the inequality \((a+b)^p\le a^p+b^p\) for \(a,b\ge0\) and \(0<p<1\): assuming \(a+b>0\) and putting \(s=a/(a+b)\), \(t=b/(a+b)\), one has \(s,t\in[0,1]\), \(s+t=1\) and, since \(u^p\ge u\) on \([0,1]\),
\begin{equation*} \frac{a^p+b^p}{(a+b)^p}=s^p+t^p\ge s+t=1 .\tag{\(\dagger\)} \end{equation*}
(i) By \((\dagger)\), \(|f+g|^p\le|f|^p+|g|^p\) and \(|cf|^p=|c|^p|f|^p\), so \(L^p(\mu)\) is a linear space and \(d_p(f,g)<\infty\); \(d_p\) is symmetric, vanishes exactly when \(f=g\) a.e., is translation invariant, and \((\dagger)\) with \(a=|f-h|\), \(b=|h-g|\) integrates to \(d_p(f,g)\le d_p(f,h)+d_p(h,g)\).
Completeness: let \((f_n)\) be \(d_p\)-fundamental and choose \(n_1<n_2<\cdots\) with \(d_p(f_{n_{k+1}},f_{n_k})\le2^{-k}\). By the monotone convergence theorem \(\int\sum_k|f_{n_{k+1}}-f_{n_k}|^pd\mu<\infty\), so \(a_k:=|f_{n_{k+1}}(x)-f_{n_k}(x)|\) has \(\sum_ka_k^p<\infty\) for a.e. \(x\); then \(a_k\le1\) for large \(k\), hence \(a_k\le a_k^p\) and \(\sum_ka_k<\infty\), so \((f_{n_k}(x))\) converges a.e. to a measurable \(f\) (put \(f=0\) on the exceptional null set). By Fatou’s theorem,
\begin{equation*} \int|f-f_{n_k}|^p\,d\mu\le\liminf_{j\to\infty}\int|f_{n_j}-f_{n_k}|^p\,d\mu \le\sup_{j\ge k}d_p(f_{n_j},f_{n_k})\longrightarrow0 , \end{equation*}
so \(f=(f-f_{n_k})+f_{n_k}\in L^p(\mu)\) and \(d_p(f_{n_k},f)\to0\); a fundamental sequence with a convergent subsequence converges. (Check!)
(ii) Addition is jointly continuous, since \(d_p(f+g,f_0+g_0)\le d_p(f,f_0)+d_p(g,g_0)\) by \((\dagger)\), and so is scalar multiplication:
\begin{equation*} d_p(cf,c_0f_0)\le|c|^pd_p(f,f_0)+|c-c_0|^p\int|f_0|^p\,d\mu\longrightarrow0 \end{equation*}
as \(c\to c_0\) and \(d_p(f,f_0)\to0\), the factor \(|c|^p\) staying bounded. With (i) and the translation invariance of \(d_p\), the topology is a vector topology and \(L^p(\mu)\) is a complete metrizable topological vector space.
(iii) Only \(L\equiv0\). Continuity at \(0\) gives \(r>0\) with \(|L(g)|<1\) whenever \(d_p(g,0)<r\). Let \(f\in L^p[a,b]\) with \(I:=\int_a^b|f|^pdx>0\) (otherwise \(f=0\) and \(L(f)=0\)). The function \(F(t)=\int_a^t|f|^pdx\) is continuous and nondecreasing from \(0\) to \(I\), so for a given \(n\) the intermediate value theorem provides \(a=t_0\le\cdots\le t_n=b\) with \(F(t_k)=kI/n\), i.e. \(\int_{J_k}|f|^pdx=I/n\) for \(J_k=[t_{k-1},t_k)\). With \(f_k=nfI_{J_k}\) one has \(f=\frac1n(f_1+\cdots+f_n)\) a.e. and
\begin{equation*} d_p(f_k,0)=n^p\int_{J_k}|f|^p\,dx=n^{p-1}I\longrightarrow0\qquad(n\to\infty) \end{equation*}
because \(p-1<0\). Choosing \(n\) with \(n^{p-1}I<r\) gives \(|L(f)|\le\frac1n\sum_{k=1}^{n}|L(f_k)|<1\), and applying this to \(tf\), \(t\to+\infty\), gives \(L(f)=0\).
Consequently no norm describes \(d_p\)-convergence: picking \(f_0\ne0\) (the space contains all bounded functions) and extending \(\ell(\lambda f_0)=\lambda\|f_0\|\) by the Hahn–Banach theorem to \(\ell\) with \(|\ell|\le\|\cdot\|\), the nonzero \(\ell\) would be sequentially, hence topologically, \(d_p\)-continuous, contradicting the previous paragraph.
Exercises 4.7.63–4.7.69
(\(\circ\)) Show that a probability measure \(\mu\) on a \(\sigma\)-algebra \(\mathcal{A}\) is separable if and only if all spaces \(L^p(\mu)\), \(p \in (0,+\infty)\), are separable, and the separability of either of these spaces is sufficient.
All three properties are equivalent: separability of \(\mu\) gives separability of every \(L^p(\mu)\), and separability of one \(L^p(\mu)\) gives back separability of \(\mu\). Write \(\rho_p\) for the metric of \(L^p(\mu)\), so \(\rho_p(f,g) = \Phi_p\big(\int|f-g|^p d\mu\big)\) with \(\Phi_p(u)=u^{1/p}\) for \(p \ge 1\) and \(\Phi_p(u)=u\) for \(0<p<1\) (Exercise 4.7.62 supplies the metric in the latter range); \(\Phi_p\) is an increasing homeomorphism of \([0,+\infty)\) fixing \(0\). Separability of \(\mu\) means that \((\mathcal{A}/\mu,d)\), \(d(A,B)=\mu(A\triangle B)\), is separable, i.e. there is a countable \(\{A_n\}\subset\mathcal{A}\) with \(\inf_n\mu(A\triangle A_n)=0\) for every \(A\in\mathcal{A}\).
(i) \(\mu\) separable \(\Rightarrow\) \(L^p(\mu)\) separable for every \(p\in(0,+\infty)\). Simple functions are dense: \(f I_{\{|f|\le n\}}\to f\) in \(L^p\) by dominated convergence (dominant \(|f|^p\)), and a bounded \(g\), \(|g|\le M\), satisfies \(|g-s|\le\delta\) for the simple \(s=\sum_i c_iI_{g^{-1}(J_i)}\) built on a partition of \([-M,M]\) into intervals \(J_i\) of length \(<\delta\), whence \(\int|g-s|^p d\mu\le\delta^p\) (for \(p\ge1\) this is Lemma 4.2.1). Let \(\mathcal{D}\) be the countable set of functions \(\sum_{i\le k}q_iI_{A_{n_i}}\), \(q_i\in\mathbb{Q}\). Given simple \(s=\sum_{i\le k}c_iI_{B_i}\), decompose \(c_iI_{B_i}-q_iI_{A_{n_i}} = (c_i-q_i)I_{B_i}+q_i(I_{B_i}-I_{A_{n_i}})\) and use \(|I_{B_i}-I_{A_{n_i}}|=I_{B_i\triangle A_{n_i}}\):
\begin{equation*} \int \big| c_i I_{B_i} - q_i I_{A_{n_i}} \big|^p d\mu \le 2^{p}\Big( |c_i-q_i|^p \mu(B_i) + |q_i|^p \mu(B_i \triangle A_{n_i}) \Big). \end{equation*}
Choosing \(q_i\) close to \(c_i\) and \(A_{n_i}\) close to \(B_i\) makes each summand as small as we please, so summing the \(k\) terms by the triangle inequality for \(\rho_p\) puts an element of \(\mathcal{D}\) within \(\varepsilon\) of \(s\). Hence \(\mathcal{D}\) is dense.
(ii) \(L^p(\mu)\) separable for a single \(p\) \(\Rightarrow\) \(\mu\) separable. The map \(J\colon\mathcal{A}/\mu\to L^p(\mu)\), \(A\mapsto I_A\), is injective and \(|I_A-I_B|^p = I_{A\triangle B}\), so
\begin{equation*} \rho_p(I_A, I_B) = \Phi_p\big(\mu(A \triangle B)\big) = \Phi_p\big(d(A,B)\big), \end{equation*}
i.e. \(J\) is a homeomorphism onto \(\{I_A: A\in\mathcal{A}\}\subset L^p(\mu)\). Every subspace of a separable metric space is separable (the trace of a countable base is a countable base), so \((\mathcal{A}/\mu,d)\) is separable.
Let \(\mathcal{A}\) be a countably generated \(\sigma\)-algebra (i.e., generated by a countable family of sets) and let \(\mu_t\), \(t \in T\), be some family of probability measures on \(\mathcal{A}\). Prove that this family is separable in the variation norm precisely when there exists a probability measure \(\mu\) on \(\mathcal{A}\) such that \(\mu_t \ll \mu\) for all \(t \in T\).
The dominating measure is \(\mu=\sum_n 2^{-n}\mu_{t_n}\) along a countable variation-dense subfamily, and conversely a dominating \(\mu\) embeds \(\mathcal{M}:=\{\mu_t: t\in T\}\) isometrically into the separable space \(L^1(\mu)\).
(i) A dominating \(\mu\) makes \(\mathcal{M}\) separable. First, \(\mu\) is a separable measure: with \(\mathcal{A}=\sigma(\{E_j\})\), the algebra \(\mathcal{A}_0\) generated by the \(E_j\) is countable, and the class of \(A\in\mathcal{A}\) approximable by members of \(\mathcal{A}_0\) in the Frechet–Nikodym metric is closed under complements (\(A^c\triangle A_0^c = A\triangle A_0\)), finite unions (\((A\cup B)\triangle(A_0\cup B_0)\subset(A\triangle A_0)\cup(B\triangle B_0)\)) and countable unions (continuity of \(\mu\) reduces \(\bigcup_n A_n\) to a finite subunion within \(\varepsilon/2\)), hence is a \(\sigma\)-algebra containing every \(E_j\) and so equals \(\mathcal{A}\); this is Exercise 1.12.102. By Exercise 4.7.63, \(L^1(\mu)\) is separable. Radon–Nikodym (applicable since each \(\mu_t\ll\mu\) with \(\mu\) finite) gives densities \(f_t\ge0\), \(\mu_t=f_t\cdot\mu\), and with \(P=\{f_t-f_s\ge0\}\) a Hahn positive set for \(\mu_t-\mu_s\),
\begin{equation*} \|\mu_t - \mu_s\| = (\mu_t-\mu_s)(P) - (\mu_t-\mu_s)(X\setminus P) = \|f_t-f_s\|_{L^1(\mu)} . \end{equation*}
So \(\mu_t\mapsto f_t\) is an isometry of \(\mathcal{M}\) onto a subspace of the separable space \(L^1(\mu)\), and subspaces of separable metric spaces are separable (Exercise 4.7.63).
(ii) Separability of \(\mathcal{M}\) produces a dominating \(\mu\). A countable dense subset of \(\mathcal{M}\) has the form \(\{\mu_{t_n}\}_{n\ge1}\); put
\begin{equation*} \mu := \sum_{n=1}^{\infty} 2^{-n} \mu_{t_n}, \end{equation*}
a probability measure by term-by-term integration of nonnegative series. If \(\mu(A)=0\) then \(\mu_{t_n}(A)=0\) for all \(n\), so for \(t \in T\) and \(\varepsilon>0\), picking \(n\) with \(\|\mu_t-\mu_{t_n}\|<\varepsilon\),
\begin{equation*} \mu_t(A) = |(\mu_t - \mu_{t_n})(A)| \le \|\mu_t - \mu_{t_n}\| < \varepsilon . \end{equation*}
Thus \(\mu_t(A)=0\), i.e. \(\mu_t \ll \mu\) for every \(t \in T\).
(\(\circ\)) Let \(\mu\) be a probability measure and let \(f \in L^p(\mu)\). Show that the function \(\theta\colon r \mapsto \ln \|f\|^r_{L^r(\mu)}\) is convex on \([1,p]\), i.e., \(\theta\big(tr+(1-t)s\big) \le t\theta( r) + (1-t)\theta(s)\) for all \(0 < t < 1\) and \(r,s \in [1,p]\).
Convexity is Hoelder’s inequality with the conjugate exponents \(1/t\) and \(1/(1-t)\) applied to the splitting \(|f|^{u} = |f|^{tr}\cdot|f|^{(1-t)s}\), where \(u := tr+(1-t)s \in [1,p]\).
Here \(\theta( r) = \ln\int_X|f|^r\,d\mu\), finite for \(r \in [1,p]\): Hoelder with the exponents \(p/r\), \(p/(p-r)\) and \(\mu(X)=1\) give \(\int|f|^r d\mu \le \|f\|^r_{L^p(\mu)}<\infty\). If \(f=0\) a.e. then \(\theta\equiv-\infty\) and the inequality is trivial in the extended sense; otherwise \(\int|f|^r d\mu>0\) for all \(r\) and \(\theta\) is real-valued. Hoelder now gives
\begin{equation*} \int_X |f|^{u}\,d\mu \le \Big( \int_X |f|^{r} d\mu \Big)^{t} \Big( \int_X |f|^{s} d\mu \Big)^{1-t} , \end{equation*}
and taking logarithms of the positive finite sides,
\begin{equation*} \theta\big(tr+(1-t)s\big) \le t\,\theta( r) + (1-t)\,\theta(s), \qquad 0<t<1,\ r,s \in [1,p]. \end{equation*}
Let \(\psi\) be a positive function on \([1,+\infty)\) increasing to the infinity. Prove that there exists a positive measurable function \(f\) on \([0,1]\) such that \(\|f\|_p \le \psi(p)\) for all \(p \ge 1\) and \(\lim_{p \to \infty}\|f\|_p = \infty\).
Take pairwise disjoint intervals \(A_k \subset [0,1]\), \(k \ge 1\), of lengths
\begin{equation*} a_k := 2^{-k-2} \min_{0 \le n \le k-2} \Big(\frac{\psi(n+1)}{\psi(k)}\Big)^{n+2} \qquad (\text{empty minimum} := 1), \end{equation*}
put \(A_0 := [0,1]\setminus\bigcup_{k\ge1}A_k\) (legitimate since \(\sum_k a_k \le \sum_k 2^{-k-2} = 1/4\)) and set
\begin{equation*} f := \frac{\psi(1)}{2}\,I_{A_0} + \sum_{k \ge 1} \psi(k)\, I_{A_k} > 0 . \end{equation*}
Since \(\lambda\) is a probability measure on \([0,1]\), Hoelder’s inequality with the exponents \(q/p\), \(q/(q-p)\) gives \(\|g\|_p \le \|g\|_q\) for \(1\le p\le q\), so \(p\mapsto\|f\|_p\) is nondecreasing and its limit exists in \([0,+\infty]\).
Key estimate: \(\|f\|_{n+2}\le\psi(n+1)\) for every integer \(n\ge0\). With \(M:=\psi(n+1)\) and \(q:=n+2\ge2\), split \(\|f\|_q^q = (\psi(1)/2)^q\lambda(A_0) + \sum_{k\ge1}\psi(k)^q a_k\) into three pieces.
- (i) \(\psi(1)\le M\) and \(2^{-q}\le 1/4\), so \((\psi(1)/2)^q\lambda(A_0) \le M^q/4\).
- (ii) For \(k \le n+1\), \(\psi(k)\le M\), so \(\sum_{k\le n+1}\psi(k)^qa_k \le M^q\sum_k a_k \le M^q/4\).
- (iii) For \(k \ge n+2\) the index \(n\) occurs in the minimum defining \(a_k\), so \(a_k \le 2^{-k-2}M^q/\psi(k)^q\) and \(\sum_{k\ge n+2}\psi(k)^qa_k \le M^q 2^{-n-3} \le M^q/8\).
Adding, \(\|f\|_{n+2}^{n+2} \le M^{n+2}(\tfrac14+\tfrac14+\tfrac18) < M^{n+2}\). Hence for \(p \in [n+1,n+2]\), monotonicity in \(p\) and monotonicity of \(\psi\) give
\begin{equation*} \|f\|_p \le \|f\|_{n+2} \le \psi(n+1) \le \psi(p) , \end{equation*}
so \(f \in L^p[0,1]\) with the required bound for every \(p \ge 1\). Finally \(f=\psi(k)\) on \(A_k\), \(\lambda(A_k)=a_k>0\), whence \(\|f\|_p \ge \psi(k)a_k^{1/p} \to \psi(k)\) as \(p\to\infty\); letting \(k \to \infty\) gives \(\lim_{p\to\infty}\|f\|_p=\infty\).
Prove Corollary 4.5.5.
For \(f_n,f\in\mathcal{L}^1(\mu)\) with \(f_n \to f\) a.e., convergence in \(L^1(\mu)\) is equivalent to (a) uniform absolute continuity of the integrals \(\int_E|f_n|\,d\mu\) (Definition 4.5.2) together with (b) the existence, for each \(\varepsilon>0\), of \(X_\varepsilon\) with \(\mu(X_\varepsilon)<\infty\) and \(\sup_n\int_{X\setminus X_\varepsilon}|f_n|\,d\mu<\varepsilon\). Both halves use the corresponding properties of a single \(g \in \mathcal{L}^1(\mu)\) and \(\eta>0\): (i) some \(\delta>0\) has \(\mu(E)<\delta \Rightarrow \int_E|g|\,d\mu<\eta\), since monotone convergence supplies \(m\) with \(\int_{\{|g|>m\}}|g|\,d\mu<\eta/2\) and then \(\int_E|g|\,d\mu \le \eta/2+m\mu(E)\); (ii) some \(Y\) has \(\mu(Y)<\infty\) and \(\int_{X\setminus Y}|g|\,d\mu<\eta\), since \(Y_m:=\{|g|>1/m\}\) has \(\mu(Y_m)\le m\int|g|\,d\mu<\infty\) by Chebyshev and \(\int_{X\setminus Y_m}|g|\,d\mu\to0\) by dominated convergence.
Necessity. Let \(\|f_n-f\|_{L^1(\mu)}\to0\), fix \(\varepsilon>0\) and \(N\) with \(\|f_n-f\|_{L^1(\mu)}<\varepsilon/2\) for \(n>N\). Applying (i) with \(\eta=\varepsilon/2\) to the finitely many functions \(f,f_1,\dots,f_N\) yields \(\delta>0\) handling the indices \(n \le N\) directly, while for \(n>N\) and \(\mu(E)<\delta\),
\begin{equation*} \int_E |f_n|\,d\mu \le \|f_n-f\|_{L^1(\mu)} + \int_E|f|\,d\mu < \varepsilon , \end{equation*}
which is (a). Likewise (ii) with \(\eta=\varepsilon/4\) gives \(Y_f,Y_1,\dots,Y_N\) of finite measure; for \(X_\varepsilon := Y_f\cup Y_1\cup\dots\cup Y_N\) the indices \(n\le N\) give at most \(\varepsilon/4\) and for \(n>N\),
\begin{equation*} \int_{X\setminus X_\varepsilon}|f_n|\,d\mu \le \|f_n-f\|_{L^1(\mu)} + \int_{X\setminus Y_f}|f|\,d\mu < \tfrac{\varepsilon}{2}+\tfrac{\varepsilon}{4} < \varepsilon, \end{equation*}
which is (b).
Sufficiency. Fix \(\varepsilon>0\) and take \(S\) from (b), so \(\mu(S)<\infty\) and \(\sup_n\int_{X\setminus S}|f_n|\,d\mu<\varepsilon\); Fatou’s theorem applied to \(|f_n|I_{X\setminus S}\to|f|I_{X\setminus S}\) a.e. gives \(\int_{X\setminus S}|f|\,d\mu\le\varepsilon\), whence
\begin{equation*} \int_X |f_n-f|\,d\mu \le 2\varepsilon + \int_S |f_n-f|\,d\mu . \end{equation*}
Choose \(\delta>0\) so that \(\mu(E)<\delta\) forces both \(\sup_n\int_E|f_n|\,d\mu\le\varepsilon\) (by (a)) and \(\int_E|f|\,d\mu\le\varepsilon\) (by (i)). Since \(\mu|_S\) is finite and \(f_n\to f\) a.e. on \(S\), Egorov’s theorem gives \(A \subset S\) with \(\mu(A)<\delta\) and uniform convergence on \(S\setminus A\), so
\begin{equation*} \int_S |f_n - f|\,d\mu \le \mu(S)\sup_{S\setminus A}|f_n-f| + \int_A |f_n|\,d\mu + \int_A|f|\,d\mu . \end{equation*}
The first term vanishes as \(n\to\infty\) and the others are at most \(\varepsilon\) each, so \(\limsup_n\int_X|f_n-f|\,d\mu \le 4\varepsilon\); as \(\varepsilon\) was arbitrary, \(\|f_n-f\|_{L^1(\mu)}\to0\).
(\(\circ\)) Suppose that a function \(f \in L^1[0,2\pi]\) satisfies Dini’s condition at a point \(x\) (see Theorem 3.8.8). Prove that its Fourier series at \(x\) converges to \(f(x)\).
Dini’s condition makes the difference \(S_n(x)-f(x)\) an integral against \(\sin((n+\frac12)u)\) of an \(L^1\) function, so it tends to \(0\) by Riemann–Lebesgue. Fix an integrable representative of \(f\), extend it \(2\pi\)-periodically, and let \(\eta \in (0,\pi]\) be such that \(u\mapsto (f(x+u)-f(x))/u\) is integrable on \([-\eta,\eta]\) (Theorem 3.8.8).
By formula (4.3.6) the partial sums are given by the Dirichlet kernel; the integrand there is \(2\pi\)-periodic in \(t\), so the integral over \([0,2\pi]\) equals that over \([x-\pi,x+\pi]\) and \(t=x+u\) gives
\begin{equation*} S_n(x) = \frac{1}{2\pi}\int_{-\pi}^{\pi} f(x+u)\, \frac{\sin\big((n+\tfrac12)u\big)}{\sin\frac{u}{2}}\,du . \end{equation*}
Applying this to \(f \equiv 1\), whose partial sums are all \(1\), normalizes the kernel to total mass \(2\pi\), so subtraction yields
\begin{equation*} S_n(x) - f(x) = \frac{1}{2\pi}\int_{-\pi}^{\pi} g(u)\,\sin\big((n+\tfrac12)u\big)\,du , \qquad g(u) := \frac{f(x+u)-f(x)}{\sin\frac{u}{2}} . \end{equation*}
Here \(g \in \mathcal{L}^1[-\pi,\pi]\): on \([-\eta,\eta]\) it is the Dini quotient times \(h(u)=u/\sin\frac u2\), which is bounded there (it extends continuously with \(h(0)=2\)); on \(\eta\le|u|\le\pi\) one has \(|\sin\frac u2|\ge\sin\frac\eta2>0\), so \(|g(u)|\le(|f(x+u)|+|f(x)|)/\sin\frac\eta2\), integrable since \(f\) is integrable over a period. Expanding \(\sin((n+\tfrac12)u) = \sin(nu)\cos\frac u2 + \cos(nu)\sin\frac u2\),
\begin{equation*} S_n(x)-f(x) = \frac{1}{2\pi}\int_{-\pi}^{\pi} G(u)\sin(nu)\,du + \frac{1}{2\pi}\int_{-\pi}^{\pi} H(u)\cos(nu)\,du \end{equation*}
with \(G:=g\cos\frac u2\) and \(H:=g\sin\frac u2 = f(x+\cdot)-f(x)\), both in \(\mathcal{L}^1[-\pi,\pi]\). By the Riemann–Lebesgue theorem (Exercise 4.7.79(i)) both integrals tend to \(0\), i.e. \(S_n(x)\to f(x)\).
(W. Orlicz) Let \(\{e_n\}\) be an orthonormal basis in the space \(L^2[a,b]\).
(i) Prove that
\begin{equation*} \sum_{n=1}^{\infty}\int_A |e_n(x)|^2\,dx = \infty \end{equation*}
for every set \(A \subset [a,b]\) of positive measure.
(ii) Prove that \(\sum_{n=1}^{\infty}|e_n(x)|^2 = \infty\) a.e.
(i) The sum equals the number of elements of an orthonormal basis of \(L^2(A)\), hence is infinite. Indeed, \(L^2(A)\) is infinite-dimensional (Exercise 4.7.46) and separable (Corollary 4.2.2), so Gram–Schmidt on a countable dense sequence yields a countably infinite orthonormal basis \(\{\varphi_k\}\); extend each \(\varphi_k\) by zero, so that \(I_A\varphi_k = \varphi_k\) and \(\{\varphi_k\}\) is orthonormal in \(L^2[a,b]\) with \(\|\varphi_k\|=1\). Since \(I_A\) is a real bounded multiplier, \((I_Ae_n,\varphi_k) = (e_n,I_A\varphi_k) = (e_n,\varphi_k)\), and Parseval’s equality in \(L^2(A)\) (Theorem 4.3.6) gives
\begin{equation*} \int_A |e_n(x)|^2\,dx = \|I_Ae_n\|^2 = \sum_{k=1}^{\infty}\big|(e_n,\varphi_k)\big|^2 . \end{equation*}
Summing over \(n\) and interchanging the two summations (all terms are nonnegative), Parseval’s equality for the basis \(\{e_n\}\) applied to each \(\varphi_k\) gives
\begin{equation*} \sum_{n=1}^{\infty}\int_A |e_n(x)|^2\,dx = \sum_{k=1}^{\infty}\sum_{n=1}^{\infty}\big|(\varphi_k,e_n)\big|^2 = \sum_{k=1}^{\infty}\|\varphi_k\|^2 = \infty . \end{equation*}
(ii) Put \(S := \sum_{n\ge1}|e_n|^2\) and \(A_M := \{S \le M\}\). If \(\lambda(A_M)>0\) for some \(M\), then term-by-term integration of the nonnegative series gives
\begin{equation*} \sum_{n=1}^{\infty}\int_{A_M}|e_n(x)|^2\,dx = \int_{A_M} S\,dx \le M\,\lambda(A_M) < \infty , \end{equation*}
contradicting (i). So every \(A_M\) is null, and \(\{S<\infty\} = \bigcup_M A_M\) is null.
Exercises 4.7.70–4.7.76
Let \(\mathcal{R}\) be a semiring in a \(\sigma\)-algebra \(\mathcal{A}\) with a probability measure \(\mu\). Show that the set of linear combinations of the indicator functions of sets in \(\mathcal{R}\) is everywhere dense in \(L^1(\mu)\) precisely when, for every \(A \in \mathcal{A}\) and \(\varepsilon > 0\), there exists a set \(B\) that is a union of finitely many sets in \(\mathcal{R}\) such that \(\mu(A \triangle B) < \varepsilon\).
Both directions run through the identity \(\|I_A-I_B\|_{L^1(\mu)} = \mu(A\triangle B)\); the nontrivial one converts an \(L^1\)-approximation \(g\) of \(I_A\) into the set \(B := \{g>1/2\}\). Write \(L\) for the linear span of \(\{I_R : R\in\mathcal{R}\}\) and \(\mathcal{R}_\sigma\) for the finite unions of members of \(\mathcal{R}\); by Lemma 1.2.14, \(\mathcal{R}_\sigma\) is a ring and each of its members is a finite disjoint union of sets of \(\mathcal{R}\), so \(I_B\in L\) for every \(B\in\mathcal{R}_\sigma\).
(i) Approximation \(\Rightarrow\) density. If each \(I_A\), \(A\in\mathcal{A}\), lies in the closure \(\overline{L}\), then so does every simple function, and simple functions are dense in \(L^1(\mu)\) by Lemma 4.2.1; hence \(\overline{L} = L^1(\mu)\).
(ii) Density \(\Rightarrow\) approximation. Fix \(A\in\mathcal{A}\), \(\varepsilon>0\) and \(g = \sum_{i\le m}c_iI_{R_i} \in L\) with \(\int_X|I_A-g|\,d\mu<\varepsilon/2\). The sets
\begin{equation*} E_S := \Big(\bigcap_{i \in S} R_i\Big) \setminus \bigcup_{i \notin S} R_i , \qquad \emptyset \ne S \subset \{1,\dots,m\}, \end{equation*}
together with \(E_\emptyset := X\setminus\bigcup_i R_i\) partition \(X\), and \(g\) is constant on each, with value \(0\) on \(E_\emptyset\). Each \(E_S\), \(S\ne\emptyset\), lies in \(\mathcal{R}_\sigma\): the set \(R:=\bigcap_{i\in S}R_i\) is in \(\mathcal{R}\) (semirings are stable under finite intersections), \(E_S = R\setminus\bigcup_{i\notin S}(R\cap R_i)\), and \(\mathcal{R}_\sigma\) is a ring. Hence \(B := \{g>1/2\}\), being the union of those \(E_S\) with \(S \ne \emptyset\) on which \(g>1/2\), belongs to \(\mathcal{R}_\sigma\). On \(A\triangle B\) one has \(|I_A-g|\ge 1/2\) (either \(I_A=1\), \(g\le1/2\), or \(I_A=0\), \(g>1/2\)), so \[ \tfrac{1}{2}\,\mu(A \triangle B) \le \int_X |I_A - g|\,d\mu < \tfrac{\varepsilon}{2}. \]
Suppose that a sequence of \(\mu\)-integrable functions \(f_n\) (where \(\mu\) takes values in \([0, +\infty]\)) converges almost everywhere to a function \(f\) and that there exist integrable functions \(g_n\) such that \(|f_n| \le g_n\) almost everywhere. Prove that if the sequence \(\{g_n\}\) converges in \(L^1(\mu)\) (or the measure \(\mu\) is finite and \(\{g_n\}\) is uniformly integrable), then \(f\) is integrable and \(\{f_n\}\) converges to \(f\) in \(L^1(\mu)\).
Both hypotheses give the conclusion through generalized dominated convergence applied to \(0 \le |f_n-f| \le g_n+g\).
(i) \(g_n \to g\) in \(L^1(\mu)\). Choosing indices with \(\|g_{n_k}-g\|_{L^1(\mu)}\le2^{-k}\), term-by-term integration (legitimate for \([0,+\infty]\)-valued \(\mu\) by Corollary 2.8.6) gives \(\int_X\sum_k|g_{n_k}-g|\,d\mu\le1\), so \(g_{n_k}\to g\) a.e.; letting \(k\to\infty\) in \(|f_{n_k}|\le g_{n_k}\) yields \(|f|\le g\) a.e., and \(f\) is integrable. Suppose first that \(g_n\to g\) a.e. as well. Then \(h_n := g_n+g-|f_n-f| \ge 0\) a.e., \(h_n \to 2g\) a.e., and Fatou’s theorem (Corollary 2.8.4, extended by Corollary 2.8.6) with \(\int g_n\,d\mu \to \int g\,d\mu\) gives
\begin{equation*} 2\int_X g\,d\mu \le \liminf_{n} \int_X h_n\,d\mu = 2\int_X g\,d\mu - \limsup_{n} \int_X |f_n - f|\,d\mu ; \end{equation*}
cancelling the finite \(\int g\,d\mu\) leaves \(\|f_n-f\|_{L^1(\mu)}\to0\) (this step is Young’s theorem 2.8.8). In general, if \(\|f_{n_j}-f\|_{L^1(\mu)}\ge c>0\) along a subsequence, the extraction above applied to \(\{g_{n_j}\}\) gives a further subsequence with \(g_{n_{j_i}}\to g\) a.e., along which the preceding paragraph forces \(\|f_{n_{j_i}}-f\|_{L^1(\mu)}\to0\) – a contradiction.
(ii) \(\mu\) finite, \(\{g_n\}\) uniformly integrable. Then \(\{f_n\}\) is uniformly integrable too, since \(\{|f_n|>C\}\subset\{g_n>C\}\) a.e. gives \[ \int_{\{|f_n| > C\}} |f_n|\,d\mu \le \int_{\{g_n > C\}} g_n\,d\mu \to 0 \quad (C \to +\infty) \] uniformly in \(n\) by Definition 4.5.1. Egorov’s theorem (Theorem 2.2.1, applicable since \(\mu\) is finite) upgrades \(f_n\to f\) a.e. to convergence in measure, so Theorem 4.5.4 gives integrability of \(f\) and \(f_n\to f\) in \(L^1(\mu)\).
(\(\circ\)) Let \((X, \mathcal{A}, \mu)\) be a probability space and let integrable functions \(f_n\) converge in measure to an integrable function \(f\) such that \[ \lim_{n \to \infty} \int_X \sqrt{1 + f_n^2}\,d\mu = \int_X \sqrt{1 + f^2}\,d\mu . \] Prove that \(f_n \to f\) in \(L^1(\mu)\).
The majorants \(F_n := \sqrt{1+f_n^2} \ge |f_n|\) converge to \(F := \sqrt{1+f^2}\) in \(L^1(\mu)\), which makes \(\{f_n\}\) uniformly integrable, and Theorem 4.5.4 finishes.
Since \(|\Phi^{\prime}|\le1\) for \(\Phi(t)=\sqrt{1+t^2}\), the map \(\Phi\) is \(1\)-Lipschitz, so \(|F_n-F|\le|f_n-f|\) and \(F_n \to F\) in measure; also \(1 \le \Phi(t)\le 1+|t|\), so all \(F_n, F\) are integrable. From \(0 \le (F-F_n)^+ \le F\) and \((F-F_n)^+\to0\) in measure, the dominated convergence theorem for convergence in measure (Theorem 2.8.5) gives \(\int(F-F_n)^+d\mu\to0\); combining with the hypothesis via \(|F_n-F| = (F_n-F)+2(F-F_n)^+\),
\begin{equation*} \int_X |F_n - F|\,d\mu = \Big(\int_X F_n\,d\mu - \int_X F\,d\mu\Big) + 2\int_X (F - F_n)^{+}\,d\mu \to 0 . \end{equation*}
Hence \(\{F_n\}\) has uniformly absolutely continuous integrals: given \(\varepsilon>0\), take \(N\) with \(\|F_n-F\|_{L^1(\mu)}<\varepsilon/2\) for \(n>N\) and apply Theorem 2.5.7 to the finitely many functions \(F,F_1,\dots,F_N\) to get \(\delta>0\) handling \(n \le N\) and giving \(\int_AF\,d\mu<\varepsilon/2\), whence \(\int_AF_n\,d\mu<\varepsilon\) for \(n>N\) too. As \(\mu\) is finite and \(\{F_n\}\) is \(L^1\)-bounded, Proposition 4.5.3 makes \(\{F_n\}\) uniformly integrable. Now \(\{|f_n|>C\}\subset\{F_n>C\}\) gives
\begin{equation*} \int_{\{|f_n| > C\}} |f_n|\,d\mu \le \int_{\{F_n > C\}} F_n\,d\mu \xrightarrow[C \to +\infty]{} 0 \quad \text{uniformly in } n , \end{equation*}
so \(\{f_n\}\) is uniformly integrable and converges to \(f\) in measure; Theorem 4.5.4 yields \(f_n \to f\) in \(L^1(\mu)\).
(Klei, Miyara) Let \((X, \mathcal{A}, \mu)\) be a probability space and let \(M\) be a norm bounded set in \(L^1(\mu)\). The modulus of uniform integrability of \(M\) is the function \[ \eta(M, \varepsilon) := \sup\Big\{ \int_A |f|\,d\mu :\ f \in M,\ A \in \mathcal{A},\ \mu(A) \le \varepsilon \Big\} . \] Set \(\eta(M) := \lim_{\varepsilon \to 0} \eta(M, \varepsilon)\). It is clear that the equality \(\eta(M) = 0\) is equivalent to the uniform integrability of \(M\). Let \(f_n \in L^1(\mu)\), \(f_n \ge 0\), be such that the sequence of the integrals of \(f_n\) is convergent. Prove that \[ \int_X \liminf_{n \to \infty} f_n\,d\mu \le \lim_{n \to \infty} \int_X f_n\,d\mu - \eta(\{f_n\}) . \] Show that under the above conditions the equality occurs precisely when \(\{f_n\}\) contains a subsequence convergent a.e. to the function \(\liminf_{n \to \infty} f_n\).
The inequality is Fatou’s theorem applied to \(h_k := f_{n_k}I_{X\setminus A_k}\) along a subsequence whose sets \(A_k\) nearly realize \(\eta\) while shrinking fast enough for Borel–Cantelli. Write \(M := \{f_n\}\), \(\eta := \eta(M)\), \(L := \lim_n\int_X f_n\,d\mu\) and \(f := \liminf_n f_n\). Here \(\varepsilon\mapsto\eta(M,\varepsilon)\) is nondecreasing and bounded, so \(\eta = \inf_{\varepsilon>0}\eta(M,\varepsilon) \le \eta(M,\varepsilon)\), and \(f\ge0\) is integrable with \(\int_X f\,d\mu\le L<\infty\) by Fatou’s theorem (Corollary 2.8.4); in particular \(f<\infty\) a.e.
Selection. There are \(n_1<n_2<\dots\) and \(A_k\in\mathcal{A}\) with \[ \mu(A_k) \le 2^{-k} \qquad \text{and} \qquad \int_{A_k} f_{n_k}\,d\mu \ge \eta - \frac{1}{k} . \tag{2} \] Indeed, given \(n_1<\dots<n_{k-1}\), put \(N := n_{k-1}\). If \(\eta-1/k\le0\), take \(n_k := N+1\) and \(A_k := \emptyset\). Otherwise Theorem 2.5.7, applied to the finitely many integrable \(f_1,\dots,f_N\) with \(\eta-1/k>0\) in place of \(\varepsilon\), gives \(\delta_0>0\) with \(\int_Df_i\,d\mu<\eta-1/k\) whenever \(\mu(D)<\delta_0\); for \(\delta := \min(\delta_0/2, 2^{-k})\) this reads \(\eta(\{f_1,\dots,f_N\},\delta)\le\eta-1/k\). Since
\begin{equation*} \eta(M,\delta) = \max\big(\eta(\{f_1,\dots,f_N\},\delta),\ \eta(\{f_n\}_{n > N},\delta)\big) \ge \eta > \eta - \tfrac1k , \end{equation*}
the maximum is attained at the tail term, which therefore supplies \(n_k>N\) and \(A_k\) with \(\mu(A_k)\le\delta\le2^{-k}\) satisfying (2).
The inequality. With \(h_k := f_{n_k}I_{X\setminus A_k} \ge 0\), the bound \(\sum_k\mu(A_k)<\infty\) and Borel–Cantelli give, for a.e. \(x\), that \(h_k(x) = f_{n_k}(x)\) for all large \(k\), hence \[ \liminf_{k} h_k \ge \liminf_{n} f_n = f \quad \text{a.e.} \tag{3} \] By (2), \(\int_X h_k\,d\mu \le \int_X f_{n_k}\,d\mu - \eta + 1/k\), so Fatou’s theorem (Corollary 2.8.4) yields \[ \int_X f\,d\mu \le \int_X \liminf_{k} h_k\,d\mu \le \liminf_{k}\int_X h_k\,d\mu \le L - \eta . \tag{4} \]
(i) Equality gives an a.e. convergent subsequence. Then every inequality in (4) is an equality, so \(\lim_k\int_Xh_k\,d\mu = \int_Xf\,d\mu\) and, with (3), \(\liminf_kh_k = f\) a.e. As \(f<\infty\) a.e. this says \((f-h_k)^+\to0\) a.e., dominated by the integrable \(f\); hence \(\int(f-h_k)^+d\mu\to0\) and, via \(|h_k-f| = (h_k-f)+2(f-h_k)^+\),
\begin{equation*} \int_X |h_k - f|\,d\mu = \Big(\int_X h_k\,d\mu - \int_X f\,d\mu\Big) + 2\int_X (f - h_k)^{+}\,d\mu \to 0 . \end{equation*}
So \(h_k \to f\) in \(L^1(\mu)\), hence in measure, and Theorem 2.2.5(i) gives \(h_{k_j}\to f\) a.e.; since \(h_k = f_{n_k}\) eventually a.e., \(f_{n_{k_j}}\to f\) a.e.
(ii) An a.e. convergent subsequence gives equality. Let \(f_{n_k}\to f\) a.e. and fix \(\delta>0\). Egorov’s theorem (Theorem 2.2.1) gives \(A\) with \(\mu(A)<\delta\) and uniform convergence on \(X\setminus A\), so \(\int_{X\setminus A}f_{n_k}\,d\mu\to\int_{X\setminus A}f\,d\mu\) (\(\mu\) is a probability measure) and, using \(\int_Xf_{n_k}\,d\mu\to L\) and \(f\ge0\),
\begin{equation*} \eta(M,\delta) \ge \lim_{k} \int_A f_{n_k}\,d\mu = L - \int_{X \setminus A} f\,d\mu \ge L - \int_X f\,d\mu . \end{equation*}
Letting \(\delta\to0\) gives \(\eta \ge L-\int_Xf\,d\mu\), which with (4) forces equality.
(Farrell) (i) Let \((X, \mathcal{A}, \mu)\) be a probability space and let \(\mathcal{F}\) be an algebra of bounded measurable functions such that, for every measurable set \(A\), there exists \(f \in \mathcal{F}\) with \(f > 0\) a.e. on \(A\) and \(f \le 0\) a.e. on \(X \setminus A\). Prove that for all \(p \in [1, \infty)\) the algebra \(\mathcal{F}\) is dense in \(L^p(\mu)\). Moreover, the same is true if the hypothesis is fulfilled for every set \(A\) in some family \(\mathcal{E} \subset \mathcal{A}\) with the property that the linear space generated by \(I_E\), \(E \in \mathcal{E}\), is dense in \(L^1(\mu)\).
(ii) Let \(\mu\) be a Borel probability measure on the real line and let \(f\) be a strictly increasing bounded function. Show that the algebra of functions generated by \(f\) and \(1\) is dense in \(L^p(\mu)\), \(1 \le p < \infty\).
(i) The separating function \(g\) for \(A\) is turned into \(I_A\) by composing with polynomials \(P_n\) that vanish at \(0\) and approximate the indicator of \((0,\infty)\). Write \(\overline{\mathcal{F}}\) for the closure of \(\mathcal{F}\) in \(L^p(\mu)\); since \(\mathcal{F}\) need not contain constants, every polynomial composed with a member of \(\mathcal{F}\) is taken with zero constant term, so that \(P\circ g\in\mathcal{F}\).
Fix \(N>0\) and let \(\varphi_n:[-N,N]\to[0,1]\) be continuous with \(\varphi_n = 0\) on \([-N,0]\), \(\varphi_n(t)=nt\) on \([0,1/n]\) and \(\varphi_n=1\) on \([1/n,N]\). Weierstrass gives polynomials \(Q_n\) with \(\sup_{[-N,N]}|Q_n-\varphi_n|<1/n\); putting \(P_n := Q_n - Q_n(0)\) and using \(|Q_n(0)|<1/n\) yields \(P_n(0)=0\), \(\sup_{[-N,N]}|P_n - \varphi_n| < 2/n\), hence \(|P_n|\le3\) on \([-N,N]\) with \(P_n\to1\) on \((0,N]\) and \(P_n\to0\) on \([-N,0]\).
Now if \(g\in\mathcal{F}\), \(|g|<N\), satisfies \(g>0\) a.e. on \(A\) and \(g\le0\) a.e. off \(A\), then \(P_n\circ g\in\mathcal{F}\) is bounded by \(3\) and converges a.e. to \(I_A\), hence in \(L^p(\mu)\) by dominated convergence: \(I_A\in\overline{\mathcal{F}}\). If this applies to every \(A\in\mathcal{A}\), then \(\overline{\mathcal{F}}\) contains all simple functions and Lemma 4.2.1 gives \(\overline{\mathcal{F}} = L^p(\mu)\).
For the version with \(\mathcal{E}\), let \(G\) be the span of \(\{I_E : E\in\mathcal{E}\}\) and \(\Theta_N(t) := \max(-N,\min(t,N))\), which is \(1\)-Lipschitz with \(\Theta_N(0)=0\). For \(g = \sum_{i\le m}c_iI_{E_i}\in G\), the previous paragraph gives \(u_n^{(i)}\in\mathcal{F}\), \(|u_n^{(i)}|\le3\), with \(u_n^{(i)}\to I_{E_i}\) a.e.; then \(u_n := \sum_i c_iu_n^{(i)} \in \mathcal{F}\) satisfies \(u_n\to g\) a.e. and \(|u_n|\le C:=3\sum_i|c_i|\). Choosing polynomials \(R_j\) with \(R_j(0)=0\) (legitimate as \(\Theta_N(0)=0\)) and \(\sup_{[-C,C]}|R_j-\Theta_N|\le\varepsilon_j\to0\),
\begin{equation*} \|R_j \circ u_n - \Theta_N(g)\|_{L^p(\mu)} \le \varepsilon_j + \|\Theta_N(u_n) - \Theta_N(g)\|_{L^p(\mu)} , \end{equation*}
whose second term vanishes as \(n\to\infty\) by dominated convergence (\(|\Theta_N(u_n)|\le N\)); so \(\Theta_N(g)\in\overline{\mathcal{F}}\). For bounded measurable \(\varphi\) with \(N>\sup|\varphi|\), density of \(G\) in \(L^1(\mu)\) gives \(g_k\) with \(\|g_k-\varphi\|_{L^1(\mu)}\to0\), and \(|\Theta_N(g_k)-\varphi|\le\min(|g_k-\varphi|,2N)\) gives \[ \int_X |\Theta_N(g_k) - \varphi|^p\,d\mu \le (2N)^{p-1}\|g_k - \varphi\|_{L^1(\mu)} \to 0 , \] so \(\varphi\in\overline{\mathcal{F}}\). Finally \(\Theta_N(h)\to h\) in \(L^p(\mu)\) for every \(h\in L^p(\mu)\) (dominated convergence, \(|\Theta_N(h)-h|\le2|h|\)), so \(\overline{\mathcal{F}} = L^p(\mu)\).
(ii) Apply (i) with \(\mathcal{E} := \{(c,+\infty): c\in\mathbb{R}\}\) and the separating function \(h := f - f( c)\cdot1\), which lies in the algebra \(\mathcal{F} = \{P(f)\}\) and is \(>0\) on \((c,\infty)\), \(\le 0\) on \((-\infty,c]\) because \(f\) is strictly increasing (and Borel, being monotone). The span \(G\) of \(\{I_{(c,\infty)}\}\) is dense in \(L^1(\mu)\): its closure contains \(I_{(a,b]} = I_{(a,\infty)}-I_{(b,\infty)}\) and \(I_{\mathbb{R}} = \lim_{c\to-\infty}I_{(c,\infty)}\), hence all indicators of the semiring \(\mathcal{R}\) of intervals \((a,b]\), and by Exercise 4.7.70 it suffices that every Borel set be approximable in \(\mu(A\triangle B)\) by a finite union of such intervals – which is the \(\sigma\)-algebra argument of Exercise 4.7.64 applied to \(\mathcal{A}_0 := \mathcal{R}_\sigma\) (Lemma 1.2.14). Hence the algebra generated by \(f\) and \(1\) is dense in \(L^p(\mu)\) for every \(p\in[1,\infty)\).
(G. Hardy) Let \(f \in L^p(0, +\infty)\), where \(p > 1\). Show that the functions \[ \varphi(x) = \frac{1}{x}\int_0^x f(t)\,dt, \qquad \psi(x) = \int_x^{\infty} \frac{f(t)}{t}\,dt \] belong to \(L^p(0, +\infty)\) as well.
Hardy’s inequalities \[ \|\varphi\|_p \le \frac{p}{p-1}\,\|f\|_p , \qquad \|\psi\|_p \le p\,\|f\|_p \tag{1} \] hold, and in particular \(\varphi,\psi \in L^p(0,+\infty)\). Let \(q := p/(p-1) \in (1,\infty)\). Hoelder’s inequality gives \(\int_0^x|f|\,dt \le \|f\|_px^{1/q}\) and \(\int_x^\infty|f(t)|t^{-1}dt \le \|f\|_p(q-1)^{-1/q}x^{(1-q)/q}\) (finite as \(q>1\)), so \(\varphi,\psi\) are finite and continuous on \((0,\infty)\); replacing \(f\) by \(|f|\) we may assume \(f \ge 0\), whence \(\varphi,\psi\ge0\). (Modifying \(f\) on a null set is harmless: by Tonelli the sets \(\{(x,s): xs\in Z\}\) and \(\{(x,s): x/s\in Z\}\) are null for null \(Z\).)
The substitutions \(t=xs\) and \(t=x/s\), \(s\in(0,1)\), give \[ \varphi(x) = \int_0^1 f(xs)\,ds , \qquad \psi(x) = \int_0^1 \frac{f(x/s)}{s}\,ds , \tag{2} \] and the same substitutions in the \(x\)-integral give, for fixed \(s\in(0,1)\),
\begin{equation*} \Big(\int_0^{\infty} f(xs)^p\,dx\Big)^{1/p} = s^{-1/p}\|f\|_p , \qquad \Big(\int_0^{\infty} f(x/s)^p\,dx\Big)^{1/p} = s^{1/p}\|f\|_p . \tag{3} \end{equation*}
Let \(F\) be \(\varphi\) or \(\psi\) and \(F_n := FI_{[1/n,n]}\), which is bounded with support of finite measure, so \(\|F_n\|_p<\infty\); assuming \(\|F_n\|_p>0\) and setting \(g := F_n^{p-1}/\|F_n\|_p^{p-1} \ge 0\) one checks \(\|g\|_q = 1\) and \[ \|F_n\|_p = \int_0^{\infty} F_n g\,dx \le \int_0^{\infty} F g\,dx . \tag{4} \] By (2), Tonelli (nonnegative integrands), Hoelder’s inequality and (3),
\begin{equation*} \begin{aligned} \int_0^{\infty} \varphi(x)g(x)\,dx &= \int_0^{\infty}\!\!\Big(\int_0^1 f(xs)\,ds\Big)g(x)\,dx = \int_0^1\!\!\Big(\int_0^{\infty} f(xs)g(x)\,dx\Big)ds \\ &\le \int_0^1 \Big(\int_0^\infty f(xs)^p dx\Big)^{1/p}\|g\|_q\,ds = \|f\|_p\int_0^1 s^{-1/p}\,ds = \frac{p}{p-1}\,\|f\|_p , \end{aligned} \end{equation*}
where \(\int_0^1 s^{-1/p}ds = p/(p-1)\) is finite because \(p>1\). By (4) and monotone convergence (\(\varphi_n^p\uparrow\varphi^p\)), \(\|\varphi\|_p \le \frac{p}{p-1}\|f\|_p\). Similarly, by (2), Tonelli, Hoelder and (3),
\begin{equation*} \begin{aligned} \int_0^{\infty}\psi(x)g(x)\,dx &= \int_0^1 \frac{1}{s}\Big(\int_0^{\infty} f(x/s)g(x)\,dx\Big)ds \le \int_0^1 \frac{1}{s}\Big(\int_0^{\infty}f(x/s)^p dx\Big)^{1/p}\|g\|_q\,ds \\ &= \|f\|_p\int_0^1 s^{1/p - 1}\,ds = p\,\|f\|_p , \end{aligned} \end{equation*}
since \(\int_0^1 s^{1/p-1}ds = p\). By (4) and monotone convergence, \(\|\psi\|_p \le p\|f\|_p\), so (1) holds and \(\varphi,\psi \in L^p(0,+\infty)\).
(\(\circ\)) Let \(G\) be an everywhere dense set in \(L^q(\mu)\), \(p^{-1} + q^{-1} = 1\), \(q > 1\), and let a sequence \(\{f_n\}\) be bounded in the norm of \(L^p(\mu)\). Prove that this sequence weakly converges to \(f \in L^p(\mu)\) precisely when the integrals of \(f_n g\) converge to the integral of \(fg\) for every \(g \in G\).
Norm boundedness lets a \(3\varepsilon\)-argument transfer the convergence from \(G\) to all of \(L^q(\mu)\). Since \(p = q/(q-1)\in(1,\infty)\), every continuous linear functional on \(L^p(\mu)\) is generated by an element of \(L^q(\mu)\) (Theorem 4.4.1 for \(\sigma\)-finite \(\mu\), Exercise 4.7.87 in general), so weak convergence \(f_n \to f\) means \[ \lim_{n \to \infty} \int_X f_n h\,d\mu = \int_X f h\,d\mu \qquad \text{for every } h \in L^q(\mu) . \tag{1} \] Necessity is immediate from \(G\subset L^q(\mu)\). For sufficiency, put \(C := \sup_n\|f_n\|_{L^p(\mu)}<\infty\), fix \(h\in L^q(\mu)\) and \(\varepsilon>0\), and pick \(g\in G\) with \(\|h-g\|_{L^q(\mu)}<\varepsilon\); all integrals below are finite by Hoelder’s inequality. For every \(n\),
\begin{equation*} \begin{aligned} \Big| \int_X f_n h\,d\mu - \int_X f h\,d\mu \Big| &\le \Big| \int_X f_n (h - g)\,d\mu \Big| + \Big| \int_X f_n g\,d\mu - \int_X f g\,d\mu \Big| + \Big| \int_X f (g - h)\,d\mu \Big| \\ &\le \|f_n\|_{L^p}\,\|h-g\|_{L^q} + \Big| \int_X f_n g\,d\mu - \int_X f g\,d\mu \Big| + \|f\|_{L^p}\,\|h-g\|_{L^q} \\ &\le \big(C + \|f\|_{L^p(\mu)}\big)\,\varepsilon + \Big| \int_X f_n g\,d\mu - \int_X f g\,d\mu \Big| . \end{aligned} \end{equation*}
Letting \(n\to\infty\) and using the hypothesis on \(G\), \[ \limsup_{n \to \infty}\Big| \int_X f_n h\,d\mu - \int_X f h\,d\mu \Big| \le \big(C + \|f\|_{L^p(\mu)}\big)\,\varepsilon , \] and as \(\varepsilon>0\) was arbitrary, (1) holds for every \(h \in L^q(\mu)\).
Exercises 4.7.77–4.7.83
(\(\circ\)) Give an example of a sequence of functions \(f \in L^1[0,1]\) that is bounded in the norm of \(L^1[0,1]\) and converges a.e. to \(0\), but has no subsequence convergent in the weak topology of \(L^1[0,1]\).
Take \(f_n := n\,I_{[0,1/n]}\) on \([0,1]\) with Lebesgue measure \(\lambda\). Then \(\|f_n\|_{L^1}=1\) for all \(n\), and \(f_n(t)=0\) as soon as \(n>1/t\), so \(f_n \to 0\) at every \(t \in (0,1]\).
Since \(\lambda\) is finite, Theorem 4.4.1 identifies the dual of \(L^1[0,1]\) with \(L^\infty[0,1]\), so weak convergence \(f_{n_k}\to f\) would give, testing against \(g \equiv 1\) and against \(g = I_{[c,1]}\) (on which \(f_{n_k}\) vanishes once \(n_k>1/c\)),
\begin{equation*} \int_0^1 f \, d\lambda = 1 \qquad\text{and}\qquad \int_c^1 f\,d\lambda = 0 \ \ \text{for every } c \in (0,1) . \end{equation*}
Letting \(c\downarrow0\) in the second, by dominated convergence with dominant \(|f|\), forces \(\int_0^1f\,d\lambda=0\) – a contradiction. Hence no subsequence converges weakly.
(\(\circ\)) Let \(1 < p < \infty\). Construct an example of a sequence of functions \(f_n\) that weakly converges to zero in the space \(L^p[0,1]\) and converges to zero almost everywhere on \([0,1]\), but does not converge in the norm of \(L^p[0,1]\).
Take \(f_n := n^{1/p}I_{[0,1/n]}\) on \([0,1]\). Then \(\|f_n\|_{L^p}^p = n\cdot n^{-1} = 1\) for all \(n\), and \(f_n(x)=0\) once \(n>1/x\), so \(f_n \to 0\) a.e.
For weak convergence, note that \(L^\infty[0,1]\) is dense in the dual space \(L^q[0,1]\), \(q=p/(p-1)\) (truncations \(gI_{\{|g|\le N\}}\) converge in \(L^q\) by dominated convergence), and \(\{f_n\}\) is \(L^p\)-bounded, so by Exercise 4.7.76 it suffices to test against \(g \in L^\infty[0,1]\):
\begin{equation*} \left| \int_0^1 f_n g \, d\lambda \right| \le \|g\|_{L^\infty}\, n^{1/p - 1} \xrightarrow[n \to \infty]{} 0 , \end{equation*}
since \(1/p-1<0\). Norm convergence fails: the only possible norm limit is the weak limit \(0\) (norm convergence implies weak convergence, and weak limits are unique), yet \(\|f_n\|_{L^p}=1\).
(\(\circ\)) (i) (Riemann–Lebesgue theorem) Show that
\begin{equation*} \lim_{n \to \infty} \int_0^{2\pi} f(x) \sin nx \, dx = 0 \end{equation*}
for every Lebesgue integrable function \(f\).
(ii) Let \(\mu\) be a probability measure and let \(\{\varphi_n\}\) be an orthonormal system in \(L^2(\mu)\) such that \(|\varphi_n| \le M\), where \(M\) is a number. Show that
\begin{equation*} \lim_{n \to \infty} \int f \varphi_n \, d\mu = 0 \end{equation*}
for every \(\mu\)-integrable function \(f\).
(i) On an interval the integral is \(O(1/n)\), and the indicators of intervals span a dense subspace of \(L^1[0,2\pi]\). Indeed,
\begin{equation*} \left| \int_\alpha^\beta \sin nx \, dx \right| = \left| \frac{\cos n\alpha - \cos n\beta}{n} \right| \le \frac{2}{n} \to 0 , \end{equation*}
so the assertion holds for every step function; and step functions are dense in \(L^1[0,2\pi]\), since simple functions are (Lemma 4.2.1) and every measurable \(A\) satisfies \(\lambda(A \triangle B) < \delta\) for some finite union \(B\) of intervals. Hence, given \(f \in L^1[0,2\pi]\) and a step function \(s\) with \(\|f-s\|_{L^1}<\varepsilon/2\), the bound \(|\sin nx|\le1\) gives
\begin{equation*} \left| \int_0^{2\pi} f(x)\sin nx\,dx \right| \le \frac{\varepsilon}{2} + \left| \int_0^{2\pi} s(x)\sin nx\,dx\right| < \varepsilon \end{equation*}
for all sufficiently large \(n\).
(ii) For bounded \(f\), say \(|f| \le N\) a.e., one has \(f \in L^2(\mu)\) because \(\mu(X)=1\), so Bessel’s inequality (Theorem 4.3.6) gives \(\sum_n\big(\int_X f\varphi_n\,d\mu\big)^2 \le \|f\|^2_{L^2(\mu)}<\infty\) and the terms tend to \(0\). For general \(f \in L^1(\mu)\) put \(f_N := fI_{\{|f|\le N\}}\), so that \(\|f-f_N\|_{L^1(\mu)}\to0\) by dominated convergence; choosing \(N\) with \(\|f-f_N\|_{L^1(\mu)}<\varepsilon/(2(M+1))\) and using \(|\varphi_n|\le M\),
\begin{equation*} \left| \int_X f \varphi_n \, d\mu \right| \le M\|f - f_N\|_{L^1(\mu)} + \left| \int_X f_N \varphi_n\,d\mu\right| < \varepsilon \end{equation*}
for all large \(n\) by the bounded case.
Give an example of a sequence of nonnegative functions \(f_n\) that weakly converges in \(L^1[0,1]\) to a function \(f\) and \(\|f_n\|_{L^1} \to \|f\|_{L^1}\), but \(\{f_n\}\) does not converge in the norm of \(L^1[0,1]\).
Take \(f_n(x) := 1 + \sin(2\pi nx) \ge 0\) and \(f \equiv 1\) on \([0,1]\).
Weak convergence: the dual of \(L^1[0,1]\) is \(L^\infty[0,1]\) (Theorem 4.4.1), and for \(g \in L^\infty[0,1] \subset L^1[0,1]\) the Riemann–Lebesgue theorem of Exercise 4.7.79(i) gives \(\int_0^1 g(x)\sin(2\pi nx)\,dx \to 0\), i.e. \(\int_0^1 f_ng\,d\lambda \to \int_0^1 g\,d\lambda\). The norms agree exactly:
\begin{equation*} \|f_n\|_{L^1} = \int_0^1\bigl(1+\sin(2\pi nx)\bigr)dx = 1 + \frac{1-\cos(2\pi n)}{2\pi n} = 1 = \|f\|_{L^1} . \end{equation*}
But the substitution \(u = 2\pi nx\) and \(\int_0^{2\pi}|\sin u|\,du = 4\) give
\begin{equation*} \|f_n - f\|_{L^1} = \int_0^1 |\sin(2\pi nx)|\,dx = \frac{1}{2\pi n}\cdot n\int_0^{2\pi}|\sin u|\,du = \frac{2}{\pi} , \end{equation*}
and since norm convergence would force the limit to be the weak limit \(f\), the sequence has no \(L^1\)-limit.
Show that there exists a sequence of positive continuous functions \(f_n\) on \([0,1]\) and a continuous function \(f\) such that for all \(a,b \in [0,1]\) one has
\begin{equation*} \lim_{n \to \infty} \int_a^b f_n(t)\,dt = \int_a^b f(t)\,dt , \end{equation*}
but there is a measurable set \(E\) such that the integrals of \(f_n\) over \(E\) do not converge.
Take \(f \equiv 1\) and let \(E\) be a set splitting every interval, i.e.
\begin{equation*} 0 < \lambda(E \cap I) < \lambda(I) \quad \text{for every interval } I \subset [0,1] \text{ of positive length;} \tag{1} \end{equation*}
the \(f_n\) will be continuous approximations of densities concentrated alternately on \(E\) and on its complement.
Construction of \(E\). Enumerate by \(\{J_j\}\) the intervals with rational endpoints in \([0,1]\) and build inductively pairwise disjoint nowhere dense compact sets \(A_j, B_j \subset J_j\) of positive measure: the union \(F\) of the sets already built is closed and nowhere dense, so \(J_j\setminus F\) contains an interval, inside two disjoint subintervals of which one places Cantor-type sets of positive measure (the usual construction removing middle intervals of total length \(\le\) half). Then \(E := \bigcup_j A_j\) satisfies (1), since any interval \(I\) of positive length contains some \(J_j\), giving \(\lambda(E\cap I)\ge\lambda(A_j)>0\) and \(\lambda(I\setminus E)\ge\lambda(B_j)>0\).
Densities. Let \(I_{n,k} := [(k-1)/n, k/n]\), \(k \le n\); by (1) all \(\lambda(E\cap I_{n,k})\) and \(\lambda(I_{n,k}\setminus E)\) are positive, so
\begin{equation*} \varphi_n := \sum_{k=1}^n \frac{1/n}{\lambda(E\cap I_{n,k})} \, I_{E \cap I_{n,k}}, \qquad \psi_n := \sum_{k=1}^n \frac{1/n}{\lambda(I_{n,k}\setminus E)}\, I_{I_{n,k}\setminus E} \end{equation*}
are nonnegative with mass exactly \(1/n = \lambda(I_{n,k})\) on each \(I_{n,k}\), while \(\int_E\varphi_n\,d\lambda = 1\) and \(\int_E\psi_n\,d\lambda = 0\). Counting the at most two partition intervals meeting \([a,b]\) without lying in it, that equidistribution of mass gives, for \(\theta_n \in \{\varphi_n,\psi_n\}\),
\begin{equation*} \left| \int_a^b \theta_n \, d\lambda - (b - a) \right| \le \frac{2}{n} . \tag{2} \end{equation*}
Continuous versions. Put \(\theta_n := \varphi_n\) for odd \(n\) and \(\theta_n := \psi_n\) for even \(n\). Continuous functions are dense in \(L^1[0,1]\) (Corollary 4.2.2 applied on \(\mathbb{R}^1\) after extension by zero), so there is a continuous \(g_n\) with \(\|g_n-\theta_n\|_{L^1}<1/(2n)\); since \(\theta_n\ge0\), the function \(|g_n|\) does as well, and \(f_n := |g_n| + 1/(2n)\) is continuous, strictly positive, with
\begin{equation*} \|f_n - \theta_n\|_{L^1} \le \| |g_n| - \theta_n\|_{L^1} + \frac{1}{2n} < \frac1n . \tag{3} \end{equation*}
Combining (2) and (3), \(|\int_a^b f_n\,dt - \int_a^b f\,dt| < 3/n \to 0\) for all \(a,b\). But by (3) again, \(\int_E f_n\,d\lambda > 1 - 1/n\) for odd \(n\) and \(\int_E f_n\,d\lambda < 1/n\) for even \(n\), so \(\int_E f_n\,d\lambda\) does not converge.
Let \(\mu\) be a measure with values in \([0,+\infty]\) on a space \((X,\mathcal{A})\). The following terminology is used in the books Hunt [448] and Bauer [70]: a set \(M\) in \(\mathcal{L}^1(\mu)\) (or in \(L^1(\mu)\)) is called uniformly integrable if
\begin{equation*} \forall\, \varepsilon > 0 \ \ \exists\, g \in \mathcal{L}^1(\mu): \quad \int_{\{|f| > g\}} |f|\,d\mu \le \varepsilon, \quad \forall\, f \in M. \tag{4.7.18} \end{equation*}
With such a definition, any integrable function is uniformly integrable.
(i) Show that (4.7.18) yields the existence of a measurable set \(E\) such that the measure \(\mu\) on \(E\) is \(\sigma\)-finite and every function \(f \in M\) vanishes a.e. outside \(E\).
(ii) Show that for finite measures (4.7.18) is equivalent to the uniform integrability in our sense.
(iii) Show that (4.7.18) is equivalent to the following property: the set \(M\) is bounded in the norm of \(L^1(\mu)\) and, for every \(\varepsilon > 0\), there exist a nonnegative integrable function \(h\) and a number \(\delta > 0\) such that, whenever \(A \in \mathcal{A}\) and
\begin{equation*} \int_A h \, d\mu \le \delta , \end{equation*}
one has
\begin{equation*} \int_A |f|\,d\mu \le \varepsilon \quad \text{for all } f \in M . \end{equation*}
(iv) Let the measure \(\mu\) be \(\sigma\)-finite and let \(h > 0\) be a \(\mu\)-integrable function. Show that (4.7.18) is equivalent to the property that, for every \(\varepsilon > 0\), there exists \(C > 0\) such that
\begin{equation*} \int_{\{|f| > Ch\}} |f| \, d\mu \le \varepsilon, \quad \forall\, f \in M . \end{equation*}
In addition, (4.7.18) is equivalent to the following: the set \(M\) is bounded in the norm of \(L^1(\mu)\) and, for every \(\varepsilon > 0\), there exists a number \(\delta > 0\) such that if \(A \in \mathcal{A}\) and
\begin{equation*} \int_A h \, d\mu \le \delta , \end{equation*}
then
\begin{equation*} \int_A |f| \, d\mu \le \varepsilon \quad \text{for all } f \in M . \end{equation*}
(v) Prove that (4.7.18) is equivalent to the following: \(M\) is bounded in \(L^1(\mu)\), the functions in \(M\) have uniformly absolutely continuous integrals and, for every \(\varepsilon > 0\), there exists a measurable set \(X_\varepsilon\) such that \(\mu(X_\varepsilon) < \infty\) and
\begin{equation*} \int_{X \setminus X_\varepsilon} |f| \, d\mu \le \varepsilon \quad \text{for all } f \in M . \end{equation*}
Every part turns on replacing the majorant \(g\) of (4.7.18) by a multiple of a fixed integrable function, via Chebyshev’s inequality. Four preliminaries, used throughout: (a) one may assume \(g\ge0\), since \(\{|f|>|g|\}\subset\{|f|>g\}\); (b) a single integrable \(g\ge0\) has absolutely continuous integral even for infinite \(\mu\) (dominated convergence gives \(C\) with \(\int_{\{g>C\}}g\,d\mu<\varepsilon/2\), then \(\int_Ag\,d\mu\le\varepsilon/2+C\mu(A)\)); (c) \(\int_{\{g\le\eta\}}g\,d\mu\to0\) as \(\eta\downarrow0\), again by dominated convergence; (d) (4.7.18) makes \(M\) bounded in \(L^1(\mu)\), since with \(\varepsilon=1\),
\begin{equation*} \|f\|_1 = \int_{\{|f| > g\}} |f|\,d\mu + \int_{\{|f| \le g\}} |f| \, d\mu \le 1 + \|g\|_1 =: S . \end{equation*}
(i) Take \(g_n \ge 0\) realizing (4.7.18) with \(\varepsilon = 1/n\) and put \(E := \bigcup_n\{g_n>0\}\). Each \(\{g_n>1/k\}\) has finite measure by Chebyshev, so \(\mu\) is \(\sigma\)-finite on \(E\). All \(g_n\) vanish off \(E\), so \((X\setminus E)\cap\{|f|>0\} \subset \{|f|>g_n\}\) and
\begin{equation*} \int_{X \setminus E} |f| \, d\mu \le \int_{\{|f| > g_n\}} |f| \, d\mu \le \frac1n \to 0 . \end{equation*}
(ii) Given (4.7.18) with \(\varepsilon/2\) and \(C_0\) with \(\int_{\{g>C_0\}}g\,d\mu\le\varepsilon/2\) (remark (b) applied to the single function \(g\)), any \(C \ge C_0\) satisfies \(\{|f|>C\} \subset \{|f|>g\} \cup (\{g>C_0\}\cap\{|f|\le g\})\), whence
\begin{equation*} \int_{\{|f| > C\}} |f| \, d\mu \le \int_{\{|f|>g\}} |f| \, d\mu + \int_{\{g > C_0\}} g\,d\mu \le \varepsilon \end{equation*}
for all \(f\in M\): this is Definition 4.5.1 (and holds for any \(\mu\)). Conversely, if \(\mu\) is finite and \(M\) is uniformly integrable in that sense, then \(g :\equiv C\) is integrable and \(\{|f|>g\} = \{|f|>C\}\).
(iii) Given (4.7.18) with \(\varepsilon/2\), put \(h := g\) and \(\delta := \varepsilon/2\); if \(\int_A h\,d\mu \le \delta\) then
\begin{equation*} \int_A |f|\,d\mu \le \int_{\{|f|>g\}} |f|\,d\mu + \int_A g \, d\mu \le \varepsilon . \end{equation*}
Conversely, let \(S := \sup_{f\in M}\|f\|_1<\infty\) (assume \(S>0\)) and take \(h,\delta\) for \(\varepsilon/2\). Setting \(C := S/\delta\) and \(g := Ch\), the set \(A_f := \{|f|>Ch\}\) satisfies \(C\int_{A_f}h\,d\mu \le \int_{A_f}|f|\,d\mu \le S\), i.e. \(\int_{A_f}h\,d\mu\le\delta\), so \(\int_{\{|f|>g\}}|f|\,d\mu \le \varepsilon/2\) for every \(f \in M\).
(iv) Write (a) for the property with \(Ch\) and (b) for the \(L^1\)-bounded property with \(\delta\); we show (a) \(\Rightarrow\) (4.7.18) \(\Rightarrow\) (b) \(\Rightarrow\) (a). The first is immediate with \(g := Ch\). For the second, boundedness is remark (d); with \(g\ge0\) realizing \(\varepsilon/2\) and \(\nu := h\cdot\mu\) (a finite measure, and \(h>0\) makes \(\nu\) and \(\mu\) have the same null sets), one has \(\int_Ag\,d\mu = \int_A(g/h)\,d\nu\) with \(g/h \in \mathcal{L}^1(\nu)\), so absolute continuity of that single integral gives \(\delta>0\) with \(\nu(A)\le\delta \Rightarrow \int_A(g/h)\,d\nu\le\varepsilon/2\), and then
\begin{equation*} \int_A |f|\,d\mu \le \int_{\{|f|>g\}}|f|\,d\mu + \int_A g \, d\mu \le \varepsilon . \end{equation*}
The third is the Chebyshev argument of (iii) verbatim, with \(C := S/\delta\).
(v) Assume (4.7.18); boundedness is remark (d), and with \(g\ge0\) realizing \(\varepsilon/2\), remark (b) gives \(\delta>0\) with \(\int_Ag\,d\mu<\varepsilon/2\) for \(\mu(A)<\delta\), whence \(\int_A|f|\,d\mu<\varepsilon\): the integrals are uniformly absolutely continuous (Definition 4.5.2). For the third condition, remark (c) gives \(\eta>0\) with \(\int_{\{g\le\eta\}}g\,d\mu\le\varepsilon/2\); then \(X_\varepsilon := \{g>\eta\}\) has \(\mu(X_\varepsilon)\le\eta^{-1}\|g\|_1<\infty\) by Chebyshev and
\begin{equation*} \int_{X\setminus X_\varepsilon} |f|\,d\mu \le \int_{\{|f| > g\}} |f|\,d\mu + \int_{\{g\le\eta\}} g \, d\mu \le \varepsilon . \end{equation*}
Conversely, given the three conditions, take \(X_\varepsilon\) with \(\int_{X\setminus X_\varepsilon}|f|\,d\mu\le\varepsilon/2\), then \(\delta>0\) from uniform absolute continuity for \(\varepsilon/2\), and put \(C := 2S/\delta\), \(g := CI_{X_\varepsilon}\) (integrable since \(\mu(X_\varepsilon)<\infty\)). Splitting
\begin{equation*} \{|f| > g\} = \bigl( \{|f| > C\}\cap X_\varepsilon \bigr) \cup \bigl( \{|f| > 0\}\setminus X_\varepsilon\bigr) , \end{equation*}
the second piece contributes at most \(\varepsilon/2\), and the first has measure \(\le S/C = \delta/2<\delta\) by Chebyshev, hence also contributes at most \(\varepsilon/2\). So (4.7.18) holds.
Let \(0 < p < q < \infty\) and let \(\mu\) be a countably additive measure with values in \([0,+\infty]\). (i) Prove that \(L^p(\mu) \not\subset L^q(\mu)\) precisely when there exist sets of arbitrarily small positive \(\mu\)-measure. (ii) Prove that \(L^q(\mu)\not\subset L^p(\mu)\) precisely when there exist sets of arbitrarily large finite \(\mu\)-measure.
Write \(m := \inf\{\mu(A): \mu(A)>0\}\) and \(R := \sup\{\mu(A): \mu(A)<\infty\}\), so the two conditions read \(m=0\) and \(R=+\infty\); in both parts the counterexample is a weighted sum of indicators of a disjoint sequence extracted from sets of nearly extremal measure.
(i) If \(m>0\), then \(f \in L^p(\mu)\) is essentially bounded: Chebyshev gives \(\mu(\{|f|>t\}) \le t^{-p}\|f\|_p^p \to 0\), so \(\mu(\{|f|>t_0\})<m\), hence \(=0\), for some \(t_0\), and then
\begin{equation*} \int_X |f|^q \, d\mu \le t_0^{\,q-p}\int_X |f|^p\,d\mu < \infty . \end{equation*}
If \(m=0\), choose \(A_1\) with \(0<\mu(A_1)\le1/2\) and inductively \(A_{n+1}\) with \(0<\mu(A_{n+1})\le\min\{\frac14\mu(A_n), 2^{-(n+1)}\}\), so that \(\sum_{k>n}\mu(A_k)\le\mu(A_n)/3\). The sets \(B_n := A_n\setminus\bigcup_{k>n}A_k\) are pairwise disjoint with
\begin{equation*} c_n := \mu(B_n) \ge \mu(A_n) - \sum_{k>n}\mu(A_k) \ge \tfrac{2}{3}\mu(A_n) > 0 \end{equation*}
and \(c_n \le 2^{-n} \to 0\). With \(\theta := 1-p/q \in (0,1)\), pick \(n_1<n_2<\cdots\) with \(c_{n_k}^\theta \le 2^{-k}\) and set \(f := \sum_k c_{n_k}^{-1/q}I_{B_{n_k}}\). Then
\begin{equation*} \int_X |f|^q \, d\mu = \sum_{k} c_{n_k}^{-1}c_{n_k} = +\infty , \qquad \int_X |f|^p \, d\mu = \sum_{k} c_{n_k}^{\,\theta} \le 1 , \end{equation*}
so \(f \in L^p(\mu)\setminus L^q(\mu)\).
(ii) If \(R<\infty\), then \(\{|f|>0\} = \bigcup_n\{|f|>1/n\}\) is an increasing union of sets of finite measure (Chebyshev), so \(\mu(\{|f|>0\})\le R\), and splitting at \(|f|=1\), where \(|f|^p\le1\) resp. \(|f|^p\le|f|^q\),
\begin{equation*} \int_X |f|^p\,d\mu \le \mu(\{|f|>0\}) + \int_X |f|^q\,d\mu \le R + \int_X |f|^q\,d\mu < \infty . \end{equation*}
If \(R=+\infty\), choose \(A_1\) with \(1 \le \mu(A_1)<\infty\) and inductively \(A_{n+1}\) of finite measure with \(\mu(A_{n+1}) \ge 2\sum_{k\le n}\mu(A_k) + 2^{n+1}\). The sets \(B_n := A_n\setminus\bigcup_{k<n}A_k\) are pairwise disjoint with
\begin{equation*} c_n := \mu(B_n) \ge \mu(A_n) - \sum_{k<n}\mu(A_k) \ge \tfrac{1}{2}\mu(A_n) \ge 2^{n-1} \end{equation*}
and \(c_n \le \mu(A_n)<\infty\). With \(\sigma := q/p - 1 > 0\), pick \(n_1<n_2<\cdots\) with \(c_{n_k}^{-\sigma}\le2^{-k}\) and set \(f := \sum_k c_{n_k}^{-1/p}I_{B_{n_k}}\). Then
\begin{equation*} \int_X |f|^p\,d\mu = \sum_{k} c_{n_k}^{-1}c_{n_k} = +\infty , \qquad \int_X |f|^q \, d\mu = \sum_{k} c_{n_k}^{-\sigma} \le 1 , \end{equation*}
so \(f \in L^q(\mu)\setminus L^p(\mu)\).
Exercises 4.7.84–4.7.90
Let \(f\) and \(g\) be integrable on \([0,1]\) and let \(|f(x)| \le g(x)\). Prove that there exists a sequence of integrable functions \(f_n\) such that, for every measurable set \(E \subset [0,1]\), one has
\begin{equation*} \lim_{n\to\infty} \int_E f_n\,dx = \int_E f\,dx, \qquad \lim_{n\to\infty} \int_E |f_n|\,dx = \int_E g\,dx . \end{equation*}
Take \(f_n := \varepsilon_n g\) with \(\varepsilon_n\colon[0,1]\to\{-1,1\}\) chosen so that \(\varepsilon_n \to \alpha\) in the weak sense against \(L^1\), where \(\alpha := f/g\) on \(\{g>0\}\) and \(\alpha := 0\) elsewhere. Then \(|\alpha|\le1\) and \(\alpha g = f\) a.e. (on \(\{g=0\}\) both sides vanish since \(|f|\le g\)), while \(|f_n| = g\), which already gives \(\int_E|f_n|\,dx = \int_Eg\,dx\) for every \(E\) and every \(n\).
Construction of \(\varepsilon_n\). On the dyadic interval \(I_{n,k} := [k2^{-n},(k+1)2^{-n})\) of generation \(n\), put \(\alpha_{n,k} := 2^n\int_{I_{n,k}}\alpha\,dx \in [-1,1]\) and let \(\varepsilon_n := 1\) on the left portion of \(I_{n,k}\) of length \(\ell_{n,k} := 2^{-n}(1+\alpha_{n,k})/2\) and \(\varepsilon_n := -1\) on the rest. Then
\begin{equation*} \int_{I_{n,k}} \varepsilon_n \,dx = 2\ell_{n,k} - 2^{-n} = 2^{-n}\alpha_{n,k} = \int_{I_{n,k}} \alpha \,dx , \end{equation*}
and summing over the generation-\(n\) subintervals of a dyadic \(J\) of generation \(m \le n\),
\begin{equation*} \int_J \varepsilon_n\,dx = \int_J \alpha\,dx \qquad \text{for every dyadic } J \text{ of generation } m \le n. \tag{1} \end{equation*}
Given \(\varphi \in L^1[0,1]\) and \(\delta>0\), Corollary 4.2.2 gives a continuous \(u\) with \(\|\varphi-u\|_{L^1}<\delta\), and uniform continuity makes the dyadic step function \(\psi := \sum_k u(k2^{-m})I_{I_{m,k}}\) satisfy \(\|u-\psi\|_{L^1}<\delta\) for a suitable \(m\); so \(\|\varphi-\psi\|_{L^1}<2\delta\). By (1), \(\int_0^1\varepsilon_n\psi\,dx = \int_0^1\alpha\psi\,dx\) for \(n \ge m\), whence, using \(|\varepsilon_n|\le1\) and \(|\alpha|\le1\),
\begin{equation*} \Bigl|\int_0^1 (\varepsilon_n - \alpha)\varphi \,dx\Bigr| = \Bigl|\int_0^1 (\varepsilon_n - \alpha)(\varphi - \psi) \,dx\Bigr| \le 2\|\varphi - \psi\|_{L^1} < 4\delta . \end{equation*}
Applying this to \(\varphi := gI_E \in L^1[0,1]\),
\begin{equation*} \int_E f_n \,dx = \int_0^1 \varepsilon_n\, g I_E \,dx \xrightarrow[n\to\infty]{} \int_E \alpha g\,dx = \int_E f \,dx . \end{equation*}
(\(\circ\)) Suppose that functions \(f_n\) weakly converge in \(L^p(\mu)\) to a function \(f\), where \(p \ge 1\). Show that \(\|f\|_p \le \liminf\limits_{n\to\infty}\|f_n\|_p\).
Test against a norm-one functional attaining the norm at \(f\); assume \(\|f\|_p>0\), as otherwise there is nothing to prove.
- (i) \(1<p<\infty\), \(q = p/(p-1)\): take \(g := \|f\|_p^{1-p}\operatorname{sign} f\cdot|f|^{p-1}\). Then \(|g|^q = \|f\|_p^{-p}|f|^p\) (as \((p-1)q = p\)), so \(\|g\|_q = 1\), and \(\int fg\,d\mu = \|f\|_p^{1-p}\|f\|_p^p = \|f\|_p\).
- (ii) \(p=1\): take \(g := \operatorname{sign} f\), so \(\|g\|_\infty\le1\) and \(\int fg\,d\mu = \|f\|_1\).
In both cases \(\Lambda(h) := \int hg\,d\mu\) is a continuous linear function with \(|\Lambda(h)| \le \|h\|_p\|g\|_q \le \|h\|_p\) by Hoelder’s inequality, and \(\Lambda(f) = \|f\|_p\). Weak convergence therefore gives
\begin{equation*} \|f\|_p = \lim_{n\to\infty}\Lambda(f_n) \le \liminf_{n\to\infty}|\Lambda(f_n)| \le \liminf_{n\to\infty}\|f_n\|_p . \end{equation*}
(\(\circ\)) Let \(1 < p < \infty\), \(p^{-1} + q^{-1} = 1\), let \(\mu \ge 0\) be a \(\sigma\)-finite measure, and let \(\Psi\) be a continuous linear function on \(L^p(\mu)\). Let \(f \in L^p(\mu)\) be a function such that \(\|f\|_p = 1\) and \(\Psi(f) = \|\Psi\| > 0\). Prove that \(\Psi\) is given by the function \(g = \operatorname{sign} f \cdot |f|^{p-1} \in L^q(\mu)\) by formula (4.4.1) and that \(g\) is a unique function generating \(\Psi\).
The generating function is \(\widetilde g = \|\Psi\|\,\operatorname{sign} f\cdot|f|^{p-1}\), obtained from the equality cases in Hoelder’s and Young’s inequalities; the printed statement omits the factor \(\|\Psi\|\), which is forced since a generating function has \(L^q\)-norm \(\|\Psi\|\) while \(\|g\|_q = 1\).
Put \(B := \|\Psi\|>0\) and \(g := \operatorname{sign} f\cdot|f|^{p-1}\); then \((p-1)q = p\) gives \(\|g\|_q^q = \int|f|^p d\mu = 1\). Since \(\mu\) is \(\sigma\)-finite and \(1<p<\infty\), Theorem 4.4.1 supplies \(\widetilde g \in L^q(\mu)\) with \(\Psi(h) = \int h\widetilde g\,d\mu\) and \(\|\widetilde g\|_q = B\). Hence
\begin{equation*} B = \Psi(f) = \int f \widetilde g \,d\mu \le \int |f|\,|\widetilde g|\,d\mu \le \|f\|_p\|\widetilde g\|_q = B , \end{equation*}
so both inequalities are equalities. The first forces the nonnegative integrand \(|f\widetilde g| - f\widetilde g\) to vanish, i.e.
\begin{equation*} f\widetilde g \ge 0 \quad \text{a.e.} \tag{1} \end{equation*}
For the second, Young’s inequality \(st \le s^p/p + t^q/q\) (\(s,t\ge0\)), with equality exactly when \(s^p = t^q\), applied pointwise with \(s = |f|\), \(t = |\widetilde g|/B\) and integrated gives
\begin{equation*} \frac{1}{B}\int |f|\,|\widetilde g|\,d\mu \le \frac{1}{p}\int |f|^p\,d\mu + \frac{1}{q B^q}\int |\widetilde g|^q\,d\mu = 1 , \end{equation*}
whose two sides are both \(1\); so the nonnegative function \(|f|^p/p + |\widetilde g|^q/(qB^q) - |f||\widetilde g|/B\) vanishes a.e. and the equality case yields
\begin{equation*} |\widetilde g| = B |f|^{p/q} = B|f|^{p-1} \quad \text{a.e.} \tag{2} \end{equation*}
On \(\{f\ne0\}\), (2) makes \(\widetilde g\) nonzero and (1) matches its sign with that of \(f\); on \(\{f=0\}\), (2) gives \(\widetilde g = 0\). Hence \(\widetilde g = Bg\) a.e.
Uniqueness: if \(g_1,g_2\) both generate \(\Psi\), then \(u := g_1-g_2\) satisfies \(\int hu\,d\mu = 0\) for all \(h\in L^p(\mu)\), and \(h := \operatorname{sign} u\cdot|u|^{q-1}\) lies in \(L^p(\mu)\) (since \(|h|^p = |u|^q\)), giving \(\int|u|^q d\mu = 0\), i.e. \(g_1 = g_2\).
Let \(\mu\) be a countably additive measure on a \(\sigma\)-algebra with values in \([0,+\infty]\). (i) Show that for any nonzero continuous linear function \(\Psi\) on \(L^p(\mu)\) with \(1 < p < \infty\), there exists \(f \in L^p(\mu)\) with \(\|f\|_p = 1\) and \(\Psi(f) = \|\Psi\|\).
(ii) Prove that in the case \(1 < p < \infty\) the dual to \(L^p(\mu)\) can be identified with \(L^q(\mu)\), \(q = p/(p-1)\), in the same sense as in Theorem 4.4.1.
(iii) Extend the assertion of Exercise 4.7.86 to the case of an arbitrary (not necessarily \(\sigma\)-finite) countably additive measure with values in \([0,+\infty]\).
(i) Uniform convexity of \(L^p(\mu)\) (Theorem 4.7.15, stated for measures with values in \([0,+\infty]\)) makes any maximizing sequence fundamental. Normalize \(\|\Psi\|=1\) and take \(f_n\) with \(\|f_n\|_p = 1\) and \(\Psi(f_n)\to1\) (rescale a sequence with \(\|f_n\|_p\le1\), \(\Psi(f_n)\to1\), which forces \(\|f_n\|_p\to1\)). Given \(\varepsilon>0\), let \(\delta>0\) come from uniform convexity (Definition 4.7.14) and take \(N\) with \(\Psi(f_n)>1-\delta\) for \(n \ge N\); then for \(n,m\ge N\),
\begin{equation*} \Bigl\|\frac{f_n + f_m}{2}\Bigr\|_p \ge \Psi\Bigl(\frac{f_n+f_m}{2}\Bigr) = \frac{\Psi(f_n)+\Psi(f_m)}{2} > 1-\delta , \end{equation*}
so \(\|f_n-f_m\|_p\le\varepsilon\). By completeness (Theorem 4.1.3) \(f_n \to f\) in norm, and continuity gives \(\|f\|_p = 1\), \(\Psi(f) = 1 = \|\Psi\|\).
(ii) Each \(g \in L^q(\mu)\) gives \(\Phi_g(h) := \int hg\,d\mu\) with \(\|\Phi_g\| = \|g\|_q\): Hoelder gives \(\le\), and \(h_0 := \|g\|_q^{1-q}\operatorname{sign} g\cdot|g|^{q-1}\) has \(\|h_0\|_p = 1\) and \(\Phi_g(h_0) = \|g\|_q\); the map \(g\mapsto\Phi_g\) is linear and isometric, hence injective. For surjectivity let \(\Psi \ne 0\) with \(\|\Psi\|=1\), take \(f\) from (i) and put \(g := \operatorname{sign} f\cdot|f|^{p-1}\), so \(\|g\|_q = 1\) and \(\Phi_g(f) = 1\). Fix \(h \in L^p(\mu)\) and set \(\Omega := \{f\ne0\}\cup\{h\ne0\}\). Chebyshev makes each \(\{|f|>1/n\}\) and \(\{|h|>1/n\}\) of finite measure, so
\begin{equation*} \Omega = \bigcup_{n=1}^{\infty}\bigl(\{|f| > 1/n\}\cup\{|h|>1/n\}\bigr) \end{equation*}
carries a \(\sigma\)-finite restriction \(\mu_\Omega\), and extension by zero identifies \(L^p(\mu_\Omega)\) with the closed subspace of functions vanishing off \(\Omega\), which contains \(f,h\) and (as \(\{f\ne0\}\subset\Omega\)) supports \(g\). The restriction \(\Psi_\Omega\) has \(\|\Psi_\Omega\| = 1\), attained at \(f\), so Exercise 4.7.86 on \((\Omega,\mathcal{A}_\Omega,\mu_\Omega)\) gives \(\Psi_\Omega = \Phi_{g|_\Omega}\) and hence
\begin{equation*} \Psi(h) = \int_\Omega h g \,d\mu = \int_X h g\,d\mu = \Phi_g(h) , \end{equation*}
since \(hg\) vanishes off \(\Omega\). For general \(\Psi\ne0\) this gives \(\Psi = \Phi_{\|\Psi\|g}\) with \(\|\,\|\Psi\|g\,\|_q = \|\Psi\|\), so \(g\mapsto\Phi_g\) is an isometric linear bijection of \(L^q(\mu)\) onto \(L^p(\mu)^{*}\).
(iii) Part (ii) supplies the unique \(\widetilde g \in L^q(\mu)\) generating \(\Psi\), with \(\|\widetilde g\|_q = \|\Psi\|\), for an arbitrary \(\mu\); the remainder of the solution of Exercise 4.7.86 uses only Hoelder’s and Young’s inequalities together with \(\int f\widetilde g\,d\mu = \|f\|_p\|\widetilde g\|_q\), so it applies verbatim and yields
\begin{equation*} \widetilde g = \|\Psi\|\,\operatorname{sign} f\cdot|f|^{p-1} \quad \text{a.e.} \end{equation*}
Let \(\mu\) be a nonnegative measure, \(1 \le p \le \infty\), and let \(L\) be a linear function on \(L^p(\mu)\) such that \(L(f) \ge 0\) whenever \(f \ge 0\). Prove the continuity of \(L\).
Positivity forces boundedness, hence continuity. Positivity gives monotonicity: \(u \le v\) a.e. implies \(L(v) - L(u) = L(v-u) \ge 0\).
If \(L\) were unbounded, choose \(v_n\) with \(\|v_n\|_p \le 1\) and \(|L(v_n)| \ge 4^{n}\), and put \(\varphi_n := 4^{-n}|v_n| \ge 0\). Since \(\varphi_n \ge 4^{-n}(\operatorname{sign}L(v_n))v_n\), monotonicity yields
\begin{equation*} \|\varphi_n\|_p \le 4^{-n}, \qquad L(\varphi_n) \ge 4^{-n}|L(v_n)| \ge 1. \tag{1} \end{equation*}
We produce a single \(G \in L^p(\mu)\) dominating every partial sum \(\sum_{n\le k}\varphi_n\); then monotonicity, linearity and (1) give the absurdity
\begin{equation*} L(G) \ge \sum_{n=1}^{k}L(\varphi_n) \ge k \qquad \text{for all } k \in \mathbb{N}. \end{equation*}
(i) \(p < \infty\). By Beppo Levi’s theorem on term-by-term integration of series of nonnegative functions, \(g := \sum_{n\ge1}2^{np}\varphi_n^{p}\) is integrable, since
\begin{equation*} \sum_{n=1}^{\infty}\int 2^{np}\varphi_n^{p}\,d\mu \le \sum_{n=1}^{\infty} 2^{np}4^{-np} = \sum_{n=1}^{\infty}2^{-np} < \infty \end{equation*}
(convergent as \(p \ge 1\)). Set \(G := g^{1/p}\), so \(\int G^p\,d\mu < \infty\) and \(G \in L^p(\mu)\). Nonnegativity of the terms gives \(2^{np}\varphi_n^{p} \le g\), i.e. \(\varphi_n \le 2^{-n}G\) a.e. for each \(n\); off the union of these countably many null sets, \(\sum_{n\le k}\varphi_n \le G\sum_{n\le k}2^{-n} \le G\) for all \(k\).
(ii) \(p = \infty\). By (1) we may take representatives with \(0 \le \varphi_n \le 4^{-n}\) everywhere, so \(G := \sum_{n\ge1}\varphi_n\) satisfies \(0 \le G \le 1/3\), whence \(G \in L^\infty(\mu)\) and \(\sum_{n\le k}\varphi_n \le G\) for every \(k\).
Construct an example of a countably additive measure \(\mu\) with values in \([0,+\infty]\) defined on a \(\sigma\)-algebra \(\mathcal{A}\) such that there exists a continuous linear function \(\Psi\) on \(L^1(\mu)\) that cannot be written in the form indicated in Theorem 4.4.1.
Take counting measure \(\mu\) on the countable-cocountable \(\sigma\)-algebra
\begin{equation*} \mathcal{A} := \{A \subset [0,1]: A \text{ or } [0,1]\setminus A \text{ is at most countable}\} \end{equation*}
(a \(\sigma\)-algebra, and \(\mu\) is countably additive, with \(\emptyset\) its only null set; Check!), and \(\Psi(f) := \sum_{t \le 1/2}f(t)\).
If \(f\) is \(\mathcal{A}\)-measurable with \(\int|f|\,d\mu < \infty\), then each \(\{|f| \ge 1/n\}\) is finite, so \(\{f \ne 0\}\) is at most countable; conversely any \(f\) vanishing off an at most countable set is \(\mathcal{A}\)-measurable. Hence
\begin{equation*} L^1(\mu) = \Bigl\{f : \{f\ne 0\}\text{ at most countable}, \ \sum_t|f(t)| < \infty\Bigr\}, \quad \|f\|_1 = \sum_t|f(t)|, \end{equation*}
so the defining sum for \(\Psi\) converges absolutely and \(|\Psi(f)| \le \|f\|_1\): \(\Psi\) is continuous, with \(\Psi(I_{\{0\}}) = 1\) giving \(\|\Psi\| = 1\).
Suppose \(\Psi(f) = \int fg\,d\mu = \sum_t f(t)g(t)\) for some \(\mathcal{A}\)-measurable \(g\) as in Theorem 4.4.1 (\(p=1\), \(q=\infty\)). Testing on \(f = I_{\{s\}} \in L^1(\mu)\) gives \(g(s) = \Psi(I_{\{s\}})\), i.e. \(g = I_{[0,1/2]}\) everywhere, since \(\emptyset\) is the only null set. But \(g^{-1}(\{1\}) = [0,1/2]\) is neither at most countable nor co-countable, so \(g\) is not \(\mathcal{A}\)-measurable: contradiction. (Theorem 4.4.1 is not applicable, since sets of finite \(\mu\)-measure are finite and so \(\mu\) is not \(\sigma\)-finite.)
(i) Construct a space \((X,\mathcal{A},\mu)\) with a countably additive measure \(\mu\) with values in \([0,+\infty]\) and an \(\mathcal{A}\)-measurable function \(f\) that belongs to no \(L^p(\mu)\) with \(p \in [1,+\infty)\), but \(fg \in L^1(\mu)\) for every function \(g \in \bigcup_{q \ge 1}L^q(\mu)\).
(ii) Show that if a space \((X,\mathcal{A},\mu)\) with a countably additive measure \(\mu\) with values in \([0,+\infty]\) and an \(\mathcal{A}\)-measurable function \(f\) are such that \(\mu\) is \(\sigma\)-finite on the set \(\{f \ne 0\}\) and \(fg \in L^1(\mu)\) for every function \(g \in L^q(\mu)\), where \(1 < q \le \infty\), then \(f \in L^p(\mu)\), where \(p^{-1}+q^{-1} = 1\).
(iii) Let a measure \(\mu\) on a measurable space \((X,\mathcal{A})\) be semifinite in the sense of Exercise 1.12.132, let \(f\) be an \(\mathcal{A}\)-measurable function, and let \(p^{-1}+q^{-1}=1\), where \(1 \le p < \infty\). Suppose that \(fg \in L^1(\mu)\) for every function \(g \in L^q(\mu)\). Show that \(f \in L^p(\mu)\).
In (i) the union is over finite exponents \(q \in [1,+\infty)\).
(i) Take \(X = [0,1]\), \(\mathcal{A} = 2^X\), \(\mu(A) = +\infty\) for \(A \ne \emptyset\) and \(\mu(\emptyset) = 0\), and \(f \equiv 1\). Then \(\mu\) is countably additive with \(\emptyset\) its only null set (Check!), and for measurable \(h \ge 0\) not identically \(0\) there is \(c>0\) with \(\{h>c\} \ne \emptyset\), so \(\int h\,d\mu \ge c\,\mu(\{h>c\}) = +\infty\). Hence \(L^p(\mu) = \{0\}\) for every \(p \in [1,+\infty)\): in particular \(f \notin L^p(\mu)\), while every \(g \in \bigcup_{q\ge1}L^q(\mu)\) is the zero class, so \(fg = 0 \in L^1(\mu)\).
(ii) Put \(S := \{f \ne 0\}\). For \(q = \infty\), \(p = 1\), no \(\sigma\)-finiteness is needed: \(g := \operatorname{sign} f \in L^\infty(\mu)\), so \(|f| = fg \in L^1(\mu)\).
Let \(1 < q < \infty\). Write \(\mathcal{A}_S := \{A \in \mathcal{A}: A \subset S\}\) and let \(\mu_S\) be the restriction of \(\mu\), \(\sigma\)-finite by hypothesis. Extension by zero (legitimate since \(S \in \mathcal{A}\)) maps \(L^q(\mu_S)\) isometrically into \(L^q(\mu)\), so the hypothesis yields
\begin{equation*} (f|_S)\,u \in L^1(\mu_S) \qquad \text{for every } u \in L^q(\mu_S), \end{equation*}
because \(fg\) vanishes off \(S\). Corollary 4.4.5 applied on \((S,\mathcal{A}_S,\mu_S)\) — whose \(\sigma\)-finiteness and \(1<p<\infty\) are just what it requires — gives \(f|_S \in L^p(\mu_S)\), and \(\int_X|f|^p d\mu = \int_S|f|^p d\mu_S < \infty\).
(iii) By (ii) it suffices to prove \(\mu\) \(\sigma\)-finite on \(\{f \ne 0\}\); the case \(q = \infty\) is as in (ii), so let \(1 < q < \infty\).
First, semifiniteness makes finite-measure subsets exhaust infinite measure: if \(\mu(E) = \infty\) then \(s(E) := \sup\{\mu(B): B \subset E,\ \mu(B) < \infty\} = \infty\). Otherwise pick \(B_n \subset E\) of finite measure with \(\mu(B_n) \to s(E)\); the increasing unions \(C_N := \bigcup_{n \le N}B_n\) have \(\mu(C_N) \le s(E)\), so by continuity from below \(B := \bigcup_n B_n\) has \(\mu(B) = s(E) < \infty\); then \(\mu(E\setminus B) = \infty\), and a semifinite \(C \subset E \setminus B\) with \(0 < \mu( C) < \infty\) makes \(\mu(B \cup C) = s(E) + \mu( C) > s(E)\), a contradiction.
Now suppose \(E := \{|f| \ge c\}\) had \(\mu(E) = \infty\) for some \(c > 0\). Choose inductively disjoint \(E_n \subset E\) with \(n^{2p/q} \le \mu(E_n) =: a_n < \infty\), possible because \(E \setminus \bigcup_{k<n}E_k\) still has infinite measure and the previous paragraph applies to it. With \(b_n := (n^2 a_n)^{-1/q}\) and \(g := \sum_n b_n I_{E_n} \ge 0\), disjointness gives
\begin{equation*} \int |g|^q\,d\mu = \sum_{n=1}^{\infty} b_n^q a_n = \sum_{n=1}^{\infty} n^{-2} < \infty , \end{equation*}
so \(g \in L^q(\mu)\), whereas \(|f| \ge c\) on each \(E_n\) forces
\begin{equation*} \int |fg|\,d\mu \ge c\sum_{n=1}^{\infty} b_n a_n = c\sum_{n=1}^{\infty} a_n^{1/p}\,n^{-2/q} \ge c\sum_{n=1}^{\infty} 1 = \infty , \end{equation*}
using \(1 - 1/q = 1/p\) and \(a_n^{1/p} \ge n^{2/q}\). This contradicts \(fg \in L^1(\mu)\). Hence every \(\{|f| \ge 1/n\}\) has finite measure and \(\{f \ne 0\} = \bigcup_n \{|f| \ge 1/n\}\) is \(\sigma\)-finite.
Exercises 4.7.91–4.7.97
(Segal) (i) Let \(\mu\) be a measure with values in \([0,+\infty]\). Prove that \(\mu\) is semifinite precisely when the embedding \(L^\infty(\mu) \to \bigl(L^1(\mu)\bigr)^*\) is injective.
(ii) Let \(\mu\) be a semifinite measure. Prove that \(\mu\) is Maharam (or localizable) in the sense of Exercise 1.12.134 precisely when, for every \(L \in L^1(\mu)^*\), there exists a unique element \(g_L \in L^\infty(\mu)\) with
\begin{equation*} L(f) = \int f g_L \, d\mu \quad \text{for all } f \in L^1(\mu). \end{equation*}
In this case, \(L \mapsto g_L\) is an isometry between \(L^1(\mu)^*\) and \(L^\infty(\mu)\).
Both parts describe the map \(T\colon L^\infty(\mu) \to (L^1(\mu))^*\), \((Tg)(f) := \int_X fg\,d\mu\), which satisfies \(\|Tg\| \le \|g\|_\infty\): (i) asks when \(T\) is injective, (ii) when it is onto. Write \(\mathcal{A}^f := \{A \in \mathcal{A}: \mu(A) < \infty\}\); scalars are real (in the complex case replace \(\operatorname{sign}h\) by \(\overline{h}/|h|\) throughout). Note that \(\{f \ne 0\}\) is \(\sigma\)-finite for \(f \in L^1(\mu)\), since \(\mu(\{|f|>1/n\}) \le n\|f\|_{L^1}\).
(i) Let \(\mu\) be semifinite and \(Tg = 0\). If \(\|g\|_\infty > 0\), some \(E := \{|g| > c\}\) has \(\mu(E) > 0\), and semifiniteness gives \(E_0 \subset E\) with \(0 < \mu(E_0) < \infty\); then \(f := (\operatorname{sign}g)I_{E_0} \in L^1(\mu)\) and
\begin{equation*} 0 = (Tg)(f) = \int_{E_0}|g|\,d\mu \ge c\,\mu(E_0) > 0 , \end{equation*}
so \(g = 0\) and \(T\) is injective. If \(\mu\) is not semifinite, take \(E\) with \(\mu(E) = \infty\) all of whose measurable subsets are null or of infinite measure, and \(g := I_E\), so \(\|g\|_\infty = 1\). For \(f \in L^1(\mu)\) each \(\{x \in E: |f(x)| > 1/n\}\) has finite measure, hence is null, so \(fI_E = 0\) a.e. and \(Tg = 0\): \(T\) is not injective.
(ii) For semifinite \(\mu\), \(T\) is isometric, so by (i) only surjectivity is at issue. Indeed, given \(0 < \varepsilon < \|g\|_\infty\), semifiniteness supplies \(B_0 \subset \{|g| > \|g\|_\infty - \varepsilon\}\) with \(0 < \mu(B_0) < \infty\), and \(f := \mu(B_0)^{-1}(\operatorname{sign}g)I_{B_0}\) has \(\|f\|_{L^1} = 1\) with
\begin{equation*} (Tg)(f) = \frac{1}{\mu(B_0)}\int_{B_0}|g|\,d\mu \ge \|g\|_\infty - \varepsilon . \end{equation*}
Maharam implies onto. Let \(L \in L^1(\mu)^*\) and \(M := \|L\|\). For \(A \in \mathcal{A}^f\), extension by zero embeds \(L^1(\mu|_A)\) isometrically in \(L^1(\mu)\), so Theorem 4.4.1 applied to the finite measure \(\mu|_A\) with \(p=1\) gives \(g_A\) on \(A\) with \(|g_A| \le M\) and
\begin{equation*} L(f) = \int_A f g_A\,d\mu \quad\text{for } f \in L^1(\mu) \text{ vanishing off } A. \tag{\(*\)} \end{equation*}
Applying \((*)\) for \(A\) and for \(B\) to \(f := \operatorname{sign}(g_A-g_B)I_{A\cap B}\) gives \(\int_{A\cap B}|g_A-g_B|\,d\mu = 0\), i.e. \(g_A = g_B\) a.e. on \(A \cap B\) \((**)\).
Glue by localizability: for rational \(c \in [-M,M]\) let \(E_c\) be an essential supremum (Exercise 1.12.134) of \(\mathcal{M}_c := \{\{x \in A: g_A(x) > c\} : A \in \mathcal{A}^f\}\). Then for each \(A \in \mathcal{A}^f\),
\begin{equation*} \mu\bigl((E_c \cap A)\,\triangle\,\{x \in A: g_A(x) > c\}\bigr) = 0 . \end{equation*}
One inclusion is the definition of \(E_c\); for the other, \(E’ := E_c \setminus (A \cap \{g_A \le c\})\) is again an essential upper bound, because for \(B \in \mathcal{A}^f\) and \(N_B := \{x\in B: g_B(x)>c\}\),
\begin{equation*} N_B \setminus E’ \subset (N_B \setminus E_c) \cup (N_B \cap A \cap \{g_A \le c\}), \end{equation*}
the first set null by definition and the second null by \((**)\) (up to a null set \(N_B \cap A = \{x \in A\cap B: g_A(x)>c\}\)); minimality then gives \(\mu(E_c \setminus E’) = 0\), which is the assertion.
Put \(g(x) := \sup\{c \in \mathbb{Q}\cap[-M,M] : x \in E_c\}\), with \(\sup\emptyset := -M\). Then \(\{g>t\}\) is the countable union of the \(E_c\) with \(c > t\), so \(g \in L^\infty(\mu)\) with \(\|g\|_\infty \le M\); and for \(A \in \mathcal{A}^f\), off the null set collecting the countably many symmetric differences above, \(x \in E_c\) iff \(g_A(x) > c\) for every rational \(c\), whence \(g = g_A\) a.e. on \(A\) by density of \(\mathbb{Q}\). Finally, for \(f \in L^1(\mu)\) write the \(\sigma\)-finite set \(\{f \ne 0\}\) as an increasing union of \(A_n \in \mathcal{A}^f\); then \(fI_{A_n} \to f\) in \(L^1(\mu)\) and \(|fI_{A_n}g| \le M|f|\), so dominated convergence and \((*)\) give
\begin{equation*} L(f) = \lim_n L(fI_{A_n}) = \lim_n \int_{A_n} fg\,d\mu = \int_X fg\,d\mu . \end{equation*}
Onto implies Maharam. Let \(\mathcal{M} \subset \mathcal{A}\); we produce an essential supremum. Replacing \(\mathcal{M}\) by \(\{M \cap A: M \in \mathcal{M},\ A \in \mathcal{A}^f\}\) preserves the class of essential upper bounds (if \(\mu(M\setminus E) > 0\), semifiniteness gives \(A \subset M\setminus E\) with \(0 < \mu(A) < \infty\), and \((M\cap A)\setminus E = A\) is non-null), and closing under finite unions preserves it too; so assume \(\mathcal{M}\) is upward directed with all members in \(\mathcal{A}^f\) (for \(\mathcal{M} = \emptyset\) take \(E = \emptyset\)). For \(f \ge 0\) in \(L^1(\mu)\) the net \(M \mapsto \int_M f\,d\mu\) is nondecreasing and bounded by \(\int f\,d\mu\), so with \(f = f^+-f^-\) the limit
\begin{equation*} \Lambda(f) := \lim_{M \in \mathcal{M}} \int_M f\,d\mu \end{equation*}
exists; \(\Lambda\) is linear with \(|\Lambda(f)| \le \|f\|_{L^1}\), so \(\Lambda = Tg\) for some \(g \in L^\infty(\mu)\) by hypothesis. From \(\Lambda(I_A) = \lim_M \mu(M \cap A) \in [0,\mu(A)]\),
\begin{equation*} 0 \le \int_A g\,d\mu \le \mu(A) \qquad \text{for all } A \in \mathcal{A}^f, \tag{\(\dagger\)} \end{equation*}
which forces \(0 \le g \le 1\) a.e.: a non-null set \(\{g < -1/n\}\) or \(\{g>1\}\) contains, by semifiniteness, a set of finite positive measure violating \((\dagger)\). Take a representative with \(0 \le g \le 1\) and put \(E := \{g = 1\}\).
(a) Upper bound: for \(M_0 \in \mathcal{M}\) directedness makes \(\mu(M \cap M_0)\) eventually \(\mu(M_0)\), so \(\int_{M_0}g\,d\mu = \mu(M_0) < \infty\), and \(g \le 1\) gives \(g = 1\) a.e. on \(M_0\), i.e. \(\mu(M_0\setminus E) = 0\).
(b) Minimality: if \(\mu(M \setminus E’)= 0\) for all \(M \in \mathcal{M}\) yet \(\mu(E \setminus E’) > 0\), semifiniteness gives \(A \subset E\setminus E’\) with \(0 < \mu(A) < \infty\); then \(\int_A g\,d\mu = \mu(A) > 0\), while \(\mu(M \cap A) = 0\) for every \(M\) gives \(\int_A g\,d\mu = \Lambda(I_A) = 0\).
Let \(X = \mathbb{R}^2\), \(\mu(A) = +\infty\) if \(A\) is uncountable, \(\mu(A) = \delta_0(A)\) if \(A\) is at most countable, where \(\delta_0\) is Dirac’s measure at the origin. Show that \(\mu\) is a countably additive measure on the \(\sigma\)-algebra of all sets in \(\mathbb{R}^2\) with values in \([0,+\infty]\) that is neither localizable nor semifinite. Verify that \(L^1(\mu) = L^p(\mu) \ne L^\infty(\mu)\) for all \(p \in [1,+\infty)\) and \(\|f\|_{L^p(\mu)} = |f(0)|\) for all \(f \in \mathcal{L}^p(\mu)\).
Writing \(0\) for the origin, the measure is \(\mu(A) = +\infty\) for uncountable \(A\), and \(\mu(A) = I_A(0)\) for at most countable \(A\); its null sets are exactly the at most countable sets missing \(0\).
Countable additivity: for disjoint \(A_n\) with union \(A\), either all \(A_n\) are at most countable, so \(A\) is too and \(\mu(A) = \delta_0(A) = \sum_n\delta_0(A_n)\); or some \(A_{n_0}\) is uncountable, and then both \(\mu(A)\) and \(\sum_n\mu(A_n)\) equal \(+\infty\).
Not semifinite: \(E := \mathbb{R}^2 \setminus \{0\}\) has \(\mu(E) = +\infty\), yet every \(B \subset E\) has \(\mu(B) \in \{0,+\infty\}\). Not localizable either, since a Maharam measure is semifinite by the definition in Exercise 1.12.134.
Every function on \(\mathbb{R}^2\) is measurable, and for \(h \ge 0\) with \(S := \{h>0\}\setminus\{0\}\):
(i) \(S\) uncountable. Then \(\{h > 1/n\}\setminus\{0\}\) is uncountable for some \(n\), and \(\varphi := n^{-1}I_{\{h>1/n\}\setminus\{0\}} \le h\) has \(\int\varphi\,d\mu = +\infty\), so \(\int h\,d\mu = +\infty\).
(ii) \(S\) at most countable. Then \(S\) is null, so from \(h = hI_{\{0\}} + hI_S\) we get \(\int h\,d\mu = h(0)\mu(\{0\}) = h(0)\).
Taking \(h = |f|^p\), for every \(p \in [1,+\infty)\)
\begin{equation*} \mathcal{L}^p(\mu) = \{f : \{x \ne 0: f(x) \ne 0\} \text{ at most countable}\}, \qquad \|f\|_{L^p(\mu)} = |f(0)| , \end{equation*}
a description independent of \(p\), so \(L^1(\mu) = L^p(\mu)\) for all finite \(p\), one-dimensional via \(f \mapsto f(0)\) (two such \(f\) are equivalent exactly when they agree at \(0\)).
But \(I_E \in \mathcal{L}^\infty(\mu)\) with \(\|I_E\|_{L^\infty(\mu)} = 1\), since deleting an at most countable set leaves uncountably many points of \(E\), whereas \(I_E \notin \mathcal{L}^p(\mu)\). Indeed \(I_{\{0\}}, I_E, I_{\{0\}\cup E} = I_{\{0\}}+I_E\) are distinct nonzero classes, so \(\dim L^\infty(\mu) \ge 2 > 1 = \dim L^p(\mu)\).
Let \((X,\mathcal{A},\mu)\) be a space with a complete countably additive measure \(\mu\) with values in \([0,+\infty]\). Denote by \(\mathcal{N}_{loc}(\mu)\) the class of locally zero sets, i.e., sets \(E\) such that \(\mu(E \cap A) = 0\) for all \(A \in \mathcal{A}\) with \(\mu(A) < \infty\). Next, denote by \(L^\infty_{loc}(\mu)\) the class of all \(\mu\)-measurable functions \(f\) with \(\|f\|_{\infty,loc} < \infty\), where we set \(\|f\|_{\infty,loc} = \inf\bigl\{ a : \{x : |f(x)| > a\} \in \mathcal{N}_{loc}(\mu) \bigr\}\) and identify functions that are not equal only on a set from \(\mathcal{N}_{loc}(\mu)\).
(i) Prove that \(L^\infty_{loc}(\mu)\) is a Banach space with the norm \(\|\cdot\|_{\infty,loc}\).
(ii) Prove that for all \(f \in L^\infty_{loc}(\mu)\) one has
\begin{equation*} \|f\|_{\infty,loc} = \sup\Bigl\{ \Bigl| \int_X f g \, d\mu \Bigr| ,\ \|g\|_{L^1(\mu)} = 1 \Bigr\}, \end{equation*}
and the mapping \(L^\infty_{loc}(\mu) \to \bigl(L^1(\mu)\bigr)^*\) is injective and preserves the distances.
(iii) Let \(\mathcal{P}\) be the class of all simple \(\mu\)-integrable functions and let a \(\mu\)-measurable function \(f\) be such that \(fg \in L^1(\mu)\) for all \(g \in \mathcal{P}\) and
\begin{equation*} \sup\Bigl\{ \Bigl| \int_X f g \, d\mu \Bigr| : g \in \mathcal{P},\ \|g\|_{L^1(\mu)} = 1 \Bigr\} < \infty . \end{equation*}
Prove that \(f \in L^\infty_{loc}(\mu)\).
(iv) Let a measure \(\mu\) be decomposable in the sense of Exercise 1.12.131. Prove that every continuous linear functional on \(L^1(\mu)\) is generated by a function from the class \(L^\infty_{loc}(\mu)\), i.e., \(\bigl(L^1(\mu)\bigr)^*\) is naturally isomorphic to \(L^\infty_{loc}(\mu)\).
Write \(\mathcal{A}^f := \{A : \mu(A) < \infty\}\) and \(\mathcal{N} := \mathcal{N}_{loc}(\mu)\) (completeness makes \(\mu\)-measurability the same as \(\mathcal{A}\)-measurability). Three facts are used throughout:
(a) \(\mathcal{N}\) is a \(\sigma\)-ideal containing the null sets, since \(\mu(\bigcup_n E_n \cap A) \le \sum_n \mu(E_n \cap A)\) for \(A \in \mathcal{A}^f\);
(b) \(E \in \mathcal{N}\) exactly when every measurable \(B \subset E\) with \(\mu(B) < \infty\) is null (take \(B = E \cap A\));
(c) the infimum is attained, i.e. \(\{|f| > \|f\|_{\infty,loc}\} = \bigcup_n\{|f| > \|f\|_{\infty,loc}+1/n\} \in \mathcal{N}\) by (a). The norm is well defined on classes, as \(\{|f|>a\}\,\triangle\,\{|h|>a\} \subset \{f \ne h\}\). Also \(\{g \ne 0\}\) is \(\sigma\)-finite for \(g \in L^1(\mu)\), so \(\mu(E \cap \{g \ne 0\}) = 0\) whenever \(E \in \mathcal{N}\).
(i) Homogeneity follows from \(\{|\lambda f| > a\} = \{|f| > a/|\lambda|\}\), the triangle inequality from \(\{|f+h| > a+b\} \subset \{|f|>a\} \cup \{|h|>b\}\) with (a), and \(\|f\|_{\infty,loc} = 0\) means \(\{f \ne 0\} \in \mathcal{N}\) by (c), i.e. \(f\) is the zero class.
For completeness, let \((f_n)\) be Cauchy. By (c) and (a),
\begin{equation*} N := \bigcup_{n,m} \{|f_n - f_m| > \|f_n - f_m\|_{\infty,loc}\} \in \mathcal{N}, \end{equation*}
and off \(N\) the sequence is uniformly Cauchy; put \(f := \lim_n f_n I_{X\setminus N}\), measurable as a pointwise limit. If \(\|f_n - f_m\|_{\infty,loc} \le \varepsilon\) for \(n,m \ge n_0\), letting \(m \to \infty\) gives \(|f_n - f| \le \varepsilon\) off \(N\), so \(\{|f_n-f| > \varepsilon\} \subset N\) and \(\|f_n - f\|_{\infty,loc} \le \varepsilon\); in particular \(f \in L^\infty_{loc}(\mu)\).
(ii) Put \(c := \|f\|_{\infty,loc}\). Since \(\{|f| > c\} \in \mathcal{N}\) by (c), it meets \(\{g \ne 0\}\) in a null set for each \(g \in L^1(\mu)\), so \(|fg| \le c|g|\) a.e. and
\begin{equation*} \Bigl|\int_X fg\,d\mu\Bigr| \le c\,\|g\|_{L^1(\mu)} : \end{equation*}
\(Tf := (g \mapsto \int fg\,d\mu)\) lies in \((L^1(\mu))^*\), is unchanged when \(f\) is altered on a set of \(\mathcal{N}\), and the supremum is at most \(c\). Conversely, for \(0 < \varepsilon < c\) the set \(B := \{|f| > c-\varepsilon\}\) is not in \(\mathcal{N}\), so (b) gives \(B_0 \subset B\) with \(0 < \mu(B_0) < \infty\), and \(g := \mu(B_0)^{-1}(\operatorname{sign}f)I_{B_0}\) has \(\|g\|_{L^1(\mu)} = 1\) with
\begin{equation*} \int_X fg\,d\mu = \frac{1}{\mu(B_0)}\int_{B_0}|f|\,d\mu \ge c - \varepsilon \end{equation*}
(for \(c = 0\) both sides vanish). Hence \(\|Tf\| = \|f\|_{\infty,loc}\): \(T\) is a linear isometry, so injective and distance-preserving.
(iii) \(\|f\|_{\infty,loc} \le M\), the stated supremum. Otherwise some \(c > M\) has \(B := \{|f| > c\} \notin \mathcal{N}\), and (b) gives \(B_0 \subset B\) with \(0 < \mu(B_0) < \infty\); then \(g := \mu(B_0)^{-1}(\operatorname{sign}f)I_{B_0}\) is simple and integrable, i.e. \(g \in \mathcal{P}\) with \(\|g\|_{L^1(\mu)} = 1\) (as \(f \ne 0\) on \(B_0\)), while \(\int fg\,d\mu \ge c > M\).
(iv) Let \(\{X_\alpha\}\) be a decomposition as in Exercise 1.12.131: disjoint sets of finite measure covering \(X\) with (a) \(E \cap X_\alpha \in \mathcal{A}\) for all \(\alpha\) implying \(E \in \mathcal{A}\), and (b) \(\mu(E) = \sum_\alpha \mu(E \cap X_\alpha)\). Given \(L \in (L^1(\mu))^*\), \(M := \|L\|\), Theorem 4.4.1 with \(p=1\) for the finite measure \(\mu|_{X_\alpha}\) (into which \(L\) restricts with norm \(\le M\) via extension by zero) gives \(g_\alpha\) on \(X_\alpha\) with \(|g_\alpha| \le M\) and
\begin{equation*} L(f) = \int_{X_\alpha} f g_\alpha \, d\mu \quad\text{for } f \in L^1(\mu) \text{ vanishing off } X_\alpha . \end{equation*}
Set \(g := g_\alpha\) on \(X_\alpha\); then \(\{g>t\} \cap X_\alpha = \{g_\alpha > t\} \in \mathcal{A}\), so \(\{g>t\} \in \mathcal{A}\) by (a), and \(g \in L^\infty(\mu) \subset L^\infty_{loc}(\mu)\).
Given \(f \in L^1(\mu)\) and \(S := \{f \ne 0\}\), each \(\{|f| > 1/n\}\) has finite measure, so by (b) only countably many \(\alpha\) have \(\mu(S \cap X_\alpha) > 0\), say \(\alpha_1,\alpha_2,\dots\); the remainder \(R := S \setminus \bigcup_k X_{\alpha_k}\) meets every \(X_\alpha\) in a null set, so \(\mu( R) = 0\) by (b). With \(f_n := fI_{X_{\alpha_1}\cup\cdots\cup X_{\alpha_n}}\) we have \(f_n \to f\) a.e. and \(|f_n| \le |f|\), so \(f_n \to f\) in \(L^1(\mu)\); since \(|f_n g| \le M|f|\), dominated convergence twice gives
\begin{equation*} L(f) = \lim_n L(f_n) = \lim_n \int_X f_n g\,d\mu = \int_X f g\,d\mu . \end{equation*}
So the isometry \(T\) of (ii) is onto, hence an isometric isomorphism.
Let \(\mu\) and \(\gamma\) be the measures with values in \([0,+\infty]\) defined in Exercise 1.12.137 and let
\begin{equation*} l(f) = \int f \, d\gamma, \qquad f \in L^1(\mu). \end{equation*}
Prove that \(l\) is a continuous linear functional on \(L^1(\mu)\), but there is no function \(g \in L^\infty_{loc}(\mu)\) such that
\begin{equation*} l(f) = \int f g \, d\mu \quad \text{for all } f \in L^1(\mu). \end{equation*}
The functional is continuous, but no \(\mu\)-measurable \(g\) whatever represents it: a cardinality clash between the horizontal and the vertical lines rules it out.
In Exercise 1.12.137, \(\operatorname{card}X = \mathfrak{c} < \operatorname{card}Y\), \(\mathcal{A}\) consists of the \(A \subset X\times Y\) all of whose horizontal sections \(A_b\) and vertical sections \(A^a\) are at most countable or co-countable in their lines, \(\gamma(A)\) counts the \(b\) with co-countable \(A_b\), \(v(A)\) counts the \(a\) with co-countable \(A^a\), and \(\mu = \gamma + v\). Here \(\mu(N) = 0\) forces every section of \(N\) to be at most countable, so all subsets of \(N\) lie in \(\mathcal{A}\) and are null: \(\mu\) is complete.
Continuity. From \(\mu \ge \gamma\) every \(\mu\)-null set is \(\gamma\)-null (so \(l\) depends only on the \(\mu\)-class) and \(\int h\,d\gamma \le \int h\,d\mu\) for measurable \(h \ge 0\) (clear for simple \(h\), then take suprema), whence \(|l(f)| \le \int|f|\,d\gamma \le \|f\|_{L^1(\mu)}\) and \(\|l\| \le 1\).
Put \(H_b := X \times \{b\}\) and \(V_a := \{a\}\times Y\); both lie in \(\mathcal{A}\), and inspecting the sections of their complements gives \(\gamma(H_b) = 1 = v(V_a)\), \(v(H_b) = 0 = \gamma(V_a)\) (Check!), so
\begin{equation*} \mu(H_b) = \mu(V_a) = 1, \qquad l(I_{H_b}) = 1, \qquad l(I_{V_a}) = 0. \tag{\(*\)} \end{equation*}
For \(S \subset X\), the same inspection shows \(S \times \{b\} \in \mathcal{A}\) exactly when \(S\) is at most countable or co-countable, with \(\mu(S\times\{b\}) = 1\) in the second case and \(0\) in the first. So the trace of \(\mathcal{A}\) on \(H_b\) is the countable-cocountable \(\sigma\)-algebra of \(X\), and any function \(h\) measurable for it is constant off an at most countable set: \(T := \{t : \{h>t\}\text{ co-countable}\}\) is a half-line, nonempty (else \(X = \bigcup_n\{h>-n\}\) would be at most countable) and bounded above (else \(\emptyset = \bigcap_n\{h>n\}\) would be co-countable), and for \(c := \sup T\) each \(\{h \le c - 1/n\}\) and each \(\{h \ge c+1/n\}\) is at most countable, hence so is \(\{h \ne c\}\). Symmetrically on each \(V_a\), with \(Y\) in place of \(X\).
Suppose now \(l(f) = \int fg\,d\mu\) for all \(f \in L^1(\mu)\), with \(g\) merely \(\mu\)-measurable. On \(H_b\), \(g\) equals a constant \(c_b\) off an at most countable \(N_b \subset X\), and \(N_b \times \{b\}\) is \(\mu\)-null, so \((*)\) gives \(1 = c_b\,\mu(H_b) = c_b\):
\begin{equation*} g(x,b) = 1 \quad\text{for all } x \in X \setminus N_b. \tag{1} \end{equation*}
Symmetrically \(g = d_a\) off an at most countable \(M_a \subset Y\) on \(V_a\), with \(\{a\}\times M_a\) \(\mu\)-null, and \(0 = l(I_{V_a}) = d_a\):
\begin{equation*} g(a,y) = 0 \quad\text{for all } y \in Y \setminus M_a. \tag{2} \end{equation*}
Since \(\operatorname{card}\bigcup_{a \in X}M_a \le \mathfrak{c}\cdot\aleph_0 = \mathfrak{c} < \operatorname{card}Y\), some \(b\) lies in no \(M_a\); then (2) gives \(g(\cdot,b) \equiv 0\) on \(X\), so (1) forces \(N_b = X\), impossible for an at most countable set.
(\(\circ\)) Let \(f_n, f \in L^\infty[a,b]\). Prove that the following conditions are equivalent:
(i) one has
\begin{equation*} \int_a^b f_n(x) g(x) \, dx \to \int_a^b f(x) g(x) \, dx, \qquad \forall\, g \in L^1[a,b]; \end{equation*}
(ii) one has \(\sup_n \|f_n\|_{L^\infty} < \infty\) and
\begin{equation*} \int_a^z f_n(x) \, dx \to \int_a^z f(x) \, dx, \qquad \forall\, z \in [a,b]. \end{equation*}
This is uniform boundedness plus density of step functions. Since Lebesgue measure on \([a,b]\) is finite, Theorem 4.4.1 (\(p=1\), \(q=\infty\)) makes \(\Lambda_h(g) := \int_a^b hg\,dx\) a continuous functional on \(L^1[a,b]\) with \(\|\Lambda_h\| = \|h\|_{L^\infty}\); write \(\Lambda_n := \Lambda_{f_n}\), \(\Lambda := \Lambda_f\), so (i) says \(\Lambda_n \to \Lambda\) pointwise on \(L^1[a,b]\).
(i) implies (ii). Each sequence \((\Lambda_n(g))_n\) converges, hence is bounded, so the Banach-Steinhaus theorem 4.4.3(i) on the Banach space \(L^1[a,b]\) gives \(\sup_n\|f_n\|_{L^\infty} = \sup_n\|\Lambda_n\| < \infty\). The displayed convergence is (i) for \(g = I_{[a,z]} \in L^1[a,b]\).
(ii) implies (i). Put \(C := \max\{\sup_n\|f_n\|_{L^\infty}, \|f\|_{L^\infty}\} < \infty\), so \(\|\Lambda_n\|, \|\Lambda\| \le C\). Since \(I_{[u,v]} = I_{[a,v]} - I_{[a,u]}\) a.e., linearity upgrades the hypothesis to
\begin{equation*} \Lambda_n(h) \to \Lambda(h) \qquad \text{for every step function } h , \tag{\(*\)} \end{equation*}
and step functions are dense in \(L^1[a,b]\) (Lemma 4.2.1: simple functions are dense, and indicators of measurable sets are approximated in \(L^1\) by indicators of finite unions of intervals). Given \(g \in L^1[a,b]\) and \(\varepsilon > 0\), pick a step function \(h\) with \(\|g-h\|_{L^1} < \varepsilon\); then
\begin{equation*} |\Lambda_n(g) - \Lambda(g)| \le |\Lambda_n(g-h)| + |\Lambda_n(h)-\Lambda(h)| + |\Lambda(h-g)| \le 2C\varepsilon + |\Lambda_n(h)-\Lambda(h)| , \end{equation*}
so \(\limsup_n|\Lambda_n(g)-\Lambda(g)| \le 2C\varepsilon\) by \((*)\), and \(\varepsilon\) was arbitrary.
(\(\circ\)) Let \(f\) be a measurable function on the real line with a period 1.
(i) Prove that if \(f \in L^1[0,1]\), then
\begin{equation*} \lim_{n \to \infty} \int_0^1 g(x) f(nx) \, dx = \int_0^1 g(x) \, dx \int_0^1 f(x) \, dx \tag{4.7.19} \end{equation*}
for all \(g \in C[0,1]\) (where \(n \in \mathbb{N}\)).
(ii) Prove that if \(f\) is bounded, then the above relation is true for all \(g \in L^1[0,1]\).
The engine is the bound \((*)\) below; both parts then follow by approximating \(g\) with step functions.
Subtracting the constant \(c := \int_0^1 f\,dx\) (legitimate, as \(\int_0^1 gc\,dx = c\int_0^1 g\), and periodicity and boundedness are preserved), assume \(\int_0^1 f\,dx = 0\) and prove \(\int_0^1 g(x)f(nx)\,dx \to 0\). Periodicity and \(y = nx\) also give \(\|f(n\,\cdot)\|_{L^1[0,1]} = \|f\|_{L^1[0,1]}\).
The primitive \(\Phi(z) := \int_0^z f(y)\,dy\) is continuous and \(1\)-periodic, since periodicity of \(f\) gives \(\Phi(z+1)-\Phi(z) = \int_z^{z+1}f = \int_0^1 f = 0\); hence \(K := \sup_{z \ge 0}|\Phi(z)| \le \|f\|_{L^1[0,1]} < \infty\), and for \(0 \le u \le v \le 1\) the substitution \(y = nx\) yields
\begin{equation*} \int_u^v f(nx)\,dx = \frac{\Phi(nv)-\Phi(nu)}{n}, \qquad \Bigl|\int_u^v f(nx)\,dx\Bigr| \le \frac{2K}{n} . \tag{\(*\)} \end{equation*}
Summing \((*)\) over the intervals of a step function \(h = \sum_{j\le m}c_jI_{J_j}\),
\begin{equation*} \Bigl|\int_0^1 h(x)f(nx)\,dx\Bigr| \le \frac{2K}{n}\sum_{j=1}^{m}|c_j| \longrightarrow 0 . \tag{\(**\)} \end{equation*}
(i) Given \(g \in C[0,1]\) and \(\varepsilon > 0\), uniform continuity provides a step function \(h\) with \(\sup_{[0,1]}|g-h| \le \varepsilon\). Then
\begin{equation*} \Bigl|\int_0^1 (g-h)(x)f(nx)\,dx\Bigr| \le \varepsilon\int_0^1|f(nx)|\,dx = \varepsilon\|f\|_{L^1[0,1]} , \end{equation*}
so \((**)\) gives \(\limsup_n|\int_0^1 g(x)f(nx)\,dx| \le \varepsilon\|f\|_{L^1[0,1]}\) for every \(\varepsilon > 0\).
(ii) Now \(|f| \le M\). Step functions are dense in \(L^1[0,1]\) (Lemma 4.2.1 and the remark following it), so for \(g \in L^1[0,1]\) choose \(h\) with \(\|g-h\|_{L^1[0,1]} \le \varepsilon\); then
\begin{equation*} \Bigl|\int_0^1 (g-h)(x)f(nx)\,dx\Bigr| \le M\|g-h\|_{L^1[0,1]} \le M\varepsilon , \end{equation*}
and \((**)\) gives \(\limsup_n|\int_0^1 g(x)f(nx)\,dx| \le M\varepsilon\).
(\(\circ\)) Let \(f\) be a bounded measurable function on the real line with a period 1. Show that if a sequence of functions \(f(nx)\) has a subsequence convergent on a set of positive measure, then \(f\) a.e. equals some constant.
The constant is \(c := \int_0^1 f(x)\,dx\), and the proof runs Exercise 4.7.96(ii) twice: once on \(f\) to identify the limit, once on \((f-c)^2\) to kill the variance.
Say \(|f| \le M\) and \(f(n_kx) \to \varphi(x)\) on a set of positive measure. Choosing \(m\) with \(\lambda(E_1 \cap [m,m+1)) > 0\) and translating, we get \(E \subset [0,1]\) with \(\lambda(E) > 0\) and \(f(n_kx) \to \psi(x) := \varphi(x+m)\) on \(E\), since \(f(n(x+m)) = f(nx)\) by \(1\)-periodicity and \(nm \in \mathbb{Z}\); \(\psi\) is measurable with \(|\psi| \le M\).
For measurable \(A \subset E\), Exercise 4.7.96(ii) with \(g = I_A\) (legitimate: \(f\) is bounded, measurable, \(1\)-periodic) and dominated convergence (\(|f(n_kx)| \le M\), \(\lambda(A) < \infty\)) give
\begin{equation*} \int_A \psi \,dx = \lim_k \int_A f(n_kx)\,dx = c\,\lambda(A) . \end{equation*}
Taking \(A = \{\psi > c\}\) and \(A = \{\psi < c\}\) makes both null, so \(f(n_kx) \to c\) a.e. on \(E\).
Apply 4.7.96(ii) to the bounded \(1\)-periodic \(h := (f-c)^2\) with \(g = I_E\), and compare with dominated convergence along \(n_k\) (where \(0 \le h(n_kx) \to 0\) a.e. on \(E\)):
\begin{equation*} \lambda(E)\int_0^1 (f(x)-c)^2\,dx = \lim_k \int_E h(n_kx)\,dx = 0 . \end{equation*}
As \(\lambda(E) > 0\), \(f = c\) a.e. on \([0,1]\), hence a.e. on \(\mathbb{R}\) by periodicity.
Exercises 4.7.98–4.7.104
Prove that the functions \(|\sin \pi n x|\) converge weakly in \(L^2[0,1]\) and find their limit.
They converge weakly to the constant \(2/\pi\), the mean of \(\varphi(x) := |\sin\pi x|\):
\begin{equation*} \int_0^1 \varphi(x)\,dx = \int_0^1 \sin\pi x\,dx = \frac{2}{\pi} \end{equation*}
(as \(\sin\pi x \ge 0\) on \([0,1]\)), and \(\varphi\) is continuous, \(1\)-periodic with \(0 \le \varphi \le 1\).
Put \(h := \varphi - 2/\pi\), so \(h\) is \(1\)-periodic and bounded with \(\int_0^1 h = 0\); as in Exercise 4.7.96 its primitive \(H(x) := \int_0^x h\) is continuous and \(1\)-periodic, \(M := \sup_{\mathbb{R}}|H| < \infty\), and for \(0 \le a < b \le 1\)
\begin{equation*} \Bigl|\int_a^b h(nx)\,dx\Bigr| = \frac{|H(nb)-H(na)|}{n} \le \frac{2M}{n} , \end{equation*}
whence \(\int_0^1 s(x)h(nx)\,dx \to 0\) for every step function \(s\). Step functions are dense in \(L^2[0,1]\) (Lemma 4.2.1 with regularity of Lebesgue measure), and \(\|h(n\,\cdot)\|_{L^2} \le \|h\|_\infty \le 1\), so for \(g \in L^2[0,1]\) and \(s\) with \(\|g-s\|_{L^2} < \varepsilon\) the Cauchy–Bunyakovskii inequality gives
\begin{equation*} \Bigl|\int_0^1 g(x)h(nx)\,dx\Bigr| \le \|g-s\|_{L^2}\|h(n\,\cdot)\|_{L^2} + \Bigl|\int_0^1 s(x)h(nx)\,dx\Bigr| \le \varepsilon + o(1). \end{equation*}
Since \(|\sin\pi nx| = h(nx) + 2/\pi\), this is exactly \(\int_0^1 g|\sin\pi n\,\cdot| \to (2/\pi)\int_0^1 g\) for all \(g \in L^2[0,1]\).
Suppose a sequence of functions \(f_n\) converges weakly in \(L^1[0,1]\) to a function \(f\). Is it true that the functions \(|f_n|\) converge weakly to \(|f|\)?
No: take \(f_n(x) = \sin\pi nx\), which converges weakly to \(f = 0\), while \(\int_0^1|f_n|\,dx = 2/\pi\) for every \(n\) by periodicity of \(|\sin\pi y|\), so the test function \(g \equiv 1 \in L^\infty[0,1]\) already defeats \(|f_n| \to |f| = 0\).
For the weak convergence, Theorem 4.4.1 with \(p=1\) makes it convergence of \(\int_0^1 gf_n\,dx\) for every \(g \in L^\infty[0,1] \subset L^2[0,1]\). The functions \(e_n(x) := \sqrt{2}\sin(\pi nx)\) are orthonormal in \(L^2[0,1]\), since
\begin{equation*} \int_0^1 2\sin(\pi nx)\sin(\pi mx)\,dx = \int_0^1 \bigl[\cos(\pi(n-m)x) - \cos(\pi(n+m)x)\bigr]\,dx \end{equation*}
and \(\int_0^1\cos(\pi kx)\,dx = 0\) for every integer \(k \ne 0\), while the first integrand is \(1\) for \(n = m\). Bessel’s inequality then gives \(\sum_n|\int_0^1 ge_n\,dx|^2 \le \|g\|_{L^2}^2 < \infty\), so \(\int_0^1 gf_n\,dx \to 0\).
Prove that for every irrational number \(\alpha\), there exist infinitely many rational numbers \(p/q\), where \(p, q\) are integers, such that \(|\alpha - p/q| < q^{-2}\).
Dirichlet’s pigeonhole gives, for each \(n\), integers \(p,q\) with \(1 \le q \le n\) and
\begin{equation*} \Bigl|\alpha - \frac{p}{q}\Bigr| < \frac{1}{nq} \le \frac{1}{q^{2}} , \tag{1} \end{equation*}
and that suffices. Indeed, the \(n+1\) fractional parts \(\{k\alpha\}\), \(k = 0,\dots,n\), lie in the \(n\) intervals \([j/n,(j+1)/n)\), so two of them, say for \(n_1 < n_2 \le n\), share an interval; with \(q := n_2-n_1 \in \{1,\dots,n\}\) and \(p := [n_2\alpha]-[n_1\alpha]\),
\begin{equation*} |q\alpha - p| = \bigl|\{n_2\alpha\} - \{n_1\alpha\}\bigr| < \frac{1}{n} , \end{equation*}
and dividing by \(q \le n\) gives (1) (with \(\alpha \ne p/q\), \(\alpha\) being irrational).
If only finitely many rationals \(r_1,\dots,r_m\) admitted such a representation, then \(\varepsilon := \min_i|\alpha - r_i| > 0\) by irrationality; choosing \(n > 1/\varepsilon\), the pair from (1) has \(|\alpha - p/q| < 1/n < \varepsilon\) and \(|\alpha - p/q| < q^{-2}\), so \(p/q\) is one of the \(r_i\) yet lies within \(\varepsilon\) of \(\alpha\) — a contradiction.
Let \(f\) be a measurable function on \([0,1)\), extended periodically to the whole real line and having the integral \(I(f)\) over \([0,1]\). For every \(n \in \mathbb{N}\), we consider the Riemannian sum
\begin{equation*} S_n f(x) := n^{-1} \sum_{k=0}^{n} f(x + k/n), \qquad x \in [0,1). \end{equation*}
(i) Prove that \(\|S_n f\|_{L^p[0,1)} \le \|f\|_{L^p[0,1)}\) for all \(f \in L^p[0,1)\), \(p \in [1,\infty)\), and that \(\|I(f) - S_n f\|_{L^p[0,1)} \to 0\) as \(n \to \infty\).
(ii) Show that, for every function \(f \in L^1[0,1)\), there exists a sequence \(n_m \to \infty\) such that \(S_{n_m} f(x) \to I(f)\) for almost all \(x \in [0,1)\) (in fact, one can take \(n_m = 2^m\), see Example 10.3.18 in Chapter 10).
(iii) Give an example of an integrable function \(f\) with a period \(1\) such that \(S_n f(x) \to I(f)\) only on a measure zero set. Verify that if \(f(x) = x^{-r}\) for \(x \in (0,1)\), where \(r \in (1/2,1)\), then one has the equality \(\limsup_{n\to\infty} S_n f(x) = +\infty\) almost everywhere.
(iv) Show that in (iii) one can take for \(f\) the indicator of an open set.
The printed sum’s \(k=n\) term repeats its \(k=0\) term by periodicity (for \(f \equiv 1\) it gives \((n+1)/n\)), so we read \(S_nf(x) := n^{-1}\sum_{k=0}^{n-1}f(x+k/n)\), which differs from it by \(n^{-1}f(x) \to 0\) both in \(L^p[0,1)\) and a.e.
Throughout, a \(1\)-periodic integrable \(g\) satisfies \(\int_0^1 g(x+h)\,dx = \int_0^1 g\,dx\) for every \(h\) (split \(h\) into integer and fractional parts; Check!).
(i) Applied to \(|f|^p\), that identity makes each translation \(T_hf := f(\cdot+h)\) an isometry of \(L^p[0,1)\), so Minkowski’s inequality gives
\begin{equation*} \|S_nf\|_{L^p[0,1)} \le \frac{1}{n}\sum_{k=0}^{n-1}\|T_{k/n}f\|_{L^p[0,1)} = \|f\|_{L^p[0,1)} . \end{equation*}
With \(Pf := I(f)\mathbf{1}\), Hölder on the probability space \([0,1)\) gives \(\|Pf\|_{L^p} = |I(f)| \le \|f\|_{L^p}\), so the operators \(A_n := S_n - P\) satisfy \(\|A_n\| \le 2\) and it suffices to show \(A_ng \to 0\) for \(g\) in a dense set. Take the continuous \(1\)-periodic functions, dense since step functions are (Lemma 4.2.1 with regularity of Lebesgue measure) and each step function is approximated in \(L^p\) by a continuous function vanishing near \(0\) and \(1\). For such \(g\), with modulus of continuity \(\omega\) and \(I(g) = \int_0^1 g(x+t)\,dt\),
\begin{equation*} |S_ng(x) - I(g)| = \Bigl|\sum_{k=0}^{n-1}\int_{k/n}^{(k+1)/n}\bigl[g(x+k/n)-g(x+t)\bigr]dt\Bigr| \le \omega(1/n) , \end{equation*}
so \(\|S_ng - I(g)\|_{L^p[0,1)} \le \omega(1/n) \to 0\).
(ii) By (i) with \(p=1\), \(S_nf \to I(f)\) in \(L^1[0,1)\), hence in measure (Chebyshev), so Theorem 2.2.5 supplies \(n_m \to \infty\) with \(S_{n_m}f \to I(f)\) a.e.
(iii) Take \(f(y) := \{y\}^{-r}\) with \(r \in (1/2,1)\), so \(f \ge 0\) is \(1\)-periodic and integrable with \(I(f) = 1/(1-r)\); then \(\limsup_n S_nf(x) = +\infty\) at every irrational \(x\), so \(S_nf \to I(f)\) only on a subset of \(\mathbb{Q}\).
Fix \(\alpha\) irrational and \(q \in \mathbb{N}\), and put \(\theta := \alpha - [q\alpha]/q\), so \(q\theta = \{q\alpha\} \in (0,1)\). Writing \([q\alpha]+k = qm_k + j_k\) with \(j_k \in \{0,\dots,q-1\}\),
\begin{equation*} \alpha + \frac{k}{q} = m_k + \frac{j_k + q\theta}{q}, \qquad \Bigl\{\alpha + \frac{k}{q}\Bigr\} = \theta + \frac{j_k}{q} , \end{equation*}
and \(j_k\) runs over all residues as \(k\) does, so \(\min_{0 \le k < q}\{\alpha + k/q\} = \{q\alpha\}/q\). All summands being nonnegative,
\begin{equation*} S_qf(\alpha) \ge \frac{1}{q}\Bigl(\frac{\{q\alpha\}}{q}\Bigr)^{-r} = q^{\,r-1}\{q\alpha\}^{-r} . \tag{2} \end{equation*}
So infinitely many \(q\) with \(\{q\alpha\} < 1/q\) give \(S_qf(\alpha) \ge q^{2r-1} \to \infty\), as \(2r-1 > 0\).
Exercise 4.7.100 supplies infinitely many \(q\) with \(\|q\alpha\| < 1/q\) (\(\|t\|\) = distance to the nearest integer), but (2) needs closeness from above; the one-sided form also holds. Writing \(\|j\alpha\| > 0\) and pairwise distinct (irrationality), call \(Q\) a record if \(\|Q\alpha\| < \|j\alpha\|\) for \(1 \le j < Q\), so that
\begin{equation*} \min_{1 \le j \le n}\|j\alpha\| = \|Q(n)\alpha\|, \qquad Q(n) := \max\{\text{records } Q \le n\} , \tag{3} \end{equation*}
and records \(Q_1 < Q_2 < \cdots\) are infinite in number, since \(\min_{j \le n}\|j\alpha\| < 1/n \to 0\) by Exercise 4.7.100. Let \(p_k\) be nearest to \(Q_k\alpha\) and \(\delta_k := Q_k\alpha - p_k\), so \(|\delta_{k+1}| < |\delta_k| < 1/2\).
(a) \(\|Q_k\alpha\| < 1/Q_k\): with \(n := Q_{k+1}-1 \ge Q_k\) one has \(Q(n) = Q_k\), so (3) and Exercise 4.7.100 give \(\|Q_k\alpha\| < 1/n \le 1/Q_k\).
(b) The \(\delta_k\) alternate in sign: were \(\delta_k,\delta_{k+1}\) of one sign, then \(q := Q_{k+1}-Q_k \le Q_{k+1}-1\) and \(p := p_{k+1}-p_k\) would give \(q\alpha - p = \delta_{k+1}-\delta_k\) with \(0 < |\delta_{k+1}-\delta_k| < |\delta_k| < 1/2\), so \(\|q\alpha\| < \|Q_k\alpha\|\), contradicting (3).
By (b), \(\delta_k > 0\) for infinitely many \(k\), and then \(p_k = [Q_k\alpha]\), so \(\{Q_k\alpha\} = \|Q_k\alpha\| < 1/Q_k\) by (a). With (2) this gives \(\limsup_n S_nf(x) = +\infty\) at every irrational \(x\).
(iv) Yes: what is needed is an open \(1\)-periodic \(U\) with \(m := |U \cap [0,1)| < 1\) and
\begin{equation*} \limsup_{n\to\infty} S_n\mathbf{1}_U(x) > m \qquad \text{for a.e. } x , \tag{4} \end{equation*}
and such a set is furnished by the theorem of Besicovitch and Rudin cited in the book’s hint (A. S. Besicovitch [84]; W. Rudin [833]). Three proved constraints delimit any construction.
(a) By (i) with \(p=1\), \(S_n\mathbf{1}_U \to m\) in measure, so \(|D_n( c)| \to 0\) for \(D_n( c) := \{|S_n\mathbf{1}_U - m| \ge c\}\); a.e. divergence means \(|\limsup_n D_n( c)| > 0\) for some rational \(c > 0\), which by Borel–Cantelli forces \(\sum_n|D_n( c)| = \infty\).
(b) The extreme value is unattainable: \(S_n\mathbf{1}_U(x) = 1\) forces \(x \in U\) (the term \(k=0\)), so the set of \(x\) attaining it for infinitely many \(n\) has measure \(m < 1\); the excess in (4) must come from progressions only partly inside \(U\).
(c) The set \(U_1 := \bigcup_{q \ge 1}\bigcup_{a \in \mathbb{Z}}(a/q,\,a/q+q^{-2})\) suggested by Exercise 4.7.100 is useless: by (iii) the condition \(0 < \{qx\} < 1/q\) both puts the whole progression \(\{x+k/q\}_{k<q}\) in \(U_1\) and puts \(x\) in \((a/q, a/q+q^{-2})\) with \(a = [qx]\), and it holds for infinitely many \(q\) at every irrational \(x\); so \(|U_1 \cap [0,1)| = 1\). Shrinking the intervals to length \(\varepsilon q^{-2}\) keeps measure \(1\) by Khintchine’s theorem (\(\sum_q \varepsilon/q = \infty\)), while thinning the scales so that \(\sum_i|G_i| < \infty\) makes Borel–Cantelli kill (4) off a null set.
(A fully rigorous proof of the remaining step is beyond the scope of this page; see the reference given in the book.)
Let \(\mu\) be a probability measure and let \(f \in L^1(\mu)\). Prove that \(f\) belongs to \(L^p(\mu)\) with some \(p \in (1,\infty)\) precisely when there exists \(C > 0\) such that
\begin{equation*} \sum_{k=1}^{n} \mu(A_k)^{1-p} \Big| \int_{A_k} f\,d\mu \Big|^p \le C \end{equation*}
for every finite partition of the space into disjoint measurable sets \(A_k\) of positive measure. In addition, the smallest possible constant \(C\) equals \(\|f\|_p^p\).
Necessity is one application of Hölder, sufficiency a partition of \(X\) by the level sets of \(|f|\); the two together pin the optimal constant at \(\|f\|_p^p\). Fix \(p \in (1,\infty)\), \(q := p/(p-1)\), so \(p/q = p-1\).
If \(f \in L^p(\mu)\), Hölder on \(A_k\) applied to \(f\) and \(1\) gives
\begin{equation*} \Bigl|\int_{A_k}f\,d\mu\Bigr|^p \le \Bigl(\int_{A_k}|f|^p\,d\mu\Bigr)\mu(A_k)^{p-1} , \end{equation*}
so summing over a partition, \(\sum_k \mu(A_k)^{1-p}|\int_{A_k}f\,d\mu|^p \le \|f\|_p^p\): the constant \(C = \|f\|_p^p\) is admissible.
Conversely, let \(f \in L^1(\mu)\) satisfy the bound with constant \(C\). Null members of a partition may be absorbed into another set without changing any term, so the hypothesis holds for every finite disjoint family of sets of positive measure covering \(X\) up to a null set. Put \(g := |f|\) and \(X^{\pm} := \{f \ge 0\}, \{f<0\}\); for \(A\) inside \(X^+\) or \(X^-\) we have \(|\int_A f\,d\mu| = \int_A g\,d\mu\), so for all partitions of that type
\begin{equation*} \sum_{k}\mu(A_k)^{1-p}\Bigl(\int_{A_k}g\,d\mu\Bigr)^{p} \le C . \tag{1} \end{equation*}
Fix \(N\), put \(g_N := \min(g,N)\), fix \(\varepsilon > 0\), and let \(\mathcal{P}\) consist of the sets \(E_k \cap X^{\pm}\) of positive measure, where \(E_k := \{c_k \le g_N < c_k + \varepsilon\}\), \(c_k := k\varepsilon\), \(k < K\), \(K\varepsilon > N\). On \(A \in \mathcal{P}\) with \(A \subset E_{k}\) we have \(g \ge c_{k}\), so \(\mu(A)^{1-p}(\int_A g\,d\mu)^p \ge c_{k}^p\mu(A)\), and (1) gives \(\sum_{A}c_{k(A)}^p\mu(A) \le C\). Since \(g_N < c_{k(A)}+\varepsilon\) on \(A\), the mean value theorem (\(0 \le c_{k(A)} \le N\)) and \(\sum_A \mu(A) = 1\) yield
\begin{equation*} \|g_N\|_p^p = \sum_{A \in \mathcal{P}}\int_A g_N^p\,d\mu \le \sum_{A \in \mathcal{P}} c_{k(A)}^p\mu(A) + p(N+\varepsilon)^{p-1}\varepsilon \le C + p(N+\varepsilon)^{p-1}\varepsilon . \end{equation*}
Letting \(\varepsilon \to 0\) gives \(\|g_N\|_p^p \le C\) for each \(N\), and \(g_N^p \uparrow |f|^p\) with monotone convergence gives \(\|f\|_p^p \le C\).
Let \(f \in L^1[0,1]\) and
\begin{equation*} F(x) = \int_0^x f(t)\,dt . \end{equation*}
Prove that \(f \in L^p[0,1]\) with some \(p \in (1,+\infty)\) precisely when there exists \(C > 0\) such that
\begin{equation*} \sum_{k=1}^{n} \frac{|F(x_k) - F(x_{k-1})|^p}{(x_k - x_{k-1})^{p-1}} \le C \end{equation*}
for every finite partition \(0 = x_0 < x_1 < \cdots < x_n = 1\), and the smallest possible \(C\) coincides with \(\|f\|_p^p\).
The sums are those of Exercise 4.7.102 for \(\lambda\) on \([0,1]\) and the partition \(A_k := [x_{k-1},x_k)\), with \(\lambda(A_k) = \Delta x_k\) and \(\int_{A_k}f\,d\lambda = \Delta F_k\); so its necessity half already gives \(\sum_k|\Delta F_k|^p(\Delta x_k)^{1-p} \le \|f\|_p^p\). Fix \(q := p/(p-1)\). Here only interval partitions are allowed, so sufficiency is reproved.
Let \(f \in \mathcal{L}^1[0,1]\) satisfy the bound with constant \(C\). For a step function \(g = \sum_k c_k\mathbf{1}_{[x_{k-1},x_k)}\), Hölder for finite sums with exponents \(q,p\) applied to the factorization \(c_k\Delta F_k = (c_k(\Delta x_k)^{1/q})(\Delta F_k(\Delta x_k)^{-1/q})\) gives
\begin{equation*} \Bigl|\int_0^1 fg\,dx\Bigr| \le \|g\|_q\Bigl(\sum_{k}\frac{|\Delta F_k|^p}{(\Delta x_k)^{p-1}}\Bigr)^{1/p} \le C^{1/p}\|g\|_q . \tag{1} \end{equation*}
This extends to all \(g \in L^\infty[0,1]\): let \(g_j\) be the average of \(g\) over the \(2^j\) dyadic intervals of rank \(j\), so \(\|g_j\|_q \le \|g\|_q\) by Jensen and \(g_j \to g\) in \(L^1[0,1]\) (true for continuous \(g\) by uniform continuity, the averagings are \(L^1\)-contractions, and \(C[0,1]\) is dense by Corollary 4.2.2); pass to a subsequence converging a.e. (Theorem 2.2.5) and use dominated convergence with \(|fg_{j_m}| \le \|g\|_\infty|f|\).
Now put \(f_N := \min(|f|,N)\) and \(g := (\operatorname{sign}f)f_N^{p-1} \in L^\infty[0,1]\), so \(fg \ge f_N^p\) and \(|g|^q = f_N^p\). With \(J_N := \int_0^1 f_N^p\,dx \le N^p\), (1) gives
\begin{equation*} J_N \le \int_0^1 fg\,dx \le C^{1/p}\|g\|_q = C^{1/p}J_N^{1/q} , \end{equation*}
so \(J_N \le C\) (divide by \(J_N^{1/q}\) when \(J_N > 0\), using \(1-1/q = 1/p\)), and \(f_N^p \uparrow |f|^p\) with monotone convergence gives \(\|f\|_p^p \le C\).
(\(\circ\)) Let \(f \in L^1(\mathbb{R}^1)\).
(i) Show that if \(\varepsilon_n \to 0\), then
\begin{equation*} \lim_{n\to\infty} \int_{-\infty}^{+\infty} |f(x+\varepsilon_n) - f(x)|\,dx = 0 . \end{equation*}
(ii) Show that
\begin{equation*} \lim_{|t|\to\infty} \int_{-\infty}^{+\infty} |f(x+t)-f(x)|\,dx = 2\int_{-\infty}^{+\infty}|f(x)|\,dx . \end{equation*}
(iii) Let \(f_n \to f\) in \(L^1(\mathbb{R}^1)\) and \(a_n \to a\) in \(\mathbb{R}^1\). Show that
\begin{equation*} \lim_{n\to\infty}\int_{-\infty}^{+\infty} |f_n(x+a_n) - f(x+a)|\,dx = 0, \qquad \lim_{n\to\infty}\int_{-\infty}^{+\infty} |f_n(x+a_n)|\,dx = \int_{-\infty}^{+\infty}|f(x+a)|\,dx . \end{equation*}
All three parts run on translation invariance \(\|u_h\|_1 = \|u\|_1\), where \(u_h(x) := u(x+h)\), plus density of \(C_0^\infty(\mathbb{R})\) in \(L^1(\mathbb{R})\) (Corollary 4.2.2).
(i) Continuity of translation in \(L^1\) (Lemma 4.2.3 with \(p=1\)): given \(\varepsilon > 0\), take \(\varphi \in C_0^\infty(\mathbb{R})\), vanishing off \([-R,R]\), with \(\|f-\varphi\|_1 < \varepsilon\); then \(\|f_h - \varphi_h\|_1 = \|f-\varphi\|_1\), so
\begin{equation*} \|f_h - f\|_1 \le 2\varepsilon + \|\varphi_h - \varphi\|_1 \le 2\varepsilon + 2(R+1)\sup_x|\varphi(x+h)-\varphi(x)| \end{equation*}
for \(|h| \le 1\), and the last supremum tends to \(0\) by uniform continuity of \(\varphi\).
(ii) With the same \(\varphi\), the supports of \(\varphi_t\) and \(\varphi\) are disjoint once \(|t| > 2R\), so \(\|\varphi_t - \varphi\|_1 = 2\|\varphi\|_1\). Then for \(|t| > 2R\),
\begin{equation*} \bigl|\,\|f_t-f\|_1 - 2\|f\|_1\bigr| \le \|(f-\varphi)_t - (f-\varphi)\|_1 + 2\bigl|\,\|\varphi\|_1 - \|f\|_1\bigr| < 4\varepsilon , \end{equation*}
the first term being at most \(2\|f-\varphi\|_1\) by translation invariance.
(iii) With \(\varepsilon_n := a_n - a \to 0\), the substitution \(y = x+a\) and the triangle inequality give
\begin{equation*} \int|f_n(x+a_n)-f(x+a)|\,dx \le \|f_n - f\|_1 + \|f_{\varepsilon_n} - f\|_1 \longrightarrow 0 \end{equation*}
by hypothesis and (i). The second relation follows since \(\|f_n(\cdot+a_n)\|_1 = \|f_n\|_1\) and \(\|f(\cdot+a)\|_1 = \|f\|_1\), with \(|\|f_n\|_1 - \|f\|_1| \le \|f_n-f\|_1 \to 0\).
Exercises 4.7.105–4.7.111
(\(\circ\)) (Young [1035]) Suppose that integrable functions \(f_n\) on a space with a finite measure \(\mu\) converge a.e. to a function \(f\) and
\begin{equation*} \int_E f_n\,d\mu \to 0 \quad \text{as } n\to\infty,\ \mu(E)\to 0 . \end{equation*}
Prove that \(f\) is integrable.
The hypothesis makes the integrals of the \(f_n\) uniformly absolutely continuous; Fatou transfers that to \(f\), and truncation then bounds \(\int|f|\,d\mu\).
Read the hypothesis as: for every \(\varepsilon>0\) there are \(N\) and \(\delta_0>0\) with \(|\int_E f_n\,d\mu| < \varepsilon\) for all \(n \ge N\) and all \(\mu(E) < \delta_0\). Absolute continuity of the integral (Theorem 2.5.7) for the finitely many \(f_1,\dots,f_{N-1}\) supplies \(\delta_1,\dots,\delta_{N-1}\) with \(\int_E|f_j|\,d\mu < \varepsilon\) when \(\mu(E) < \delta_j\); put \(\delta := \min_j \delta_j\). For \(n \ge N\) and \(\mu(E) < \delta\), splitting \(E\) by the sign of \(f_n\) gives
\begin{equation*} \int_E|f_n|\,d\mu = \Bigl|\int_{E \cap \{f_n \ge 0\}}f_n\,d\mu\Bigr| + \Bigl|\int_{E \cap \{f_n<0\}}f_n\,d\mu\Bigr| < 2\varepsilon , \end{equation*}
so \(\sup_n\int_E|f_n|\,d\mu \le 2\varepsilon\) whenever \(\mu(E) < \delta\). Since \(|f_n|\mathbf{1}_E \to |f|\mathbf{1}_E\) a.e., Fatou’s theorem 2.8.3 gives
\begin{equation*} \int_E|f|\,d\mu \le \liminf_n \int_E|f_n|\,d\mu \le 2\varepsilon \qquad (\mu(E) < \delta) . \end{equation*}
Take \(\varepsilon = 1\) with its \(\delta\). The sets \(E_k := \{|f| > k\}\) decrease to the null set \(\{|f| = \infty\}\), so \(\mu(E_k) \to 0\) by countable additivity of the finite \(\mu\); choosing \(k\) with \(\mu(E_k) < \delta\),
\begin{equation*} \int_X|f|\,d\mu = \int_{E_k}|f|\,d\mu + \int_{X \setminus E_k}|f|\,d\mu \le 2 + k\,\mu(X) < \infty . \end{equation*}
Construct a sequence \(f_n\in L^1[0,1]\) with \(\|f_n\|_{L^1[0,1]}\le 1\) that is uniformly integrable on no set \(E\) of positive measure (in particular, the closure of this sequence in the weak topology of \(L^1(E)\) is not compact).
Take \(f_n := \mathbf{1}_{I_n}/\lambda(I_n)\), where \(I_1,I_2,\dots\) enumerates the countable family of dyadic intervals \([j2^{-m},(j+1)2^{-m}]\) of \([0,1]\); then \(\|f_n\|_{L^1[0,1]} = 1\), and no restriction to a set of positive measure is uniformly integrable.
Let \(\lambda(E) > 0\). By the Lebesgue differentiation theorem 5.6.2 applied to \(\mathbf{1}_E\), a.e. point of \(E\) is a density point, so fix a density point \(x \in (0,1)\) of \(E\) that is not a dyadic rational, and let \(D_m(x)\) be the unique dyadic interval of generation \(m\) containing \(x\). Since \(D_m(x) \subset \overline{B}(x,2^{-m})\) and \(\lambda(D_m(x)) = 2^{-m} = \tfrac12\lambda(\overline{B}(x,2^{-m}))\),
\begin{equation*} \frac{\lambda(D_m(x) \cap E)}{\lambda(D_m(x))} \ge 1 - \frac{\lambda(\overline{B}(x,2^{-m})\setminus E)}{2^{-m}} \longrightarrow 1 . \end{equation*}
With \(n_m\) the index of \(D_m(x)\) we have \(f_{n_m} = 2^m\mathbf{1}_{D_m(x)}\), so for any \(C>0\), taking \(m\) with \(2^m > C\) and the above ratio at least \(1/2\),
\begin{equation*} \int_{E \cap \{|f_{n_m}|>C\}}|f_{n_m}|\,d\lambda = \frac{\lambda(E \cap D_m(x))}{\lambda(D_m(x))} \ge \frac12 . \end{equation*}
Thus \(\sup_n\int_{E\cap\{|f_n|>C\}}|f_n|\,d\lambda \ge 1/2\) for every \(C\), and \(\lambda|_E\) being finite, Theorem 4.7.18 denies weak compactness of the closure in \(L^1(E)\).
(\(\circ\)) Let \((X,\mathcal{A},\mu)\) be a probability space. Prove that a set \(F\subset L^1(\mu)\) is uniformly integrable precisely when
\begin{equation*} \lim_{M\to+\infty}\ \sup_{f\in F}\int_X\max(|f|-M,0)\,d\mu=0 . \qquad (4.7.20) \end{equation*}
The two quantities squeeze each other: with \(\Phi(M) := \sup_{f \in F}\int_X\max(|f|-M,0)\,d\mu\) and \(\Psi( C) := \sup_{f\in F}\int_{\{|f|>C\}}|f|\,d\mu\), both nonincreasing,
\begin{equation*} \Phi(M) \le \Psi(M), \qquad \Psi(2M) \le 2\Phi(M) , \end{equation*}
so \(\Phi(M) \to 0\) iff \(\Psi( C) \to 0\), the latter being uniform integrability (Definition 4.5.1).
The first inequality holds because \(\max(|f|-M,0)\) vanishes off \(\{|f|>M\}\) and is at most \(|f|\) there. For the second, \(M \le |f|/2\) on \(\{|f| \ge 2M\}\) gives \(\max(|f|-M,0) \ge |f|/2\) there, so
\begin{equation*} \int_{\{|f| \ge 2M\}}|f|\,d\mu \le 2\int_{\{|f|\ge2M\}}\max(|f|-M,0)\,d\mu \le 2\Phi(M) \end{equation*}
by nonnegativity of the integrand; take the supremum over \(f\) and use \(\{|f|>2M\} \subset \{|f| \ge 2M\}\).
(see Bourgain [120]) Show that a set \(F\subset L^1[0,1]\) has compact closure in the weak topology if and only if, for every \(\varepsilon>0\), there exists a number \(C\) such that, for every function \(f\in F\), there is a measurable set \(S_f\subset[0,1]\) such that
\begin{equation*} \int_{S_f}|f(t)|\,dt\le\varepsilon\quad\text{and}\quad |f(t)|\le C\ \text{ for all } t\in[0,1]\setminus S_f . \end{equation*}
Lebesgue measure on \([0,1]\) being finite, Theorem 4.7.18 identifies weak relative compactness with uniform integrability, so it suffices to match the latter with the stated condition (B).
If \(F\) is uniformly integrable, Definition 4.5.1 gives, for \(\varepsilon > 0\), a \(C\) with \(\int_{\{|f|>C\}}|f|\,dt \le \varepsilon\) for all \(f \in F\), and \(S_f := \{|f| > C\}\) does the job.
Conversely assume (B). By Proposition 4.5.3 it is enough to have boundedness plus uniformly absolutely continuous integrals. Applying (B) with \(\varepsilon = 1\) and its \(C_1\),
\begin{equation*} \|f\|_{L^1} = \int_{S_f}|f|\,dt + \int_{[0,1]\setminus S_f}|f|\,dt \le 1 + C_1 . \end{equation*}
Given \(\varepsilon > 0\), apply (B) with \(\varepsilon/2\), obtaining \(C\), and put \(\delta := \varepsilon/(2C+2)\); then \(\lambda(A) < \delta\) gives, for every \(f \in F\),
\begin{equation*} \int_A|f|\,dt \le \int_{S_f}|f|\,dt + C\lambda(A) < \frac{\varepsilon}{2} + \frac{C\varepsilon}{2C+2} < \varepsilon , \end{equation*}
which is Definition 4.5.2.
Let \(A\) be a nonempty set. Suppose that for every \(n\in\mathbb{N}\) and \(\alpha\in A\), we are given a function \(f_{n,\alpha}\in L^2[0,1]\) such that, for every function \(g\) in \(L^2[0,1]\), one has \(\lim_{n\to\infty}(f_{n,\alpha},g)=0\) uniformly in \(\alpha\in A\). Prove that, for every \(g\in L^2[0,1]\) and \(\varepsilon>0\), there exists \(N\) such that for every interval \(I\subset[0,1]\) one has
\begin{equation*} \Big|\int_I g(x)f_{n,\alpha}(x)\,dx\Big|<\varepsilon,\qquad \forall\, n\ge N,\ \alpha\in A . \end{equation*}
Prove the analogous assertion for functions on a cube in \(\mathbb{R}^n\).
The hypothesis applied to \(g\mathbf{1}_I\) settles each fixed interval; the content is that \(N\) can be chosen uniformly in \(I\), and it comes from a Banach–Steinhaus bound plus one fixed finite partition, an arbitrary \(I\) differing from a union of its cells by two cells of small measure. Here \((u,v) = \int uv\,dx\).
There are \(N_0\) and \(R < \infty\) with \(\|f_{n,\alpha}\|_2 \le R\) for all \(n \ge N_0\), \(\alpha \in A\) (only large \(n\) is claimed, which is all the assertion needs). Indeed, the sets
\begin{equation*} C_{m,k} := \{g \in L^2[0,1] : |(f_{n,\alpha},g)| \le k \text{ for all } n \ge m,\ \alpha \in A\} \end{equation*}
are closed (intersections of preimages of \([-k,k]\) under continuous functionals) and cover \(L^2[0,1]\) (the hypothesis with \(\eta = 1\) puts each \(g\) into some \(C_{m,1}\)), so Baire’s theorem in the complete space \(L^2[0,1]\) (Theorem 4.1.3) gives \(g_0\), \(\rho>0\), \(m\), \(k\) with \(\{\|g-g_0\|_2 \le \rho\} \subset C_{m,k}\); then \(\|h\|_2 \le \rho\) forces \(|(f_{n,\alpha},h)| \le |(f_{n,\alpha},g_0+h)| + |(f_{n,\alpha},g_0)| \le 2k\) for \(n \ge m\), i.e. \(\|f_{n,\alpha}\|_2 \le 2k/\rho\).
Fix \(g\) and \(\varepsilon > 0\). Absolute continuity of the integral of \(|g|^2\) (Theorem 2.5.7) gives \(\delta>0\) with \(\int_S|g|^2dx \le (\varepsilon/(4R+4))^2\) whenever \(\lambda(S) < \delta\), so Cauchy–Bunyakowsky yields
\begin{equation*} \Bigl|\int_S gf_{n,\alpha}\,dx\Bigr| \le \|g\mathbf{1}_S\|_2\|f_{n,\alpha}\|_2 < \frac{\varepsilon}{4} \qquad (\lambda(S)<\delta,\ n \ge N_0) . \tag{2} \end{equation*}
Take \(k\) with \(1/k < \delta\), let \(J_1,\dots,J_k\) be the consecutive intervals of length \(1/k\), and apply the hypothesis to each \(g\mathbf{1}_{J_i}\) to get \(N \ge N_0\) with \(|\int_{J_i}gf_{n,\alpha}dx| < \varepsilon/(2k)\) for all \(n \ge N\), \(i\), \(\alpha\). For an interval \(I\) with endpoints in \(J_{i_0}\) and \(J_{i_1}\): if \(i_0 = i_1\) then \(\lambda(I) < \delta\) and (2) applies; otherwise \(I\) is, up to finitely many points, the disjoint union of \(E_1 := I \cap J_{i_0}\), the at most \(k\) cells \(J_i\) with \(i_0<i<i_1\), and \(E_2 := I \cap J_{i_1}\), where \(\lambda(E_1),\lambda(E_2) < \delta\), so
\begin{equation*} \Bigl|\int_I gf_{n,\alpha}\,dx\Bigr| < k\cdot\frac{\varepsilon}{2k} + \frac{\varepsilon}{4} + \frac{\varepsilon}{4} = \varepsilon , \end{equation*}
with \(N\) independent of \(I\).
For a cube, normalized to \(Q = [0,1]^d\), the last two paragraphs give \(N_0, R, \delta\) verbatim. Partition \(Q\) into the \(k^d\) subcubes \(Q_i\) of edge \(h := 1/k\) with \(2d^{3/2}h < \delta\), and choose \(N \ge N_0\) with \(|\int_{Q_i}gf_{n,\alpha}dx| < \varepsilon/(2k^d)\) for all \(i \le k^d\), \(n \ge N\), \(\alpha \in A\). For a box \(I = \prod_j[a_j,b_j] \subset Q\) put \(R_I := \bigcup\{Q_i : Q_i \subset I\}\). A point of \(I\) farther than the subcube diameter \(\sqrt{d}h\) from \(\mathbb{R}^d \setminus I\) lies in \(R_I\), so
\begin{equation*} I \setminus R_I \subset I \setminus \prod_{j=1}^{d} \bigl(a_j + \sqrt{d}h,\ b_j - \sqrt{d}h\bigr) , \end{equation*}
using \(\operatorname{dist}(x,\mathbb{R}^d\setminus I) = \min_j\min(x_j-a_j,b_j-x_j)\) and reading an empty interval as \(\emptyset\). With \(v_j := b_j-a_j\), \(u_j := \max(v_j - 2\sqrt{d}h,0)\) and \(\prod_j v_j - \prod_j u_j \le \sum_j(v_j-u_j)\) for \(0 \le u_j \le v_j \le 1\),
\begin{equation*} \lambda_d(I \setminus R_I) \le \sum_{j=1}^{d}2\sqrt{d}\,h = 2d^{3/2}h < \delta , \end{equation*}
so summing over the (essentially disjoint) \(Q_i \subset I\) and applying (2) to \(I \setminus R_I\),
\begin{equation*} \Bigl|\int_I gf_{n,\alpha}\,dx\Bigr| < k^{d}\cdot\frac{\varepsilon}{2k^{d}} + \frac{\varepsilon}{4} < \varepsilon . \end{equation*}
(\(\circ\)) Let \(\mu\) be a finite nonnegative measure and let \(1\le p<\infty\). Prove that a set \(K\subset L^p(\mu)\) has compact closure in \(L^p(\mu)\) precisely when the set \(\{|f|^p:\ f\in K\}\) is uniformly integrable and every sequence in \(K\) contains a subsequence convergent in measure.
Since \(L^p(\mu)\) is complete (Theorem 4.1.3), compact closure means total boundedness, equivalently that every sequence in \(K\) has an \(L^p\)-convergent subsequence; the two stated conditions are exactly what the Lebesgue–Vitali theorem 4.5.4 needs.
Necessity. An \(L^p\)-convergent subsequence converges in measure by Chebyshev, \(\mu(|f_j - f| \ge c) \le c^{-p}\|f_j-f\|_p^p\). For uniform integrability of \(\{|f|^p\}\) it suffices, \(\mu\) being finite, to check boundedness and uniform absolute continuity (Proposition 4.5.3). Given \(\varepsilon>0\), put \(\varepsilon’ := \tfrac12\varepsilon^{1/p}\) and take a finite \(\varepsilon’\)-net \(f_1,\dots,f_m \in K\); boundedness is \(\|f\|_p \le \max_i\|f_i\|_p + \varepsilon’\), and Theorem 2.5.7 gives \(\delta>0\) with \(\int_A|f_i|^p d\mu \le (\varepsilon’)^p\) for all \(i\) when \(\mu(A)<\delta\), whence by Minkowski (Lemma 4.1.1 on \(\mu|_A\))
\begin{equation*} \Bigl(\int_A|f|^p d\mu\Bigr)^{1/p} \le \Bigl(\int_A|f-f_i|^p d\mu\Bigr)^{1/p} + \Bigl(\int_A|f_i|^p d\mu\Bigr)^{1/p} \le 2\varepsilon’ , \end{equation*}
i.e. \(\int_A|f|^p d\mu \le \varepsilon\).
Sufficiency. Given \(\{f_j\} \subset K\), pass to a subsequence converging in measure to some \(f\) and then (Theorem 2.2.5) to a further one converging a.e. Then \(|f|^p\) is integrable by Fatou’s theorem 2.8.3 (the \(L^1\)-bound coming from uniform integrability with Proposition 4.5.3). Put \(h_k := |f_{j_k} - f|^p\).
(i) \(h_k \to 0\) in measure, since \(\{h_k > c\} = \{|f_{j_k}-f| > c^{1/p}\}\).
(ii) \(\{h_k\}\) is uniformly integrable: convexity of \(t \mapsto t^p\) gives \(h_k \le G_k := 2^{p-1}(|f_{j_k}|^p + |f|^p)\), and \(\{G_k\}\) is uniformly integrable because boundedness and uniform absolute continuity pass through sums and constants (Proposition 4.5.3); finally \(0 \le h \le G\) forces \(\int_{\{h>C\}}h\,d\mu \le \int_{\{G>C\}}G\,d\mu\).
By Theorem 4.5.4 applied to \(h_k\) with limit \(0\), \(\|f_{j_k}-f\|_p^p = \int h_k\,d\mu \to 0\).
Let \(1\le p<\infty\) and let \(K\) be a bounded set in \(L^p(\mathbb{R}^n)\).
(i) (A.N. Kolmogorov; for \(p=1\), A.N. Tulaikov) Prove that the closure of \(K\) in \(L^p(\mathbb{R}^n)\) is compact precisely when the following conditions are fulfilled:
(a) one has
\begin{equation*} \sup_{f\in K}\ \lim_{C\to\infty}\int_{|x|>C}|f(x)|^p\,dx=0 , \end{equation*}
(b) for every \(\varepsilon>0\), there exists \(r>0\) such that \(\sup_{f\in K}\|f-S_rf\|_p\le\varepsilon\), where \(S_rf\) is Steklov’s function defined by the equality
\begin{equation*} S_rf(x):=\lambda_n\big(B(x,r)\big)^{-1}\int_{B(x,r)}f(y)\,dy , \end{equation*}
\(B(x,r)\) is the ball of radius \(r\) centered at \(x\).
(ii) (M. Riesz) Show that the compactness of the closure of \(K\) is equivalent also to condition (a) combined with
(b’) one has
\begin{equation*} \sup_{f\in K}\ \lim_{h\to0}\int_{\mathbb{R}^n}|f(x+h)-f(x)|^p\,dx=0 . \end{equation*}
(iii) (V.N. Sudakov) Show that conditions (a) and (b) (or (a) and (b’)) yield the boundedness of \(K\) in \(L^p(\mathbb{R}^n)\), hence there is no need to require boundedness in advance.
As printed, the suprema in (a) and (b’) stand outside the limits, which makes them vacuous (each single \(f\) satisfies both, by dominated convergence and Lemma 4.2.3); we read them uniformly in \(f \in K\), as (b) is stated.
Write \(b_r := \lambda_n(B_r)\), \(B_r := B(0,r)\), \(g_r := b_r^{-1}\mathbf{1}_{B_r}\), so \(\|g_r\|_1 = 1\) and \(S_rf = f*g_r\) (\(g_r\) is even); \(q\) is conjugate to \(p\). Two lemmas carry everything.
(A) Young: \(\|f*g\|_p \le \|f\|_p\|g\|_1\). Jensen for \(t \mapsto t^p\) against the probability measure \(|g(z)|dz/\|g\|_1\) gives \(|f*g(x)|^p \le \|g\|_1^{p-1}\int|f(x-z)|^p|g(z)|dz\), and Tonelli with translation invariance finishes. In particular \(\|S_rf\|_p \le \|f\|_p\).
(B) \(\|S_rf - f\|_p \le \sup_{|z| \le r}\|f(\cdot - z) - f\|_p\). Jensen against \(b_r^{-1}dz\) on \(B_r\) applied to \(S_rf(x)-f(x) = b_r^{-1}\int_{B_r}[f(x-z)-f(x)]dz\), then Tonelli.
(i) Necessity. Compactness gives, for \(\varepsilon>0\), a finite \(\varepsilon/3\)-net \(f_1,\dots,f_m \in K\). For (a), dominated convergence gives \(C_0\) with \((\int_{|x|>C}|f_i|^pdx)^{1/p} \le \varepsilon/3\) for all \(i\) and \(C \ge C_0\), and Minkowski transfers it:
\begin{equation*} \Bigl(\int_{|x|>C}|f|^p dx\Bigr)^{1/p} \le \|f-f_i\|_p + \Bigl(\int_{|x|>C}|f_i|^p dx\Bigr)^{1/p} \le \frac{2\varepsilon}{3} . \end{equation*}
For (b), Lemma A gives \(\|f - S_rf\|_p \le 2\|f-f_i\|_p + \|f_i - S_rf_i\|_p\), and Lemma B with continuity of translation (Lemma 4.2.3) sends each of the finitely many \(\|f_i - S_rf_i\|_p\) to \(0\) as \(r \to 0\).
(i) Sufficiency. Let \(M := \sup_K\|f\|_p < \infty\) and let (a), (b) hold; by completeness (Theorem 4.1.3) it suffices to exhibit a finite \(\varepsilon\)-net.
Pick \(r\) with \(\sup_K\|f - f*g_r\|_p \le \varepsilon/5\) by (b), and \(\varphi \in C_0^\infty(\mathbb{R}^n)\) with \(\|\varphi - g_r\|_1 \le \varepsilon/(5(M+1))\) by density in \(L^1\) (Corollary 4.2.2); Lemma A then gives
\begin{equation*} \sup_{f \in K}\|f - f*\varphi\|_p \le \frac{2\varepsilon}{5} , \tag{3} \end{equation*}
so a finite \(2\varepsilon/5\)-net for \(\{f*\varphi\}\) yields a \(4\varepsilon/5\)-net for \(K\).
Let \(\varphi\) vanish off \(B(0,\rho)\), take \(C\) with \(\sup_K\|f\mathbf{1}_{\{|y|>C\}}\|_p \le \varepsilon/(5(\|\varphi\|_1+1))\) by (a), and put \(R := C+\rho\). For \(|x| > R\), \(\varphi(x-y) \ne 0\) forces \(|y| > C\), so \(f*\varphi = (f\mathbf{1}_{\{|y|>C\}})*\varphi\) there and Lemma A gives
\begin{equation*} \Bigl(\int_{|x|>R}|f*\varphi|^p dx\Bigr)^{1/p} \le \|f\mathbf{1}_{\{|y|>C\}}\|_p\|\varphi\|_1 \le \frac{\varepsilon}{5} . \tag{4} \end{equation*}
On \(\overline{B}(0,R)\), Hölder gives \(|f*\varphi| \le M\|\varphi\|_q\) and \(|f*\varphi(x) - f*\varphi(x’)| \le M\omega(|x-x’|)\) with \(\omega(t) := \sup_{|h| \le t}\|\varphi(\cdot+h)-\varphi\|_q \to 0\) (Lemma 4.2.3 for \(q<\infty\); uniform continuity of \(\varphi\) for \(q=\infty\)). So the restrictions are uniformly bounded and equicontinuous, hence totally bounded in \(C(\overline{B}(0,R))\) by Arzela–Ascoli: choosing \(\sigma\) with \(\sigma\lambda_n(B(0,R))^{1/p} \le \varepsilon/5\) and \(f_1,\dots,f_m \in K\) that \(\sigma\)-approximate in the sup-norm there, the functions \(v_i := (f_i*\varphi)\mathbf{1}_{B(0,R)}\) satisfy, with (4),
\begin{equation*} \|f*\varphi - v_i\|_p \le \sigma\lambda_n(B(0,R))^{1/p} + \Bigl(\int_{|x|>R}|f*\varphi|^pdx\Bigr)^{1/p} \le \frac{2\varepsilon}{5} . \end{equation*}
(ii) If the closure is compact, then with an \(\varepsilon/3\)-net and translation invariance of the norm,
\begin{equation*} \|f(\cdot+h)-f\|_p \le \frac{2\varepsilon}{3} + \max_{i \le m}\|f_i(\cdot+h)-f_i\|_p , \end{equation*}
the maximum tending to \(0\) by Lemma 4.2.3: this is (b’). Conversely (b’) gives (b) via Lemma B, since \(\sup_K\|f - S_rf\|_p \le \sup_{|z|\le r}\sup_K\|f(\cdot-z)-f\|_p\), so (i) applies.
(iii) Assume (a) and (b) (the case of (b’) is included, by (b’)\(\Rightarrow\)(b) above). Take \(C_0\) with \(\int_{|x|>C_0}|f|^pdx \le 1\) and \(r\) with \(\|f - S_rf\|_p \le 1\) for all \(f \in K\), and put \(U := \overline{B}(0,C_0)\), \(u := f\mathbf{1}_U\), \(v := f - u\), so \(\|v\|_p \le 1\) and, by Lemma A,
\begin{equation*} \|u - S_ru\|_p \le \|f - S_rf\|_p + \|v\|_p + \|S_rv\|_p \le 3 . \tag{6} \end{equation*}
Identify \(L^p(U)\) with the functions vanishing off \(U\), write \(\tilde u\) for extension by zero, and let \(Tu := (\tilde u * g_r)|_U\), of norm \(\le 1\) by Lemma A.
\(T\) is compact: for \(\|u\|_p \le 1\) the functions \(\tilde u * g_r\) are bounded in \(L^p\) and vanish off \(B(0,C_0+r)\), so (a) is trivial for them, while
\begin{equation*} \|(\tilde u*g_r)(\cdot+h) - \tilde u*g_r\|_p \le \|\tilde u\|_p\|g_r(\cdot+h)-g_r\|_1 \longrightarrow 0 \quad (h \to 0) \end{equation*}
by Lemma 4.2.3 for the fixed \(g_r \in L^1\), which is (b’) uniformly; part (ii) makes their closure compact, and restriction to \(U\) is bounded.
\(1\) is not an eigenvalue: let \(u\) vanish off \(U\) with \(S_ru = u\) a.e. on \(U\). Then \(u \in L^1\) (\(U\) bounded, Hölder), and \(w := S_ru\) is continuous, since \(|w(x)-w(x’)| \le b_r^{-1}\int_{B(x,r)\triangle B(x’,r)}|u|\,dy \to 0\) by absolute continuity of the integral (Theorem 2.5.7). Let \(S := \max_U|w|\), let \(x_0\) maximize \(|x|\) on the nonempty compact set \(F := \{x \in U : |w(x)| = S\}\). Since \(|u| \le S\) a.e. on \(U\) and \(u = 0\) off \(U\),
\begin{equation*} S = |w(x_0)| \le b_r^{-1}\int_{B(x_0,r) \cap U}|u|\,dy \le \frac{\lambda_n(B(x_0,r) \cap U)}{b_r}\,S \le S . \end{equation*}
If \(S>0\), equality throughout forces the open set \(B(x_0,r)\setminus U\) to be null, hence empty, and \(|u| = S\) a.e. on \(B(x_0,r)\); as \(u = w\) a.e. there with \(w\) continuous, \(|w| \equiv S\) on \(B(x_0,r) \subset U\), which contains points of \(F\) of norm exceeding \(|x_0|\) — contradiction. So \(S = 0\) and \(u = 0\).
Hence \(I-T\) is injective with \(T\) compact, so the Fredholm alternative makes it bijective and the open mapping theorem makes \((I-T)^{-1}\) bounded. By (6), \(\|(I-T)u\|_{L^p(U)} \le \|u - S_ru\|_p \le 3\), so \(\|u\|_p \le 3\|(I-T)^{-1}\|\) and
\begin{equation*} \|f\|_p^p = \|u\|_p^p + \|v\|_p^p \le \bigl(3\|(I-T)^{-1}\|\bigr)^{p} + 1 \qquad \text{for all } f \in K . \end{equation*}
Exercises 4.7.112–4.7.118
(\(\circ\)) Let \(\mu\) be a signed measure on a measurable space \((X,\mathcal{A})\) such that \(\mu(X) = 0\). Prove that \(\|\mu\| = 2\sup_{A\in\mathcal{A}}|\mu(A)|\). In particular, for probability measures \(\mu_1\) and \(\mu_2\), we have \(\|\mu_1-\mu_2\| = 2\sup_{A\in\mathcal{A}}|(\mu_1-\mu_2)(A)|\).
The supremum is attained at the positive part \(X^+\) of the Hahn decomposition \(X = X^+ \cup X^-\) (Theorem 3.1.1), for which Corollary 3.1.2 and Definition 3.1.4 give \(\mu^{+}(A) = \mu(A \cap X^+)\), \(\mu^{-}(A) = -\mu(A \cap X^-)\), so that
\begin{equation*} \|\mu\| = |\mu|(X) = \mu(X^+) - \mu(X^-) = 2\mu(X^+) , \end{equation*}
the last step because \(\mu(X^+) + \mu(X^-) = \mu(X) = 0\).
For arbitrary \(A\), splitting \(\mu(A) = \mu(A \cap X^+) + \mu(A \cap X^-)\) and dropping the nonpositive (resp. nonnegative) term gives
\begin{equation*} \mu(X^-) \le \mu(A \cap X^-) \le \mu(A) \le \mu(A \cap X^+) \le \mu(X^+) , \end{equation*}
the outer inequalities because \(\mu(X^-\setminus A) \le 0 \le \mu(X^+ \setminus A)\). Hence \(|\mu(A)| \le \mu(X^+) = \|\mu\|/2\), with equality at \(A = X^+\).
For probability measures, \(\mu := \mu_1 - \mu_2\) has \(\mu(X) = 0\).
Construct a sequence of bounded signed countably additive measures on some algebra such that this sequence is uniformly bounded on every set in this algebra, but is not bounded in the variation norm.
Take \(X = \mathbb{N}\), let \(\mathcal{A}\) be the algebra of finite and cofinite subsets of \(\mathbb{N}\), set \(c_k = (-1)^k/k\) (so \(\sum_k c_k\) converges, not absolutely), and put
\begin{equation*} \mu_n(A) := \sum_{k\in A,\ k\le n} c_k = \sum_{k=1}^{n} c_k\,\delta_k(A),\qquad A\in\mathcal{A}. \end{equation*}
Each \(\mu_n\) is a finite linear combination of Dirac measures, hence a bounded countably additive signed measure on \(\mathcal{A}\).
Pointwise boundedness, with \(S_n := \sum_{k\le n}c_k\) and \(S := \sup_m|S_m|<\infty\) (the series converges):
(i) \(A\) finite: \(|\mu_n(A)|\le\sum_{k\in A}|c_k|<\infty\), a bound free of \(n\).
(ii) \(A\) cofinite, \(F := \mathbb{N}\setminus A\) finite:
\begin{equation*} |\mu_n(A)| = \Big|S_n - \sum_{k\in F,\ k\le n}c_k\Big|\ \le\ S + \sum_{k\in F}|c_k|<\infty . \end{equation*}
Unboundedness in variation: the sets \(\{1\},\dots,\{n\},\ \mathbb{N}\setminus\{1,\dots,n\}\) form a finite \(\mathcal{A}\)-partition of \(\mathbb{N}\), so
\begin{equation*} \|\mu_n\| = |\mu_n|(\mathbb{N})\ \ge\ \sum_{k=1}^{n}|\mu_n(\{k\})| = \sum_{k=1}^n \frac1k \longrightarrow \infty . \end{equation*}
Find a sequence of nonnegative countably additive measures that has a finite limit on every set in some algebra \(\mathcal{A}\), but does not converge on some set in \(\sigma(\mathcal{A})\).
On \(X = [0,1]\) with \(\mathcal{A}\) the algebra of finite unions of intervals (so \(\sigma(\mathcal{A}) = \mathcal{B}([0,1])\)), take
\begin{equation*} \mu_{2m-1} := \lambda,\qquad \mu_{2m} := f_m\cdot\lambda\quad (m\in\mathbb{N}), \end{equation*}
with \(f_m\) built as follows. Enumerate the rationals of \((0,1)\) as \(\{q_j\}\) and put
\begin{equation*} E := \bigcup_{j\ge1}\big(q_j - 2^{-j-3},\,q_j+2^{-j-3}\big)\cap(0,1), \end{equation*}
an open dense subset of \([0,1]\) with \(\lambda(E)\le\sum_j 2^{-j-2} = 1/4\). For \(k\le m\) the open set \(E\cap\big(\frac{k-1}{m},\frac{k}{m}\big)\) is nonempty by density, so it contains a nondegenerate closed interval; let \(\varphi_{m,k}\ge0\) be a continuous tent supported there with \(\int_0^1\varphi_{m,k} = 1/m\), and set \(f_m := \sum_{k\le m}\varphi_{m,k}\). Then \(f_m\ge0\) is continuous, vanishes off \(E\), and \(\int_0^1 f_m = 1\); in particular each \(\mu_n\) is a finite nonnegative countably additive measure.
Convergence on \(\mathcal{A}\). For \(0\le a\le b\le1\), counting the grid intervals \(\big(\frac{k-1}{m},\frac{k}{m}\big)\) inside \([a,b]\) and those meeting \([a,b]\) gives \(N_-/m\le\int_a^b f_m\le N_+/m\) with \(m(b-a)-2\le N_-\le N_+\le m(b-a)+2\), so
\begin{equation*} \Big|\int_a^b f_m(t)\,dt - (b-a)\Big|\le \frac{2}{m}\ \xrightarrow[m\to\infty]{}\ 0 . \end{equation*}
Hence \(\mu_n(J)\to\lambda(J)\) for every interval \(J\), and by additivity over a decomposition of \(A\in\mathcal{A}\) into finitely many disjoint intervals, \(\mu_n(A)\to\lambda(A)\), a finite limit.
Divergence on \(E\in\sigma(\mathcal{A})\). Since \(f_m = 0\) off \(E\),
\begin{equation*} \mu_{2m}(E) = \int_E f_m\,dt = 1\quad\text{for all }m,\qquad \mu_{2m-1}(E) = \lambda(E)\le \tfrac14 , \end{equation*}
so \(\{\mu_n(E)\}\) has two distinct limit points and does not converge.
Prove that if a \(\sigma\)-algebra \(\mathcal{A}\) is infinite, then the topology of convergence of measures on all sets in \(\mathcal{A}\) cannot be generated by a norm.
No norm generates this topology \(\tau\), because every basic \(\tau\)-neighbourhood of \(0\) contains a whole line. Here \(\mathcal{M} = \mathcal{M}(X,\mathcal{A})\) is the space of bounded countably additive real measures, and by 4.7(v) a base at \(0\) is given by
\begin{equation*} W_{A_1,\dots,A_n,\varepsilon} = \{\mu\in\mathcal{M}:\ |\mu(A_i)|<\varepsilon,\ i\le n\}, \qquad A_i\in\mathcal{A},\ \varepsilon>0 . \end{equation*}
Claim: for any \(A_1,\dots,A_n\in\mathcal{A}\) there is \(\mu\in\mathcal{M}\), \(\mu\ne0\), with \(\mu(A_i) = 0\) for all \(i\le n\). The \(A_i\) generate the finite \(\mathcal{A}\)-partition of \(X\) into the nonempty atoms
\begin{equation*} P_\theta = \bigcap_{i=1}^{n}A_i^{\theta_i},\qquad \theta\in\{0,1\}^n, \quad A^1 := A,\ A^0 := X\setminus A, \end{equation*}
listed as \(P_1,\dots,P_m\); each \(P_j\) lies inside \(A_i\) or is disjoint from it. Since \(A = \bigcup_{j\le m}(A\cap P_j)\), the map \(A\mapsto(A\cap P_1,\dots,A\cap P_m)\) is injective on \(\mathcal{A}\), so infiniteness of \(\mathcal{A}\) forces some trace \(\{A\cap P_j: A\in\mathcal{A}\}\) to be infinite; in particular there is \(A\in\mathcal{A}\) with \(\emptyset\ne A\subsetneq P_j\). Pick \(x\in A\), \(y\in P_j\setminus A\) and set \(\mu := \delta_x-\delta_y\in\mathcal{M}\). Then \(\mu(A_i) = 1_{A_i}(x)-1_{A_i}(y) = 0\) as \(x,y\in P_j\), while \(\mu(A) = 1\), so \(\mu\ne0\).
If \(\tau\) were generated by a norm \(p\), the ball \(\{p<1\}\) would contain some \(W_{A_1,\dots,A_n,\varepsilon}\), which by the claim contains \(t\mu\) for every \(t>0\); then \(t\,p(\mu)<1\) for all \(t>0\) forces \(p(\mu) = 0\) with \(\mu\ne0\), a contradiction.
Let \(\mathcal{A}\) be the Borel \(\sigma\)-algebra of \([0,1]\). Show that on the space \(\mathcal{M}\) of all countably additive measures on \(\mathcal{A}\) all three topologies considered in 4.7(v), i.e., the topology of setwise convergence, the topology generated by the duality with the space of all bounded \(\mathcal{A}\)-measurable functions, and the topology \(\sigma(\mathcal{M},\mathcal{M}^{*})\), are distinct, although the collections of convergent countable sequences in these topologies are the same.
The three topologies are
\begin{equation*} \tau_1 = \sigma(\mathcal{M},\mathcal{F})\ \subsetneq\ \tau_2 = \sigma(\mathcal{M},\mathcal{B}) \ \subsetneq\ \tau_3 = \sigma(\mathcal{M},\mathcal{M}^{*}), \end{equation*}
yet a sequence convergent in the weakest converges in the strongest. Here \(\mathcal{M} = \mathcal{M}([0,1],\mathcal{A})\), \(\mathcal{A} = \mathcal{B}([0,1])\), consists of bounded measures (Corollary 3.1.3) and is Banach in the variation norm (Theorem 4.6.1); \(\mathcal{F}\subset\mathcal{B}\) are the simple and the bounded \(\mathcal{A}\)-measurable functions, embedded in \(\mathcal{M}^{*}\) by \(f\mapsto L_f\), \(L_f(\mu) = \int f\,d\mu\), which is norm-continuous since \(|L_f(\mu)|\le\sup_x|f(x)|\,\|\mu\|\) and injective since \(L_f(\delta_x) = f(x)\). All three spaces separate \(\mathcal{M}\), already through the \(L_{1_A}\), and \(\tau_1\) is setwise convergence.
Distinctness reduces to \(\mathcal{F}\ne\mathcal{B}\ne\mathcal{M}^{*}\), by the duality fact that for a separating subspace \(G\) of the algebraic dual of \(E\) the dual of \((E,\sigma(E,G))\) is \(G\): a \(\sigma(E,G)\)-continuous \(\varphi\) satisfies \(|\varphi|<1\) on some \(\{|g_i|<\varepsilon,\ i\le n\}\), hence vanishes on \(\bigcap_{i\le n}\ker g_i\) (replace \(x\) by \(tx\), \(t\to\infty\)), hence factors through \(x\mapsto(g_1(x),\dots,g_n(x))\) and is a linear combination of the \(g_i\). So \(\sigma(E,G) = \sigma(E,H)\) with \(G\subset H\) forces \(G = H\).
(i) \(\mathcal{F}\ne\mathcal{B}\): \(f(x) = x\) is bounded Borel but not simple, and \(g\mapsto L_g\) is injective.
(ii) \(\mathcal{B}\ne\mathcal{M}^{*}\): put \(\ell(\mu) := \mu_s([0,1])\), where \(\mu = \mu_a+\mu_s\), \(\mu_a\ll\lambda\), \(\mu_s\perp\lambda\), is the Lebesgue decomposition (Theorem 3.2.3, applicable as \(\mu\) and \(\lambda\) are bounded). It is unique, since a measure \(\sigma\) both \(\ll\lambda\) and carried by a \(\lambda\)-null \(N\) has \(\sigma(A) = \sigma(A\cap N) = 0\); uniqueness gives linearity of \(\mu\mapsto\mu_s\) (a sum of two \(\lambda\)-singular measures is carried by the union of their null carriers). With \(N\) a \(\lambda\)-null carrier of \(|\mu_s|\) we get \(\mu_s(A) = \mu(A\cap N)\), so \(|\ell(\mu)|\le\|\mu\|\) and \(\ell\in\mathcal{M}^{*}\). Were \(\ell = L_f\), then \(\delta_x\perp\lambda\) would give \(f(x) = \ell(\delta_x) = 1\) for every \(x\), whence \(L_f(\lambda) = \lambda([0,1]) = 1\), while \(\lambda_s = 0\) gives \(\ell(\lambda) = 0\).
Same convergent sequences: \(\tau_1\subset\tau_2\subset\tau_3\) gives one direction, and if \(\mu_n\to\mu\) setwise then \(\mu_n\to\mu\) in \(\tau_3\) by Corollary 4.7.26 (whose hypothesis is exactly setwise convergence of a sequence in \(\mathcal{M}\)), hence in \(\tau_2\) as well.
Let \(\mathcal{A}\) be an algebra of sets and let \(\{\mu_n\}\) be a uniformly countably additive sequence of bounded measures on the generated \(\sigma\)-algebra \(\sigma(\mathcal{A})\). Prove that if, for every \(A\in\mathcal{A}\), there exists a finite limit \(\lim_{n\to\infty}\mu_n(A)\), then the same is true for every \(A\in\sigma(\mathcal{A})\).
Each \(\{\mu_n(A)\}\), \(A\in\sigma(\mathcal{A})\), is Cauchy, by approximating \(A\) in the dominating measure
\begin{equation*} \nu := \sum_{n=1}^{\infty}2^{-n}\big(1+\|\mu_n\|\big)^{-1}|\mu_n| , \end{equation*}
a nonnegative measure on \(\sigma(\mathcal{A})\) with \(\nu(X)\le1\) and \(\mu_n\ll\nu\) for all \(n\). Uniform countable additivity is hypothesis (i) of Lemma 4.6.5, whose condition (iv) for this \(\nu\) says: for every \(\varepsilon>0\) there is \(\delta>0\) with
\begin{equation*} B\in\sigma(\mathcal{A}),\ \nu(B)\le\delta\ \Longrightarrow\ |\mu_n(B)|\le\varepsilon\ \text{ for all }n. \tag{\(*\)} \end{equation*}
Approximation: every \(A\in\sigma(\mathcal{A})\) admits \(A_0\in\mathcal{A}\) with \(\nu(A\triangle A_0)<\delta\). The family \(\mathcal{D}\) of such \(A\) contains \(\mathcal{A}\) and is closed under complements, since \((X\setminus A)\triangle(X\setminus A_0) = A\triangle A_0\). For \(A = \bigcup_{k\ge1}A^k\) with \(A^k\in\mathcal{D}\), continuity from below of \(\nu\) gives \(N\) with \(\nu\big(A\setminus\bigcup_{k\le N}A^k\big)<\delta/2\); choosing \(B_k\in\mathcal{A}\) with \(\nu(A^k\triangle B_k)<\delta2^{-k-1}\) and \(A_0 := \bigcup_{k\le N}B_k\in\mathcal{A}\),
\begin{equation*} A\,\triangle\,A_0\ \subset\ \Big(A\setminus\bigcup_{k\le N}A^k\Big) \cup\bigcup_{k\le N}\big(A^k\triangle B_k\big) \end{equation*}
has \(\nu\)-measure \(<\delta\). So \(\mathcal{D}\) is a \(\sigma\)-algebra containing \(\mathcal{A}\), i.e. \(\mathcal{D} = \sigma(\mathcal{A})\).
Now fix \(A\in\sigma(\mathcal{A})\) and \(\varepsilon>0\), take \(\delta\) from \((*)\) and \(A_0\in\mathcal{A}\) with \(\nu(A\triangle A_0)<\delta\). Both \(A\setminus A_0\) and \(A_0\setminus A\) then have \(\nu\)-measure \(<\delta\), so \((*)\) gives \(|\mu_n(A)-\mu_n(A_0)| = |\mu_n(A\setminus A_0)-\mu_n(A_0\setminus A)|\le2\varepsilon\) for all \(n\); as \(\lim_n\mu_n(A_0)\) exists, there is \(N\) with \(|\mu_n(A_0)-\mu_m(A_0)|\le\varepsilon\) for \(n,m\ge N\), whence
\begin{equation*} |\mu_n(A)-\mu_m(A)|\le 2\varepsilon+\varepsilon+2\varepsilon = 5\varepsilon\qquad (n,m\ge N). \end{equation*}
(Drewnowski) (i) Let \(\mathcal{A}\) be a \(\sigma\)-algebra and let \(\mu\colon\mathcal{A}\to\mathbb{R}^1\) be a bounded additive function. Suppose that \(A_n\in\mathcal{A}\) are disjoint sets. Prove that there exists a sequence \(\{n_k\}\) such that \(\mu\) is countably additive on the \(\sigma\)-algebra generated by \(\{A_{n_k}\}\).
(ii) Show that if in (i) we are given a sequence of bounded additive functions \(\mu_i\) on \(\mathcal{A}\), then one can choose a common sequence \(\{n_k\}\) for all \(\mu_i\).
(i) Thin out \(\{A_n\}\) so that the semivariation of the tails decays geometrically. Write \(M := \sup_{A\in\mathcal{A}}|\mu(A)|<\infty\) and
\begin{equation*} \eta(A) := \sup\{|\mu(B)|:\ B\in\mathcal{A},\ B\subset A\}\ \le\ M , \end{equation*}
a monotone set function. For pairwise disjoint \(C_k\in\mathcal{A}\), grouping \(k\le N\) by the sign of \(\mu(C_k)\) and using finite additivity gives \(\sum_{k\le N}|\mu(C_k)|\le2M\), hence \(\sum_k|\mu(C_k)|\le2M\); consequently \(\eta\) is exhaustive, since for disjoint \(D_j\) the choice of \(B_j\subset D_j\) with \(|\mu(B_j)|\ge\eta(D_j)/2\) makes \(\sum_j|\mu(B_j)|<\infty\), so \(\eta(D_j)\to0\).
Criterion: for pairwise disjoint \(B_k\in\mathcal{A}\) with \(R := X\setminus\bigcup_kB_k\), the \(\sigma\)-algebra \(\Sigma := \sigma(\{B_k\})\) consists of the sets \(\bigcup_{k\in S}B_k\) and \(R\cup\bigcup_{k\in S}B_k\), \(S\subset\mathbb{N}\) (that family contains each \(B_k\) and is closed under complements and countable unions), and \(\mu\) is countably additive on \(\Sigma\) iff
\begin{equation*} \mu\Big(\bigcup_{k\in S}B_k\Big) = \sum_{k\in S}\mu(B_k) \qquad\text{for every }S\subset\mathbb{N}. \tag{\(**\)} \end{equation*}
Necessity is clear. Conversely, a disjoint decomposition \(C = \bigcup_iC_i\) in \(\Sigma\) has \(C_i = \bigcup_{k\in S_i}B_k\) with the \(S_i\) disjoint of union \(S\), at most one \(C_i\) also carrying \(R\), so by finite additivity and \((**)\)
\begin{equation*} \sum_i\mu(C_i) = \mu( R)+\sum_{k\in S}\mu(B_k) = \mu( R)+\mu\Big(\bigcup_{k\in S}B_k\Big) = \mu( C), \end{equation*}
the regrouping of the double series being licit since \(\sum_k|\mu(B_k)|\le2M\) (omit \(\mu( R)\) throughout if \(R\not\subset C\)).
Construction: build infinite sets \(\mathbb{N} = M_0\supset M_1\supset\cdots\) with \(n_k := \min M_k\) strictly increasing and \(\eta\big(\bigcup_{n\in M_k}A_n\big)<2^{-k}\). Given \(M_{k-1}\), partition it into infinitely many disjoint infinite \(P_1,P_2,\dots\) and put \(D_j := \bigcup_{n\in P_j}A_n\in\mathcal{A}\) (here \(\mathcal{A}\) is a \(\sigma\)-algebra); the \(D_j\) are disjoint, so \(\eta(D_j)\to0\) by exhaustivity, and for \(j\) with \(\eta(D_j)<2^{-k}\) set \(M_k := P_j\setminus\{1,\dots,n_{k-1}\}\), whose \(A\)-union lies in \(D_j\).
Since \(\{n_j: j>k\}\subset M_k\), monotonicity of \(\eta\) gives
\begin{equation*} \Big|\mu\Big(\bigcup_{j\in S,\ j>k}A_{n_j}\Big)\Big|\le\eta\Big(\bigcup_{n\in M_k}A_n\Big)<2^{-k} \qquad (S\subset\mathbb{N}), \end{equation*}
so letting \(N\to\infty\) in the finitely additive splitting
\begin{equation*} \mu\Big(\bigcup_{j\in S}A_{n_j}\Big) = \sum_{j\in S,\ j\le N}\mu(A_{n_j}) + \mu\Big(\bigcup_{j\in S,\ j>N}A_{n_j}\Big) \end{equation*}
yields \((**)\) for \(B_j := A_{n_j}\), i.e. countable additivity of \(\mu\) on \(\sigma(\{A_{n_k}\})\).
(ii) Iterate and diagonalize. With \(N_0 := \mathbb{N}\), apply (i) to \(\mu_i\) and the disjoint family \(\{A_n\}_{n\in N_{i-1}}\) to get an infinite \(N_i\subset N_{i-1}\) with
\begin{equation*} \mu_i\Big(\bigcup_{n\in T}A_n\Big) = \sum_{n\in T}\mu_i(A_n)\qquad\text{for every }T\subset N_i , \end{equation*}
and pick \(n_1<n_2<\cdots\) with \(n_k\in N_k\). Fix \(i\) and \(S\subset\mathbb{N}\); as \(\{n_k: k\ge i\}\subset N_i\), splitting off the finitely many \(k<i\) by finite additivity and applying the last display to \(T = \{n_k: k\in S,\ k\ge i\}\) gives \(\mu_i\big(\bigcup_{k\in S}A_{n_k}\big) = \sum_{k\in S}\mu_i(A_{n_k})\), which is \((**)\) for \(\mu_i\) and \(B_k = A_{n_k}\).
Exercises 4.7.119–4.7.125
(P. Antosik and J. Mikusinski) Suppose that for all \(i,j \in \mathbb{N}\) we have numbers \(x_{ij}\) such that, for every \(j\), there exists a finite limit \(x_j = \lim\limits_{i\to\infty} x_{ij}\), and that every sequence of natural numbers \(m_j\) possesses a subsequence \(\{k_j\}\) such that the sequence \(\sum_{j=1}^{\infty} x_{ik_j}\) converges to a finite limit as \(i \to \infty\). Prove that \(x_j = \lim\limits_{i\to\infty} x_{ij}\) uniformly in \(j \in \mathbb{N}\), \(\lim\limits_{j\to\infty} x_{ij} = 0\) uniformly in \(i \in \mathbb{N}\), and \(\lim\limits_{j\to\infty} x_{jj} = 0\).
All three assertions come from the diagonal lemma (b) below. Write (I) for the hypothesis that \(x_j = \lim_ix_{ij}\) exists finitely and (II) for the subsequence hypothesis; (II) is invoked only for strictly increasing \((m_j)\) (a sequence repeating a value has a constant subsequence, for which (II) merely forces that column to vanish), and the \((k_j)\) it returns is then strictly increasing.
(a) For fixed \(i\), \(x_{ij}\to0\) as \(j\to\infty\): otherwise \(|x_{i_0m_j}|\ge\delta>0\) along a strictly increasing \((m_j)\), whereas (II) makes \(\sum_jx_{i_0k_j}\) converge along a subsequence of it.
(b) Diagonal lemma: if \((w_{nm})\) satisfies (a\(^{\prime}\)) \(\lim_nw_{nm}=0\) for every \(m\), and (b\(^{\prime}\)) every strictly increasing \((m_p)\) has a subsequence \((k_p)\) with all \(\sum_pw_{nk_p}\) convergent and \(\lim_n\sum_pw_{nk_p}\) finite, then \(w_{nn}\to0\).
Both hypotheses pass to a submatrix \(w^{\prime}_{rs} := w_{n_rn_s}\) along strictly increasing \((n_r)\): (a\(^{\prime}\)) is clear, and for \((m_p)\) one applies (b\(^{\prime}\)) to \((n_{m_p})_p\) and restricts the row index to \((n_r)\). So if \(w_{nn}\not\to0\), passing to such a submatrix lets us assume \(|w_{nn}|\ge2\varepsilon\) for all \(n\); as in (a), (b\(^{\prime}\)) gives \(w_{nm}\to0\) as \(m\to\infty\) for each fixed \(n\). Choose inductively \(n_1<n_2<\cdots\) with
\begin{equation*} |w_{n_pn_q}|<\varepsilon\,4^{-p-q}\qquad (p\ne q), \end{equation*}
possible since at stage \(q\) only the finitely many columns \(n_p\) (use (a\(^{\prime}\))) and rows \(n_p\), \(p<q\), are constrained, each tending to \(0\). Hence \(\sum_{q\ne r}|w_{n_rn_q}|\le\varepsilon4^{-r}\sum_{q\ge1}4^{-q} = \tfrac{\varepsilon}{3}4^{-r}\). Applying (b\(^{\prime}\)) to \((n_{2p})_p\) gives even indices \(q_1<q_2<\cdots\) with \(S_i := \sum_pw_{in_{q_p}}\) convergent and \(L := \lim_iS_i\) finite. For odd \(r\) all terms of \(S_{n_r}\) are off-diagonal, so \(|S_{n_r}|\le\tfrac{\varepsilon}{3}4^{-r}\to0\), forcing \(L = 0\); but
\begin{equation*} |S_{n_{q_s}}|\ \ge\ |w_{n_{q_s}n_{q_s}}|-\sum_{p\ne s}|w_{n_{q_s}n_{q_p}}| \ \ge\ 2\varepsilon-\tfrac{\varepsilon}{3} = \tfrac{5\varepsilon}{3}, \end{equation*}
so \(|L|\ge5\varepsilon/3\), a contradiction.
(c) \(x_{ij}\to x_j\) uniformly in \(j\). If not, there is \(\varepsilon>0\) with \(|x_{ij}-x_j|\ge3\varepsilon\) for arbitrarily large \(i\) and suitable \(j\). By (I) each set \(\{i: |x_{ij}-x_j|\ge3\varepsilon\}\) is finite; let \(I_j\) be its largest element (\(0\) if empty). Construct \(i_1<l_1<i_2<l_2<\cdots\) and \(j_1<j_2<\cdots\): given \(l_{n-1},j_{n-1}\), put \(N_n := 1+\max(l_{n-1},\max\{I_j: j\le j_{n-1}\})\) and pick \(i_n\ge N_n\) and \(j_n\) with \(|x_{i_nj_n}-x_{j_n}|\ge3\varepsilon\); then \(i_n>l_{n-1}\), and \(i_n>I_j\) for \(j\le j_{n-1}\) forces \(j_n>j_{n-1}\). By (I) on the finitely many columns \(j_m\), \(m\le n\), pick \(l_n>i_n\) with \(|x_{l_nj_m}-x_{j_m}|<\varepsilon\) for all \(m\le n\). Put \(w_{nm} := x_{i_nj_m}-x_{l_nj_m}\): then (a\(^{\prime}\)) holds because \(i_n,l_n\to\infty\), and (b\(^{\prime}\)) holds because (II) applied to the strictly increasing \((j_{m_p})_p\) yields \((k_p)\) with \(\sigma_i := \sum_px_{ij_{k_p}}\) convergent and \(\sigma_i\to L\), whence \(\sum_pw_{nk_p} = \sigma_{i_n}-\sigma_{l_n}\to0\). Yet
\begin{equation*} |w_{nn}|\ \ge\ |x_{i_nj_n}-x_{j_n}|-|x_{l_nj_n}-x_{j_n}|\ \ge\ 3\varepsilon-\varepsilon = 2\varepsilon , \end{equation*}
contradicting (b).
(d) Given \(\varepsilon>0\), take \(i_0\) with \(\sup_j|x_{ij}-x_j|<\varepsilon/2\) for \(i\ge i_0\) by (c), then \(J_0\) with \(|x_{i_0j}|<\varepsilon/2\) for \(j\ge J_0\) by (a); thus \(|x_j|<\varepsilon\) and \(|x_{ij}|\le|x_j|+|x_{ij}-x_j|<\tfrac32\varepsilon\) for \(i\ge i_0\), \(j\ge J_0\), while (a) applied to the finitely many \(i<i_0\) gives \(J\ge J_0\) with \(|x_{ij}|<\varepsilon\) there for \(j\ge J\). Hence \(\sup_i|x_{ij}|\to0\) as \(j\to\infty\), and in particular \(x_{jj}\to0\).
(i) Deduce Theorem 4.6.3 from Exercise 4.7.119.
(ii) Prove that Corollary 4.6.4 remains valid in the case where \(\mu_n\) is a bounded finitely additive set function on a \(\sigma\)-algebra \(\mathcal{A}\).
(i) Apply Exercise 4.7.119 to \(x_{ij} := \mu_i(A_j)\), where \(\mu_n\in\mathcal{M}(X,\mathcal{A})\) converges on every \(A\in\mathcal{A}\), \(\mu(A) := \lim_n\mu_n(A)\), and \((A_j)\) is any disjoint sequence in \(\mathcal{A}\). Hypothesis (I) holds with \(x_j = \mu(A_j)\), and (II) holds with \(k_j = m_j\), since countable additivity of \(\mu_i\) gives \(\sum_jx_{im_j} = \mu_i\big(\bigcup_jA_{m_j}\big)\), which has a finite limit in \(i\) by assumption. Hence
\begin{equation*} \lim_{j\to\infty}\ \sup_n|\mu_n(A_j)| = 0 \end{equation*}
for every disjoint sequence, which is condition (ii) of Lemma 4.6.5, whose condition (i) is uniform countable additivity of \(\{\mu_n\}\).
Countable additivity of \(\mu\) (which is finitely additive and real-valued): for disjoint \(A_j\) with union \(A\),
\begin{equation*} \Big|\mu(A)-\sum_{j\le N}\mu(A_j)\Big| = \Big|\lim_n\mu_n\Big(\bigcup_{j>N}A_j\Big)\Big| \le\sup_n\Big|\sum_{j>N}\mu_n(A_j)\Big|\xrightarrow[N\to\infty]{}0 \end{equation*}
by that uniform countable additivity; being real-valued on a \(\sigma\)-algebra, \(\mu\) is bounded (Corollary 3.1.3), so \(\mu\in\mathcal{M}(X,\mathcal{A})\).
Uniform boundedness. Call \(E\in\mathcal{A}\) large if \(\sup_n|\mu_n|(E) = \infty\), and note \(c_E := \sup_n|\mu_n(E)|<\infty\) because \(\mu_n(E)\) converges. For \(E\) large and \(M>0\), choose \(n\) with \(|\mu_n|(E)>2(M+c_E)+1\) and a finite \(\mathcal{A}\)-partition \(E = \bigcup_{i\le k}A_i\) with \(\sum_i|\mu_n(A_i)|>2(M+c_E)\); for \(B := \bigcup\{A_i: \mu_n(A_i)\ge0\}\),
\begin{equation*} \mu_n(B)+\big(-\mu_n(E\setminus B)\big) = \sum_i|\mu_n(A_i)|>2(M+c_E), \end{equation*}
both summands being nonnegative, so one set \(D_0\in\{B,E\setminus B\}\) has \(|\mu_n(D_0)|>M+c_E\), and the other, \(D_1\), has \(|\mu_n(D_1)| = |\mu_n(E)-\mu_n(D_0)|\ge|\mu_n(D_0)|-c_E>M\); thus \(|\mu_n(D_i)|>M\) for both \(i\). Since \(|\mu_k|(E) = |\mu_k|(D_0)+|\mu_k|(D_1)\), one of them, say \(C\), is large, and the other, \(D = E\setminus C\), has \(|\mu_n(D)|>M\). Starting from \(E_0 = X\) (large, if the claim fails) with \(M = k\) at stage \(k\) produces disjoint \(D_k\subset E_{k-1}\) and \(n_k\) with \(|\mu_{n_k}(D_k)|>k\), contradicting the display above. So \(\sup_n\|\mu_n\|<\infty\).
Assertions (ii), (iii) of Theorem 4.6.3: with \(\nu := \sum_n2^{-n}(1+\|\mu_n\|)^{-1}|\mu_n|\), a finite nonnegative measure with \(\mu_n\ll\nu\), the function
\begin{equation*} \alpha(t) := \sup\big\{|\mu_n(A)|:\ n\in\mathbb{N},\ A\in\mathcal{A},\ \nu(A)\le t\big\} \end{equation*}
is nondecreasing, bounded by \(\sup_n\|\mu_n\|\), and satisfies \(\sup_n|\mu_n(A)|\le\alpha(\nu(A))\), i.e. (4.6.2), while \(\alpha(t)\to0\) as \(t\to0\) by condition (iv) of Lemma 4.6.5 (available since \(\{\mu_n\}\) is uniformly countably additive with \(\mu_n\ll\nu\)). The same condition (iv) applied to any bounded \(\lambda\ge0\) with \(\mu_n\ll\lambda\) gives (iii).
(ii) Suppose \(\sup_n|\mu_n|(X) = \infty\) for bounded finitely additive \(\mu_n\) with \(\sup_n|\mu_n(A)|<\infty\) for each \(A\). The splitting argument of part (i) uses only finite additivity of \(\mu_n\) and of \(|\mu_n|\), finiteness of \(c_E\), and a finite partition of \(E\) (no Hahn decomposition), so it again yields disjoint \(D_k\in\mathcal{A}\) and \(n_k\) with \(|\mu_{n_k}(D_k)|>k\). By Exercise 4.7.118(ii) there is a subsequence \((D_{k_p})\) such that every \(\mu_n\) is countably additive on \(\mathcal{A}_0 := \sigma(\{D_{k_p}\})\); the restrictions \(\mu_n|_{\mathcal{A}_0}\) are then bounded countably additive measures with \(\sup_n|\mu_n(A)|<\infty\) on \(\mathcal{A}_0\), so Corollary 4.6.4 gives \(\sup_n\sup_{A\in\mathcal{A}_0}|\mu_n(A)|<\infty\), contradicting \(|\mu_{n_{k_p}}(D_{k_p})|>k_p\to\infty\).
Prove Proposition 4.7.39.
The map \(m\mapsto l_m\), \(l_m(f) = \int_Xf\,dm\), is a linear isometry of \(ba(X,\mathcal{A})\), normed by \(\|m\|_1 = |m|(X)\), onto \(B(X,\mathcal{A})^{*}\).
On the space \(S\) of simple functions set \(I(f) := \sum_{i\le n}c_im(A_i)\) for \(f = \sum_ic_iI_{A_i}\) with disjoint \(A_i\in\mathcal{A}\) covering \(X\). Passing to the common refinement \(\{A_i\cap B_j\}\) of two such representations, where \(c_i = d_j\) whenever \(A_i\cap B_j\ne\emptyset\), finite additivity gives
\begin{equation*} \sum_ic_im(A_i) = \sum_{i,j}c_im(A_i\cap B_j) = \sum_{i,j}d_jm(A_i\cap B_j) = \sum_jd_jm(B_j), \end{equation*}
so \(I\) is well defined, and linear on \(S\) by the same refinement device, with
\begin{equation*} |I(f)|\le\|f\|_\infty\sum_i|m(A_i)|\le\|f\|_\infty|m|(X) = \|f\|_\infty\|m\|_1 . \end{equation*}
\(S\) is dense in the Banach space \(B(X,\mathcal{A})\): for \(|f|\le M\) and \(\varepsilon>0\), split \([-M,M]\) into disjoint intervals \(J_i\) of length \(<\varepsilon\), pick \(c_i\in J_i\), and \(g = \sum_ic_iI_{f^{-1}(J_i)}\in S\) has \(\|f-g\|_\infty\le\varepsilon\). So \(I\), uniformly continuous on \(S\) by the last estimate, extends uniquely to \(l_m\in B(X,\mathcal{A})^{*}\) with \(\|l_m\|\le\|m\|_1\).
Isometry: for a finite \(\mathcal{A}\)-partition \(X = \bigcup_{i\le n}A_i\) and \(\varepsilon_i := 1\) if \(m(A_i)\ge0\), \(\varepsilon_i := -1\) otherwise, the function \(f = \sum_i\varepsilon_iI_{A_i}\) has \(\|f\|_\infty\le1\) and \(l_m(f) = \sum_i|m(A_i)|\); taking the supremum over partitions gives \(\|l_m\|\ge\|m\|_1\). In particular the (clearly linear) map \(m\mapsto l_m\) is injective.
Surjectivity: for \(l\in B(X,\mathcal{A})^{*}\) put \(m(A) := l(I_A)\), finitely additive since \(I_{A\cup B} = I_A+I_B\) for disjoint \(A,B\). With \(\varepsilon_i\) as above,
\begin{equation*} \sum_{i\le n}|m(A_i)| = l\Big(\sum_i\varepsilon_iI_{A_i}\Big) \le\|l\|\,\Big\|\sum_i\varepsilon_iI_{A_i}\Big\|_\infty\le\|l\| , \end{equation*}
so \(\|m\|_1\le\|l\|<\infty\) and \(m\in ba(X,\mathcal{A})\); and \(l = l_m\), since they agree on indicators, hence on \(S\) by linearity, hence on \(B(X,\mathcal{A})\) by density and continuity.
Prove Lemma 4.7.40.
Extract the indices by driving down the functional \(\Phi\) below. Lemma 4.7.40 asserts, for uniformly bounded \(\{m_n\}\subset ba(X,\mathcal{A})\), disjoint \(A_i\in\mathcal{A}\) and \(\varepsilon>0\), the existence of \(k_1<k_2<\cdots\) with \(|m_{k_n}|\big(\bigcup_{j\ne n}A_{k_j}\big)<\varepsilon\) for all \(n\). If \(M := \sup_n|m_n|(X) = 0\) take \(k_n = n\); otherwise replace \(m_n\) by \(m_n/M\) and \(\varepsilon\) by \(\varepsilon/M\), so that \(|m_n|\big(\bigcup_jA_j\big)\le1\) for all \(n\). For infinite \(S\subset\mathbb{N}\) the conclusion reads
\begin{equation*} |m_k|\Big(\bigcup_{j\in S\setminus\{k\}}A_j\Big)<\varepsilon\qquad\text{for all }k\in S, \tag{\(*\)} \end{equation*}
and we put \(\Phi(S) := \sup_{k\in S}|m_k|\big(\bigcup_{j\in S}A_j\big)\in[0,1]\).
Main step: an infinite \(S\) has an infinite subset \(S^{\prime}\) satisfying \((*)\) or with \(\Phi(S^{\prime})\le\Phi(S)-\varepsilon\). Partition \(S\) into disjoint infinite parts \(\Sigma_p\). If some \(\Sigma_p\) satisfies \((*)\), take \(S^{\prime} = \Sigma_p\). Otherwise choose \(k_p\in\Sigma_p\) with \(|m_{k_p}|\big(\bigcup_{j\in\Sigma_p\setminus\{k_p\}}A_j\big)\ge\varepsilon\) and set \(S^{\prime} := \{k_p: p\in\mathbb{N}\}\), infinite since the \(\Sigma_p\) are disjoint. As \(S^{\prime}\cap\Sigma_p = \{k_p\}\) and the \(A_i\) are disjoint, \(\bigcup_{j\in S^{\prime}}A_j\) and \(\bigcup_{j\in\Sigma_p\setminus\{k_p\}}A_j\) are disjoint subsets of \(\bigcup_{j\in S}A_j\), so finite additivity and monotonicity of the variation give, using \(k_p\in S\),
\begin{equation*} |m_{k_p}|\Big(\bigcup_{j\in S^{\prime}}A_j\Big)\ \le\ |m_{k_p}|\Big(\bigcup_{j\in S}A_j\Big)-\varepsilon \ \le\ \Phi(S)-\varepsilon , \end{equation*}
and the supremum over \(p\), i.e. over \(S^{\prime}\), is \(\Phi(S^{\prime})\le\Phi(S)-\varepsilon\).
Iterating from \(S_0 = \mathbb{N}\) with \(\Phi(S_0)\le1\), the second alternative can occur at most \(\lceil1/\varepsilon\rceil\) times, since \(\Phi\ge0\); hence at some stage the first alternative gives an infinite \(S\) with \((*)\), and its increasing enumeration is the required \(k_1<k_2<\cdots\).
Prove Lemma 4.7.41.
Suppose the conclusion \(\lim_n\sum_j|m_n(\{j\})| = 0\) of Lemma 4.7.41 fails: there are \(\varepsilon>0\) and \(n_1<n_2<\cdots\) with \(\sum_j|m_{n_i}(\{j\})|>5\varepsilon\); write \(\nu_i := m_{n_i}\). Each such series converges, since \(\sum_{j\in F}|m(\{j\})|\le|m|(\mathbb{N})\) for finite \(F\), and \(M := \sup_n\|m_n\|_1<\infty\) by Exercise 4.7.120(ii), whose hypothesis \(\sup_n|m_n(A)|<\infty\) holds because \(m_n(A)\to0\) for every \(A\subset\mathbb{N}\).
Disjointification: there are \(i_1<i_2<\cdots\) and pairwise disjoint finite \(E_r\subset\mathbb{N}\) with \(\sum_{j\in E_r}|\nu_{i_r}(\{j\})|>4\varepsilon\). Take \(i_1 := 1\) and finite \(E_1\) with \(\sum_{j\in E_1}|\nu_{i_1}(\{j\})|>5\varepsilon\); given \(E_1,\dots,E_r\) with finite union \(G_r\), use \(\nu_i(\{j\})\to0\) for the finitely many \(j\in G_r\) to pick \(i_{r+1}>i_r\) with \(\sum_{j\in G_r}|\nu_{i_{r+1}}(\{j\})|<\varepsilon\), then finite \(F\) with \(\sum_{j\in F}|\nu_{i_{r+1}}(\{j\})|>5\varepsilon\), and set \(E_{r+1} := F\setminus G_r\), whose sum exceeds \(5\varepsilon-\varepsilon\).
Put \(\lambda_r := \nu_{i_r}\), uniformly bounded by \(M\) with \(\lambda_r(A)\to0\) for every \(A\). By Lemma 4.7.40 (Exercise 4.7.122), applied to \(\{\lambda_r\}\), the disjoint \(E_r\) and \(\varepsilon\), there are \(r_1<r_2<\cdots\) with \(|\lambda_{r_s}|\big(\bigcup_{t\ne s}E_{r_t}\big)<\varepsilon\). Splitting the finite set \(E_{r_s}\) into \(B^{+} = \{j\in E_{r_s}: \lambda_{r_s}(\{j\})\ge0\}\) and \(B^{-}\), finite additivity gives
\begin{equation*} |\lambda_{r_s}(B^{+})|+|\lambda_{r_s}(B^{-})| = \sum_{j\in E_{r_s}}|\lambda_{r_s}(\{j\})|>4\varepsilon , \end{equation*}
so one of these two sets, \(B_s\), has \(|\lambda_{r_s}(B_s)|>2\varepsilon\). Put \(B := \bigcup_sB_s\); then \(B\setminus B_s\subset\bigcup_{t\ne s}E_{r_t}\), so \(|\lambda_{r_s}(B\setminus B_s)|\le|\lambda_{r_s}|\big(\bigcup_{t\ne s}E_{r_t}\big)<\varepsilon\) and, by finite additivity,
\begin{equation*} |\lambda_{r_s}(B)|\ \ge\ |\lambda_{r_s}(B_s)|-|\lambda_{r_s}(B\setminus B_s)| \ >\ 2\varepsilon-\varepsilon = \varepsilon \end{equation*}
for every \(s\), contradicting \(\lambda_{r_s}(B)\to0\) (here \(\lambda_{r_s} = m_{n_{i_{r_s}}}\) with \(n_{i_{r_s}}\to\infty\)).
(Kaczmarz, Nikliborc) Let \(\varphi\) be a continuous even function on the real line with the following properties (\(\alpha\)): \(\varphi(t)>0\) if \(t\ne0\) and there exist \(A\) and \(a\) such that \(\varphi(t)\ge A\) if \(|t|\ge a\). Let \(\mu\) be a probability measure on \((X,\mathcal{A})\) and let \(f_n\) be \(\mu\)-measurable functions.
(i) Suppose that \[ \int_X\varphi(f_n-f_m)\,d\mu\to0\quad\text{as }n,m\to\infty. \] Prove that there exists a \(\mu\)-measurable function \(f\) such that \[ \int_X\varphi(f-f_n)\,d\mu\to0. \]
(ii) Let \(\varphi\) satisfy the following additional condition (\(\beta\)): there is \(N\) such that \(\varphi(t+s)\le N\varphi(t)+N\varphi(s)\). Suppose that the functions \(\varphi\circ f_n\) are integrable. Show that in (i) one has \[ \int_X\varphi(f_n)\,d\mu\to\int_X\varphi(f)\,d\mu. \]
(iii) Suppose that the functions \(f_n\) converge a.e. to some function \(f\) and there exists a function \(\varphi\) with the properties (\(\alpha\)) and (\(\beta\)) and finite integrals \(\varphi\circ f_n\). Show that there exists a continuous even function \(\psi\) with the properties (\(\alpha\)) and (\(\beta\)) such that \(\lim\limits_{|t|\to\infty}\psi(t)/\varphi(t)=0\) and \[ \int_X\psi(f-f_n)\,d\mu\to0. \] In particular, since one can always take a bounded function for \(\varphi\), there exists an unbounded function \(\psi\) with the aforementioned properties.
Throughout, (\(\alpha\)) is read with \(A>0\) and \(\varphi(0) = 0\), since for \(A\le0\) the condition at infinity is vacuous and for \(\varphi(0)>0\) one gets \(\inf_{\mathbb{R}}\varphi>0\), making the hypotheses of (i) and (ii) unsatisfiable.
(i) The \(f_n\) are fundamental in measure, because \(c_\delta := \inf\{\varphi(t): |t|\ge\delta\}>0\) for every \(\delta>0\) (as \(\varphi\ge A>0\) off \([-a,a]\) and \(\varphi\) is continuous and positive on the compact \(\{\delta\le|t|\le a\}\)), so Chebyshev’s inequality gives
\begin{equation*} \mu\big(x:\ |f_n(x)-f_m(x)|\ge\delta\big)\le\frac{1}{c_\delta}\int_X\varphi(f_n-f_m)\,d\mu \xrightarrow[n,m\to\infty]{}0 . \end{equation*}
By the Riesz theorem (Theorem 2.2.5(ii)) there is a \(\mu\)-measurable \(f\) with \(f_n\to f\) in measure, and \(f_{n_k}\to f\) a.e. along a subsequence (Theorem 2.2.5(i)). Fix \(\varepsilon>0\), take \(N_\varepsilon\) with \(\int_X\varphi(f_n-f_m)\,d\mu\le\varepsilon\) for \(n,m\ge N_\varepsilon\), and let \(n\ge N_\varepsilon\); continuity of \(\varphi\) gives \(\varphi(f_n-f_{n_k})\to\varphi(f_n-f)\) a.e., so Fatou’s theorem applied to these nonnegative functions yields
\begin{equation*} \int_X\varphi(f-f_n)\,d\mu = \int_X\varphi(f_n-f)\,d\mu \le\liminf_{k\to\infty}\int_X\varphi(f_n-f_{n_k})\,d\mu\le\varepsilon \end{equation*}
(the \(n_k\) are eventually \(\ge N_\varepsilon\), and \(\varphi\) is even).
(ii) Set \(\eta_n := \int_X\varphi(f-f_n)\,d\mu\to0\) by (i), and take \(N\ge1\) in (\(\beta\)). By (\(\beta\)) and evenness, \(\varphi(f)\le N\varphi(f_n)+N\varphi(f-f_n)\), whose right side is integrable for a fixed \(n\), so \(\varphi\circ f\in L^1(\mu)\); symmetrically \(\varphi(f_n)\le N\varphi(f)+N\varphi(f_n-f)\), whence
\begin{equation*} \int_E\varphi(f_n)\,d\mu\ \le\ N\int_E\varphi(f)\,d\mu+N\eta_n\qquad (E\in\mathcal{A}). \end{equation*}
Taking \(E = X\) bounds \(\{\varphi\circ f_n\}\) in \(L^1(\mu)\); and given \(\varepsilon>0\), choosing \(n_0\) with \(N\eta_n<\varepsilon/2\) for \(n\ge n_0\) and then \(\delta>0\) so that \(\mu(E)<\delta\) forces \(N\int_E\varphi(f)\,d\mu<\varepsilon/2\) and \(\int_E\varphi(f_n)\,d\mu<\varepsilon\) for the finitely many \(n<n_0\) (absolute continuity of the integral) makes the integrals uniformly absolutely continuous, so \(\{\varphi\circ f_n\}\) is uniformly integrable by Proposition 4.5.3 (\(\mu\) is finite). Moreover \(\varphi\circ f_n\to\varphi\circ f\) in measure, since every subsequence has a further subsequence with \(f_{p_{k_r}}\to f\) a.e. (Theorem 2.2.5(i)), hence \(\varphi\circ f_{p_{k_r}}\to\varphi\circ f\) a.e. and in measure. The Lebesgue–Vitali theorem (Theorem 4.5.4) now gives \(\varphi\circ f_n\to\varphi\circ f\) in \(L^1(\mu)\), hence the convergence of the integrals.
(iii) Take \(\psi := \theta\circ\varphi\) with \(\theta(x) := \int_0^x\sigma(u)\,du\) for the weight \(\sigma\) built as follows. Put \(H_n := \varphi(f-f_n)\), so \(H_n\to0\) a.e. and each \(H_n\) is finite; as \(\mu\) is a probability measure,
\begin{equation*} \lim_{n\to\infty}\mu(H_n>T) = 0\ \ (T>0), \qquad \Theta(T) := \sup_n\mu(H_n>T)\xrightarrow[T\to\infty]{}0 , \end{equation*}
the second because the first with \(T = 1\) gives \(n_0\) with \(\mu(H_n>T)<\varepsilon\) for \(n>n_0\), \(T\ge1\), while each of the finitely many \(H_n\), \(n\le n_0\), is finite. Note \(\Theta\) is nonincreasing with \(\Theta\le1\). Now let \(\sigma \equiv s_0 := 1\) on \([0,T_1)\), where \(\Theta(T_1)\le4^{-1}\), and inductively choose \(T_{k+1}>T_k\) with
\begin{equation*} \Theta(T_{k+1})\le4^{-(k+1)}\qquad\text{and}\qquad T_{k+1}-T_k\ \ge\ \frac{2^k}{\min(s_{k-1},2^{-k})} , \end{equation*}
possible since \(\Theta(T)\to0\), putting \(s_k := 2^k/(T_{k+1}-T_k)\le\min(s_{k-1},2^{-k})\) and \(\sigma := s_k\) on \([T_k,T_{k+1})\). Then \(\sigma\) is positive, nonincreasing, \(\sigma(u)\to0\), and
\begin{equation*} \begin{aligned} \int_0^\infty\sigma(u)\,du &= s_0T_1+\sum_{k\ge1}s_k(T_{k+1}-T_k) = s_0T_1+\sum_{k\ge1}2^k = \infty,\\ \int_0^\infty\sigma(u)\Theta(u)\,du &\le T_1+\sum_{k\ge1}s_k4^{-k}(T_{k+1}-T_k) = T_1+\sum_{k\ge1}2^k4^{-k}<\infty, \end{aligned} \end{equation*}
the second line using \(\Theta\le4^{-k}\) on \([T_k,T_{k+1})\) and \(\Theta\le1\) on \([0,T_1)\). Hence \(\theta\) is continuous, strictly increasing, concave, \(\theta(0) = 0\), unbounded, with \(\theta(x)/x\to0\) (it is the mean of \(\sigma\) over \([0,x]\)), and concavity with \(\theta(0) = 0\) gives \(\theta(Nx)\le N\theta(x)\) for \(N\ge1\) and \(\theta(x+y)\le\theta(x)+\theta(y)\).
Thus \(\psi\) is continuous and even, has (\(\alpha\)) (\(\psi(t)>0\) for \(t\ne0\) and \(\psi\ge\theta(A)>0\) for \(|t|\ge a\), by monotonicity of \(\theta\)) and (\(\beta\)) with the same \(N\):
\begin{equation*} \psi(t+s)\le\theta\big(N\varphi(t)+N\varphi(s)\big) \le\theta\big(N\varphi(t)\big)+\theta\big(N\varphi(s)\big) \le N\psi(t)+N\psi(s), \end{equation*}
while \(\psi(t)/\varphi(t) = \theta(\varphi(t))/\varphi(t)\to0\) as \(|t|\to\infty\) whenever \(\varphi(t)\to\infty\), and \(\psi\) is unbounded if \(\varphi\) is. Finally, by Tonelli’s theorem and then dominated convergence (with \(\mu(H_n>u)\le\Theta(u)\), \(\sigma\Theta\) integrable, and \(\mu(H_n>u)\to0\) for each \(u>0\)),
\begin{equation*} \int_X\psi(f-f_n)\,d\mu = \int_0^\infty\sigma(u)\,\mu(H_n>u)\,du\ \longrightarrow\ 0 . \end{equation*}
The ratio requirement is consistent with (\(\alpha\)) for \(\psi\) only when \(\varphi(t)\to\infty\), and the last sentence of the exercise follows by taking \(\varphi(t) = |t|\), which has (\(\alpha\)) and (\(\beta\)) with \(N = 1\), so that \(\psi = \theta(|\cdot|)\) is unbounded with \(\psi(t)/|t|\to0\).
Let \(\varphi:[0,\infty)\to[0,\infty)\) with \(\varphi(0)=0\) be either an increasing concave function or a convex function with \(\varphi(2x)\le C\varphi(x)\). Let \((X,\mathcal{A},\mu)\) be a probability space and let measurable functions \(f_n\) converge in measure to \(f\). Suppose that \(\varphi\circ|f|,\varphi\circ|f_n|\in L^1(\mu)\) and \[ \int_X\varphi\circ|f_n|\,d\mu\to\int_X\varphi\circ|f|\,d\mu . \] Prove that \[ \int_X\varphi\circ|f_n-f|\,d\mu\to0 \] and that the functions \(\varphi\circ|f_n|\) are uniformly integrable.
The printed hypotheses must be read with \(\varphi\) continuous at \(0\), since otherwise the increasing concave \(\varphi(0) = 0\), \(\varphi(x) = 1+x\) for \(x>0\), together with \(f\equiv1\), \(f_n\equiv1+1/n\), satisfies them while \(\int\varphi\circ|f_n-f| = 1+1/n\not\to0\); on \((0,\infty)\) continuity is automatic for a finite convex or concave function.
In both cases \(\varphi\) is nondecreasing (hypothesis in the concave case; in the convex case \(\varphi(x)\le\frac{x}{y}\varphi(y)\le\varphi(y)\) for \(0\le x\le y\)) and
\begin{equation*} \varphi(x+y)\le C_1\big[\varphi(x)+\varphi(y)\big],\qquad C_1 := \max(C/2,1), \end{equation*}
because concavity with \(\varphi(0) = 0\) gives \(\varphi(x)\ge\frac{x}{x+y}\varphi(x+y)\) and \(\varphi(y)\ge\frac{y}{x+y}\varphi(x+y)\), which add up, while convexity plus doubling gives \(\varphi(x+y)\le\frac12[\varphi(2x)+\varphi(2y)]\le\frac{C}{2}[\varphi(x)+\varphi(y)]\).
Now \(g_n := \varphi\circ|f_n|\to g := \varphi\circ|f|\) in measure: \(\big||f_n|-|f|\big|\le|f_n-f|\) gives \(|f_n|\to|f|\) in measure, and every subsequence has a further one with \(|f_{p_{k_r}}|\to|f|\) a.e. (Theorem 2.2.5(i)), hence with \(g_{p_{k_r}}\to g\) a.e. and so in measure (\(\mu\) is finite). The same subsequence device upgrades this to \(L^1\): a further subsequence \(g_{m_k}\to g\) a.e. consists of nonnegative integrable functions with \(\int g_{m_k}\,d\mu\to\int g\,d\mu\), so \(\int_X|g_{m_k}-g|\,d\mu\to0\) by the Vitali–Scheffe theorem (Theorem 2.8.9); hence \(\|g_n-g\|_{L^1(\mu)}\to0\).
Uniform integrability of \(\{g_n\}\): the family is \(L^1\)-bounded since \(\|g_n\|_{L^1}\le\|g\|_{L^1}+\|g_n-g\|_{L^1}\), and for \(\varepsilon>0\), choosing \(n_0\) with \(\|g_n-g\|_{L^1}<\varepsilon/2\) for \(n\ge n_0\) and \(\delta>0\) with \(\mu(E)<\delta\) forcing \(\int_Eg\,d\mu<\varepsilon/2\) and \(\int_Eg_n\,d\mu<\varepsilon\) for the finitely many \(n<n_0\) (absolute continuity of the integral), we get \(\int_Eg_n\,d\mu\le\int_Eg\,d\mu+\|g_n-g\|_{L^1}<\varepsilon\) for all \(n\); Proposition 4.5.3 applies since \(\mu\) is finite.
Finally \(h_n := \varphi\circ|f_n-f|\) obeys
\begin{equation*} 0\le h_n\le\varphi\big(|f_n|+|f|\big)\le C_1\big[g_n+g\big] \end{equation*}
by monotonicity and the displayed subadditivity, so \(\{h_n\}\) inherits \(L^1\)-boundedness and uniform absolute continuity and is uniformly integrable (Proposition 4.5.3); and \(h_n\to0\) in measure, since continuity of \(\varphi\) at \(0\) gives \(\delta>0\) with \(\varphi(\delta)<\varepsilon\), whence \(\mu(h_n>\varepsilon)\le\mu(|f_n-f|>\delta)\to0\). The Lebesgue–Vitali theorem (Theorem 4.5.4) then yields \(\int_X\varphi\circ|f_n-f|\,d\mu\to0\).
Exercises 4.7.126–4.7.132
Let \(\mu\) be a nonnegative measure and let \(\varphi\colon [0,+\infty)\to[0,+\infty)\) be a continuous increasing convex function such that \(\varphi(0)=0\), \(\varphi(x)>0\) if \(x>0\). For any measurable function \(f\), we set
\begin{equation*} \|f\|_\varphi := \inf\Big\{\alpha>0\colon \int \varphi(|f|/\alpha)\,d\mu\le 1\Big\} \end{equation*}
and denote by \(\mathcal{L}^\varphi(\mu)\) the set of all \(f\) with \(\|f\|_\varphi<\infty\). Show that: (i) \(\mathcal{L}^\varphi\) is closed under sums and multiplication by scalars and the corresponding linear space \(L^\varphi(\mu)\) of the equivalence classes is complete with respect to the norm \(\|\cdot\|_\varphi\) (the Orlicz space); (ii) if \(f\) and \(g\) are equimeasurable, then \(\|f\|_\varphi=\|g\|_\varphi\).
Everything follows from three properties of \(\Phi(f,\alpha) := \int_X\varphi(|f|/\alpha)\,d\mu\), which is nonincreasing in \(\alpha>0\), so that \(\|f\|_\varphi\) is the left endpoint of \(\{\alpha>0: \Phi(f,\alpha)\le1\}\) (with \(\inf\emptyset := +\infty\)). Convexity with \(\varphi(0) = 0\) gives \(\varphi(s)\le\frac st\varphi(t)\) for \(0<s<t\), so \(t\mapsto\varphi(t)/t\) is nondecreasing, \(\varphi\) is strictly increasing, and
\begin{equation*} \varphi(t)\ge\varphi(1)\,t\ \ (t\ge1),\qquad\text{hence}\ \varphi(t)\to+\infty\ \text{as}\ t\to+\infty; \tag{A} \end{equation*}
being also continuous, \(\varphi\) is a homeomorphism of \([0,+\infty)\) onto itself, with inverse \(\varphi^{-1}\).
(1) If \(0<\|f\|_\varphi<\infty\), then \(\Phi(f,\|f\|_\varphi)\le1\): for \(\alpha_j\downarrow\alpha_0 := \|f\|_\varphi\) with \(\Phi(f,\alpha_j)\le1\) one has \(\varphi(|f|/\alpha_j)\uparrow\varphi(|f|/\alpha_0)\), so monotone convergence (Theorem 2.8.2, Corollary 2.8.6) gives \(\Phi(f,\alpha_0)\le1\). Hence \(\Phi(f,\alpha)\le1\) whenever \(\alpha\ge\|f\|_\varphi\), \(\alpha>0\).
(2) \(\|f\|_\varphi = 0\) iff \(f = 0\) a.e.: if \(\mu(|f|>c) =: \delta>0\) for some \(c>0\), then \(1\ge\Phi(f,\alpha)\ge\varphi(c/\alpha)\delta\) for all \(\alpha>0\), contradicting \(\varphi(c/\alpha)\to\infty\) as \(\alpha\to0\) by (A).
(3) If \(\|f\|_\varphi\le\alpha\), then \(\mu(|f|\ge\varepsilon)\le1/\varphi(\varepsilon/\alpha)\), since \(\Phi(f,\alpha)\le1\) by (1) while \(\varphi(|f|/\alpha)\ge\varphi(\varepsilon/\alpha)\) on \(\{|f|\ge\varepsilon\}\).
(i) Homogeneity is the substitution \(\alpha = |c|\beta\) in \(\Phi(cf,\alpha) = \Phi(f,\beta)\). For the triangle inequality take \(\alpha>\|f\|_\varphi\), \(\beta>\|g\|_\varphi\) and \(\lambda := \alpha/(\alpha+\beta)\in(0,1)\); since \(\frac{|f+g|}{\alpha+\beta}\le\lambda\frac{|f|}{\alpha}+(1-\lambda)\frac{|g|}{\beta}\) pointwise and \(\varphi\) is increasing and convex,
\begin{equation*} \varphi\Big(\frac{|f+g|}{\alpha+\beta}\Big) \le \lambda\,\varphi\Big(\frac{|f|}{\alpha}\Big)+(1-\lambda)\,\varphi\Big(\frac{|g|}{\beta}\Big), \end{equation*}
whose integral is \(\Phi(f+g,\alpha+\beta)\le\lambda+(1-\lambda) = 1\) by (1); so \(\|f+g\|_\varphi\le\alpha+\beta\), and \(\alpha\downarrow\|f\|_\varphi\), \(\beta\downarrow\|g\|_\varphi\) give subadditivity. With (2), \(\|\cdot\|_\varphi\) is a norm on \(L^\varphi(\mu)\).
Completeness: for a Cauchy \(\{f_n\}\) pick \(n_1<n_2<\cdots\) with \(\|f_{n_{k+1}}-f_{n_k}\|_\varphi\le2^{-k}\); by (3) with \(\alpha = 2^{-k}\), \(\varepsilon = 2^{-k/2}\), and by (A),
\begin{equation*} \mu(A_k)\le\frac1{\varphi(2^{k/2})}\le\frac{1}{\varphi(1)\,2^{k/2}}, \qquad A_k := \{|f_{n_{k+1}}-f_{n_k}|\ge2^{-k/2}\}, \end{equation*}
so \(\sum_k\mu(A_k)<\infty\) and \(Z := \bigcap_m\bigcup_{k\ge m}A_k\) is \(\mu\)-null; off \(Z\) the reals \(f_{n_k}(x)\) form a Cauchy sequence, with measurable limit \(f\) (set \(f := 0\) on \(Z\)). Given \(\varepsilon>0\), take \(N\) with \(\|f_n-f_m\|_\varphi\le\varepsilon\) for \(n,m\ge N\), so \(\Phi(f_n-f_m,\varepsilon)\le1\) by (1); letting \(m = n_k\to\infty\), continuity of \(\varphi\) and Fatou’s theorem (Theorem 2.8.3, Corollary 2.8.6) give
\begin{equation*} \Phi(f_n-f,\varepsilon)\le\liminf_{k\to\infty}\Phi(f_n-f_{n_k},\varepsilon)\le1 , \end{equation*}
i.e. \(\|f_n-f\|_\varphi\le\varepsilon\) for \(n\ge N\); in particular \(f = f_N-(f_N-f)\in\mathcal{L}^\varphi\).
(ii) \(\Phi(f,\alpha)\) depends only on the distribution function \(s\mapsto\mu(|f|>s)\). Indeed, with \(\psi := \varphi(\cdot/\alpha)\), \(s_{n,j} := \alpha\varphi^{-1}(j2^{-n})\) and \(\psi_n := \sum_{j\le n2^n}2^{-n}\mathbf{1}_{(s_{n,j},+\infty)}\), the equivalence \(\psi(u)>j2^{-n}\iff u>s_{n,j}\) turns \(\psi_n(u)\) into \(\min\big(n,2^{-n}(\lceil2^n\psi(u)\rceil-1)\big)\), which increases in \(n\) to \(\psi(u)\) (Check!), so monotone convergence gives
\begin{equation*} \Phi(f,\alpha) = \lim_{n\to\infty}\sum_{j=1}^{n2^n}2^{-n}\,\mu(\{|f|>s_{n,j}\}). \end{equation*}
Equimeasurable \(f\) and \(g\) thus have \(\Phi(f,\alpha) = \Phi(g,\alpha)\) for every \(\alpha>0\), hence \(\|f\|_\varphi = \|g\|_\varphi\).
Let \(\mu\) be a finite nonnegative measure. For every measurable function \(f\), we set
\begin{equation*} f^*(t)=\inf\big\{s\ge0\colon \mu(\{x\colon |f(x)|>s\})\le t\big\} \end{equation*}
and for all \(p,q\in[1,\infty)\) we define the Lorentz space \(L^{p,q}(\mu)\) as the set of all equivalence classes of measurable functions \(f\) such that
\begin{equation*} \int_0^\infty t^{1/p-1}\,[f^*(t)]^q\,dt<\infty . \end{equation*}
Show that \(L^{p,p}(\mu)=L^p(\mu)\). On Lorentz classes, see Stein, Weiss [908], Nielsen [714], Zaanen [1043].
The printed exponent \(1/p-1\) is a misprint for \(q/p-1\), the standard normalization \(\int_0^\infty[t^{1/p}f^*(t)]^q\,dt/t\) (for \(\mu\) Lebesgue on \((0,1)\) and \(f(x) = x^{-1/3}\), \(p = q = 2\), the printed integral diverges although \(f\in L^2\)), so for \(q = p\) the condition reads \(\int_0^\infty[f^*(t)]^p\,dt<\infty\) and the assertion to prove is
\begin{equation*} \int_0^\infty [f^*(t)]^p\,dt = \int_X|f|^p\,d\mu\qquad (1\le p<\infty). \end{equation*}
The functions \(f^*\) on \(([0,\infty),\lambda)\) and \(f\) on \((X,\mu)\) are equimeasurable (Exercise 3.10.102(iv)). Indeed, \(F(s) := \mu(\{|f|>s\})\) is nonincreasing and right-continuous, since \(\mu\) is finite and \(\{|f|>s\} = \bigcup_{s^{\prime}>s}\{|f|>s^{\prime}\}\) is an increasing union, and
\begin{equation*} f^*(t)>s\iff t<F(s): \end{equation*}
if \(t\ge F(s)\) then \(s\in\{s^{\prime}: F(s^{\prime})\le t\}\), so \(f^*(t)\le s\); and if \(t<F(s)\), right-continuity gives \(s^{\prime\prime}>s\) with \(F(s^{\prime\prime})>t\), so \(\{s^{\prime}: F(s^{\prime})\le t\}\subset(s^{\prime\prime},+\infty)\) and \(f^*(t)\ge s^{\prime\prime}>s\). Hence \(\{t\ge0: f^*(t)>s\} = [0,F(s))\), i.e. \(\lambda(\{f^*>s\}) = F(s)\).
Tonelli’s theorem applied to \(\{(y,s): 0<s<h(y)\}\) gives the layer-cake formula \(\int_Yh^p\,d\nu = \int_0^\infty p\,s^{p-1}\nu(\{h>s\})\,ds\) for \(\nu\) \(\sigma\)-finite, which covers both the finite \(\mu\) and \(\lambda\); applying it to \(h = |f|\) and to \(h = f^*\),
\begin{equation*} \begin{aligned} \int_X|f|^p\,d\mu &= \int_0^\infty p\,s^{p-1}\mu(\{|f|>s\})\,ds\\ &= \int_0^\infty p\,s^{p-1}\lambda(\{f^*>s\})\,ds = \int_0^\infty[f^*(t)]^p\,dt . \end{aligned} \end{equation*}
One side is finite exactly when the other is, and \(f = g\) a.e. gives the same \(F\), hence \(f^* = g^*\); so \(L^{p,p}(\mu) = L^p(\mu)\) with equality of norms.
(\(\circ\)) (Tagamlickii [930]) Let \(\mu\) be a probability measure and let a sequence of \(\mu\)-integrable functions \(f_n\) converge in measure to a function \(f\). Prove the equivalence of the following conditions: (i) \(f\in L^1(\mu)\) and \(f_n\to f\) in \(L^1(\mu)\); (ii) for every subsequence \(\{f_{n_k}\}\), there exists a function \(\varphi\in L^1(\mu)\) such that, for infinitely many values \(k\), one has \(|f_{n_k}(x)|\le\varphi(x)\) a.e.
(i) \(\Rightarrow\) (ii). Given a subsequence \(\{f_{n_k}\}\), choose \(k_1<k_2<\cdots\) with \(\|f_{n_{k_j}}-f\|_{L^1(\mu)}\le2^{-j}\) and put
\begin{equation*} \varphi:=|f|+\sum_{j=1}^\infty |f_{n_{k_j}}-f| . \end{equation*}
The partial sums increase with integrals at most \(\|f\|_{L^1}+1\), so \(\varphi\in L^1(\mu)\) by monotone convergence (Theorem 2.8.2), and \(|f_{n_{k_j}}|\le|f|+|f_{n_{k_j}}-f|\le\varphi\) a.e. for the infinitely many \(k = k_j\).
(ii) \(\Rightarrow\) (i). Applied to the whole sequence, (ii) yields \(\varphi\in L^1(\mu)\) and an infinite \(K\) with \(|f_n|\le\varphi\) a.e. for \(n\in K\); since \(\{f_n\}_{n\in K}\) still converges to \(f\) in measure, some subsequence of it converges a.e. (Theorem 2.2.5(i)), so \(|f|\le\varphi\) a.e. and \(f\in L^1(\mu)\). If \(f_n\not\to f\) in \(L^1(\mu)\), take \(\varepsilon>0\) and a subsequence with \(\|f_{n_k}-f\|_{L^1(\mu)}\ge\varepsilon\) for all \(k\), and apply (ii) to it: there are \(\varphi\in L^1(\mu)\) and infinitely many \(k\), enumerated so that \(g_j := f_{n_{k_j}}\), with \(|g_j|\le\varphi\) a.e. Since \(g_j\to f\) in measure, a further subsequence has \(g_{j_i}\to f\) a.e. (Theorem 2.2.5(i)), whence \(|g_{j_i}-f|\le2\varphi\in L^1(\mu)\) and dominated convergence (Theorem 2.8.1) gives \(\|g_{j_i}-f\|_{L^1(\mu)}\to0\), a contradiction.
(Fréchet [316], Veress [974]) Let \(\mu\) be a probability measure on a space \(X\) and let \(M\) be some set of \(\mu\)-measurable functions. Prove the equivalence of the following conditions: (i) the set \(M\) has compact closure in the metric of convergence in measure (Exercise 4.7.60); (ii) every sequence in \(M\) contains an a.e. convergent subsequence; (iii) for every \(\varepsilon>0\) and \(\alpha>0\), there exists a finite collection of measurable functions \(\psi_1,\dots,\psi_n\) such that, for every function \(f\in M\), one can find an index \(i\le n\) with \(\mu(\{x\colon |f(x)-\psi_i(x)|\ge\varepsilon\})<\alpha\); (iv) for every \(\varepsilon>0\), there exist a number \(C>0\) and a finite partition of the space into disjoint measurable parts \(E_1,\dots,E_n\) such that, for every function \(f\in M\), there exists a measurable set \(E_f\) with the following properties:
\begin{equation*} \mu(E_f)<\varepsilon,\qquad \sup_{x\in X\setminus E_f}|f(x)|<C,\qquad \sup_{x,y\in E_i\setminus E_f}|f(x)-f(y)|<\varepsilon \end{equation*}
for all \(f\in M\) and \(i=1,\dots,n\).
Work in \(L^0(\mu)\), the a.e. finite measurable functions with the complete metric \(d(f,g) := \int_X\min(|f-g|,1)\,d\mu\) of Exercise 4.7.60, whose convergence is convergence in measure, and use for \(\eta>0\)
\begin{equation*} \begin{aligned} \mu(\{|f-g|\ge\eta\})&\le d(f,g)/\min(\eta,1),\\ d(f,g)&\le\eta+\mu(\{|f-g|\ge\eta\}), \end{aligned}\tag{E} \end{equation*}
the first since \(\min(|f-g|,1)\ge\min(\eta,1)\) on \(\{|f-g|\ge\eta\}\), the second since \(\mu(X) = 1\). In (iii) the \(\psi_i\) may be taken everywhere finite: with \(N_i := \{|\psi_i| = \infty\}\), an \(f\in M\) with \(\mu(\{|f-\psi_i|\ge\varepsilon\})<\alpha\) forces \(\mu(N_i)<\alpha\) (there \(|f-\psi_i| = \infty\) a.e.), so \(\psi_i\mathbf{1}_{X\setminus N_i}\) turns (iii) for \((\varepsilon,\alpha/2)\) into (iii) for \((\varepsilon,\alpha)\).
(i) \(\Leftrightarrow\) (iii). Since \(L^0(\mu)\) is complete, (i) is total boundedness of \(M\), centres allowed outside \(M\) (a \(\delta\)-ball meeting \(M\) sits in the \(2\delta\)-ball about a point of \(M\)). Given (i) and \(\varepsilon,\alpha>0\), a finite \(\delta\)-net \(\psi_1,\dots,\psi_n\) with \(\delta := \alpha\min(\varepsilon,1)\) gives, by the first line of (E), \(\mu(\{|f-\psi_i|\ge\varepsilon\})\le d(f,\psi_i)/\min(\varepsilon,1)<\alpha\). Conversely, for \(\delta>0\), (iii) with \(\varepsilon = \alpha = \delta/2\) and the second line of (E) give \(d(f,\psi_i)\le\delta/2+\delta/2 = \delta\), so \(M\) is totally bounded.
(i) \(\Rightarrow\) (ii): \(\overline M\) is compact metric, hence sequentially compact, so a subsequence converges in measure and a further one a.e. (Theorem 2.2.5(i)).
(ii) \(\Rightarrow\) (i): an a.e. convergent subsequence converges in measure (Theorem 2.2.3, \(\mu\) finite), so every sequence in \(M\) has a \(d\)-convergent subsequence; for \(g_n\in\overline M\), choosing \(f_n\in M\) with \(d(f_n,g_n)<1/n\) gives \(g_{n_k}\to f\in\overline M\), so \(\overline M\) is sequentially compact, hence compact.
(iii) \(\Rightarrow\) (iv). Apply (iii) with \((\varepsilon/8,\varepsilon/8)\) to get finite-valued \(g_1,\dots,g_N\) such that each \(f\in M\) has a \(k\le N\) with \(\mu(\{|f-g_k|\ge\varepsilon/8\})<\varepsilon/8\). Each \(g_k\) being finite, choose \(C_0\) with \(\mu(\{|g_k|\ge C_0\})<\varepsilon/(8N)\) for all \(k\), so \(B := \bigcup_{k\le N}\{|g_k|\ge C_0\}\) has \(\mu(B)<\varepsilon/8\); split \([-C_0,C_0]\) into disjoint intervals \(I_1,\dots,I_m\) of length \(<\varepsilon/8\) with \(t_j\in I_j\) and set
\begin{equation*} s_k:=\sum_{j=1}^m t_j\,\mathbf{1}_{g_k^{-1}(I_j)\setminus B},\qquad k\le N , \end{equation*}
so \(|g_k-s_k|<\varepsilon/8\) off \(B\). Let \(E_1,\dots,E_n\) be the atoms of the finite algebra generated by \(B\) and the \(g_k^{-1}(I_j)\) (each \(s_k\) is constant on each \(E_i\), and \(B\) is a union of some \(E_i\)), and \(C := C_0+\varepsilon/8\); both depend only on \(\varepsilon\). For \(f\in M\) with its \(k\), the set \(E_f := B\cup\{|f-g_k|\ge\varepsilon/8\}\) has \(\mu(E_f)<\varepsilon\) and \(|f|\le|g_k|+\varepsilon/8<C\) off \(E_f\), while for \(x,y\in E_i\setminus E_f\), where \(s_k(x) = s_k(y)\),
\begin{equation*} \begin{aligned} |f(x)-f(y)|\le\ &|f(x)-g_k(x)|+|g_k(x)-s_k(x)|\\ &+|s_k(y)-g_k(y)|+|g_k(y)-f(y)|<4\cdot\tfrac\varepsilon8<\varepsilon . \end{aligned} \end{equation*}
(iv) \(\Rightarrow\) (iii). Given \(\varepsilon,\alpha>0\), apply (iv) with \(\varepsilon^{\prime} := \min(\varepsilon/2,\alpha)\), obtaining \(C\) and \(E_1,\dots,E_n\), and let \(c_1,\dots,c_m\) be an \(\varepsilon^{\prime}\)-net in \([-C,C]\); the finite family
\begin{equation*} \psi_{(j_1,\dots,j_n)}:=\sum_{i=1}^n c_{j_i}\mathbf{1}_{E_i},\qquad (j_1,\dots,j_n)\in\{1,\dots,m\}^n , \end{equation*}
works. Given \(f\in M\), for each \(i\) with \(E_i\setminus E_f\ne\emptyset\) pick \(x_i\) there and \(j_i\) with \(|c_{j_i}-f(x_i)|<\varepsilon^{\prime}\) (possible as \(|f(x_i)|<C\)), and \(j_i := 1\) otherwise; then \(|f(x)-c_{j_i}|\le|f(x)-f(x_i)|+\varepsilon^{\prime}<2\varepsilon^{\prime}\le\varepsilon\) for \(x\in E_i\setminus E_f\), and since \(X\setminus E_f = \bigcup_i(E_i\setminus E_f)\),
\begin{equation*} \mu(\{|f-\psi_{(j_1,\dots,j_n)}|\ge\varepsilon\})\le\mu(E_f)<\varepsilon^{\prime}\le\alpha . \end{equation*}
Let \(\mathcal{A}\) be a \(\sigma\)-algebra of subsets of a space \(X\). Prove that a set \(M\) in the space of all bounded measures on \(\mathcal{A}\) has compact closure in the topology of setwise convergence precisely when for every uniformly bounded sequence of \(\mathcal{A}\)-measurable functions \(f_n\) converging pointwise to \(0\), one has the equality
\begin{equation*} \lim_{n\to\infty}\int_X f_n\,d\mu=0 \end{equation*}
uniformly in \(\mu\in M\).
Necessity. Compactness of \(\overline M\) is condition (iv) of Theorem 4.7.25, so by that theorem \(M\) is bounded in variation, \(S := \sup_{\mu\in M}\|\mu\|<\infty\), and uniformly \(\nu\)-continuous for a probability measure \(\nu\) (condition (ii); \(M\subset\{0\}\) is trivial): for every \(\eta>0\) there is \(\beta>0\) with \(|\mu(A)|\le\eta\) for all \(\mu\in M\) whenever \(\nu(A)\le\beta\). The variations inherit this, since \(\nu(A\cap X^{\pm}_\mu)\le\beta\) for a Hahn decomposition \(X = X^+_\mu\cup X^-_\mu\) and
\begin{equation*} |\mu|(A) = \mu(A\cap X^+_\mu)-\mu(A\cap X^-_\mu)\le2\eta . \end{equation*}
Let now \(|f_n|\le C\) with \(f_n\to0\) pointwise, and let \(\varepsilon>0\). Choose \(\beta\) with \(\nu(A)\le\beta\Rightarrow|\mu|(A)\le\varepsilon/(2C+2)\) for all \(\mu\in M\), then by Egorov’s theorem (Theorem 2.2.1, \(\nu\) being finite) a set \(A_\beta\) with \(\nu(X\setminus A_\beta)<\beta\) on which \(f_n\to0\) uniformly, and finally \(N\) with \(\sup_{A_\beta}|f_n|\le\varepsilon/(2S+2)\) for \(n\ge N\); then
\begin{equation*} \Big|\int_X f_n\,d\mu\Big|\le\frac{\varepsilon}{2S+2}\,\|\mu\|+C\,|\mu|(X\setminus A_\beta)\le\varepsilon \end{equation*}
for all \(n\ge N\) and \(\mu\in M\), with \(N\) independent of \(\mu\).
Sufficiency: we verify condition (iii) of Theorem 4.7.25. Applying the hypothesis to \(f_n := n^{-1}\mathbf{1}_A\) (uniformly bounded, tending to \(0\) pointwise) gives \(\sup_{\mu\in M}|\mu(A)|<\infty\) for every \(A\in\mathcal{A}\); so if \(\sup_{\mu\in M}\|\mu\| = \infty\), measures \(\mu_n\in M\) with \(\|\mu_n\|\ge n\) would violate Corollary 4.6.4, whose hypothesis \(\sup_n|\mu_n(A)|<\infty\) for all \(A\) is exactly what we just proved. And for pairwise disjoint \(A_i\) the functions \(\mathbf{1}_{A_i}\) are uniformly bounded and tend to \(0\) pointwise, so
\begin{equation*} \sup_{\mu\in M}|\mu(A_i)| = \sup_{\mu\in M}\Big|\int_X \mathbf{1}_{A_i}\,d\mu\Big|\longrightarrow0 , \end{equation*}
which is condition (ii) of Lemma 4.6.5, giving uniform countable additivity.
(Areshkin [29]) Suppose that bounded countably additive signed measures \(\mu_n\) on a \(\sigma\)-algebra \(\mathcal{A}\) in a space \(X\) converge to a measure \(\mu\) on every set in \(\mathcal{A}\). Let \(X=X^+\cup X^-\), \(X=X_n^+\cup X_n^-\) be the Hahn decompositions for \(\mu\) and \(\mu_n\). Prove that the measures \(|\mu_n|\) converge to \(|\mu|\) on every set in \(\mathcal{A}\) precisely when
\begin{equation*} \lim_{n\to\infty}\mu_n(X^+\cap X_n^-)=\lim_{n\to\infty}\mu_n(X^-\cap X_n^+)=0 . \end{equation*}
Write \(P := X^+\), \(N := X^-\), \(P_n := X_n^+\), \(N_n := X_n^-\) and
\begin{equation*} a_n:=\mu_n(P\cap N_n)\le0,\qquad b_n:=\mu_n(N\cap P_n)\ge0 , \end{equation*}
the stated condition being \(a_n\to0\) and \(b_n\to0\). Throughout, a Hahn decomposition gives \(|\mu|(A) = \mu(A\cap P)-\mu(A\cap N)\) and \(|\mu_n|(A) = \mu_n(A\cap P_n)-\mu_n(A\cap N_n)\).
Necessity. From \(\mu_n(P) = \mu_n(P\cap P_n)+a_n\) we get \(|\mu_n|(P) = \mu_n(P)-2a_n\), while \(|\mu|(P) = \mu(P)\), so
\begin{equation*} 2a_n = \mu_n(P)-|\mu_n|(P)\longrightarrow \mu(P)-|\mu|(P) = 0 ; \end{equation*}
symmetrically \(|\mu_n|(N) = 2b_n-\mu_n(N)\) and \(|\mu|(N) = -\mu(N)\) give \(2b_n = |\mu_n|(N)+\mu_n(N)\to0\).
Sufficiency. Fix \(A\in\mathcal{A}\) and split \(A\cap P_n\), \(A\cap N_n\) along \(X = P\cup N\):
\begin{equation*} \begin{aligned} |\mu_n|(A) = \ &\mu_n(A\cap P\cap P_n)+\mu_n(A\cap N\cap P_n)\\ &-\mu_n(A\cap P\cap N_n)-\mu_n(A\cap N\cap N_n). \end{aligned} \end{equation*}
As \(\mu_n\ge0\) on subsets of \(P_n\) and \(\mu_n\le0\) on subsets of \(N_n\), monotonicity bounds the two middle terms by \(0\le\mu_n(A\cap N\cap P_n)\le b_n\) and \(a_n\le\mu_n(A\cap P\cap N_n)\le0\), so
\begin{equation*} \begin{aligned} |\mu_n|(A)&=\mu_n(A\cap P\cap P_n)-\mu_n(A\cap N\cap N_n)+r_n^{\prime}\\ &=\big[\mu_n(A\cap P)-\mu_n(A\cap P\cap N_n)\big]\\ &\qquad-\big[\mu_n(A\cap N)-\mu_n(A\cap N\cap P_n)\big]+r_n^{\prime}\\ &=\mu_n(A\cap P)-\mu_n(A\cap N)+r_n , \end{aligned} \end{equation*}
with \(|r_n^{\prime}|\le|a_n|+b_n\) and \(|r_n|\le2(|a_n|+b_n)\to0\). Setwise convergence on the fixed sets \(A\cap P\) and \(A\cap N\) then gives \(|\mu_n|(A)\to\mu(A\cap P)-\mu(A\cap N) = |\mu|(A)\).
(Areshkin [31]) Suppose that bounded nonnegative countably additive measures \(\mu_n\) on a \(\sigma\)-algebra \(\mathcal{A}\) in a space \(X\) converge to a measure \(\mu\) on every set in \(\mathcal{A}\) and that we are given \(\mathcal{A}\)-measurable functions \(f_n\) and \(f\). (i) Suppose that the functions \(f_n\) converge to \(f\) \(\mu\)-a.e. Prove that, for every \(\delta>0\), one has \(\lim\limits_{n\to\infty}\mu_n(\{x\colon |f(x)-f_n(x)|\ge\delta\})=0\). (ii) Suppose that for every \(\delta>0\) one has \(\lim\limits_{n\to\infty}\mu_n(\{x\colon |f(x)-f_n(x)|\ge\delta\})=0\) and that the functions \(f_n\) are uniformly bounded. Prove that
\begin{equation*} \lim_{n\to\infty}\int_X f_n\,d\mu_n=\int_X f\,d\mu. \tag{4.7.21} \end{equation*}
(iii) Suppose that \(f_n(x)\to f(x)\) \(\mu\)-a.e., \(f_n\in L^1(\mu_n)\) and that, for every \(\varepsilon>0\), there exists \(\delta>0\) such that
\begin{equation*} \int_E f_n\,d\mu_n<\varepsilon\quad\text{whenever } E\in\mathcal{A}\text{ and }\mu_n(E)<\delta . \end{equation*}
Prove that \(f\in L^1(\mu)\) and (4.7.21) holds. (iv) Deduce from (ii) that if the functions \(f_n\) are nonnegative and converge \(\mu\)-a.e. to \(f\), then
\begin{equation*} \int_X f\,d\mu\le\liminf_{n\to\infty}\int_X f_n\,d\mu_n . \end{equation*}
(v) Suppose that the functions \(f_n\) converge \(\mu\)-a.e. to \(f\) and that there exist \(\mathcal{A}\)-measurable functions \(g_n\) convergent \(\mu\)-a.e. to a function \(g\) such that \(|f_n|\le g_n\), \(g_n\in L^1(\mu_n)\), \(g\in L^1(\mu)\), and
\begin{equation*} \int_X g\,d\mu=\lim_{n\to\infty}\int_X g_n\,d\mu_n . \end{equation*}
Deduce from (iv) that (4.7.21) holds.
Throughout, \(\mu\) is bounded, nonnegative and countably additive and \(S := \sup_n\mu_n(X)<\infty\) (Theorem 4.6.3, or directly from \(\mu_n(X)\to\mu(X)\)), setwise convergence upgrades to
\begin{equation*} \int_X h\,d\mu_n\longrightarrow\int_X h\,d\mu\quad\text{for bounded }\mathcal{A}\text{-measurable } h \tag{I} \end{equation*}
(true for simple \(h\) by hypothesis, and a uniform \(\varepsilon\)-approximation of \(h\) by a simple \(s\) costs at most \(\varepsilon(S+\mu(X))\)), and \(T_k(t) := \max(-k,\min(t,k))\) is \(1\)-Lipschitz.
(i) With \(A_n := \{|f-f_n|\ge\delta\}\) and \(E_m := \bigcup_{n\ge m}A_n\), every point of \(\bigcap_mE_m\) lies in infinitely many \(A_n\), hence is a point where \(f_n\not\to f\); so \(\mu(\bigcap_mE_m) = 0\) and \(\mu(E_m)\downarrow0\) (\(\mu\) is finite). Since \(A_n\subset E_m\) for \(n\ge m\) and \(E_m\) is a fixed set,
\begin{equation*} \limsup_{n\to\infty}\mu_n(A_n)\le\lim_{n\to\infty}\mu_n(E_m) = \mu(E_m)\xrightarrow[m\to\infty]{}0 . \end{equation*}
(ii) Let \(|f_n|\le C\). Then \(|f|\le C\) \(\mu\)-a.e.: \(B_\delta := \{|f|>C+\delta\}\subset A_n(\delta) := \{|f-f_n|\ge\delta\}\), so \(\mu(B_\delta) = \lim_n\mu_n(B_\delta)\le\lim_n\mu_n(A_n(\delta)) = 0\) by (i), and \(\{|f|>C\} = \bigcup_mB_{1/m}\). Replacing \(f\) by \(T_C(f)\) alters neither \(\int f\,d\mu\) nor the hypothesis, since \(|T_C(f)-f_n| = |T_C(f)-T_C(f_n)|\le|f-f_n|\); so assume \(|f|\le C\) everywhere. With \(\delta := \varepsilon/(3S+3)\),
\begin{equation*} \Big|\int_X(f_n-f)\,d\mu_n\Big|\le 2C\,\mu_n(A_n(\delta))+\delta\,S<\tfrac23\varepsilon \end{equation*}
for large \(n\), while \(\big|\int f\,d\mu_n-\int f\,d\mu\big|<\varepsilon/3\) for large \(n\) by (I).
(iii) The hypothesis must be read with absolute values,
\begin{equation*} \Big|\int_E f_n\,d\mu_n\Big|<\varepsilon \quad\text{for all } n\text{ and } E\in\mathcal{A}\text{ with }\mu_n(E)<\delta, \tag{K} \end{equation*}
since as printed it fails for \(\mu_n = \mu = \) Lebesgue measure on \([0,1]\) and \(f_n = -n\mathbf{1}_{[0,1/n]}\to0\) a.e., where \(\int_Ef_n\,d\mu\le0\) yet \(\int f_n\,d\mu = -1\). Applying (K) to \(E\cap\{f_n\ge0\}\) and \(E\cap\{f_n<0\}\) gives \(\int_E|f_n|\,d\mu_n<2\varepsilon\) whenever \(\mu_n(E)<\delta\).
First, \(A^{*} := \sup_n\int_X|f_n|\,d\mu_n<\infty\). Let \(\delta_1\) correspond to \(\varepsilon = 1\), so \(\mu_n(E)<\delta_1\) forces \(\int_E|f_n|\,d\mu_n<2\). As \(f\) is finite \(\mu\)-a.e., choose \(k_0\) with \(\mu(\{|f|>k_0\})<\delta_1/3\), then \(N_1\) with \(\mu_n(\{|f|>k_0\})<\delta_1/3\) for \(n\ge N_1\) (a fixed set), and by (i) \(N_2\) with \(\mu_n(\{|f-f_n|\ge1\})<\delta_1/3\) for \(n\ge N_2\); since \(\{|f_n|>k_0+1\}\subset\{|f|>k_0\}\cup\{|f-f_n|\ge1\}\) has \(\mu_n\)-measure \(<\delta_1\) there,
\begin{equation*} \int_X|f_n|\,d\mu_n\le2+(k_0+1)S\qquad\big(n\ge\max(N_1,N_2)\big), \end{equation*}
the finitely many earlier integrals being finite because \(f_n\in L^1(\mu_n)\).
Next, \(f\in L^1(\mu)\): for fixed \(k\), \(|f_n|\to|f|\) \(\mu\)-a.e., so (i) and \(1\)-Lipschitzness give \(\mu_n(\{|T_k(|f|)-T_k(|f_n|)|\ge\delta\})\to0\), and (ii) applied to the functions \(T_k(|f_n|)\le k\) yields \(\int_XT_k(|f|)\,d\mu = \lim_n\int_XT_k(|f_n|)\,d\mu_n\le A^{*}\); as \(T_k(|f|)\uparrow|f|\), monotone convergence (Theorem 2.8.2) gives \(\int_X|f|\,d\mu\le A^{*}\).
For (4.7.21), let \(\delta\) correspond to \(\varepsilon/2\) in (K), so \(\mu_n(E)<\delta\) forces \(\int_E|f_n|\,d\mu_n<\varepsilon\), and take \(k>A^{*}/\delta\) with also \(\int_{\{|f|>k\}}|f|\,d\mu<\varepsilon\) (possible as \(f\in L^1(\mu)\)). Chebyshev’s inequality gives \(\mu_n(\{|f_n|>k\})\le A^{*}/k<\delta\), whence
\begin{equation*} \Big|\int_X f_n\,d\mu_n-\int_X T_k(f_n)\,d\mu_n\Big|\le\int_{\{|f_n|>k\}}|f_n|\,d\mu_n<\varepsilon \end{equation*}
for all \(n\), and likewise \(\big|\int_Xf\,d\mu-\int_XT_k(f)\,d\mu\big|<\varepsilon\); since \(\int_XT_k(f_n)\,d\mu_n\to\int_XT_k(f)\,d\mu\) by (i) and (ii), we get \(\limsup_n\big|\int f_n\,d\mu_n-\int f\,d\mu\big|\le2\varepsilon\).
(iv) For fixed \(k\), the map \(t\mapsto\min(t,k)\) is \(1\)-Lipschitz, so (i) gives \(\mu_n(\{|\min(f,k)-\min(f_n,k)|\ge\delta\})\to0\), and (ii), applicable as these functions are bounded by \(k\), gives
\begin{equation*} \int_X\min(f,k)\,d\mu=\lim_{n\to\infty}\int_X\min(f_n,k)\,d\mu_n\le\liminf_{n\to\infty}\int_X f_n\,d\mu_n \end{equation*}
(using \(0\le\min(f_n,k)\le f_n\), with \(f\ge0\) \(\mu\)-a.e.). Since \(\min(f,k)\uparrow f\) \(\mu\)-a.e., monotone convergence (Theorem 2.8.2) gives \(\int_Xf\,d\mu\le\liminf_n\int_Xf_n\,d\mu_n\) in \([0,+\infty]\).
(v) Here \(f_n\in L^1(\mu_n)\), and passing to the limit in \(|f_n|\le g_n\) gives \(|f|\le g\in L^1(\mu)\), so all integrals are finite. Applying (iv) to the nonnegative \(g_n-f_n\to g-f\) and using \(\int g_n\,d\mu_n\to\int g\,d\mu\) finitely,
\begin{equation*} \int_X g\,d\mu-\int_X f\,d\mu\le\int_X g\,d\mu-\limsup_{n\to\infty}\int_X f_n\,d\mu_n , \end{equation*}
i.e. \(\limsup_n\int f_n\,d\mu_n\le\int f\,d\mu\); the same applied to \(g_n+f_n\to g+f\ge0\) gives \(\int f\,d\mu\le\liminf_n\int f_n\,d\mu_n\).
Exercises 4.7.133–4.7.139
(Areshkin, Klimkin) Suppose that a sequence of measures \(\mu_n\) on a \(\sigma\)-algebra \(\mathcal{A}\) converges on every set in \(\mathcal{A}\) to a measure \(\mu\) and let \(\mathcal{A}\)-measurable functions \(f_n\) converge pointwise to a function \(f\), where \(f_n \in \mathcal{L}^1(\mu_n)\). Prove that the following conditions are equivalent:
(a) \(f \in \mathcal{L}^1(\mu)\) and
\begin{equation*} \int_A f \, d\mu = \lim_{n \to \infty} \int_A f_n \, d\mu_n \quad \text{for every } A \in \mathcal{A}; \end{equation*}
(b) for every \(\varepsilon > 0\), there exists \(\delta > 0\) such that
\begin{equation*} \Big| \int_A f_n \, d\mu_n \Big| \le \varepsilon \quad \text{whenever } A \in \mathcal{A} \text{ and } |\mu_n|(A) \le \delta . \end{equation*}
Put \(\nu := \nu_0/\nu_0(X)\), where \(\nu_0 := \sum_n2^{-n}|\mu_n|/(1+\|\mu_n\|)\) (if \(\nu_0(X) = 0\) then all \(\mu_n\), hence \(\mu\), vanish and both conditions hold trivially); \(\nu\) is a probability measure with \(|\mu_n|\ll\nu\) for all \(n\), and \(|\mu|\ll\nu\), since \(\nu(B) = 0\) forces \(\mu(B^{\prime}) = \lim_n\mu_n(B^{\prime}) = 0\) for every \(B^{\prime}\subset B\) in \(\mathcal{A}\). Condition (b) is meant uniformly in \(n\); for each fixed \(n\) it is automatic from Theorem 2.5.7.
Four ingredients.
(1) Theorem 4.6.3, applied to the setwise convergent \(\{\mu_n\}\), gives \(M := \sup_n\|\mu_n\|<\infty\), and its assertion (iii) with the dominating \(\nu\) gives \(\sup\{|\mu_n(B)|: \nu(B)\le t,\ n\in\mathbb{N}\}\to0\) as \(t\to0\); since \(|\mu_n|(B) = \mu_n(B\cap X^{+}_n)-\mu_n(B\cap X^{-}_n)\) for a Hahn decomposition of \(\mu_n\) (Theorem 3.1.1), with both sets inside \(B\), this passes to variations:
\begin{equation*} \forall\,\sigma>0\ \exists\,\eta>0:\quad \nu(B)\le\eta\ \Longrightarrow\ \sup_n|\mu_n|(B)\le\sigma . \tag{1} \end{equation*}
(I) \(\int_X\psi\,d\mu_n\to\int_X\psi\,d\mu\) for every bounded \(\mathcal{A}\)-measurable \(\psi\): approximate \(\psi\) uniformly within \(\epsilon\) by a simple \(s = \sum_{i\le k}a_i1_{A_i}\), so that the difference of the integrals is at most \(\epsilon(M+\|\mu\|)+\sum_i|a_i||\mu_n(A_i)-\mu(A_i)|\).
(II) For \(\varphi\in\mathcal{L}^1(\lambda)\), \(\varphi\cdot\lambda\) is a measure with \(|\varphi\cdot\lambda| = |\varphi|\cdot|\lambda|\) (write \(\lambda = \theta\cdot|\lambda|\) with \(\theta = 1_{X^{+}}-1_{X^{-}}\) and split by the sign of \(\varphi\theta\)). Hence \(\lambda_n(A) := \int_Af_n\,d\mu_n\) is a measure with \(|\lambda_n|(A) = \int_A|f_n|\,d|\mu_n|\), and \(\lambda_n\ll\nu\).
(III) Condition (b) is equivalent to its variational form
\begin{equation*} \forall\,\varepsilon>0\ \exists\,\delta>0:\ |\mu_n|(A)\le\delta \Longrightarrow\int_A|f_n|\,d|\mu_n|\le\varepsilon\ \ \text{for all }n, \tag{b\(^{\prime}\)} \end{equation*}
because for a Hahn decomposition \(X = P_n\cup N_n\) of \(\lambda_n\) the sets \(A\cap P_n\), \(A\cap N_n\) lie in \(A\), so applying (b) with \(\varepsilon/2\) to each gives \(|\lambda_n|(A) = \lambda_n(A\cap P_n)-\lambda_n(A\cap N_n)\le\varepsilon\).
Egorov sets. As \(f_n\to f\) everywhere and \(\nu\) is finite, Egorov’s theorem (Theorem 2.2.1) gives \(Y_j\in\mathcal{A}\) with \(\nu(X\setminus Y_j)\le2^{-j-1}\) and \(f_n\to f\) uniformly on \(Y_j\); replacing \(Y_j\) by \(Y_j\cap\{|f|\le C_j\}\), where \(\nu(|f|>C_j)\le2^{-j-1}\), and putting \(X_j := Y_1\cup\dots\cup Y_j\) yields an increasing sequence with
\begin{equation*} \nu(X\setminus X_j)\le2^{-j},\quad \sup_{X_j}|f|\le C_j^{\prime} := \max_{i\le j}C_i, \quad f_n\to f\ \text{unif. on }X_j , \tag{2} \end{equation*}
whence \(|\mu|\big(X\setminus\bigcup_jX_j\big) = 0\) because \(|\mu|\ll\nu\).
(a) \(\Rightarrow\) (b). By (II), \(\lambda(A) := \int_Af\,d\mu\) is a measure and (a) says \(\lambda_n\to\lambda\) setwise; as \(\lambda_n\ll\nu\), Theorem 4.6.3(iii) plus the Hahn argument of (1) give
\begin{equation*} \forall\,\varepsilon>0\ \exists\,\eta>0:\quad \nu(B)\le\eta\ \Longrightarrow\ \sup_n\int_B|f_n|\,d|\mu_n|\le\tfrac\varepsilon2 . \tag{3} \end{equation*}
Given \(\varepsilon>0\), take such an \(\eta\), then \(j\) with \(2^{-j}\le\eta\), and write \(X_0 := X_j\), \(C := C_j^{\prime}\); by (2) there is \(n_0\) with \(\sup_{X_0}|f_n-f|\le1\), hence \(|f_n|\le C+1\) on \(X_0\), for \(n\ge n_0\). For such \(n\) and \(|\mu_n|(A)\le\delta_{*} := \varepsilon/(2(C+1))\),
\begin{equation*} \Big|\int_A f_n\,d\mu_n\Big|\le(C+1)|\mu_n|(A)+\int_{A\setminus X_0}|f_n|\,d|\mu_n| \le\tfrac\varepsilon2+\tfrac\varepsilon2 , \end{equation*}
the tail estimated by (3) with \(B = X\setminus X_0\); for the finitely many \(n<n_0\), absolute continuity of the integral of \(|f_n|\in\mathcal{L}^1(|\mu_n|)\) (Theorem 2.5.7) gives \(\delta_n\), and \(\delta := \min(\delta_{*},\delta_1,\dots,\delta_{n_0-1})\) serves all \(n\).
(b) \(\Rightarrow\) (a). Claim:
\begin{equation*} \forall\,\varepsilon>0\ \exists\,\eta>0:\quad \nu(B)\le\eta\ \Longrightarrow\ \int_B|f|\,d|\mu|\le\varepsilon . \tag{4} \end{equation*}
Take \(\delta\) from (b\(^{\prime}\)) for \(\varepsilon\) and \(\eta\) from (1) for \(\sigma = \delta\), so \(\nu(B)\le\eta\) gives \(\sup_n\int_B|f_n|\,d|\mu_n|\le\varepsilon\). Fix such a \(B\) and any \(j\), and let \(\theta := 1_{X^{+}}-1_{X^{-}}\) for a Hahn decomposition of \(\mu\), so that \(|\mu| = \theta\cdot\mu\). The function \(|f|\theta1_{B\cap X_j}\) is bounded by \(C_j^{\prime}\), so (I) gives
\begin{equation*} \int_{B\cap X_j}|f|\,d|\mu| = \lim_{n\to\infty}\int_{B\cap X_j}|f|\,\theta\,d\mu_n , \end{equation*}
and splitting \(|f| = |f_n|+(|f|-|f_n|)\) there bounds the right side by \(\varepsilon\): the second part is at most \(M\sup_{X_j}|f-f_n|\to0\) by (2), the first at most \(\int_B|f_n|\,d|\mu_n|\le\varepsilon\) since \(|\theta|\le1\). So \(\int_{B\cap X_j}|f|\,d|\mu|\le\varepsilon\) for every \(j\), and monotone convergence together with \(|\mu|\big(X\setminus\bigcup_jX_j\big) = 0\) gives (4).
Applying (4) with \(\varepsilon = 1\) and \(j\) with \(2^{-j}\le\eta\),
\begin{equation*} \int_X|f|\,d|\mu| = \int_{X_j}|f|\,d|\mu|+\int_{X\setminus X_j}|f|\,d|\mu| \le C_j^{\prime}\|\mu\|+1<\infty , \end{equation*}
so \(f\in\mathcal{L}^1(\mu)\). Finally fix \(A\in\mathcal{A}\) and \(\varepsilon>0\); with \(\delta\) from (b\(^{\prime}\)), \(\eta\) as in the Claim (so \(\nu(B)\le\eta\) also gives \(\sup_n\int_B|f_n|\,d|\mu_n|\le\varepsilon\)), \(j\) with \(2^{-j}\le\eta\), \(X_0 := X_j\) and \(C := C_j^{\prime}\),
\begin{equation*} \begin{aligned} \Big| \int_A f_n \, d\mu_n - \int_A f \, d\mu \Big| &\le \Big| \int_{A \cap X_0} (f_n - f) \, d\mu_n \Big|\\ &\quad + \Big| \int_{A \cap X_0} f \, d\mu_n - \int_{A \cap X_0} f \, d\mu \Big| + 2 \varepsilon , \end{aligned} \end{equation*}
the two tail terms \(\int_{A\setminus X_0}|f_n|\,d|\mu_n|\) and \(\int_{A\setminus X_0}|f|\,d|\mu|\) being estimated by the Claim’s bound and by (4) with \(B = X\setminus X_0\). The first term on the right is at most \(M\sup_{X_0}|f_n-f|\to0\) by (2) and (1), and the second tends to \(0\) by (I), since \(f1_{A\cap X_0}\) is bounded by \(C\). Hence \(\limsup_n\big|\int_Af_n\,d\mu_n-\int_Af\,d\mu\big|\le2\varepsilon\).
(Gowurin) Let \(\mathcal{X}\) be the space of all equivalence classes of Lebesgue measurable sets in \([0,1]\) equipped with the metric \(d(A,B) = \lambda(A \triangle B)\), where \(\lambda\) is Lebesgue measure. Let \(S(E_0,r) = \{E \in \mathcal{X} \colon d(E,E_0) = r\}\) be the sphere of radius \(r \in (0,1)\) with the center \(E_0 \in \mathcal{X}\). Suppose that this sphere does not contain the element corresponding to the empty set. Prove that if \(f_n \in L^1[0,1]\) and
\begin{equation*} \lim_{n \to \infty} \int_E f_n \, dx = 0 \quad \text{for all } E \in S(E_0,r), \end{equation*}
then the same is true for every measurable set \(E \subset [0,1]\).
Write \(\nu_n(E) := \int_Ef_n\,dx\) and \(S := S(E_0,r)\), all sets being taken modulo null sets. Two facts recur: each \(\nu_n\) is \(d\)-continuous, since \(|\nu_n(E)-\nu_n(F)|\le\int_{E\triangle F}|f_n|\,dx\to0\) as \(\lambda(E\triangle F)\to0\) by absolute continuity of the integral of \(f_n\in L^1\) (Theorem 2.5.7) (F1); and \(\lambda\) is atomless (Example 1.12.8), so Corollary 1.12.10 gives, inside any measurable \(W\), subsets of every measure in \([0,\lambda(W)]\) (F2). Moreover, for disjoint \(E,A\),
\begin{equation*} d(E \cup A, E_0) = d(E,E_0) + \lambda(A \setminus E_0) - \lambda(A \cap E_0), \tag{1} \end{equation*}
by splitting \((E\cup A)\triangle E_0 = [(E\setminus E_0)\cup(A\setminus E_0)]\cup[(E_0\setminus E)\setminus A]\) into disjoint pieces.
Put \(a := \lambda(E_0)\); as \(d(\emptyset,E_0) = a\), the hypothesis \(\emptyset\notin S\) says \(\gamma := |a-r|>0\). Let
\begin{equation*} Y := E_0 \ \text{ if } a > r, \qquad Y := [0,1] \setminus E_0 \ \text{ if } a < r, \qquad Y^{\prime} := [0,1] \setminus Y , \end{equation*}
and \(\kappa := \lambda(Y)-\gamma\), which equals \(r>0\) in the first case and \(1-r>0\) in the second, so \(\lambda(Y) = \gamma+\kappa\) with \(\gamma,\kappa>0\). Call \(A\) balanced if \(\lambda(A\cap Y) = \lambda(A\setminus Y)\); then \(\lambda(A\setminus E_0) = \lambda(A\cap E_0)\), so by (1) \(d(E\cup A,E_0) = d(E,E_0)\) for balanced \(A\) disjoint from \(E\).
Lemma 0. \(\mathcal{S}_Y := \{G\subset Y: \lambda(G) = \gamma\}\) is a nonempty subset of \(S\): for \(a>r\), \(d(G,E_0) = \lambda(E_0\setminus G) = a-\gamma = r\); for \(a<r\), \(d(G,E_0) = \lambda(G)+\lambda(E_0) = \gamma+a = r\); and \(\gamma<\lambda(Y)\) with (F2) gives existence. So the hypothesis yields
\begin{equation*} \lim_{n \to \infty} \nu_n(G) = 0 \qquad \text{for every } G \in\mathcal{S}_Y , \tag{2} \end{equation*}
and also \(\nu_n(G\cup A)\to0\) for balanced \(A\) disjoint from such a \(G\).
Lemma 1. Disjoint \(P,P^{\prime}\subset Y\) with \(\lambda(P) = \lambda(P^{\prime}) = \beta\le\min(\gamma,\kappa)\) have \(\nu_n(P)-\nu_n(P^{\prime})\to0\). Since \(\lambda(Y\setminus(P\cup P^{\prime})) = \gamma+\kappa-2\beta\ge\gamma-\beta\), (F2) gives \(G\in\mathcal{S}_Y\) with \(P^{\prime}\subset G\) and \(G\cap P = \emptyset\); then \(G^{\prime} := (G\setminus P^{\prime})\cup P\in\mathcal{S}_Y\) and \(\nu_n(P)-\nu_n(P^{\prime}) = \nu_n(G^{\prime})-\nu_n(G)\to0\) by (2).
Lemma 2. Balanced \(A\) has \(\nu_n(A)\to0\). If \(\alpha := \lambda(A\cap Y)\le\kappa\), then \(\lambda(Y\setminus A) = \gamma+\kappa-\alpha\ge\gamma\), so (F2) gives \(G\in\mathcal{S}_Y\) disjoint from \(A\), and \(\nu_n(A) = \nu_n(G\cup A)-\nu_n(G)\to0\) by Lemma 0. In general split \(A\cap Y\) and \(A\setminus Y\) into \(N\) pieces \(B_i\), \(C_i\) of measure \(\alpha/N\le\kappa\); each \(B_i\cup C_i\) is balanced of the previous kind, and \(\nu_n(A) = \sum_i\nu_n(B_i\cup C_i)\).
Lemma 3. \(P\subset Y\) with \(\lambda(P) = (m/N)\gamma\le\lambda(Y)\) has \(\nu_n(P)\to0\). Replacing \((m,N)\) by \((mj,Nj)\), assume \(\gamma/N\le\min(\gamma,\kappa)\). Splitting some \(G\in\mathcal{S}_Y\) into \(G_1,\dots,G_N\) of measure \(\gamma/N\), Lemma 1 gives \(\nu_n(G_i)-\nu_n(G_1)\to0\) while \(\sum_i\nu_n(G_i) = \nu_n(G)\to0\) by (2), so \(\nu_n(G_1)\to0\). For arbitrary \(Q\subset Y\) with \(\lambda(Q) = \gamma/N\), repeating this with \(G\subset Y\setminus Q\) (possible since \(\lambda(Y\setminus Q)\ge\gamma\)) makes \(G_1\cap Q = \emptyset\), so \(\nu_n(Q)\to0\) by Lemma 1; splitting \(P\) into \(m\) such pieces finishes.
Lemma 4 (Baire step). For every \(\varepsilon>0\) there are \(N_\varepsilon\), \(G_1\in\mathcal{S}_Y\) and \(\rho>0\) with
\begin{equation*} |\nu_k(G)| \le \varepsilon \quad \text{whenever } k \ge N_\varepsilon , \ G \in \mathcal{S}_Y , \ d(G,G_1) \le \rho . \end{equation*}
Indeed \((\mathcal{X},d)\) is complete (Theorem 1.12.6(ii)) and \(\mathcal{S}_Y = \{G: \lambda(G) = \gamma,\ \lambda(G\setminus Y) = 0\}\) is closed, both functionals being \(1\)-Lipschitz, so \(\mathcal{S}_Y\) is a nonempty complete metric space; the sets \(F_N := \{G\in\mathcal{S}_Y: |\nu_k(G)|\le\varepsilon\ \text{for } k\ge N\}\) are closed by (F1) and cover \(\mathcal{S}_Y\) by (2), so some \(F_{N_\varepsilon}\) contains a ball by the Baire category theorem.
Fix \(\varepsilon>0\) and the \(N_\varepsilon,G_1,\rho\) of Lemma 4, and set
\begin{equation*} \delta := \min\{\rho/6, \ \gamma/3, \ \kappa/3\} > 0 . \end{equation*}
Lemma 5. For \(k\ge N_\varepsilon\) and disjoint \(P,P^{\prime}\subset Y\) with \(\lambda(P) = \lambda(P^{\prime}) = \beta\le\delta\), \(|\nu_k(P)-\nu_k(P^{\prime})|\le2\varepsilon\). Put \(H := (G_1\setminus P)\cup P^{\prime}\subset Y\), so \(d(H,G_1)\le2\beta\) and \(\lambda(H) = \gamma+\eta\) with \(\eta := \lambda(P^{\prime}\setminus G_1)-\lambda(G_1\cap P)\), \(|\eta|\le\beta\). If \(\eta\ge0\), then \(\lambda(H\setminus P^{\prime})\ge\gamma+\eta-\beta\ge\eta\), so (F2) gives \(Z\subset H\setminus P^{\prime}\) of measure \(\eta\) and we put \(G := H\setminus Z\); if \(\eta<0\), then \(\lambda(Y\setminus(H\cup P))\ge\kappa-2\beta\ge\beta\ge-\eta\), so \(Z\subset Y\setminus(H\cup P)\) of measure \(-\eta\) gives \(G := H\cup Z\). Either way \(G\in\mathcal{S}_Y\), \(P^{\prime}\subset G\), \(G\cap P = \emptyset\) and \(d(G,G_1)\le|\eta|+2\beta\le3\beta\le\rho\), while \(G^{\prime} := (G\setminus P^{\prime})\cup P\in\mathcal{S}_Y\) has \(d(G^{\prime},G_1)\le5\beta\le5\rho/6<\rho\); so Lemma 4 bounds \(|\nu_k(P)-\nu_k(P^{\prime})| = |\nu_k(G^{\prime})-\nu_k(G)|\le2\varepsilon\).
Lemma 6. \(\Phi_k := \sup\{|\nu_k(M)|: M\subset Y,\ \lambda(M)\le\delta\}\le4\varepsilon\) for \(k\ge N_\varepsilon\); here \(\Phi_k\le\int_Y|f_k|\,dx<\infty\), which legitimizes the absorption below. Fix \(M\subset Y\) with \(\beta := \lambda(M)\in(0,\delta]\) and \(J := \lfloor\gamma/\beta\rfloor\ge3\) (as \(\beta\le\gamma/3\)), so \(J\beta\le\gamma\). With \(H := G_1\cup M\) one has \(\lambda(H) = \gamma+\eta\), \(\eta := \lambda(M\setminus G_1)\in[0,\beta]\), \(d(H,G_1) = \eta\), and \(\lambda(H\setminus M)\ge\eta\), so (F2) gives \(Z\subset H\setminus M\) of measure \(\eta\) and \(G := H\setminus Z\in\mathcal{S}_Y\) with \(M\subset G\), \(d(G,G_1)\le2\beta\le\rho\), whence \(|\nu_k(G)|\le\varepsilon\) by Lemma 4. Since \(\lambda(G\setminus M) = \gamma-\beta\ge(J-1)\beta\), pick disjoint \(M_2,\dots,M_J\subset G\setminus M\) of measure \(\beta\), and let \(R := G\setminus(M\cup M_2\cup\dots\cup M_J)\subset Y\), \(\lambda( R) = \gamma-J\beta<\beta\le\delta\). Additivity gives
\begin{equation*} J \nu_k(M) = \nu_k(G) - \nu_k( R) + \sum_{j=2}^{J} \bigl(\nu_k(M) - \nu_k(M_j)\bigr) , \end{equation*}
each summand being controlled by Lemma 5, so \(|\nu_k(M)|\le2\varepsilon+(\varepsilon+\Phi_k)/3\); taking the supremum over such \(M\) and absorbing \(\Phi_k/3<\infty\) gives \(\tfrac23\Phi_k\le\tfrac73\varepsilon\), i.e. \(\Phi_k\le\tfrac72\varepsilon\le4\varepsilon\).
Lemma 7. With \(\delta_1 := \min\{\delta,\lambda(Y)\}\), every \(A\) with \(\lambda(A)\le\delta_1\) has \(\limsup_k|\nu_k(A)|\le8\varepsilon\): Lemma 6 gives \(|\nu_k(A\cap Y)|\le4\varepsilon\) for \(k\ge N_\varepsilon\), and for \(C := A\setminus Y\), (F2) provides \(B\subset Y\) with \(\lambda(B) = \lambda( C)\le\delta\), so \(B\cup C\) is balanced with \(\nu_k(B\cup C)\to0\) (Lemma 2) and \(|\nu_k(B)|\le4\varepsilon\) (Lemma 6), whence \(\limsup_k|\nu_k( C)|\le4\varepsilon\).
Conclusion. Let \(A\) be measurable, \(p := \lambda(A\cap Y)\), \(q := \lambda(A\setminus Y)\).
(i) \(p\ge q\): take \(A_1\subset A\cap Y\) of measure \(q\), so that \(A = A_b\cup B\) disjointly with \(A_b := A_1\cup(A\setminus Y)\) balanced, hence \(\nu_k(A_b)\to0\) by Lemma 2, and \(B := (A\cap Y)\setminus A_1\subset Y\) of measure \(p-q\). If \(p-q\le\delta\), Lemma 6 gives \(\limsup_k|\nu_k(B)|\le4\varepsilon\); otherwise, \(\{(m/N)\gamma\}\) being dense in \((0,+\infty)\), choose \(m,N\) with \((m/N)\gamma\le p-q\le(m/N)\gamma+\delta\) and \(P\subset B\) of measure \((m/N)\gamma\le\lambda(Y)\), so that \(\nu_k(P)\to0\) by Lemma 3 and \(\limsup_k|\nu_k(B\setminus P)|\le4\varepsilon\) by Lemma 6.
(ii) \(p<q\): take \(A_2\subset A\setminus Y\) of measure \(p\), so \(A = A_b\cup C\) with \(A_b := A_2\cup(A\cap Y)\) balanced and \(C\subset Y^{\prime}\) of measure \(q-p\). If \(\lambda( C)\le\delta_1\), Lemma 7 applies. Otherwise choose \(m,N\) with \((m/N)\gamma\le q-p\le(m/N)\gamma+\delta_1\) and \(P\subset C\) of measure \((m/N)\gamma\), split \(P\) into \(\ell\) equal pieces \(P_i\) of measure \((m/(N\ell))\gamma\le\lambda(Y)\), and pick \(B_i\subset Y\) with \(\lambda(B_i) = \lambda(P_i)\); each \(B_i\cup P_i\) is balanced, so \(\nu_k(P_i) = \nu_k(B_i\cup P_i)-\nu_k(B_i)\to0\) by Lemmas 2 and 3, hence \(\nu_k(P)\to0\), while \(\limsup_k|\nu_k(C\setminus P)|\le8\varepsilon\) by Lemma 7.
In both cases \(\limsup_k|\nu_k(A)|\le8\varepsilon\), and \(\varepsilon>0\) was arbitrary.
(S. Saks) Prove that the class of all open sets is a first category set (a countable union of nowhere dense sets) in the space \(X\) from the previous exercise, i.e. in the space of all equivalence classes of Lebesgue measurable sets in \([0,1]\) equipped with the metric \(d(A,B) = \lambda(A \triangle B)\), where \(\lambda\) is Lebesgue measure.
The class \(\mathcal{G}\) of open sets satisfies \(\mathcal{G}\subset\{[\emptyset]\}\cup\bigcup_{k\ge1}M_k\), where, for a fixed enumeration \(\{U_n\}\) of the intervals of \([0,1]\) with rational endpoints, \(M_k\) is the set of classes \([U]\) of open \(U\subset[0,1]\) admitting indices \(n_1,\dots,n_k\) with
\begin{equation*} U_{n_i} \subset U \ (i \le k), \qquad \lambda\Bigl( U \setminus \bigcup_{i\le k} U_{n_i} \Bigr) \le \frac{\lambda(U)}{4} , \end{equation*}
and each of these sets is nowhere dense in the complete metric space \(X\) (Theorem 1.12.6(ii)). We use throughout that \(M\subset X\) is nowhere dense exactly when every ball \(B(E_0,r)\) contains a ball disjoint from \(M\).
The inclusion. A nonempty open \(U\) has \(\lambda(U)>0\) and countably many components \(J_s\), disjoint nondegenerate intervals relatively open in \([0,1]\); take \(m\) with \(\lambda(U\setminus\bigcup_{s\le m}J_s)<\lambda(U)/8\) and, inside each \(J_s\), a rational-endpoint interval \(R_s = U_{n_s}\) with \(\lambda(J_s\setminus R_s)<\lambda(U)/(8m)\) (possible as \(0\) and \(1\) are rational). Then \(\lambda(U\setminus\bigcup_{s\le m}U_{n_s})<\lambda(U)/8+\lambda(U)/8\), so \([U]\in M_m\).
\(\{[\emptyset]\}\) is nowhere dense, being closed with \([\emptyset]\) not isolated: \(d(\emptyset,[0,t]) = t\) for \(t\in(0,1)\).
Combs are dense. Call \(C = I_1\cup\dots\cup I_p\) a \(p\)-comb if the \(I_j\) are open intervals of one positive length with disjoint closures. Given \(E_0\), \(r>0\) and \(P\), the finite unions of intervals form an algebra generating the Lebesgue sets (Example 1.4.4, Theorem 1.5.6), so some union \(A\) of \(m\) intervals has \(\lambda(E_0\triangle A)<r/2\). Fix \(N\) with \(2m/N<r/8\), put \(D_i := ((i-1)/N,i/N)\), \(S := \{i: D_i\subset A\}\) and \(A^{\prime\prime} := \bigcup_{i\in S}D_i\); an \(i\notin S\) with \(D_i\cap A\ne\emptyset\) forces the connected \(D_i\) to contain a point of \(\partial A\), a set of at most \(2m\) points, so \(\lambda(A\triangle A^{\prime\prime}) = \lambda(A\setminus A^{\prime\prime})\le2m/N<r/8\). If \(S = \emptyset\), then \(\lambda(E_0)<r/2+r/8\) and any comb with even \(p\ge P\) and \(\lambda( C)<r/4\) serves. Otherwise fix even \(M\ge P\), subdivide each \(D_i\), \(i\in S\), into \(M\) equal intervals, and shrink each of the \(p := M|S|\ge P\) of them symmetrically by \(\beta\) with \(2\beta p<r/8\); the resulting \(p\)-comb \(C\subset A^{\prime\prime}\) (consecutive intervals separated by \(2\beta\)) has
\begin{equation*} d(E_0,C) \le \lambda(E_0 \triangle A) + \lambda(A \triangle A^{\prime\prime}) + \lambda(A^{\prime\prime} \triangle C) < \tfrac{r}{2}+\tfrac{r}{8}+\tfrac{r}{8} < r . \end{equation*}
Each \(M_k\) is nowhere dense. Fix \(k\) and \(B(E_0,r)\), take a \(p\)-comb \(C\) with \(p\ge8k\) even and \(d(E_0,C)<r/2\), and put \(\eta := \min_{j\ne j^{\prime}}\operatorname{dist}(I_j,I_{j^{\prime}})>0\) and \(\delta := \tfrac12\min\{\eta,\lambda( C)/16,r\}\), so that \(B(C,\delta)\subset B(E_0,r)\). Suppose some \([U]\in M_k\) had \(\lambda(U\triangle C)<\delta\), with \(W := \bigcup_{i\le k}U_{n_i}\subset U\) and \(\lambda(U\setminus W)\le\lambda(U)/4\).
(a) Each \(U_{n_i}\) meets at most one \(I_j\): meeting \(I_{j_1}\) at \(x\) and a later \(I_{j_2}\) at \(y\) would give \([x,y]\subset U_{n_i}\subset U\), hence the gap \(G\) between \(I_{j_1}\) and the next constituent interval satisfies \(G\subset U\), \(G\cap C = \emptyset\), \(\lambda(G)\ge\eta\), so \(\lambda(U\triangle C)\ge\eta>\delta\).
(b) Hence at least \(p-k\ge\frac78p\) of the \(I_j\), each of measure \(\lambda( C)/p\), miss \(W\), giving \(\lambda(C\setminus W)\ge\frac78\lambda( C)\).
(c) Since \(C\setminus W\subset(C\setminus U)\cup(U\setminus W)\) and \(\lambda(U)\le\lambda( C)+\delta\le(1+\tfrac1{32})\lambda( C)\),
\begin{equation*} \lambda(C \triangle U)\ \ge\ \tfrac78\lambda( C)-\tfrac14\bigl(1+\tfrac1{32}\bigr)\lambda( C) = \tfrac{79}{128}\lambda( C)\ >\ \tfrac{\lambda( C)}{32}\ \ge\ \delta , \end{equation*}
a contradiction; so \(B(C,\delta)\cap M_k = \emptyset\).
Verify the equivalence of (i) and (ii) in Theorem 4.7.27.
Both implications run by contradiction from the tail estimate available for a single measure: each \(|\mu_\alpha|\) is finite, nonnegative and countably additive on \(\mathcal{S}\) (Definition 3.1.4 and formula (3.1.3)), so for pairwise disjoint \(R_k\in\mathcal{R}\),
\begin{equation*} \Bigl|\sum_{k\ge n}\mu_\alpha(R_k)\Bigr| = \Bigl|\mu_\alpha\Bigl(\bigcup_{k\ge n}R_k\Bigr)\Bigr| \le|\mu_\alpha|\Bigl(\bigcup_{k\ge n}R_k\Bigr) = \sum_{k\ge n}|\mu_\alpha|(R_k)\to0 , \end{equation*}
a tail of a convergent series; also finite unions of the \(R_k\) lie in the ring \(\mathcal{R}\).
(i) \(\Rightarrow\) (ii). If (ii) fails, then after passing to a subsequence there are measures \(\nu_n\) of the family, disjoint \(R_n\in\mathcal{R}\) and \(\delta>0\) with \(|\nu_n(R_n)|\ge\delta\) for all \(n\). Put \(S_n := \bigcup_{k\ge n}R_k\), take \(n_1 := 1\) and, using the display for the fixed measure \(\nu_{n_k}\), indices \(n_{k+1}>n_k\) with \(|\nu_{n_k}|(S_{n_{k+1}})<\delta/2\). The sets \(Q_i := R_{n_i}\) are disjoint members of \(\mathcal{R}\) with \(\bigcup_{i>k}Q_i\subset S_{n_{k+1}}\), so
\begin{equation*} \Bigl| \sum_{i\ge k} \nu_{n_k}(Q_i) \Bigr| \ge |\nu_{n_k}(R_{n_k})| - |\nu_{n_k}|(S_{n_{k+1}}) > \delta/2 \end{equation*}
for every \(k\), contradicting (i) applied to \(\{Q_i\}\).
(ii) \(\Rightarrow\) (i). If (i) fails for disjoint \(R_j\in\mathcal{R}\), there are \(\delta>0\) and an infinite \(K\) with \(\sup_\alpha\bigl|\sum_{j\ge k}\mu_\alpha(R_j)\bigr|>\delta\) for \(k\in K\). Construct \(m_1<p_1<m_2<p_2<\cdots\): given \(m_k\in K\), pick \(\nu^{(k)}\) in the family with \(\bigl|\sum_{j\ge m_k}\nu^{(k)}(R_j)\bigr|>\delta\), then (by the display for \(\nu^{(k)}\)) an index \(p_k>m_k\) with \(\bigl|\sum_{j\ge p_k}\nu^{(k)}(R_j)\bigr|<\delta/2\), and then \(m_{k+1}\in K\) with \(m_{k+1}>p_k\). The sets \(E_k := \bigcup_{m_k\le j<p_k}R_j\in\mathcal{R}\) are pairwise disjoint, the index blocks being so, and finite additivity gives
\begin{equation*} |\nu^{(k)}(E_k)| \ge \Bigl| \sum_{j\ge m_k} \nu^{(k)}(R_j) \Bigr| - \Bigl| \sum_{j\ge p_k} \nu^{(k)}(R_j) \Bigr| > \delta/2 , \end{equation*}
contradicting (ii).
Let \(\mu_n\) be real measures of bounded variation on the \(\sigma\)-ring \(\mathcal{S}\) generated by a ring \(\mathcal{R}\). Suppose that \(\lim_{n\to\infty}\mu_n(R_n)=0\) for every infinite sequence of disjoint sets \(R_n\in\mathcal{R}\).
(i) Let \(A_k=\bigcup_{j=1}^{\infty}A_j^k\), \(B_k=\bigcup_{j=1}^{\infty}B_j^k\), where \(A_j^k,B_j^k\in\mathcal{R}\), and let the sets \(E_k=A_k\backslash B_k\) be pairwise disjoint. Prove that \(\lim_{n\to\infty}|\mu_n|(E_n)=0\).
(ii) Prove that, for every \(S\in\mathcal{S}\) and \(\varepsilon>0\), there exists a set \(R\) of the form \(R=\bigcup_{j=1}^{\infty}R_j\) with \(R_j\in\mathcal{R}\) such that \(|\mu_n|(S\,\triangle\,R)<\varepsilon\) for all \(n\).
Each \(|\mu_n|\) is a finite nonnegative countably additive measure on \(\mathcal{S}\), and (H) denotes the hypothesis that \(\mu_n(R_n)\to0\) for every disjoint sequence \(R_n\in\mathcal{R}\), empty terms being allowed. Theorem 4.7.27(iv) is proved there by means of this exercise, so only (H), the Hahn decomposition (Theorem 3.1.1) and Lemma A below may be used. For \(D\in\mathcal{S}\) write \(\mathcal{S}_D := \{S\in\mathcal{S}: S\subset D\}\); for \(D\in\mathcal{R}\) the algebra \(\mathcal{R}_D\) generates \(\mathcal{S}_D\), since \(\{S\in\mathcal{S}: S\cap D\in\sigma_D(\mathcal{R}_D)\}\) is a \(\sigma\)-ring containing \(\mathcal{R}\).
Lemma A. For a finite nonnegative \(\nu\) on \(\mathcal{S}\), every \(S\in\mathcal{S}\) and \(\delta>0\) admit \(E\in\mathcal{R}\) with \(\nu(S\triangle E)<\delta\). The class \(\mathcal{M}\) of such \(S\) contains \(\mathcal{R}\) and is closed under differences and finite unions, by
\begin{equation*} (S_1\ast S_2)\,\triangle\,(E_1\ast E_2)\subset (S_1\,\triangle\,E_1)\cup(S_2\,\triangle\,E_2), \qquad \ast\in\{\backslash,\cup\}, \end{equation*}
and under countable unions, since for \(S = \bigcup_iS_i\) finiteness of \(\nu\) gives \(m\) with \(\nu(S\backslash\bigcup_{i\le m}S_i)<\delta/2\); so \(\mathcal{M}\) is a \(\sigma\)-ring containing \(\mathcal{R}\), i.e. \(\mathcal{M} = \mathcal{S}\).
Lemma B. For \(\mu\) of bounded variation and \(T\in\mathcal{R}\) there is \(T^{\prime}\in\mathcal{R}\), \(T^{\prime}\subset T\), with \(|\mu(T^{\prime})|\ge|\mu|(T)/4\). Assume \(|\mu|(T)>0\) and set \(\delta := |\mu|(T)/4\). The Hahn decomposition on \(\mathcal{S}_T\) gives \(P\subset T\) with \(|\mu|(T) = \mu(P)-\mu(T\backslash P)\), and Lemma A (with \(\nu = |\mu|\)) gives \(E\in\mathcal{R}\) with \(|\mu|(P\triangle E)<\delta\). For \(T^{\prime} := E\cap T\in\mathcal{R}\) one has \(P\triangle T^{\prime}\subset P\triangle E\), and \((T\backslash T^{\prime})\triangle(T\backslash P) = T^{\prime}\triangle P\), so both \(|\mu(T^{\prime})-\mu(P)|\) and \(|\mu(T\backslash T^{\prime})-\mu(T\backslash P)|\) are \(<\delta\), whence
\begin{equation*} |\mu(T^{\prime})|+|\mu(T\backslash T^{\prime})|\ \ge\ |\mu|(T)-2\delta\ =\ |\mu|(T)/2 \end{equation*}
and one of the two sets, both in \(\mathcal{R}\) and inside \(T\), does it.
Lemma C. For pairwise disjoint \(T_k\in\mathcal{R}\) and arbitrary indices \((n_k)\), \(|\mu_{n_k}|(T_k)\to0\). Otherwise \(|\mu_{n_k}|(T_k)\ge\varepsilon\) along an infinite \(K\), and Lemma B gives disjoint \(T_k^{\prime}\subset T_k\) in \(\mathcal{R}\) with \(|\mu_{n_k}(T_k^{\prime})|\ge\varepsilon/4\) for \(k\in K\). (a) If \(n_k = m\) for infinitely many \(k\in K\), then \(\sum_k|\mu_m|(T_k^{\prime}) = |\mu_m|(\bigcup_kT_k^{\prime})<\infty\) forces \(\mu_m(T_k^{\prime})\to0\) along them. (b) Otherwise choose \(k_1<k_2<\cdots\) in \(K\) with \(n_{k_1}<n_{k_2}<\cdots\) and put \(G_m := T_{k_i}^{\prime}\) for \(m = n_{k_i}\), \(G_m := \emptyset\) otherwise: these are disjoint sets of \(\mathcal{R}\) with \(\mu_m(G_m)\not\to0\), contradicting (H).
Lemma D. For pairwise disjoint \(A_j\in\mathcal{R}\) and \(U_p := \bigcup_{j>p}A_j\), \(\sup_n|\mu_n|(U_p)\to0\). If not, \(\sup_n|\mu_n|(U_p)>\varepsilon\) for every \(p\); starting from \(q_0 := 0\), pick \(m_k\) with \(|\mu_{m_k}|(U_{q_{k-1}})>\varepsilon\) and then, since \(U_q\downarrow\emptyset\) and \(|\mu_{m_k}|\) is finite, \(q_k>q_{k-1}\) with \(|\mu_{m_k}|(U_{q_k})<\varepsilon/2\), so that the pairwise disjoint sets \(V_k := \bigcup_{q_{k-1}<j\le q_k}A_j\in\mathcal{R}\) satisfy
\begin{equation*} |\mu_{m_k}|(V_k)=|\mu_{m_k}|(U_{q_{k-1}})-|\mu_{m_k}|(U_{q_k})>\varepsilon/2 , \end{equation*}
contradicting Lemma C.
(i) We prove the stronger statement (i’): if \(G_k = A^{(k)}\backslash B^{(k)}\) are pairwise disjoint, with \(A^{(k)} = \bigcup_jA^k_j\) and \(B^{(k)} = \bigcup_jB^k_j\), \(A^k_j,B^k_j\in\mathcal{R}\), then \(|\mu_{n_k}|(G_k)\to0\) for every index sequence \((n_k)\); \(n_k = k\) gives (i). If (i’) fails, then after discarding terms and relabelling there are pairwise disjoint \(F_i = A^{(i)}\backslash B^{(i)}\) of the same form and measures \(\nu_i := \mu_{n_{k_i}}\) with \(|\nu_i|(F_i)\ge\varepsilon\). Replacing \(A^i_j\) by \(A^i_j\backslash(A^i_1\cup\dots\cup A^i_{j-1})\in\mathcal{R}\), and likewise for \(B^i_j\), we may take these sequences disjoint, so Lemma D gives \(p_i\) with
\begin{equation*} |\mu_n|\Bigl(\bigcup_{j>p_i}A^i_j\Bigr)<\varepsilon 2^{-i}/8, \qquad |\mu_n|\Bigl(\bigcup_{j>p_i}B^i_j\Bigr)<\varepsilon 2^{-i}/8\qquad\text{for all }n . \end{equation*}
Put \(C_i := A^{\prime}\backslash B^{\prime}\in\mathcal{R}\), where \(A^{\prime} := \bigcup_{j\le p_i}A^i_j\), \(A^{\prime\prime} := \bigcup_{j>p_i}A^i_j\) and similarly for \(B\). A point of \(C_i\backslash F_i\) lies in \(B^{\prime\prime}\) and a point of \(F_i\backslash C_i\) lies in \(A^{\prime\prime}\), so \(C_i\triangle F_i\subset A^{\prime\prime}\cup B^{\prime\prime}\) and \(|\mu_n|(C_i\triangle F_i)<\varepsilon2^{-i}/4\) for all \(n\). A point of \(C_i\cap C_j\) outside both symmetric differences would lie in \(F_i\cap F_j = \emptyset\), so \(C_i\cap C_j\subset(C_i\triangle F_i)\cup(C_j\triangle F_j)\) and \(|\mu_n|(C_i\cap C_j)<\tfrac\varepsilon4(2^{-i}+2^{-j})\) for \(i\ne j\). Hence the pairwise disjoint sets \(D_i := C_i\backslash\bigcup_{j<i}C_j\in\mathcal{R}\), for which \(C_i\triangle D_i = \bigcup_{j<i}(C_i\cap C_j)\), satisfy
\begin{equation*} |\mu_n|(C_i\,\triangle\,D_i)\le\sum_{j<i}\frac{\varepsilon}{4}\bigl(2^{-i}+2^{-j}\bigr) <\frac{\varepsilon}{4}\Bigl((i-1)2^{-i}+1\Bigr)\le\frac{5\varepsilon}{16} \end{equation*}
(using \((i-1)2^{-i}\le1/4\)), so \(|\nu_i|(F_i\triangle D_i)<\varepsilon2^{-i}/4+5\varepsilon/16\le7\varepsilon/16<3\varepsilon/4\) and therefore \(|\nu_i|(D_i)\ge|\nu_i|(F_i)-|\nu_i|(F_i\triangle D_i)>\varepsilon/4\), contradicting Lemma C for the index sequence \((n_{k_i})\).
(ii) Fix \(S\in\mathcal{S}\) and \(\varepsilon>0\), and put \(\nu_n := |\mu_1|+\dots+|\mu_n|\). Lemma A gives \(E_n\in\mathcal{R}\) with \(\nu_n(S\triangle E_n)<\varepsilon2^{-n}/4\); let \(D_n := \bigcup_{j\ge n}E_j\in\mathcal{S}\), a decreasing sequence. For \(k\le n\), using \(|\mu_k|\le\nu_n\le\nu_j\) for \(j\ge n\),
\begin{equation*} \begin{aligned} |\mu_k|(S\backslash D_n)&\le\nu_n(S\,\triangle\,E_n)<\varepsilon2^{-n}/4,\\ |\mu_k|(D_n\backslash S)&\le\sum_{j\ge n}\nu_j(S\,\triangle\,E_j)<\varepsilon2^{-n}/2, \end{aligned} \end{equation*}
so \(|\mu_k|(S\triangle D_n)<\tfrac34\varepsilon2^{-n}\) for \(k\le n\); call this \((*)\). Moreover some \(p\) has \(|\mu_n|(D_p\backslash D_n)<\varepsilon/2\) for all \(n>p\): otherwise, taking \(p_1 := 1\), then \(n_i>p_i\) with \(|\mu_{n_i}|(D_{p_i}\backslash D_{n_i})\ge\varepsilon/2\) and \(p_{i+1} := n_i\), the sets \(H_i := D_{p_i}\backslash D_{n_i}\) are pairwise disjoint (for \(i<l\), \(H_l\subset D_{p_l}\subset D_{n_i}\) while \(H_i\cap D_{n_i} = \emptyset\)) and each is a difference of two countable unions of ring sets, so (i’) with the indices \((n_i)\) forces \(|\mu_{n_i}|(H_i)\to0\). Take \(R := D_p = \bigcup_{j\ge p}E_j\), of the required form. For \(n\le p\), \((*)\) gives \(|\mu_n|(S\triangle R)<\tfrac34\varepsilon2^{-p}<\varepsilon\), and for \(n>p\), where \(D_n\triangle D_p = D_p\backslash D_n\),
\begin{equation*} |\mu_n|(S\,\triangle\,R)\le|\mu_n|(S\,\triangle\,D_n)+|\mu_n|(D_n\,\triangle\,D_p) <\tfrac34\varepsilon2^{-n}+\tfrac{\varepsilon}{2}<\varepsilon . \end{equation*}
(Dubrovskii) Let \(\{\varphi_\alpha\}\) be a uniformly bounded family of countably additive measures on a \(\sigma\)-algebra \(\mathcal{M}\) dependent on the parameter \(\alpha\) from some set \(A\). For every sequence of disjoint sets \(E_n \in \mathcal{M}\) we let
\begin{equation*} \delta(\{E_n\}) = \lim_{n\to\infty}\Big[\sup_{\alpha\in A} |\varphi_\alpha|\Big(\bigcup_{k=n+1}^{\infty} E_k\Big)\Big]. \end{equation*}
Denote by \(\Delta\) the supremum of the numbers \(\delta(\{E_n\})\) over all possible sequences of the indicated type. Suppose that there exists a nonnegative measure \(\mu\) on \(\mathcal{M}\) such that \(\varphi_\alpha \ll \mu\) for all \(\alpha\). Set \(f_\alpha := d\varphi_\alpha/d\mu\). Prove that \(\Delta\) coincides with the quantity
\begin{equation*} \lim_{N\to\infty}\Big[\sup_{\alpha\in A}\int_{\{|f_\alpha|>N\}}\big[|f_\alpha| - N\big]\,d\mu\Big]. \end{equation*}
In particular, the latter is independent of \(\mu\).
\(\Delta = \Theta\), where \(\Theta := \lim_{N\to\infty}\Theta(N)\) and \(\Theta(N) := \sup_{\alpha}\int_X (|f_\alpha| - N)^{+} d\mu\). Here \(f_\alpha \in L^1(\mu)\) by the Radon-Nikodym theorem (Theorem 3.2.2), and uniform boundedness means
\begin{equation*} C := \sup_{\alpha \in A} \|\varphi_\alpha\| = \sup_{\alpha\in A} |\varphi_\alpha|(X) < \infty . \end{equation*}
Here \(t^{+} = \max(t,0)\). Since \(\varphi_\alpha = f_\alpha\cdot\mu\) is real, \(\{f_\alpha\ge0\}\) and \(\{f_\alpha<0\}\) realize its Hahn decomposition, so \(\varphi_\alpha^{\pm} = f_\alpha^{\pm}\cdot\mu\) and, by Definition 3.1.4,
\begin{equation*} |\varphi_\alpha|(E) = \int_E |f_\alpha|\, d\mu, \qquad E \in \mathcal{M}. \tag{a} \end{equation*}
(For complex \(\varphi_\alpha\) the same follows from \(|\varphi_\alpha|(E) = \sup\sum_i |\varphi_\alpha(E_i)|\) over finite partitions.) Both limits exist by monotonicity: \(N \mapsto (|f_\alpha|-N)^{+}\) decreases pointwise, so \(\Theta(\cdot)\) is nonincreasing and bounded by \(C\), whence \(\Theta = \inf_N \Theta(N)\) and
\begin{equation*} \Theta(N) \ge \Theta \qquad \text{for every } N \ge 0 ; \tag{b} \end{equation*}
likewise \(S_n := \bigcup_{k>n}E_k\) decreases, so \(\delta(\{E_n\}) = \inf_n \sup_\alpha|\varphi_\alpha|(S_n)\), and \(0\le\Delta\le C\). We use once the absolute continuity of the integral: given \(g \in L^1(\mu)\) and \(\eta>0\), dominated convergence (Theorem 2.8.1) gives \(N_0\) with \(\int_X(|g|-N_0)^{+}d\mu < \eta/2\), and then \(|g| \le (|g|-N_0)^{+} + N_0\) makes \(\delta := \eta/(2N_0+2)\) work.
(i) \(\Delta \le \Theta\). For disjoint \(\{E_n\}\) one has \(\bigcap_n S_n = \emptyset\), so \(\mu(S_n)\to0\) by countable additivity of the finite measure \(\mu\); and by (a) with \(|f_\alpha| \le (|f_\alpha|-N)^{+} + N\),
\begin{equation*} \sup_{\alpha}|\varphi_\alpha|(S_n) \le \Theta(N) + N\,\mu(S_n) . \end{equation*}
Let \(n\to\infty\), then \(N\to\infty\), then take the supremum over all such sequences.
(ii) \(\Theta \le \Delta\). Assume \(\Theta>0\) and fix \(\varepsilon \in (0,\Theta)\); by (b) every \(N\) admits \(\alpha\) with \(\int_X(|f_\alpha|-N)^{+}d\mu > \Theta(N)-\varepsilon \ge \Theta-\varepsilon\). Take \(N_1 \ge 1\) and such an \(\alpha_1\); having chosen \(\alpha_1,\dots,\alpha_m\) and \(N_1,\dots,N_m\), let \(\delta_j>0\) (\(j\le m\)) be furnished by absolute continuity for \(g = f_{\alpha_j}\), \(\eta = \varepsilon 2^{-j}\):
\begin{equation*} \mu(B) < \delta_j \ \Longrightarrow \ |\varphi_{\alpha_j}|(B) = \int_B |f_{\alpha_j}|\,d\mu < \varepsilon 2^{-j}. \end{equation*}
then pick \(N_{m+1}\) with \(C/N_{m+1} \le 2^{-(m+1)}\min_{j\le m}\delta_j\) and \(\alpha_{m+1}\) as above for \(N = N_{m+1}\). Put \(A_m := \{|f_{\alpha_m}| > N_m\}\). Chebyshev’s inequality and (a) give \(\mu(A_m) \le |\varphi_{\alpha_m}|(X)/N_m \le C/N_m\), so \(\mu(A_k) \le 2^{-k}\delta_m\) whenever \(k>m\), whence \(\mu(\bigcup_{k>m}A_k) \le 2^{-m}\delta_m < \delta_m\) and
\begin{equation*} |\varphi_{\alpha_m}|\Big(\bigcup_{k>m} A_k\Big) < \varepsilon 2^{-m}. \tag{1} \end{equation*}
Set \(E_m := A_m \setminus \bigcup_{k>m}A_k\); these are disjoint, since \(j<m\) forces \(E_j \cap A_m = \emptyset\) while \(E_m \subset A_m\). Fix \(n\) and take any \(m>n\). Then \(E_m \subset \bigcup_{k>n}E_k\), and (1) together with (a) and \(\mathbf{1}_{A_m}|f_{\alpha_m}| \ge (|f_{\alpha_m}|-N_m)^{+}\) yields
\begin{equation*} |\varphi_{\alpha_m}|(E_m) \ge |\varphi_{\alpha_m}|(A_m) - \varepsilon 2^{-m} > \Theta - \varepsilon - \varepsilon 2^{-m} . \end{equation*}
Hence \(\sup_\alpha |\varphi_\alpha|(\bigcup_{k>n}E_k) > \Theta - 2\varepsilon\) for every \(n\), so \(\Delta \ge \delta(\{E_n\}) \ge \Theta - 2\varepsilon\); let \(\varepsilon \downarrow 0\).
Since \(\Delta\) is defined without reference to \(\mu\), the limit on the right is the same for every dominating \(\mu\).
(M.N. Bobynin, E.H. Gohman) Suppose \(\mathcal{A}\) is a \(\sigma\)-algebra of subsets of a space \(X\), \(A_n \in \mathcal{A}\), \(A_{n+1} \subset A_n\) and \(\bigcap_{n=1}^{\infty} A_n = \emptyset\). Let \(\mu_n\) be measures on \(\mathcal{A}\) (possibly signed or complex-valued) such that \(\mu_n(A_n) \neq 0\) for all \(n\). Prove that there exists a set \(A \in \mathcal{A}\) such that one has \(|\mu_n(A)| > \frac{1}{5}|\mu_n(A_n)|\) for infinitely many indices \(n\).
Take \(A\) to be the greedy union constructed below. Writing \(\mu_n = \alpha_n + i\beta_n\) with \(\alpha_n,\beta_n\) real (and \(\beta_n = 0\) in the signed case), Corollary 3.1.3 and Definition 3.1.4 make \(\nu_n := |\alpha_n| + |\beta_n|\) a finite nonnegative measure on \(\mathcal{A}\) with
\begin{equation*} |\mu_n(B)| \le |\alpha_n(B)| + |\beta_n(B)| \le \nu_n(B) \le \nu_n( C), \qquad B \subset C. \tag{1} \end{equation*}
Since \(A_m \downarrow \emptyset\) and \(\nu_n\) is finite, continuity at zero (Proposition 1.3.3) gives \(\nu_n(A_m) \to 0\) as \(m\to\infty\) for each fixed \(n\). Set \(n_1 := 1\), \(c_k := |\mu_{n_k}(A_{n_k})| > 0\), and choose inductively \(n_{k+1} > n_k\) with
\begin{equation*} \nu_{n_k}(A_{n_{k+1}}) < \tfrac{1}{10} c_k . \tag{2} \end{equation*}
Put \(E_k := A_{n_k}\setminus A_{n_{k+1}}\); these are disjoint, since \(k<l\) gives \(E_l \subset A_{n_l} \subset A_{n_{k+1}}\), which misses \(E_k\). By additivity with (1) and (2),
\begin{equation*} |\mu_{n_k}(E_k)| \ge c_k - \nu_{n_k}(A_{n_{k+1}}) > \tfrac{9}{10} c_k . \tag{3} \end{equation*}
Define \(\varepsilon_k \in \{0,1\}\) recursively by \(\varepsilon_k := 1\) if \(|H_k| \le \tfrac12 c_k\) and \(\varepsilon_k := 0\) otherwise, where \(H_k := \sum_{j<k,\ \varepsilon_j = 1}\mu_{n_k}(E_j)\) is a finite sum depending only on \(\varepsilon_1,\dots,\varepsilon_{k-1}\); put \(A := \bigcup_{k:\,\varepsilon_k=1}E_k \in \mathcal{A}\).
Fix \(k\) and split \(A\) into the disjoint pieces \(P_k := \bigcup_{j<k,\,\varepsilon_j=1}E_j\), then \(M_k := E_k\) if \(\varepsilon_k = 1\) and \(M_k := \emptyset\) otherwise, then \(T_k := \bigcup_{j>k,\,\varepsilon_j=1}E_j\). Countable additivity gives
\begin{equation*} \mu_{n_k}(A) = H_k + \varepsilon_k \mu_{n_k}(E_k) + \mu_{n_k}(T_k) , \end{equation*}
while \(T_k \subset A_{n_{k+1}}\) with (1) and (2) gives \(|\mu_{n_k}(T_k)| < \tfrac{1}{10}c_k\). Two cases:
(i) \(\varepsilon_k = 0\), so \(|H_k| > \tfrac12 c_k\) and \(|\mu_{n_k}(A)| > \tfrac12 c_k - \tfrac1{10}c_k = \tfrac25 c_k\);
(ii) \(\varepsilon_k = 1\), so \(|H_k| \le \tfrac12 c_k\) and, by (3),
\begin{equation*} |\mu_{n_k}(A)| > \tfrac{9}{10}c_k - \tfrac12 c_k - \tfrac1{10}c_k = \tfrac{3}{10}c_k . \end{equation*}
In both cases \(|\mu_{n_k}(A)| > \tfrac{3}{10}c_k > \tfrac15 |\mu_{n_k}(A_{n_k})|\), and the \(n_k\) are infinitely many.
Exercises 4.7.140–4.7.146
Let \(\mu\) and \(\nu\) be bounded measures on a \(\sigma\)-algebra \(\mathcal{A}\). Show that
\begin{equation*} \mu \vee \nu(A) = \sup\{\mu(B) + \nu(A\setminus B) : B \in \mathcal{A},\ B \subset A\}, \qquad \forall\, A \in \mathcal{A}, \end{equation*}
\begin{equation*} \mu \wedge \nu(A) = \inf\{\mu(B) + \nu(A\setminus B) : B \in \mathcal{A},\ B \subset A\}, \qquad \forall\, A \in \mathcal{A}. \end{equation*}
Both extrema are attained, at \(B_0 := A\cap\{f \ge g\}\) and \(B_1 := A\cap\{f\le g\}\) respectively, where \(f,g\) are the densities of \(\mu,\nu\) with respect to \(\lambda := |\mu|+|\nu|\). Indeed \(\mu \ll \lambda\) and \(\nu \ll \lambda\), so the Radon-Nikodym theorem (Theorem 3.2.2) supplies \(f, g \in L^1(\lambda)\) with \(\mu = f\cdot\lambda\), \(\nu = g\cdot\lambda\), and by 4.7(i)
\begin{equation*} \mu\vee\nu(A) = \int_A \max(f,g)\, d\lambda, \qquad \mu\wedge\nu(A) = \int_A \min(f,g)\, d\lambda . \end{equation*}
For \(B \in \mathcal{A}\), \(B \subset A\), the pointwise bounds \(f, g \le \max(f,g)\) and additivity over \(A = B \sqcup (A\setminus B)\) give
\begin{equation*} \mu(B) + \nu(A\setminus B) = \int_B f\, d\lambda + \int_{A\setminus B} g\, d\lambda \le \int_A \max(f,g)\,d\lambda = \mu\vee\nu(A), \end{equation*}
with equality for \(B = B_0\), since \(\max(f,g) = f\) on \(B_0\) and \(= g\) on \(A\setminus B_0 = A\cap\{f<g\}\). Symmetrically, \(\min(f,g) \le f, g\) gives \(\mu(B)+\nu(A\setminus B) \ge \mu\wedge\nu(A)\), with equality for \(B = B_1\). (Check!)
Let \(\mu\) be a nonnegative measure and let \(f, g \in L^p(\mu)\), \(1 < p < \infty\). Show that the function
\begin{equation*} F(t) = \int |f + tg|^p \, d\mu \end{equation*}
is differentiable and
\begin{equation*} F^{\prime}(0) = p \int |f|^{p-2} f g \, d\mu . \end{equation*}
\(F\) is differentiable on all of \(\mathbb{R}^1\), with
\begin{equation*} F^{\prime}(t) = p \int |f+tg|^{p-2}(f+tg)\,g\, d\mu , \end{equation*}
which at \(t = 0\) is the asserted formula. Here \(|s|^{p-2}s\) means \(|s|^{p-1}\operatorname{sgn} s\), with value \(0\) at \(s=0\), so that \(\varphi(s) := |s|^p\) is \(C^1\) on \(\mathbb{R}^1\) with \(\varphi^{\prime}(s) = p|s|^{p-2}s\) (at \(s=0\) because \(|h|^p/h \to 0\) for \(p>1\)) and \(|\varphi^{\prime}|\) nondecreasing in \(|s|\). The right-hand side is finite: \(\big||f|^{p-2}f\big|^{p^{\prime}} = |f|^p\) with \(p^{\prime} = p/(p-1)\) puts \(|f|^{p-2}f \in L^{p^{\prime}}(\mu)\), so Hoelder’s inequality (Theorem 2.11.1) with \(g \in L^p(\mu)\) gives \(|f|^{p-2}fg \in L^1(\mu)\).
Fix \(t\) and \(0 < |h| \le 1\). The mean value theorem applied to \(s \mapsto \varphi(f(x)+tg(x)+sg(x))\) on the interval between \(0\) and \(h\) produces \(\theta = \theta(x,h) \in (0,1)\) with
\begin{equation*} \Delta_h := \frac{|f+(t+h)g|^p - |f+tg|^p}{h} = \varphi^{\prime}\big(f + (t+\theta h)g\big)\,g , \end{equation*}
whence, since \(|f + (t+\theta h)g| \le G := |f| + (|t|+1)|g|\) and \(|\varphi^{\prime}|\) is nondecreasing in the modulus of its argument,
\begin{equation*} |\Delta_h| \le p\,G^{p-1}|g| =: \Phi . \end{equation*}
Now \(G \in L^p(\mu)\) gives \(G^{p-1} \in L^{p^{\prime}}(\mu)\) with \(\|G^{p-1}\|_{p^{\prime}} = \|G\|_p^{p-1}\), so Hoelder’s inequality yields \(\int\Phi\,d\mu \le p\|G\|_p^{p-1}\|g\|_p < \infty\): one integrable majorant for all \(h\). Since \(\varphi^{\prime}\) is continuous, \(\Delta_h \to \varphi^{\prime}(f+tg)g\) pointwise, so along any sequence \(h_n \to 0\) with \(0<|h_n|\le1\) dominated convergence (Theorem 2.8.1) gives
\begin{equation*} \frac{F(t+h_n)-F(t)}{h_n} = \int \Delta_{h_n}\, d\mu \longrightarrow p\int |f+tg|^{p-2}(f+tg)\,g\, d\mu . \end{equation*}
The sequence being arbitrary, the difference quotients converge as \(h \to 0\).
Let \(1 < p < \infty\). Show that, for every \(\varepsilon > 0\), there exists \(\delta > 0\) such that if \(f, g \in L^p[0,1]\), \(\|g\|_p = 1\), \(\|f\|_p \le \delta\), and
\begin{equation*} \int f(x)\, dx = 0, \end{equation*}
then
\begin{equation*} \int\!\!\int |f(x) + g(y)|^p \, dx\, dy \le 1 + \varepsilon \|f\|_p . \end{equation*}
Take \(\delta := (\varepsilon/2K_p)^{1/(p-1)}\) if \(1 < p \le 2\) and \(\delta := \min(1, \varepsilon/2K_p)\) if \(p \ge 2\), where
\begin{equation*} K_p := \max\big(\tfrac{p(p-1)}{2}\max((3/2)^{p-2}, 2^{2-p}),\ 3^p + 2^p + p\,2^{p-1}\big). \end{equation*}
Write \(\eta := \|f\|_p\) (for \(\eta = 0\) the left side is \(\|g\|_p^p = 1\)), and put, with \(\varphi(s) = |s|^p\), so \(\varphi^{\prime}(s) = p|s|^{p-2}s\) by Exercise 4.7.141 and \(\varphi^{\prime\prime}(s) = p(p-1)|s|^{p-2}\) for \(s \ne 0\),
\begin{equation*} \psi(a,b) := |a+b|^p - |a|^p - p|a|^{p-2}a\,b . \end{equation*}
The key pointwise bound is
\begin{equation*} |\psi(a,b)| \le K_p\big(|a|^{p-2}b^2\,\mathbf{1}_{\{|b| \le |a|/2\}} + |b|^p\big). \tag{1} \end{equation*}
(i) \(|b| \le |a|/2\): every \(\xi\) on the segment from \(a\) to \(a+b\) has \(|a|/2 \le |\xi| \le \tfrac32|a|\), so Taylor’s formula with Lagrange remainder gives \(\psi(a,b) = \tfrac{p(p-1)}{2}|\xi|^{p-2}b^2\), and \(|\xi|^{p-2} \le (\tfrac32|a|)^{p-2}\) if \(p \ge 2\), \(|\xi|^{p-2} \le 2^{2-p}|a|^{p-2}\) if \(p<2\).
(ii) \(|b| > |a|/2\): then \(|a| < 2|b|\), \(|a+b| \le 3|b|\), and
\begin{equation*} |\psi(a,b)| \le 3^p|b|^p + 2^p|b|^p + p\,2^{p-1}|b|^p \le K_p|b|^p . \end{equation*}
Hence \(|\psi(a,b)| \le 2K_p|b|^p\) when \(1<p\le2\), since on \(\{|b|\le|a|/2\}\) one has \(|a|^{p-2}b^2 = |b|^p(|b|/|a|)^{2-p} \le |b|^p\); and \(|\psi(a,b)| \le K_p(|a|^{p-2}b^2 + |b|^p)\) when \(p \ge 2\), dropping the indicator.
Now set \(a = g(y)\), \(b = f(x)\) and integrate over the square. Here \(|f(x)+g(y)|^p \le 2^{p-1}(|f(x)|^p + |g(y)|^p)\) is integrable, and so is \(|g(y)|^{p-1}|f(x)|\) by Tonelli’s theorem (Theorem 3.4.5), because Hoelder’s inequality on the probability space \([0,1]\) gives \(\int|g|^{p-1}dy \le 1\) while \(\int|f|dx \le \eta\). So Fubini’s theorem (Theorem 3.4.4) with \(\int f\,dx = 0\) kills the first-order term:
\begin{equation*} \int\!\!\int |g(y)|^{p-2}g(y)f(x)\,dx\,dy = \Big(\int |g|^{p-2}g\,dy\Big)\Big(\int f\,dx\Big) = 0 , \end{equation*}
leaving
\begin{equation*} \int\!\!\int |f(x)+g(y)|^p\,dx\,dy = 1 + \int\!\!\int \psi\big(g(y),f(x)\big)\,dx\,dy. \tag{2} \end{equation*}
(i) \(1 < p \le 2\): by (2), the left side is at most \(1 + 2K_p\eta^p = 1 + 2K_p\eta^{p-1}\cdot\eta \le 1 + \varepsilon\eta\) once \(\eta \le \delta\), since \(p>1\).
(ii) \(p \ge 2\): Hoelder on \([0,1]\) with exponent \(p/(p-2)\) gives \(\int|g|^{p-2}dy \le 1\) and \(\int f^2 dx \le \eta^2\) (for \(p=2\) both are immediate), so by (2) the left side is at most
\begin{equation*} 1 + K_p(\eta^2 + \eta^p) \le 1 + 2K_p\eta\cdot\eta \le 1 + \varepsilon\eta , \end{equation*}
using \(\eta \le \delta \le 1\).
(Carlen, Loss) Let \(\mu\) be a probability measure on a space \(X\) and let \(u \in L^2(\mu)\) have unit \(L^2(\mu)\)-norm and zero integral.
(i) Prove that for every \(\alpha \in [0,1]\) and \(p \ge 2\), letting \(f = \alpha u + \sqrt{1-\alpha^2}\), one has
\begin{equation*} \|f\|_{L^p(\mu)}^p \le (1-\alpha)^{p/2} + \frac{\alpha^2 p(p-1)}{2}\,\|f\|_{L^p(\mu)}^{p-2}\,\|u\|_{L^p(\mu)}^2 , \end{equation*}
provided that \(u \in L^p(\mu)\).
(ii) Let \(u^2\ln(u^2) \in L^1(\mu)\). Prove that
\begin{equation*} \int_X f^2 \ln(f^2)\, d\mu \le 2\alpha^2 + \alpha^4 + \alpha^2 \int_X u^2\ln(u^2)\, d\mu . \end{equation*}
The printed \((1-\alpha)^{p/2}\) in (i) is a misprint for \((1-\alpha^2)^{p/2}\) (at \(p=2\) both sides must equal \(1\), which is what makes (ii) follow), so we prove
\begin{equation*} \|f\|_{p}^p \le (1-\alpha^2)^{p/2} + \frac{\alpha^2 p(p-1)}{2}\,\|f\|_{p}^{p-2}\,\|u\|_{p}^2 . \tag{1} \end{equation*}
Write \(\beta := \sqrt{1-\alpha^2}\), so \(f = \alpha u + \beta\), and use \(\int u\,d\mu = 0\), \(\int u^2 d\mu = 1\), \(\mu(X)=1\).
(i) Put \(h(t) := \int_X |\beta + tu|^p d\mu\) for \(t \in [0,1]\), finite since \(|\beta+tu| \le 1+|u| \in L^p(\mu)\). The \(t\)-independent majorants \(p(1+|u|)^{p-1}|u|\) and \(p(p-1)(1+|u|)^{p-2}u^2\) are \(\mu\)-integrable by Hoelder’s inequality (Theorem 2.11.1), so Corollary 2.8.7(ii) applies twice and \(h \in C^2[0,1]\) with
\begin{equation*} h^{\prime}(t) = p\!\int |\beta+tu|^{p-2}(\beta+tu)u\, d\mu, \quad h^{\prime\prime}(t) = p(p-1)\!\int |\beta+tu|^{p-2}u^2 d\mu \ge 0 . \end{equation*}
Thus \(h\) is convex with \(h(0) = (1-\alpha^2)^{p/2}\) and \(h^{\prime}(0) = p\beta^{p-1}\int u\,d\mu = 0\), hence nondecreasing. Taylor’s formula with Lagrange remainder yields \(\theta \in (0,\alpha)\) with
\begin{equation*} h(\alpha) = (1-\alpha^2)^{p/2} + \frac{\alpha^2 p(p-1)}{2}\int |\beta+\theta u|^{p-2}u^2\, d\mu , \end{equation*}
and Hoelder with exponents \(p/(p-2)\), \(p/2\) bounds that integral by \(h(\theta)^{(p-2)/p}\|u\|_p^2\), which is at most \(\|f\|_p^{p-2}\|u\|_p^2\) because \(h(\theta) \le h(\alpha) = \|f\|_p^p\). This is (1).
(ii) Since (1) is an equality at \(p=2\), differentiate there. Assume first \(u\) bounded and set \(\Phi(p) := \int|f|^p d\mu\), \(U(p) := \int|u|^p d\mu\), \(G(p) := \Phi(p)^{(p-2)/p}U(p)^{2/p}\) and \(R(p) := (1-\alpha^2)^{p/2} + \tfrac{\alpha^2}{2}p(p-1)G(p)\). As \(t \mapsto t^p|\ln t|\) is bounded on \([0,M]\) uniformly for \(p\) near \(2\) and \(\mu\) is finite, dominated convergence gives \(\Phi^{\prime}(p) = \int|f|^p\ln|f|\,d\mu\) and likewise for \(U\); writing \(E := \int f^2\ln(f^2)d\mu\), \(S := \int u^2\ln(u^2)d\mu\), this says \(\Phi^{\prime}(2) = E/2\), \(U^{\prime}(2) = S/2\), while \(\Phi(2) = U(2) = G(2) = R(2) = 1\). From \(\Phi \le R\) on \([2,\infty)\) with equality at \(2\), comparing difference quotients as \(p \downarrow 2\) gives \(\Phi^{\prime}(2) \le R^{\prime}(2)\) (\(R\) is differentiable at \(2\) since \(\Phi(2) = U(2) = 1 > 0\)). Now
\begin{equation*} \ln G(p) = \frac{p-2}{p}\ln\Phi(p) + \frac{2}{p}\ln U(p) \end{equation*}
and \(\ln\Phi(2) = \ln U(2) = 0\) kill every term of \((\ln G)^{\prime}(2)\) but one, leaving \((\ln G)^{\prime}(2) = U^{\prime}(2) = S/2\) and so \(G^{\prime}(2) = S/2\); with \(\tfrac{d}{dp}(1-\alpha^2)^{p/2}|_{p=2} = \tfrac12(1-\alpha^2)\ln(1-\alpha^2)\) and \(\tfrac{d}{dp}\tfrac{\alpha^2}{2}p(p-1)|_{p=2} = \tfrac32\alpha^2\),
\begin{equation*} \tfrac{E}{2} \le R^{\prime}(2) = \tfrac12(1-\alpha^2)\ln(1-\alpha^2) + \tfrac32\alpha^2 + \tfrac12\alpha^2 S . \end{equation*}
Finally \((1-s)\ln(1-s) \le -s(1-s) = s^2 - s\) at \(s = \alpha^2\) turns this into \(E \le 2\alpha^2 + \alpha^4 + \alpha^2 S\).
For general \(u\), truncate: with \(u_n := \max(-n,\min(n,u))\), \(c_n := \int u_n d\mu \to 0\) and \(\sigma_n^2 := \int(u_n-c_n)^2 d\mu \to 1\) by dominated convergence, so \(v_n := (u_n-c_n)/\sigma_n\) is bounded with \(\int v_n d\mu = 0\), \(\|v_n\|_2 = 1\), \(v_n \to u\) pointwise. Put \(\Psi(t) := t\ln t\), \(\Psi(0) := 0\), and \(f_n := \alpha v_n + \beta \to f\) pointwise. Since \(\Psi \ge -e^{-1}\) and \(\mu\) is finite, Fatou’s theorem gives
\begin{equation*} \int \Psi(f^2)\,d\mu \le \liminf_n \int \Psi(f_n^2)\, d\mu . \end{equation*}
On the other side, \(\Psi(\lambda t) = \lambda\Psi(t) + \lambda t\ln\lambda\) with \(w_n := (u_n-c_n)^2\) and \(\int \sigma_n^{-2}w_n d\mu = 1\) gives
\begin{equation*} \int \Psi(v_n^2)\, d\mu = \sigma_n^{-2}\!\int \Psi(w_n)\, d\mu + \ln(\sigma_n^{-2}) . \end{equation*}
Once \(|c_n| \le 1\) we have \(w_n \le T := (|u|+1)^2\), and \(\Psi\) increasing on \([e^{-1},\infty)\), \(\le 0\) below, gives \(\Psi(w_n) \le \Psi(T) + e^{-1}\) with \(\Psi(T) \in L^1(\mu)\) (bounded by \(4\ln4\) on \(\{|u|<1\}\), and by \(4u^2\ln 4 + 4u^2\ln(u^2)\) on \(\{|u|\ge1\}\), where \(T \le 4u^2\)). So Fatou applied to \(\Psi(T)+e^{-1}-\Psi(w_n) \ge 0\) gives \(\limsup_n \int\Psi(w_n)d\mu \le S\); these integrals being bounded above, with \(\sigma_n^{-2} \to 1\) and \(\ln(\sigma_n^{-2}) \to 0\), also \(\limsup_n \int\Psi(v_n^2)d\mu \le S\). Taking \(\liminf_n\) in the bounded-case inequality for \(v_n\) now gives the claim.
(i) (Clarkson) Prove the following inequalities for \(f,g \in L^p(\mu)\):
\begin{equation*} \left\|\frac{f+g}{2}\right\|_p^p + \left\|\frac{f-g}{2}\right\|_p^p \le \frac{1}{2}\|f\|_p^p + \frac{1}{2}\|g\|_p^p, \qquad 2 \le p < \infty, \end{equation*}
\begin{equation*} \left\|\frac{f+g}{2}\right\|_p^{p^{\prime}} + \left\|\frac{f-g}{2}\right\|_p^{p^{\prime}} \le \left[\frac{1}{2}\|f\|_p^p + \frac{1}{2}\|g\|_p^p\right]^{\frac{1}{p-1}}, \qquad 1 < p \le 2,\ p^{\prime} = \frac{p}{p-1}. \end{equation*}
(ii) (Hanner) Prove the following inequalities for \(f,g \in L^p(\mu)\), where \(1 \le p \le 2\):
\begin{equation*} \|f+g\|_p^p + \|f-g\|_p^p \ge \left(\|f\|_p + \|g\|_p\right)^p + \bigl|\,\|f\|_p - \|g\|_p\,\bigr|^p, \end{equation*}
\begin{equation*} \left(\|f+g\|_p + \|f-g\|_p\right)^p + \bigl|\,\|f+g\|_p - \|f-g\|_p\,\bigr|^p \le 2^p\left(\|f\|_p + \|g\|_p\right)^p. \end{equation*}
Prove the reversed inequalities in the case \(2 \le p < \infty\).
(i) For \(2 \le p < \infty\) the inequality is pointwise. The parallelogram identity gives \(|\tfrac{a+b}{2}|^2 + |\tfrac{a-b}{2}|^2 = \tfrac12(|a|^2+|b|^2)\), so with \(r = p/2 \ge 1\) the elementary bounds \(x^r + y^r \le (x+y)^r\) (divide by \((x+y)^r\) and use \(u^r \le u\) on \([0,1]\)) and \((\tfrac{x+y}{2})^r \le \tfrac{x^r+y^r}{2}\) (convexity of \(t^r\)) yield
\begin{equation*} \Big|\frac{a+b}{2}\Big|^{p} + \Big|\frac{a-b}{2}\Big|^{p} \le \Big(\frac{|a|^2+|b|^2}{2}\Big)^{p/2} \le \frac{|a|^{p}+|b|^{p}}{2} . \end{equation*}
Substitute \(a = f(x)\), \(b = g(x)\) and integrate; all four functions lie in \(L^1(\mu)\) by Minkowski’s inequality (Theorem 2.11.9).
For \(1 < p \le 2\) and \(q := p^{\prime} \ge 2\), the scalar inequality
\begin{equation*} \big(|a+b|^{q} + |a-b|^{q}\big)^{1/q} \le 2^{1/q}\big(|a|^{p}+|b|^{p}\big)^{1/p} \tag{A} \end{equation*}
holds by Riesz-Thorin interpolation for \(T(a,b) := (a+b,\,a-b)\) on the two-point space: \(\|T\|_{\ell^1\to\ell^\infty} \le 1\) since \(\max(|a+b|,|a-b|) \le |a|+|b|\), and \(\|T\|_{\ell^2\to\ell^2} = 2^{1/2}\) by the parallelogram identity, while \(\theta := 2/q\) produces exactly the exponent pair \((p,q)\). (Riesz-Thorin is not proved in the book; see the comments to Chapter 4.) We also use, for \(r \ge 1\) and nonnegative integrable \(\varphi_1,\varphi_2\) with \(I_j := \int\varphi_j\,d\mu\),
\begin{equation*} \big(I_1^{r} + I_2^{r}\big)^{1/r} \le \int_X \big(\varphi_1^{r} + \varphi_2^{r}\big)^{1/r} d\mu , \tag{B} \end{equation*}
which follows by integrating the planar Hoelder bound for \(c_1\varphi_1 + c_2\varphi_2\) with \(c_j := I_j^{\,r-1}(I_1^r+I_2^r)^{-1/r^{\prime}}\), for which \(c_1^{r^{\prime}}+c_2^{r^{\prime}} = 1\) and \(c_1I_1+c_2I_2 = (I_1^r+I_2^r)^{1/r}\). Put \(F := (f+g)/2\), \(G := (f-g)/2\). Raising (A) to the power \(q\), dividing by \(2^{q}\) and raising to the power \(p-1\), with \(q(p-1) = p\), gives pointwise
\begin{equation*} \big(|F|^{q} + |G|^{q}\big)^{p-1} \le \frac{|f|^{p}+|g|^{p}}{2} , \end{equation*}
while (B) with \(r = 1/(p-1) \ge 1\) and \(\varphi_1 = |F|^{p}\), \(\varphi_2 = |G|^{p}\), so that \(pr = q\), gives
\begin{equation*} \big(\|F\|_p^{q} + \|G\|_p^{q}\big)^{p-1} \le \int_X \big(|F|^{q}+|G|^{q}\big)^{p-1} d\mu . \end{equation*}
Combining the two and raising to the power \(1/(p-1)\) is the second Clarkson inequality.
(ii) The printed right-hand side of the second inequality is a misprint for \(2^{p}(\|f\|_p^p + \|g\|_p^p)\); we prove that sharp form (H2), which implies the printed one because \(x^p+y^p \le (x+y)^p\), and whose reversal is the correct reversed statement for \(p \ge 2\) (the printed inequality reversed already fails for \(g = f\), \(\|f\|_p = 1\), where it reads \(2^{p+1} \ge 4^{p}\)). Write (H1) for the first inequality. Then (H1) and (H2) are equivalent, and so are their reversals: applying (H1) to \(F = f+g\), \(G = f-g\), whose sum and difference are \(2f\) and \(2g\), gives (H2), and applying (H2) to \((f+g)/2\), \((f-g)/2\) gives (H1). So only (H1) is at issue.
Pointwise, for scalars \(u,v\) and \(1 \le p \le 2\),
\begin{equation*} |u+v|^{p} + |u-v|^{p} \ge \big(|u|+|v|\big)^{p} + \big|\,|u|-|v|\,\big|^{p} , \tag{C} \end{equation*}
reversed for \(p \ge 2\): with \(S := |u|^2+|v|^2\) and \(D := 2\operatorname{Re}(u\bar v)\), so \(|u\pm v|^2 = S \pm D\) and \(|D| \le 2|u||v| \le S\), the even function \(\psi(D) := (S+D)^{p/2}+(S-D)^{p/2}\) has
\begin{equation*} \psi^{\prime}(D) = \tfrac{p}{2}\big[(S+D)^{p/2-1} - (S-D)^{p/2-1}\big], \qquad 0<D<S, \end{equation*}
which is \(\le 0\) for \(p \le 2\) and \(\ge 0\) for \(p \ge 2\); hence \(\psi(D) \ge \psi(2|u||v|)\), the right-hand side of (C).
For \(p = 1\), the pointwise bound \(|f+g|+|f-g| \ge 2\max(|f|,|g|)\) integrates to (H1). For \(1 < p \le 2\) we may assume \(A := \|f\|_p \ge B := \|g\|_p > 0\) (the case \(B = 0\) is an identity); put \(r := B/A \in (0,1]\) and
\begin{equation*} \alpha( r) := (1+r)^{p-1}+(1-r)^{p-1}, \qquad \beta( r) := \frac{(1+r)^{p-1}-(1-r)^{p-1}}{r^{p-1}} , \end{equation*}
so that \(\alpha( r) + \beta( r)r^{p} = (1+r)^{p} + (1-r)^{p}\). Then, for all \(x,y \ge 0\),
\begin{equation*} \alpha( r)x^{p} + \beta( r)y^{p} \le (x+y)^{p} + |x-y|^{p} . \tag{D} \end{equation*}
Indeed, by homogeneity take \(x = 1\), \(y = s \ge 0\) and set \(h(s) := (1+s)^{p}+|1-s|^{p}-\alpha( r)-\beta( r)s^{p}\), which is \(C^1\) on \((0,\infty)\) since \(p>1\), with
\begin{equation*} h^{\prime}(s) = p\,s^{p-1}\big[\Phi(1/s) - \Phi(1/r)\big], \qquad \Phi(w) := (1+w)^{p-1} - \operatorname{sgn}(w-1)|w-1|^{p-1}, \end{equation*}
because \(s^{p-1}\Phi(1/s) = (1+s)^{p-1}-\operatorname{sgn}(1-s)|1-s|^{p-1}\) and \(\Phi(1/r) = \beta( r)\). As \(p \le 2\) makes \(\Phi^{\prime}(w) = (p-1)[(1+w)^{p-2}-|w-1|^{p-2}] \le 0\), the map \(s \mapsto \Phi(1/s)\) is nondecreasing, so \(h\) is minimal at \(s = r\), where the identity above gives \(h( r) = 0\). Applying (D) with \(x = |f|\), \(y = |g|\), then (C), and integrating,
\begin{equation*} \alpha( r)A^{p} + \beta( r)B^{p} \le \|f+g\|_p^{p} + \|f-g\|_p^{p} , \end{equation*}
whose left-hand side equals \(A^{p}[\alpha( r)+\beta( r)r^{p}] = (A+B)^{p}+(A-B)^{p}\) since \(B = rA\). This is (H1) for \(1 \le p \le 2\).
For \(2 \le p < \infty\) put \(q := p^{\prime} \in (1,2]\), \(A := \|f\|_p\), \(B := \|g\|_p\), \(R := \|f+g\|_p\), \(S := \|f-g\|_p\), and assume \(R^p+S^p>0\). The norming function \(u := \overline{\operatorname{sgn}(f+g)}\,|f+g|^{p-1}R^{1-p}\) (and \(u := 0\) if \(R = 0\)), together with \(v\) defined the same way from \(f-g\) and \(S\), satisfies \(\|u\|_q, \|v\|_q \le 1\), \(\int(f+g)u\,d\mu = R\) and \(\int(f-g)v\,d\mu = S\), by \((p-1)q = p\). Fix \(\lambda,\nu \ge 0\) and set \(X := \|\lambda u+\nu v\|_q\), \(Y := \|\lambda u - \nu v\|_q\). Since \((f+g)\lambda u + (f-g)\nu v = f(\lambda u+\nu v) + g(\lambda u-\nu v)\), Hoelder’s inequality (Theorem 2.11.1) gives \(\lambda R + \nu S \le AX+BY\), and planar Hoelder applied to
\begin{equation*} AX+BY = \tfrac12\big[(A+B)(X+Y) + (A-B)(X-Y)\big] \end{equation*}
bounds this by \(\tfrac12[(A+B)^{p}+|A-B|^{p}]^{1/p}[(X+Y)^{q}+|X-Y|^{q}]^{1/q}\). By (H2) for the exponent \(q \in (1,2]\), applied to \(\lambda u\) and \(\nu v\), the last bracket is at most \(2^{q}(\lambda^{q}+\nu^{q})\), so
\begin{equation*} \lambda R + \nu S \le \big[(A+B)^{p}+|A-B|^{p}\big]^{1/p}\big(\lambda^{q}+\nu^{q}\big)^{1/q} . \end{equation*}
Take \(\lambda := R^{p-1}\), \(\nu := S^{p-1}\), so that \(\lambda R+\nu S = R^{p}+S^{p}\) and \((\lambda^q+\nu^q)^{1/q} = (R^p+S^p)^{1/q}\); dividing by \((R^p+S^p)^{1/q}>0\) and using \(1-1/q = 1/p\) gives the reversal of (H1), and with it, by the equivalence, the reversal of (H2).
(Douglas) Suppose that \((X,\mathcal{A})\) is a measurable space, \(M^+(\mathcal{A})\) is the set of all finite nonnegative measures on \(\mathcal{A}\), \(F\) is some linear space of real \(\mathcal{A}\)-measurable functions. Let \(\mu \in M^+(\mathcal{A})\) and \(F \subset L^1(\mu)\). Set
\begin{equation*} E^\mu := \Big\{\nu \in M^+(\mathcal{A}) : F \subset L^1(\nu),\ \int f\, d\nu = \int f\, d\mu \ \text{ for all } f \in F\Big\}. \end{equation*}
(i) Prove that \(F\) is dense in \(L^1(\mu)\) precisely when \(\mu\) is an extreme point in \(E^\mu\), i.e., there are no measures \(\mu_1,\mu_2 \in E^\mu\) and \(t \in (0,1)\) such that one has \(\mu_1 \neq \mu\) and \(\mu = t\mu_1 + (1-t)\mu_2\).
(ii) Let \(\mathcal{B}\) be a sub-\(\sigma\)-algebra in \(\mathcal{A}\). Prove that \(\mu\) is an extreme point in the set of all measures \(\nu \in M^+(\mathcal{A})\) such that \(\nu|_{\mathcal{B}} = \mu|_{\mathcal{B}}\) precisely when, for every \(A \in \mathcal{A}\), there exists \(B \in \mathcal{B}\) with \(\mu(A \bigtriangleup B) = 0\).
(i) Not dense implies not extreme: if \(\overline{F} \neq L^1(\mu)\), the Hahn-Banach theorem and the duality \(L^1(\mu)^* = L^\infty(\mu)\) (Theorem 4.4.1, \(\mu\) being finite) give \(g \in L^\infty(\mu)\), \(g \neq 0\), with \(\int fg\,d\mu = 0\) for all \(f \in F\); after scaling, \(|g| \le 1\) everywhere. Then \(\mu_{1,2} := (1 \pm g)\cdot\mu\) are finite nonnegative measures, \(F \subset L^1(\mu_i)\) since \(|f|(1\pm g) \le 2|f|\), and
\begin{equation*} \int f\, d\mu_{1,2} = \int f\, d\mu \pm \int fg\, d\mu = \int f\, d\mu , \end{equation*}
so \(\mu_1,\mu_2 \in E^\mu\) with \(\mu = \tfrac12(\mu_1+\mu_2)\); and \(\mu_1 \neq \mu\), because one of \(\{g>0\}\), \(\{g<0\}\) has positive \(\mu\)-measure and \(\mu_1(A) - \mu(A) = \int_A g\,d\mu \neq 0\) for it.
Dense implies extreme: let \(\mu = t\mu_1 + (1-t)\mu_2\) with \(t \in (0,1)\), \(\mu_i \in E^\mu\). Then \(t\mu_1 \le \mu\), so \(\mu_1 \ll \mu\) and the Radon-Nikodym theorem (Theorem 3.2.2) gives \(\mu_1 = g_1\cdot\mu\) with \(0 \le g_1 \le t^{-1}\) \(\mu\)-a.e., since integrating over \(\{g_1 > t^{-1}\}\) or \(\{g_1<0\}\) would violate \(\mu_1 \le t^{-1}\mu\); in particular \(g_1 \in L^\infty(\mu)\). The identity \(\int f\,d\mu_1 = \int fg_1\,d\mu\) holds for indicators, hence for simple functions, hence for all \(f \in L^1(\mu_1) \supset F\) by density of simple functions (Lemma 4.2.1). So
\begin{equation*} \int f\,(g_1-1)\, d\mu = \int f\, d\mu_1 - \int f\, d\mu = 0, \qquad f \in F, \end{equation*}
and the continuous functional \(h \mapsto \int h(g_1-1)d\mu\) on \(L^1(\mu)\) vanishes on the dense set \(F\), hence identically; \(h = \operatorname{sgn}(g_1-1)\) gives \(\int|g_1-1|d\mu = 0\), i.e. \(\mu_1 = \mu\).
(ii) Take for \(F\) the bounded \(\mathcal{B}\)-measurable functions, which lie in \(L^1\) of every finite measure. Then
\begin{equation*} E^\mu = \{\nu \in M^+(\mathcal{A}) : \nu|_{\mathcal{B}} = \mu|_{\mathcal{B}}\} , \end{equation*}
since \(f = \mathbf{1}_B\) gives one inclusion, and uniform approximation by \(\mathcal{B}\)-simple functions (the measures being finite of equal total mass) gives the other. By (i) the assertion reduces to: \(F\) is dense in \(L^1(\mu)\) if and only if every \(A \in \mathcal{A}\) has a \(B \in \mathcal{B}\) with \(\mu(A \bigtriangleup B) = 0\).
(a) If such \(B\) always exist, then \(\mathbf{1}_A = \mathbf{1}_B\) \(\mu\)-a.e., so every \(\mathcal{A}\)-simple function agrees \(\mu\)-a.e. with a \(\mathcal{B}\)-simple one, which lies in \(F\); simple functions are dense in \(L^1(\mu)\) (Lemma 4.2.1).
(b) Conversely, given \(A \in \mathcal{A}\) choose \(h_n \in F\) with \(h_n \to \mathbf{1}_A\) in \(L^1(\mu)\) and, along a subsequence, \(\mu\)-a.e. Then \(\varphi := \limsup_n h_n\) (set to \(0\) where that limit superior is infinite) is \(\mathcal{B}\)-measurable and equals \(\mathbf{1}_A\) \(\mu\)-a.e., so \(B := \{\varphi \ge 1/2\} \in \mathcal{B}\) satisfies \(A \bigtriangleup B \subset \{\varphi \neq \mathbf{1}_A\}\), a \(\mu\)-null set.
Let \(X\) be an infinite-dimensional normed space. (i) Prove that the weak topology on any ball is strictly weaker than the norm topology. (ii) Prove that \(X\) with the weak topology is not metrizable.
Both assertions come from one observation: every basic weak neighbourhood
\begin{equation*} U(x_0; l_1,\dots,l_k;\varepsilon) := \{x : |l_i(x-x_0)| < \varepsilon,\ i \le k\} \end{equation*}
contains the whole affine subspace \(x_0 + N\), where \(N := \bigcap_{i\le k}\ker l_i \neq \{0\}\), since otherwise \(x \mapsto (l_1(x),\dots,l_k(x))\) would embed \(X\) linearly into \(\mathbb{R}^k\). Each \(l \in X^*\) being norm-continuous, the weak topology is contained in the norm topology (also on subsets), so only strictness is at issue.
(i) Let \(B := \{\|x-x_0\| \le R\}\). The relatively norm-open set \(V := \{\|x-x_0\|<R\}\) is not weakly open in \(B\): a basic \(W = U(x_0;l_1,\dots,l_k;\varepsilon)\) with \(W \cap B \subset V\) is contradicted by \(x_0 + y\), where \(y \in N\setminus\{0\}\) is rescaled to \(\|y\| = R\), so that \(x_0+y \in W \cap B\) but \(x_0 + y \notin V\). For an open ball \(\{\|x-x_0\|<R\}\), argue the same way with \(\{\|x-x_0\|<R/2\}\) and \(\|y\| = 3R/4\).
(ii) Suppose the weak topology were generated by a metric \(d\), and choose for each \(n\) a basic \(U(0;l_{n,1},\dots,l_{n,k_n};\varepsilon_n) \subset \{x : d(x,0)<n^{-1}\}\). Given \(l \in X^*\), the weak neighbourhood \(\{|l|<1\}\) contains some \(d\)-ball, hence some \(U(0;l_{n,1},\dots,l_{n,k_n};\varepsilon_n)\), hence the subspace \(N_n := \bigcap_{i \le k_n}\ker l_{n,i}\); as \(N_n\) is invariant under all scalars, \(l \equiv 0\) on \(N_n\). With \(Tx := (l_{n,1}(x),\dots,l_{n,k_n}(x))\), the equality \(Tx = Ty\) then forces \(l(x) = l(y)\), so \(l = \lambda \circ T\) for a linear \(\lambda\) on \(T(X)\), and extending \(\lambda\) to \(\mathbb{R}^{k_n}\) gives \(l = \sum_{i\le k_n} c_i l_{n,i}\). Hence
\begin{equation*} X^* = \bigcup_{n=1}^\infty L_n, \qquad L_n := \operatorname{span}\{l_{m,i} : m \le n,\ i \le k_m\} , \end{equation*}
with each \(L_n\) finite-dimensional, hence closed. But \(X^*\) is an infinite-dimensional Banach space: for independent \(x_1,\dots,x_m\) the Hahn-Banach theorem yields \(f_i \in X^*\) with \(f_i(x_j) = \delta_{ij}\), and such families exist for every \(m\). So each \(L_n\) is a proper closed subspace and therefore nowhere dense (a subspace with interior contains a ball at \(0\), hence everything), and \(X^*\) is a countable union of nowhere dense sets, contradicting Baire’s theorem.
Exercises 4.7.147–4.7.153
Prove that every weakly compact set in \(l^1\) is norm compact.
Realize \(l^1\) as \(L^1(\mu)\) with \(\mu\) the counting measure on \(\mathbb{N}\), so \(L^\infty(\mu) = l^\infty = (l^1)^*\) and the weak topology is exactly \(\sigma(L^1(\mu),L^\infty(\mu))\). A weakly compact \(K\) equals its own weak closure, so condition (iv) of Theorem 4.7.20 holds: \(K\) is norm bounded and, for each \(\varepsilon>0\), there is a set of finite measure, i.e. a finite \(X_\varepsilon \subset \mathbb{N}\), outside which the tails are uniformly small; with \(N := \max X_\varepsilon\),
\begin{equation*} \sum_{n>N}|x_n| < \varepsilon \qquad \text{for all } x = (x_n) \in K. \tag{\(\ast\)} \end{equation*}
(Uniform absolute continuity of the integrals says nothing here, since \(\mu(A)<1\) forces \(A = \emptyset\).)
Given \(\varepsilon>0\), take \(N\) from \((\ast)\) for \(\varepsilon/3\) and let \(P_Nx := (x_1,\dots,x_N,0,0,\dots)\). Then \(P_N(K)\) is a bounded subset of an \(N\)-dimensional subspace, hence totally bounded: there are \(y^{(1)},\dots,y^{(m)}\) such that every \(P_Nx\) lies within \(\varepsilon/3\) of some \(y^{(j)}\), and then
\begin{equation*} \|x-y^{(j)}\| \le \|x - P_Nx\| + \|P_Nx - y^{(j)}\| < \tfrac{\varepsilon}{3} + \tfrac{\varepsilon}{3} < \varepsilon . \end{equation*}
So \(K\) is totally bounded in norm. Being weakly compact it is weakly closed (the weak topology is Hausdorff), hence norm closed, and a closed totally bounded subset of the complete space \(l^1\) is norm compact.
(i) Let \(\mu\) be a separable finite nonnegative measure. Show that every uniformly integrable subset of \(L^1(\mu)\) is metrizable in the weak topology. In particular, every weakly compact subset of \(L^1(\mu)\) is metrizable in the weak topology.
(ii) Let \(\mathcal{A}\) be a countably generated \(\sigma\)-algebra. Show that every compact subset of the space \(\mathcal{M}\) of all bounded measures on \(\mathcal{A}\) with the setwise convergence topology is metrizable.
Both parts follow from one lemma: a compact Hausdorff space \(K\) carrying a countable family of continuous \(h_n : K \to \mathbb{R}^1\) that separates its points is metrizable, along with every subset. Indeed \(h = (h_n)_{n\ge1} : K \to \mathbb{R}^{\mathbb{N}}\) is continuous and injective, \(\mathbb{R}^{\mathbb{N}}\) is metrized by \(\rho(a,b) = \sum_n 2^{-n}\min(1,|a_n-b_n|)\), and a continuous injection of a compact space into a Hausdorff space is a homeomorphism onto its image. (This is Exercise 6.10.24.)
(i) Let \(M \subset L^1(\mu)\) be uniformly integrable. Its closure \(K\) in \(\sigma(L^1(\mu),L^\infty(\mu))\) is compact by Theorem 4.7.18, and the topology is Hausdorff, since \(\int fg\,d\mu = 0\) for all bounded \(g\) gives \(f = 0\) a.e. on taking \(g = \operatorname{sign}f\). By separability of \(\mu\) (see 1.12(iii)) pick \(A_n \in \mathcal{A}\) such that every \(A \in \mathcal{A}\) has \(\mu(A \bigtriangleup A_n)\) arbitrarily small, and put
\begin{equation*} h_n(f) := \int_X f\,I_{A_n}\, d\mu = \int_{A_n} f\, d\mu , \end{equation*}
weakly continuous by the definition of the topology. They separate points: if \(h := f-g\) has \(\int_{A_n}h\,d\mu = 0\) for all \(n\), then for \(A \in \mathcal{A}\) and \(\varepsilon>0\), absolute continuity of \(\int|h|d\mu\) gives \(\delta>0\) with \(\int_S|h|d\mu<\varepsilon\) whenever \(\mu(S)<\delta\), and choosing \(n\) with \(\mu(A \bigtriangleup A_n)<\delta\),
\begin{equation*} \Big|\int_A h\, d\mu\Big| = \Big|\int_A h\,d\mu - \int_{A_n}h\,d\mu\Big| \le \int_{A\bigtriangleup A_n}|h|\,d\mu < \varepsilon , \end{equation*}
so \(h = 0\) a.e. The lemma metrizes \(K\), hence \(M\). If \(M\) is weakly compact it is its own weak closure, hence uniformly integrable by Theorem 4.7.18.
(ii) On \(\mathcal{M}\) with the setwise convergence topology of 4.7(v) each \(\mu \mapsto \mu(B)\), \(B \in \mathcal{A}\), is continuous, and these separate points, so the topology is Hausdorff. Let \(\mathcal{A}_0\) be the algebra generated by countably many generators \(A_n\) of \(\mathcal{A}\); it is countable, being the union of the finite algebras generated by \(A_1,\dots,A_m\). Enumerate \(\mathcal{A}_0 = \{B_1,B_2,\dots\}\) and put \(h_n(\mu) := \mu(B_n)\). These separate points of \(\mathcal{M}\): if \(\lambda := \mu-\nu\) vanishes on \(\mathcal{A}_0\), then \(\{A : \lambda(A) = 0\}\) is a monotone class (countable additivity and boundedness of \(\lambda\) handle \(A_k \uparrow A\) and \(A_k \downarrow A\)) containing the algebra \(\mathcal{A}_0\), hence contains \(\sigma(\mathcal{A}_0) = \mathcal{A}\) by the monotone class theorem (Theorem 1.9.3). The lemma applies to any compact \(K \subset \mathcal{M}\).
(i) Let \(f\in L^2(\mathbb{R}^1)\). Show that the set \(\mathcal{F}\) of all functions of the form \(\sum_{k=1}^{n}c_kf(x+\delta_k)\), where \(n\in\mathbb{N}\), \(c_k,\delta_k\in\mathbb{R}^1\), is everywhere dense in \(L^2(\mathbb{R}^1)\) precisely when the set of zeros of the Fourier transform of \(f\) has measure zero.
(ii) Let \(f\in L^1(\mathbb{R}^1)\). Show that the set \(\mathcal{F}\) indicated in (i) is everywhere dense in \(L^1(\mathbb{R}^1)\) precisely when the Fourier transform of the function \(f\) does not vanish.
(iii) (Segal) Show that if \(1<p<2\), then the a.e. positivity of the Fourier transform of a function \(f\in L^p(\mathbb{R}^1)\) does not imply that the set indicated in (i) is everywhere dense in \(L^p(\mathbb{R}^1)\).
Throughout, \(\widehat{f}\) is the Fourier transform of Definition 3.8.1 and \(Ff := (2\pi)^{1/2}\widehat f\) the unnormalized one, so that \(F(f*g) = Ff\cdot Fg\); since \(\widehat{f(\cdot+\delta)}(y) = e^{i\delta y}\widehat f(y)\), every \(h = \sum_{k\le n}c_kf(\cdot+\delta_k)\) has \(\widehat h = (\sum_k c_ke^{i\delta_ky})\widehat f\). We use complex scalars; for real \(f\) and real \(c_k\) the density statements are equivalent, since taking real parts of an approximating sum does not increase the error, and a nonzero complex annihilator has a nonzero real or imaginary part.
(i) Density holds exactly when \(\lambda(Z) = 0\) for \(Z := \{\widehat f = 0\}\), \(\lambda\) Lebesgue measure. If \(\lambda(Z)>0\), pick \(A \subset Z\) with \(0<\lambda(A)<\infty\); by Plancherel’s theorem (the transform is unitary on \(L^2\); see (3.8.3) and Exercise 3.10.76) there is \(g\) with \(\widehat g = I_A\) and \(\|g\|_2 > 0\), and for every \(h \in \mathcal{F}\), whose transform vanishes a.e. on \(A \subset Z\),
\begin{equation*} \int_{\mathbb{R}} h\,\overline{g}\, dx = \int_{\mathbb{R}} \widehat h\,\overline{\widehat g}\, dy = \int_A \widehat h\, dy = 0 , \end{equation*}
so the closure of the subspace \(\mathcal{F}\) is proper. Conversely, if \(\lambda(Z)=0\) and \(\mathcal{F}\) were not dense, some \(g \ne 0\) is orthogonal to it, so Plancherel gives
\begin{equation*} 0 = \int_{\mathbb{R}} e^{i\delta y}\,\widehat f(y)\overline{\widehat g(y)}\, dy , \qquad \delta \in \mathbb{R}^1 , \end{equation*}
i.e. the transform of \(\Phi := \widehat f\,\overline{\widehat g} \in L^1\) (a product of two \(L^2\) functions) vanishes identically; then \(\Phi = 0\) a.e. by Proposition 3.8.6, forcing \(\widehat g = 0\) a.e. and \(g = 0\).
(ii) Density holds exactly when \(Ff\) vanishes nowhere. If \(Ff(y_0) = 0\), the nonzero \(g(x) := e^{-iy_0x} \in L^\infty(\mathbb{R}^1) = (L^1(\mathbb{R}^1))^*\) annihilates \(\mathcal{F}\), since
\begin{equation*} \int_{\mathbb{R}} f(x+\delta)e^{-iy_0x}\, dx = e^{iy_0\delta}Ff(y_0) = 0 \end{equation*}
(for real scalars use \(\cos(y_0x)\) and \(\sin(y_0x)\)). The converse is Wiener’s Tauberian theorem; let \(I\) be the \(L^1\)-closure of \(\mathcal{F}\) and assume \(Ff\) nowhere zero.
(a) \(I\) is an ideal: \(y \mapsto \tau_yf\) is \(L^1\)-continuous (Lemma 4.2.3), so for a step function \(g = \sum_jc_jI_{B_j}\) with bounded intervals \(B_j\) the integrals \(\int_{B_j}\tau_yf\,dy\) are \(L^1\)-limits of Riemann sums belonging to \(\mathcal{F}\), whence \(g*f \in I\); and \(\|g_k*f-g*f\|_1 \le \|g_k-g\|_1\|f\|_1\) extends this to all \(g \in L^1\).
(b) Functions with compactly supported transform are dense in \(L^1\): fix \(\eta \in C_0^\infty\) with \(0\le\eta\le1\), \(\eta=1\) on \([-1,1]\), \(\eta=0\) off \([-2,2]\); its inverse transform \(h_1\) is Schwartz with \(Fh_1 = \eta\) (Proposition 3.8.5, Corollary 3.8.12), and \(h_\varepsilon(x) := \varepsilon h_1(\varepsilon x)\) satisfies \(Fh_\varepsilon(t) = \eta(t/\varepsilon)\), \(\|h_\varepsilon\|_1 = \|h_1\|_1\). As \(\int h_1\,dx = Fh_1(0) = 1\), Theorem 4.2.4 gives \(h_{1/\lambda}*g \to g\) in \(L^1\), while \(F(h_\varepsilon*g) = Fh_\varepsilon\,Fg\) has compact support.
(c) Local inversion: if \(Ff(t_0) \ne 0\) there are \(u \in L^1\) and \(\varepsilon>0\) with \(Fu = 1/Ff\) on \((t_0-\varepsilon,t_0+\varepsilon)\). For \(t_0 = 0\) and \(Ff(0) = \int f\,dx = 1\), put \(w_\varepsilon := h_\varepsilon - f*h_\varepsilon\), so that \(w_\varepsilon(x) = \int f(y)[h_\varepsilon(x)-h_\varepsilon(x-y)]\,dy\) and
\begin{equation*} \|w_\varepsilon\|_1 \le \int |f(y)|\,\|h_1 - \tau_{\varepsilon y}h_1\|_1\, dy \longrightarrow 0 \end{equation*}
by continuity of translation in \(L^1\) and dominated convergence with majorant \(2\|h_1\|_1|f|\). Fix \(w := w_\varepsilon\) with \(\|w\|_1<1/2\). In the unitization \(A_1 := L^1(\mathbb{R})\oplus\mathbb{C}\delta\) of the Banach algebra \((L^1,*)\), the element \(\delta - w\) is invertible with inverse \(\delta+v\), \(v = \sum_{n\ge1}w^{*n} \in L^1\), and extending \(F\) by \(F\delta \equiv 1\) keeps it multiplicative, so \((1-Fw)(1+Fv) = 1\). Hence \(u := h_\varepsilon + h_\varepsilon*v \in L^1\) has
\begin{equation*} Fu = \frac{Fh_\varepsilon}{1-Fw} = \frac{Fh_\varepsilon}{1-Fh_\varepsilon(1-Ff)} = \frac{1}{Ff} \quad \text{on } [-\varepsilon,\varepsilon], \end{equation*}
where \(Fh_\varepsilon = 1\). For general \(t_0\) with \(c := Ff(t_0) \ne 0\), apply this to \(f_{t_0}(x) := c^{-1}e^{-it_0x}f(x)\), whose transform at \(0\) is \(1\), and set \(u(x) := c^{-1}e^{it_0x}u_0(x)\).
(d) For compact \(K\), continuity and nonvanishing of \(Ff\) give finitely many open intervals \(U_j\) covering \(K\) and \(u_j \in L^1\) with \(Fu_j = 1/Ff\) on \(U_j\); take a smooth partition of unity \(\psi_j = Fp_j\) (as in (b)) with \(\operatorname{supp}\psi_j \subset U_j\) and \(\sum_j\psi_j = 1\) on an open \(W \supset K\), and put \(u := \sum_j p_j*u_j\), so that \(Fu = \sum_j\psi_jFu_j = 1/Ff\) on \(W\). If \(Fg\) is supported in \(K\), then
\begin{equation*} F\big((g*u)*f\big) = Fg\cdot Fu\cdot Ff = Fg \end{equation*}
(on \(W\) because \(Fu\,Ff = 1\) there, off \(K\) because \(Fg = 0\)), so \(g = (g*u)*f \in I\) by Proposition 3.8.6 and (a). By (b) such \(g\) are dense and \(I\) is closed, so \(I = L^1(\mathbb{R}^1)\).
(iii) No. Fix \(1<p<2\), \(p^{\prime} = p/(p-1) > 2\), choose \(2/p^{\prime} < d < 1\) and \(1/p^{\prime} < \beta < d/2\). By Salem’s theorem (R. Salem, Ark. Mat. 1951; Kahane, Some random series of functions, Ch. 17) there is a compact \(E_0\) of Hausdorff dimension \(d\) carrying a Borel probability measure \(\mu_0\) with \(|\widetilde{\mu_0}(t)| \le C(1+|t|)^{-\beta}\); passing to \(E := E_0\cup(-E_0)\) and \(\mu := \tfrac12(\mu_0+\check\mu_0)\) we may assume \(E = -E\) and that \(g := \widetilde\mu\) is real and even, with the same decay. Then \(\lambda(E) = 0\) since \(d<1\), and \(\beta p^{\prime}>1\) gives \(g \in L^{p^{\prime}}(\mathbb{R}^1) = (L^p(\mathbb{R}^1))^*\), \(g \ne 0\) because \(g(0)=1\). By Whitney’s theorem there is \(\phi \in C^\infty\), \(\phi \ge 0\), with zero set \(E\); with \(\chi \in C_0^\infty\), \(0\le\chi\le1\), \(\chi = 1\) near \(E\), the function \(\psi_1 := \chi\phi + (1-\chi)\) is smooth, nonnegative, equal to \(1\) off a compact set, with zero set \(E\), so
\begin{equation*} \psi(t) := \psi_1(t)\psi_1(-t)e^{-t^2} \end{equation*}
is an even nonnegative Schwartz function with zero set \(E\). Let \(f\) be its inverse Fourier transform, again Schwartz, so \(f \in L^1(\mathbb{R}^1)\cap L^p(\mathbb{R}^1)\) with \(\widehat f = \psi\) by Corollary 3.8.12: thus \(\widehat f \ge 0\) everywhere and \(\widehat f>0\) off the null set \(E\). Finally, for every \(\delta\), Fubini’s theorem (legitimate since \(\int\!\!\int|f(x+\delta)|\,dx\,\mu(ds) = \|f\|_1\)) and evenness of \(\psi\) give
\begin{equation*} \int_{\mathbb{R}} f(x+\delta)g(x)\, dx = \int_{\mathbb{R}} e^{-i\delta s}(2\pi)^{1/2}\psi(s)\,\mu(ds) = 0 , \end{equation*}
since \(\psi\) vanishes on \(E\) and \(\mu\) is concentrated there. So the nonzero functional \(g\) annihilates \(\mathcal{F}\), which by the Hahn-Banach theorem is not dense in \(L^p(\mathbb{R}^1)\).
(\(\circ\)) Suppose that a sequence of functions \(f_n\in L^1(\mu)\) converges weakly to a function \(f\) and a sequence of functions \(g_n\in L^1(\mu)\) converges weakly to a function \(g\) and \(|f_n(x)|\le g_n(x)\) for all \(n\). Show that \(|f(x)|\le g(x)\) a.e. Construct an example demonstrating that the estimates \(|f_n(x)|\le|g_n(x)|\) do not imply that \(|f(x)|\le|g(x)|\) a.e.
\(|f| \le g\) a.e.; and for the second part take \(f_n \equiv 1\) with \(g_n = r_n\) the Rademacher functions. For measurable \(A\), the function \(\varphi := I_A\operatorname{sign}f\) lies in \(L^\infty(\mu)\) and \(f\varphi = I_A|f|\), so weak convergence of \(f_n\) and of \(g_n\) (against \(I_A \in L^\infty(\mu)\)), together with \(|f_n| \le g_n\), gives
\begin{equation*} \int_A |f|\, d\mu = \lim_n \int_A f_n \operatorname{sign}f\, d\mu \le \limsup_n \int_A |f_n|\, d\mu \le \lim_n \int_A g_n\, d\mu = \int_A g\, d\mu . \end{equation*}
Taking \(A := \{|f|>g\}\) makes \(\int_A(|f|-g)\,d\mu \le 0\) with a strictly positive integrable integrand, so \(\mu(A)=0\).
For the example, work on \([0,1]\) with Lebesgue measure: \(f_n \equiv 1\) converges to \(f \equiv 1\) even in norm, while
\begin{equation*} g_n(x) := r_n(x) = (-1)^{\lfloor 2^n x\rfloor} \end{equation*}
satisfies \(|g_n| \equiv 1 \ge |f_n|\) at every point. The \(r_n\) are orthonormal in \(L^2[0,1]\), since for \(m<n\) the function \(r_m\) is constant on each dyadic interval of rank \(m\), over which \(r_n\) has zero integral (Check!), so Bessel’s inequality gives \(\int_0^1 r_n\varphi\,dx \to 0\) for every \(\varphi \in L^\infty[0,1] \subset L^2[0,1]\), i.e. \(g_n \to g \equiv 0\) weakly in \(L^1[0,1]\). Thus \(|f| \equiv 1 > 0 \equiv |g|\).
(\(\circ\)) Let \(\mu\) be a bounded nonnegative Borel measure on an open cube \(V\) in \(\mathbb{R}^n\). Show that the set \(C_0^\infty(V)\) of infinitely differentiable functions with support in \(V\) is everywhere dense in \(L^p(\mu)\), \(1\le p<\infty\).
It suffices that \(I_B\) lie in the \(L^p(\mu)\)-closure \(L\) of \(C_0^\infty(V)\) for every Borel \(B \subset V\): simple functions are dense in \(L^p(\mu)\) (Lemma 4.2.1), \(L\) is a closed subspace, and every \(\mu\)-measurable set agrees with a Borel set up to a \(\mu\)-null set. Fix a smooth nondecreasing \(\theta : \mathbb{R}^1 \to [0,1]\) with \(\theta = 0\) on \((-\infty,0]\) and \(\theta = 1\) on \([1,\infty)\), of Lipschitz constant \(L_\theta\); then \(\varphi \in C_0^\infty(V)\) implies \(\theta\circ\varphi \in C_0^\infty(V)\) with values in \([0,1]\), since \(\theta(0)=0\). Let
\begin{equation*} \mathcal{D} := \big\{B \in \mathcal{B}(V) : \inf\|I_B - \varphi\|_{L^p(\mu)} = 0 \text{ over } \varphi \in C_0^\infty(V),\ 0\le\varphi\le1 \big\}. \end{equation*}
(i) \(V \in \mathcal{D}\): take compact cubes \(Q_k \uparrow V\) and smooth Urysohn functions \(\psi_k \in C_0^\infty(V)\), \(0\le\psi_k\le1\), \(\psi_k = 1\) on \(Q_k\); then \(|1-\psi_k| \le I_{V\setminus Q_k}\), so \(\|I_V - \psi_k\|_{L^p(\mu)}^p \le \mu(V\setminus Q_k) \to 0\) by continuity of the finite measure \(\mu\).
(ii) Complements: given \(B \in \mathcal{D}\), choose \(\varphi\) with \(\|I_B-\varphi\|_{L^p(\mu)} < \varepsilon/(2L_\theta)\) and, by (i), \(\psi\) with \(\|1-\psi\|_{L^p(\mu)} < \varepsilon/(2L_\theta)\). Then \(\chi := \theta\circ(\psi-\varphi) \in C_0^\infty(V)\) takes values in \([0,1]\), and since \(1 - I_B\) takes only the values \(0,1\), so that \(\theta(1-I_B) = 1-I_B\),
\begin{equation*} |\chi - I_{V\setminus B}| \le L_\theta\big(|\psi - 1| + |\varphi - I_B|\big), \end{equation*}
giving \(\|\chi - I_{V\setminus B}\|_{L^p(\mu)} < \varepsilon\).
(iii) Finite intersections: with \(\|I_B-\varphi\|_{L^p(\mu)}, \|I_C-\psi\|_{L^p(\mu)} < \varepsilon/2\), the product \(\varphi\psi \in C_0^\infty(V)\) satisfies, as \(|\psi| \le 1\) and \(|I_B|\le1\),
\begin{equation*} |\varphi\psi - I_{B\cap C}| \le |\varphi - I_B| + |\psi - I_C| . \end{equation*}
(iv) Increasing unions: \(B_k \uparrow B\) gives \(\|I_B - I_{B_k}\|_{L^p(\mu)}^p = \mu(B\setminus B_k) \to 0\).
So \(\mathcal{D}\) is a \(\sigma\)-algebra. It contains every open box \(K\) with \(\overline{K} \subset V\): closed boxes \(Q_m \uparrow K\) and smooth Urysohn \(\varphi_m\) with \(\varphi_m = 1\) on \(Q_m\), \(\operatorname{supp}\varphi_m \subset K\), satisfy \(|\varphi_m - I_K| \le I_{K\setminus Q_m}\), so \(\|\varphi_m - I_K\|_{L^p(\mu)}^p \le \mu(K\setminus Q_m) \to 0\). Every open subset of the open cube \(V\) is a countable union of such boxes with rational vertices, so \(\mathcal{D} \supset \mathcal{B}(V)\).
Let \(\mu\) be a probability measure and let \(M\) be a convex set in \(L^1(\mu)\) that consists of probability densities and is closed with respect to convergence in measure. Show that \(M\) is compact in the weak topology.
\(M\) is weakly closed with relatively weakly compact closure, hence weakly compact. Closedness: norm convergence in \(L^1(\mu)\) implies convergence in measure, so \(M\) is norm closed, and a convex norm-closed set is weakly closed, by Hahn-Banach separation of \(M\) from any \(f_0 \notin M\).
Compactness of the closure: every \(f \in M\) has \(\|f\|_{L^1(\mu)} = 1\), so by Corollary 4.7.21 it suffices that \(\sup_{f\in M}\int_{A_n}f\,d\mu \to 0\) for every decreasing \(A_n\) with \(\bigcap_n A_n = \emptyset\). If this fails, the nonincreasing quantities \(s_n := \sup_{f\in M}\int_{A_n}f\,d\mu\) tend to some \(2\alpha>0\), and we may pick \(f_n \in M\) with \(\int_{A_n}f_n\,d\mu > \alpha\). Since \(\sup_n\|f_n\|_{L^1(\mu)} = 1\), Komlos’ theorem 4.7.24 gives a subsequence whose arithmetic means \(S_k := k^{-1}(f_{n_1}+\dots+f_{n_k})\) converge a.e. to some \(f \in L^1(\mu)\). Each \(S_k\) lies in \(M\) by convexity, and a.e. convergence on a probability space is convergence in measure, so \(f \in M\); in particular \(f \ge 0\) with \(\int_X f\,d\mu = 1\), and the Vitali-Scheffe theorem 2.8.9 gives \(\|S_k - f\|_{L^1(\mu)} \to 0\). Fix \(m\). For \(j \ge m\) we have \(A_{n_j} \subset A_{n_m}\) and hence \(\int_{A_{n_m}}f_{n_j}\,d\mu > \alpha\), so for \(k \ge m\)
\begin{equation*} \int_{A_{n_m}} S_k\, d\mu \ge \frac{1}{k}\sum_{j=m}^{k}\alpha = \frac{k-m+1}{k}\,\alpha , \end{equation*}
and letting \(k \to \infty\) in \(L^1(\mu)\) gives \(\int_{A_{n_m}}f\,d\mu \ge \alpha\) for every \(m\). This contradicts \(\int_{A_{n_m}}f\,d\mu \to 0\), which holds by dominated convergence since \(A_{n_m} \downarrow \emptyset\) and \(f\) is integrable.
Let \((X,\mathcal{A},\mu)\) be a probability space and let \(\{f_n\}\) be a sequence of probability densities convergent \(\mu\)-a.e. to a function \(f\). Let \(\Lambda\in L^\infty(\mu)^*\) be a limit point of \(\{f_n\}\) in the \(\ast\)-weak topology of \(L^\infty(\mu)^*\) (which exists by the Banach–Alaoglu theorem). Then \(\Lambda\) corresponds to a nonnegative additive set function \(\Lambda_0\) on \(\mathcal{A}\). Show that \(\Lambda_0=f\cdot\mu+\Lambda_a\), where \(\Lambda_a\) is a nonnegative additive function on \(\mathcal{A}\) without \(\sigma\)-additive component.
The countably additive part of \(\Lambda_0\) is exactly \(f\cdot\mu\). Put \(\Lambda_0(A) := \Lambda(I_A)\); it is finitely additive by linearity of \(\Lambda\), vanishes on \(\mu\)-null sets (there \(I_A = 0\) in \(L^\infty(\mu)\)), and is nonnegative, since for \(g \ge 0\) each weak-\(\ast\) neighbourhood \(\{\Psi : |\Psi(g)-\Lambda(g)|<\varepsilon\}\) contains some \(f_n\), whence \(\Lambda(g) > \int f_ng\,d\mu - \varepsilon \ge -\varepsilon\); and \(\Lambda_0(X) = \Lambda(1) = 1\) because \(\int f_n\,d\mu \equiv 1\). The same remark gives, for fixed \(g \in L^\infty(\mu)\),
\begin{equation*} \liminf_{n} \int_X f_n g\, d\mu \le \Lambda(g) \le \limsup_{n} \int_X f_n g\, d\mu , \tag{\(\ast\)} \end{equation*}
with \(\Lambda(g) = \lim_n \int f_ng\,d\mu\) whenever that limit exists.
By the Yosida-Hewitt decomposition of 3.10(iv) (Theorem 3.10.14 and Corollary 3.10.15), \(\Lambda_0 = \nu + \Lambda_a\) uniquely with \(\nu \ge 0\) countably additive and \(\Lambda_a \ge 0\) purely additive. Uniqueness makes \(\nu\) the largest countably additive minorant of \(\Lambda_0\): for countably additive \(0 \le \sigma \le \Lambda_0\), decomposing \(\Lambda_0 - \sigma = \tau_c + \tau_p\) exhibits \(\Lambda_0 = (\sigma+\tau_c)+\tau_p\), so \(\nu = \sigma + \tau_c \ge \sigma\). Since \(0 \le \nu \le \Lambda_0\) and \(\Lambda_0\) kills \(\mu\)-null sets, \(\nu \ll \mu\), so \(\nu = \varrho\cdot\mu\) with \(\varrho \ge 0\) by the Radon-Nikodym theorem 3.2.2.
(i) \(f \le \varrho\) a.e.: here \(f \ge 0\) a.e., and Fatou’s theorem together with \((\ast)\) at \(g = I_A\) gives
\begin{equation*} \int_A f\, d\mu \le \liminf_{n}\int_A f_n\, d\mu \le \Lambda_0(A), \qquad A \in \mathcal{A} \end{equation*}
(so \(f\) is integrable, by \(A = X\)). Thus \(f\cdot\mu\) is a countably additive minorant of \(\Lambda_0\) and \(f\cdot\mu \le \nu = \varrho\cdot\mu\).
(ii) \(\varrho \le f\) a.e.: fix \(\varepsilon>0\), take \(\delta>0\) from absolute continuity of the integral of \(f+\varrho\), and use Egoroff’s theorem 2.2.1 to get \(E\) with \(\mu(X\setminus E)<\delta\), hence \([(f+\varrho)\cdot\mu](X\setminus E)<\varepsilon\), on which \(f_n \to f\) uniformly. For \(A \subset E\) then \(\int_Af_n\,d\mu \to \int_Af\,d\mu\), so \(\Lambda_0(A) = \int_A f\,d\mu\) by \((\ast)\) and
\begin{equation*} 0 \le \Lambda_a(A) = \Lambda_0(A) - \nu(A) = \int_A (f-\varrho)\, d\mu \le 0 \end{equation*}
by (i); hence \(f = \varrho\) a.e. on \(E\). Applying this with \(\varepsilon = 1/k\) and putting \(E^* := \bigcup_k E_k\) gives \(f = \varrho\) a.e. on \(E^*\) and \(\int_{X\setminus E^*}(f+\varrho)\,d\mu = 0\), so \(f = \varrho = 0\) a.e. off \(E^*\).
Thus \(\nu = f\cdot\mu\) and \(\Lambda_0 = f\cdot\mu + \Lambda_a\).
Exercise 4.7.154
Construct probability densities \(f_n\) on \([0,1]\) with Lebesgue measure \(\lambda\) that converge to \(0\) in measure but where the constant function \(1\) belongs to the closure of \(\{f_n\}\) in the weak topology \(\sigma(L^1,L^\infty)\). In particular, in the previous exercise, one cannot replace convergence almost everywhere by convergence in measure.
Enumerate \(\bigcup_{n\ge1}\mathcal{F}_n\) as \(\{f_j\}\), taking \(\mathcal{F}_1\) first, then \(\mathcal{F}_2\), and so on, where \(\mathcal{F}_n\) is the (finite) set of functions on \([0,1]\) that are constant on each
\begin{equation*} J_{n,k} := \big[(k-1)4^{-n},\, k4^{-n}\big), \qquad k = 1,\dots,4^n \end{equation*}
(the last one closed), take all their values in the grid \(G_n := \{m\,8^{-n} : 0 \le m \le 32^n\} \subset [0,4^n]\), and satisfy
\begin{equation*} \int_0^1 f\, d\lambda = 1, \qquad \lambda(\{f>0\}) \le 2^{-n} . \end{equation*}
Each \(\mathcal{F}_n\) is nonempty: \(2^n\mathbf{1}_{[0,2^{-n}]}\) is constant on every \(J_{n,k}\) and \(2^n = 16^n8^{-n} \in G_n\). Every \(f_j\) is a probability density.
\(f_j \to 0\) in measure: for \(f \in \mathcal{F}_n\) and \(\tau>0\) we have \(\lambda(\{f>\tau\}) \le \lambda(\{f>0\}) \le 2^{-n}\), and beyond the finitely many members of \(\mathcal{F}_1\cup\dots\cup\mathcal{F}_{n_0}\) every \(f_j\) lies in some \(\mathcal{F}_n\) with \(n>n_0\).
It remains to meet every basic weak neighbourhood of \(1\),
\begin{equation*} U = \Big\{\varphi : \Big|\int_0^1 \psi_i(\varphi-1)\,d\lambda\Big| < \varepsilon, \ i \le r\Big\}, \qquad \psi_i \in L^\infty[0,1]. \end{equation*}
Replacing each \(\psi_i\) by a simple \(\widetilde\psi_i\) with \(\|\psi_i-\widetilde\psi_i\|_\infty \le \varepsilon/4\) costs at most \(2\cdot\varepsilon/4\), since \(\|\varphi-1\|_{L^1} \le 2\) for any density \(\varphi\); and refining the level sets of \(\widetilde\psi_1,\dots,\widetilde\psi_r\) gives a finite partition of \([0,1]\) into sets \(A_1,\dots,A_N\) of positive measure with \(\widetilde\psi_i = \sum_j c_{ij}\mathbf{1}_{A_j}\) (null pieces contribute nothing), so that with \(M := 1+\max_{i,j}|c_{ij}|\),
\begin{equation*} \Big|\int_0^1\widetilde\psi_i(f-1)\,d\lambda\Big| \le M\sum_{j\le N}\Big|\int_{A_j}f\,d\lambda - \alpha_j\Big|, \qquad \alpha_j := \lambda(A_j). \end{equation*}
So it suffices to find, for any such partition and any \(\eta \in (0,1)\), some \(f \in \mathcal{F}_n\) with \(|\int_{A_j}f\,d\lambda - \alpha_j| < \eta\) for all \(j\); one then takes \(\eta := \min\{1/2,\ \varepsilon/(2MN+1)\}\), making the right-hand side above less than \(\varepsilon/2\).
Put \(\theta := \eta/32\). By the Lebesgue density theorem (Chapter 5) choose a density point \(a_j \in A_j\cap(0,1)\) of each \(A_j\) and then \(\delta_0>0\) with
\begin{equation*} \lambda\big(A_j \cap [a_j-\delta, a_j+\delta]\big) > (1-\theta)\,2\delta, \qquad 0<\delta\le\delta_0 , \end{equation*}
the intervals \([a_j-\delta_0,a_j+\delta_0]\) being pairwise disjoint and contained in \((0,1)\). Choose \(n\) with
\begin{equation*} \frac{2^{-n}}{4N} \le \delta_0, \qquad 4N2^{-n} \le \theta, \qquad 8^{-n} \le \frac{\eta}{8}, \qquad 2^n \ge 4N , \end{equation*}
and set \(\delta := 2^{-n}/(4N)\), \(\Delta_j := [a_j-\delta,a_j+\delta]\). Let \(E_j\) be the union of those \(J_{n,k}\) contained in \(\Delta_j\); then \(\Delta_j\setminus E_j\) is covered by at most two intervals of rank \(n\), so \(\lambda(\Delta_j\setminus E_j) \le 2\cdot4^{-n} \le \theta\cdot2\delta\), and with \(v_j := \lambda(E_j)\), \(u_j := \lambda(A_j\cap E_j)\),
\begin{equation*} (1-\theta)2\delta \le v_j \le 2\delta, \quad u_j > (1-2\theta)2\delta, \quad \frac{v_j-u_j}{v_j} \le \frac{2\theta}{1-\theta} \le 4\theta . \tag{1} \end{equation*}
The \(E_j\) are disjoint, each \(v_j\) is an integer multiple of \(4^{-n}\), and
\begin{equation*} \sum_{j\le N} v_j \le 2N\delta = \tfrac12 2^{-n} . \tag{2} \end{equation*}
Heights: put \(\beta_j := \alpha_j/v_j\) and \(b_j := 8^{-n}\lfloor \beta_j 8^n\rfloor\), the largest grid point \(\le \beta_j\), so \(0 \le \beta_j - b_j \le 8^{-n}\) and \(b_j \in G_n\), since \(\beta_j \le 1/\delta = 4N2^n \le 4^n\) by \(2^n \ge 4N\). With \(S := \sum_j b_jv_j\), the bounds \(b_j \le \beta_j\) and \(b_jv_j \ge \alpha_j - 8^{-n}v_j\) with \(\sum_jv_j\le1\) give
\begin{equation*} 1 - 8^{-n} \le S \le 1 . \tag{3} \end{equation*}
Writing \(b_j = l_j8^{-n}\), \(v_j = m_j4^{-n}\) with integers \(l_j, m_j \ge 0\), the number \(L := 32^n - \sum_j l_jm_j\) is an integer with \(L\,32^{-n} = 1-S \in [0, 8^{-n}]\), i.e. \(0 \le L \le 4^n\). By (2) the \(E_j\) together occupy at most \(\tfrac12 2^n < 4^n\) intervals of rank \(n\), so some \(E_0 := J_{n,k_0}\) is disjoint from all of them; put \(b_0 := L\,8^{-n} \in G_n\) and
\begin{equation*} f := b_0\mathbf{1}_{E_0} + \sum_{j\le N} b_j\mathbf{1}_{E_j} . \end{equation*}
Then \(f\) is constant on each \(J_{n,k}\) with values in \(G_n\), and
\begin{equation*} \int_0^1 f\, d\lambda = b_0 4^{-n} + S = L\,32^{-n} + S = 1, \qquad \lambda(\{f>0\}) \le 4^{-n} + \tfrac12 2^{-n} \le 2^{-n} , \end{equation*}
so \(f \in \mathcal{F}_n\). Splitting \(\int_{A_j}f\,d\lambda = b_ju_j + \sum_{i\ne j}b_i\lambda(E_i\cap A_j) + b_0\lambda(E_0\cap A_j)\):
(i) \(b_ju_j \le b_jv_j \le \beta_jv_j = \alpha_j\), while (1), (3) and \(b_jv_j \ge \alpha_j - 8^{-n}\) give \(b_ju_j \ge (1-4\theta)(\alpha_j-8^{-n}) \ge \alpha_j - 4\theta - 8^{-n}\), using \(\alpha_j \le 1\);
(ii) for \(i \ne j\) disjointness gives \(E_i\cap A_j \subset E_i\setminus A_i\), so by (1) \(b_i\lambda(E_i\cap A_j) \le \beta_i(v_i-u_i) \le 4\theta\alpha_i\), and \(\sum_i \alpha_i = 1\);
(iii) \(b_0\lambda(E_0\cap A_j) \le b_0\lambda(E_0) = 1-S \le 8^{-n}\).
Hence \(|\int_{A_j}f\,d\lambda - \alpha_j| \le 8\theta + 2\cdot 8^{-n} \le \eta/4 + \eta/4 < \eta\), as required.
So \(1\) lies in the \(\sigma(L^1,L^\infty)\)-closure of \(\{f_j\}\) without being a member of it (each \(f \in \mathcal{F}_n\) vanishes on a set of measure at least \(1-2^{-n}\)), while \(f_j \to 0\) in measure. Consequently \(\Lambda(\psi) := \int_0^1\psi\,d\lambda\) is a \(\ast\)-weak limit point of \(\{f_j\}\) whose set function is \(\Lambda_0 = \lambda \ne 0\cdot\lambda\), so almost everywhere convergence in Exercise 4.7.153 cannot be weakened to convergence in measure.
Connections between the Integral and Derivative
Exercises 5.8.37–5.8.43
(\(\circ\)) Prove that if a function \(f\) has a finite derivative at every point of the line, then \(f^{\prime}\) has a dense set of continuity points (see, however, Exercise 5.8.119). Hence there exists a closed interval on which the function \(|f^{\prime}|\) is bounded. In particular, \(f\) is Lipschitzian on this interval.
\(f^{\prime}\) is the everywhere finite pointwise limit of the continuous functions
\begin{equation*} g_k(x) := k\big(f(x+1/k) - f(x)\big), \qquad k \in \mathbb{N}, \end{equation*}
continuous because \(f\), being everywhere differentiable, is continuous, and convergent to \(f^{\prime}(x)\) since \(g_k(x)\) is the difference quotient at step \(h = 1/k\). So Baire’s theorem of Exercise 2.12.73, applied on an arbitrary closed interval \([a,b]\), makes the continuity points of \(f^{\prime}|_{[a,b]}\) dense in \([a,b]\); at an interior point continuity of the restriction is continuity of \(f^{\prime}\) on \(\mathbb{R}^1\), and every nonempty open set contains such an interval, so the continuity points of \(f^{\prime}\) are dense in \(\mathbb{R}^1\).
At such a point \(x_0\) choose \(\delta>0\) with \(|f^{\prime}(x)-f^{\prime}(x_0)|<1\) for \(|x-x_0|<\delta\). Then \(|f^{\prime}| \le L := |f^{\prime}(x_0)|+1\) on \(J := [x_0-\delta/2,\, x_0+\delta/2]\), and the mean value theorem gives
\begin{equation*} |f(v)-f(u)| = |f^{\prime}(\xi)|\,|v-u| \le L|v-u|, \qquad u,v \in J . \end{equation*}
Prove that if the derivative of a function \(f\) is everywhere finite and equals almost everywhere some continuous function, then it equals that function everywhere and \(f\) is continuously differentiable.
\(f^{\prime} = g\) everywhere, so \(f \in C^1\). Fix \([a,b]\). Since \(f\) is differentiable at every point, Theorem 5.7.7 with empty exceptional set makes \(f^{\prime}\) Henstock-Kurzweil integrable with
\begin{equation*} (HK)\int_a^z f^{\prime} = f(z) - f(a), \qquad z \in [a,b]. \end{equation*}
The function \(f^{\prime}-g\) vanishes a.e., hence is Henstock-Kurzweil integrable with zero integral (Example 5.7.5), so by linearity (Proposition 5.7.6(iii), with (i) to restrict to \([a,z]\)) the same formula holds for \(g\). Being continuous, \(g\) is bounded and Lebesgue integrable on \([a,z]\), hence McShane integrable with the same value (Theorem 5.7.14) and therefore Henstock-Kurzweil integrable with that value, since every tagged partition is a free one. Thus
\begin{equation*} f(z) - f(a) = \int_a^z g(t)\, dt , \qquad z \in [a,b], \end{equation*}
an ordinary Lebesgue integral of a continuous function, and
\begin{equation*} \Big|\frac1h\int_z^{z+h} g\,dt - g(z)\Big| \le \sup_{|t-z|\le|h|}|g(t)-g(z)| \longrightarrow 0 \end{equation*}
gives \(f^{\prime}(z) = g(z)\) for \(z \in (a,b)\); \([a,b]\) was arbitrary.
Method (2): put \(h(x) := f(x) - \int_a^x g(t)\,dt\), differentiable everywhere with \(h^{\prime} = f^{\prime}-g = 0\) a.e. For \(N := \{h^{\prime} \ne 0\}\), which is \(\lambda\)-null, Proposition 5.5.4 gives \(\lambda(h(N)) \le \int_N |h^{\prime}|\,dx = 0\) and, in its last assertion with \(L = 0\), \(\lambda(h([a,b]\setminus N)) = 0\). So the interval \(h([a,b])\) has measure zero, i.e. is a point, and \(h\) is constant.
(\(\circ\)) (i) Construct a continuous strictly increasing function \(f\) on the real line such that \(f^{\prime}(x) = 0\) a.e.
(ii) Show that for such a function one can take
\begin{equation*} f(t) = P\Bigl(\omega \colon\ \sum_{n=1}^{\infty} \xi_n(\omega) 2^{-n} < t\Bigr), \end{equation*}
where \(\xi_n\) are independent random variables (see Chapter 10) on a probability space \((\Omega, P)\) such that \(P(\xi_n = 1) = p\), \(P(\xi_n = 0) = 1 - p\), where \(p \in (0,1)\) and \(p \ne 1/2\).
(i) Take
\begin{equation*} f := \sum_{n=1}^\infty u_n, \qquad u_n(x) := 2^{-n}C_0\Big(\frac{x-a_n}{b_n-a_n}\Big), \end{equation*}
where \(C_0\) is the Cantor function of Proposition 3.6.5 extended by \(0\) on \((-\infty,0]\) and \(1\) on \([1,\infty)\), and \(\{(a_n,b_n)\}\) enumerates the pairs of rationals with \(a_n<b_n\). Each \(u_n\) is continuous, nondecreasing, with values in \([0,2^{-n}]\), so the series converges uniformly and \(f\) is continuous and nondecreasing. It is strictly increasing: for \(x<y\) pick rationals \(x<a<b<y\) and the index \(n\) with \((a_n,b_n) = (a,b)\); then \(u_n(x) = 0\), \(u_n(y) = 2^{-n}\), and the other summands are nondecreasing, so \(f(y)-f(x) \ge 2^{-n}>0\). On each \([-N,N]\) the \(u_n\) are nondecreasing with everywhere convergent sum, so Proposition 5.2.8 (Fubini’s theorem on term-by-term differentiation; see Exercise 5.8.42) gives \(f^{\prime} = \sum_n u_n^{\prime}\) a.e.; and \(u_n^{\prime}(x) = 2^{-n}(b_n-a_n)^{-1}C_0^{\prime}((x-a_n)/(b_n-a_n)) = 0\) a.e., since \(C_0^{\prime} = 0\) a.e. and affine maps preserve null sets. Hence \(f^{\prime} = 0\) a.e.
(ii) Put \(X := \sum_{n\ge1}\xi_n 2^{-n} \in [0,1]\) and \(\mu := P\circ X^{-1}\), so \(f(t) = \mu((-\infty,t))\) is nondecreasing and left continuous, \(0\) for \(t\le0\) and \(1\) for \(t>1\); write \(q := \max(p,1-p) \in (0,1)\).
(a) Continuity amounts to \(P(X=t) = 0\). A \(t \in [0,1]\) has at most two binary expansions, and for a fixed digit sequence independence gives \(P(\xi_n = \varepsilon_n \ \forall n) \le q^N \to 0\).
(b) Strict increase: for \(\varepsilon_1,\dots,\varepsilon_N \in \{0,1\}\) and \(k := \sum_{n\le N}\varepsilon_n2^{N-n}\), on \(A := \{\xi_n = \varepsilon_n,\ n \le N\}\) one has \(X \in [k2^{-N},(k+1)2^{-N}]\), so
\begin{equation*} \mu\big([k2^{-N},(k+1)2^{-N}]\big) \ge P(A) = \prod_{n\le N}p^{\varepsilon_n}(1-p)^{1-\varepsilon_n} > 0 . \end{equation*}
Given \(0\le s<t\le1\), choose \(N\) with \(2^{-N} < (t-s)/2\) and \(k := \lfloor s2^N\rfloor+1\); then \([k2^{-N},(k+1)2^{-N}] \subset [s,t)\), so \(f(t)-f(s) = \mu([s,t)) > 0\).
(c) \(\mu \perp \lambda\). Let \(d_n(x)\) be the \(n\)-th binary digit of \(x \in [0,1)\) in its terminating expansion. Under \(\lambda\) the \(d_n\) are independent and uniform on \(\{0,1\}\), since \(\{d_n = \varepsilon_n,\ n\le N\}\) is a dyadic interval of length \(2^{-N}\); under \(\mu\) they have the law of the \(\xi_n\), because \(P(\xi_n = 1 \ \forall n \ge m) \le q^N \to 0\) makes \((\xi_n)\) a.s. not eventually \(1\), and then \(d_n(X) = \xi_n\) for all \(n\). By the strong law of large numbers for bounded independent variables (Borel’s theorem, from \(E(\sum_{n\le N}(\eta_n - E\eta_n))^4 = O(N^2)\) and Borel-Cantelli), the set
\begin{equation*} B := \Big\{x \in [0,1) : \lim_N \tfrac1N \sum_{n\le N} d_n(x) = p\Big\} \end{equation*}
has \(\mu(B) = 1\) and, since \(p \ne 1/2\), \(\lambda(B) = 0\).
(d) A singular \(F(t) := \mu((-\infty,t))\) has \(F^{\prime} = 0\) a.e. Indeed \(F\) is nondecreasing, hence a.e. differentiable (Theorem 5.2.6). Take Borel \(A\) with \(\mu(A) = 0\), \(\lambda(\mathbb{R}^1\setminus A) = 0\); by regularity of the finite Borel measure \(\mu\) (the sets approximable from outside by open and from inside by closed sets form a \(\sigma\)-algebra containing the closed ones) choose, for \(c,R,\varepsilon>0\), an open \(V \supset A\cap(-R,R)\) with \(\mu(V)<\varepsilon\). The intervals \([x,x+h] \subset V\) with \(\mu([x,x+h])>ch\) form a Vitali cover of
\begin{equation*} E := \Big\{x \in A\cap(-R,R) : \limsup_{h\to0+}\frac{\mu([x,x+h])}{h} > c\Big\}, \end{equation*}
so Vitali’s theorem 5.5.1 yields a disjoint countable subfamily covering \(E\) up to a null set, whence
\begin{equation*} \lambda^*(E) \le \sum_j h_j \le \frac1c\sum_j \mu\big([x_j,x_j+h_j]\big) \le \frac{\mu(V)}{c} < \frac{\varepsilon}{c} . \end{equation*}
Letting \(\varepsilon \to 0\), then \(c \to 0\) along a sequence and \(R \to \infty\), gives \(\mu([x,x+h])/h \to 0\) as \(h\to0+\) for a.e. \(x\), and at a point of differentiability \(0 \le (F(x+h)-F(x))/h \le \mu([x,x+h])/h\). So \(f^{\prime} = 0\) a.e.
Finally \(f(0) = 0\) and \(f(1^-) = 1\) by (a), so
\begin{equation*} F(x) := \lfloor x\rfloor + f\big(x - \lfloor x \rfloor\big) \end{equation*}
is continuous at the integers, strictly increasing on all of \(\mathbb{R}^1\) by (b), with \(F^{\prime} = 0\) a.e.
(\(\circ\)) (Riesz) Prove that a nonnegative function \(f\) on \([a,b]\) is Lebesgue integrable precisely when there exists a nondecreasing function \(F\) on \([a,b]\) such that \(F^{\prime}(x) = f(x)\) a.e. In addition, the integral of \(f\) equals the infimum of the differences \(F(b) - F(a)\) over all such functions \(F\).
\(f \ge 0\) is Lebesgue integrable exactly when the class \(\mathcal{F}\) of nondecreasing \(F\) on \([a,b]\) with \(F^{\prime} = f\) a.e. is nonempty, and then the infimum is attained at \(F_0(x) := \int_a^x f(t)\,dt\). If \(f\) is integrable, \(F_0\) is nondecreasing (as \(f\ge0\)) and absolutely continuous (Theorem 5.3.6) with \(F_0^{\prime} = f\) a.e. (Theorem 5.4.2), so \(F_0 \in \mathcal{F}\) and
\begin{equation*} F_0(b) - F_0(a) = \int_a^b f(t)\, dt . \end{equation*}
Conversely, for \(F \in \mathcal{F}\) Corollary 5.2.7 gives \(F^{\prime}\) finite a.e. and integrable with \(\int_a^b F^{\prime}\,dx \le F(b)-F(a)\); since \(f = F^{\prime}\) a.e., \(f\) is integrable and
\begin{equation*} \int_a^b f(x)\, dx = \int_a^b F^{\prime}(x)\, dx \le F(b) - F(a) . \end{equation*}
Taking the infimum over \(\mathcal{F}\) and combining the two displays gives \(\int_a^b f\,dx = \inf_{F\in\mathcal{F}}(F(b)-F(a))\).
(\(\circ\)) Let \(f\) be an absolutely continuous function on \([0,1]\). For every \(h > 0\) let \(f_h(x) := h^{-1}\bigl( f(x+h) - f(x) \bigr)\), where \(f(x+h) = f(1)\) if \(x + h > 1\). Show that
\begin{equation*} \lim_{h \to 0} \| f_h - f^{\prime} \|_{L^1[0,1]} = 0 . \end{equation*}
\(f_h\) is the running average of \(g\), the function equal to \(f^{\prime}\) on \([0,1]\) and to \(0\) elsewhere, and the claim is continuity of translation in \(L^1\). Absolute continuity gives \(f^{\prime} \in L^1[0,1]\) and the Newton-Leibniz formula (Theorem 5.3.6), so \(g \in L^1(\mathbb{R}^1)\) and, for \(x \in [0,1]\) and \(h \in (0,1)\),
\begin{equation*} f_h(x) = \frac1h\int_x^{x+h} g(t)\, dt = \frac1h\int_0^h g(x+s)\, ds =: (A_hg)(x) , \end{equation*}
both when \(x+h \le 1\) and, since \(g\) vanishes on \((1,\infty)\) and \(f(x+h) = f(1)\) there, when \(x+h>1\). Hence, by Fubini-Tonelli applied to \((x,s) \mapsto |g(x+s)-g(x)|\) on \(\mathbb{R}^1\times[0,h]\),
\begin{equation*} \|f_h - f^{\prime}\|_{L^1[0,1]} \le \int_{\mathbb{R}} |A_hg - g|\, dx \le \frac1h\int_0^h \|g(\cdot+s)-g\|_{L^1(\mathbb{R})}\, ds , \end{equation*}
which is at most \(\sup_{0\le s\le h}\|g(\cdot+s)-g\|_{L^1(\mathbb{R})} \to 0\) as \(h \to 0\), by continuity of translation in \(L^1\) (Lemma 4.2.3).
Method (2): the family \(\{f_h : h \in (0,1)\}\) is uniformly integrable. For measurable \(E \subset [0,1]\) with \(\lambda(E) \le \varepsilon\) and \(M>0\), Tonelli’s theorem applied to \((x,t) \mapsto |g(t)|I_{[x,x+h]}(t)I_E(x)\) and the splitting of \(|g|\) at level \(M\) give
\begin{equation*} \int_E |f_h(x)|\, dx \le M\varepsilon + \int_{\{|f^{\prime}|>M\}} |f^{\prime}(t)|\, dt , \end{equation*}
using \(\int_0^1 I_{[x,x+h]}(t)\,dt \le h\) for the bounded part and \(h^{-1}\int_0^1 I_{[x,x+h]}(t)I_E(x)\,dx \le h^{-1}\lambda([t-h,t]) = 1\) for the other. The bound is independent of \(h\), and choosing \(M\) then \(\varepsilon\) makes it small, so the integrals are uniformly absolutely continuous and Proposition 4.5.3 (Lebesgue measure being atomless) gives uniform integrability. Since \(f_h \to f^{\prime}\) at every Lebesgue point of \(g\) in \([0,1)\), hence a.e. by Theorem 5.6.2, Theorem 4.5.4 gives convergence in \(L^1[0,1]\).
(\(\circ\)) Prove Proposition 5.2.8.
Proposition 5.2.8 asserts that for nondecreasing \(f_n\) on \([a,b]\) with \(f(x) = \sum_{n\ge1}f_n(x)\) convergent at every \(x\), one has \(f^{\prime} = \sum_n f_n^{\prime}\) a.e. Replacing \(f_n\) by \(f_n - f_n(a)\) alters \(f\) by the constant \(\sum_n f_n(a)\) and no derivative, so assume \(f_n(a) = 0\), making all \(f_n\) and \(f\) nonnegative and nondecreasing. Let \(Z\) be a null set off which \(f\), every \(f_n\) and every partial sum \(s_k := \sum_{n\le k}f_n\) has a finite derivative; these are countably many nondecreasing, hence bounded-variation, functions, so Theorem 5.2.6 applies to each.
(i) Since \(f - s_k = \sum_{n>k}f_n\) is again nondecreasing, the difference quotients satisfy \(h^{-1}(s_k(x+h)-s_k(x)) \le h^{-1}(f(x+h)-f(x))\), so letting \(h \to 0+\) off \(Z\),
\begin{equation*} \sum_{n\le k} f_n^{\prime}(x) = s_k^{\prime}(x) \le f^{\prime}(x), \qquad x \in [a,b]\setminus Z . \end{equation*}
The terms \(f_n^{\prime}(x)\) are nonnegative off \(Z\), so the partial sums increase to a finite limit and
\begin{equation*} g(x) := \sum_{n\ge1} f_n^{\prime}(x) \le f^{\prime}(x), \qquad x \in [a,b]\setminus Z . \tag{1} \end{equation*}
(ii) Since \(\sum_n f_n(b)\) converges with nonnegative terms, choose \(n_1<n_2<\cdots\) with \(\sum_{n>n_k}f_n(b) < 2^{-k}\) and put \(\varphi_k := f - s_{n_k} = \sum_{n>n_k}f_n\). Each \(\varphi_k\) is nondecreasing and nonnegative with
\begin{equation*} 0 \le \varphi_k(x) \le \varphi_k(b) < 2^{-k}, \qquad x \in [a,b], \end{equation*}
so \(\sum_k\varphi_k\) converges at every point and (i) applied to it makes \(\sum_k\varphi_k^{\prime}(x)\) convergent a.e.; hence \(\varphi_k^{\prime}(x) \to 0\) a.e. But off \(Z\) we have \(\varphi_k^{\prime}(x) = f^{\prime}(x) - s_{n_k}^{\prime}(x) \to f^{\prime}(x) - g(x)\) by (1), since \(n_k \to \infty\). So \(f^{\prime} = g\) a.e.
Let \(\varrho \in \mathcal{L}(\mathbb{R}^1)\) be absolutely continuous on bounded intervals. Suppose that a function \(f\) is either absolutely continuous on bounded intervals or everywhere differentiable. Let \(f\varrho^{\prime}\) and \(f^{\prime}\varrho\) be in \(\mathcal{L}^1(\mathbb{R}^1)\). Prove the equality
\begin{equation*} \int_{-\infty}^{+\infty} f^{\prime}(t)\varrho(t)\,dt = -\int_{-\infty}^{+\infty} f(t)\varrho^{\prime}(t)\,dt . \end{equation*}
It suffices to treat bounded \(f\). Fix an even \(\chi \in C^\infty(\mathbb{R}^1)\) with \(0\le\chi\le1\), \(\chi = 1\) on \([-1,1]\), \(\chi = 0\) off \([-2,2]\), and put \(\theta_n(t) := \int_0^t \chi(s/n)\,ds\), so \(\theta_n \in C^1\) is bounded with \(|\theta_n^{\prime}| \le 1\), \(|\theta_n(t)| \le |t|\) and \(\theta_n(t) = t\) for \(|t|\le n\). Being \(1\)-Lipschitz and \(C^1\), \(\theta_n\) preserves either hypothesis on \(f\), and \(g_n := \theta_n\circ f\) is bounded with \(g_n^{\prime} = \theta_n^{\prime}(f)f^{\prime}\) wherever \(f^{\prime}\) exists, so
\begin{equation*} |g_n^{\prime}\varrho| \le |f^{\prime}\varrho| \in L^1(\mathbb{R}^1), \qquad |g_n\varrho^{\prime}| \le |f\varrho^{\prime}| \in L^1(\mathbb{R}^1) . \end{equation*}
Since \(\theta_n^{\prime}(f(t)) = 1\) once \(n \ge |f(t)|\), dominated convergence carries the identity from each \(g_n\) to \(f\). Note that \(\varrho\) is continuous with \(\varrho^{\prime}\) existing a.e. and the Newton-Leibniz formula valid on bounded intervals (Theorem 5.3.6).
(i) \(f\) bounded and absolutely continuous on bounded intervals. As \(\varrho \in L^1(\mathbb{R}^1)\) is continuous, \(\liminf_{t\to+\infty}|\varrho(t)| = 0\) (otherwise \(|\varrho| \ge c>0\) past some \(T\)), so there are \(b_n \to +\infty\) and \(a_n \to -\infty\) with \(\varrho(b_n), \varrho(a_n) \to 0\), whence \(f(b_n)\varrho(b_n) - f(a_n)\varrho(a_n) \to 0\) by boundedness of \(f\). Integration by parts on \([a_n,b_n]\), where both functions are absolutely continuous (Corollary 5.4.3), gives
\begin{equation*} \int_{a_n}^{b_n} f^{\prime}\varrho\, dt + \int_{a_n}^{b_n} f\varrho^{\prime}\, dt = f(b_n)\varrho(b_n) - f(a_n)\varrho(a_n) , \end{equation*}
and dominated convergence with the majorants \(|f^{\prime}\varrho|, |f\varrho^{\prime}|\) gives the identity.
(ii) \(f\) bounded and everywhere differentiable (it need not be absolutely continuous: \(f(x) = x^2\sin x^{-4}\), \(f(0)=0\)). On a closed interval \([c,d]\) where \(\varrho\) does not vanish, \(m := \min_{[c,d]}|\varrho|>0\) gives \(|f^{\prime}| \le |f^{\prime}\varrho|/m \in L^1[c,d]\); Theorem 5.7.7 makes \(f^{\prime}\) Henstock-Kurzweil integrable there with \((HK)\int_c^z f^{\prime} = f(z)-f( c)\), while Lebesgue integrability gives McShane and hence Henstock-Kurzweil integrability with the Lebesgue value (Theorem 5.7.14), so by uniqueness \(f(z) = f( c)+\int_c^z f^{\prime}\,dt\) and \(f \in AC[c,d]\) by Theorem 5.3.6. Write the open set \(U := \{\varrho \ne 0\}\) as a countable disjoint union of open intervals and fix a component \(I = (\alpha,\beta)\). Take \(c_k \downarrow \alpha\), \(d_k \uparrow \beta\) inside \(I\), choosing \(\varrho(c_k) \to 0\) when \(\alpha = -\infty\) as in (i), while for finite \(\alpha\) one has \(\alpha \notin U\), so \(\varrho(\alpha) = 0\) and \(\varrho(c_k) \to 0\) by continuity; similarly at \(\beta\). Integration by parts on \([c_k,d_k] \subset I\), where \(f\) and \(\varrho\) are absolutely continuous, has boundary terms tending to \(0\), so
\begin{equation*} \int_{\alpha}^{\beta} f^{\prime}\varrho\, dt = -\int_{\alpha}^{\beta} f\varrho^{\prime}\, dt . \end{equation*}
Off \(U\) we have \(\varrho = 0\), hence \(f^{\prime}\varrho = 0\), and also \(\varrho^{\prime} = 0\) a.e.: at a non-isolated point \(t\) of \(\{\varrho = 0\}\) where \(\varrho^{\prime}(t)\) exists, taking zeros \(t_i \to t\) gives \(\varrho^{\prime}(t) = \lim_i(\varrho(t_i)-\varrho(t))/(t_i-t) = 0\), while the isolated points form a countable set and \(\varrho^{\prime}\) exists a.e. Summing over the components, by countable additivity of the integrals of \(f^{\prime}\varrho\) and \(f\varrho^{\prime}\), gives the identity.
Exercises 5.8.44–5.8.50
Let \(\mu\) be a probability measure on a space \(X\) and let \(f\) be a nonnegative \(\mu\)-measurable function. Suppose that \(\varphi\) is a locally absolutely continuous increasing function on \([0,+\infty)\). Prove that
\begin{equation*} \int_X \varphi\big(f(x)\big)\,\mu(dx) = \varphi(0) + \int_0^{\infty} \varphi^{\prime}(t)\,\mu\big(x:\ f(x)>t\big)\,dt, \end{equation*}
where both integrals are finite or infinite simultaneously.
The identity is Tonelli’s theorem applied to \(H(x,t) := g(t)I_{\{f(x)>t\}} \ge 0\). Local absolute continuity gives \(\varphi(s) = \varphi(0)+\int_0^s\varphi^{\prime}(t)\,dt\) for \(s \ge 0\) (Theorem 5.3.6), with \(\varphi^{\prime} \ge 0\) wherever it exists since \(\varphi\) increases; let \(g \ge 0\) be a Borel version of \(\varphi^{\prime}\), extended by \(0\) on the null set where the derivative fails to exist, so that
\begin{equation*} \varphi(s)-\varphi(0) = \int_0^{\infty} g(t)\, I_{\{t<s\}}\, dt , \qquad s \ge 0 , \tag{1} \end{equation*}
and, by monotone convergence, also for \(s = +\infty\) with \(\varphi(+\infty) := \lim_{s\to+\infty}\varphi(s)\), so \(f\) may take the value \(+\infty\). The set
\begin{equation*} \{(x,t) : f(x)>t\} = \bigcup_{q \in \mathbb{Q},\, q\ge0} \{x : f(x)>q\}\times[0,q) \end{equation*}
lies in \(\mathcal{A}_\mu \otimes \mathcal{B}([0,+\infty))\), so \(H\) is measurable and nonnegative, and both measures are \(\sigma\)-finite; Tonelli’s theorem then equates the two iterated integrals, whose inner integrals are \(\varphi(f(x))-\varphi(0)\) by (1) and \(g(t)\mu(x : f(x)>t)\) respectively:
\begin{equation*} \int_X \big[\varphi(f(x))-\varphi(0)\big]\,\mu(dx) = \int_0^{\infty}\varphi^{\prime}(t)\,\mu\big(x : f(x)>t\big)\, dt . \tag{2} \end{equation*}
Both sides of (2) integrate nonnegative functions, hence are finite or infinite together, and adding \(\varphi(0) = \varphi(0)\mu(X)\) gives the stated identity.
Construct a continuous function \(F\) on the interval \([0,1]\) such that, at every point in the interval, it has a finite or infinite derivative \(f\) that is almost everywhere finite and integrable, but the function
\begin{equation*} \Phi(x) = F(0) + \int_0^x f(t)\,dt \end{equation*}
has no finite or infinite derivative at infinitely many points (in particular, \(\Phi\) does not coincide with \(F\)).
Take \(F=\Phi_0+\psi\) with \(\psi\) singular and \(\Phi_0\) the indefinite integral built below; both summands are nondecreasing, so an infinite derivative of either forces \(F^{\prime}=+\infty\).
Let \(K=\bigcap_n K_n\), where \(K_n\) keeps in each of the \(2^{n-1}\) intervals of \(K_{n-1}\) (length \(8^{-(n-1)}\)) the two extreme subintervals of length \(8^{-n}\), deleting a middle gap of length \(6\cdot 8^{-n}\). Then \(\lambda(K)=\lim_n 2^n8^{-n}=0\), and with \(2^{n-1}\) gaps at level \(n\),
\begin{equation*} \Sigma:=\sum_{\text{gaps } G}\sqrt{\lambda(G)} =\frac{\sqrt6}{2}\sum_{n\ge1}2^{-n/2}<\infty. \tag{3} \end{equation*}
Let \(\mu\) be the Borel probability measure with \(\mu(I)=2^{-n}\) on each of the \(2^n\) intervals \(I\) of \(K_n\); it is atomless, so \(\psi(x):=\mu([0,x])\) is continuous and nondecreasing, and it is constant on every gap, whence \(\psi^{\prime}=0\) off \(K\) and \(\psi\) is singular. Code \(x\in K\) by \(\omega(x)\in\{0,1\}^{\mathbb{N}}\), \(\omega_n\) telling which child of \(I_{n-1}(x)\) is \(I_n(x)\), with endpoints \(l_n(x)<r_n(x)\); \(|I_n(x)|=8^{-n}\) and \(\mu(I_n(x))=2^{-n}\), and eventually constant codes are exactly the gap endpoints together with \(0,1\).
Put \(P:=\{x\in K:\ \text{every block of equal consecutive symbols of } \omega(x)\text{ has length}\le2\}\). Its complement in \(K\) is the union of the relatively clopen cylinder sets \(\{\omega_n=\omega_{n+1}=\omega_{n+2}\}\), so \(P\) is compact; it is uncountable (concatenate blocks from \(\{01,0011\}\)) and contains no gap endpoint, so each of its points is a two-sided limit point of \(K\). Fix \(x\in P\) and \(n\ge0\): among \(\omega_{n+1},\omega_{n+2},\omega_{n+3}\) some symbol is \(0\) and some is \(1\), and if \(\omega_{n+j}=0\) then the sibling \(J\) of \(I_{n+j}(x)\) lies to the right of \(x\) inside \(I_n(x)\), giving \(\psi(r_n(x))-\psi(x)\ge \mu(J)=2^{-(n+j)}\); symmetrically on the left. Hence for all \(n\),
\begin{equation*} \psi(r_n(x))-\psi(x)\ \ge\ 2^{-(n+3)},\qquad \psi(x)-\psi(l_n(x))\ \ge\ 2^{-(n+3)} . \tag{4} \end{equation*}
If \(8^{-(n+2)}<h\le 8^{-(n+1)}\) then \(r_{n+2}(x)-x\le 8^{-(n+2)}<h\), so by monotonicity and (4) at level \(n+2\),
\begin{equation*} \frac{\psi(x+h)-\psi(x)}{h}\ \ge\ \frac{2^{-(n+5)}}{8^{-(n+1)}} =2^{2n-2}\xrightarrow[n\to\infty]{}+\infty , \end{equation*}
and the same with \(l_{n+2}(x)\) for \(h<0\); thus \(\psi^{\prime}=+\infty\) on \(P\). Choose distinct \(p_k\in P\) with \(p_k\to q\in P\) and set \(\rho_k:=\tfrac13\min(1-p_k,|p_k-q|,\inf_{j\ne k}|p_k-p_j|)>0\), \(B_k:=(p_k-\rho_k,p_k+\rho_k)\); these are pairwise disjoint subintervals of \((0,1)\) accumulating only at \(q\). (Check!)
Let \(f_1:=\operatorname{dist}(\cdot,P)^2\operatorname{dist}(\cdot,K)^{-1/2}\) off \(K\) and \(f_1:=0\) on \(K\); it is Borel and continuous off \(K\). Since \(\int_G\operatorname{dist}(t,K)^{-1/2}dt\le 2\sqrt{2\lambda(G)}\) on each gap \(G\), summation and (3) give
\begin{equation*} \int_0^1\operatorname{dist}(t,K)^{-1/2}\,dt\ \le\ 2\sqrt2\,\Sigma=:C_0<\infty, \tag{5} \end{equation*}
so \(f_1\le\operatorname{dist}(\cdot,K)^{-1/2}\) is integrable.
Fix \(k\), write \(p:=p_k\), and choose recursively \(t_1>s_1>t_2>s_2>\dots\to0\) with \(t_1<\rho_k\), \(p+t_j,p+s_j\in K\), \(s_j\le t_j/2\) and \(t_{j+1}\le s_j/(j+1)\) (possible since \(p\) is a two-sided limit point of \(K\)). As \(\lambda((p+s_j,p+t_j)\setminus K)=t_j-s_j\ge t_j/2\), select finitely many gaps inside \((p+s_j,p+t_j)\) of total length \(\ge0.4\,t_j\) and place on each a continuous trapezoidal bump \(0\le b\le1\) vanishing at its endpoints with integral \(\ge0.9\) times its length. Let \(f_2^{(k)}\le1\) be the sum of these bumps over \(j\ge1\); it vanishes outside \(B_k\), and
\begin{equation*} \int_p^{p+t_j}f_2^{(k)}\ \ge\ 0.36\,t_j\ \ge\ \frac{t_j}{4},\qquad \int_p^{p+s_j}f_2^{(k)}\ \le\ 2t_{j+1}\ \le\ \frac{2s_j}{j+1}, \tag{6} \end{equation*}
the second because levels \(i\le j\) lie to the right of \(p+s_j\), levels \(i>j\) lie in \((p,p+t_{j+1})\), and \(t_{i+1}\le t_i/2\). Put \(f:=f_1+\sum_kf_2^{(k)}\): nonnegative, Borel, integrable by (5), and continuous off \(K\), since near \(y\notin K\) (so \(y\ne q\)) only finitely many \(B_k\) intrude and only finitely many intervals \((p_k+s_j,p_k+t_j)\) meet a neighbourhood of \(y\), these shrinking to \(p_k\in K\).
Set \(\Phi_0(x):=\int_0^xf(t)\,dt\) and \(F:=\Phi_0+\psi\); both are continuous and nondecreasing, so all their difference quotients are nonnegative.
(i) \(x\notin K\): \(\psi\) is locally constant at \(x\) and \(f\) is continuous there, so \(F^{\prime}(x)=f(x)\) is finite.
(ii) \(x\in K\setminus P\): with \(d:=\operatorname{dist}(x,P)>0\) and \(0<|h|\le d/2\), every \(t\) between \(x\) and \(x+h\) has \(\operatorname{dist}(t,K)\le|h|\) and \(\operatorname{dist}(t,P)\ge d/2\), so \(f(t)\ge \tfrac{d^2}{4}|h|^{-1/2}\) and the quotient of \(\Phi_0\) is at least \(\tfrac{d^2}{4}|h|^{-1/2}\to+\infty\); hence \(F^{\prime}(x)=+\infty\).
(iii) \(x\in P\): \(\psi^{\prime}(x)=+\infty\), hence \(F^{\prime}(x)=+\infty\).
So \(F^{\prime}\) exists everywhere, equals \(f\) off the null set \(K\), and \(F(0)=\psi(0)\), whence \(\Phi=\psi(0)+\Phi_0\). At \(p=p_k\): both \(p\) and \(p+s_j\) lie in \(K\), so every gap meeting \((p,p+s_j)\) lies in \([p,p+s_j]\), where \(\operatorname{dist}(t,P)\le s_j\); thus \(\int_p^{p+s_j}f_1\le C_0s_j^2\) by (5), and with (6) and the disjointness of the \(B_i\),
\begin{equation*} \frac{\Phi_0(p+t_j)-\Phi_0(p)}{t_j}\ \ge\ \frac14,\qquad \frac{\Phi_0(p+s_j)-\Phi_0(p)}{s_j}\ \le\ \frac{2}{j+1}+C_0s_j\to0 . \end{equation*}
Hence \(\Phi\) has no derivative, finite or infinite, at any of the points \(p_1,p_2,\dots\), and \(F-\Phi=\psi-\psi(0)\) is nonconstant.
Show that given \(E\subset[0,1]\) with \(\lambda(E)=0\), there exists a continuous nondecreasing function \(\psi\) on \([0,1]\) with \(\psi^{\prime}(x)=+\infty\) for all \(x\in E\).
Take \(\psi:=\sum_{n\ge1}\varphi_n\) with \(\varphi_n(x):=\lambda\big(G_n\cap[0,x]\big)\), where \(G_n\supset E\) is open with \(\lambda(G_n)<2^{-n}\) (outer regularity, applicable since \(\lambda(E)=0\)). Each \(\varphi_n\) is nondecreasing and \(1\)-Lipschitz, because \(\varphi_n(y)-\varphi_n(x)=\lambda(G_n\cap(x,y])\le y-x\) for \(x<y\), and \(0\le\varphi_n\le\lambda(G_n)<2^{-n}\); so the series converges uniformly (Weierstrass test) and \(\psi\) is continuous and nondecreasing.
Fix \(x_0\in E\) and \(k\in\mathbb{N}\). The sets \(G_1,\dots,G_k\) are open and contain \(x_0\), so some \(\delta>0\) has \((x_0-\delta,x_0+\delta)\subset G_n\) for all \(n\le k\); for \(0<|h|<\delta\) with \(x_0+h\in[0,1]\) the interval between \(x_0\) and \(x_0+h\) lies in each such \(G_n\), so \(\varphi_n(x_0+h)-\varphi_n(x_0)=h\). Since every \(\varphi_n\) is nondecreasing, all its difference quotients are nonnegative, whence
\begin{equation*} \frac{\psi(x_0+h)-\psi(x_0)}{h}\ \ge\ \sum_{n=1}^{k}\frac{\varphi_n(x_0+h)-\varphi_n(x_0)}{h}\ =\ k . \end{equation*}
So \(\liminf_{h\to0}\big(\psi(x_0+h)-\psi(x_0)\big)/h\ge k\) for every \(k\), i.e. \(\psi^{\prime}(x_0)=+\infty\) (one-sided at an endpoint of \([0,1]\)).
Construct an example of a continuous function \(f\) on \((0,1)\) that at no point has the usual derivative, but is approximately differentiable almost everywhere.
Take \(f:=\sum_{n\ge1}g_n\), a uniformly convergent series of periodic tent functions with huge slopes on supports of measure \(2^{-n}\). Put \(\delta_n:=100^{-n}\); given \(g_1,\dots,g_{n-1}\) and \(S_{n-1}:=\sum_{m<n}\operatorname{Lip}(g_m)<\infty\), choose
\begin{equation*} 0<\eta_n\le\min\Big(\delta_n^2,\ \frac{\delta_n}{8(S_{n-1}+1)}\Big),\qquad \ell_n:=2^{-n}\eta_n , \end{equation*}
and let \(g_n\) be the \(\eta_n\)-periodic continuous function
\begin{equation*} g_n(x):=\delta_n\max\Big(0,\ 1-\frac{2}{\ell_n} \operatorname{dist}\big(x,\eta_n\mathbb{Z}\big)\Big). \end{equation*}
So \(0\le g_n\le\delta_n\), \(g_n=\delta_n\) on \(\eta_n\mathbb{Z}\), \(\operatorname{Lip}(g_n)=2\delta_n/\ell_n\), and \(g_n\) vanishes outside \(U_n:=\{\operatorname{dist}(\cdot,\eta_n\mathbb{Z})<\ell_n/2\}\), where \(\lambda(U_n\cap(0,1))\le\ell_n/\eta_n+\ell_n\le2^{-n+1}\); being piecewise linear, \(g_n\) is differentiable off a countable set. As \(\sum_n\delta_n<\infty\), \(f\) is continuous on \((0,1)\). For \(|y-x|\le\eta_n\) the terms \(m<n\) contribute at most \(\eta_nS_{n-1}\le\delta_n/8\) and the terms \(m>n\) at most \(2\sum_{m>n}\delta_m=\frac{2}{99}\delta_n<\delta_n/8\), so
\begin{equation*} \Big|\big(f(y)-f(x)\big)-\big(g_n(y)-g_n(x)\big)\Big|\ \le\ \frac{\delta_n}{4} . \tag{7} \end{equation*}
No derivative anywhere: fix \(x\in(0,1)\) and \(n\) with \([x-\eta_n,x+\eta_n]\subset(0,1)\).
(i) \(g_n(x)\le\delta_n/2\): let \(u^{\pm}\) be the nearest points of \(\eta_n\mathbb{Z}\) on either side of \(x\), where \(g_n=\delta_n\); by (7), \(f(u^{\pm})-f(x)\ge\delta_n-\delta_n/2-\delta_n/4=\delta_n/4>0\).
(ii) \(g_n(x)>\delta_n/2\): let \(u^{\pm}\) be the nearest points of \((\tfrac12+\mathbb{Z})\eta_n\) on either side, where \(g_n=0\) since \(\ell_n<\eta_n\); by (7), \(f(u^{\pm})-f(x)\le-\delta_n/2+\delta_n/4<0\).
Either way \(0<|u^{\pm}-x|\le\eta_n\) and the increments have equal signs, so the two difference quotients have opposite signs and modulus at least \(\delta_n/(4\eta_n)\ge1/(4\delta_n)\to+\infty\), using \(\eta_n\le\delta_n^2\). Hence
\begin{equation*} \limsup_{h\to0}\frac{f(x+h)-f(x)}{h}=+\infty,\qquad \liminf_{h\to0}\frac{f(x+h)-f(x)}{h}=-\infty . \end{equation*}
Approximate differentiability a.e.: put \(V_N:=\bigcup_{n\ge N}U_n\), so \(\lambda(V_N\cap(0,1))\le2^{-N+2}\), and let \(Z\) be the union of the null sets \(\bigcap_NV_N\), the points failing to be density points of some \([0,1]\setminus V_N\) (Lebesgue density theorem, from Theorem 5.6.2 applied to the indicators of these countably many sets), and the countable set where some \(g_m\) is not differentiable. For \(x\in(0,1)\setminus Z\) choose \(N\) with \(x\notin U_n\) for \(n\ge N\) and put \(L:=\sum_{m<N}g_m^{\prime}(x)\). If \(y\notin V_N\) then \(g_n(y)=0=g_n(x)\) for \(n\ge N\), so by ordinary differentiability of the finite sum at \(x\),
\begin{equation*} f(y)-f(x)=\sum_{m<N}\big(g_m(y)-g_m(x)\big)=L(y-x)+o(|y-x|)\quad(y\to x). \end{equation*}
Hence for each \(\varepsilon>0\) the set \(A_\varepsilon:=\{y:\ |f(y)-f(x)-L(y-x)|>\varepsilon|y-x|\}\) is contained in \(V_N\) near \(x\), and \(V_N\) has density \(0\) at \(x\); so every \(A_\varepsilon\) has density \(0\) at \(x\), i.e. \(f\) is approximately differentiable at \(x\) with \(\operatorname{ap}f^{\prime}(x)=L\). Since \(\lambda(Z)=0\), this holds almost everywhere.
(Lusin [633, 46]) (i) Let \(\psi\) be a continuous function on \([0,1]\) such that one has \(\psi^{\prime}(x)=0\) a.e. Show that there exists a set \(E\subset[0,1]\) of measure \(1\) such that \(\psi(E)\) has measure zero.
(ii) Let \(\psi\) be a non-constant continuous function on \([0,1]\) such that \(\psi^{\prime}(x)=0\) a.e. Show that there exists a set \(M\) of measure zero such that \(\psi(M)\) has positive measure.
(i) Take \(E=\bigcup_nE_n\) \(\sigma\)-compact with \(E_n\) compact, \(\lambda(E)=1\) and \(\psi^{\prime}=0\) on \(E\); such an \(E\) exists by inner regularity inside the measurable set \(\{\psi^{\prime}=0\}\) of measure \(1\). Since \(\psi\) is differentiable at every point of \(E\), Proposition 5.5.4 gives that \(\psi(E)\) is measurable with
\begin{equation*} \lambda\big(\psi(E)\big)\ \le\ \int_E|\psi^{\prime}(x)|\,dx=0 . \end{equation*}
(ii) Take \(M:=\big([0,1]\setminus E\big)\cap\psi^{-1}(B)\), where \(E\) is as in (i) and \(B:=\psi([0,1])\setminus\psi(E)\). Here \(\psi(E)=\bigcup_n\psi(E_n)\) is \(\sigma\)-compact, hence Borel, and null by (i), while \(\psi\) is continuous and non-constant on \([0,1]\), so \(\psi([0,1])=[\alpha,\beta]\) with \(\alpha<\beta\) and
\begin{equation*} \lambda(B)=(\beta-\alpha)-\lambda\big(\psi(E)\cap[\alpha,\beta]\big) =\beta-\alpha>0 . \end{equation*}
Then \(M\subset[0,1]\setminus E\) is Borel with \(\lambda(M)=0\), and \(\psi(M)=B\): the inclusion \(\subset\) holds by definition, and if \(b\in B\) then any \(x\) with \(\psi(x)=b\) has \(x\notin E\) (otherwise \(b\in\psi(E)\)), so \(x\in M\).
(\(\circ\)) Show that every absolutely continuous function has Lusin’s property (N), i.e., takes all measure zero sets to measure zero sets.
Absolute continuity controls oscillations, not just increments. Let \(f\) be absolutely continuous on \([a,b]\), let \(\lambda(Z)=0\) with \(Z\subset(a,b)\) (the two endpoints add at most two points to \(f(Z)\)), fix \(\varepsilon>0\) and take the corresponding \(\delta>0\). If \((\alpha_i,\beta_i)\), \(i\le N\), are pairwise disjoint with \(\sum_i(\beta_i-\alpha_i)<\delta\), put \(\omega_i:=\operatorname{osc}(f,[\alpha_i,\beta_i])\) and choose \(c_i<d_i\) in \([\alpha_i,\beta_i]\) with \(|f(d_i)-f(c_i)|=\omega_i\) (the extrema are attained, \(f\) being continuous; discard \(\omega_i=0\)); the intervals \((c_i,d_i)\) are disjoint of total length \(<\delta\), so
\begin{equation*} \sum_{i=1}^{N}\omega_i=\sum_{i=1}^N|f(d_i)-f(c_i)|<\varepsilon . \tag{8} \end{equation*}
Choose an open \(U\) with \(Z\subset U\subset(a,b)\) and \(\lambda(U)<\delta\), and write \(U=\bigcup_j(\alpha_j,\beta_j)\) with the intervals pairwise disjoint. The continuous image of an interval is an interval of length its oscillation, so \(\lambda^*\big(f((\alpha_j,\beta_j))\big)\le\omega_j\), while (8) applied to the first \(N\) intervals and \(N\to\infty\) gives \(\sum_j\omega_j\le\varepsilon\). Since \(f(Z)\subset\bigcup_jf\big((\alpha_j,\beta_j)\big)\), subadditivity yields \(\lambda^*(f(Z))\le\varepsilon\), and \(\varepsilon>0\) was arbitrary.
Suppose that a function \(f\) on \([a,b]\) is differentiable at all points of some set \(E\). Show that \(f(E)\) has measure zero precisely when \(f^{\prime}(x)=0\) a.e. on \(E\).
Proposition 5.5.4 gives one implication and Lemma 5.8.13 the other. Since \(E\) is not assumed measurable, both conditions are read with outer measure: \(\lambda^{*}(f(E))=0\) and \(\lambda^{*}(\{x\in E:\ f^{\prime}(x)\ne0\})=0\). By Theorem 5.8.12 the set \(D\) of points with a finite derivative is Borel and \(f^{\prime}|_D\) is Borel, so \(D_0:=\{x\in D:\ f^{\prime}(x)=0\}\) is Borel; by hypothesis \(E\subset D\).
Sufficiency: if \(A:=\{x\in E:\ f^{\prime}(x)\ne0\}\) is null, then \(E\subset D_0\cup A\), so \(f(E)\subset f(D_0)\cup f(A)\). Since \(D_0\) is measurable and \(f\) is differentiable at each of its points, Proposition 5.5.4 gives
\begin{equation*} \lambda\big(f(D_0)\big)\ \le\ \int_{D_0}|f^{\prime}(x)|\,dx = 0 , \end{equation*}
and writing \(A=\bigcup_nA_n\) with \(A_n:=A\cap\{|f^{\prime}|\le n\}\) null, the last assertion of the same proposition gives \(\lambda(f(A_n))\le n\lambda(A_n)=0\). Hence \(\lambda^{*}(f(E))=0\).
Necessity: if \(\lambda^{*}(f(E))=0\), choose a Borel \(Z\supset f(E)\) with \(\lambda(Z)=0\) and let \(E^{*}\) be the set where \(f^{\prime}\) exists and is nonzero. Lemma 5.8.13 applied to the null set \(Z\) gives \(\lambda(f^{-1}(Z)\cap E^{*})=0\), while
\begin{equation*} \{x\in E:\ f^{\prime}(x)\ne 0\}\ \subset\ f^{-1}(Z)\cap E^{*} , \end{equation*}
since such an \(x\) lies in \(E^{*}\) and has \(f(x)\in Z\). Thus \(f^{\prime}=0\) a.e. on \(E\).
Exercises 5.8.51–5.8.57
(S. Banach, M.A. Zareckii) Prove that a function \(f\) on \([0,1]\) is absolutely continuous precisely when it is continuous, is of bounded variation and possesses Lusin’s property (N).
Necessity is immediate: an absolutely continuous \(f\) is continuous (one interval in the definition), of bounded variation by Proposition 5.3.4, and has Lusin’s property (N) by Exercise 5.8.49.
Sufficiency. Let \(f\) be continuous, of bounded variation and with property (N). By Theorem 5.2.6 it has a finite derivative a.e.; let \(D\subset(0,1)\) be the set of such points, Borel with \(f^{\prime}\) Borel on it by Theorem 5.8.12, and \(\lambda([0,1]\setminus D)=0\). Writing \(f=f_1-f_2\) with \(f_1(x):=V(f,[0,x])\) and \(f_2:=f_1-f\) both nondecreasing (Proposition 5.2.2(i)), Corollary 5.2.7 makes \(f_1^{\prime},f_2^{\prime}\) integrable with \(|f^{\prime}|\le f_1^{\prime}+f_2^{\prime}\) a.e., so \(f^{\prime}\), extended by \(0\) off \(D\), lies in \(L^1[0,1]\).
Fix \(0\le a<b\le1\). Then \(f([a,b])\) is a compact interval containing \(f(a),f(b)\) by continuity, property (N) kills the image of the null set \([a,b]\setminus D\), and Proposition 5.5.4 applies to the measurable set \([a,b]\cap D\), at each point of which \(f\) is differentiable; hence
\begin{equation*} |f(b)-f(a)|\ \le\ \lambda\big(f([a,b])\big)\ \le\ \lambda\big(f([a,b]\cap D)\big)\ \le\ \int_a^b|f^{\prime}(x)|\,dx . \end{equation*}
Given \(\varepsilon>0\), absolute continuity of the Lebesgue integral supplies \(\delta>0\) with \(\int_S|f^{\prime}|\,dx<\varepsilon\) whenever \(\lambda(S)<\delta\); so for pairwise disjoint \((a_i,b_i)\subset[0,1]\) of total length \(<\delta\),
\begin{equation*} \sum_{i=1}^n|f(b_i)-f(a_i)|\ \le\ \int_{\bigcup_i(a_i,b_i)}|f^{\prime}(x)|\,dx\ <\ \varepsilon . \end{equation*}
Let \(f\) be an absolutely continuous function on \([0,1]\) such that for a.e. \(x\) one has \(f^{\prime}(x) > 0\). Prove that \(f\) is strictly increasing and the inverse function is absolutely continuous on \([f(0), f(1)]\).
By Theorems 5.3.6 and 5.4.2, \(f(t)-f(s)=\int_s^tf^{\prime}(x)\,dx\) for \(0\le s\le t\le1\); for \(s<t\) the integrand is nonnegative a.e. and cannot vanish a.e. on \([s,t]\), since \(f^{\prime}>0\) a.e. and \(t-s>0\), so \(f(t)>f(s)\) and \(f\) is strictly increasing. Hence \(f\) is a homeomorphism of \([0,1]\) onto \([c,d]:=[f(0),f(1)]\), \(c<d\), and \(g:=f^{-1}\) is continuous and strictly increasing, so of bounded variation with \(V(g,[c,d])=1\). By the Banach–Zareckii theorem (Exercise 5.8.51) it remains to check property (N) for \(g\).
Let \(D\) be the set of points of \((0,1)\) where \(f\) has a finite nonzero derivative; it is measurable by Theorem 5.8.12, and \(\lambda([0,1]\setminus D)=0\) since \(f\) is differentiable a.e. with positive derivative a.e. Let \(Z\subset[c,d]\) with \(\lambda(Z)=0\) and put \(A:=g(Z)\), so \(f(A)=Z\) and \(A\subset f^{-1}(Z)\). Lemma 5.8.13 applied to \(f\) and \(Z\) gives \(\lambda(f^{-1}(Z)\cap E)=0\), where \(E\supset D\) is the set of points with a nonzero derivative; hence
\begin{equation*} \lambda^{*}(A)\ \le\ \lambda\big(f^{-1}(Z)\cap D\big) +\lambda\big([0,1]\setminus D\big)\ =\ 0 . \end{equation*}
So \(g\) has property (N) and is absolutely continuous by Exercise 5.8.51.
Let \(f\) be a continuous function on \([0,1]\) and let \(D\) be the set of all points of differentiability of \(f\) on \((0,1)\). Prove that \(f\) is absolutely continuous precisely when \(f^{\prime}\) is integrable on \(D\) and \(f([0,1] \setminus D)\) has measure zero. In particular, if \(f\) is differentiable everywhere in \((0,1)\) and \(f^{\prime} \in L^1[0,1]\), then \(f\) is absolutely continuous.
Both directions run on Proposition 5.5.4 and Lusin’s property (N); here \(D\) is the set of points of \((0,1)\) with a finite derivative, Borel with \(f^{\prime}\) Borel on it by Theorem 5.8.12.
Necessity. If \(f\) is absolutely continuous, Theorem 5.3.6 gives an integrable \(h\) with \(f(x)=f(0)+\int_0^xh(t)\,dt\), and \(f^{\prime}=h\) a.e. by Theorem 5.4.2; so \(\lambda([0,1]\setminus D)=0\) and \(\int_D|f^{\prime}|=\int_0^1|h|<\infty\), while property (N) (Exercise 5.8.49) applied to the null set \([0,1]\setminus D\) makes \(f([0,1]\setminus D)\) null.
Sufficiency. Extend \(f^{\prime}\) by \(0\) off \(D\), so \(|f^{\prime}|\mathbf{1}_D\in L^1[0,1]\), and fix \(0\le a<b\le1\). Continuity makes \(f([a,b])\) a compact interval containing \(f(a),f(b)\); \(f([a,b]\setminus D)\subset f([0,1]\setminus D)\) is null by hypothesis; and Proposition 5.5.4 applies to the measurable set \([a,b]\cap D\), at each point of which \(f\) is differentiable. Hence
\begin{equation*} |f(b)-f(a)|\ \le\ \lambda\big(f([a,b])\big)\ \le\ \int_{[a,b]\cap D}|f^{\prime}(x)|\,dx . \end{equation*}
Given \(\varepsilon>0\), absolute continuity of the integral of \(|f^{\prime}|\mathbf{1}_D\) gives \(\delta>0\) with \(\int_{S\cap D}|f^{\prime}|\,dx<\varepsilon\) whenever \(\lambda(S)<\delta\); for pairwise disjoint \((a_i,b_i)\) of total length \(<\delta\) and \(S:=\bigcup_i(a_i,b_i)\), the sets \([a_i,b_i]\cap D\) being disjoint up to finitely many points,
\begin{equation*} \sum_{i=1}^n|f(b_i)-f(a_i)|\ \le\ \int_{S\cap D}|f^{\prime}(x)|\,dx\ <\ \varepsilon . \end{equation*}
If \(f\) is differentiable everywhere in \((0,1)\) with \(f^{\prime}\in L^1[0,1]\), then \(D=(0,1)\) and \(f([0,1]\setminus D)=\{f(0),f(1)\}\) is null, so \(f\) is absolutely continuous.
(M.A. Zareckii) Let \(f\) be a continuous strictly increasing function on an interval \([a,b]\). (i) Prove that \(f\) is absolutely continuous precisely when \(f\) takes the set \(\{x \colon f^{\prime}(x) = +\infty\}\) to a measure zero set.
(ii) Let \(g\) be the inverse function for \(f\). Prove that \(g\) is absolutely continuous precisely when the set \(\{x \colon f^{\prime}(x) = 0\}\) has measure zero.
(i) Both directions go through the Banach–Zareckii theorem (Exercise 5.8.51). Here \(f\) is a homeomorphism of \([a,b]\) onto \([c,d]:=[f(a),f(b)]\) with inverse \(g\), and \([a,b]=D\cup A_\infty\cup E\), where \(D\) is the set of finite differentiability, \(A_\infty:=\{x:\ f^{\prime}(x)=+\infty\}\) (the value \(-\infty\) is excluded by monotonicity) and \(E\) is the rest; these are measurable by Theorem 5.8.12, and \(\lambda(A_\infty)=\lambda(E)=0\) because a monotone function is a.e. finitely differentiable (Theorem 5.2.6). All difference quotients of \(f\) are positive, so all Dini derivates are nonnegative; \(\overline Df,\underline Df\) denote the two-sided ones.
Lemma A (Proposition 5.5.4 for a derivate). If \(h\) is continuous and strictly increasing on \([a,b]\), \(L\ge0\), and \(\underline Dh\le L\) on \(S\subset(a,b)\), then \(\lambda^{*}(h(S))\le L\lambda^{*}(S)\). Indeed, fix \(\varepsilon>0\) and open \(S\subset U\subset(a,b)\) with \(\lambda(U)\le\lambda^{*}(S)+\varepsilon\). Each \(x\in S\) admits \(t_n\to0\) with \(|h(x+t_n)-h(x)|\le(L+\varepsilon)|t_n|\) and with the closed interval \(\Delta\) from \(x\) to \(x+t_n\) inside \(U\); the images \(I=h(\Delta)\) are nondegenerate closed intervals containing \(h(x)\) of length at most \((L+\varepsilon)|t_n|\to0\), so they form a Vitali cover of \(h(S)\) and Theorem 5.5.1 yields a disjoint subfamily \(I_j=h(\Delta_j)\) with \(\lambda(h(S)\setminus\bigcup_jI_j)=0\). Disjointness of the \(I_j\) forces disjointness of the \(\Delta_j\subset U\) (as \(h(\Delta_j\cap\Delta_k)\subset I_j\cap I_k\)), whence
\begin{equation*} \lambda^{*}(h(S))\le\sum_j|I_j|\le(L+\varepsilon)\sum_j|\Delta_j| \le(L+\varepsilon)\big(\lambda^{*}(S)+\varepsilon\big) , \end{equation*}
and \(\varepsilon\to0\) finishes.
Lemma B. \(\lambda(f(E))=0\). For \(x\in E\) one has \(\underline Df(x)<+\infty\) (otherwise \(\overline Df(x)=+\infty\) too and \(f^{\prime}(x)=+\infty\) would exist), so \(E\setminus\{a,b\}=\bigcup_nS_n\) with \(S_n:=\{x\in E\cap(a,b):\ \underline Df(x)\le n\}\) null, and Lemma A gives \(\lambda^{*}(f(S_n))\le n\cdot0=0\).
If \(f\) is absolutely continuous it is a.e. finitely differentiable, so \(\lambda(A_\infty)=0\), and property (N) (Exercise 5.8.49) gives \(\lambda(f(A_\infty))=0\). Conversely, if \(\lambda(f(A_\infty))=0\) and \(\lambda(S)=0\), then
\begin{equation*} f(S)\ \subset\ f(S\cap D)\cup f(A_\infty)\cup f(E) \end{equation*}
is null: the first piece by Proposition 5.5.4, applicable since the null set \(S\cap D\) is measurable with \(f\) differentiable on it, so \(\lambda(f(S\cap D))\le\int_{S\cap D}|f^{\prime}|\,dx=0\); the second by hypothesis; the third by Lemma B. Thus \(f\) is continuous, monotone (hence of bounded variation) and has property (N), so Exercise 5.8.51 applies.
(ii) Apply (i) to \(g\), again continuous and strictly increasing. For \(y\in(c,d)\) and \(x=g(y)\), the substitution \(s:=f(x+t)-f(x)\) is a bijection between punctured neighbourhoods of \(0\), with \(t\to0\) iff \(s\to0\), and
\begin{equation*} \frac{g(y+s)-g(y)}{s}=\left(\frac{f(x+t)-f(x)}{t}\right)^{-1} , \end{equation*}
both quotients being positive; so \(g^{\prime}(y)=+\infty\) exactly when \(f^{\prime}(x)=0\). Hence \(g(\{y:\ g^{\prime}(y)=+\infty\})\) differs from \(Z:=\{x:\ f^{\prime}(x)=0\}\) by at most the two points \(a,b\), and by (i) applied to \(g\) the function \(g\) is absolutely continuous precisely when \(\lambda(Z)=0\).
Let \(f\) be a continuous function with property (N) on \([a,b]\). Prove that for almost every \(y\) the set \(f^{-1}(y)\) is at most countable.
The witness is a measurable set \(E_0\) of maximal measure on which \(f\) is countable-to-one and which catches almost all values.
Selection. For compact \(K\subset[a,b]\) put
\begin{equation*} E_K:=\{x\in K:\ f(z)\ne f(x)\ \text{for all}\ z\in K\ \text{with}\ z<x\} , \end{equation*}
the leftmost points of the level sets. Then \(K\setminus E_K=\bigcup_nB_n\) with \(B_n:=\{x\in K:\ \exists z\in K,\ z\le x-1/n,\ f(z)=f(x)\}\) closed (take \(x_k\to x\) in \(B_n\), extract a convergent subsequence from the witnesses \(z_k\in K\), and use continuity of \(f\)), so \(E_K\) is Borel; \(f(E_K)=f(K)\), since \(\min(f^{-1}(y)\cap K)\in E_K\) for \(y\in f(K)\); and \(f\) is injective on \(E_K\) by construction.
Let \(\mathcal{C}\) consist of the measurable \(E\subset[a,b]\) with (a) \(\lambda^{*}(f([a,b])\setminus f(E))=0\) (outer measure, as images of measurable sets need not be measurable) and (b) \(f^{-1}(y)\cap E\) at most countable for every \(y\). It contains \(E_{[a,b]}\) and is stable under countable unions, so for \(\beta:=\sup\{\lambda(E):\ E\in\mathcal{C}\}\) the union \(E_0\) of a maximizing sequence lies in \(\mathcal{C}\) with \(\lambda(E_0)=\beta\).
Then \(\lambda^{*}(f([a,b]\setminus E_0))=0\). Otherwise set \(A:=[a,b]\setminus E_0\); property (N) excludes \(\lambda(A)=0\), and inner regularity gives \(A=N\cup\bigcup_nK_n\) with \(K_n\) compact, \(\lambda(N)=0\), so
\begin{equation*} 0<\lambda^{*}(f(A))\le\lambda\big(f(N)\big)+\sum_n\lambda\big(f(K_n)\big) =\sum_n\lambda\big(f(K_n)\big) \end{equation*}
and \(\lambda(f(K))>0\) for some compact \(K\subset A\). Then \(E_1:=E_K\) has \(\lambda(f(E_1))=\lambda(f(K))>0\), hence \(\lambda(E_1)>0\) by property (N), while \(E_0\cup E_1\in\mathcal{C}\) is a disjoint union with \(\lambda(E_0\cup E_1)=\beta+\lambda(E_1)>\beta\) – a contradiction.
So \(Y_0:=f([a,b]\setminus E_0)\) is null, and for \(y\notin Y_0\) one has \(f^{-1}(y)=f^{-1}(y)\cap E_0\), at most countable by (b).
Let \(f\) be a continuous function on \([a,b]\) with property (N). Let \(P\) be the set of all points where \(f\) has a finite nonnegative derivative and let \(N\) be the set of all points where \(f\) has a finite nonpositive derivative. Prove that
\begin{equation*} -\lambda\big(f(N)\big) \le f(b) - f(a) \le \lambda\big(f(P)\big) . \end{equation*}
Deduce the existence of points of differentiability of \(f\).
It suffices to prove \(f(b)-f(a)\le\lambda(f(P))\) for every continuous \(f\) with property (N): applied to \(-f\), which is continuous with property (N) and whose set of points with a finite nonnegative derivative is exactly \(N\) with \(\lambda(-f(N))=\lambda(f(N))\), it gives the left-hand inequality. The sets \(P,N\) are Borel by Theorem 5.8.12, and their images are measurable, since \(P=M\cup\bigcup_nS_n\) with \(S_n\) compact and \(\lambda(M)=0\) makes \(f(P)=f(M)\cup\bigcup_nf(S_n)\) measurable by property (N), likewise \(f(N)\). We may assume \(f(a)<f(b)\), the claim being trivial otherwise. Write \(\overline Df,\underline Df\) for the two-sided derivates and \(D^{\pm}f, D_{\pm}f\) for the one-sided ones of Theorem 5.8.12.
Lemma. A compact at most countable \(C\subset\mathbb{R}\) with at least two points has isolated points \(x_1<x_2\) with \((x_1,x_2)\cap C=\varnothing\). Let \(F\subset C\) be the closed set of non-isolated points. If \(F=\varnothing\), then \(C\) is compact without accumulation points, hence finite, and two consecutive points work. Otherwise \(F\) is compact and at most countable, hence not perfect, so some \(z\) has \((z-\delta,z+\delta)\cap F=\{z\}\) and every point of \(C\) in that interval other than \(z\) is isolated in \(C\). Say \(C\) accumulates at \(z\) from the right (the other case is symmetric); pick \(x_2\in C\cap(z,z+\delta)\) and put \(x_1:=\sup(C\cap[z,x_2))\in C\). Then \(x_1<x_2\) as \(x_2\) is isolated, \(x_1>z\) by right-accumulation, so \(x_1\) is isolated, and \((x_1,x_2)\cap C=\varnothing\) by definition of the supremum.
Selection. By Exercise 5.8.55 the \(y\) with \(f^{-1}(y)\) uncountable form a null set, so
\begin{equation*} Y:=\{y\in(f(a),f(b)):\ E_y:=f^{-1}(y)\ \text{is at most countable}\} \end{equation*}
is measurable with \(\lambda(Y)=f(b)-f(a)\); each \(E_y\) is compact, nonempty by the intermediate value theorem, and misses \(a,b\). For \(y\in Y\) there is \(x_y\in E_y\), isolated in \(E_y\), with \(\overline Df(x_y)\ge0\):
(i) \(E_y=\{x_0\}\): since \(f-y\) has no zero on \([a,x_0)\) or on \((x_0,b]\) and \(f(a)<y<f(b)\), we get \(f<y\) to the left and \(f>y\) to the right, so every difference quotient at \(x_0\) is positive; take \(x_y:=x_0\).
(ii) \(E_y\) has at least two points: the Lemma applied to \(C:=E_y\) gives isolated \(x_1<x_2\) with \(f-y\) of constant sign on \((x_1,x_2)\). If \(f>y\) there, the quotients \([f(t)-f(x_1)]/(t-x_1)\) are positive for \(t\in(x_1,x_2)\) and \(x_y:=x_1\) works; if \(f<y\) there, the quotients at \(x_2\) are positive and \(x_y:=x_2\) works.
Put \(X:=\{x_y:\ y\in Y\}\), so \(f(X)=Y\) and \(f\) is injective on \(X\).
Differentiability on \(X\). For \(x=x_y\in X\), isolation supplies \(\delta>0\) with \(f\ne f(x)\) on \(0<|t-x|<\delta\), where \(f-f(x)\) has constant sign on each side. If the signs agree, \(x\) is a strict local extremum, and such points are at most countable: assign to each a rational pair \(p<x<q\) inside its interval of strictness, injectively, since two strict maxima sharing a pair would each dominate the other. If the signs differ, either all quotients near \(x\) are positive, giving \(D_{+}f(x),D_{-}f(x)\ge0\), or all are negative, giving \(D^{+}f(x),D^{-}f(x)\le0\). Alternatives (b), (c), (d) of Theorem 5.8.12 each require a one-sided lower derivate \(-\infty\) and a one-sided upper derivate \(+\infty\), so off the null set \(Z\) of that theorem both possibilities leave only alternative (a): \(f^{\prime}(x)\) exists finitely. Hence \(X_0:=\{x\in X:\ f^{\prime}(x)\ \text{exists and is finite}\}\) has \(\lambda(X\setminus X_0)=0\), and \(f^{\prime}(x)=\overline Df(x)\ge0\) on \(X_0\), i.e. \(X_0\subset P\). Property (N) gives \(\lambda(f(X\setminus X_0))=0\), so
\begin{equation*} f(b)-f(a)=\lambda(Y)=\lambda^{*}\big(f(X)\big)\le \lambda^{*}\big(f(X_0)\big)+\lambda\big(f(X\setminus X_0)\big) \le\lambda\big(f(P)\big) . \end{equation*}
Points of differentiability: if \(f\) is constant, every point is one. Otherwise take \(c\in(a,b]\) with \(f( c)\ne f(a)\); the two inequalities on \([a,c]\), whose own sets \(P,N\) sit inside \(P\cup\{a,c\}\) and \(N\cup\{a,c\}\), give \(-\lambda(f(N))\le f( c)-f(a)\le\lambda(f(P))\) with \(f( c)-f(a)\ne0\), so \(\lambda(f(P))>0\) or \(\lambda(f(N))>0\); in particular \(P\cup N\ne\varnothing\).
(i) Let \(f\) be a continuous function on \([a,b]\) with property (N) and let \(g\) be an integrable function on \([a,b]\) such that \(f^{\prime}(x) \le g(x)\) at almost every point \(x\) where \(f^{\prime}(x)\) exists. Show that \(f\) is absolutely continuous.
(ii) Show that a continuous function \(f\) on \([a,b]\) is absolutely continuous precisely when it has property (N) and the function \(f^{\prime}(x)\) is integrable over the set \(P\) of all points at which it exists and is finite and nonnegative. In particular, if \(f\) is continuous, has property (N), is a.e. differentiable and \(f^{\prime}\) is integrable, then \(f\) is absolutely continuous.
(iii) Show that every continuous function \(f\) on \([a,b]\) with property (N) is differentiable on a set of positive measure (but not necessarily a.e.).
(i) Exercise 5.8.56 bounds each increment by \(\int|g|\), which forces bounded variation, and Banach–Zareckii finishes. Let \(D\) be the set of finite differentiability and \(P:=\{x\in D:\ f^{\prime}(x)\ge0\}\), both Borel with \(f^{\prime}\) Borel on \(D\) by Theorem 5.8.12. For \(a\le\alpha<\beta\le b\), the restriction of \(f\) to \([\alpha,\beta]\) is continuous with property (N) and its own \(P\)-set differs from \(P\cap[\alpha,\beta]\) by at most the two endpoints, so Exercise 5.8.56 there, then Proposition 5.5.4 (applicable, since \(P\cap[\alpha,\beta]\) is measurable and \(f\) is differentiable on it), then \(0\le f^{\prime}\le g\) a.e. on \(P\), give
\begin{equation*} f(\beta)-f(\alpha)\ \le\ \lambda\big(f(P\cap[\alpha,\beta])\big)\ \le\ \int_{P\cap[\alpha,\beta]}f^{\prime}(x)\,dx\ \le\ \int_{\alpha}^{\beta}|g(x)|\,dx . \tag{9} \end{equation*}
For a partition \(a=x_0<\dots<x_n=b\) put \(\Delta_k:=f(x_{k+1})-f(x_k)\), \(S^{+}:=\sum_{\Delta_k\ge0}\Delta_k\) and \(S^{-}:=\sum_{\Delta_k<0} (-\Delta_k)\); (9) and the disjointness of the intervals give \(S^{+}\le\|g\|_{L^1}\), while \(S^{+}-S^{-}=f(b)-f(a)\), so
\begin{equation*} \sum_k|\Delta_k|=S^{+}+S^{-}\ \le\ 2\|g\|_{L^1}+|f(b)-f(a)| . \end{equation*}
This bound is partition-free, so \(f\) is of bounded variation; continuous, of bounded variation and with property (N), it is absolutely continuous by Exercise 5.8.51.
(ii) Necessity: an absolutely continuous \(f\) has property (N) by Exercise 5.8.49 and \(f^{\prime}\in L^1[a,b]\) by Theorems 5.3.6 and 5.4.2, hence is integrable over \(P\). Sufficiency: set \(g:=f^{\prime}\) on \(P\) and \(g:=0\) off \(P\), measurable by Theorem 5.8.12 and integrable because \(\int_a^b|g|=\int_P|f^{\prime}|<\infty\). At every point of finite differentiability \(f^{\prime}\le g\): if \(f^{\prime}(x)\ge0\) then \(x\in P\) and \(g(x)=f^{\prime}(x)\); otherwise \(g(x)=0>f^{\prime}(x)\). That is the hypothesis of (i), whose proof uses only such points, so \(f\) is absolutely continuous. In particular a continuous \(f\) with property (N) that is a.e. differentiable with \(f^{\prime}\in L^1[a,b]\) has \(\int_P|f^{\prime}|<\infty\) and is absolutely continuous.
(iii) If \(\lambda(D)=0\), with \(D\) measurable by Theorem 5.8.12, then \(P\subset D\) is null and \(\int_P|f^{\prime}|\,dx=0<\infty\), so (ii) makes \(f\) absolutely continuous, hence differentiable a.e. (Proposition 5.3.4 and Theorem 5.2.6) and \(\lambda(D)=b-a>0\) – a contradiction; thus \(\lambda(D)>0\). Almost everywhere differentiability can fail: Ruziewicz [836] constructed a continuous function with property (N) that is not differentiable a.e.
Exercises 5.8.58–5.8.64
(Menchoff) Let \(\psi\) be a continuous function on \([0,1]\) that is not a constant and let \(\psi^{\prime}(x) = 0\) a.e. Then, for every absolutely continuous function \(\varphi\) on \([0,1]\), the function \(\psi + \varphi\) has no property (N). In particular, the sum of any absolutely continuous function with the Cantor function has no property (N).
Suppose \(h:=\psi+\varphi\) had property (N). By Theorems 5.3.6 and 5.4.2 the function \(\varphi\) is differentiable a.e. with \(\varphi^{\prime}\in L^1[0,1]\), so off a null set \(Z\) both \(\varphi^{\prime}(x)\) and \(\psi^{\prime}(x)=0\) exist and
\begin{equation*} h^{\prime}(x)=\psi^{\prime}(x)+\varphi^{\prime}(x)=\varphi^{\prime}(x) \le|\varphi^{\prime}(x)|=:g(x) , \end{equation*}
with \(g\) integrable. So \(h\) is continuous, has property (N), and satisfies \(h^{\prime}\le g\) at almost every point of differentiability: Exercise 5.8.57(i) makes \(h\) absolutely continuous. Then \(\psi=h-\varphi\) is absolutely continuous, so the Newton–Leibniz formula and \(\psi^{\prime}=0\) a.e. give \(\psi(x)=\psi(0)+\int_0^x\psi^{\prime}(t)\,dt=\psi(0)\) for all \(x\), contradicting non-constancy.
The Cantor function (Example 3.6.6) qualifies as \(\psi\): it is continuous, non-constant since \(\psi(0)=0\) and \(\psi(1)=1\), and constant on each interval adjacent to the Cantor set, whose union has full measure, so \(\psi^{\prime}=0\) a.e.
(\(\circ\)) Let \(f\) be an absolutely continuous monotone function on an interval \([a,b]\) and let \(\varphi\) be an absolutely continuous function on an interval \([c,d]\) containing \(f([a,b])\). Show that \(\varphi(f)\) is absolutely continuous on \([a,b]\).
Given \(\varepsilon>0\), take \(\delta>0\) from the absolute continuity of \(\varphi\) on \([c,d]\), then \(\tau>0\) from that of \(f\) with \(\delta\) in place of \(\varepsilon\); this \(\tau\) works for \(\varphi\circ f\). We may assume \(f\) increasing, replacing \((f,\varphi)\) by \((-f,\varphi(-\cdot))\) otherwise: reflection is an isometry, so it preserves absolute continuity, and the composition is unchanged.
Let \((a_i,b_i)\subset[a,b]\), \(i\le n\), be pairwise disjoint with \(\sum_i(b_i-a_i)<\tau\), relabelled so that \(a\le a_1<b_1\le a_2<\dots<b_n \le b\), and put \(\alpha_i:=f(a_i)\le\beta_i:=f(b_i)\), all lying in \(f([a,b])\subset[c,d]\). The choice of \(\tau\) gives
\begin{equation*} \sum_{i=1}^n(\beta_i-\alpha_i)=\sum_{i=1}^n|f(b_i)-f(a_i)|<\delta , \end{equation*}
and the nondegenerate intervals \((\alpha_i,\beta_i)\) are pairwise disjoint, since \(i<j\) gives \(b_i\le a_j\), hence \(\beta_i\le\alpha_j\) by monotonicity – this is where monotonicity is essential. Degenerate indices contribute nothing, so the choice of \(\delta\) gives
\begin{equation*} \sum_{i=1}^n\big|\varphi(f(b_i))-\varphi(f(a_i))\big| =\sum_{\alpha_i<\beta_i}|\varphi(\beta_i)-\varphi(\alpha_i)|<\varepsilon , \end{equation*}
which is Definition 5.3.1 for \(\varphi\circ f\).
(\(\circ\)) Find two absolutely continuous functions \(f, g \colon [0,1] \to [0,1]\) such that their composition is not absolutely continuous.
Take \(g(x)=x^2\sin^2(1/x)\) with \(g(0)=0\), and \(f(y)=\sqrt y\); both map \([0,1]\) into \([0,1]\). The function \(g\) is Lipschitz with constant \(3\), since on \((0,1]\)
\begin{equation*} g^{\prime}(x)=2x\sin^2(1/x)-\sin(2/x),\qquad |g^{\prime}(x)|\le 2x+1\le3 , \end{equation*}
and \(|g(x)-g(0)|\le x^2\le x\); hence \(g\) is absolutely continuous. So is \(f\), being the indefinite integral \(\sqrt y=\int_0^y(2\sqrt t)^{-1}dt\) of an integrable function (Theorem 5.3.6). But \(h:=f\circ g\) satisfies \(h(x)=x|\sin(1/x)|\), and for \(y_k:=1/(\pi k)\), \(x_k:=1/(\pi(k+1/2))\), interlacing as \(x_{k+1}<y_{k+1}<x_k<y_k\le1/\pi\), we have \(h(y_k)=0\) and \(h(x_k)=x_k\), so the partition through \(x_n<y_n<\dots<x_1<y_1\) gives
\begin{equation*} V(h)\ \ge\ \sum_{k=1}^{n}\big(|h(x_k)-h(y_k)|+|h(x_k)-h(y_{k+1})|\big) =\frac{2}{\pi}\sum_{k=1}^{n}\frac{1}{k+1/2}\ \xrightarrow[n\to\infty]{}\ +\infty . \end{equation*}
Thus \(h\) is not of bounded variation, hence not absolutely continuous by Proposition 5.3.4.
(\(\circ\)) (Fichtenholz) (i) Let a function \(F\) on \([a,b]\) be such that the composition \(F \circ f\) is absolutely continuous for every absolutely continuous function \(f\) with values in \([a,b]\). Prove that \(F\) is Lipschitzian.
(ii) Let functions \(f \colon [a,b] \to [c,d]\) and \(F \colon [c,d] \to \mathbb{R}^1\) be absolutely continuous. Suppose that \(f\) satisfies the following Fichtenholz condition: there is a natural number \(k\) such that for every \(y\), the set \(f^{-1}(y)\) consists of at most \(k\) intervals (possibly degenerating to points). Prove that the function \(F \circ f\) is absolutely continuous.
(iii) Suppose that a function \(f \colon [a,b] \to [c,d]\) is continuous, but does not satisfy the Fichtenholz condition indicated in (ii). Show that there exists an absolutely continuous function \(F\) on \([c,d]\) such that the function \(F \circ f\) is not absolutely continuous.
(i) Taking \(f\) the identity shows \(F\) is absolutely continuous, hence continuous and bounded, \(|F|\le M\). If \(F\) were not Lipschitz, there would be \(u_n<v_n\) in \([a,b]\) with
\begin{equation*} d_n:=|F(v_n)-F(u_n)|>n\,\ell_n,\qquad \ell_n:=v_n-u_n>0 , \end{equation*}
and \(d_n\le2M\) forces \(\ell_n<2M/n\). Let \(u^{*}\) be a limit point of \((u_n)\) and pick \(n_1<n_2<\dots\) with \(n_k\ge4^k\) and \(|u_{n_k}-u^{*}|\le2^{-k-2}\), so \(|p_{k+1}-p_k|\le2^{-k-1}\) for \(p_k:=u_{n_k}\); write \(q_k:=v_{n_k}\), \(L_k:=q_k-p_k\le 2M\cdot4^{-k}\), \(D_k:=d_{n_k}\ge4^kL_k\), and \(m_k:=\lceil2^{-k}/L_k\rceil\), so that \(2^{-k}\le m_kL_k\le2^{-k}+L_k\).
Build \(f\) on \([0,1]\) (an affine change of variable preserves absolute continuity, so the domain is immaterial): divide \(I_k:=[2^{-k},2^{-k+1}]\) into \(2m_k+1\) equal closed subintervals, let \(f\) be piecewise linear on the first \(2m_k\), running \(m_k\) times through the cycle \(p_k\to q_k\to p_k\), and linear from \(p_k\) to \(p_{k-1}\) on the last one (with \(p_0:=p_1\)); set \(f(0):=u^{*}\). The two definitions agree at each shared endpoint \(2^{-k+1}\), all values are convex combinations of points of \([a,b]\), and on \(I_k\) they stay within \(2^{-k-2}+L_k+2^{-k-1}\to0\) of \(u^{*}\), so \(f\) is continuous on \([0,1]\). It is absolutely continuous:
\begin{equation*} V(f)=\sum_{k\ge1}\big(2m_kL_k+|p_{k-1}-p_k|\big) \le\sum_{k\ge1}\big(3\cdot2^{-k}+2L_k\big)<\infty , \end{equation*}
and on \([t,1]\), \(t>0\), only finitely many \(I_k\) intrude, so there \(f\) is piecewise linear with \(f(x)-f(t)=\int_t^xf^{\prime}(s)\,ds\); since \(\int_0^1|f^{\prime}|=V(f)<\infty\), letting \(t\to0^{+}\) gives \(f(x)=f(0)+\int_0^xf^{\prime}(s)\,ds\) and Theorem 5.3.6 applies. But at the \(2m_k+1\) division points in \(I_k\) the values of \(f\) alternate \(p_k,q_k,\dots,p_k\), so
\begin{equation*} V(F\circ f;[0,1])\ \ge\ \sum_{k\ge1}2m_kD_k\ \ge\ \sum_{k\ge1}2\cdot4^km_kL_k\ \ge\ \sum_{k\ge1}2\cdot4^k2^{-k}=\infty , \end{equation*}
so \(F\circ f\) is not of bounded variation and, by Proposition 5.3.4, not absolutely continuous – contradicting the hypothesis.
(ii) The composition \(g:=F\circ f\) is continuous and has property (N), since both \(f\) and \(F\) have it (Exercise 5.8.49) and \(f(E)\subset[c,d]\) for \(E\subset[a,b]\). For bounded variation, fix a partition \(a=t_0<\dots<t_N=b\), let \(J_i\) be the closed interval with endpoints \(f(t_{i-1}),f(t_i)\) and \(J_i^{\circ}\) its interior; the Newton–Leibniz formula for the absolutely continuous \(F\) gives \(|g(t_i)-g(t_{i-1})|\le\int_{J_i^{\circ}}|F^{\prime}|\), so
\begin{equation*} \sum_{i=1}^{N}|g(t_i)-g(t_{i-1})|\ \le\ \int_c^d|F^{\prime}(y)|\,\Theta(y)\,dy, \qquad \Theta:=\sum_{i=1}^N\mathbf{1}_{J_i^{\circ}} . \end{equation*}
Here \(\Theta\le k\): for each \(i\) with \(y\in J_i^{\circ}\) the intermediate value theorem gives \(x_i\in(t_{i-1},t_i)\) with \(f(x_i)=y\), and two such \(x_i<x_j\) lie in distinct components of \(f^{-1}(y)\), since otherwise \([x_i,x_j]\subset f^{-1}(y)\) would give \(f(t_i)=y\), contradicting \(y\in J_i^{\circ}\); the Fichtenholz condition permits at most \(k\) components. Hence \(V(g;[a,b])\le k\int_c^d|F^{\prime}|=k\,V(F;[c,d])<\infty\), and Exercise 5.8.51 makes \(g\) absolutely continuous.
(iii) Failure of the condition gives, for each \(n\), a level \(y_0\) whose fibre has at least \(n\) components; we claim that for every \(n\ge2\) there are \(c\le\alpha_n<\beta_n\le d\) and an integer \(k_n\ge(n-1)/2\) with
\begin{equation*} V(F\circ f;[a,b])\ \ge\ 2k_n\big(F(\beta_n)-F(\alpha_n)\big) \quad\text{for every nondecreasing } F . \tag{10} \end{equation*}
Order \(n\) components \(C_1<\dots<C_n\) of the closed set \(f^{-1}(y_0)\) (each a point or a closed interval) and put \(r_i:=\max C_i<s_{i+1}:=\min C_{i+1}\); some \(z_i\in(r_i,s_{i+1})\) has \(f(z_i)\ne y_0\), else \(C_i\) and \(C_{i+1}\) would lie in one component. With \(\eta:=\min_i|f(z_i)-y_0|>0\), one of \(P:=\{i:\ f(z_i)>y_0\}\), \(M:=\{i:\ f(z_i)<y_0\}\) has at least \((n-1)/2\) elements. If it is \(P\), put \(\alpha_n:=y_0\), \(\beta_n:=y_0+\eta\le d\), \(k_n:=|P|\): since \(f(r_i)=f(s_{i+1})=y_0\) and \(f(z_i)\ge\beta_n\), each \(i\in P\) contributes two increments of \(F\circ f\) of size at least \(F(\beta_n)-F(\alpha_n)\), and the points \(r_i<z_i<s_{i+1}\), \(i\in P\), increase strictly (for \(i<j\) in \(P\), \(s_{i+1}\le r_j\)), so they serve as partition points and give (10). The case \(|M|\ge(n-1)/2\) is symmetric, with \(\alpha_n:=y_0-\eta\ge c\), \(\beta_n:=y_0\).
Now take \(n_j:=2\cdot4^j+1\), write \(\alpha_j^{\prime},\beta_j^{\prime}, k_j^{\prime}\ge4^j\) for the resulting data and \(\ell_j:=\beta_j^{\prime} -\alpha_j^{\prime}>0\), and set
\begin{equation*} F(y):=\int_c^y\varrho(t)\,dt,\qquad \varrho:=\sum_{j\ge1} \frac{2^{-j}}{\ell_j}\mathbf{1}_{[\alpha_j^{\prime},\beta_j^{\prime}]} . \end{equation*}
Monotone convergence gives \(\int_c^d\varrho\le\sum_j2^{-j}=1\), so \(\varrho\in L^1[c,d]\), and \(F\) is absolutely continuous by Theorem 5.3.6 and nondecreasing. Since \(F(\beta_j^{\prime})-F(\alpha_j^{\prime})\ge2^{-j}\), (10) gives \(V(F\circ f;[a,b])\ge2k_j^{\prime}2^{-j}\ge2^{j+1}\) for every \(j\); so \(F\circ f\) is not of bounded variation, hence not absolutely continuous (Proposition 5.3.4).
(G.M. Fichtenholz) Let \(f\) be an absolutely continuous function on \([a,b]\) and let \(\varphi\) be an absolutely continuous function on an interval \([c,d]\) containing \(f([a,b])\). Show that \(\varphi \circ f\) is absolutely continuous on \([a,b]\) precisely when it is of bounded variation.
Necessity is Proposition 5.3.4. Conversely, \(\varphi\circ f\) is continuous and has property (N): for \(E\subset[a,b]\) with \(\lambda(E)=0\), Exercise 5.8.49 applied to \(f\) gives \(\lambda(f(E))=0\), and applied to \(\varphi\) on \([c,d]\supset f([a,b])\) it gives \(\lambda(\varphi(f(E)))=0\). So if \(\varphi\circ f\) is also of bounded variation, the Banach–Zareckii theorem (Exercise 5.8.51) makes it absolutely continuous.
(i) (Lebesgue) Show that there exist two functions with property (N) such that their sum does not have this property.
(ii) (Mazurkiewicz) There exists a continuous function \(f\) with property (N) such that \(f(x) + cx\) has no property (N) whenever \(c \ne 0\).
(iii) Construct two continuous functions \(f\) and \(g\) with property (N) on \([0,1]\) such that their product \(fg\) has no property (N).
(i) Take for \(f,g\) the two coordinates of a homeomorphism of the Cantor set \(C=\{\sum_{n\ge1}2a_n3^{-n}:\ a_n\in\{0,1\}\}\) onto \(C\times C\), extended affinely across the gaps. Such a homeomorphism exists: \((a_n)\mapsto \sum_n2a_n3^{-n}\) is a continuous bijection of the compact \(\{0,1\}^{\mathbb{N}}\) onto \(C\), hence a homeomorphism, and splitting a sequence into odd- and even-indexed terms gives \(\psi=(\psi_1,\psi_2)\) with
\begin{equation*} \psi_1\Big(\sum_n2a_n3^{-n}\Big)=\sum_n2a_{2n-1}3^{-n},\qquad \psi_2\Big(\sum_n2a_n3^{-n}\Big)=\sum_n2a_{2n}3^{-n} , \end{equation*}
so \(\psi( C)=C\times C\). Since \(0,1\in C\), each interval \((p,q)\) adjacent to \(C\) has \(p,q\in C\); set \(f:=\psi_1\), \(g:=\psi_2\) on \(C\) and affine on each \([p,q]\).
\(f\) is continuous: clear off \(C\), and at \(x\in C\) take \(\varepsilon\in(0,1/2)\) with \(\eta\in(0,1)\) from the uniform continuity of \(\psi_1\) on \(C\), so that \(|\psi_1(u)-\psi_1(v)|\le\varepsilon\) for \(u,v\in C\) with \(|u-v|\le\eta\), and let \(|t-x|<\varepsilon\eta\). If \(t\in C\), then \(|f(t)-f(x)|\le\varepsilon\). Otherwise \(t\in(p,q)\) with, say, \(x\le p\), so \(p-x\le\eta\) and \(|\psi_1(p)-\psi_1(x)|\le\varepsilon\); if \(q-x\le\eta\) then \(f(t)\) lies between two values within \(\varepsilon\) of \(f(x)\), while if \(q-x>\eta\) then \(q-p\ge q-t>\eta/2\) and \(t-p<\varepsilon\eta\), whence
\begin{equation*} |f(t)-\psi_1(p)|=\frac{t-p}{q-p}|\psi_1(q)-\psi_1(p)| \le\frac{\varepsilon\eta}{\eta/2}\cdot1=2\varepsilon . \end{equation*}
In every case \(|f(t)-f(x)|\le3\varepsilon\); likewise for \(g\).
Property (N): for \(\lambda(E)=0\) the image \(f(E\cap C)\subset\psi_1( C)=C\) is null and \(f\) is affine, hence Lipschitz, on each closed gap \([p_k,q_k]\), so \(f(E)\) is a countable union of null sets; likewise \(g\). But \((f+g)( C)=\{s+t:\ s,t\in C\}=C+C=[0,2]\), because \(\psi( C)=C\times C\) and any \(z\in[0,2]\) has \(z/2=\sum_nc_n3^{-n}\) with \(c_n\in\{0,1,2\}\), each \(c_n=a_n+b_n\) for some \(a_n,b_n\in\{0,1\}\). So \(f+g\) maps the null set \(C\) onto a set of measure \(2\).
(ii) The construction reduces to a plane statement that we cite. First, such an \(f\) is necessarily of unbounded variation: a continuous function of bounded variation with property (N) is absolutely continuous (Exercise 5.8.51), so \(f(x)+cx\) would be too, and would have property (N) by Exercise 5.8.49. Nor can \(f=\psi+\varphi\) with \(\varphi\) absolutely continuous and \(\psi\) continuous nonconstant with \(\psi^{\prime}=0\) a.e.: Exercise 5.8.58 applied to the absolutely continuous \(\varphi+cx\) would then deny property (N) also for \(c=0\), which \(f\) must have.
Reduction: it suffices to produce a compact \(D\subset[0,1]\) with \(\lambda(D)=0\) and a continuous \(h\) on \(D\) with \(\lambda(h(D))=0\) and
\begin{equation*} \lambda\big(\{h(x)+cx:\ x\in D\}\big)>0\qquad\text{for every } c\ne0 . \tag{11} \end{equation*}
Extending \(h\) affinely across the gaps of \(D\) and constantly outside \([\min D,\max D]\) gives, by the two arguments of (i), a continuous \(f\) with property (N), while \(f(x)+cx\) carries the null set \(D\) onto the set in (11).
The natural candidate fails. For \(D:=\{\sum_n\varepsilon_ns_n:\ \varepsilon_n\in\{0,1\}\}\) with \(s_n>S_n:=\sum_{m>n}s_m\) and \(2^nS_n\to0\), and \(h(\sum_n\varepsilon_ns_n):=\sum_n\varepsilon_nt_n\), the set in (11) is \(\{\sum_n\varepsilon_nu_n\}\) with \(u_n:=t_n+cs_n\); the only elementary criterion for such sum-sets – if \(u_n\le\sum_{m>n}u_m\) for all \(n\) then the sum-set is \([0,\sum_nu_n]\) – demands \(t_n<T_n:=\sum_{m>n}t_m\) for all \(n\), which by the same criterion forces \(h(D)=[0,T_0]\), contradicting \(\lambda(h(D))=0\). Nor does freezing digits help: on a separated subfamily \(M\) the sub-sum-set is a Cantor set of measure \(\lim_k2^k\tilde U_{n_k}\le \lim_k2^k(\tilde T_{n_k}+|c|\tilde S_{n_k})=0\), so positive measure can only come from the overlapping, non-uniquely-coded part of the sums, exactly as in \(C+C=[0,2]\). And the alternating-scale scheme that would meet the interval criterion for every \(c>0\) needs runs with \(S_n>s_n\), where the coding is no longer injective and a digitwise \(h\) is neither well defined nor continuous (two arbitrarily close points of \(D\) would get \(h\)-values differing by about \(t_n-T_n\)). The existence of \(D\) and \(h\) satisfying (11) is cited from Mazurkiewicz [664]. (A fully rigorous proof of the remaining step is beyond the scope of this page; see the reference given in the book.)
(iii) Take \(F:=\exp f\) and \(G:=\exp g\) with \(f,g\) from (i); both are continuous with property (N), since \(f,g\) have it and \(\exp\) is Lipschitz on \([0,1]\), where \(f\) and \(g\) take their values. But \(FG=\exp(f+g)\), so
\begin{equation*} (FG)( C)=\exp\big((f+g)( C)\big)=\exp\big([0,2]\big)=[1,e^2] , \end{equation*}
a set of positive measure, while \(\lambda( C)=0\).
(Burenkov) (i) Construct an absolutely continuous function \(\Phi\) on the real line and an infinitely differentiable function \(f\) such that the function \(\Phi\bigl(f(x)\bigr)\) is not absolutely continuous on \([0,1]\).
(ii) Let \(\Phi\) be a function of bounded variation on \([c,d]\) and let \(f\) be a differentiable function on \([a,b]\) such that \(f^{\prime}\) is of bounded variation and \(f([a,b]) \subset [c,d]\). Prove that the function \(\Phi\bigl(f(x)\bigr) f^{\prime}(x)\) is of bounded variation on \([a,b]\).
(iii) Let \(\Phi\) be an absolutely continuous function on \([c,d]\) and let \(f\) be a differentiable function on \([a,b]\) such that \(f^{\prime}\) is absolutely continuous and \(f([a,b]) \subset [c,d]\). Prove that the function \(\Phi\bigl(f(x)\bigr) f^{\prime}(x)\) has property (N) and is absolutely continuous on \([a,b]\).
(i) Take \(f(x):=e^{-1/x}\sin^2(1/x)\) for \(x>0\), \(f:=0\) for \(x\le0\), and \(\Phi\) piecewise affine with \(\Phi(y):=0\) for \(y\le0\), \(\Phi(\alpha_k):=1/k\), \(\Phi(y):=1\) for \(y\ge\alpha_1\), where \(\alpha_k:=e^{-\pi(k+1/2)}\). Every derivative of \(f\) on \((0,\infty)\) is a finite sum of terms \(e^{-1/x}q(1/x)\theta(1/x)\) with \(q\) polynomial and \(\theta\) bounded trigonometric, and \(e^{-t}q(t)\to0\) as \(t\to+\infty\), so induction gives \(f\in C^{\infty}(\mathbb{R})\) with \(f^{(n)}(0)=0\) and \(0\le f\le1\). Both \(\alpha_k\) and \(1/k\) decrease to \(0\), so \(\Phi\) is continuous, bounded and nondecreasing, and it is absolutely continuous on every bounded interval: it is piecewise affine with finitely many pieces on \([t,\alpha_1]\), \(t>0\), so \(\Phi(\alpha_1)-\Phi(t)=\int_t^{\alpha_1} \Phi^{\prime}\), and \(t\to0^{+}\) with monotone convergence gives \(\Phi(y)-\Phi(0)=\int_0^y\Phi^{\prime}\), whence Theorem 5.3.6 applies. With \(x_k:=1/(\pi(k+1/2))\) and \(z_k:=1/(\pi k)\), interlacing as \(z_{k+1}<x_k<z_k\), the composition \(g:=\Phi\circ f\) has \(g(z_k)=0\) and \(g(x_k)=\Phi(\alpha_k)=1/k\), so
\begin{equation*} V(g;[0,1])\ \ge\ \sum_{k=1}^{n}\big(|g(x_k)-g(z_k)|+|g(x_k)-g(z_{k+1})|\big) =2\sum_{k=1}^{n}\frac1k\ \xrightarrow[n\to\infty]{}\ \infty , \end{equation*}
so \(g\) is not of bounded variation, hence not absolutely continuous (Proposition 5.3.4).
(ii) Put \(h:=(\Phi\circ f)f^{\prime}\), \(A:=\sup|\Phi|\), \(B:=\sup_{[a,b]}|f^{\prime}|\) (both finite, bounded variation implying boundedness), \(W:=V(f^{\prime};[a,b])\), \(L:=2B+2W\), and \(v(y):=V(\Phi;[c,y])\), so that \(|\Phi(\beta)-\Phi(\alpha)|\le v(\beta)-v(\alpha)\) for \(\alpha\le\beta\). Fix a partition \(a=t_0<\dots<t_N=b\), write \(c_i:=|f^{\prime}(t_{i-1})|\) and let \(J_i\) be the closed interval with endpoints \(f(t_{i-1}),f(t_i)\). The Abel splitting
\begin{equation*} h(t_i)-h(t_{i-1})=\Phi(f(t_i))\big[f^{\prime}(t_i)-f^{\prime}(t_{i-1})\big] +f^{\prime}(t_{i-1})\big[\Phi(f(t_i))-\Phi(f(t_{i-1}))\big] \end{equation*}
gives
\begin{equation*} \sum_{i=1}^N|h(t_i)-h(t_{i-1})|\ \le\ AW+\sum_{i=1}^Nc_i \big(v(\max J_i)-v(\min J_i)\big) . \tag{12} \end{equation*}
The counting estimate is \(\Sigma(y):=\sum_{i:\ y\in J_i}c_i\le L\) for every \(y\in[c,d]\). Write \(\{i:\ y\in J_i\}=\{i_1<\dots<i_r\}\); the intermediate value theorem gives \(x_j\in[t_{i_j-1},t_{i_j}]\) with \(f(x_j)=y\), and \(i_j\le i_{j+1}-1\) yields \(x_j\le t_{i_j}\le t_{i_{j+1}-1}\le x_{j+1}\) and \(\tau_j:=t_{i_j-1}\in[x_{j-1},x_j]\) for \(j\ge2\). For such \(j\):
(a) \(x_{j-1}<x_j\): from \(f(x_{j-1})=f(x_j)=y\) and differentiability, Rolle gives \(\xi\in(x_{j-1},x_j)\) with \(f^{\prime}(\xi)=0\), and \(\tau_j\) lies in the same interval, so \(c_{i_j}=|f^{\prime}(\tau_j)-f^{\prime}(\xi)|\le V(f^{\prime};[x_{j-1},x_j])\).
(b) \(x_{j-1}=x_j\) and \(j<r\): the chain \(x_{j-1}\le t_{i_{j-1}}\le\tau_j \le x_j\) collapses, so \(\tau_j=t_{i_{j-1}}=x_j\), while \(x_{j+1}\ge t_{i_{j+1}-1}\ge t_{i_j}>t_{i_{j-1}}=x_j\); Rolle on \([x_j,x_{j+1}]\) gives \(c_{i_j}\le V(f^{\prime};[x_j,x_{j+1}])\).
(c) \(j=r\) with \(x_{r-1}=x_r\): use \(c_{i_r}\le B\).
Two consecutive collapses \(x_{j-1}=x_j=x_{j+1}\) are impossible, since they would force \(t_{i_{j-1}}=x_j\) and \(t_{i_j}=x_{j+1}=x_j\) against \(t_{i_{j-1}}<t_{i_j}\); so each interval \([x_{m-1},x_m]\), \(2\le m\le r\), with pairwise disjoint interiors, is charged at most twice, giving \(\sum_{j\ge2}c_{i_j}\le2W+B\), and \(c_{i_1}\le B\) yields \(\Sigma(y)\le L\).
Extend \(v\) constantly outside \([c,d]\), let \(\bar v\) be its right-continuous modification and \(\mu\) the Lebesgue–Stieltjes measure with \(\mu((\alpha,\beta])=\bar v(\beta)-\bar v(\alpha)\), so \(v(\beta)-v(\alpha)\le\mu([\alpha,\beta])\). Integrating the finite sum of indicators,
\begin{equation*} \sum_{i=1}^Nc_i\big(v(\max J_i)-v(\min J_i)\big)\le\sum_{i=1}^Nc_i\mu(J_i) =\int_{[c,d]}\Sigma(y)\,\mu(dy)\le L\,V(\Phi;[c,d]) , \end{equation*}
so by (12) every partition sum is at most \(AW+L\,V(\Phi;[c,d])\).
(iii) The function \(h:=(\Phi\circ f)f^{\prime}\) is continuous (\(f^{\prime}\) being absolutely continuous, hence continuous), of bounded variation by (ii), since \(\Phi\) and \(f^{\prime}\) are of bounded variation by Proposition 5.3.4, and it has property (N); Exercise 5.8.51 then gives absolute continuity. For property (N): \(U:=\{f^{\prime}\ne0\}\) is relatively open, so it is a countable union of components, on each of which \(f^{\prime}\) has constant sign, and each component is a countable union of closed intervals \([\alpha,\beta]\) on which \(f\) is strictly monotone and absolutely continuous. There \(\Phi\circ f\) is absolutely continuous by Exercise 5.8.59, hence so is \(h\) as a product of absolutely continuous functions (Corollary 5.3.3), and \(h\) has property (N) on \([\alpha,\beta]\) by Exercise 5.8.49. Since \(h\equiv0\) off \(U\), a null set \(E\) has
\begin{equation*} h(E)\ \subset\ \{0\}\cup\bigcup_m h\big(E\cap[\alpha_m,\beta_m]\big) , \end{equation*}
a countable union of null sets.
Exercises 5.8.65–5.8.71
(i) (Banach, Saks [58], Bary, Menchoff [67]) A continuous function \(f\) has the form \(f = \varphi\circ\psi\), where \(\varphi\) and \(\psi\) are absolutely continuous functions, precisely when \(f\) has the following property (S): for every \(\varepsilon > 0\), there exists \(\delta > 0\) such that the measure of the set \(f(E)\) does not exceed \(\varepsilon\) whenever the measure of \(E\) does not exceed \(\delta\).
(ii) (Bary, Menchoff [67]) A continuous function \(f\) is the composition of two absolutely continuous functions precisely when \(f\) takes the set of all points \(x\) where there is no finite derivative to a measure zero set.
(i) Necessity. Property (S), read with outer measures, holds for an absolutely continuous \(g\) on \([c,d]\): let \(D\) be its set of differentiability, so \(\lambda([c,d]\setminus D)=0\) and \(g^{\prime}\in L^1\) (Theorem 5.3.6), and let \(E^{\prime}\supset E\) be a measurable hull, \(\lambda(E^{\prime})=\lambda^{*}(E)\); then
\begin{equation*} \lambda^{*}\big(g(E)\big)\le\lambda\big(g(E^{\prime}\cap D)\big) +\lambda\big(g(E^{\prime}\setminus D)\big) \le\int_{E^{\prime}\cap D}|g^{\prime}(x)|\,dx , \end{equation*}
by Proposition 5.5.4 on the measurable set \(E^{\prime}\cap D\), where \(g\) is differentiable, and by property (N) (Exercise 5.8.49) on the null set \(E^{\prime}\setminus D\); absolute continuity of the integral of \(|g^{\prime}|\) supplies \(\delta\) for a given \(\varepsilon\). Property (S) survives composition – chain the two \(\delta\)’s – so \(f=\varphi\circ\psi\) has it.
Sufficiency. Property (S) gives (N) at once (let \(\varepsilon\downarrow0\) on a null set) and also \(N_f(y):=\#f^{-1}(y)<\infty\) a.e. Suppose \(\lambda^{*}(S)=c>0\) for \(S:=\{N_f=\infty\}\). Put \(h_m:=(b-a)/m\), \(I_{m,k}:=[a+(k-1)h_m,a+kh_m]\), \(B_{m,k}:=f(I_{m,k})\) (compact intervals) and \(S^n_m:=\{y:\ \#\{k\le m:\ y\in B_{m,k}\}\ge n\}\), Borel. Fix \(n\): each \(y\in S\) has \(n\) preimages at mutual distance at least some \(\eta>0\), hence in \(n\) distinct \(I_{m,k}\) once \(h_m<\eta\), so \(S\subset\bigcup_MT^n_M\) with \(T^n_M:=\bigcap_{m\ge M}S^n_m\) increasing in \(M\); continuity of \(\lambda\) from below gives \(M\) with \(\lambda(S^n_m)>c/2\) for all \(m\ge M\). Fix \(m\ge\max(M,n)\) and select intervals greedily: for measurable \(R_j\subset S^n_m\),
\begin{equation*} \sum_{k=1}^{m}\lambda(B_{m,k}\cap R_j)=\int_{R_j}\#\{k:\ y\in B_{m,k}\}\,dy \ \ge\ n\,\lambda(R_j) , \end{equation*}
so some \(k_{j+1}\) has \(\lambda(B_{m,k_{j+1}}\cap R_j)\ge(n/m)\lambda(R_j)\); with \(R_0:=S^n_m\) and \(R_{j+1}:=R_j\setminus B_{m,k_{j+1}}\) this gives \(\lambda(R_j)\le(1-n/m)^j\lambda(S^n_m)\le e^{-jn/m}\lambda(S^n_m)\). Taking \(j:=\lceil m\ln2/n\rceil\) and \(E_n:=\bigcup_{i\le j}I_{m,k_i}\), whose image contains \(S^n_m\setminus R_j\),
\begin{equation*} \lambda\big(f(E_n)\big)\ge\tfrac12\lambda(S^n_m)>\frac c4,\qquad \lambda(E_n)\le j\,h_m\le\frac{(b-a)(\ln2+1)}{n}\ \xrightarrow[n\to\infty]{}0 , \end{equation*}
contradicting (S) with \(\varepsilon:=c/4\).
The indicatrix has a Borel majorant: with \(I_{j,k}\) the \(2^j\) dyadic subintervals of \([a,b]\), \(B_{j,k}:=f(I_{j,k})\) and \(c_j:=\sum_k\mathbf{1}_{B_{j,k}}\), each \(I_{j,k}\) is the union of its halves, so \(c_{j+1}\ge c_j\) and \(\widetilde N:=\lim_jc_j\) is Borel. If \(N_f(y)\ge n\), then \(n\) preimages lie in distinct dyadic intervals for large \(j\), so \(\widetilde N\ge N_f\); conversely each \(y\in B_{j,k}\) has a preimage in \(I_{j,k}\), and a point lies in at most two closed \(I_{j,k}\), so \(c_j\le2N_f\). Thus \(N_f\le\widetilde N\le2N_f\) and \(\widetilde N<\infty\) a.e.
Assume \(f\) nonconstant, put \(J:=f([a,b])=[m_0,M_0]\) and
\begin{equation*} w:=\frac{1}{1+\widetilde N}\ \ (w:=0\ \text{where}\ \widetilde N=\infty), \qquad \eta(y):=\int_{m_0}^{y}w(s)\,ds , \end{equation*}
and take \(\psi:=\eta\circ f\), \(\varphi:=\eta^{-1}\), so \(\varphi\circ\psi=f\). Here \(0\le w\le1\) is Borel with \(w>0\) a.e., so \(\eta\) is absolutely continuous and strictly increasing, a homeomorphism of \(J\) onto \([0,\eta(M_0)]\). Both factors are absolutely continuous by Exercise 5.8.51. Indeed, for \(t_1\le\dots\le t_{p+1}\) in \([a,b]\),
\begin{equation*} \sum_{i=1}^{p}|\psi(t_{i+1})-\psi(t_i)|=\int_Jw(y)\kappa(y)\,dy,\qquad \kappa(y):=\#\{i:\ y\in\Delta_i\} , \end{equation*}
where \(\Delta_i\) is the open interval between \(f(t_i)\) and \(f(t_{i+1})\) (endpoints form a null set); the intermediate value theorem puts a preimage of each \(y\in\Delta_i\) inside \((t_i,t_{i+1})\), and these intervals are pairwise disjoint, so \(\kappa\le N_f\le\widetilde N\) and \(V_a^b(\psi)\le\int_J\widetilde N/(1+\widetilde N)\,dy\le\lambda(J)<\infty\); and \(\psi\) has property (N), since \(f\) has it and \(\eta\) has it by Exercise 5.8.49. As for \(\varphi\), it is continuous and nondecreasing, hence of bounded variation, and for Borel \(B^{\prime}\) with \(\lambda(B^{\prime})=0\) the change of variables for monotone absolutely continuous maps (Corollary 5.4.4) applied to \(\mathbf{1}_{B^{\prime}}\) gives
\begin{equation*} \int_{\eta^{-1}(B^{\prime})}w(y)\,dy=\int_J\mathbf{1}_{B^{\prime}} \big(\eta(y)\big)\eta^{\prime}(y)\,dy=\lambda(B^{\prime})=0 , \end{equation*}
which forces \(\lambda(\eta^{-1}(B^{\prime}))=0\) because \(w>0\) a.e.
The construction used only property (N) and \(N_f<\infty\) a.e.; call that implication the Lemma.
(ii) Let \(Z\) be the set of points where \(f\) has no finite derivative. If \(f=\varphi\circ\psi\), let \(A\) and \(B\) be the null sets (Theorem 5.3.6) where \(\psi\), respectively \(\varphi\), has no finite derivative; the chain rule differentiates \(f\) at every \(x\notin A\) with \(\psi(x)\notin B\), so \(Z\subset A\cup\psi^{-1}(B)\) and
\begin{equation*} f(Z)\subset\varphi\big(\psi(A)\big)\cup\varphi(B) , \end{equation*}
which is null because \(\psi\) and \(\varphi\) have property (N) (Exercise 5.8.49).
Conversely let \(\lambda(f(Z))=0\). Then \(f\) has property (N): for \(\lambda(A)=0\), Proposition 5.5.4 gives \(\lambda(f(A\setminus Z))\le\int_{A\setminus Z}|f^{\prime}|\,dx=0\), while \(f(A\cap Z)\subset f(Z)\). And \(N_f<\infty\) a.e.: if \(N_f(y)=\infty\), the infinite closed set \(f^{-1}(y)\) has an accumulation point \(p\) lying in it, so \(\{N_f=\infty\}\subset f(P)\) where \(P\) is the set of points accumulated by their own level set; each \(p\in P\) with a finite \(f^{\prime}(p)\) has \(f^{\prime}(p)=0\) (take \(x_k\to p\), \(x_k\ne p\), \(f(x_k)=f(p)\)), so \(P\subset\{f^{\prime}=0\}\cup Z\) and, by Proposition 5.5.4,
\begin{equation*} \lambda\big(f(P)\big)\le\int_{\{f^{\prime}=0\}}|f^{\prime}(x)|\,dx +\lambda\big(f(Z)\big)=0 . \end{equation*}
The Lemma applies.
(i) (Fichtenholz [293]) Show that property (S) in the previous exercise does not follow from property (N).
(ii) (Banach [51]) Prove that a continuous function \(f\) on an interval has property (S) precisely when it has property (N) and assumes almost every value only at finitely many points.
(i) Take the identity on a fat Cantor set, spiked on the deleted intervals. Let \(K\subset[0,1]\) arise by deleting at step \(k\) the concentric open middle interval of length \(\eta_k:=4^{-k}\) from each of the \(2^{k-1}\) intervals of step \(k-1\); the deleted length totals \(\sum_k2^{k-1}4^{-k}=1/2\), so \(K\) is compact, nowhere dense, \(\lambda(K)=1/2\), and the \(2^k\) surviving basic intervals of rank \(k\) have common length \(\ell_k\) with \(2^k\ell_k=\tfrac12+2^{-k-1}\), i.e. \(\ell_k\le2^{-k}\). Write \(\mathcal{I}_k\) for the rank-\(k\) deleted intervals and \(\varepsilon_k:=2^{-(k-1)}>\eta_k\). Set \(f(x):=x\) on \(K\), and on \((u,v)\in\mathcal{I}_k\) let \(f\) be piecewise linear with nodes
\begin{equation*} f(u)=u,\quad f\Big(u+\frac{v-u}{3}\Big)=u+\varepsilon_k,\quad f\Big(u+\frac{2(v-u)}{3}\Big)=u-\varepsilon_k,\quad f(v)=v , \end{equation*}
so that \(f([u,v])=[u-\varepsilon_k,u+\varepsilon_k]\), since \(u<v\le u+\eta_k<u+\varepsilon_k\).
\(f\) is continuous: off \(K\), and one-sidedly at the endpoints of deleted intervals, \(f\) is piecewise linear with nodal values matching the identity. At \(p\in K\) approached from a side on which \(p\) is not such an endpoint, fix \(k_0\): the deleted intervals of rank \(\le k_0\) are finitely many and \(p\) is interior to none, so some \(\rho<\varepsilon_{k_0+1}\) makes every \(x\) with \(|x-p|<\rho\) on that side either a point of \(K\), where \(|f(x)-f(p)|=|x-p|<\rho\), or a point of a deleted \((u,v)\) of rank \(k>k_0\), where \(|u-x|\le\eta_k\) and hence
\begin{equation*} |f(x)-f(p)|\le\varepsilon_k+|u-p|\le\varepsilon_k+\eta_k+|x-p| \le\varepsilon_{k_0+1}+\eta_{k_0+1}+\rho\ \xrightarrow[k_0\to\infty]{}\ 0 . \end{equation*}
\(f\) has property (N), since \([0,1]\) is the countable union of \(K\) and the closed deleted intervals, on each of which \(f\) is Lipschitz (the identity on \(K\); three linear pieces elsewhere), and Lipschitz maps preserve nullity. But (S) fails: for \(E_k:=\bigcup_{(u,v)\in\mathcal{I}_k}(u,v)\) we have \(\lambda(E_k)=2^{k-1}4^{-k}=\tfrac12 2^{-k}\to0\), while \(f(E_k)\supset K\). Indeed both interior nodes lie in \((u,v)\), so \(f((u,v))=[u-\varepsilon_k,u+\varepsilon_k]\) by the intermediate value theorem; and each \(y\in K\) lies in a rank-\(k-1\) basic interval \(B\) whose middle deleted interval \((u,v)\in\mathcal{I}_k\) has \(u\in B\), whence \(|y-u|\le\ell_{k-1}\le\varepsilon_k\). So \(\lambda(f(E_k))\ge1/2\) for all \(k\).
(ii) Necessity: (S) gives (N) at once (a null set has outer measure below every \(\delta\)), and \(\lambda(\{N_f=\infty\})=0\), where \(N_f(y):=\#f^{-1}(y)\), is the greedy-covering argument in the solution of Exercise 5.8.65.
Sufficiency: let \(f\) be continuous with property (N) and \(N_f<\infty\) a.e., and suppose (S) failed for some \(\varepsilon>0\). Taking \(\delta:=2^{-k-1}\) gives \(E_k\) with \(\lambda^{*}(E_k)\le2^{-k-1}\) and \(\lambda^{*}(f(E_k))>\varepsilon\); by regularity choose \(U_k\supset E_k\) open in \([a,b]\) with \(\lambda(U_k)\le2^{-k}\), so \(f(U_k)\) is an at most countable union of intervals and points, hence measurable with \(\lambda(f(U_k))>\varepsilon\). Put \(A:=\limsup_kf(U_k)\) and \(G:=\limsup_kU_k\). The sets \(\bigcup_{k\ge K}f(U_k)\) decrease, lie in the bounded set \(f([a,b])\), and have measure at least \(\varepsilon\), so \(\lambda(A)\ge\varepsilon\); and \(\lambda(\bigcup_{k\ge K}U_k)\le \sum_{k\ge K}2^{-k}\to0\) gives \(\lambda(G)=0\), hence \(\lambda(f(G))=0\) by property (N). For \(y\in A\setminus f(G)\), the set \(f^{-1}(y)\) meets infinitely many \(U_k\) while no point of it lies in \(G\), so each of its points lies in only finitely many \(U_k\) – impossible for a finite \(f^{-1}(y)\). Hence \(N_f=\infty\) on \(A\setminus f(G)\) and
\begin{equation*} \lambda\big(\{y:\ N_f(y)=\infty\}\big)\ \ge\ \lambda(A)-\lambda\big(f(G)\big)\ \ge\ \varepsilon>0 , \end{equation*}
a contradiction.
(\(\circ\)) (i) Show that if a sequence of increasing functions \(\psi_n\) on the real line converges to an increasing function \(\psi\) at all points of an everywhere dense set, then it converges to \(\psi\) at every point of continuity of \(\psi\).
(ii) Let \(\{\psi_n\}\) be a uniformly bounded sequence of increasing functions on \([a,b]\). Show that \(\{\psi_n\}\) contains a pointwise convergent subsequence.
(i) Fix \(\varepsilon>0\) and a continuity point \(\tau\) of \(\psi\) (increasing means nondecreasing). Continuity gives \(\rho>0\) with \(|\psi(t)-\psi(\tau)|<\varepsilon/2\) for \(|t-\tau|<\rho\), and density of \(D\) gives \(\alpha\in D\cap(\tau-\rho,\tau)\) and \(\beta\in D\cap(\tau,\tau+\rho)\), so \(\psi(\beta)-\psi(\alpha)<\varepsilon\). Choose \(m\) with \(|\psi(\alpha)-\psi_n(\alpha)|<\varepsilon\) and \(|\psi(\beta)-\psi_n(\beta)|<\varepsilon\) for \(n\ge m\); then monotonicity of \(\psi_n\) and of \(\psi\) gives, for such \(n\),
\begin{equation*} \psi(\tau)-2\varepsilon\le\psi(\alpha)-\varepsilon<\psi_n(\alpha) \le\psi_n(\tau)\le\psi_n(\beta)<\psi(\beta)+\varepsilon \le\psi(\tau)+2\varepsilon . \end{equation*}
As \(\varepsilon>0\) was arbitrary, \(\psi_n(\tau)\to\psi(\tau)\).
(ii) Put \(C:=\sup_n\sup_{[a,b]}|\psi_n|<\infty\) and \(D:=(\mathbb{Q}\cap[a,b])\cup\{a,b\}\), countable and dense. Enumerating \(D\) and extracting diagonally (Bolzano–Weierstrass at each point, the values being bounded by \(C\)) gives a subsequence \(\psi_{n_k}\) converging on \(D\) to a bounded nondecreasing \(\psi_0\). Set
\begin{equation*} \psi(x):=\sup\{\psi_0(t):\ t\in D,\ t\le x\},\qquad x\in[a,b] , \end{equation*}
which is nondecreasing, bounded, and equal to \(\psi_0\) on \(D\) by monotonicity of \(\psi_0\). Extend every function by its value at \(a\) on \((-\infty,a)\) and by its value at \(b\) on \((b,\infty)\): monotonicity persists and, as \(a,b\in D\), convergence holds on a dense subset of \(\mathbb{R}\), so (i) gives \(\psi_{n_k}\to\psi\) at every continuity point of \(\psi\). The discontinuity set \(S\) of the bounded monotone \(\psi\) is at most countable (Corollary 5.2.4), so one further diagonal extraction over \(S\) yields a subsequence converging at every point of \([a,b]\).
(\(\circ\)) Let \(f_1,\dots,f_n\) be functions of bounded variation on the interval \([a,b]\) such that \(\bigl(f_1(x),\dots,f_n(x)\bigr)\in U\subset\mathbb{R}^n\) for all \(x\in[a,b]\). Suppose that a function \(\varphi\colon U\to\mathbb{R}^1\) satisfies the Lipschitz condition. Show that the composition \(\varphi(f_1,\dots,f_n)\) is a function of bounded variation on \([a,b]\).
The variation is bounded by \(C\sum_jV_a^b(f_j)\), where \(C\) is a Lipschitz constant for \(\varphi\). Writing \(F:=(f_1,\dots,f_n)\) and \(g:=\varphi\circ F\), and using \(\|u-v\|\le\sum_{j=1}^n|u_j-v_j|\), any \(t_1\le t_2\le\dots\le t_{m+1}\) in \([a,b]\) give
\begin{equation*} \sum_{i=1}^{m}|g(t_{i+1})-g(t_i)|\le C\sum_{j=1}^{n}\sum_{i=1}^{m} |f_j(t_{i+1})-f_j(t_i)|\le C\sum_{j=1}^{n}V_a^b(f_j)<\infty , \end{equation*}
a bound independent of the points chosen; take the supremum over them.
(\(\circ\)) Let \(f\) and \(g\) be functions of bounded variation on \([a,b]\). Show that \(fg\) is a function of bounded variation on \([a,b]\), and if \(g\ge c>0\), then so is the function \(f/g\).
Apply Exercise 5.8.68 with \(n=2\), \((f_1,f_2)=(f,g)\) and \(\varphi(x,y)=xy\), respectively \(x/y\), on a rectangle containing the range of the pair (the printed hint’s \([a,b]\times[a,b]\) is a slip). Functions of bounded variation are bounded: \(|f|\le|f(a)|+V_a^b(f)=:M_1\) and \(|g|\le|g(a)|+V_a^b(g)=:M_2\).
Product: on \(U:=[-M_1,M_1]\times[-M_2,M_2]\),
\begin{equation*} |xy-x^{\prime}y^{\prime}|\le|x||y-y^{\prime}|+|y^{\prime}||x-x^{\prime}| \le\max(M_1,M_2)\big(|x-x^{\prime}|+|y-y^{\prime}|\big) , \end{equation*}
which is at most \(\sqrt2\max(M_1,M_2)\|(x,y)-(x^{\prime},y^{\prime})\|\), so \(fg\) is of bounded variation with \(V_a^b(fg)\le2\max(M_1,M_2) \big(V_a^b(f)+V_a^b(g)\big)\).
Quotient: for \(g\ge c>0\), on \(U:=[-M_1,M_1]\times[c,M_2]\),
\begin{equation*} \Big|\frac{x}{y}-\frac{x^{\prime}}{y^{\prime}}\Big| =\frac{|x(y^{\prime}-y)+y(x-x^{\prime})|}{yy^{\prime}} \le\frac{\max(M_1,M_2)}{c^{2}}\big(|x-x^{\prime}|+|y-y^{\prime}|\big) , \end{equation*}
so \(f/g\) is of bounded variation, with \(V_a^b(f/g)\le2\max(M_1,M_2)c^{-2} \big(V_a^b(f)+V_a^b(g)\big)\).
(\(\circ\)) Show that the space \(BV[a,b]\) of all functions of bounded variation on \([a,b]\) is a Banach space with respect to the norm \(\|f\|_{BV}=|f(a)|+V_a^b(f)\).
\(\|\cdot\|_{BV}\) is a norm and \(BV[a,b]\) is complete for it. Homogeneity and subadditivity come from Definition 5.2.1 together with \(V_a^b(\alpha f+\beta g)\le|\alpha|V_a^b(f)+|\beta|V_a^b(g)\), and \(\|f\|_{BV}=0\) forces \(f(a)=0\) and \(V_a^b(f)=0\), i.e. \(f\equiv0\). Since \(|f(x)|\le|f(a)|+|f(x)-f(a)|\le\|f\|_{BV}\),
\begin{equation*} \sup_{x\in[a,b]}|f(x)|\le\|f\|_{BV}. \tag{\(\ast\)} \end{equation*}
Let \(\{f_n\}\) be Cauchy in \(BV[a,b]\), so bounded, say \(\|f_n\|_{BV}\le C\). Applying \((\ast)\) to \(f_n-f_m\) makes \(\{f_n\}\) uniformly Cauchy, hence \(f_n\to f\) uniformly; pointwise convergence passes each finite sum to the limit, so for any \(t_1\le\dots\le t_{m+1}\) in \([a,b]\),
\begin{equation*} \sum_{i=1}^{m}|f(t_{i+1})-f(t_i)| =\lim_{n\to\infty}\sum_{i=1}^{m}|f_n(t_{i+1})-f_n(t_i)|\le C , \end{equation*}
whence \(V_a^b(f)\le C\) and \(f\in BV[a,b]\). Given \(\varepsilon>0\) choose \(n_0\) with \(\|f_n-f_k\|_{BV}\le\varepsilon\) for \(n,k\ge n_0\); the same limit passage applied to \(f_k-f_n\) gives, for \(n\ge n_0\),
\begin{equation*} \sum_{i=1}^{m}\bigl|(f-f_n)(t_{i+1})-(f-f_n)(t_i)\bigr| =\lim_{k\to\infty}\sum_{i=1}^{m}\bigl|(f_k-f_n)(t_{i+1})-(f_k-f_n)(t_i)\bigr| \le\varepsilon . \end{equation*}
Taking suprema over collections yields \(V_a^b(f-f_n)\le\varepsilon\), and \(|f(a)-f_n(a)|=\lim_k|f_k(a)-f_n(a)|\le\varepsilon\), so \(\|f-f_n\|_{BV}\le2\varepsilon\). Thus \(f_n\to f\) in \(BV[a,b]\).
(\(\circ\)) (i) Show that every bounded nondecreasing function \(f\) on a set \(T\subset\mathbb{R}^1\) is of bounded variation and \(V(f,T)\le2\sup_{t\in T}|f(t)|\).
(ii) Let \(f\) be a function of bounded variation on a set \(T\subset\mathbb{R}^1\) and let \(V(x)=V\bigl(f,(-\infty,x]\cap T\bigr)\), \(x\in T\). Show that \(V\) and \(V-f\) are nondecreasing functions on \(T\) and that the set of points of continuity of \(V\) coincides with the set of points of continuity of \(f\).
(iii) Show that if a function \(f\) is of bounded variation on a set \(T\subset\mathbb{R}^1\), then there exist two nondecreasing functions \(f_1\) and \(f_2\) on the whole real line such that \(f=f_1-f_2\) on \(T\) and \(V(f,T)=V(f_1,\mathbb{R}^1)+V(f_2,\mathbb{R}^1)\).
(i) \(V(f,T)=\sup_Tf-\inf_Tf\le2\sup_{t\in T}|f(t)|<\infty\) by Fact 1 below; continuity on \(T\) means continuity relative to \(T\) throughout.
Two facts follow from Definition 5.2.1. Fact 1: for bounded nondecreasing \(h\) on \(T\), \(V(h,T)=\sup_Th-\inf_Th\), since the sums telescope to \(h(t_{m+1})-h(t_1)\) and two points nearly attaining \(\inf_Th\), \(\sup_Th\) (in that order, by monotonicity) give the reverse bound. (Check!) Fact 2: for \(x\in T\),
\begin{equation*} V(f,T)=V\bigl(f,T\cap(-\infty,x]\bigr)+V\bigl(f,T\cap[x,+\infty)\bigr), \end{equation*}
because concatenating collections from the two halves at \(x\) gives \(\ge\), while inserting \(x\) into a collection in \(T\) does not decrease its sum (triangle inequality) and splits it at \(x\), giving \(\le\). Applying Fact 2 inside \(T\cap(-\infty,y]\) at a point \(x\le y\) of \(T\),
\begin{equation*} V(y)-V(x)=V\bigl(f,T\cap[x,y]\bigr)\ge|f(y)-f(x)|,\qquad x\le y\text{ in }T, \tag{1} \end{equation*}
the last bound coming from the two-point collection \(x\le y\).
(ii) By (1), \(V\) is nondecreasing and \(V(y)-V(x)\ge\pm(f(y)-f(x))\), so \(V-f\) (and \(V+f\)) is nondecreasing; also \(|f(y)-f(x)|\le|V(y)-V(x)|\), so \(f\) is continuous at every continuity point of \(V\). Conversely, let \(f\) be continuous at \(x_0\in T\) and set \[ \alpha:=\inf\bigl\{V(f,T\cap[x_0,y]):\ y\in T,\ y>x_0\bigr\}, \] so that by (1) right continuity of \(V\) at \(x_0\) says exactly \(\alpha=0\) (nothing to prove if \(x_0\) is not a limit point of \(T\) from the right; left continuity is symmetric). Suppose \(\alpha>0\). Choose \(\rho>0\) with \(|f(t)-f(x_0)|<\alpha/4\) for \(t\in T\), \(|t-x_0|<\rho\), pick \(y_0\in T\cap(x_0,x_0+\rho)\), put \(W:=V(f,T\cap[x_0,y_0])\in[\alpha,\infty)\), and take a collection \(t_1\le\dots\le t_{m+1}\) in \(T\cap[x_0,y_0]\) with \(\Sigma:=\sum_i|f(t_{i+1})-f(t_i)|>W-\alpha/4\). Adjoining \(x_0\) and deleting repetitions we may assume \(t_1=x_0<t_2\) (all \(t_i=x_0\) would give \(\Sigma=0\le W-\alpha/4\)); then \(|f(t_2)-f(x_0)|<\alpha/4\) and, by Fact 2 applied at \(t_2\) inside \(T\cap[x_0,y_0]\),
\begin{equation*} W-\frac{\alpha}{4}<\Sigma<\frac{\alpha}{4}+V\bigl(f,T\cap[t_2,y_0]\bigr) =\frac{\alpha}{4}+W-V\bigl(f,T\cap[x_0,t_2]\bigr), \end{equation*}
so \(V(f,T\cap[x_0,t_2])<\alpha/2\), contradicting the definition of \(\alpha\) since \(t_2\in T\), \(t_2>x_0\). Hence the continuity points of \(V\) and of \(f\) are the same.
(iii) Take \(f_1:=(V+f)/2\) and \(f_2:=(V-f)/2\), bounded and nondecreasing on \(T\) by (ii), with \(f_1-f_2=f\). By (1), \(V(y)-V(x)\le V(f,T)\) for \(x\le y\), while any collection \(t_1\le\dots\le t_{m+1}\) satisfies \(V(t_{m+1})-V(t_1)\ge\sum_i|f(t_{i+1})-f(t_i)|\); hence \(\sup_TV-\inf_TV=V(f,T)\), i.e. \(V(V,T)=V(f,T)\) by Fact 1. Computing the suprema (resp. infima) of the nondecreasing functions \(f_1\), \(f_2\), \(f_1+f_2=V\) along one common sequence in \(T\) tending to \(\sup T\) (resp. \(\inf T\)) gives \(\sup_TV=\sup_Tf_1+\sup_Tf_2\) and likewise for infima, so Fact 1 yields
\begin{equation*} V(f_1,T)+V(f_2,T)=\sup_TV-\inf_TV=V(f,T). \tag{2} \end{equation*}
Extend by \(\widetilde f_i(x):=\sup\{f_i(t):t\in T,\ t\le x\}\), with \(\widetilde f_i(x):=\inf_Tf_i\) when \(T\cap(-\infty,x]=\emptyset\): each \(\widetilde f_i\) is nondecreasing on \(\mathbb{R}^1\), agrees with \(f_i\) on \(T\) (monotonicity), and has the same supremum and infimum as \(f_i\), so \(V(\widetilde f_i,\mathbb{R}^1)=V(f_i,T)\) by Fact 1. Thus \(\widetilde f_1-\widetilde f_2=f\) on \(T\) and, by (2), \(V(\widetilde f_1,\mathbb{R}^1)+V(\widetilde f_2,\mathbb{R}^1)=V(f,T)\).
Exercises 5.8.72–5.8.78
(\(\circ\)) Suppose that a function \(f\) on a set \(T \subset \mathbb{R}^1\) is of bounded variation. Show that \(f\) can be extended to \(\mathbb{R}^1\) in such a way that \(V(f, \mathbb{R}^1) = V(f, T)\).
Take \(F:=f_1-f_2\), where \(f_1,f_2\) are the nondecreasing functions on \(\mathbb{R}^1\) furnished by Exercise 5.8.71(iii), so that \(F=f\) on \(T\) and \(V(f_1,\mathbb{R}^1)+V(f_2,\mathbb{R}^1)=V(f,T)\). Subadditivity of the variation (inequality (5.2.1), whose proof applies verbatim over an arbitrary set) gives \[ V(F,\mathbb{R}^1)\le V(f_1,\mathbb{R}^1)+V(f_2,\mathbb{R}^1)=V(f,T), \] while every admissible collection of points of \(T\) is admissible for \(\mathbb{R}^1\) and \(F=f\) on \(T\), so \(V(F,\mathbb{R}^1)\ge V(F,T)=V(f,T)\). Hence \(V(F,\mathbb{R}^1)=V(f,T)\).
(\(\circ\)) Let \(f\) be a function of bounded variation on \([a,b]\). We redefine \(f\) at all discontinuity points, making it left continuous (the discontinuities of \(f\) are jumps). Show that the obtained function \(f_0\) is of bounded variation and the following estimate holds: \(V(f_0,[a,b]) \le V(f,[a,b])\).
The modified function is \(f_0(x)=f(x-)\) for \(x\in(a,b]\) and \(f_0(a)=f(a)\), all one-sided limits existing and being finite because \(f\) is a difference of two nondecreasing functions (Proposition 5.2.2). Fix \(a\le t_0<t_1<\dots<t_m\le b\) and \(\varepsilon>0\), and choose \(s_0<s_1<\dots<s_m\) in \([a,b]\) with \(s_0:=a\) if \(t_0=a\) and otherwise \(s_i\in(\max\{s_{i-1},t_{i-1}\},t_i)\) so close to \(t_i\) that \[ |f(s_i)-f_0(t_i)|<\frac{\varepsilon}{2m+2}, \] which is possible since \(f(s)\to f(t_i-)\) as \(s\uparrow t_i\) and \(\max\{s_{i-1},t_{i-1}\}<t_i\). Then \(|f_0(t_i)-f_0(t_{i-1})|\le|f(s_i)-f(s_{i-1})|+\varepsilon/(m+1)\) for each \(i\), and since \(s_0<\dots<s_m\) is an admissible collection for \(f\),
\begin{equation*} \sum_{i=1}^{m}|f_0(t_i)-f_0(t_{i-1})| \le\sum_{i=1}^{m}|f(s_i)-f(s_{i-1})|+\varepsilon\le V(f,[a,b])+\varepsilon . \end{equation*}
Letting \(\varepsilon\to0\) and taking the supremum over collections (repeated points add nothing) gives \(V(f_0,[a,b])\le V(f,[a,b])<\infty\).
Let \(f_n\) be functions on \([a,b]\) such that \(\sup_n V_a^b(f_n) \le C < \infty\) and \(f_n \to f\) in \(L^1[a,b]\). Show that \(f\) coincides almost everywhere on \([a,b]\) with a function of bounded variation. In this case, we shall say that \(f\) is of essentially bounded variation defined by the formula \(\|f\|_{BV} := \inf V_a^b(g)\), where \(\inf\) is taken over all functions \(g\) of bounded variation that are equal almost everywhere to \(f\).
\(f\) agrees a.e. with the function \(g\) of bounded variation obtained below, and \(\|f\|_{BV}\le C\). Convergence in \(L^1[a,b]\) implies convergence in measure, so the Riesz theorem (Theorem 2.2.5(i)) gives a subsequence \(f_{n_k}\to f\) a.e.; let
\begin{equation*} T := \{ x \in [a,b] : f(x) \text{ is finite and } f_{n_k}(x) \to f(x) \}, \end{equation*}
a set of full Lebesgue measure in \([a,b]\). Any \(t_0\le\dots\le t_m\) in \(T\) is an admissible collection in \([a,b]\), so
\begin{equation*} \sum_{i=1}^{m}|f(t_i)-f(t_{i-1})| =\lim_{k\to\infty}\sum_{i=1}^{m}|f_{n_k}(t_i)-f_{n_k}(t_{i-1})|\le C , \end{equation*}
whence \(V(f,T)\le C\). By Exercise 5.8.72 applied to \(f|_T\) there is \(g\) on \(\mathbb{R}^1\) with \(g=f\) on \(T\) and \(V(g,\mathbb{R}^1)=V(f,T)\le C\); in particular \(V_a^b(g)\le C\) and \(g=f\) off the null set \([a,b]\setminus T\). Hence \(f\) coincides a.e. with a function of bounded variation and \(\|f\|_{BV}\le V_a^b(g)\le\sup_nV_a^b(f_n)\).
Show that a measurable function \(f\) on \([a,b]\) is of essentially bounded variation if the following quantity is finite:
\begin{equation*} \mathrm{ess}\,V_a^b(f) := \sup \Bigl\{ \sum_{i=1}^{m} |f(t_i) - f(t_{i-1})| \Bigr\}, \end{equation*}
where \(\sup\) is taken over all \(m \in \mathbb{N}\) and all points \(a < t_0 < t_1 < \dots < t_m < b\) that are points of the approximate continuity of \(f\).
If \(A:=\mathrm{ess}\,V_a^b(f)<\infty\), then \(f\) coincides a.e. with a function of bounded variation and \(\|f\|_{BV}\le A\). Let
\begin{equation*} T := \{ x \in (a,b) : x \text{ is a point of approximate continuity of } f \}, \end{equation*}
a set of full measure in \([a,b]\), since the finite measurable \(f\) is approximately continuous a.e. (Theorem 5.8.9). Discarding repetitions, every collection \(t_0\le\dots\le t_m\) in \(T\) becomes an admissible \(a<t_0<\dots<t_m<b\) of approximate continuity points, so \(\sum_{i=1}^{m}|f(t_i)-f(t_{i-1})|\le A\) by the definition of \(A\), and therefore \(V(f,T)\le A\). By Exercise 5.8.72 there is \(g\) on \(\mathbb{R}^1\) with \(g=f\) on \(T\) and \(V(g,\mathbb{R}^1)=V(f,T)\le A\); then \(V_a^b(g)\le A\) and \(f=g\) a.e. on \([a,b]\), so \(f\) is of essentially bounded variation in the sense of Exercise 5.8.74 with \(\|f\|_{BV}\le V_a^b(g)\le A\).
(i) Show that an integrable function \(f\) coincides almost everywhere on \([a,b]\) with some function of bounded variation precisely when
\begin{equation*} \int_a^b |f(x+h) - f(x)| \, dx = O(h) \quad \text{as } h \to 0, \end{equation*}
where we set \(f = 0\) outside \([a,b]\).
(ii) Show that if
\begin{equation*} \int_a^b |f(x+h) - f(x)| \, dx = o(h) \quad \text{as } h \to 0, \end{equation*}
then \(f\) almost everywhere on \([a,b]\) coincides with some constant.
(i) The condition is necessary and sufficient, and in (ii) the constant is \(0\). Extend \(f\) by \(0\) outside \([a,b]\) and write
\begin{equation*} \omega(h) := \int_a^b |f(x+h)-f(x)| \, dx, \qquad \Omega(h) := \int_{\mathbb{R}} |f(x+h)-f(x)| \, dx . \end{equation*}
Clearly \(\omega\le\Omega\), while \(\Omega(-h)=\Omega(h)\) by the substitution \(y=x+h\). For \(0<h<b-a\), vanishing of \(f\) off \([a,b]\) gives \(\Omega(h)=\int_a^{a+h}|f|+\omega(h)\); with \(k:=\lfloor(b-a)/h\rfloor\ge1\), telescoping \(|f(x)|\le\sum_{j=0}^{k-1}|f(x+jh)-f(x+(j+1)h)|+|f(x+kh)|\) and integrating over \(x\in[a,a+h]\) (substituting \(u=x+jh\)),
\begin{equation*} \int_a^{a+h}|f|\le\int_a^{a+kh}|f(u)-f(u+h)|\,du+\int_{a+kh}^{a+(k+1)h}|f| \le2\,\omega(h), \end{equation*}
since \(a+kh\le b\) for the first term and, for the second, \(a+kh>b-h\) with \(f=0\) to the right of \(b\), so it is at most \(\int_{b-h}^{b}|f|=\int_{b-h}^{b}|f(x)-f(x+h)|\,dx\le\omega(h)\). Hence \(\omega\le\Omega\le3\omega\) on \((0,b-a)\), and \(\omega(h)=O(h)\) iff \(\Omega(\delta)=O(|\delta|)\), \(\omega(h)=o(h)\) iff \(\Omega(\delta)=o(|\delta|)\).
(i) Necessity: let \(f=g\) a.e. with \(V_a^b(g)<\infty\) and put \(G:=g\mathbf{1}_{[a,b]}\), so \(f=G\) a.e. on \(\mathbb{R}^1\) and \(W:=V(G,\mathbb{R}^1)\le V_a^b(g)+|g(a)|+|g(b)|<\infty\). Let \(P,N\) be the positive and negative variations of \(G\) on \((-\infty,x]\): they are bounded nondecreasing, \(G=P-N\) (as \(G\) vanishes near \(-\infty\), so \(P(-\infty)=N(-\infty)=0\)), and their total increments add up to \(W\). For bounded nondecreasing \(M\) with finite limits at \(\pm\infty\) and \(h>0\),
\begin{equation*} \int_{\mathbb{R}}\bigl(M(x+h)-M(x)\bigr)dx =\lim_{R\to\infty}\Bigl(\int_R^{R+h}M-\int_{-R}^{-R+h}M\Bigr) =h\bigl(M(+\infty)-M(-\infty)\bigr), \end{equation*}
so applying this to \(P\) and \(N\) and using \(|G(x+h)-G(x)|\le(P(x+h)-P(x))+(N(x+h)-N(x))\) yields \(\omega(h)\le\Omega(h)\le hW\) for \(h>0\).
Sufficiency: if \(\omega(h)\le Mh\) for small \(h>0\), the comparison gives \(K<\infty\) with \(\Omega(\delta)\le K|\delta|\) for all small \(\delta\ne0\). Let \(g\in C^\infty(\mathbb{R})\) be a probability density supported in \([0,1]\), \(g_h(x):=h^{-1}g(x/h)\) and \(f_h:=f*g_h\), which is smooth and integrable with \(f_h^{\prime}=f*g_h^{\prime}\). Fubini’s theorem and translation invariance give \[ \int_{\mathbb{R}}|f_h(x+\delta)-f_h(x)|\,dx\le\Omega(\delta)\le K|\delta| , \] so Fatou’s lemma applied to \(\delta^{-1}(f_h(x+\delta)-f_h(x))\to f_h^{\prime}(x)\) yields \(\int_{\mathbb{R}}|f_h^{\prime}|\le K\) and hence \(V_a^b(f_h)=\int_a^b|f_h^{\prime}|\le K\). Since \(\|f_h-f\|_{L^1(\mathbb{R})}\to0\) (Theorem 4.2.4), the functions \(f_{1/n}\) have variations bounded by \(K\) and converge to \(f\) in \(L^1[a,b]\), so \(f\) coincides a.e. with a function of bounded variation by Exercise 5.8.74.
(ii) Here \(\Omega(\delta)=o(|\delta|)\), so the same Fatou bound gives, for each fixed \(h>0\),
\begin{equation*} \int_{\mathbb{R}}|f_h^{\prime}| \le\liminf_{\delta\to0^{+}}\frac{\Omega(\delta)}{\delta}=0 . \end{equation*}
Thus \(f_h^{\prime}\equiv0\) by continuity, \(f_h\) is constant, and the constant is \(0\) because \(f_h\in L^1(\mathbb{R})\); letting \(h\to0\) in \(\|f_h-f\|_{L^1}\to0\) gives \(f=0\) a.e. on \([a,b]\).
Let \(f\) be a function of bounded variation on \([a,b]\) and let \(g\) be a nonnegative measurable function on the real line with unit integral. Show that the function
\begin{equation*} f * g(x) = \int_{-\infty}^{+\infty} f(x-y) g(y) \, dy, \end{equation*}
where \(f(x) = f(a)\) if \(x \le a\) and \(f(x) = f(b)\) if \(x \ge b\), is of bounded variation and \(V(f*g, \mathbb{R}^1) \le V(f,[a,b])\).
\(V(f*g,\mathbb{R}^1)\le W:=V(f,[a,b])\), because the extension already satisfies \(V(f,\mathbb{R}^1)=W\): projecting a collection \(t_0\le\dots\le t_m\) onto \([a,b]\) by \(\tau_i:=\min\{\max\{t_i,a\},b\}\) leaves all values unchanged, \(f(t_i)=f(\tau_i)\), since \(f\) is constant on \((-\infty,a]\) and on \([b,+\infty)\), so \(\sum_i|f(t_i)-f(t_{i-1})|=\sum_i|f(\tau_i)-f(\tau_{i-1})|\le W\). In particular \(\sup_{\mathbb{R}}|f|\le|f(a)|+W=:M\), so \(y\mapsto f(x-y)g(y)\) is dominated by \(Mg\in L^1(\mathbb{R}^1)\) and \(f*g\) is defined at every \(x\). Given \(x_0\le\dots\le x_m\), the shifted collection \(x_0-y\le\dots\le x_m-y\) is admissible for each \(y\), whence
\begin{equation*} \begin{aligned} \sum_{i=1}^{m}|f*g(x_i)-f*g(x_{i-1})| &\le\int_{-\infty}^{+\infty} g(y)\sum_{i=1}^{m}|f(x_i-y)-f(x_{i-1}-y)|\,dy\\ &\le W\int_{-\infty}^{+\infty} g(y)\,dy=W, \end{aligned} \end{equation*}
the interchange of the finite sum with the integral being legitimate because each integrand is dominated by \(2Mg\). Taking the supremum over collections gives \(V(f*g,\mathbb{R}^1)\le W\).
(i) Prove that a Borel measure \(\mu\) on \(\mathbb{R}^n\) is absolutely continuous with respect to Lebesgue measure if and only if \(\lim_{t \to 0} \|\mu_{th} - \mu\| = 0\) for every \(h \in \mathbb{R}^n\), where \(\mu_h(B) := \mu(B-h)\).
(ii) Prove that if a Borel measure \(\mu\) on \(\mathbb{R}^n\) is differentiable along \(n\) linearly independent vectors, then it is absolutely continuous with respect to Lebesgue measure.
(i) Necessity: \(\mu=p\cdot\lambda_n\) with \(p\in L^1\) by the Radon–Nikodym theorem, so \(\mu_{th}-\mu\) has density \(p(\cdot-th)-p\) and \[ \|\mu_{th}-\mu\|=\int_{\mathbb{R}^n}|p(x-th)-p(x)|\,dx\xrightarrow[t\to0]{}0 \] by Lemma 4.2.3. Below we use \(\|\eta_h\|=\|\eta\|\) and \((\eta_u)_v=\eta_{u+v}\), both immediate from \(\eta_h(B)=\eta(B-h)\).
Sufficiency: with \(w_0:=0\), \(w_i:=v_1e_1+\dots+v_ie_i\) in the standard basis and \(\mu_{w_i}-\mu_{w_{i-1}}=(\mu_{v_ie_i}-\mu)_{w_{i-1}}\), \[ \|\mu_v-\mu\|\le\sum_{i=1}^{n}\|\mu_{v_ie_i}-\mu\|\xrightarrow[v\to0]{}0 , \] so \(h\mapsto\mu(B-h)\) is continuous for every Borel \(B\), since \(|\mu_h(B)-\mu_{h^{\prime}}(B)|\le\|\mu_{h-h^{\prime}}-\mu\|\). Let \(p\) be a smooth probability density with \(\operatorname{supp}p\subset\{|x|\le R\}\), \(p_j(x):=j^np(jx)\), \(\sigma_j:=p_j\cdot\lambda_n\), and set \(\mu*\sigma_j(B):=\int\int\mathbf{1}_B(x+y)p_j(y)\,dy\,\mu(dx)\). The double integral converges absolutely (inner integral at most \(1\), \(|\mu|\) finite), so Fubini’s theorem gives both
\begin{equation*} \mu*\sigma_j(B)=\int_{\mathbb{R}^n} p_j(y)\,\mu(B-y)\,dy =\int_B\Bigl(\int_{\mathbb{R}^n} p_j(z-x)\,\mu(dx)\Bigr)dz , \end{equation*}
the second expression exhibiting \(\mu*\sigma_j=u_j\cdot\lambda_n\ll\lambda_n\) with \(u_j\in L^1\). From the first, for any \(\delta>0\),
\begin{equation*} |\mu*\sigma_j(B)-\mu(B)|\le\sup_{|y|\le\delta}|\mu(B-y)-\mu(B)| +2\|\mu\|\int_{|y|>\delta}p_j(y)\,dy , \end{equation*}
whose last integral vanishes for \(j>R/\delta\) while the supremum tends to \(0\) as \(\delta\to0\) by the continuity just proved; hence \(\mu*\sigma_j(B)\to\mu(B)\) for every \(B\). If \(\lambda_n(B)=0\) then \(\mu*\sigma_j(B)=\int_Bu_j\,d\lambda_n=0\) for all \(j\), so \(\mu(B)=0\) and \(\mu\ll\lambda_n\).
(ii) It suffices to show \(\|\mu_{th}-\mu\|\le|t|\,\|d_h\mu\|\) whenever the generalized derivative \(\nu:=d_h\mu\) of Definition 5.8.24 exists: writing an arbitrary \(h=\sum_ic_ih_i\) in the basis \(h_1,\dots,h_n\) and telescoping as in (i),
\begin{equation*} \|\mu_{th}-\mu\|\le\sum_{i=1}^{n}\|\mu_{tc_ih_i}-\mu\| \le|t|\sum_{i=1}^{n}|c_i|\,\|d_{h_i}\mu\|\xrightarrow[t\to0]{}0 , \end{equation*}
so \(\mu\ll\lambda_n\) by (i). Keeping \(p_j,\sigma_j\), put \(\mu^j:=\mu*\sigma_j=u_j\cdot\lambda_n\) with \(u_j(x)=\int p_j(x-y)\mu(dy)\), smooth by differentiation under the integral sign (\(p_j\in C_0^\infty\), \(|\mu|\) finite). Applying the definition of \(\nu\) to \(\psi(y):=p_j(x-y)\), for which \(\partial_h\psi(y)=-(\partial_hp_j)(x-y)\), yields \(\partial_hu_j=w_j\), where \(w_j\) is the density of \(\nu*\sigma_j\), so \(\int|w_j|\,d\lambda_n=\|\nu*\sigma_j\|\le\|\nu\|\) and
\begin{equation*} \int_{\mathbb{R}^n}|u_j(x-th)-u_j(x)|\,dx =\int_{\mathbb{R}^n}\Bigl|t\int_0^1(\partial_hu_j)(x-sth)\,ds\Bigr|dx \le|t|\,\|\nu\| , \end{equation*}
that is, \(\|(\mu^j)_{th}-\mu^j\|\le|t|\,\|\nu\|\) for every \(j\). For \(\varphi\in C_b(\mathbb{R}^n)\) the dominated convergence theorem gives \(\int\varphi\,d\mu^j\to\int\varphi\,d\mu\), and likewise \(\int\varphi\,d(\mu^j)_{th}\to\int\varphi\,d\mu_{th}\) because \(\int\varphi\,d\eta_c=\int\varphi(x+c)\,\eta(dx)\); hence \(\int\varphi\,d(\mu_{th}-\mu)\le|t|\,\|\nu\|\) for all \(|\varphi|\le1\). Since \(\|\eta\|=\sup\{\int\varphi\,d\eta:\varphi\in C_b(\mathbb{R}^n),\ |\varphi|\le1\}\) for bounded Borel \(\eta\) (Hahn decomposition, inner regularity by Theorem 1.4.8, and Urysohn’s lemma), this gives \(\|\mu_{th}-\mu\|\le|t|\,\|d_h\mu\|\).
Exercises 5.8.79–5.8.85
Prove that a Borel measure \(\mu\) on \((a,b)\) has a bounded measure as the generalized derivative along \(1\) precisely when \(\mu\) has a density \(\varrho\) with respect to Lebesgue measure on \((a,b)\) such that \(\varrho\) coincides a.e. with a function of bounded variation. In addition, in this case \(\mu\) is given by the density \(\nu\big((a,x]\big)\), where \(\nu\) is the generalized derivative of \(\mu\).
\(\mu\) has a bounded generalized derivative \(\nu\) along \(1\) precisely when \(\mu=\varrho\,dx\) with \(\varrho\) equal a.e. to a function of bounded variation, and then \(\varrho(x)=\nu\big((a,x]\big)+c\) a.e.; by Definition 5.8.24 the derivative means \(\int\psi^{\prime}\,d\mu=-\int\psi\,d\nu\) for all \(\psi\in C_0^\infty\big((a,b)\big)\). The last assertion of the exercise must be read up to this additive constant, since \(\int\psi^{\prime}\,dx=0\) makes the derivative blind to constants (on \((0,1)\), \(\mu=\lambda\) has \(\nu=0\) while \(\nu((0,x])\equiv0\)), the constant being forced to vanish when \(a=-\infty\).
Two facts are used. (1) A bounded right continuous \(G\) of bounded variation on \(\mathbb{R}\) with \(G(-\infty)=0\) is \(G(x)=\nu\big((-\infty,x]\big)\) for a unique bounded signed Borel measure \(\nu\), with \(\|\nu\|\le3V(G,\mathbb{R})\). Indeed, for \(W(x):=V\big(G,(-\infty,x]\big)\) the functions \(G_1:=W\) and \(G_2:=W-G\) are bounded nondecreasing (Proposition 5.2.2(i)), vanish at \(-\infty\) by additivity of the variation (Proposition 5.2.2(iii)), and are right continuous (the argument of Exercise 5.8.71(ii) applies verbatim to one-sided continuity); Theorem 1.8.1 applied to the left continuous versions \(t\mapsto G_i(t-)\) gives nonnegative \(\nu_i\) with \(\nu_i\big((-\infty,t)\big)=G_i(t-)\), hence \(\nu_i\big((-\infty,t]\big)=G_i(t)\) by right continuity, and \(\nu:=\nu_1-\nu_2\) works; uniqueness holds because the rays form a \(\pi\)-system generating \(\mathcal{B}(\mathbb{R})\), and \(\|\nu\|\le\nu_1(\mathbb{R})+\nu_2(\mathbb{R})\le3V(G,\mathbb{R})\) since \(|G|\le V(G,\mathbb{R})\). (2) (du Bois-Reymond) A signed Borel measure \(\sigma\) on \((a,b)\), finite on compacta, with \(\int\psi^{\prime}\,d\sigma=0\) for all \(\psi\in C_0^\infty\big((a,b)\big)\) equals \(c\,\lambda\): fix \(\eta\in C_0^\infty\) with \(\int\eta\,dx=1\), put \(c:=\int\eta\,d\sigma\), and note that for \(\varphi\in C_0^\infty\) the function \(\psi(x):=\int_a^x[\varphi(t)-(\int\varphi\,ds)\eta(t)]\,dt\) lies in \(C_0^\infty\big((a,b)\big)\) with \(\psi^{\prime}=\varphi-(\int\varphi\,dx)\eta\), so \(\int\varphi\,d\tau=0\) for \(\tau:=\sigma-c\lambda\); approximating \(\mathbf{1}_K\) boundedly by smooth functions supported in a compact neighbourhood of \(K\) gives \(\tau(K)=0\) on compacta, whence \(\tau=0\) by inner regularity (Theorem 1.4.8).
Sufficiency. Let \(\mu=\varrho\,dx\) with \(\varrho=g\) a.e., \(g\) of bounded variation on \((a,b)\). The right limits \(\hat g(x):=g(x+)\) exist (Proposition 5.2.2 makes \(g\) a difference of bounded monotone functions), \(\hat g=g\) off the at most countable discontinuity set, so \(\hat g=\varrho\) a.e., and \(\hat g\) is right continuous with
\begin{equation*} \sum_i|\hat g(t_{i+1})-\hat g(t_i)| =\lim_{\epsilon\downarrow0}\sum_i|g(t_{i+1}+\epsilon)-g(t_i+\epsilon)| \le V\big(g,(a,b)\big). \end{equation*}
Extend \(\hat g\) by its finite limits \(\hat g(a+)\), \(\hat g(b-)\) to the two rays; then \(G:=\hat g-\hat g(a+)\) satisfies (1) and the resulting \(\nu\) is concentrated on \((a,b)\) (\(G\) is constant on the rays and continuous at \(a\) and \(b\)) with \(\|\nu\|\le3V\big(g,(a,b)\big)<\infty\). For \(\psi\in C_0^\infty\big((a,b)\big)\), using \(\int\psi^{\prime}\,dx=0\) and Fubini’s theorem (legitimate as \(\psi^{\prime}\) is bounded with compact support and \(|\nu|\) is finite),
\begin{equation*} \begin{aligned} \int_{(a,b)}\psi^{\prime}\,d\mu&=\int_a^b\psi^{\prime}(x)\hat g(x)\,dx =\int_a^b\psi^{\prime}(x)G(x)\,dx\\ &=\int_{(a,b)}\Big(\int_t^b\psi^{\prime}(x)\,dx\Big)\nu(dt) =-\int_{(a,b)}\psi(t)\,\nu(dt), \end{aligned} \end{equation*}
so \(\nu=d_1\mu\) is bounded and \(\hat g(x)=\hat g(a+)+\nu\big((a,x]\big)\).
Necessity. Given such a \(\nu\), put \(G(x):=\nu\big((a,x]\big)\), which is bounded, right continuous and of bounded variation because \(\sum_i|G(t_{i+1})-G(t_i)|=\sum_i|\nu((t_i,t_{i+1}])|\le\|\nu\|\). The computation above gives \(\int_a^b\psi^{\prime}G\,dx=-\int\psi\,d\nu\), so \(\sigma:=\mu-G\,dx\), finite on compacta, satisfies \(\int\psi^{\prime}\,d\sigma=0\); by (2), \(\sigma=c\,\lambda\) and \(\mu=(G+c)\,dx\) with \(G+c\) of bounded variation.
(\(\circ\)) Let \(\mu\) be a Borel measure on the real line (possibly signed) and let \(F_\mu(t)=\mu\big((-\infty,t)\big)\). Show that the measure \(\mu\) is mutually singular with Lebesgue measure if and only if \(F_\mu^{\prime}(t)=0\) a.e.
Both directions follow from the identity \(F_\mu^{\prime}=\varrho\) a.e., where \(\mu=\varrho\cdot\lambda+\nu\), \(\nu\perp\lambda\), is the Lebesgue decomposition of \(\mu\) with respect to Lebesgue measure (Theorem 3.2.3, applied to \(\mu^{\pm}\) and the equivalent finite measure \(e^{-x^2}\lambda\), which has the same null and singular sets as \(\lambda\)). Indeed \(F_{\varrho\cdot\lambda}(t)=F_{\varrho\cdot\lambda}(-m)+\int_{-m}^{t}\varrho(x)\,dx\) on each \([-m,m]\), so \(F_{\varrho\cdot\lambda}^{\prime}=\varrho\) a.e. by Theorem 5.4.2. For the singular part, \(|\nu|\perp\lambda\) has zero absolutely continuous component, so Theorem 5.8.8 gives
\begin{equation*} \lim_{r\to0}\frac{|\nu|\big(B(t,r)\big)}{2r}=D_\lambda|\nu|(t)=0 \qquad\text{for a.e. }t, \end{equation*}
while \(|F_\nu(t+h)-F_\nu(t)|\le|\nu|\big(B(t,|h|)\big)\) for both signs of \(h\), since \([t,t+h)\) and \([t+h,t)\) lie in \(B(t,|h|)\); hence \(F_\nu^{\prime}=0\) a.e. and \(F_\mu^{\prime}=\varrho\) a.e. Thus \(\mu\perp\lambda\) forces \(\varrho=0\) a.e. and \(F_\mu^{\prime}=0\) a.e., while \(F_\mu^{\prime}=0\) a.e. forces \(\varrho=0\) a.e., i.e. \(\mu=\nu\perp\lambda\).
(\(\circ\)) Let \(\mu\) be a signed Borel measure on the real line. Show that for all \(x\) one has \(V\big(F_\mu,(-\infty,x]\big)=V\big(F_\mu,(-\infty,x)\big)=|\mu|\big((-\infty,x)\big)\) and \(\|\mu\|=V(F_\mu,\mathbb{R}^1)\), where \(F_\mu(t)=\mu\big((-\infty,t)\big)\).
All three quantities equal \(|\mu|\big((-\infty,x)\big)\), the computations resting on \(F_\mu(s)-F_\mu(t)=\mu\big([t,s)\big)\) for \(t\le s\).
Upper bound: for \(t_1\le\dots\le t_{m+1}\) in \((-\infty,x]\) the intervals \([t_i,t_{i+1})\) are pairwise disjoint subsets of \((-\infty,x)\), so
\begin{equation*} \sum_{i=1}^{m}|F_\mu(t_{i+1})-F_\mu(t_i)| \le\sum_{i=1}^{m}|\mu|\big([t_i,t_{i+1})\big) \le|\mu|\big((-\infty,x)\big), \end{equation*}
whence \(V\big(F_\mu,(-\infty,x]\big)\le|\mu|\big((-\infty,x)\big)\), while \(V\big(F_\mu,(-\infty,x)\big)\le V\big(F_\mu,(-\infty,x]\big)\) since \((-\infty,x)\subset(-\infty,x]\).
Lower bound: fix \(\varepsilon>0\), take a Hahn decomposition \(\mathbb{R}=X^+\cup X^-\) for \(\mu\) (Theorem 3.1.1), put \(A^{\pm}:=X^{\pm}\cap(-\infty,x)\), so that \(|\mu|\big((-\infty,x)\big)=\mu(A^+)-\mu(A^-)\), and use inner regularity of the finite measure \(|\mu|\) (Theorem 1.4.8) to pick compacta \(K_1\subset A^+\), \(K_2\subset A^-\) with \(|\mu|(A^{\pm}\setminus K_i)<\varepsilon/2\), so that \(|\mu|\big((-\infty,x)\big)\le\mu(K_1)-\mu(K_2)+\varepsilon\). Separate them by disjoint open sets \(U_1\supset K_1\), \(U_2\supset K_2\) with \(U_1\cup U_2\subset(-\infty,x)\) and, shrinking by Theorem 1.4.8 again, \(|\mu|(U_i\setminus K_i)<\varepsilon\). Covering each \(K_i\) by finitely many of the open intervals composing \(U_i\) and replacing the part of \(K_i\) in such an interval \((\alpha,\beta)\) by a half-open \([p,q)\) with \(\alpha<p<q<\beta\le x\), we obtain pairwise disjoint intervals \([a_j,b_j)\subset U_1\) covering \(K_1\) and \([c_i,d_i)\subset U_2\) covering \(K_2\), with all endpoints in \((-\infty,x)\); since the symmetric differences lie in \(U_i\setminus K_i\),
\begin{equation*} \mu(K_1)\le\sum_{j}\big|\mu\big([a_j,b_j)\big)\big|+\varepsilon,\qquad -\mu(K_2)\le\sum_{i}\big|\mu\big([c_i,d_i)\big)\big|+\varepsilon . \end{equation*}
Relabelling all these (pairwise disjoint, as \(U_1\cap U_2=\emptyset\)) intervals as \([p_1,q_1),\dots,[p_N,q_N)\) with \(q_l\le p_{l+1}\), the points \(p_1\le q_1\le\dots\le p_N\le q_N\) form an admissible collection in \((-\infty,x)\) whose variation sum contains the terms \(|F_\mu(q_l)-F_\mu(p_l)|=|\mu([p_l,q_l))|\) and further nonnegative ones, so
\begin{equation*} |\mu|\big((-\infty,x)\big)\le\mu(K_1)-\mu(K_2)+\varepsilon \le V\big(F_\mu,(-\infty,x)\big)+3\varepsilon ; \end{equation*}
let \(\varepsilon\to0\). Finally every finite collection of points of \(\mathbb{R}\) lies in some \((-\infty,x)\), so by continuity of \(|\mu|\) along \((-\infty,m)\uparrow\mathbb{R}\),
\begin{equation*} V(F_\mu,\mathbb{R}^1)=\sup_{x}|\mu|\big((-\infty,x)\big) =|\mu|(\mathbb{R}^1)=\|\mu\| . \end{equation*}
(\(\circ\)) (i) Let \(f\) be a function of bounded variation on \([a,b]\) that vanishes outside of an at most countable set. Show that \(V(f,[a,x])^{\prime}=0\) a.e.
(ii) Let \(f\) be a function of bounded variation on \([a,b]\). Show that
\begin{equation*} V(f,[a,x])^{\prime}=|f^{\prime}(x)|\quad\text{a.e.} \end{equation*}
(i) \(u(x):=V(f,[a,x])\) is an increasing pointwise limit of step functions, so \(u^{\prime}=0\) a.e. For any \(g\) vanishing outside an at most countable set and \(x\in(a,b]\),
\begin{equation*} V\big(g,[a,x]\big)=2\sum_{y\in(a,x)}|g(y)|+|g(a)|+|g(x)| , \end{equation*}
the sum running over the \(y\) with \(g(y)\ne0\): in a strictly increasing collection (repetitions contribute nothing) every interior point occurs in two terms and each endpoint in one, which gives \(\le\); conversely, given a finite \(\Phi\subset\{g\ne0\}\cap[a,x]\), interleaving points at which \(g\) vanishes (each gap is uncountable, \(\{g\ne0\}\) is not) produces a collection whose sum is exactly the right-hand side with the sums restricted to \(\Phi\). Apply this to \(f_n:=f\cdot\mathbf{1}_{\{x_1,\dots,x_n\}}\), where \(\{x_i\}\supset\{f\ne0\}\): the functions \(u_n(x):=V(f_n,[a,x])\) are nondecreasing in \(x\), increase to \(u(x)\) pointwise (all terms are nonnegative, and the total sum is finite as \(f\) has bounded variation), and each is constant off the finite set \(\{x_1,\dots,x_n\}\), so \(u_n^{\prime}=0\) there. Writing \(u=u_1+\sum_{n\ge1}(u_{n+1}-u_n)\) as a pointwise convergent series of nondecreasing functions, Proposition 5.2.8 gives \(u^{\prime}=\lim_nu_n^{\prime}=0\) a.e.
(ii) Assume first that \(f\) is left continuous. Then \(V_f(x):=V(f,[a,x])\) and \(V_f-f\) are bounded nondecreasing (Proposition 5.2.2(i)) and left continuous (the argument of Exercise 5.8.71(ii) in its one-sided form), so extending each by constants off \([a,b]\) (thus bounded, nondecreasing, left continuous, vanishing at \(-\infty\)), applying Theorem 1.8.1 to get nonnegative measures with \(\mu_\varphi\big([a,x)\big)=\varphi(x)-\varphi(a)\), and subtracting, we obtain a bounded signed Borel measure \(\mu\) concentrated on \([a,b)\) with
\begin{equation*} f(x)-f(a)=\mu\big([a,x)\big)=F_\mu(x),\qquad x\in[a,b]. \end{equation*}
Since \(F_\mu\) is constant on \((-\infty,a]\) and equals \(f-f(a)\) on \([a,x]\), Exercise 5.8.81 gives
\begin{equation*} V\big(f,[a,x]\big)=V\big(F_\mu,(-\infty,x]\big) =|\mu|\big([a,x)\big)=F_{|\mu|}(x),\qquad x\in[a,b). \end{equation*}
Let \(\mu=g\cdot\lambda+\nu\), \(\nu\perp\lambda\), be the Lebesgue decomposition and let \(N\) be a \(\lambda\)-null Borel set carrying \(\nu\); splitting any Borel \(A\) along \(N\) gives \(|\mu|=|g|\cdot\lambda+|\nu|\), since \(|g\cdot\lambda|=|g|\cdot\lambda\) (the sets \(\{g\ge0\}\), \(\{g<0\}\) form a Hahn decomposition for \(g\cdot\lambda\)). Hence the identity of Exercise 5.8.80, applied to \(\mu\) and then to the measure \(|\mu|\), whose absolutely continuous part has density \(|g|\), yields \(f^{\prime}=F_\mu^{\prime}=g\) a.e. and
\begin{equation*} V\big(f,[a,x]\big)^{\prime}=F_{|\mu|}^{\prime}(x) =|g(x)|=|f^{\prime}(x)|\qquad\text{a.e. on }[a,b). \end{equation*}
In general put \(g(a):=f(a)\) and \(g(x):=f(x-)\) for \(x\in(a,b]\), the left limits existing as \(f\) is a difference of two bounded nondecreasing functions. Then \(g\) is left continuous (for nondecreasing \(\varphi\), \(\lim_{y\uparrow x}\varphi(y-)=\sup_{y<x}\varphi(y)=\varphi(x-)\)), of bounded variation, since for \(t_1<\dots<t_{m+1}\)
\begin{equation*} \sum_i|g(t_{i+1})-g(t_i)| =\lim_{\epsilon\downarrow0}\sum_i|f(t_{i+1}-\epsilon)-f(t_i-\epsilon)| \le V\big(f,[a,b]\big) \end{equation*}
with \(\epsilon\) small enough that \(t_1-\epsilon>a\) (a collection starting at \(t_1=a\) is handled by \(|g(t_2)-g(a)|=\lim_{\epsilon\downarrow0}|f(t_2-\epsilon)-f(a)|\)), and \(g=f\) off the at most countable discontinuity set \(D\) of \(f\). So \(h:=f-g\) is of bounded variation and vanishes outside \(D\), and \(H(x):=V(h,[a,x])\) satisfies \(H^{\prime}=0\) a.e. by (i). For \(y<x\), inequality (5.2.1) on \([y,x]\) together with \(V(f,[y,x])=V(f,[a,x])-V(f,[a,y])\) (Proposition 5.2.2(iii)) and the same for \(g\) give
\begin{equation*} |w(x)-w(y)|\le H(x)-H(y),\qquad |h(x)-h(y)|\le H(x)-H(y), \end{equation*}
where \(w(x):=V(f,[a,x])-V(g,[a,x])\); hence \(w^{\prime}=h^{\prime}=0\) at every point where \(H^{\prime}=0\), i.e. a.e. With the left continuous case applied to \(g\),
\begin{equation*} V\big(f,[a,x]\big)^{\prime}=V\big(g,[a,x]\big)^{\prime} =|g^{\prime}(x)|=|f^{\prime}(x)|\qquad\text{a.e.} \end{equation*}
Let \(F=(F_1,\dots,F_n)\colon\mathbb{R}^n\to\mathbb{R}^n\) be a continuous mapping. Suppose that the functions \(F_i\) belong to the Sobolev class \(W^{p,r}(\mathbb{R}^n)\), where \(pr>n\). Prove that \(F\) has Lusin’s property (N), i.e., takes all measure zero sets to measure zero sets.
Morrey’s inequality on dyadic cubes gives \(\lambda_n^{*}\big(F(E)\big)=0\) whenever \(\lambda_n(E)=0\). First, \(|DF|\in L^{q}_{\mathrm{loc}}\) for some \(q>n\): for \(r=1\) take \(q=p>n\); for \(r>1\) the components of \(\nabla F_i\) lie in \(W^{p,r-1}(\mathbb{R}^n)\), and the Sobolev embedding theorem gives \(W^{p,r-1}\subset L^{q}\) with \(1/q=1/p-(r-1)/n\) when \(p(r-1)<n\), where
\begin{equation*} q=\frac{pn}{n-p(r-1)}>n\iff p>n-p(r-1)\iff pr>n , \end{equation*}
while for \(p(r-1)\ge n\) every finite \(q\) is locally admissible. Write \(p\) for this exponent from now on; \(F\) is continuous, hence bounded on cubes, so \(F_i\in W^{p,1}(Q)\) for every cube \(Q\).
Since \(p>n\), Morrey’s inequality provides \(C=C(n,p)\) with
\begin{equation*} |u(x)-u(y)|\le C\,\|\nabla u\|_{L^p(Q)}\,|x-y|^{\alpha}, \qquad \alpha:=1-\frac np\in(0,1), \end{equation*}
for the continuous representative of \(u\in W^{p,1}(Q)\) and all \(x,y\) in a cube \(Q\), with \(C\) independent of the position and size of \(Q\): translation invariance is clear, and under \(u_\rho(z):=u(\rho z)\) one has \(\|\nabla u_\rho\|_{L^p(Q/\rho)}=\rho^{1-n/p}\|\nabla u\|_{L^p(Q)}\) while \(|x/\rho-y/\rho|^{\alpha}=\rho^{-\alpha}|x-y|^{\alpha}\), and the factors cancel as \(\alpha=1-n/p\). Applying this to each \(F_i\) with \(|\nabla F_i|\le|DF|\),
\begin{equation*} |F(x)-F(y)|\le C\sqrt{n}\,\|DF\|_{L^p(Q)}\,|x-y|^{\alpha},\qquad x,y\in Q. \tag{1} \end{equation*}
By countable subadditivity of \(\lambda_n^{*}\) we may assume \(E\subset K\) for a fixed open cube \(K\) with compact closure; put \(M:=\|DF\|_{L^p(K)}<\infty\). Given \(\varepsilon>0\), take open \(E\subset U\subset K\) with \(\lambda_n(U)<\varepsilon\) and decompose \(U\) into at most countably many closed dyadic cubes \(Q_j\subset U\) with pairwise disjoint interiors (maximal dyadic cubes with closure in \(U\)) and edges \(r_j\), so that \(\sum_jr_j^{n}=\lambda_n(U)<\varepsilon\). Since \(\operatorname{diam}Q_j=\sqrt n\,r_j\), (1) puts \(F(Q_j)\) into a ball of radius \(C\sqrt n\,n^{\alpha/2}\|DF\|_{L^p(Q_j)}r_j^{\alpha}\), whence with \(C_1:=\omega_n\big(C\sqrt n\,n^{\alpha/2}\big)^n\) and Holder’s inequality for the conjugate exponents \(p/n\), \(p/(p-n)\),
\begin{equation*} \begin{aligned} \lambda_n^{*}\big(F(E)\big)&\le\sum_j\lambda_n^{*}\big(F(Q_j)\big) \le C_1\sum_j\|DF\|_{L^p(Q_j)}^{\,n}r_j^{\alpha n}\\ &\le C_1\Big(\sum_j\|DF\|_{L^p(Q_j)}^{\,p}\Big)^{n/p} \Big(\sum_jr_j^{\alpha np/(p-n)}\Big)^{(p-n)/p} \le C_1M^{\,n}\varepsilon^{(p-n)/p}, \end{aligned} \end{equation*}
because \(\alpha n\,p/(p-n)=n\) and the cubes \(Q_j\) have disjoint interiors inside \(K\). Letting \(\varepsilon\to0\) gives \(\lambda_n^{*}\big(F(E)\big)=0\).
(Sierpinski) Let \(f\colon\mathbb{R}\to\mathbb{R}\) be an arbitrary function and let \(\{h_n\}\) be a sequence of nonzero numbers approaching zero. Prove that there exists a function \(F\colon\mathbb{R}\to\mathbb{R}\) such that
\begin{equation*} \lim_{n\to\infty}\frac{F(x+h_n)-F(x)}{h_n}=f(x)\quad\text{for all }x . \end{equation*}
Construct \(F\) so that \(F(x+h_n)-F(x)=h_nf(x)\) for all \(n\ge N(x)\); no regularity is asserted, and this \(F\) is in general non-measurable. Since \(x\) and \(x+h_n\) lie in the same coset of the countable additive group \(G\) generated by \(\{h_n\}\), it suffices, choosing representatives of the cosets by the axiom of choice, to define \(F\) on each coset \(C=t+G\) separately; \(C\) is countably infinite and \(c+h_n\in C\) for \(c\in C\).
Collision lemma: for \(c\ne c^{\prime}\) in \(C\) only finitely many \(n\) satisfy \(c+h_n\in\{c^{\prime}\}\cup\{c^{\prime}+h_m:m\in\mathbb{N}\}\). Indeed, put \(d:=c^{\prime}-c\ne0\): the equality \(c+h_n=c^{\prime}\) means \(h_n=d\), which holds for finitely many \(n\) as \(h_n\to0\); the equality \(c+h_n=c^{\prime}+h_m\) means \(h_n-h_m=d\), impossible when \(n,m>I\), where \(|h_i|<|d|/2\) for \(i>I\), while \(n\le I\) leaves finitely many \(n\), and for each of the finitely many \(m\le I\) the requirement \(h_n=d+h_m\) has no solution if \(d+h_m=0\) (as \(h_n\ne0\)) and finitely many otherwise.
Enumerate \(C=\{c_1,c_2,\dots\}\) and proceed in stages, stage \(k\) assigning values on \(T_k:=\{c_k+h_n:n\ge N_k\}\) and setting \(F(c_k):=0\) if it is still undefined:
\begin{equation*} F(c_k+h_n):=F(c_k)+h_nf(c_k),\qquad n\ge N_k . \end{equation*}
The collision lemma applied to the finitely many \(j\le k\) (together with \(c_k+h_n\ne c_k\)) permits a choice of \(N_k\) so large that
\begin{equation*} T_k\cap\big(P_{k-1}\cup\{c_1,\dots,c_k\}\big)=\emptyset, \qquad P_{k-1}:=T_1\cup\dots\cup T_{k-1}. \tag{2} \end{equation*}
The recursion is then consistent: the prescription is unambiguous, since \(c_k+h_n=c_k+h_{n^{\prime}}\) forces \(h_n=h_{n^{\prime}}\); by (2) stage \(k\) overwrites no earlier value; and for \(j>k\) we have \(T_j\cap T_k=\emptyset\) and \(c_k\notin T_j\), while \(F(c_j)\) is set to \(0\) only when undefined, so no later stage overwrites stage \(k\). Every point of \(C\) is some \(c_k\), so \(F\) is defined on \(C\), and for \(x=c_k\) and every \(n\ge N_k\),
\begin{equation*} \frac{F(x+h_n)-F(x)}{h_n}=f(x). \end{equation*}
Carrying this out on every coset of \(G\) produces \(F\colon\mathbb{R}\to\mathbb{R}\) whose difference quotients converge to \(f(x)\) at every \(x\).
Prove that for every sequence of numbers \(h_n>0\) decreasing to zero, there exists a continuous function \(F\colon[0,1]\to\mathbb{R}\) such that, for every Lebesgue measurable function \(f\) on \([0,1]\), there exists a subsequence \(\{h_{n_k}\}\) with
\begin{equation*} \lim_{k\to\infty}\frac{F(x+h_{n_k})-F(x)}{h_{n_k}}=f(x)\qquad\text{a.e. on }[0,1]. \end{equation*}
Build a continuous \(F\) on \(\mathbb{R}\) vanishing outside \([0,1]\) whose difference quotients \(Q_n(x):=[F(x+h_n)-F(x)]/h_n\) are dense, for convergence in measure, in the set of a.e. finite measurable functions on \([0,1]\); then for a given \(f\) pick \(n_k\) with \(Q_{n_k}\to f\) in measure and pass to an a.e. convergent subsequence by Riesz’s Theorem 2.2.5(i). Convergence in measure is metrized by \(d(u,v):=\int_0^1\min(|u-v|,1)\,dx\), since \(\lambda(|u-v|>\varepsilon)\le d(u,v)/\min(\varepsilon,1)\) and \(d(u,v)\le\varepsilon+\lambda(|u-v|>\varepsilon)\).
Dense family: let \(\mathcal{D}\) be the countable set of step functions \(\sum_{k\le m}q_k\mathbf{1}_{[t_{k-1},t_k)}\) with rational values and rational partition points of \([0,1]\), extended by \(0\). Given an a.e. finite measurable \(f\) and \(\varepsilon>0\), truncate at a level \(M\) with \(\lambda(|f|>M)<\varepsilon/3\), replace the truncation by a continuous \(u\) off a set of measure \(<\varepsilon/3\) (Lusin’s theorem), and approximate the uniformly continuous \(u\) within \(\varepsilon\) by some \(q\in\mathcal{D}\) off a further set of measure \(<\varepsilon/3\); then \(\lambda(|f-q|>\varepsilon)<\varepsilon\).
Building block: for a step function \(G\) vanishing outside \([0,1]\) and constant on each \([t_{k-1},t_k)\), \(k\le m\), and any \(s,\delta>0\) there are a continuous piecewise linear \(\phi\) with finitely many breakpoints, vanishing outside \([0,1]\), with \(\|\phi\|_\infty\le s\), and \(h^{*}>0\) with
\begin{equation*} \lambda\Big(\Big\{x\in[0,1]:\ \frac{\phi(x+h)-\phi(x)}{h}\ne G(x)\Big\}\Big)\le\delta \qquad(0<h\le h^{*}). \end{equation*}
Indeed, let \(a_k\) be the value of \(G\) on \([t_{k-1},t_k)\), \(A:=1+\max_k|a_k|\), and choose \(P>0\), \(\theta\in(0,1/2)\) with \(AP\le s\), \(mP\le\delta/2\), \(2\theta\le\delta/2\). Split each \([t_{k-1},t_k)\) into consecutive blocks of length \(P\) (leaving a piece of length \(<P\)); on a block \([u,u+P]\) let \(\phi\) be affine with slope \(a_k\) on \([u,u+(1-\theta)P]\) and return linearly to \(0\) on \([u+(1-\theta)P,u+P]\), and set \(\phi:=0\) on the leftovers and outside \([0,1]\). Then \(\phi\) is continuous, \(\|\phi\|_\infty\le AP\le s\), and for \(h\le h^{*}:=\theta P\) the quotient equals \(G(x)=a_k\) whenever \([x,x+h]\) stays inside an affine stretch; the remaining \(x\) fill the leftovers (total length \(\le mP\le\delta/2\)) and one set of length \(\theta P+h\le2\theta P\) per block, with at most \(1/P\) blocks, of total length \(\le2\theta\le\delta/2\).
Construction: enumerate \(\mathcal{D}\) as \((g_j)\) with every element occurring infinitely often, put \(\varepsilon_j:=2^{-j}\), \(n_0:=0\), \(h_{n_0}:=1\), and recursively, given \(\phi_1,\dots,\phi_{j-1}\) and \(n_1<\dots<n_{j-1}\), set \(\Psi_{j-1}:=\sum_{i<j}\phi_i\) (piecewise linear with \(N_{j-1}\) breakpoints) and \(G_j:=g_j-\Psi_{j-1}^{\prime}\), which after modification on the finite set of breakpoints of \(\Psi_{j-1}\) and jump points of \(g_j\) (this changes the exceptional set above by a finite set only) is a step function of the required form. Apply the building block to \(G_j\) with \(s_j:=\min(2^{-j},2^{-j}h_{n_{j-1}})\), \(\delta_j:=\varepsilon_j/3\), obtaining \(\phi_j\) and \(h_j^{*}\), and choose \(n_j>n_{j-1}\) with
\begin{equation*} h_{n_j}\le h^{*}_j,\qquad N_{j-1}h_{n_j}\le\varepsilon_j/3, \qquad h_{n_j}\le\varepsilon_j/3 , \end{equation*}
possible since \(h_n\downarrow0\). As \(\|\phi_i\|_\infty\le2^{-i}\), the series \(F:=\sum_{i\ge1}\phi_i\) converges uniformly, so \(F\) is continuous and vanishes outside \([0,1]\).
Fix \(j\), write \(h:=h_{n_j}\), and let \(x\in[0,1]\) lie outside the exceptional set of the building block for \(\phi_j\) at scale \(h\le h_j^{*}\) (measure \(\le\varepsilon_j/3\)), outside the set of \(x\) for which \([x,x+h]\) contains a breakpoint of \(\Psi_{j-1}\) (measure \(\le N_{j-1}h\le\varepsilon_j/3\)), and outside \((1-h,1]\) (measure \(\le\varepsilon_j/3\)). For such \(x\) the function \(\Psi_{j-1}\) is affine on \([x,x+h]\), so
\begin{equation*} \sum_{i\le j}\frac{\phi_i(x+h)-\phi_i(x)}{h}=\Psi_{j-1}^{\prime}(x)+G_j(x)=g_j(x), \end{equation*}
while \(h_{n_{i-1}}\le h\) and \(\|\phi_i\|_\infty\le2^{-i}h_{n_{i-1}}\) for \(i>j\) give
\begin{equation*} \Big|\sum_{i>j}\frac{\phi_i(x+h)-\phi_i(x)}{h}\Big|\le\sum_{i>j}2^{1-i}=2^{1-j}. \end{equation*}
Hence \(|Q_{n_j}-g_j|\le2^{1-j}\) outside a set of measure \(\le\varepsilon_j\), and \(d(Q_{n_j},g_j)\le3\cdot2^{-j}\to0\).
Now given a.e. finite measurable \(f\) and \(\varepsilon>0\), take \(q\in\mathcal{D}\) with \(d(f,q)\le\varepsilon\); since \(q=g_j\) for infinitely many \(j\), we get \(d(Q_{n_j},f)\le2\varepsilon\) for infinitely many \(j\). Letting \(\varepsilon\downarrow0\) selects a strictly increasing sequence of indices along which \(Q_n\to f\) in measure, and Riesz’s Theorem 2.2.5(i) extracts from it a subsequence \(\{h_{n_k}\}\) along which the convergence holds a.e. on \([0,1]\).
Exercises 5.8.86–5.8.92
(Fichtenholz) Let \(\varphi\) be a nondecreasing function on \([c,d]\), let \(a=\varphi( c)\), \(b=\varphi(d)\), and let a function \(f\) be integrable on the interval \([a,b]\). Let
\begin{equation*} F(x)=\int_a^x f(t)\,dt . \end{equation*}
Suppose that the function \(F\circ\varphi\) is absolutely continuous. Prove the equality
\begin{equation*} \int_a^b f(x)\,dx=\int_c^d f\bigl(\varphi(y)\bigr)\varphi^{\prime}(y)\,dy . \end{equation*}
Everything reduces to \(\int_{E_0}G^{\prime}\,dy=0\), where \(G:=F\circ\varphi\) and \(E_1\), \(E_0\) are the sets of \(y\in(c,d)\) at which \(\varphi^{\prime}(y)\) exists and is positive, respectively zero. Indeed \(G\) is absolutely continuous by hypothesis, so the Newton–Leibniz formula gives
\begin{equation*} \int_c^dG^{\prime}(y)\,dy=G(d)-G( c)=F(b)-F(a)=\int_a^bf(x)\,dx , \end{equation*}
while \(E_0\cup E_1\) has full measure in \([c,d]\), the monotone \(\varphi\) being differentiable a.e. with finite nonnegative derivative and both sets being measurable (Theorem 5.8.12).
On \(E_1\): \(F\) is absolutely continuous as an indefinite integral and \(F^{\prime}=f\) off a null set \(Z\) (Theorem 5.4.2), and Lemma 5.8.13 applied to \(\varphi\), whose derivative is nonzero on \(E_1\), gives \(\lambda\big(\varphi^{-1}(Z)\cap E_1\big)=0\). So for a.e. \(y\in E_1\) the function \(\varphi\) is differentiable at \(y\) and \(F\) at \(\varphi(y)\) with \(F^{\prime}(\varphi(y))=f(\varphi(y))\), and the chain rule gives \(G^{\prime}(y)=f(\varphi(y))\varphi^{\prime}(y)\) there. On \(E_0\) we have \(\varphi^{\prime}(y)=0\) with \(f(\varphi(y))\) finite, so \(f(\varphi(y))\varphi^{\prime}(y)=0\); hence \(y\mapsto f(\varphi(y))\varphi^{\prime}(y)\) agrees a.e. with the integrable function \(G^{\prime}\mathbf 1_{E_1}\) (no separate measurability discussion for \(f\circ\varphi\) is needed) and
\begin{equation*} \int_c^df\big(\varphi(y)\big)\varphi^{\prime}(y)\,dy=\int_{E_1}G^{\prime}\,dy =\int_c^dG^{\prime}\,dy-\int_{E_0}G^{\prime}\,dy . \end{equation*}
Fix \(\varepsilon>0\). By absolute continuity of the Lebesgue integral pick \(\eta>0\) with \(\int_A|G^{\prime}|<\varepsilon\) whenever \(\lambda(A)<\eta\), and an open \(U\) with \(E_0\subset U\subset(c,d)\) and \(\lambda(U\setminus E_0)<\eta\), so that \(\int_{U\setminus E_0}|G^{\prime}|<\varepsilon\). By absolute continuity of \(F\) pick \(\delta>0\) such that disjoint intervals \((\alpha_i,\beta_i)\subset[a,b]\) with \(\sum_i(\beta_i-\alpha_i)\le\delta\) satisfy \(\sum_i|F(\beta_i)-F(\alpha_i)|\le\varepsilon\) (countable collections included, by passing to the limit). For \(u\in E_0\), \(\varphi^{\prime}(u)=0\) gives, for all small \(r>0\),
\begin{equation*} \varphi(u+r)-\varphi(u-r)\le\frac{2\delta r}{d-c},\qquad [u-r,u+r]\subset U , \end{equation*}
so these intervals form a Vitali cover of \(E_0\) and Theorem 5.5.1 provides pairwise disjoint \([a_i,b_i]\subset U\) covering \(E_0\) up to measure zero with \(\varphi(b_i)-\varphi(a_i)\le\delta(b_i-a_i)/(d-c)\), whence \(\sum_i(\varphi(b_i)-\varphi(a_i))\le\delta\) because \(\sum_i(b_i-a_i)\le d-c\). As \(\varphi\) is nondecreasing, the intervals \((\varphi(a_i),\varphi(b_i))\subset[a,b]\) are pairwise disjoint, so the choice of \(\delta\) and the absolute continuity of \(G\) give
\begin{equation*} \sum_i\Bigl|\int_{a_i}^{b_i}G^{\prime}\,dy\Bigr| =\sum_i\big|F\big(\varphi(b_i)\big)-F\big(\varphi(a_i)\big)\big|\le\varepsilon . \end{equation*}
Since the disjoint intervals \([a_i,b_i]\subset U\) cover \(E_0\) up to a null set,
\begin{equation*} \int_{E_0}G^{\prime}\,dy=\sum_i\int_{a_i}^{b_i}G^{\prime}\,dy -\sum_i\int_{[a_i,b_i]\setminus E_0}G^{\prime}\,dy , \end{equation*}
and the last sum is bounded in absolute value by \(\int_{U\setminus E_0}|G^{\prime}|<\varepsilon\); thus \(|\int_{E_0}G^{\prime}|\le2\varepsilon\) for every \(\varepsilon>0\), and \(\int_a^bf(x)\,dx=\int_c^df(\varphi(y))\varphi^{\prime}(y)\,dy\).
Let \(f\) be an increasing absolutely continuous function on \([0,1]\) such that \(f(0)=0\), \(f(1)=1\). Set
\begin{equation*} D:=\{t:\ 0<f^{\prime}(t)<+\infty\},\qquad g(t):=\inf\{s\in[0,1]:\ f(s)\le t\},\quad t\in[0,1]. \end{equation*}
Prove that \(f\bigl(g(t)\bigr)=t\), \(g\) is strictly increasing and for every bounded Borel function \(\varphi\) the following equality holds:
\begin{equation*} \int_D\varphi(t)\,dt=\int_0^1\varphi\bigl(g(s)\bigr)g^{\prime}(s)\,ds . \end{equation*}
The printed definition of \(g\) is a misprint, since \(f(0)=0\le t\) makes that infimum identically \(0\); the function meant is the right inverse \(g(t):=\sup\{s\in[0,1]:f(s)\le t\}\), and \(f\) is continuous nondecreasing with \(f([0,1])=[0,1]\).
\(f(g(t))=t\): the set \(S_t:=\{s\in[0,1]:f(s)\le t\}\) contains \(0\), and \(f(g(t))=\lim f(s_n)\le t\) for \(s_n\in S_t\), \(s_n\to g(t)\); if \(g(t)<1\) then \(f(s)>t\) for \(s>g(t)\), so \(f(g(t))\ge t\) as \(s\downarrow g(t)\), while \(g(t)=1\) gives \(1=f(1)=f(g(t))\le t\le1\). Also \(S_{t_1}\subset S_{t_2}\) for \(t_1<t_2\) makes \(g\) nondecreasing, and \(g(t_1)=g(t_2)\) would give \(t_1=t_2\) on applying \(f\), so \(g\) is strictly increasing.
\(f\) is injective on \(D\) with \(g(f(t))=t\) there: if \(f(t)=f(t^{\prime})\) for some \(t^{\prime}\ne t\), then \(f\) is constant on the segment between them, so the one-sided difference quotient at \(t\) vanishes and \(f^{\prime}(t)=0\), impossible for \(t\in D\); hence \(f(s)>f(t)\) for \(s>t\), i.e. \(S_{f(t)}=[0,t]\) and \(g(f(t))=t\).
\(\lambda\big(f([0,1]\setminus D)\big)=0\): write \([0,1]\setminus D=A_0\cup N\) with \(A_0:=\{f^{\prime}=0\}\) and \(N\) the null set where \(f\) has no finite derivative. Proposition 5.5.4 gives \(\lambda(f(A_0))\le\int_{A_0}|f^{\prime}|\,dt=0\), while \(\lambda(f(N))=0\) by Lusin’s property (N) of \(f\): for \(\varepsilon>0\) take the \(\delta\) of absolute continuity and an open \(U=\bigcup_i(\alpha_i,\beta_i)\subset(0,1)\) containing \(N\cap(0,1)\) with \(\lambda(U)<\delta\), so that \(\sum_i(f(\beta_i)-f(\alpha_i))\le\varepsilon\) (finite subsums, then \(k\to\infty\)) while \(f\big((\alpha_i,\beta_i)\big)\subset[f(\alpha_i),f(\beta_i)]\) by monotonicity and continuity. Since \(t=f(g(t))\), it follows that
\begin{equation*} g(t)\in D\qquad\text{for a.e. }t\in[0,1]. \tag{\(**\)} \end{equation*}
Being increasing, \(g\) is differentiable a.e. with finite nonnegative derivative and \(\int_0^1g^{\prime}\le g(1)-g(0)\le1\). At a point \(t\) where \(g^{\prime}(t)\) exists finitely and \((**)\) holds, \(f\) is differentiable at \(g(t)\) with \(0<f^{\prime}(g(t))<\infty\), so the chain rule applied to the identity \(f(g(t))=t\) gives
\begin{equation*} f^{\prime}\big(g(t)\big)g^{\prime}(t)=1\qquad\text{for a.e. }t\in[0,1]. \tag{\(\dagger\)} \end{equation*}
Put \(h(s):=\varphi(g(s))g^{\prime}(s)\), with \(g^{\prime}:=0\) on the null set where the derivative fails to exist or is infinite; then \(|h|\le(\sup|\varphi|)g^{\prime}\) is integrable, \(\varphi\circ g\) being Borel by monotonicity of \(g\). Since \(f\) is monotone and absolutely continuous with \(f(0)=0\), \(f(1)=1\), Corollary 5.4.4 gives (and also furnishes the integrability of the right-hand integrand)
\begin{equation*} \int_0^1h(x)\,dx=\int_0^1h\big(f(t)\big)f^{\prime}(t)\,dt . \tag{\(\ddagger\)} \end{equation*}
Let \(Z\) be the null set of points where \((\dagger)\) fails; by Lemma 5.8.13 applied to \(f\), whose derivative is nonzero on \(D\), we get \(\lambda(f^{-1}(Z)\cap D)=0\), so for a.e. \(t\in D\), using \(g(f(t))=t\) and \((\dagger)\) at \(u=f(t)\), \(f^{\prime}(t)g^{\prime}(f(t))=1\) and hence
\begin{equation*} h\big(f(t)\big)f^{\prime}(t) =\varphi\big(g(f(t))\big)g^{\prime}\big(f(t)\big)f^{\prime}(t) =\varphi(t). \end{equation*}
On \([0,1]\setminus D\) we have \(f^{\prime}=0\) a.e. and \(h\) finite, so that integrand vanishes a.e. there, and \((\ddagger)\) becomes \(\int_0^1\varphi(g(s))g^{\prime}(s)\,ds=\int_D\varphi(t)\,dt\).
Prove the following generalization of Vitali’s Theorem 5.5.2 that was indicated by Lebesgue. Let \(A\) be a set in \(\mathbb{R}^n\) and let \(\mathcal{F}\) be a family of closed sets with the following property: for every \(x\in A\), there exist a number \(\alpha(x)>0\), a sequence of sets \(F_k(x)\in\mathcal{F}\) and a sequence of cubes \(Q_k(x)\) such that \(x\in Q_k(x)\), \(F_k(x)\subset Q_k(x)\), \(\lambda_n\bigl(F_k(x)\bigr)>\alpha(x)\lambda_n\bigl(Q_k(x)\bigr)\) and \(\operatorname{diam}Q_k(x)\to0\). Then \(\mathcal{F}\) contains an at most countable subfamily of pairwise disjoint sets \(F_j\) whose union covers \(A\) up to a set of measure zero.
The selection argument of Theorem 5.5.2 works verbatim with balls replaced by the pairs \((F,Q)\); write \(\lambda:=\lambda_n\), \(\ell(Q)\) for the edge of a cube, so \(\operatorname{diam}Q=\sqrt n\,\ell(Q)\).
Selection lemma: let \(W\) be bounded open, \(\alpha\in(0,1)\), and let \(B\subset W\) be such that every \(x\in B\) and \(\varepsilon>0\) admit \(F\in\mathcal F\) and a cube \(Q\) with \(x\in Q\subset W\), \(F\subset Q\), \(\lambda(F)>\alpha\lambda(Q)\) and \(\operatorname{diam}Q<\varepsilon\); then for every \(\sigma>0\) there are finitely many pairwise disjoint \(F_1,\dots,F_N\in\mathcal F\) contained in \(W\) with \(\lambda^{*}\big(B\setminus\bigcup_{i\le N}F_i\big)<\sigma\). Indeed, let \(\mathcal S\) consist of all pairs \((F,Q)\) with \(F\in\mathcal F\), \(F\subset Q\subset W\), \(\lambda(F)>\alpha\lambda(Q)\) (so \(\lambda(Q)>0\)); if \(\mathcal S=\varnothing\) then \(B=\varnothing\). Otherwise choose recursively \((F_k,Q_k)\in\mathcal S_{k-1}:=\{(F,Q)\in\mathcal S: F\cap(F_1\cup\dots\cup F_{k-1})=\varnothing\}\) with
\begin{equation*} \lambda(Q_k)>\tfrac12\sup\{\lambda(Q):(F,Q)\in\mathcal S_{k-1}\}, \end{equation*}
the suprema being finite as all the cubes lie in the bounded set \(W\). If \(\mathcal S_K=\varnothing\) at some step, then \(B\subset F_1\cup\dots\cup F_K\) and we are done with \(\sigma=0\): a point \(x\in B\) outside this closed set has \(\rho:=\operatorname{dist}(x,F_1\cup\dots\cup F_K)>0\) and admits \((F,Q)\in\mathcal S\) with \(x\in Q\), \(\operatorname{diam}Q<\rho\), so that \((F,Q)\in\mathcal S_K\), a contradiction. Otherwise the \(F_k\) are disjoint subsets of \(W\), so
\begin{equation*} \sum_{k=1}^{\infty}\lambda(Q_k)<\frac1\alpha\sum_{k=1}^{\infty}\lambda(F_k) \le\frac{\lambda(W)}{\alpha}<\infty , \end{equation*}
whence \(\lambda(Q_k)\to0\). Let \(z_k\) be the centre of \(Q_k\) and \(V_k:=\overline B\big(z_k,\sqrt n(2^{1/n}+1)\ell(Q_k)\big)\), so that \(\lambda(V_k)=C_n\lambda(Q_k)\) with \(C_n=\omega_nn^{n/2}(2^{1/n}+1)^n\). Then \(B\setminus\bigcup_{k\le m}F_k\subset\bigcup_{k>m}V_k\): for such an \(x\) pick, as above, \((F,Q)\in\mathcal S_m\) with \(x\in Q\subset W\) and \(\operatorname{diam}Q<\operatorname{dist}(x,F_1\cup\dots\cup F_m)\); the set \(F\) must meet some \(F_l\), since otherwise \((F,Q)\in\mathcal S_k\) for all \(k\) and \(\lambda(Q_{k+1})>\frac12\lambda(Q)>0\), contradicting \(\lambda(Q_k)\to0\); for the least such \(l\) we have \(l>m\) and \((F,Q)\in\mathcal S_{l-1}\), so \(\lambda(Q_l)>\frac12\lambda(Q)\), i.e. \(\ell(Q)<2^{1/n}\ell(Q_l)\), and as \(Q\) and \(Q_l\) share a point \(w\),
\begin{equation*} |x-z_l|\le\operatorname{diam}Q+\operatorname{diam}Q_l \le\sqrt n\,(2^{1/n}+1)\ell(Q_l), \end{equation*}
that is, \(x\in V_l\). Hence \(\lambda^{*}\big(B\setminus\bigcup_{k\le m}F_k\big)\le C_n\sum_{k>m}\lambda(Q_k)\to0\), and \(m=N\) large enough proves the lemma.
Bounded case: let \(A\subset V\) with \(V\) bounded open and \(A_m:=\{x\in A:\alpha(x)>1/m\}\uparrow A\). Put \(K_0:=\varnothing\) and, given the closed union \(K_{m-1}\) of the finitely many disjoint sets already selected, apply the lemma with \(B:=R_m:=A_m\setminus K_{m-1}\), \(\alpha:=1/m\) and \(W_m:=O\cap(V\setminus K_{m-1})\), where \(O\supset R_m\) is open; its hypothesis holds because for \(x\in R_m\) the given \(F_k(x)\subset Q_k(x)\ni x\) satisfy \(\lambda(F_k(x))>\frac1m\lambda(Q_k(x))\) and \(Q_k(x)\subset W_m\) for large \(k\), as \(\operatorname{diam}Q_k(x)\to0\) and \(W_m\) is open. This yields finitely many disjoint sets of \(\mathcal F\) inside \(W_m\) (hence disjoint from all earlier ones) with union \(S_m\) such that \(\lambda^{*}(A_m\setminus K_m)=\lambda^{*}(R_m\setminus S_m)<2^{-m}\), where \(K_m:=K_{m-1}\cup S_m\). For the at most countable disjoint family \(\mathcal G\) of all selected sets, with union \(K_\infty\), every \(m\ge m_0\) gives \(\lambda^{*}(A_{m_0}\setminus K_\infty)\le\lambda^{*}(A_m\setminus K_m)<2^{-m}\), so \(\lambda^{*}(A_{m_0}\setminus K_\infty)=0\) and \(A\setminus K_\infty=\bigcup_{m_0}(A_{m_0}\setminus K_\infty)\) is null.
General case: partition \(\mathbb{R}^n\) into the open unit cubes \(U_j\) of the integer lattice and the null grid \(Z\) of the corresponding hyperplanes, and apply the bounded case to \(A\cap U_j\) inside \(V=U_j\) (for \(x\in A\cap U_j\) all but finitely many \(Q_k(x)\) lie in \(U_j\), with the same \(\alpha(x)\)), obtaining disjoint \(\mathcal G_j\subset\mathcal F\) inside \(U_j\) covering \(A\cap U_j\) up to measure zero. Sets from different \(\mathcal G_j\) are disjoint, and
\begin{equation*} A\setminus\bigcup\mathcal G \subset Z\cup\bigcup_j\Big((A\cap U_j)\setminus\bigcup\mathcal G_j\Big), \qquad \mathcal G:=\bigcup_j\mathcal G_j , \end{equation*}
is a countable union of null sets, so \(\lambda\big(A\setminus\bigcup\mathcal G\big)=0\).
Let \((X,\mathcal A,\mu)\) be a space with a finite nonnegative measure. A family \(\mathcal D\subset\mathcal A\) is called a Vitali system if it satisfies the following conditions: (a) \(\varnothing, X\in\mathcal D\), all nonempty sets in \(\mathcal D\) have positive measures, (b) if a set \(A\subset X\) is covered by a collection \(\mathcal E\subset\mathcal D\) in such a way that whenever \(x\in A\), \(B\in\mathcal D\) and \(x\in B\), there exists \(D\in\mathcal E\) with \(x\in D\subset B\), then one can find an at most countable subcollection of disjoint sets \(E_n\in\mathcal E\) with \(\mu\bigl(A\setminus\bigcup_{n=1}^\infty E_n\bigr)=0\). Suppose that \(\mathcal A\) contains all singletons and that \(\mu\) vanishes on them. Suppose we are given a sequence of countable partitions \(\Pi_n\) of the space \(X\) into measurable disjoint parts such that \(\Pi_{n+1}\) is a refinement of \(\Pi_n\). Finally, suppose that the collection \(\Pi=\bigcup_{n=1}^\infty\Pi_n\) is dense in the measure algebra \(\mathcal A/\mu\) and, for every set \(Z\) of measure zero and every \(\varepsilon>0\), there exists \(E_\varepsilon\in\Pi\) such that \(Z\subset E_\varepsilon\) and \(\mu(E_\varepsilon)<\varepsilon\). Prove that \(\Pi\) is a Vitali system. Show also that if a measure \(\nu\) is absolutely continuous with respect to \(\mu\), then
\begin{equation*} \frac{d\nu}{d\mu}(x)=\lim_{k\to\infty}\frac{\nu\bigl(B_k(x)\bigr)}{\mu\bigl(B_k(x)\bigr)}, \end{equation*}
where \(B_k(x)\in\Pi\) are chosen such that \(x\in B_k(x)\), \(0<\mu\bigl(B_k(x)\bigr)<k^{-1}\).
\(\Pi\) is a Vitali system because the maximal elements of a covering subfamily are disjoint with the same union, and the differentiation formula is the a.e. convergence of the conditional expectations \(P_nf=\mathbb E_\mu[f\mid\mathcal A_n]\). Write \(\Pi_n(x)\) for the element of \(\Pi_n\) containing \(x\), so that \(\Pi_1(x)\supset\Pi_2(x)\supset\cdots\) and any two members of the at most countable family \(\Pi\) are disjoint or nested; \(\mathcal A_n:=\sigma(\Pi_n)\) consists of the unions of members of \(\Pi_n\), and density of \(\Pi\) in \(\mathcal A/\mu\) is used in the form: for \(A\in\mathcal A\) and \(\varepsilon>0\) there are \(n\) and \(V\in\mathcal A_n\) with \(\mu(A\triangle V)<\varepsilon\). In (a) take \(\mathcal D=\{\varnothing,X\}\cup\{E\in\Pi:\mu(E)>0\}\); the members of \(\Pi\) of measure zero form an at most countable union \(Y\) with \(\mu(Y)=0\).
(b): let \(\mathcal E\subset\Pi\) cover \(A\) in the stated sense. Applying the hypothesis with \(B=X\in\mathcal D\) gives \(A\subset\bigcup\mathcal E\). Every \(E\in\mathcal E\cap\Pi_m\) lies in a maximal member of \(\mathcal E\): if \(B\in\Pi_k\) and \(B\supsetneq E\), then \(k>m\) would force \(B\subset E\), so \(k\le m\) and \(B\) is one of \(\Pi_1(E),\dots,\Pi_m(E)\), a finite chain. Distinct maximal elements are disjoint (meeting members of \(\Pi\) are nested), there are at most countably many of them, and their union is \(\bigcup\mathcal E\supset A\); enumerating them as \(E_1,E_2,\dots\) gives disjoint sets with \(A\setminus\bigcup_nE_n=\varnothing\).
Differentiation: the hypothesis on null sets applied to \(Z=\{x\}\) gives, for each \(\varepsilon>0\), some \(E\in\Pi\) with \(x\in E\) and \(\mu(E)<\varepsilon\), necessarily \(E=\Pi_n(x)\), so \(\mu(\Pi_n(x))\downarrow0\) for every \(x\). For \(x\notin Y\) all \(\mu(\Pi_n(x))\) are positive, so sets \(B_k(x)=\Pi_{n_k}(x)\) as in the statement exist and \(n_k\to\infty\), since a fixed \(n\) cannot satisfy \(0<\mu(\Pi_n(x))<k^{-1}\) for arbitrarily large \(k\). Let \(f:=d\nu/d\mu\in L^1(\mu)\) (Radon–Nikodym) and, for \(g\in L^1(\mu)\),
\begin{equation*} (P_ng)(x):=\frac{1}{\mu\big(\Pi_n(x)\big)}\int_{\Pi_n(x)}g\,d\mu\ \ (x\notin Y), \qquad (P_ng)(x):=0\ \ (x\in Y), \end{equation*}
so that \(P_ng=\mathbb E_\mu[g\mid\mathcal A_n]\) and \(\nu(\Pi_n(x))/\mu(\Pi_n(x))=(P_nf)(x)\) off \(Y\); the assertion is \(P_nf\to f\) a.e.
Maximal inequality: with \(Mg:=\sup_n|P_ng|\) and \(\mathcal E_\lambda:=\{E\in\Pi:\mu(E)>0,\ |\int_Eg\,d\mu|>\lambda\mu(E)\}\) we have \(\{Mg>\lambda\}=\bigcup\mathcal E_\lambda\) up to \(Y\), and the maximal elements \(E_j\) of \(\mathcal E_\lambda\) are pairwise disjoint with the same union, so
\begin{equation*} \mu(Mg>\lambda)\le\sum_j\mu(E_j)\le\frac1\lambda\sum_j\int_{E_j}|g|\,d\mu \le\frac1\lambda\|g\|_{L^1(\mu)} . \end{equation*}
If \(h\in L^1(\mu)\) is \(\mathcal A_m\)-measurable, then \(P_nh=h\) for \(n\ge m\); such \(h\) are dense in \(L^1(\mu)\), since simple functions are dense and each \(\mathbf 1_A\) is approximated in \(L^1(\mu)\) by \(\mathbf 1_V\) with \(V\in\mathcal A_m\), by the density of \(\Pi\). Hence, for \(\Lambda g:=\limsup_n|P_ng-g|\), the bound \(\Lambda g\le M(g-h)+|g-h|\) a.e. together with the maximal and Chebyshev inequalities gives
\begin{equation*} \mu(\Lambda g>\lambda)\le\mu\big(M(g-h)>\lambda/2\big)+\mu\big(|g-h|>\lambda/2\big) \le\frac4\lambda\|g-h\|_{L^1(\mu)} \end{equation*}
for every \(\lambda>0\), so \(\Lambda g=0\) a.e. Applying this to \(f\) and evaluating along \(n_k\to\infty\), for a.e. \(x\),
\begin{equation*} \frac{\nu\big(B_k(x)\big)}{\mu\big(B_k(x)\big)}=(P_{n_k}f)(x) \xrightarrow[k\to\infty]{}f(x)=\frac{d\nu}{d\mu}(x). \end{equation*}
(i) Let \(f\) be a bounded measurable function on a cube in \(\mathbb{R}^n\) with Lebesgue measure \(\lambda\). Prove that the set of points of the approximate continuity of \(f\) coincides with the set of its Lebesgue points. In particular, if \(f\) is a bounded measurable function on \([0,1]\), then the derivative of the function
\begin{equation*} \int_0^x f(t)\,dt \end{equation*}
equals \(f(x)\) at every point \(x\) of the approximate continuity of \(f\).
(ii) Prove that if a function \(f\) is integrable on a cube, then every Lebesgue point of \(f\) is a point of the approximate continuity, but the converse is not true.
For bounded measurable \(f\) the two sets coincide; in general a Lebesgue point is a point of approximate continuity, but not conversely. Assume \(f(x_0)=0\) (replace \(f\) by \(f-f(x_0)\)), let \(B_r:=B(x_0,r)\) with \(x_0\) interior to the cube, and use the reformulation of approximate continuity from Exercise 5.8.91: \(\operatorname{ap\,lim}_{y\to x_0}f(y)=f(x_0)\) means that \(\{|f-f(x_0)|<\varepsilon\}\) has \(x_0\) as a density point for every \(\varepsilon>0\).
(ii), first assertion (only integrability is used): for \(\varepsilon>0\) and \(A_\varepsilon:=\{|f|\ge\varepsilon\}\), Chebyshev’s inequality on \(B_r\) gives
\begin{equation*} \frac{\lambda(A_\varepsilon\cap B_r)}{\lambda(B_r)} \le\frac{1}{\varepsilon\,\lambda(B_r)}\int_{B_r}|f(y)|\,dy\xrightarrow[r\to0]{}0 , \end{equation*}
so \(x_0\) is a density point of \(\{|f|<\varepsilon\}\) for every \(\varepsilon>0\).
(i), the converse for \(|f|\le C\): given \(\varepsilon>0\) put \(A:=\{|f|<\varepsilon\}\) and take \(r_0\) with \(\lambda(B_r\setminus A)\le\varepsilon\lambda(B_r)\) for \(0<r<r_0\); then
\begin{equation*} \int_{B_r}|f(y)|\,dy\le\varepsilon\lambda(B_r)+C\lambda(B_r\setminus A) \le(1+C)\varepsilon\lambda(B_r), \end{equation*}
so the Lebesgue limit is \(0\) and the two sets coincide. Consequently, for bounded measurable \(f\) on \([0,1]\), \(F(x):=\int_0^xf(t)\,dt\) and \(x\) a point of approximate continuity,
\begin{equation*} \Bigl|\frac{F(x+h)-F(x)}{h}-f(x)\Bigr| \le2\cdot\frac{1}{2|h|}\int_{x-|h|}^{x+|h|}|f(t)-f(x)|\,dt\xrightarrow[h\to0]{}0 , \end{equation*}
i.e. \(F^{\prime}(x)=f(x)\).
(ii), the counterexample: let \(f\) be even on \([-1,1]\) with \(f(0)=0\), \(f=4^n\) on \((2^{-n}-8^{-n},2^{-n})\) and \(f=0\) elsewhere on \((0,1]\); the listed intervals are nonempty (\(8^{-n}<2^{-n-1}\)) and \(\int_{-1}^1|f|\,dt=2\sum_n4^n8^{-n}=2\). It is approximately continuous at \(0\): for \(\varepsilon>0\) the set \(A_\varepsilon=\{|f|\ge\varepsilon\}\) lies in the union of those intervals and their reflections, so for \(2^{-m-1}\le r\le2^{-m}\),
\begin{equation*} \frac{\lambda\big(A_\varepsilon\cap[-r,r]\big)}{2r} \le\frac{2\sum_{n\ge m}8^{-n}}{2^{-m}}=\frac{16}{7}\,4^{-m}\xrightarrow[r\to0]{}0 . \end{equation*}
But \(0\) is not a Lebesgue point: for \(r=2^{-m}\), \(\int_0^{2^{-m}}f=\sum_{n\ge m}2^{-n}=2^{-m+1}\), whence
\begin{equation*} \frac{1}{2r}\int_{-r}^{r}|f(t)-f(0)|\,dt=\frac{2\cdot2^{-m+1}}{2\cdot2^{-m}}=2 . \end{equation*}
Prove that the approximate continuity of a function \(f\) on \(\mathbb{R}^n\) at a point \(x\) is equivalent to the equality \(\operatorname{ap\,lim}_{y\to x}f(y)=f(x)\).
The two conditions are equivalent. Translating, assume \(x=0\) and \(f(0)=0\), write \(C( r):=[-r,r]^n\), and note that having \(0\) as a density point is the same for cubes and for balls, since \(B(0,r)\subset C( r)\subset B(0,r\sqrt n)\) with volumes comparable by a factor depending only on \(n\).
Necessity: let \(0\) be a density point of a measurable \(E_0\) with \(f(y)\to0\) as \(y\to0\) along \(E_0\). Given \(\varepsilon>0\) there is \(\delta>0\) with \(E_0\cap B(0,r)\subset A_\varepsilon:=\{|f|<\varepsilon\}\) for \(r<\delta\) (up to the point \(0\)), so
\begin{equation*} \frac{\lambda\big(A_\varepsilon\cap B(0,r)\big)}{\lambda\big(B(0,r)\big)} \ge\frac{\lambda\big(E_0\cap B(0,r)\big)}{\lambda\big(B(0,r)\big)} \xrightarrow[r\to0]{}1 , \end{equation*}
i.e. \(\operatorname{ap\,lim}_{y\to0}f(y)=0=f(0)\).
Sufficiency: let \(E_k:=\{|f|<1/k\}\), each having \(0\) as a density point, and choose \(\varepsilon_k\le\varepsilon_{k-1}/2\) with \(\varepsilon_k\downarrow0\) and
\begin{equation*} \lambda\big(C(\rho)\setminus E_{k+1}\big)\le2^{-k}\lambda\big(C(\rho)\big) \qquad\text{for all }0<\rho\le\varepsilon_k. \tag{1} \end{equation*}
With \(C_k:=C(\varepsilon_k)\) put \(E:=\{0\}\cup\bigcup_{k\ge1}\big(E_{k+1}\cap(C_k\setminus C_{k+1})\big)\). If \(y\in E\setminus\{0\}\) and \(|y|_\infty\le\varepsilon_k\), then \(y\) lies in \(C_j\setminus C_{j+1}\) for exactly one \(j\), necessarily \(j\ge k\), and then \(y\in E_{j+1}\), so \(|f(y)|<1/(j+1)\le1/(k+1)\); hence \(f(y)\to0=f(0)\) along \(E\). For the density, \((C_j\setminus C_{j+1})\setminus E\subset C_j\setminus E_{j+1}\) with \(\lambda(C_j\setminus E_{j+1})\le2^{-j}\lambda(C_j)\) by (1), while \(\bigcap_jC_j=\{0\}\subset E\), so
\begin{equation*} \lambda\big(C_m\setminus E\big)\le\sum_{j\ge m}2^{-j}\lambda(C_j) \le2^{1-m}\lambda(C_m). \tag{2} \end{equation*}
Given \(0<r\le\varepsilon_1\) choose \(k\) with \(\varepsilon_{k+1}<r\le\varepsilon_k\) and split \(C( r)\setminus E\subset(C( r)\setminus E_{k+1})\cup(C_{k+1}\setminus E)\), using that \(C( r)\setminus C_{k+1}\subset C_k\setminus C_{k+1}\), where \(E\supset E_{k+1}\); then (1) with \(\rho=r\) and (2) with \(\lambda(C_{k+1})\le\lambda(C( r))\) give
\begin{equation*} \frac{\lambda\big(C( r)\setminus E\big)}{\lambda\big(C( r)\big)} \le2^{-k}+2^{-k}=2^{1-k}\xrightarrow[r\to0]{}0 , \end{equation*}
since \(k\to\infty\) as \(r\to0\). Thus \(0\) is a density point of the measurable set \(E\) along which \(f\) tends to \(f(0)\).
Let us consider in \([0,1]\) the class \(\Delta\) of all measurable sets every point of which is a density point, and the empty set.
(i) Prove that \(\Delta\) is a topology that is strictly stronger than the usual topology of the interval. This topology is called the density topology.
(ii) Show that a function is continuous in the topology \(\Delta\) precisely when it is Lebesgue measurable.
(i) \(\Delta\) is a topology strictly stronger than the usual one; (ii) \(\Delta\)-continuity at a point is exactly approximate continuity there, so that \(\Delta\)-continuous means everywhere approximately continuous, and Lebesgue measurable means \(\Delta\)-continuous almost everywhere. At the endpoints the density is computed relative to \([0,1]\) (otherwise \([0,1]\notin\Delta\)); write \(I_r:=[0,1]\cap(x-r,x+r)\), and \(A^d\) for the set of density points of a measurable \(A\), so that \(\lambda(A\triangle A^d)=0\) by Theorem 5.6.2 applied to \(\mathbf 1_A\).
(i) Clearly \(\varnothing,[0,1]\in\Delta\). If \(A,B\in\Delta\) and \(x\in A\cap B\), then
\begin{equation*} \lambda\big(I_r\setminus(A\cap B)\big) \le\lambda(I_r\setminus A)+\lambda(I_r\setminus B) =o\big(\lambda(I_r)\big), \end{equation*}
so \(A\cap B\in\Delta\). For \(A:=\bigcup_\alpha A_\alpha\) with \(A_\alpha\in\Delta\), measurability follows from Lemma 5.8.10, and any \(x\in A\) lies in some \(A_\alpha\subset A\), so it is a density point of \(A_\alpha\) and hence of \(A\). A relatively open \(U\) satisfies \(I_r\subset U\) for small \(r\), so \(U\in\Delta\); the inclusion is strict, since \([0,1]\setminus\mathbb{Q}\in\Delta\) has empty interior. Moreover, for every measurable \(A\),
\begin{equation*} A^d\cap A\in\Delta , \tag{\(*\)} \end{equation*}
because it differs from \(A\) by the null set \(A\setminus A^d\), hence has the same density points as \(A\), and each of its points is one of them.
(ii) \(f\) is \(\Delta\)-continuous at \(x\) exactly when it is approximately continuous at \(x\): if \(x\in A\subset A_\varepsilon:=\{y:|f(y)-f(x)|<\varepsilon\}\) for some \(A\in\Delta\), then \(x\), being a density point of \(A\), is one of \(A_\varepsilon\), so \(\operatorname{ap\,lim}_{y\to x}f(y)=f(x)\), which is approximate continuity by Exercise 5.8.91; conversely each \(A_\varepsilon\) is then measurable with \(x\) as a density point, and \(A_\varepsilon^d\cap A_\varepsilon\in\Delta\) by \((*)\) contains \(x\) and lies in \(A_\varepsilon\). Hence a \(\Delta\)-continuous \(f\) is Lebesgue measurable, since \(\{f<r\}=\bigcup_{m}f^{-1}\big((-m,r)\big)\) is a countable union of \(\Delta\)-open, hence measurable, sets, while a finite measurable \(f\) is approximately continuous a.e. (Theorem 5.8.9), i.e. \(\Delta\)-continuous a.e. The printed assertion must be read in this almost everywhere sense: \(f=\mathbf 1_{[0,1/2]}\) is measurable, but \(f^{-1}\big((1/2,3/2)\big)=[0,1/2]\) has density \(1/2\) at the point \(1/2\), so \(f\) is not \(\Delta\)-continuous there, and no change on a null set repairs this, since \(\{|f-c|<\varepsilon\}\) has density at most \(1/2\) at \(1/2\) for every constant \(c\) and small \(\varepsilon\).
Exercises 5.8.93–5.8.99
Prove Theorem 5.8.5.
The Fejer kernel and its two-sided bound give both equalities at a Lebesgue point \(x\) of the \(2\pi\)-periodically extended \(f\). With \(S_k\) the partial sums of the Fourier series (4.3.5) and \(\sigma_n:=(S_0+\dots+S_{n-1})/n\) (the averaging in the printed (4.3.7); an index shift is immaterial for the limit), formula (4.3.6) together with \(\sum_{k=0}^{n-1}\sin(2k+1)z=\sin^2nz/\sin z\) at \(z=(t-x)/2\) yields
\begin{equation*} \sigma_n(x)=\int_0^{2\pi}f(x+z)\Phi_n(z)\,dz,\qquad \Phi_n(z)=\frac{1}{2\pi n}\Big(\frac{\sin\frac{nz}{2}}{\sin\frac z2}\Big)^{2}. \end{equation*}
Here \(\Phi_n\ge0\) is even and \(2\pi\)-periodic with \(\int_{-\pi}^{\pi}\Phi_n=1\) (apply the formula to \(f\equiv1\), for which \(\sigma_n\equiv1\)), and since \(|\sin(z/2)|\ge|z|/\pi\) and \(|\sin(nz/2)|\le\min\{n|z|/2,1\}\),
\begin{equation*} \Phi_n(z)\le C\min\Big\{n,\ \frac{1}{nz^2}\Big\},\qquad 0<|z|\le\pi,\quad C:=\pi . \end{equation*}
Put \(g(z):=f(x+z)+f(x-z)-2f(x)\) and \(G( r):=\int_0^r|g(z)|\,dz\); then \(G( r)\le\int_{-r}^{r}|f(x+z)-f(x)|\,dz=o( r)\) by the Lebesgue point property, and \(G\) is nondecreasing and absolutely continuous with \(G^{\prime}=|g|\) a.e. (Theorem 5.4.2). Evenness of \(\Phi_n\) and \(\int_{-\pi}^{\pi}\Phi_n=1\) give
\begin{equation*} \sigma_n(x)-f(x)=\int_{-\pi}^{\pi}\big[f(x+z)-f(x)\big]\Phi_n(z)\,dz =\int_0^{\pi}g(z)\Phi_n(z)\,dz . \end{equation*}
Fix \(\varepsilon>0\), choose \(\delta\in(0,\pi)\) with \(G( r)\le\varepsilon r\) on \((0,\delta]\), and let \(n>1/\delta\).
(i) \(\int_0^{1/n}|g|\Phi_n\,dz\le Cn\,G(1/n)\le C\varepsilon\).
(ii) Integrating by parts against the absolutely continuous \(G\), then discarding the nonpositive term \(-n^2G(1/n)\) and using \(G(z)\le\varepsilon z\) with \(n\delta>1\),
\begin{equation*} \int_{1/n}^{\delta}|g|\Phi_n\,dz\le\frac Cn\Big[\frac{G(\delta)}{\delta^2}-n^2G(1/n) +2\int_{1/n}^{\delta}\frac{G(z)}{z^3}dz\Big] \le\frac{C\varepsilon}{n\delta}+2C\varepsilon\le3C\varepsilon . \end{equation*}
(iii) \(\int_{\delta}^{\pi}|g|\Phi_n\,dz\le\frac{C}{n\delta^2}G(\pi)\to0\), as \(G(\pi)<\infty\).
Hence \(\limsup_n|\sigma_n(x)-f(x)|\le4C\varepsilon\) for every \(\varepsilon>0\), i.e. \(\sigma_n(x)\to f(x)\). Finally put \(\alpha_0:=a_0/2\) and \(\alpha_k:=a_k\cos kx+b_k\sin kx\), so that \(S_n=\sum_{k\le n}\alpha_k\) and \(|\alpha_k|\le2\pi^{-1}\int_0^{2\pi}|f|\,dt\), whence \(S( r):=\sum_{k\ge0}\alpha_kr^k\) converges absolutely for \(0\le r<1\). The series \(\sum_k\alpha_k\) is Cesaro summable to \(f(x)\) by the above, hence Abel summable to \(f(x)\) by Exercise 4.7.51:
\begin{equation*} \frac{a_0}{2}+\lim_{r\to1^-}\sum_{n=1}^{\infty}\big[a_n\cos nx+b_n\sin nx\big]r^n =\lim_{r\to1^-}S( r)=f(x). \end{equation*}
Prove that the spaces \(W^{p,1}(\Omega)\) and \(BV(\Omega)\) with the indicated norms are Banach spaces.
Both spaces are complete, because \(L^p(\Omega)\), \(L^1(\Omega)\) and the space \(\mathcal M(\Omega)\) of bounded Borel measures with the variation norm are, and the defining identities pass to the limit. Fix a mollifier \(\zeta\in C_0^\infty\), \(\zeta\ge0\), supported in the unit ball with \(\int\zeta=1\), and \(\zeta_\varepsilon(x)=\varepsilon^{-n}\zeta(x/\varepsilon)\).
Uniqueness of the generalized objects rests on two facts. (A) If \(g\in L^1_{loc}(\Omega)\) and \(\int_\Omega\psi g\,dx=0\) for all \(\psi\in C_0^\infty(\Omega)\), then \(g=0\) a.e.: for compact \(K\subset\Omega\), \(\varepsilon_0:=\operatorname{dist}(K,\mathbb{R}^n\setminus\Omega)/3\) and \(h:=gI_{K^{\prime}}\), where \(K^{\prime}\) is the compact \(2\varepsilon_0\)-neighbourhood of \(K\), we get \(h*\zeta_\varepsilon=0\) on \(K\) for \(\varepsilon<\varepsilon_0\) (as \(y\mapsto\zeta_\varepsilon(x-y)\in C_0^\infty(\Omega)\)), while \(h*\zeta_\varepsilon\to h\) in \(L^1\) by Theorem 4.2.4; exhaust \(\Omega\) by compacta. (B) A bounded Borel measure \(\nu\) on \(\Omega\) with \(\int_\Omega\psi\,d\nu=0\) for all \(\psi\in C_0^\infty(\Omega)\) vanishes: for continuous \(\varphi\) with compact support in \(\Omega\), \(\varphi*\zeta_\varepsilon\in C_0^\infty(\Omega)\) converges to \(\varphi\) uniformly, and \(|\int\varphi\,d\nu|\le\|\varphi-\varphi*\zeta_\varepsilon\|_\infty\|\nu\|\to0\); applying this to \(\varphi_k:=\max\{0,1-k\operatorname{dist}(\cdot,F)\}\to I_F\) with dominated convergence for \(|\nu|\) gives \(\nu(F)=0\) for every compact \(F\subset\Omega\), hence \(\nu(\Omega)=\lim_m\nu(K_m)=0\) along an exhaustion, and since the compacta form a class closed under finite intersections generating \(\mathcal B(\Omega)\), the parts \(\nu^{+}\), \(\nu^{-}\) of the Hahn–Jordan decomposition agree there and have equal masses, so \(\nu^{+}=\nu^{-}\) (Lemma 1.9.4).
\(W^{p,1}(\Omega)\): identity (5.8.8) of Definition 5.8.23 involves only the values on \(\operatorname{supp}\psi\), so it is read for \(f,g_i\in L^1_{loc}(\Omega)\supset L^p(\Omega)\) (Holder’s inequality), and by (A) the derivatives \(\partial_{x_i}f\) are unique and depend linearly on \(f\). Hence \(\|f\|_{W^{p,1}}=\|f\|_{L^p}+\||\nabla f|\|_{L^p}\) is a norm (subadditive by \(|\nabla(f+g)|\le|\nabla f|+|\nabla g|\) a.e., and vanishing only for \(f=0\)), equivalent to \(\|f\|_{L^p}+\sum_i\|\partial_{x_i}f\|_{L^p}\) since \(\max_i\|\partial_{x_i}f\|_{L^p}\le\||\nabla f|\|_{L^p}\le\sum_i\|\partial_{x_i}f\|_{L^p}\). If \((f_j)\) is Cauchy, the Riesz–Fischer theorem (Theorem 4.1.3) yields \(f=\lim_jf_j\) and \(g_i=\lim_j\partial_{x_i}f_j\) in \(L^p(\Omega)\); on \(K:=\operatorname{supp}\psi\) Holder gives \(\int_K|u|\,dx\le\|u\|_{L^p}\lambda_n(K)^{1-1/p}\), so \(L^p\)-convergence implies \(L^1(K)\)-convergence and
\begin{equation*} \int_\Omega\partial_{x_i}\psi\,f_j\,dx=-\int_\Omega\psi\,\partial_{x_i}f_j\,dx \end{equation*}
passes to the limit: \(f\in W^{p,1}(\Omega)\) with \(\partial_{x_i}f=g_i\) and \(\|f_j-f\|_{W^{p,1}}\to0\).
\(BV(\Omega)\): by (B) the bounded measures \(D_if\) in \(\int_\Omega\partial_{x_i}\psi\,f\,dx=-\int_\Omega\psi\,D_if(dx)\), \(\psi\in C_0^\infty(\Omega)\), are unique and linear in \(f\), so \(\|f\|_{BV}=\|f\|_{L^1(\Omega)}+\|Df\|\) is a norm, and
\begin{equation*} \max_{i\le n}\|D_if\|\le\|Df\|\le\sum_{i=1}^{n}\|D_if\| \end{equation*}
whether \(\|Df\|\) means \(\sup_{|e|\le1}\|(e,Df)\|\) or the total variation of the vector measure \(Df\); hence \(\|\cdot\|_{BV}\) is equivalent to \(\|f\|_{L^1(\Omega)}+\sum_i\|D_if\|\), for which completeness suffices. A Cauchy sequence \((f_j)\) is then Cauchy in \(L^1(\Omega)\) and each \((D_if_j)_j\) in the variation norm, both spaces being complete (Theorems 4.1.3 and 4.6.1), so \(f_j\to f\) in \(L^1(\Omega)\) and \(D_if_j\to\nu_i\) in variation. Since
\begin{equation*} \Bigl|\int_\Omega\partial_{x_i}\psi\,(f_j-f)\,dx\Bigr| \le\sup|\partial_{x_i}\psi|\,\|f_j-f\|_{L^1(\Omega)}, \qquad \Bigl|\int_\Omega\psi\,d(D_if_j-\nu_i)\Bigr|\le\sup|\psi|\,\|D_if_j-\nu_i\| , \end{equation*}
the defining identity passes to the limit, so \(f\in BV(\Omega)\) with \(D_if=\nu_i\), and \(D_i(f_j-f)=D_if_j-\nu_i\) gives \(f_j\to f\) in \(BV(\Omega)\).
Prove that \(f\in BV(\mathbb{R}^n)\) precisely when there exists a sequence of functions \(f_j\in C_0^\infty(\mathbb{R}^n)\) such that \(f_j\to f\) in \(L^1(\mathbb{R}^n)\) and \(\sup_j\bigl\||\nabla f_j|\bigr\|_{L^1(\mathbb{R}^n)}<\infty\).
Mollify and cut off in one direction, represent a bounded functional in the other. Fix \(\zeta\in C_0^\infty\), \(\zeta\ge0\), supported in the unit ball, \(\int\zeta\,dx=1\), put \(\zeta_\varepsilon(x)=\varepsilon^{-n}\zeta(x/\varepsilon)\) and \(f_\varepsilon:=f*\zeta_\varepsilon\); by 5.8(ix), \(f\in BV(\mathbb{R}^n)\) means \(f\in L^1\) and each \(D_if\) is a bounded Borel measure with
\begin{equation*} \int_{\mathbb{R}^n}\partial_{x_i}\psi\,f\,dx=-\int_{\mathbb{R}^n}\psi\,D_if(dx), \qquad \psi\in C_0^\infty(\mathbb{R}^n). \tag{\(*\)} \end{equation*}
Necessity. Applying \((*)\) to \(\psi(y)=\zeta_\varepsilon(x-y)\in C_0^\infty\), for which \(\partial_{x_i}\zeta_\varepsilon(x-y)=-\partial_{y_i}\psi(y)\), and then Tonelli’s theorem for the nonnegative kernel and the finite measure \(|D_if|\),
\begin{equation*} \begin{aligned} \partial_{x_i}f_\varepsilon(x)&=\int_{\mathbb{R}^n}\zeta_\varepsilon(x-y)\,D_if(dy),\\ \int_{\mathbb{R}^n}|\partial_{x_i}f_\varepsilon|\,dx &\le\int_{\mathbb{R}^n}\Bigl(\int_{\mathbb{R}^n}\zeta_\varepsilon(x-y)\,dx\Bigr)|D_if|(dy)=\|D_if\| , \end{aligned} \end{equation*}
so \(\||\nabla f_\varepsilon|\|_{L^1}\le N:=\sum_{i\le n}\|D_if\|\) for every \(\varepsilon>0\), while \(\|f_\varepsilon-f\|_{L^1}\to0\) by Theorem 4.2.4. Now cut off: with \(\chi\in C_0^\infty\), \(0\le\chi\le1\), \(\chi=1\) on \(B(0,1)\), \(\chi=0\) off \(B(0,2)\), \(c=\sup|\nabla\chi|\) and \(\chi_R(x)=\chi(x/R)\), the function \(\chi_Rf_\varepsilon\) lies in \(C_0^\infty(\mathbb{R}^n)\) and the product rule with \(|\nabla\chi_R|\le c/R\) gives
\begin{equation*} \bigl\||\nabla(\chi_Rf_\varepsilon)|\bigr\|_{L^1}\le N+\frac{c\|f\|_{L^1}}{R}, \qquad \|\chi_Rf_\varepsilon-f_\varepsilon\|_{L^1}\xrightarrow[R\to\infty]{}0 \end{equation*}
(the second by dominated convergence). Choosing \(R_j\ge j\) with \(\|\chi_{R_j}f_{1/j}-f_{1/j}\|_{L^1}\le1/j\) and \(c\|f\|_{L^1}/R_j\le1\), the functions \(f_j:=\chi_{R_j}f_{1/j}\in C_0^\infty\) satisfy \(f_j\to f\) in \(L^1\) and \(\||\nabla f_j|\|_{L^1}\le N+1\).
Sufficiency. Let \(f_j\in C_0^\infty\) with \(f_j\to f\) in \(L^1\) and \(M:=\sup_j\||\nabla f_j|\|_{L^1}<\infty\); in particular \(f\in L^1(\mathbb{R}^n)\). Classical integration by parts for the smooth compactly supported \(\psi,f_j\) and the boundedness of \(\partial_{x_i}\psi\) give, for each \(\psi\in C_0^\infty\),
\begin{equation*} L_i(\psi):=\lim_{j\to\infty}\int_{\mathbb{R}^n}\psi\,\partial_{x_i}f_j\,dx =-\int_{\mathbb{R}^n}\partial_{x_i}\psi\,f\,dx,\qquad |L_i(\psi)|\le M\sup_x|\psi(x)| . \end{equation*}
Since \(C_0^\infty\) is sup-norm dense in \(C_0(\mathbb{R}^n)\) (truncate by \(\chi_R\), then mollify: uniform continuity), \(L_i\) extends to \(C_0(\mathbb{R}^n)\) with norm \(\le M\); Hahn–Banach extends it norm-preservingly to \(C(S)\) for the metrizable compactification \(S=\mathbb{R}^n\cup\{\infty\}\), and the Riesz representation theorem (Theorem 7.10.4 of Volume 2) produces a bounded Borel measure whose restriction \(\nu_i\) to \(\mathcal{B}(\mathbb{R}^n)\) has \(\|\nu_i\|\le M\) and \(L_i(\psi)=\int\psi\,d\nu_i\) for \(\psi\in C_0^\infty\) (such \(\psi\) vanish at \(\infty\)). This is \((*)\) with \(D_if=\nu_i\) bounded for every \(i\le n\), i.e. \(f\in BV(\mathbb{R}^n)\).
Prove that \(f\in BV(\mathbb{R}^n)\) precisely when for every \(i\le n\), the functions
\begin{equation*} \psi_i(x_1,\dots,x_{n-1})(t)=f(x_1,\dots,x_{i-1},t,x_i,\dots,x_{n-1}) \end{equation*}
have bounded essential variations \(\|\psi_i(x_1,\dots,x_{n-1})(\,\cdot\,)\|_{BV}\) for a.e. \((x_1,\dots,x_{n-1})\) in \(\mathbb{R}^{n-1}\) (see Exercise 5.8.74) and
\begin{equation*} \int_{\mathbb{R}^{n-1}}\|\psi_i(x_1,\dots,x_{n-1})(\,\cdot\,)\|_{BV}\,dx_1\cdots dx_{n-1}<\infty . \end{equation*}
Everything follows from the identity \(\int_{\mathbb{R}^{n-1}}\|\psi_i(x^{\prime})\|_{BV}\,dx^{\prime}=\|D_if\|\) (both sides possibly \(+\infty\)), which we prove through difference quotients. Write \(x^{\prime}=(x_1,\dots,x_{n-1})\), let \(e_i\) be the \(i\)-th basis vector, and for \(u\in L^1(\mathbb{R})\), \(h>0\) put
\begin{equation*} \Lambda_h(u):=\frac1h\int_{\mathbb{R}}|u(t+h)-u(t)|\,dt,\qquad V(u):=\sup_{h>0}\Lambda_h(u)\in[0,+\infty], \end{equation*}
and for \(f\in L^1(\mathbb{R}^n)\), \(i\le n\), put
\begin{equation*} H^i_h(f):=\frac1h\int_{\mathbb{R}^n}|f(x+he_i)-f(x)|\,dx,\qquad V_i(f):=\sup_{h>0}H^i_h(f). \end{equation*}
Throughout, \(\zeta_\varepsilon\) is the mollifier of Exercise 5.8.95.
(i) Monotonicity. Telescoping \(u(t+mh)-u(t)\) over the \(m\) steps of length \(h\) gives \(\Lambda_{mh}(u)\le\Lambda_h(u)\), so \(k\mapsto\Lambda_{2^{-k}}(u)\) is nondecreasing; \(h\mapsto\Lambda_h(u)\) is continuous on \((0,\infty)\) since translation is \(L^1\)-continuous (Lemma 4.2.3), so taking \(m=\lfloor h_02^k\rfloor\ge1\) in \(\Lambda_{m2^{-k}}(u)\le\Lambda_{2^{-k}}(u)\) and letting \(k\to\infty\) yields \(\Lambda_{h_0}(u)\le\lim_k\Lambda_{2^{-k}}(u)\) for every \(h_0>0\). Hence, and verbatim in \(\mathbb{R}^n\),
\begin{equation*} V(u)=\lim_{k\to\infty}\Lambda_{2^{-k}}(u),\qquad V_i(f)=\lim_{k\to\infty}H^i_{2^{-k}}(f), \tag{1} \end{equation*}
both increasing limits.
(ii) \(V(u)=\|u\|_{BV}\), the essential variation of Exercise 5.8.74, with \(V(u)=\infty\) exactly when no version of \(u\) has bounded variation. Indeed, if \(g=u\) a.e. has bounded variation, then splitting \(\mathbb{R}\) into the intervals \([kh,(k+1)h)\) and substituting \(t=s+kh\), \(s\in[0,h)\),
\begin{equation*} \int_{\mathbb{R}}|g(t+h)-g(t)|\,dt=\int_0^h\sum_{k\in\mathbb{Z}}|g(s+(k+1)h)-g(s+kh)|\,ds\le\int_0^hV(g,\mathbb{R})\,ds=h\,V(g,\mathbb{R}), \end{equation*}
since for fixed \(s\) the points \(s+kh\) increase, so the inner sum is a variation sum; hence \(V(u)\le V(g,\mathbb{R})\) and, taking the infimum over such \(g\), \(V(u)\le\|u\|_{BV}\). Conversely let \(M:=V(u)<\infty\) and \(u_\varepsilon:=u*\zeta_\varepsilon\). Fubini’s theorem applied to \(u_\varepsilon(t+h)-u_\varepsilon(t)=\int\zeta_\varepsilon(s)[u(t+h-s)-u(t-s)]\,ds\) gives \(\Lambda_h(u_\varepsilon)\le\Lambda_h(u)\le M\), so Fatou’s theorem on the difference quotients yields \(\int|u_\varepsilon^{\prime}|\,dt\le M\) and therefore \(V(u_\varepsilon,\mathbb{R})\le M\), every variation sum being \(\sum_j|\int_{t_{j-1}}^{t_j}u_\varepsilon^{\prime}|\le\int|u_\varepsilon^{\prime}|\). By Theorem 4.2.4, \(u_\varepsilon\to u\) in \(L^1\); pick \(\varepsilon_k\to0\) with \(u_{\varepsilon_k}\to u\) pointwise on a full-measure set \(T\) (Theorem 2.2.5) and let \(g_0\) be that limit, so \(V(g_0,T)\le M\). Exercise 5.8.72 extends \(g_0\) to \(g\) on \(\mathbb{R}\) with \(V(g,\mathbb{R})=V(g_0,T)\le M\), and \(g=u\) a.e., so \(\|u\|_{BV}\le M\).
(iii) \(D_if\) is a bounded measure iff \(V_i(f)<\infty\), and then \(\|D_if\|=V_i(f)\). If \(D_if\) is bounded, then, by the computation of Exercise 5.8.95, \(\partial_{x_i}f_\varepsilon(x)=\int\zeta_\varepsilon(x-y)D_if(dy)\) with \(\int|\partial_{x_i}f_\varepsilon|\,dx\le\|D_if\|\), whence
\begin{equation*} \int_{\mathbb{R}^n}|f_\varepsilon(x+he_i)-f_\varepsilon(x)|\,dx=\int_{\mathbb{R}^n}\Bigl|\int_0^h\partial_{x_i}f_\varepsilon(x+se_i)\,ds\Bigr|dx\le h\|D_if\| . \end{equation*}
Letting \(\varepsilon\to0\) (Theorem 4.2.4) gives \(V_i(f)\le\|D_if\|\). Conversely let \(M:=V_i(f)<\infty\) and \(\psi\in C_0^\infty\). Since \(h^{-1}[\psi(\cdot+he_i)-\psi]\to\partial_{x_i}\psi\) uniformly with supports in a fixed compact set, and since after the change of variables
\begin{equation*} \int_{\mathbb{R}^n}f(x)\frac{\psi(x+he_i)-\psi(x)}{h}\,dx=\int_{\mathbb{R}^n}\frac{f(x-he_i)-f(x)}{h}\,\psi(x)\,dx , \end{equation*}
we obtain
\begin{equation*} \Bigl|\int_{\mathbb{R}^n}f\,\partial_{x_i}\psi\,dx\Bigr|=\lim_{h\to0}\Bigl|\int_{\mathbb{R}^n}\frac{f(x-he_i)-f(x)}{h}\psi(x)\,dx\Bigr|\le M\sup_x|\psi(x)| , \end{equation*}
the functional \(\psi\mapsto-\int f\,\partial_{x_i}\psi\,dx\) is sup-norm bounded by \(M\); so exactly as in Exercise 5.8.95 (extension to \(C_0(\mathbb{R}^n)\), Hahn–Banach, Riesz on the one-point compactification) there is a bounded \(\nu_i\) with \(\|\nu_i\|\le M\) representing it; thus \(D_if=\nu_i\) and \(\|D_if\|\le V_i(f)\).
(iv) Fubini and conclusion. Since \(f\in L^1(\mathbb{R}^n)\), Fubini’s theorem (Theorem 3.4.4), applied to the integrable \((x^{\prime},t)\mapsto|f(x+he_i)-f(x)|\), puts \(\psi_i(x^{\prime})\) in \(L^1(\mathbb{R})\) for a.e. \(x^{\prime}\), makes \(x^{\prime}\mapsto\Lambda_h(\psi_i(x^{\prime}))\) measurable, and gives \(H^i_h(f)=\int_{\mathbb{R}^{n-1}}\Lambda_h(\psi_i(x^{\prime}))\,dx^{\prime}\). Taking \(h=2^{-k}\), the integrands increase to \(V(\psi_i(x^{\prime}))\) by (1), so by monotone convergence and (1) again,
\begin{equation*} \begin{aligned} \int_{\mathbb{R}^{n-1}}V\bigl(\psi_i(x^{\prime})\bigr)\,dx^{\prime} &=\lim_{k\to\infty}\int_{\mathbb{R}^{n-1}}\Lambda_{2^{-k}}\bigl(\psi_i(x^{\prime})\bigr)\,dx^{\prime}\\ &=\lim_{k\to\infty}H^i_{2^{-k}}(f)=V_i(f), \end{aligned} \end{equation*}
where both sides may be \(+\infty\). By (ii) the integrand equals \(\|\psi_i(x^{\prime})(\,\cdot\,)\|_{BV}\) (and is \(+\infty\) exactly when that section has no bounded-variation version), so with (iii),
\begin{equation*} \int_{\mathbb{R}^{n-1}}\|\psi_i(x^{\prime})(\,\cdot\,)\|_{BV}\,dx^{\prime}=V_i(f)=\|D_if\| . \tag{2} \end{equation*}
If \(f\in BV(\mathbb{R}^n)\), each \(D_if\) is bounded, so \(V_i(f)<\infty\) by (iii) and (2) makes the sections finite a.e. with finite integral; conversely, finiteness of the integrals in the statement forces \(V_i(f)<\infty\) for all \(i\), hence every \(D_if\) bounded, i.e. \(f\in BV(\mathbb{R}^n)\).
Suppose that a compact set \(E\) has a smooth boundary with the outer normal \(n\). Show that \(DI_E=n\cdot\sigma_{\partial E}\), where \(\sigma_{\partial E}\) is the surface measure on the boundary of \(E\). In addition, the perimeter of \(E\) equals the surface measure of the boundary of \(E\).
\(DI_E=-n\,\sigma_{\partial E}\) and \(\mathcal{P}(E)=\sigma_{\partial E}(\partial E)\); the printed sign \(DI_E=n\cdot\sigma_{\partial E}\) belongs to the convention opposite to (5.8.9), and only \(|n|=1\) is used below. Here \(E=\overline{E^{\circ}}\) with \(\partial E\) a compact \(C^1\) hypersurface carrying a continuous outer unit normal \(n\), so \(\sigma_{\partial E}\) is a finite Borel measure.
Since \(E\) is bounded, \(I_E\in L^1(\mathbb{R}^n)\). For \(\psi\in C_0^\infty(\mathbb{R}^n)\) and \(i\le n\), the Gauss–Ostrogradskii theorem applied to the \(C^1\) field \(F=\psi e_i\) on \(E\) (legitimate: \(\partial E\) is a compact \(C^1\) hypersurface), with \(\mathrm{div}\,F=\partial_{x_i}\psi\), gives
\begin{equation*} \int_E\partial_{x_i}\psi(x)\,dx=\int_{\partial E}\bigl(F(x),n(x)\bigr)\,\sigma_{\partial E}(dx)=\int_{\partial E}\psi(x)\,n_i(x)\,\sigma_{\partial E}(dx). \end{equation*}
Thus, putting \(\nu_i(A):=-\int_{A\cap\partial E}n_i\,d\sigma_{\partial E}\), a bounded Borel measure since \(|n_i|\le1\) and \(\sigma_{\partial E}(\partial E)<\infty\), we have
\begin{equation*} \int_{\mathbb{R}^n}\partial_{x_i}\psi\,I_E\,dx=-\int_{\mathbb{R}^n}\psi\,\nu_i(dx),\qquad\psi\in C_0^\infty(\mathbb{R}^n), \end{equation*}
which is Definition 5.8.24 for \(I_E\,dx\) in the direction \(e_i\); by the uniqueness proved as Lemma B in Exercise 5.8.94, \(D_iI_E=\nu_i\). Hence \(DI_E=\theta\,\sigma_{\partial E}\) with \(\theta:=-n\) continuous and \(|\theta|\equiv1\) on \(\partial E\), and \(I_E\in BV(\mathbb{R}^n)\).
For the perimeter, \(\mathcal{P}(E)=\|DI_E\|\) is the total variation of this vector measure, i.e. the supremum of \(\sum_j|DI_E(A_j)|\) over finite Borel partitions; the sup over fields \(|e|\le1\) in the text must be read over measurable \(e\), not constant vectors (for the closed unit disc the constant-vector reading gives \(4\), not \(2\pi\)). For any partition \(\{A_j\}\),
\begin{equation*} \sum_{j}|DI_E(A_j)|=\sum_j\Bigl|\int_{A_j\cap\partial E}\theta\,d\sigma_{\partial E}\Bigr| \le\sigma_{\partial E}(\partial E). \end{equation*}
Conversely, fix \(\eta\in(0,1)\); by uniform continuity of \(\theta\) on the compact set \(\partial E\) pick \(\delta>0\) with \(|\theta(x)-\theta(y)|\le\eta\) for \(|x-y|\le\delta\), cover \(\partial E\) by finitely many balls of radius \(\delta/2\) and disjointify to a Borel partition \(\{A_j\}\) of \(\partial E\) of diameters \(\le\delta\) (adjoining \(A_0=\mathbb{R}^n\setminus\partial E\), where \(DI_E\) vanishes), and choose \(x_j\in A_j\). Since \(|\theta(x_j)|=1\),
\begin{equation*} \Bigl|\int_{A_j}\theta\,d\sigma_{\partial E}\Bigr|\ge\sigma_{\partial E}(A_j)-\int_{A_j}|\theta-\theta(x_j)|\,d\sigma_{\partial E} \ge(1-\eta)\,\sigma_{\partial E}(A_j), \end{equation*}
so \(\|DI_E\|\ge(1-\eta)\sigma_{\partial E}(\partial E)\); letting \(\eta\to0\) gives \(\mathcal{P}(E)=\sigma_{\partial E}(\partial E)\).
(i) Verify that \(\ln|x|\in BMO(\mathbb{R}^n)\);
(ii) Prove that \(\ln|G(x)|\in BMO(\mathbb{R}^n)\) for every polynomial \(G\) on \(\mathbb{R}^n\) of degree \(d\ge1\).
(iii) Prove that for every positive number \(\alpha<d^{-1}\), there exists a constant \(c_{\alpha,d}\) such that for every polynomial \(G\) on \(\mathbb{R}^n\) of degree \(d\ge1\) one has
\begin{equation*} \int_B|G(x)|^{-\alpha}\,dx\le c_{\alpha,d}\Bigl(\int_B|G(x)|\,dx\Bigr)^{-\alpha}, \end{equation*}
where \(B\) is the unit ball.
All three hold, with \(\|\ln|G|\|_{BMO}\) bounded by a constant depending only on \(n\) and \(d\). Write \(B:=B(0,1)\), \(f_{B^{\prime}}=\lambda(B^{\prime})^{-1}\int_{B^{\prime}}f\,dx\); throughout we use that a bound \(\lambda(B^{\prime})^{-1}\int_{B^{\prime}}|f-c_{B^{\prime}}|\,dx\le A\) for arbitrary constants \(c_{B^{\prime}}\) forces \(|f_{B^{\prime}}-c_{B^{\prime}}|\le A\) and hence \(\|f\|_{BMO}\le2A\).
(i) Let \(f(x)=\ln|x|\), locally integrable since \(\int_0^R|\ln\rho|\rho^{n-1}d\rho<\infty\), and let \(B^{\prime}=B(x_0,r)\).
(a) \(|x_0|\le2r\), so \(B^{\prime}\subset B(0,3r)\); take \(c_{B^{\prime}}=\ln r\) and substitute \(x=rz\):
\begin{equation*} \begin{aligned} \frac{1}{\lambda(B^{\prime})}\int_{B^{\prime}}\bigl|\ln|x|-\ln r\bigr|\,dx &\le\frac{1}{\lambda(B(0,r))}\int_{B(0,3r)}\Bigl|\ln\frac{|x|}{r}\Bigr|\,dx\\ &=\frac{1}{\lambda(B)}\int_{B(0,3)}\bigl|\ln|z|\bigr|\,dz=:A_1<\infty , \end{aligned} \end{equation*}
a constant depending only on \(n\).
(b) \(|x_0|>2r\): then \(|x_0|/2<|x_0|-r\le|x|\le|x_0|+r<\tfrac32|x_0|\) on \(B^{\prime}\), so \(c_{B^{\prime}}=\ln|x_0|\) gives mean oscillation at most \(\ln2\).
Hence \(\ln|x|\in BMO(\mathbb{R}^n)\) with \(A=\max\{A_1,\ln2\}\).
(iii) We prove (iii) first, since (ii) follows from it; note a polynomial of degree \(d\ge1\) is not identically zero, so \(|G|>0\) off a closed null set.
Lemma 1. If \(I\subset\mathbb{R}\) is a bounded interval, \(p\) a real polynomial of degree \(\le d\), \(E\subset I\) measurable with \(\lambda_1(E)=\mu>0\) and \(|p|\le\tau\) on \(E\), then
\begin{equation*} \sup_{I}|p|\le(d+1)\Bigl(\frac{(2d+2)|I|}{\mu}\Bigr)^{d}\tau . \end{equation*}
Proof. Let \(F(s)=\lambda_1\bigl(E\cap(-\infty,s]\bigr)\), continuous and nondecreasing from \(0\) to \(\mu\). Choose \(\tau_0<\tau_1<\cdots<\tau_{2d+2}\) with \(F(\tau_i)=i\mu/(2d+2)\) and put \(E_i=E\cap(\tau_{i-1},\tau_i]\), so \(\lambda_1(E_i)=\mu/(2d+2)=:\delta\). Pick \(s_j\in E_{2j+1}\) for \(j=0,\dots,d\). If \(j>k\) then between \(s_k\) and \(s_j\) lie the sets \(E_{2k+2},\dots,E_{2j}\), whence \(s_j-s_k\ge(j-k)\delta\). By the Lagrange interpolation formula, for \(s\in I\),
\begin{equation*} p(s)=\sum_{j=0}^{d}p(s_j)\prod_{k\ne j}\frac{s-s_k}{s_j-s_k},\qquad \Bigl|\prod_{k\ne j}\frac{s-s_k}{s_j-s_k}\Bigr|\le\frac{|I|^{d}}{\delta^{d}\,j!\,(d-j)!}\le\Bigl(\frac{|I|}{\delta}\Bigr)^{d}, \end{equation*}
because \(|s-s_k|\le|I|\) and \(|s_j-s_k|\ge|j-k|\delta\). Since \(|p(s_j)|\le\tau\), the claim follows. \(\square\)
Lemma 2. There is \(C=C(n,d)\) such that for every polynomial \(G\) of degree \(\le d\) with \(\max_{\overline{B}}|G|=1\) and every \(t>0\),
\begin{equation*} \lambda\bigl\{x\in B:|G(x)|\le t\bigr\}\le C\,t^{1/d}. \end{equation*}
Proof. Pick \(x_*\in\overline{B}\) with \(|G(x_*)|=1\); by convexity of \(\overline B\), for each unit \(\omega\) the set \(\{s>0:x_*+s\omega\in B\}\) is an interval \((0,\rho(\omega))\) with \(\rho(\omega)\le2\), so in polar coordinates centred at \(x_*\) (Jacobian \(s^{n-1}\le2^{n-1}\)),
\begin{equation*} \lambda\{x\in B:|G(x)|\le t\}\le2^{n-1}\int_{S^{n-1}}\lambda_1\bigl\{s\in I_\omega:|p_\omega(s)|\le t\bigr\}\,d\omega , \end{equation*}
where \(I_\omega=[0,\rho(\omega)]\) and \(p_\omega(s)=G(x_*+s\omega)\) is a polynomial in \(s\) of degree at most \(d\). Since \(0\in I_\omega\), we have \(\sup_{I_\omega}|p_\omega|\ge|G(x_*)|=1\). Applying Lemma 1 with \(E=\{s\in I_\omega:|p_\omega(s)|\le t\}\), \(\tau=t\) and \(|I_\omega|\le2\) gives
\begin{equation*} 1\le(d+1)\Bigl(\frac{4d+4}{\mu}\Bigr)^{d}t,\qquad\text{i.e.}\qquad \mu=\lambda_1(E)\le(4d+4)(d+1)^{1/d}\,t^{1/d}. \end{equation*}
Integrating over \(S^{n-1}\) yields the assertion with \(C=2^{n-1}\sigma(S^{n-1})(4d+4)(d+1)^{1/d}\). \(\square\)
Now let \(0<\alpha<1/d\), \(\deg G=d\ge1\), \(m:=\max_{\overline{B}}|G|>0\) and \(H:=G/m\), so \(\max_{\overline B}|H|=1\) and \(|H|^{-\alpha}\ge1\) on \(B\). By the layer-cake formula (Theorem 3.4.7),
\begin{equation*} \int_B|H|^{-\alpha}dx=\int_0^{\infty}\lambda\bigl\{x\in B:|H(x)|^{-\alpha}>u\bigr\}\,du =\lambda(B)+\int_1^{\infty}\lambda\bigl\{x\in B:|H(x)|<u^{-1/\alpha}\bigr\}\,du , \end{equation*}
since for \(u<1\) the set in question is all of \(B\). By Lemma 2 with \(t=u^{-1/\alpha}\),
\begin{equation*} \int_1^{\infty}\lambda\{|H|<u^{-1/\alpha}\}\,du\le C\int_1^{\infty}u^{-1/(\alpha d)}\,du=\frac{C\alpha d}{1-\alpha d}<\infty , \end{equation*}
the integral converging precisely because \(\alpha d<1\). Hence
\begin{equation*} \int_B|H|^{-\alpha}dx\le c^{\prime}(n,d,\alpha):=\lambda(B)+\frac{C\alpha d}{1-\alpha d}. \end{equation*}
Finally, \(\int_B|G|\,dx\le m\lambda(B)\), i.e. \(m^{-\alpha}\le\lambda(B)^{\alpha}\bigl(\int_B|G|\,dx\bigr)^{-\alpha}\), so
\begin{equation*} \int_B|G|^{-\alpha}dx=m^{-\alpha}\int_B|H|^{-\alpha}dx\le c^{\prime}\lambda(B)^{\alpha}\Bigl(\int_B|G|\,dx\Bigr)^{-\alpha}. \end{equation*}
This is (iii) with \(c_{\alpha,d}=c^{\prime}(n,d,\alpha)\lambda(B)^{\alpha}\).
(ii) Fix \(\alpha:=1/(2d)\in(0,1/d)\), let \(B^{\prime}=B(x_0,r)\) be any ball and set
\begin{equation*} m_{B^{\prime}}:=\frac{1}{\lambda(B^{\prime})}\int_{B^{\prime}}|G|\,dx>0,\qquad u:=\frac{|G|}{m_{B^{\prime}}},\qquad c_{B^{\prime}}:=\ln m_{B^{\prime}} . \end{equation*}
Then \(\bigl|\ln|G|-c_{B^{\prime}}\bigr|=(\ln u)^{+}+(\ln u)^{-}\). Since \(\ln u\le u\), the average of \((\ln u)^{+}\) over \(B^{\prime}\) is at most \(\lambda(B^{\prime})^{-1}\int_{B^{\prime}}u\,dx=1\); and \(\ln s\le\alpha^{-1}s^{\alpha}\) for \(s\ge1\), taken at \(s=1/u\), gives
\begin{equation*} (\ln u)^{-}\le\frac{1}{\alpha}u^{-\alpha}=\frac{m_{B^{\prime}}^{\alpha}}{\alpha}|G|^{-\alpha}. \end{equation*}
Transfer (iii) to \(B^{\prime}\) by \(x=x_0+rz\): since \(G_{B^{\prime}}(z):=G(x_0+rz)\) again has degree \(d\),
\begin{equation*} \int_{B^{\prime}}|G|^{-\alpha}dx=r^{n}\int_B|G_{B^{\prime}}|^{-\alpha}dz\le r^{n}c_{\alpha,d}\Bigl(\int_B|G_{B^{\prime}}|\,dz\Bigr)^{-\alpha} =r^{n}c_{\alpha,d}\Bigl(r^{-n}\int_{B^{\prime}}|G|\,dx\Bigr)^{-\alpha}. \end{equation*}
As \(\lambda(B^{\prime})=r^{n}\lambda(B)\) and \(\int_{B^{\prime}}|G|\,dx=\lambda(B^{\prime})m_{B^{\prime}}\), the right side is \(r^{n}c_{\alpha,d}(\lambda(B)m_{B^{\prime}})^{-\alpha}\), so
\begin{equation*} \frac{1}{\lambda(B^{\prime})}\int_{B^{\prime}}(\ln u)^{-}dx\le\frac{m_{B^{\prime}}^{\alpha}}{\alpha\lambda(B^{\prime})}\int_{B^{\prime}}|G|^{-\alpha}dx \le\frac{m_{B^{\prime}}^{\alpha}\,r^{n}c_{\alpha,d}\lambda(B)^{-\alpha}m_{B^{\prime}}^{-\alpha}}{\alpha\,r^{n}\lambda(B)}=\frac{c_{\alpha,d}}{\alpha\,\lambda(B)^{1+\alpha}} , \end{equation*}
a constant depending only on \(n,d\). Adding the two estimates,
\begin{equation*} \frac{1}{\lambda(B^{\prime})}\int_{B^{\prime}}\bigl|\ln|G|-c_{B^{\prime}}\bigr|\,dx \le A:=1+\frac{c_{\alpha,d}}{\alpha\lambda(B)^{1+\alpha}} \end{equation*}
for every ball, so \(\ln|G|\) is locally integrable and \(\|\ln|G|\|_{BMO}\le2A\).
Verify that if \(\omega\in A_p\), then the measure \(\omega(x)\,dx\) has the doubling property.
The doubling constant is \(c=2^{np}/C^{\prime}\), where \(C^{\prime}>0\) is the constant in the equivalent form (5.8.7) of the \(A_p\) condition recorded in 5.8(vii): for all nonnegative bounded measurable \(f\) and all balls \(B\),
\begin{equation*} (f_B)^p\le\frac{1}{C^{\prime}}\Bigl(\int_B\omega\,dx\Bigr)^{-1}\int_Bf^p\omega\,dx, \qquad f_B=\frac{1}{\lambda(B)}\int_Bf\,dx. \tag{5.8.7} \end{equation*}
Fix \(x_0\) and \(r>0\), put \(\mu:=\omega(x)\,dx\) (a nonnegative Borel measure, finite on balls since \(\omega\) is locally integrable), and apply (5.8.7) with \(B=B(x_0,2r)\) and the nonnegative bounded measurable \(f=I_{B(x_0,r)}\), for which \(f^p=f\) and \(f_B=2^{-n}\):
\begin{equation*} 2^{-np}\le\frac{1}{C^{\prime}}\,\frac{\mu\bigl(B(x_0,r)\bigr)}{\mu\bigl(B(x_0,2r)\bigr)}, \qquad\text{i.e.}\qquad \mu\bigl(B(x_0,2r)\bigr)\le\frac{2^{np}}{C^{\prime}}\,\mu\bigl(B(x_0,r)\bigr), \end{equation*}
the denominator being finite so that clearing it is legitimate (and the claim trivial if \(\mu(B(x_0,2r))=0\)).
Exercises 5.8.100–5.8.106
Let \(\mu\) be a nonnegative bounded Borel measure on \(\mathbb{R}^n\) and let \(B(x,r)\) be the closed ball of radius \(r > 0\) centered at \(x\). Prove that
\begin{equation*} \limsup_{y \to x} \mu\big(B(y,r)\big) \le \mu\big(B(x,r)\big). \end{equation*}
Squeeze \(B(y,r)\) between closed balls about \(x\): if \(|x-y|<1/k\) and \(z\in B(y,r)\), then \(|z-x|\le|z-y|+|y-x|<r+1/k\), so \(B(y,r)\subset B(x,r+1/k)\) and
\begin{equation*} \sup_{0<|y-x|<1/k}\mu\big(B(y,r)\big) \le \mu\big(B(x,r+1/k)\big). \end{equation*}
The closed balls \(B(x,r+1/k)\) decrease with intersection exactly \(B(x,r)\) (here closedness is essential), and \(\mu\) is bounded, so continuity from above gives
\begin{equation*} \limsup_{y \to x} \mu\big(B(y,r)\big)=\inf_{k}\sup_{0<|y-x|<1/k}\mu\big(B(y,r)\big) \le\lim_{k\to\infty}\mu\big(B(x,r+1/k)\big)=\mu\big(B(x,r)\big). \end{equation*}
(\(\circ\)) Prove Lemma 5.7.4.
For convenience we recall the statement. A function \(f\) is Henstock–Kurzweil integrable on \([a,b]\) precisely when for every \(\varepsilon > 0\) there exists a positive function \(\delta\) on \([a,b]\) such that \(|I(f,\mathcal{P}_1) - I(f,\mathcal{P}_2)| < \varepsilon\) for all tagged partitions \(\mathcal{P}_1\) and \(\mathcal{P}_2\) subordinate to \(\delta\). The same assertion with free tagged partitions in place of tagged partitions is true for the McShane integral.)
Necessity is the triangle inequality: if \(|I(f,\mathcal{P})-I|<\varepsilon/2\) for every \(\mathcal{P}\) subordinate to the \(\delta\) furnished by Definition 5.7.3 for \(\varepsilon/2\), then \(|I(f,\mathcal{P}_1)-I(f,\mathcal{P}_2)|<\varepsilon\) for any two such partitions. (Here \(I(f,\mathcal{P})=\sum_if(x_i)(d_i-c_i)\) as in Section 5.7.)
For sufficiency, apply the hypothesis with \(\varepsilon=2^{-n}\) to get positive \(\gamma_n\) with
\begin{equation*} |I(f,\mathcal{P}_1) - I(f,\mathcal{P}_2)| < 2^{-n}\quad\text{for }\mathcal{P}_1,\mathcal{P}_2 \text{ subordinate to } \gamma_n , \tag{1} \end{equation*}
and put \(\delta_n:=\min\{\gamma_1,\dots,\gamma_n\}\), again positive, with \(\delta_{n+1}\le\delta_n\); anything subordinate to \(\delta_n\) is subordinate to every \(\gamma_m\), \(m\le n\). By Lemma 5.7.2 choose a tagged partition \(\mathcal{P}_n\) subordinate to \(\delta_n\). For \(m\ge n\) both \(\mathcal{P}_n\) and \(\mathcal{P}_m\) are subordinate to \(\gamma_n\), so (1) gives \(|I(f,\mathcal{P}_n)-I(f,\mathcal{P}_m)|<2^{-n}\); hence \(I:=\lim_nI(f,\mathcal{P}_n)\) exists and \(|I(f,\mathcal{P}_n)-I|\le2^{-n}\). Given \(\varepsilon>0\), take \(n\) with \(2^{-n+1}<\varepsilon\) and \(\delta:=\delta_n\): any \(\mathcal{P}\) subordinate to \(\delta_n\) is subordinate to \(\gamma_n\), so by (1),
\begin{equation*} |I(f,\mathcal{P}) - I| \le |I(f,\mathcal{P}) - I(f,\mathcal{P}_n)| + |I(f,\mathcal{P}_n) - I| < 2^{-n+1} < \varepsilon , \end{equation*}
i.e. \(f\) is Henstock–Kurzweil integrable with integral \(I\) (Definition 5.7.3). The McShane case is verbatim the same with free tagged partitions, the only external ingredient, Lemma 5.7.2, applying unchanged since every tagged partition is a free one and subordination is the same condition.
(\(\circ\)) Prove Proposition 5.7.6.
(For convenience we recall the statement. (i) If \(f\) is Henstock–Kurzweil integrable on \([a,b]\), then \(f\) is integrable in the same sense on every interval \([\alpha,\beta] \subset [a,b]\). (ii) If \(f\) is Henstock–Kurzweil integrable on \([a,c]\) and \([c,b]\) for some point \(c \in (a,b)\), then \(f\) is integrable in the same sense on \([a,b]\) and
\begin{equation*} (HK)\int_a^b f = (HK)\int_a^c f + (HK)\int_c^b f . \end{equation*}
(iii) The set \(\mathcal{L}_{HK}[a,b]\) of all functions on \([a,b]\) that are Henstock–Kurzweil integrable is a linear space, on which the Henstock–Kurzweil integral is linear. (iv) If \(f, g \in \mathcal{L}_{HK}[a,b]\) and \(f \le g\) a.e., then \((HK)\int_a^b f \le (HK)\int_a^b g\).)
Each part is the Cauchy criterion of Lemma 5.7.4 (Exercise 5.8.101) or Definition 5.7.3 applied to a suitably built gauge; Lemma 5.7.2 supplies partitions subordinate to any positive \(\delta\), and a collection subordinate to \(\delta_1\le\delta_2\) is subordinate to \(\delta_2\).
(i) Let \([\alpha,\beta]\subset[a,b]\), \(\alpha<\beta\), and \(\varepsilon>0\), and take \(\delta>0\) on \([a,b]\) with
\begin{equation*} |I(f,\mathcal{Q}_1) - I(f,\mathcal{Q}_2)| < \varepsilon \quad\text{for }\mathcal{Q}_1,\mathcal{Q}_2 \text{ partitions of } [a,b] \text{ subordinate to } \delta \tag{1} \end{equation*}
(Lemma 5.7.4). Fix once and for all tagged partitions \(\mathcal{R}_1\) of \([a,\alpha]\) and \(\mathcal{R}_2\) of \([\beta,b]\) subordinate to \(\delta\) (Lemma 5.7.2 on those intervals; subordination does not refer to the ambient interval), empty if \(\alpha=a\) resp. \(\beta=b\). For any tagged partition \(\mathcal{P}\) of \([\alpha,\beta]\) subordinate to \(\delta\), the union \(\mathcal{R}_1\cup\mathcal{P}\cup\mathcal{R}_2\) is a tagged partition of \([a,b]\) subordinate to \(\delta\) with sum \(I(f,\mathcal{R}_1)+I(f,\mathcal{P})+I(f,\mathcal{R}_2)\); so for two such \(\mathcal{P}_1,\mathcal{P}_2\) the outer contributions cancel and (1) gives \(|I(f,\mathcal{P}_1)-I(f,\mathcal{P}_2)|<\varepsilon\). By Lemma 5.7.4, \(f\) is integrable on \([\alpha,\beta]\).
(ii) Let the integrals over \([a,c]\), \([c,b]\) be \(I_1,I_2\), let \(\varepsilon>0\), and pick \(\delta_1,\delta_2\) positive on \([a,c]\), \([c,b]\) giving \(\varepsilon/2\)-approximation of \(I_1,I_2\) (Definition 5.7.3). Put
\begin{equation*} \delta(x) := \begin{cases} \min\{\delta_1(x), (c-x)/2\}, & a \le x < c,\\[2pt] \min\{\delta_1( c), \delta_2( c)\}, & x = c, \\[2pt] \min\{\delta_2(x), (x-c)/2\}, & c < x \le b. \end{cases} \end{equation*}
and let \(\mathcal{P}=\{(x_i,[c_i,d_i])\}\) be a tagged partition of \([a,b]\) subordinate to \(\delta\). If \(x_i<c\) then \(x_i+\delta(x_i)\le(x_i+c)/2<c\), so \([c_i,d_i]\subset[a,c)\), and symmetrically \([c_i,d_i]\subset(c,b]\) when \(x_i>c\); hence every interval of \(\mathcal{P}\) containing \(c\) (there are one or two) carries the tag \(c\). Split each such \((c,[u,v])\) into \((c,[u,c])\) and \((c,[c,v])\), discarding degenerate pieces: this preserves subordination and, since \(f( c)(v-u)=f( c)(c-u)+f( c)(v-c)\), the Riemann sum. The collection now splits into
\begin{equation*} \mathcal{P}_1 := \{ (x_i,[c_i,d_i]) \in \mathcal{P} : [c_i,d_i] \subset [a,c] \}, \qquad \mathcal{P}_2 := \{ (x_i,[c_i,d_i]) \in \mathcal{P} : [c_i,d_i] \subset [c,b] \}, \end{equation*}
tagged partitions of \([a,c]\) and \([c,b]\) with \(I(f,\mathcal{P})=I(f,\mathcal{P}_1)+I(f,\mathcal{P}_2)\), and \(\mathcal{P}_1\) is subordinate to \(\delta_1\) (as \(\delta\le\delta_1\) on \([a,c)\) and \(\delta( c)\le\delta_1( c)\)), \(\mathcal{P}_2\) to \(\delta_2\). Hence
\begin{equation*} \big| I(f,\mathcal{P}) - (I_1 + I_2) \big| \le |I(f,\mathcal{P}_1) - I_1| + |I(f,\mathcal{P}_2) - I_2| < \varepsilon , \end{equation*}
so \(f\) is integrable on \([a,b]\) with integral \(I_1+I_2\).
(iii) For \(s,t\in\mathbb{R}\) not both \(0\) and \(\varepsilon^{\prime}:=\varepsilon/(|s|+|t|)\), take \(\delta:=\min\{\delta_f,\delta_g\}\) for gauges achieving \(\varepsilon^{\prime}\)-approximation of \(I_f,I_g\). Since \(I(sf+tg,\mathcal{P})=sI(f,\mathcal{P})+tI(g,\mathcal{P})\) termwise,
\begin{equation*} \big| I(sf+tg,\mathcal{P}) - (s I_f + t I_g) \big| \le |s|\,|I(f,\mathcal{P}) - I_f| + |t|\,|I(g,\mathcal{P}) - I_g| < \varepsilon \end{equation*}
for every \(\mathcal{P}\) subordinate to \(\delta\); thus \(\mathcal{L}_{HK}[a,b]\) is linear and so is the integral.
(iv) If \(u\ge0\) everywhere and \(u\in\mathcal{L}_{HK}[a,b]\) with integral \(J\), then for each \(\varepsilon>0\) any \(\mathcal{P}\) subordinate to the corresponding \(\delta\) (Lemma 5.7.2) has \(I(u,\mathcal{P})\ge0\), so \(J>-\varepsilon\) and \(J\ge0\). Now let \(f\le g\) a.e. and \(h:=g-f\), integrable with integral \(\int g-\int f\) by (iii). The set \(N=\{h<0\}\) lies in a null set, hence is measurable with \(\lambda(N)=0\) by completeness, so \(hI_N\) vanishes a.e. and is integrable with integral \(0\) by Example 5.7.5. By (iii), \(u:=h-hI_N\) is integrable with the same integral as \(h\), and \(u\ge0\) everywhere (it is \(0\) on \(N\) and \(h\ge0\) off \(N\)); therefore \((HK)\int_a^bg-(HK)\int_a^bf\ge0\).
(\(\circ\)) Prove Lemma 5.7.9.
(For convenience we recall the statement. Suppose that a function \(f\) on \([a,b]\) is Henstock–Kurzweil integrable and \(F\) is defined by (5.7.1), i.e. \(F(x) = (HK)\int_a^x f\). Let \(\varepsilon > 0\) and let \(\delta\) be a positive function such that \(|I(f,\mathcal{P}) - F(b)| < \varepsilon\) for every tagged partition \(\mathcal{P}\) of \([a,b]\) subordinate to \(\delta\). Then, for every finite collection \(\mathcal{P}_0 = \{(x_i,[c_i,d_i]),\ i = 1,\ldots,n\}\) of non-overlapping tagged intervals subordinate to \(\delta\), one has
\begin{equation*} \Big| I(f,\mathcal{P}_0) - \sum_{i=1}^n [F(d_i) - F(c_i)] \Big| \le \varepsilon , \qquad \sum_{i=1}^n \big| f(x_i)(d_i - c_i) - [F(d_i) - F(c_i)] \big| \le 2\varepsilon . \end{equation*}
Fill the gaps of \(\mathcal{P}_0\) by partitions whose sums nearly equal the corresponding increments of \(F\). By Proposition 5.7.6 (Exercise 5.8.102), \(f\) is integrable on every subinterval and \(F(\beta)-F(\alpha)=(HK)\int_\alpha^\beta f\) for \(a\le\alpha\le\beta\le b\); hence, \(F(a)=0\), for non-overlapping intervals covering \([a,b]\) the increments of \(F\) telescope to \(F(b)\).
Relabel \(\mathcal{P}_0=\{(x_i,[c_i,d_i])\}_{i\le n}\) so that \(a\le c_1<d_1\le\cdots\le c_n<d_n\le b\), and let \(K_1=[a_1,b_1],\dots,K_m=[a_m,b_m]\) be the non-degenerate intervals among \([a,c_1],[d_1,c_2],\dots,[d_n,b]\). The \([c_i,d_i]\) and the \(K_j\) are non-overlapping with union \([a,b]\), so
\begin{equation*} \sum_{i=1}^{n} [F(d_i) - F(c_i)] + \sum_{j=1}^{m} [F(b_j) - F(a_j)] = F(b). \tag{3} \end{equation*}
Fix \(\varepsilon_1>0\). Since the integral of \(f\) over \(K_j\) is \(F(b_j)-F(a_j)\), there is a gauge on \(K_j\) whose subordinate partitions approximate it within \(\varepsilon_1/m\); taking its minimum with \(\delta|_{K_j}\) and invoking Lemma 5.7.2 gives a tagged partition \(\mathcal{P}_j\) of \(K_j\), subordinate to \(\delta\), with
\begin{equation*} \big| F(b_j) - F(a_j) - I(f,\mathcal{P}_j) \big| < \varepsilon_1/m . \tag{4} \end{equation*}
Then \(\mathcal{P}:=\mathcal{P}_0\cup\mathcal{P}_1\cup\cdots\cup\mathcal{P}_m\) is a tagged partition of \([a,b]\) subordinate to \(\delta\), so \(|I(f,\mathcal{P})-F(b)|<\varepsilon\) by hypothesis, and since \(I(f,\mathcal{P})=I(f,\mathcal{P}_0)+\sum_jI(f,\mathcal{P}_j)\), (3) and (4) give
\begin{equation*} \begin{aligned} \Big| I(f,\mathcal{P}_0) - \sum_{i=1}^{n} [F(d_i) - F(c_i)] \Big| &= \Big| I(f,\mathcal{P}) - \sum_{j=1}^{m} I(f,\mathcal{P}_j) - F(b) + \sum_{j=1}^{m}[F(b_j) - F(a_j)] \Big| \\ &\le \big| I(f,\mathcal{P}) - F(b) \big| + \sum_{j=1}^{m} \big| F(b_j) - F(a_j) - I(f,\mathcal{P}_j) \big| \\ &< \varepsilon + m \cdot \frac{\varepsilon_1}{m} = \varepsilon + \varepsilon_1 . \end{aligned} \end{equation*}
Letting \(\varepsilon_1\to0\) gives the first inequality
\begin{equation*} \Big| I(f,\mathcal{P}_0) - \sum_{i=1}^{n} [F(d_i) - F(c_i)] \Big| \le \varepsilon . \tag{6} \end{equation*}
For the second, put \(\Delta_i:=f(x_i)(d_i-c_i)-[F(d_i)-F(c_i)]\) and split \(\{1,\dots,n\}\) into \(S_{+}=\{\Delta_i\ge0\}\) and \(S_{-}=\{\Delta_i<0\}\). Any subcollection of \(\mathcal{P}_0\) is again a finite non-overlapping collection subordinate to \(\delta\), so (6) applies to both, whence
\begin{equation*} \sum_{i=1}^{n} |\Delta_i| = \sum_{i \in S_{+}} \Delta_i - \sum_{i \in S_{-}} \Delta_i \le 2\varepsilon . \end{equation*}
Prove Proposition 5.8.34.
(For convenience we recall the statement. Let \(f\) be a Lipschitzian function on \(\mathbb{R}^n\) such that \(|\nabla f(x)| \ge c > 0\) a.e. If a function \(g\) is integrable on \(\mathbb{R}^n\), then
\begin{equation*} \int_{\{f > t\}} g(x)\, dx = \int_t^{\infty} \left( \int_{\{f = s\}} \frac{g(y)}{|\nabla f(y)|}\, H^{n-1}(dy) \right) ds \end{equation*}
for all \(t \in \mathbb{R}\).)
Apply the coarea formula of Theorem 5.8.29 with \(k=1\) to \(u:=gI_{\{f>t\}}/|\nabla f|\). Let \(N_0\) be the null set where \(f\) fails to be differentiable (Rademacher) or \(|\nabla f|<c\), and set \(|\nabla f|:=0\) there; for \(k=1\) the Jacobian is \(|J_f|=|\nabla f|\), so (5.8.15) reads: for measurable \(A\), the map \(s\mapsto H^{n-1}(A\cap f^{-1}(s))\) is measurable and
\begin{equation*} \int_A |\nabla f(x)|\, dx = \int_{\mathbb{R}} H^{n-1}\big( A \cap f^{-1}(s) \big)\, ds . \tag{1} \end{equation*}
(For \(n=1\) this is (5.8.14) with \(n=k=1\), where \(\mathrm{Card}=H^0\).)
(i) Null sets meet a.e. level set in an \(H^{n-1}\)-null set: for \(\lambda_n(N)=0\), (1) with \(A=N\) gives \(\int_{\mathbb{R}}H^{n-1}(N\cap f^{-1}(s))\,ds=0\), so
\begin{equation*} H^{n-1}\big( N \cap f^{-1}(s) \big) = 0 \qquad \text{for a.e. } s \in \mathbb{R}. \tag{2} \end{equation*}
With \(N=N_0\) this makes \(|\nabla f|\ge c\) hold \(H^{n-1}\)-a.e. on \(\{f=s\}\) for a.e. \(s\), so the inner integral is meaningful; and applied to \(N=\{u_1\ne u_2\}\) it shows the level-set integrals below do not depend on the choice of version (take a Borel version, which is \(H^{n-1}\)-measurable on the Borel set \(f^{-1}(s)\)).
(ii) Coarea with an integrand: for measurable \(u\ge0\), \(s\mapsto\int_{f^{-1}(s)}u\,dH^{n-1}\) is measurable and
\begin{equation*} \int_{\mathbb{R}^n} u(x)\, |\nabla f(x)|\, dx = \int_{\mathbb{R}} \left( \int_{f^{-1}(s)} u\, dH^{n-1} \right) ds . \tag{3} \end{equation*}
For \(u=I_A\) this is (1) (with the measurability in \(s\) part of Theorem 5.8.29), hence by linearity for nonnegative simple \(u\), hence for all \(u\ge0\) by monotone convergence applied on \(f^{-1}(s)\) and then in \(s\) and \(x\). If \(u\) is real-valued with \(u|\nabla f|\) integrable, apply this to \(u^{\pm}\) and subtract (both sides finite).
(iii) Fix \(t\) and put \(u:=gI_{\{f>t\}}/|\nabla f|\) off \(N_0\) and \(u:=0\) on \(N_0\); then \(|u|\le|g|/c\), so \(u\) is integrable, and \(u|\nabla f|=gI_{\{f>t\}}\) a.e. Hence (3) gives
\begin{equation*} \int_{\{f>t\}} g(x)\, dx = \int_{\mathbb{R}^n} u(x)\,|\nabla f(x)|\, dx = \int_{\mathbb{R}} \left( \int_{f^{-1}(s)} u\, dH^{n-1} \right) ds . \tag{4} \end{equation*}
For \(s\) outside the exceptional null set of (2) with \(N=N_0\), the indicator \(I_{\{f>t\}}\) equals \(1\) on \(f^{-1}(s)\) if \(s>t\) and \(0\) if \(s\le t\), while \(u=g/|\nabla f|\) on \(f^{-1}(s)\setminus N_0\) and \(H^{n-1}(N_0\cap f^{-1}(s))=0\). So the inner integral is \(\int_{\{f=s\}}g/|\nabla f|\,dH^{n-1}\) for \(s>t\) and \(0\) otherwise, and (4) becomes
\begin{equation*} \int_{\{f>t\}} g(x)\, dx = \int_t^{\infty} \left( \int_{\{f = s\}} \frac{g(y)}{|\nabla f(y)|}\, H^{n-1}(dy) \right) ds . \end{equation*}
(\(\circ\)) Let \(F\) be a closed set in \(\mathbb{R}^n\). Prove that if \(x\) is a density point of \(F\), then one has
\begin{equation*} \lim_{|y| \to 0} \frac{\mathrm{dist}\,(x+y, F)}{|y|} = 0 . \end{equation*}
A point far from \(F\) carries an \(F\)-free ball whose relative volume in \(B(x,r)\) is bounded below, contradicting density. Write \(\omega_n:=\lambda_n(B(0,1))\). If the limit fails, there are \(\varepsilon\in(0,1]\) and \(y_k\ne0\), \(|y_k|\to0\), with \(\mathrm{dist}\,(x+y_k,F)\ge\varepsilon|y_k|\); then the open ball \(B^{\circ}(z_k,\rho_k)\), \(z_k:=x+y_k\), \(\rho_k:=\varepsilon|y_k|\), misses \(F\), and it lies in \(B(x,r_k)\) with \(r_k:=(1+\varepsilon)|y_k|\to0\), since \(|w-x|\le|w-z_k|+|y_k|<r_k\) there. Hence \(\lambda_n(B(x,r_k)\setminus F)\ge\omega_n\varepsilon^n|y_k|^n\) and
\begin{equation*} \frac{\lambda_n\big( F \cap B(x,r_k) \big)}{\lambda_n\big( B(x,r_k) \big)} \le \frac{(1+\varepsilon)^n - \varepsilon^n}{(1+\varepsilon)^n} = 1 - \Bigl( \frac{\varepsilon}{1+\varepsilon} \Bigr)^{\! n} < 1 \end{equation*}
for every \(k\), contradicting the fact that \(x\) is a density point of \(F\).
Let \(Z \subset \mathbb{R}^1\) be a set of measure zero. Show that there exists a measurable set \(E\) that has density at no point \(z \in Z\), i.e., as \(r \to 0\), there is no limit of the ratio \(\lambda_n\big(E \cap K(z,r)\big) / \lambda_n\big(K(z,r)\big)\), where \(K(z,r)\) is the ball of radius \(r\) centered at \(z\).
Take \(E:=\bigcup_{j\ge0}(U_{2j}\setminus U_{2j+1})\) for a sequence of open sets \(U_0\supset U_1\supset\cdots\supset Z\), each so thin in the previous one that at the scale \(d_{U_m}(z)\) it occupies a fraction \(\le1/4\); then \(E\) nearly fills \(K(z,d_{U_{2k}}(z))\) and nearly misses \(K(z,d_{U_{2k+1}}(z))\). Here \(n=1\), \(K(z,r)=[z-r,z+r]\), \(\lambda:=\lambda_1\), and \(d_U(z):=\mathrm{dist}\,(z,\mathbb{R}^1\setminus U)\); assume \(Z\ne\varnothing\).
Lemma. Let \(U\subset\mathbb{R}^1\) be nonempty open with \(\lambda(U)<\infty\), let \(Z\subset U\) with \(\lambda(Z)=0\), and let \(\eta\in(0,1)\). Then there is an open \(V\) with
\begin{equation*} Z \subset V \subset U, \qquad \lambda(V) \le \eta\, \lambda(U), \qquad \lambda\big( V \cap K(z, d_U(z)) \big) \le \eta\, d_U(z) \ \ \text{for } z \in Z . \end{equation*}
Proof. Write \(U\) as the at most countable disjoint union of its components, bounded open intervals. Let \(I=(\alpha,\beta)\) be one, \(\ell:=\beta-\alpha\); since \(\alpha,\beta\notin U\) by maximality, for \(z\in I\)
\begin{equation*} d_U(z) = \min\{ z - \alpha, \beta - z \} =: d(z) \in (0, \ell/2], \qquad K\big(z, d(z)\big) \subset [\alpha,\beta] = \overline{I} . \tag{1} \end{equation*}
and \(\overline{I}\) meets no other component. Since \(d\) is \(1\)-Lipschitz on \(I\) and continuous, the sets
\begin{equation*} A_m := \{ x \in I : 2^{-m-1}\ell < d(x) \le 2^{-m}\ell \}, \qquad G_m := \{ x \in I : 2^{-m-2}\ell < d(x) < 2^{-m+1}\ell \} \end{equation*}
satisfy \(A_m\subset G_m\), \(G_m\) open, and \(\bigcup_{m\ge1}A_m=I\). As \(\lambda(Z\cap A_m)=0\), choose open \(O_m\supset Z\cap A_m\) with \(\lambda(O_m)<\eta\ell2^{-m-4}\) and put \(W_m:=O_m\cap G_m\) and \(V_I:=\bigcup_{m\ge1}W_m\); then \(V_I\) is open, \(Z\cap I\subset V_I\subset I\), and \(\lambda(V_I)\le\eta\ell2^{-4}\le\eta\lambda(I)\). For \(z\in Z\cap I\) pick \(n\) with \(z\in A_n\). Every \(x\in K(z,d(z))\cap I\) has \(d(x)\le d(z)+|x-z|\le2d(z)\le2^{-n+1}\ell\), whereas \(x\in W_m\subset G_m\) forces \(d(x)>2^{-m-2}\ell\); hence \(W_m\) meets \(K(z,d(z))\) only for \(m\ge n-2\) and
\begin{equation*} \lambda\big( V_I \cap K(z,d(z)) \big) \le \sum_{m \ge n-2} \eta\, \ell\, 2^{-m-4} = \eta\, \ell\, 2^{-n-1} < \eta\, d(z), \end{equation*}
the last step by \(d(z)>2^{-n-1}\ell\). Put \(V:=\bigcup_IV_I\): it is open, \(Z\subset V\subset U\), \(\lambda(V)\le\eta\lambda(U)\), and for \(z\in Z\cap I\) the ball \(K(z,d_U(z))\subset\overline I\) meets no \(V_{I^{\prime}}\), \(I^{\prime}\ne I\), by (1), so \(\lambda(V\cap K(z,d_U(z)))<\eta\,d_U(z)\). \(\square\)
Start from open \(U_0\supset Z\) with \(\lambda(U_0)<1\) and iterate the Lemma with \(\eta=1/4\), \(U_{m+1}:=V\), obtaining
\begin{equation*} U_0 \supset U_1 \supset U_2 \supset \cdots, \qquad Z \subset U_m \ \text{ for all } m, \end{equation*}
with
\begin{equation*} \lambda(U_{m+1}) \le \tfrac14 \lambda(U_m), \quad \text{hence} \quad \lambda(U_m) \le 4^{-m}, \tag{2} \end{equation*}
and, writing \(d_m(z) := d_{U_m}(z) = \mathrm{dist}\,(z, \mathbb{R}^1 \setminus U_m) > 0\) for \(z \in Z\),
\begin{equation*} \lambda\big( U_{m+1} \cap K(z, d_m(z)) \big) \le \tfrac14\, d_m(z) \qquad \text{for every } z \in Z, \ m \ge 0 . \tag{3} \end{equation*}
By (1) applied to \(U_m\), for \(z\in Z\) the ball \(K(z,d_m(z))\) lies in the closure of the component \(I\ni z\) of \(U_m\), so \(2d_m(z)\le\lambda(I)\le4^{-m}\), whence
\begin{equation*} d_m(z) \longrightarrow 0 \quad (m \to \infty) \qquad \text{uniformly in } z \in Z , \tag{4} \end{equation*}
and \(K(z,d_m(z))\setminus U_m\subset\overline I\setminus I\) is at most two points, so
\begin{equation*} \lambda\big( K(z,d_m(z)) \setminus U_m \big) = 0 . \tag{5} \end{equation*}
For the Borel set \(E:=\bigcup_{j\ge0}(U_{2j}\setminus U_{2j+1})\) we have \(U_{2k}\setminus U_{2k+1}\subset E\) and
\begin{equation*} E \cap U_{2k+1} \subset U_{2k+2} \qquad (k \ge 0) , \tag{7} \end{equation*}
since for \(j\le k\) the set \(U_{2j}\setminus U_{2j+1}\) misses \(U_{2k+1}\subset U_{2j+1}\), while for \(j\ge k+1\) it lies in \(U_{2j}\subset U_{2k+2}\).
(a) Fix \(z\in Z\), \(k\ge0\) and put \(\sigma_k:=d_{2k}(z)>0\). By (5) with \(m=2k\), \(\lambda(K(z,\sigma_k)\setminus U_{2k})=0\), so
\begin{equation*} \begin{aligned} \lambda\big( E \cap K(z,\sigma_k) \big) &\ge \lambda\big( (U_{2k} \setminus U_{2k+1}) \cap K(z,\sigma_k) \big) \\ &\ge \lambda\big( K(z,\sigma_k) \big) - \lambda\big( K(z,\sigma_k) \setminus U_{2k} \big) - \lambda\big( U_{2k+1} \cap K(z,\sigma_k) \big) \\ &\ge 2\sigma_k - 0 - \tfrac14 \sigma_k = \tfrac74 \sigma_k , \end{aligned} \end{equation*}
the last inequality by (3) with \(m=2k\); hence that ratio is \(\ge(7\sigma_k/4)/(2\sigma_k)=7/8\).
(b) Put \(\rho_k:=d_{2k+1}(z)>0\). By (5) with \(m=2k+1\), then (7), then (3) with \(m=2k+1\),
\begin{equation*} \begin{aligned} \lambda\big( E \cap K(z,\rho_k) \big) &\le \lambda\big( E \cap U_{2k+1} \cap K(z,\rho_k) \big)\\ &\le \lambda\big( U_{2k+2} \cap K(z,\rho_k) \big) \le \tfrac14 \rho_k , \end{aligned} \end{equation*}
so that ratio is at most \((\rho_k/4)/(2\rho_k)=1/8\). Since \(\sigma_k,\rho_k>0\) tend to \(0\) by (4),
\begin{equation*} \liminf_{r \to 0} \frac{\lambda\big(E \cap K(z,r)\big)}{\lambda\big(K(z,r)\big)} \le \frac18 < \frac78 \le \limsup_{r \to 0} \frac{\lambda\big(E \cap K(z,r)\big)}{\lambda\big(K(z,r)\big)} , \end{equation*}
so \(E\) has density at no point of \(Z\).
Exercises 5.8.107–5.8.113
Let \(A\) be a convex compact set of positive measure in \(\mathbb{R}^n\) and let \(B\) be the closed unit ball. (i) Prove that the limit \(\lim_{r\to 0} r^{-1}[\lambda_n(A+rB)-\lambda_n(A)]\) exists and equals the surface measure of the boundary of the set \(A\). (ii) Prove that the same limit equals the mixed volume \(v_{n-1,1}(A,B)\).
The limit is \(H^{n-1}(\partial A)\), the surface measure of 5.8(xi), and it equals \(n\,v_{n-1,1}(A,B)\); the printed statement of (ii) omits the factor \(n\). Translate so that \(0\in\operatorname{int}A\), whence \(\rho B\subset A\subset RB\) for some \(0<\rho\le R\); also \(\lambda_n(\partial A)=0\), since \(cA\subset\operatorname{int}A\) for \(0<c<1\) gives \(c^n\lambda_n(A)\le\lambda_n(\operatorname{int}A)\). Throughout we use \(H^s(f(E))\le\Lambda^sH^s(E)\) for \(\Lambda\)-Lipschitz \(f\) (Lemma 3.10.12) and \(H^s(cE)=c^sH^s(E)\), both from 3.10(iii).
Monotonicity: \(K\subset L\) convex compact implies \(H^{n-1}(\partial K)\le H^{n-1}(\partial L)\). Let \(p_K\) be the (1-Lipschitz) nearest-point projection onto \(K\). Given \(x\in\partial K\) with unit outer normal \(\nu\) of a supporting hyperplane there, for \(z\in K\) and \(t\ge0\), \[ |x+t\nu-z|^2=t^2+2t\langle \nu,x-z\rangle+|x-z|^2\ge t^2=|x+t\nu-x|^2, \] so \(p_K(x+t\nu)=x\); the ray leaves the compact \(L\), hence meets \(\partial L\), giving \(\partial K\subset p_K(\partial L)\) and the claim by Lemma 3.10.12. In particular \(H^{n-1}(\partial A)\le R^{n-1}H^{n-1}(\partial B)<\infty\).
Put \(A_s:=A+sB=\{x:\operatorname{dist}(x,A)\le s\}\) (equality since \(A\) is compact and \(B\) the closed unit ball) and \[ S(s):=H^{n-1}(\partial A_s). \] Monotonicity gives \(S(s)\ge H^{n-1}(\partial A)\), while \(sB\subset(s/\rho)A\) and convexity give \(A_s\subset(1+s/\rho)A\), so by monotonicity and scaling \[ H^{n-1}(\partial A)\le S(s)\le \Bigl(1+\frac{s}{\rho}\Bigr)^{n-1}H^{n-1}(\partial A), \] whence \(S(s)\to H^{n-1}(\partial A)\) as \(s\to0^{+}\), and \(S\) is nondecreasing, hence measurable. Let \[ u(x):=\operatorname{dist}(x,A)-\operatorname{dist}(x,\mathbb{R}^n\setminus A) \] be the signed distance, \(1\)-Lipschitz (each summand is, and one vanishes at every point), \(\le0\) on \(A\), \(>0\) off \(A\).
Then \(|\nabla u|=1\) a.e. For \(x\notin A\) put \(\nu=(x-p_A(x))/|x-p_A(x)|\), an outer normal of a supporting hyperplane of \(A\) at \(p_A(x)\) (expand \(|x-p_A(x)-s(z-p_A(x))|^2\ge|x-p_A(x)|^2\) and let \(s\downarrow0\)); the computation above then gives \(p_A(x+t\nu)=p_A(x)\) and so \(u(x+t\nu)=u(x)+t\) for \(t\ge-u(x)\), a two-sided neighbourhood of \(0\). Near \(x\), \(u=\operatorname{dist}(\cdot,A)\) is convex, so \(u^{\prime}(x;\pm\nu)=\pm1\), forcing \(\langle q,\nu\rangle=1\) for every subgradient \(q\in\partial u(x)\); as \(|q|\le1\), Cauchy–Bunyakowsky gives \(q=\nu\), and a convex function with singleton subdifferential is differentiable there, with \(|\nabla u(x)|=1\). If \(x\in\operatorname{int}A\) with \(d:=\operatorname{dist}(x,\partial A)\) and \(\nu=(z-x)/d\) for a nearest \(z\in\partial A\), then \(u(x+t\nu)=u(x)+t\) for \(0\le t<d\) (upper bound from \(|x+t\nu-z|=d-t\), lower from \(1\)-Lipschitzness), so at any point of differentiability (Rademacher) \(\langle\nabla u,\nu\rangle=1\) and hence \(|\nabla u|=1\). Since \(\lambda_n(\partial A)=0\), this holds a.e.
Moreover \(\{u\le s\}=A_s\) and \(\{u=s\}=\partial A_s\) for \(s>0\): \(\{u<s\}\) is open and inside \(A_s\), while \(u(x)=s\) gives \(u(x+\varepsilon\nu)>s\), so \(x\notin\operatorname{int}A_s\). Now Proposition 5.8.34 applies to \(f=u\) (Lipschitz with \(|\nabla u|=1\) a.e.), \(t=0\) and \(g=I_{\{0<u\le r\}}\):
\begin{equation*} \begin{aligned} \lambda_n(A_r)-\lambda_n(A) &=\lambda_n\bigl(\{0<u\le r\}\bigr)\\ &=\int_0^{\infty}\!\!\int_{\{u=s\}}\frac{g(y)}{|\nabla u(y)|}\,H^{n-1}(dy)\,ds =\int_0^r H^{n-1}(\partial A_s)\,ds, \end{aligned} \end{equation*}
using that \(u\) is differentiable with \(|\nabla u|=1\) everywhere on \(\{u=s\}\), \(s>0\). Thus \[ \lambda_n(A+rB)-\lambda_n(A)=\int_0^r S(s)\,ds,\qquad r>0, \tag{2} \] and since \(S(s)\to H^{n-1}(\partial A)\) as \(s\to0^{+}\) and \(S\) is bounded near \(0\), the averages \(r^{-1}\int_0^rS\) converge to the same limit: \[ \lim_{r\to0}\frac{\lambda_n(A+rB)-\lambda_n(A)}{r}=H^{n-1}(\partial A), \] which proves (i).
For (ii), by 3.10(vii) the function \((\alpha,\beta)\mapsto\lambda_n(\alpha A+\beta B)\) is the polynomial \[ \lambda_n(\alpha A+\beta B)=\sum_{k=0}^n \alpha^{n-k}\beta^k\binom{n}{k}v_{n-k,k}(A,B), \] with \(v_{n,0}(A,B)=\lambda_n(A)\); taking \(\alpha=1\), \(\beta=r\) and dividing by \(r\), \[ \lim_{r\to0}\frac{\lambda_n(A+rB)-\lambda_n(A)}{r}=n\,v_{n-1,1}(A,B), \] so \(H^{n-1}(\partial A)=n\,v_{n-1,1}(A,B)\). The factor \(n\) is forced by the normalization of 3.10(vii): for \(A=B\) the limit is \(n\lambda_n(B)=H^{n-1}(\partial B)\) while \(v_{n-1,1}(B,B)=\lambda_n(B)\).
Deduce Theorem 5.8.22 from Theorem 5.8.21.
Average the doubling inequality of Theorem 5.8.21(iii) over the finite set of centres. The one point to watch is uniformity: the constant \(C(n)\) of Theorem 5.8.22 must not depend on the family of balls. It does not, because the quantitative content of Theorem 5.8.21 is that \(X\in\Phi_\gamma\) with constant \(N_0\) yields a probability measure obeying \(D_\gamma\) with a constant depending only on \(N_0\) and \(\gamma\), and every \(X\subset\mathbb{R}^n\) lies in \(\Phi_n\) with the same \(N_0=3^n\): if \(y_1,\dots,y_M\in B(x,kR)\) have mutual distances \(\ge R\), the disjoint open balls \(B(y_i,R/2)\subset B(x,kR+R/2)\) give \[ M\,\omega_n(R/2)^n\le \omega_n(kR+R/2)^n,\qquad M\le (2k+1)^n\le 3^nk^n. \] So the constant \(C_0=C_0(n)\) in \(D_n\) below depends only on \(n\).
Given the family, let \(X=\{x_1,\dots,x_N\}\) with the induced metric, a nonempty compact subset of \(\mathbb{R}^n\), and let \(\mu\) be a Borel probability measure on \(X\) satisfying \(D_n\) with constant \(C_0(n)\) (Theorem 5.8.21(iii)); put \(C(n):=2^nC_0(n)\) and \(m_i:=\mu(\{x_i\})\ge0\), \(\sum_im_i=1\). Balls of \(X\) are traces \(B(x,r)\cap X\), so \(D_n\) with \(k=2\) at \(x_j\), \(R_j\) reads \[ \mu\bigl(B(x_j,2R_j)\cap X\bigr)\le C_0 2^{n}\mu\bigl(B(x_j,R_j)\cap X\bigr)=C(n)\,\mu\bigl(B(x_j,R_j)\cap X\bigr) \] for every \(j\). Summing over \(j\) and interchanging the finite sums,
\begin{equation*} \begin{aligned} \sum_{i=1}^N m_iN_i^{\prime} &=\sum_{i=1}^N m_i\sum_{j=1}^N I_{B(x_j,2R_j)}(x_i) =\sum_{j=1}^N\mu\bigl(B(x_j,2R_j)\cap X\bigr)\\ &\le C(n)\sum_{j=1}^N\mu\bigl(B(x_j,R_j)\cap X\bigr)\\ &=C(n)\sum_{i=1}^N m_i\sum_{j=1}^N I_{B(x_j,R_j)}(x_i) =C(n)\sum_{i=1}^N m_iN_i. \end{aligned} \end{equation*}
Therefore \(\sum_im_i[N_i^{\prime}-C(n)N_i]\le0\); since the \(m_i\ge0\) sum to \(1\), at least one is positive, and the bracket cannot be positive at every such index. Hence \(N_{i_0}^{\prime}\le C(n)N_{i_0}\) for some \(i_0\), with \(C(n)=2^nC_0(n)\).
(\(\circ\)) (Burstin [152]) Let \(f\) be an a.e. finite Lebesgue measurable function with periods \(\pi_n\to0\) (or, more generally, \(f(x+\pi_n)=f(x)\) a.e.). Show that \(f\) coincides a.e. with some constant.
Mollify: the a.e. periodicity of \(f\) becomes exact periodicity of \(f_\varepsilon\), forcing \(f_\varepsilon\) constant, and \(f_\varepsilon\to f\) locally in \(L^1\). Write \(N_n:=\{u:f(u+\pi_n)\ne f(u)\}\), a null set.
Replacing \(f\) by \(\arctan f\) (with the values \(\pm\pi/2\) on the null set where \(f=\pm\infty\)) we may assume \(|f|\le K\): \(\arctan\) is injective on \([-\infty,+\infty]\) and \(f\) is a.e. finite, so a constant value \(c\in(-\pi/2,\pi/2)\) for the transform returns \(f=\tan c\) a.e. Fix \(g\in C^\infty\), \(g\ge0\), supported in \([0,1]\), \(\int g=1\), and put \[ f_\varepsilon(x):=\int_0^1 f(x-\varepsilon y)g(y)\,dy=\int_{\mathbb{R}}f(t)g_\varepsilon(x-t)\,dt,\qquad g_\varepsilon(u)=\varepsilon^{-1}g(u/\varepsilon), \] which is smooth (the difference quotients of \(g_\varepsilon\) converge uniformly and are bounded, while \(f\) is bounded on the relevant compact set, so dominated convergence licenses differentiation under the integral).
Each \(\pi_n\) is an exact period of \(f_\varepsilon\): the set \(\{y:x-\varepsilon y\in N_n\}=\varepsilon^{-1}(x-N_n)\) is null, so \(f(x+\pi_n-\varepsilon y)=f(x-\varepsilon y)\) for a.e. \(y\) and \(f_\varepsilon(x+\pi_n)=f_\varepsilon(x)\) for every \(x\). Hence \[ f_\varepsilon^{\prime}(x)=\lim_{n\to\infty}\frac{f_\varepsilon(x+\pi_n)-f_\varepsilon(x)}{\pi_n}=0 \] (the numerators vanish and \(0\ne\pi_n\to0\)), so \(f_\varepsilon\equiv c_\varepsilon\) with \(|c_\varepsilon|\le K\).
Fix \(M>0\), \(I=[-M,M]\) and \(\widetilde f:=fI_{[-M-1,M+1]}\in L^1(\mathbb{R})\), so that \(f_\varepsilon(x)=\int_0^1\widetilde f(x-\varepsilon y)g(y)\,dy\) and \(f=\widetilde f\) on \(I\) for \(0<\varepsilon<1\). By Fubini and \(\int_0^1g=1\), \[ \int_I|f_\varepsilon-f|\,dx\le\sup_{|s|\le\varepsilon}\bigl\|\widetilde f(\cdot-s)-\widetilde f\bigr\|_{L^1(\mathbb{R})}\xrightarrow[\varepsilon\to0]{}0 \] by continuity of translation in \(L^1\) (Lemma 4.2.3). Since \(\|f_\varepsilon-f_{\varepsilon^{\prime}}\|_{L^1(I)}=2M|c_\varepsilon-c_{\varepsilon^{\prime}}|\), the constants \(c_\varepsilon\) converge to some \(c\) independent of \(M\), and \(\int_I|f-c|\,dx=0\). As \(M\) was arbitrary, \(f=c\) a.e.
(\(\circ\)) (Lusin [633]) Let \(E\) be a measurable set in the unit circle \(S\) equipped with the linear Lebesgue measure. Suppose that \(E\) has infinitely many centers of symmetry, i.e., points \(c\in S\) such that along with every point \(e\in E\) the set \(E\) contains the point \(e^{\prime}\in S\) symmetric to \(e\) with respect to the straight line passing through \(c\) and the origin. Prove that the measure of \(E\) equals either \(0\) or \(2\pi\).
Two nearby axes of symmetry compose to a small rotation, so \(I_E\) has arbitrarily small periods and Exercise 5.8.109 applies. Parameterize \(S\) by the angle, identifying \(S\) with \(\mathbb{R}/2\pi\mathbb{Z}\) and \(\lambda\) with Lebesgue measure on \([0,2\pi)\). Reflection in the line through the origin and the point of angle \(\gamma\) acts by \(R_\gamma:\theta\mapsto2\gamma-\theta\), depends only on \(\gamma\bmod\pi\), and \[ R_{\gamma_1}\bigl(R_{\gamma_2}(\theta)\bigr)=\theta+2(\gamma_1-\gamma_2)\pmod{2\pi}, \] a rotation \(\rho_\tau\) by \(\tau=2(\gamma_1-\gamma_2)\). If \(\gamma\) is a centre of symmetry then \(R_\gamma(E)\subset E\), and since \(R_\gamma\) is an involution, \(R_\gamma(E)=E\).
Infinitely many centres give infinitely many distinct lines, i.e. an infinite \(\Gamma\subset[0,\pi)\) of such \(\gamma\); by compactness \(\Gamma\) has an accumulation point, so for each \(m\) there are \(\gamma_1,\gamma_2\in\Gamma\) with \(0<|\gamma_1-\gamma_2|<1/m\), and \(\tau_m:=2(\gamma_1-\gamma_2)\) satisfies \(0<|\tau_m|<2/m\) and \(\rho_{\tau_m}(E)=E\). Thus \(F:=I_E\), viewed as a \(2\pi\)-periodic measurable function on \(\mathbb{R}\), has the nonzero periods \(\tau_m\to0\), so by Exercise 5.8.109 it equals a constant a.e.; that constant is \(0\) or \(1\), i.e. \(\lambda(E)=0\) or \(\lambda(E)=2\pi\).
Method (2): the set \(G\) of \(\tau\) with \(\lambda(E\triangle\rho_\tau(E))=0\) is a subgroup of \(\mathbb{R}/2\pi\mathbb{Z}\) containing every \(\tau_m\), hence dense; \(\tau\mapsto\int_0^{2\pi}F(x)F(x+\tau)\,dx\) is continuous (translation is \(L^2\)-continuous) and equals \(\lambda(E)\) on \(G\), hence everywhere, so integrating in \(\tau\) and applying Fubini gives \(2\pi\lambda(E)=\lambda(E)^2\).
Show that a function \(f\) on \((a,b)\) is convex precisely when for every \([c,d]\subset(a,b)\), one has \[ f(x)=f( c)+\int_c^x g(t)\,dt,\qquad x\in[c,d], \] where \(g\) is a nondecreasing function on \([c,d]\).
Both directions run through the three-chord inequality: \(f\) is convex on \((a,b)\) iff for all \(u<v<w\) in \((a,b)\) \[ \frac{f(v)-f(u)}{v-u}\le\frac{f(w)-f(u)}{w-u}\le\frac{f(w)-f(v)}{w-v}, \tag{\(*\)} \] each inequality being a rearrangement of \((w-u)f(v)\le(w-v)f(u)+(v-u)f(w)\), i.e. of convexity for \(v=\theta u+(1-\theta)w\), \(\theta=(w-v)/(w-u)\).
Sufficiency. Given \(u<v<w\) in \((a,b)\), pick \([c,d]\subset(a,b)\) containing \(u,w\) and the corresponding nondecreasing \(g\) (bounded and Borel, hence integrable), so \(f(y)-f(x)=\int_x^yg\) on \([c,d]\). Monotonicity of \(g\) gives \(\int_u^vg\le(v-u)g(v)\) and \(\int_v^wg\ge(w-v)g(v)\), whence \[ \frac{f(v)-f(u)}{v-u}\le g(v)\le\frac{f(w)-f(v)}{w-v}, \] which is \((*)\) for \(u,v,w\); so \(f\) is convex.
Necessity. Let \(f\) be convex and \([c,d]\subset(a,b)\). First, \(f\) is Lipschitz on \([c,d]\): choose \(a<c^{\prime}<c\le d<d^{\prime}<b\); for \(c\le x<y\le d\), applying \((*)\) to \(c^{\prime}<x<y\) and to \(x<y<d^{\prime}\), \[ \frac{f( c)-f(c^{\prime})}{c-c^{\prime}}\le\frac{f(x)-f(c^{\prime})}{x-c^{\prime}}\le\frac{f(y)-f(x)}{y-x}\le\frac{f(d^{\prime})-f(y)}{d^{\prime}-y}\le\frac{f(d^{\prime})-f(d)}{d^{\prime}-d}, \] the outer inequalities being \((*)\) again, so \(|f(y)-f(x)|\le M|y-x|\) with \(M\) the larger of the two extreme quotients in absolute value; in particular \(f\) is absolutely continuous on \([c,d]\). Next put \(g(x):=f_+’(x)\), which exists and is finite because by \((*)\) the quotient \(h\mapsto(f(x+h)-f(x))/h\) is nondecreasing in \(h>0\) and bounded below by \((f(x)-f(x-h_0))/h_0\); and \(g\) is nondecreasing, since for \(x<y\) and small \(h>0\), \((*)\) gives \[ \frac{f(x+h)-f(x)}{h}\le\frac{f(y)-f(x)}{y-x}\le\frac{f(y+h)-f(y)}{h}, \] and one lets \(h\downarrow0\). Being absolutely continuous, \(f(x)=f( c)+\int_c^x\varphi\) for some integrable \(\varphi\) (Theorem 5.3.6) with \(\varphi=f^{\prime}\) a.e. (Theorem 5.4.2), and \(f^{\prime}=g\) wherever the two-sided derivative exists; hence \(\varphi=g\) a.e. and \[ f(x)=f( c)+\int_c^x g(t)\,dt,\qquad x\in[c,d], \] with \(g\) nondecreasing.
(\(\circ\)) Let \(f\) be an absolutely continuous function on \([a,b]\), let \(\mu\) be a bounded Borel measure on \([a,b]\), and let \(\Phi_\mu(t)=\mu\bigl([a,t)\bigr)\), \(\Phi_\mu(a)=0\).
(i) Prove the equality \[ \int_{[a,b]}f(t)\,\mu(dt)=f(b)\Phi_\mu(b+)-\int_a^b f^{\prime}(t)\Phi_\mu(t)\,dt. \]
(ii) Let \(\nu\) be a bounded Borel measure on \([a,b]\). Suppose that the function \(\Phi_\nu(t)=\nu\bigl([a,t)\bigr)\) is continuous. Prove the equality \[ \int_{[a,b]}\Phi_\nu(t)\,\mu(dt)=\Phi_\mu(b+)\Phi_\nu(b)-\int_{[a,b]}\Phi_\mu(t)\,\nu(dt). \]
(iii) Let \(\mu\) be a probability measure on \([0,+\infty)\) and \(\Psi_\mu(t)=\mu\bigl([t,+\infty)\bigr)\). Suppose that a function \(f\) is absolutely continuous on every closed interval, \(f(0)=0\) and \(f^{\prime}\ge0\). Prove the equality \[ \int_{[0,+\infty)}f(t)\,\mu(dt)=\int_0^{+\infty}f^{\prime}(t)\Psi_\mu(t)\,dt. \] Prove the same equality if the condition \(f^{\prime}\ge0\) is replaced by the following condition: \(f\in L^1(\mu)\) and \(f^{\prime}\Psi_\mu\in L^1(\mathbb{R}^1)\).
All three are Fubini applied to \(\{s<t\}\) in a product of Lebesgue measure (or \(\nu\)) with \(\mu\). Note \(\Phi_\mu(b+)=\mu([a,b])\), and that \(\Phi_\mu\) is bounded Borel (a difference of bounded nondecreasing left continuous functions, by the Jordan decomposition), while \(f\) is continuous, hence \(\mu\)-integrable.
(i) Absolute continuity gives \(f(t)=f(b)-\int_a^bI_{\{s>t\}}f^{\prime}(s)\,ds\) with \(f^{\prime}\) integrable (Theorems 5.3.6, 5.4.2), and \[ \int_{[a,b]}\int_a^b I_{\{s>t\}}|f^{\prime}(s)|\,ds\,|\mu|(dt)\le \|f^{\prime}\|_{L^1[a,b]}\,|\mu|\bigl([a,b]\bigr)<\infty, \] so Fubini applies to \(\lambda\otimes\mu\) (separately to \(\mu^{\pm}\)). Integrating in \(\mu\) and using \(\mu(\{t<s\})=\Phi_\mu(s)\), \[ \int_{[a,b]}f(t)\,\mu(dt)=f(b)\Phi_\mu(b+)-\int_a^b f^{\prime}(s)\Phi_\mu(s)\,ds. \]
(ii) Continuity of \(\Phi_\nu\) on the line means \(\nu\) is atomless, since \(\Phi_\nu\) is left continuous and \[ \Phi_\nu(t+)-\Phi_\nu(t)=\nu\bigl([a,t]\bigr)-\nu\bigl([a,t)\bigr)=\nu(\{t\}) \] so that \(\Phi_\nu(b)=\nu([a,b])\). On the bounded product \(\nu\otimes\mu\) (finite, as \(|\nu\otimes\mu|\le|\nu|\otimes|\mu|\)), Fubini applied to the indicator of \(\{s<t\}\) gives \[ \int_{[a,b]}\Phi_\nu(t)\,\mu(dt)=\int_{[a,b]}\nu\bigl([a,t)\bigr)\mu(dt)=(\nu\otimes\mu)\bigl(\{s<t\}\bigr), \] and symmetrically, for the indicator of \(\{s>t\}\), \[ \int_{[a,b]}\Phi_\mu(t)\,\nu(dt)=\int_{[a,b]}\mu\bigl([a,s)\bigr)\nu(ds)=(\nu\otimes\mu)\bigl(\{s>t\}\bigr) \] (the \(\nu\)-variable being \(s\)). The sets \(\{s<t\}\), \(\{s>t\}\) and \(\Delta=\{s=t\}\) partition the square, and \((\nu\otimes\mu)(\Delta)=\int\nu(\{t\})\,\mu(dt)=0\) since \(\nu\) is atomless. Hence
\begin{equation*} \begin{aligned} \int_{[a,b]}\Phi_\nu(t)\,\mu(dt)+\int_{[a,b]}\Phi_\mu(t)\,\nu(dt) &=(\nu\otimes\mu)\bigl([a,b]\times[a,b]\bigr)\\ &=\nu\bigl([a,b]\bigr)\mu\bigl([a,b]\bigr)=\Phi_\nu(b)\Phi_\mu(b+), \end{aligned} \end{equation*}
which is the asserted equality. (If \(\nu\) has an atom at \(b\), the same computation adds the term \(\mu([a,b))\nu(\{b\})\) on the right.)
(iii) Here \(f(t)=\int_0^{+\infty}I_{\{s<t\}}f^{\prime}(s)\,ds\), and \(\mu((s,+\infty))=\Psi_\mu(s)-\mu(\{s\})\) coincides with \(\Psi_\mu(s)\) for Lebesgue-a.e. \(s\), a finite measure having at most countably many atoms.
(a) If \(f^{\prime}\ge0\), the integrand \(I_{\{s<t\}}f^{\prime}(s)\) is nonnegative, so Tonelli applies to the \(\sigma\)-finite product \(\lambda\otimes\mu\) (both sides possibly \(+\infty\)) and \[ \int_{[0,+\infty)}f(t)\,\mu(dt)=\int_0^{+\infty}f^{\prime}(s)\,\mu\bigl((s,+\infty)\bigr)\,ds=\int_0^{+\infty}f^{\prime}(s)\Psi_\mu(s)\,ds. \]
(b) If \(f\in L^1(\mu)\) and \(f^{\prime}\Psi_\mu\in L^1(\mathbb{R}^1)\), then \(\int_0^{\infty}|f^{\prime}|\Psi_\mu\,ds<\infty\) says exactly that \(I_{\{s<t\}}f^{\prime}(s)\) is \(\lambda\otimes\mu\)-integrable, and Fubini gives the same identity with both sides absolutely convergent.
(i) (Lusin [633]) Construct a measurable set \(E\subset[0,1]\) such that, letting \(f=I_E\), one has \[ \int_0^1\left|\frac{f(x+t)-f(x-t)}{t}\right|dt=\infty\quad\text{for almost all }x\in[0,1]. \]
(ii) (Titchmarsh [946]) Let \(\varphi>0\) be a continuous function on \((0,1)\) such that the function \(1/\varphi\) has an infinite integral. Prove that:
(a) there exists a continuous function \(f\) such that \[ \int_0^1\frac{|f(x+t)-f(x-t)|}{\varphi(t)}\,dt=\infty\quad\text{for a.e. }x; \]
(b) there exists a continuous function \(g\) such that the integral of the function \([g(x+t)-g(x)]/\varphi(t)\) in \(t\) diverges for a.e. \(x\).
(i) Take \(E\) the symmetric difference of periodic sets \(C(p_k,\theta_k)\cap[0,1]\) with \(\theta_k=2^{-k^2}\) and \(p_k\) shrinking fast; each scale contributes \(\ge1\) to the integral over a disjoint range \((p_k,R_k)\). Write \(\Delta_A(x,t):=|I_A(x+t)-I_A(x-t)|\in\{0,1\}\), which is \(1\) exactly when one of \(x\pm t\) lies in \(A\) and which adds mod \(2\) under symmetric differences (as \(I_A=\sum_kI_{A_k}\) mod \(2\)); the integral to be made infinite is \(\int_0^1\Delta_E(x,t)t^{-1}dt\). For \(p>0\) and \(\theta\in(0,1)\) let \[ C(p,\theta):=\bigcup_{n\in\mathbb{Z}}\,[np,\,np+\theta p], \] a closed \(p\)-periodic set of density \(\theta\). For fixed \(x\), \(t\mapsto\Delta_C(x,t)\) is \(p\)-periodic, and substituting \(u=\{x+t\}_p\) (measure-preserving on \([0,p)\)), with \(\{x-t\}_p=\{2x-u\}_p\), \[ \overline{\Delta}_C(x):=\frac1p\int_0^p\Delta_C(x,t)\,dt=\frac1p\,\lambda\bigl(A\,\triangle\,\sigma A\bigr),\qquad A=[0,\theta p],\ \ \sigma u=\{2x-u\}_p, \] on the circle \(\mathbb{R}/p\mathbb{Z}\), \(\sigma\) being a reflection, so \(\sigma A\) is an arc of length \(\theta p\) and \(\lambda(A\triangle\sigma A)=2\theta p-2\lambda(A\cap\sigma A)\). The arcs are disjoint unless their centres are within \(\theta p\), i.e. unless \(\{2x-\theta p\}_p\) lies in an arc of length \(2\theta p\); as \(x\mapsto\{2x\}_p\) is two-to-one and doubles measure, that happens for \(x\) in a set of measure \(\le2\theta p\) per period. Hence \[ \overline{\Delta}_{C(p,\theta)}(x)=2\theta\quad\text{for all }x\text{ outside a }p\text{-periodic set of density }\le2\theta. \tag{1} \] Consequently, for \(M\in\mathbb{N}\) and such good \(x\),
\begin{equation*} \begin{aligned} \int_p^{Mp}\frac{\Delta_C(x,t)}{t}\,dt &\ge\ \sum_{m=1}^{M-1}\frac{1}{(m+1)p}\int_{mp}^{(m+1)p}\Delta_C(x,t)\,dt\\ &=2\theta\sum_{m=2}^{M}\frac1m\ \ge\ 2\theta(\ln M-1). \end{aligned}\tag{2} \end{equation*}
Put \(\theta_k:=2^{-k^2}\), so \(\sum_k\theta_k<\infty\), choose \(p_1>p_2>\cdots\) recursively and set \[ C_k:=C(p_k,\theta_k),\qquad R_k:=p_k\exp(1/\theta_k),\qquad M_k:=\lfloor R_k/p_k\rfloor . \] Having chosen \(p_1,\dots,p_{k-1}\) (hence \(C_1,\dots,C_{k-1}\), whose boundaries have finitely many points in \([0,1]\)), we take \(p_k>0\) so small that
- (a) \(R_k<\min(p_{k-1},2^{-k})\) (so the intervals \((p_k,R_k)\) are disjoint and shrink to \(0\));
- (b) \(\lambda\bigl(\{x\in[0,1]:\ \operatorname{dist}(x,\partial C_j)\le R_k\ \text{for some}\ j<k\}\bigr)\le2^{-k}\);
- (c) \(p_k\le\theta_k\,p_{k-1}\).
possible since \(R_k\to0\) as \(p_k\to0\) and \(\partial C_1\cup\cdots\cup\partial C_{k-1}\) meets \([0,1]\) in a finite set. Define \[ E:=\Bigl\{x\in[0,1]:\ x\ \text{belongs to finitely many and an odd number of the sets}\ C_k\Bigr\}, \] i.e. the symmetric difference of the \(C_k\cap[0,1]\); since \(\sum_k\lambda(C_k\cap[0,1])\le\sum_k(\theta_k+p_k)<\infty\), Borel–Cantelli puts almost every \(x\) in only finitely many \(C_k\), so \(E\) is measurable. Put \(f=I_E\), extended by \(0\).
Let \(B_k\) be the exceptional set of (1) for \(C_k\) (density \(\le2\theta_k\)) and \(B_k^{\prime}\) the set in (b) (measure \(\le2^{-k}\)). As \(\sum_k(2\theta_k+2^{-k})<\infty\), Borel–Cantelli gives, for a.e. \(x\in(0,1)\), an index \(k_0(x)\) beyond which \(x\notin B_k\cup B_k^{\prime}\) and \(R_k<\min(x,1-x)\), with \(x\) in only finitely many \(C_k\). Fix such \(x\) and \(k\ge k_0(x)\); we show \[ \int_{p_k}^{R_k}\frac{\Delta_E(x,t)}{t}\,dt\ \ge\ 1 . \tag{3} \] which, the intervals \((p_k,R_k)\) being disjoint and the integrand nonnegative, yields \(\int_0^1\Delta_E(x,t)t^{-1}dt=\infty\). For \(t<R_k<\min(x,1-x)\) both \(x\pm t\) lie in \((0,1)\), so there \[ \Delta_E(x,t)=\bigoplus_j\Delta_{C_j}(x,t)\quad\text{for a.e. }t \] (for a.e. \(t\), by Borel–Cantelli again, since \(\lambda\{t\in(0,1):x\pm t\in C_j\}\le\theta_j+p_j\) is summable).
Coarse levels do not interfere: for \(j<k\) and \(t\le R_k\), (b) gives \(\operatorname{dist}(x,\partial C_j)>t\), so \(I_{C_j}\) is constant on \((x-t,x+t)\) and \(\Delta_{C_j}(x,t)=0\).
Fine levels interfere little: for \(j>k\), (c) gives \(p_j\le\theta_jp_k\), and \(\{t:x+t\in C_j\}\) is \(p_j\)-periodic of density \(\theta_j\), so its measure in an interval of length \(L\) is \(\le\theta_jL+p_j\). Splitting \((p_k,R_k)\) into \(\le\log_2(R_k/p_k)+1\) dyadic pieces and using \(\Delta_{C_j}(x,t)\le I_{C_j}(x+t)+I_{C_j}(x-t)\),
\begin{equation*} \begin{aligned} \int_{p_k}^{R_k}\frac{\Delta_{C_j}(x,t)}{t}\,dt &\le\sum_i\frac{2(\theta_j2^ip_k+p_j)}{2^ip_k}\\ &\le 2(\theta_j+\theta_j)\bigl(\log_2(R_k/p_k)+1\bigr)\le\frac{8\theta_j}{\theta_k\ln2}, \end{aligned} \end{equation*}
since \(\log_2(R_k/p_k)=1/(\theta_k\ln2)\). Summing over \(j>k\) and using \(\sum_{j>k}\theta_j\le2\theta_{k+1}=2\cdot2^{-(k+1)^2}\), \[ \sum_{j>k}\int_{p_k}^{R_k}\frac{\Delta_{C_j}(x,t)}{t}\,dt\le\frac{16}{\ln2}\,\frac{\theta_{k+1}}{\theta_k}=\frac{16}{\ln2}\,2^{-2k-1}\longrightarrow0 . \]
Since \(\Delta_E\ge\Delta_{C_k}-\sum_{j\ne k}\Delta_{C_j}\) pointwise, the last three displays and (2) give, for good \(x\) and large \(k\), \[ \int_{p_k}^{R_k}\frac{\Delta_E(x,t)}{t}\,dt\ \ge\ 2\theta_k(\ln M_k-1)-\frac{16}{\ln2}2^{-2k-1}\ \ge\ 2\theta_k\Bigl(\frac1{\theta_k}-2\Bigr)-o(1)\ \ge\ 1, \] since \(\ln M_k\ge\ln(R_k/p_k)-1=1/\theta_k-1\). This is (3), hence the assertion.
(ii) Since \(1/\varphi\) is continuous on \((0,1)\), the divergence of \(\int_0^1dt/\varphi\) is at an endpoint, so at least one of \(\int_0^{1/2}dt/\varphi\), \(\int_{1/2}^1dt/\varphi\) is infinite; three cases, the last by Baire category in \(C(\mathbb{T})\), the (complete) space of continuous \(1\)-periodic functions with the sup-norm. Throughout we use:
Lemma. If the \(1\)-periodic \(W_\alpha\) are uniformly bounded and equi-Lipschitz with means \(\overline W_\alpha\), then for every finite \([u_0,u_1]\) and \(\psi\in L^1[u_0,u_1]\), \(\int\psi(t)W_\alpha(Nt)\,dt-\overline W_\alpha\int\psi\to0\) as \(N\to\infty\), uniformly in \(\alpha\). (For \(\psi\) an interval indicator the error is \(O(N^{-1})\) uniformly, each full period of length \(1/N\) contributing \(\overline W_\alpha/N\) and at most two periods being incomplete; the general case follows by \(L^1\)-approximation by step functions, controlled by \(\sup_\alpha\|W_\alpha\|_\infty\).)
Case A: \(\int_{1/2}^1 dt/\varphi=\infty\). Take \(f(x)=g(x)=x\). Then \(|f(x+t)-f(x-t)|/\varphi(t)=2t/\varphi(t)\ge1/\varphi(t)\) and \([g(x+t)-g(x)]/\varphi(t)=t/\varphi(t)\ge1/(2\varphi(t))\) for \(t\ge1/2\), so both integrals diverge (to \(+\infty\)) for every \(x\). So we may assume from now on that \(\int_{1/2}^1dt/\varphi<\infty\) and \(\int_0^{1/2}dt/\varphi=\infty\).
Case B1: \(\int_0^{1/2}t\,\varphi(t)^{-1}dt=\infty\). Take \(f(x)=g(x)=\sin2\pi x\). If \(f^{\prime}(x)\neq0\), then there is \(\tau>0\) with \(|f(x+t)-f(x-t)|\ge t|f^{\prime}(x)|\) and \(|f(x+t)-f(x)|\ge\frac12t|f^{\prime}(x)|\) for \(0<t<\tau\) (Taylor’s formula with the uniformly bounded second derivative), whence both integrals are at least \(\frac{|f^{\prime}(x)|}{2}\int_0^{\tau}t\varphi(t)^{-1}dt=\infty\), the divergence of the latter being at \(t=0\). Since \(f^{\prime}(x)=2\pi\cos2\pi x\) vanishes only at countably many points, this holds for a.e. \(x\). For (b) the divergence is genuine and not merely a failure of absolute convergence: shrinking \(\tau\) we may assume that \(g(x+t)-g(x)\) has the constant sign of \(g^{\prime}(x)\) for \(0<t<\tau\), so the integral over \((0,\tau)\) equals \(+\infty\) or \(-\infty\), while over \([\tau,1]\) it is finite because \(1/\varphi\) is bounded on \([\tau,1/2]\) and integrable on \([1/2,1]\).
Case B2: \(\int_0^{1/2}dt/\varphi=\infty\), \(\int_0^{1}t\,\varphi(t)^{-1}dt<\infty\) (recall \(\int_{1/2}^1dt/\varphi<\infty\), so the last integral is finite as soon as it is finite near \(0\)). Here every Lipschitz function \(P\) satisfies
\begin{equation*} \begin{aligned} \int_0^1\frac{|P(x+t)-P(x-t)|}{\varphi(t)}dt &\le2\operatorname{Lip}(P)\int_0^1\frac{t\,dt}{\varphi(t)}=:L(P)<\infty,\\ \int_0^1\frac{|P(x+t)-P(x)|}{\varphi(t)}dt&\le L(P), \end{aligned} \end{equation*}
for all \(x\), so no Lipschitz function serves and we argue by category. For \(f\in C(\mathbb{T})\) put \[ I_f(x):=\int_0^1\frac{|f(x+t)-f(x-t)|}{\varphi(t)}dt,\qquad J_f(x):=\int_0^1\frac{|f(x+t)-f(x)|}{\varphi(t)}dt, \] and let \(I_f^{(n)},J_f^{(n)}\) be the same integrals over \([1/n,1]\). Since \(1/\varphi\) is bounded on \([1/n,1]\) (Case B assumptions give integrability near \(1\); boundedness on \([1/n,1-\varepsilon]\) is clear, and near \(1\) we only need integrability), the functions \(I^{(n)}_f,J^{(n)}_f\) are continuous in \(x\) and \[ |I^{(n)}_f(x)-I^{(n)}_h(x)|\le2\|f-h\|_\infty c_n,\qquad c_n:=\int_{1/n}^1\frac{dt}{\varphi(t)}<\infty, \] and likewise for \(J\). Moreover \(I^{(n)}_f\uparrow I_f\) and \(J^{(n)}_f\uparrow J_f\) pointwise, so \(I_f\) and \(J_f\) are Borel (lower semicontinuous) functions with values in \([0,+\infty]\).
For \(m,k\in\mathbb{N}\) set \[ A_{m,k}:=\bigl\{f\in C(\mathbb{T}):\ \lambda(\{x\in[0,1]:I_f(x)\le m\})\ge1/k\bigr\}. \] These sets are closed. Indeed, let \(f_i\to f\) uniformly with \(f_i\in A_{m,k}\), and fix \(n\) and \(\varepsilon>0\). On the set where \(I_{f_i}\le m\) we have \(I^{(n)}_{f_i}\le m\), hence \(I^{(n)}_f\le m+2\|f-f_i\|_\infty c_n\le m+\varepsilon\) for large \(i\); therefore \(\lambda(\{I^{(n)}_f\le m+\varepsilon\})\ge1/k\). Letting \(\varepsilon\downarrow0\) and using continuity of measure along the decreasing sets \(\{I^{(n)}_f\le m+\varepsilon\}\) we get \(\lambda(\{I^{(n)}_f\le m\})\ge1/k\); finally the sets \(\{I^{(n)}_f\le m\}\) decrease in \(n\) to \(\{I_f\le m\}\), so \(\lambda(\{I_f\le m\})\ge1/k\), i.e. \(f\in A_{m,k}\).
They have empty interior. Given \(f\in C(\mathbb{T})\) and \(\delta>0\), choose a piecewise linear \(P\) with \(\|f-P\|_\infty<\delta/2\), and put \(A:=\delta/4\), \(\eta_N(x):=A\sin(2\pi Nx)\), \(h_N:=P+\eta_N\), so \(\|h_N-f\|_\infty<\delta\). Since \(\eta_N(x+t)-\eta_N(x-t)=2A\cos(2\pi Nx)\sin(2\pi Nt)\), \[ I_{\eta_N}(x)=2A|\cos2\pi Nx|\,S_N,\qquad S_N:=\int_0^1\frac{|\sin2\pi Nt|}{\varphi(t)}\,dt. \] By the Lemma applied on \([\varepsilon,1/2]\) with \(W(u)=|\sin2\pi u|\), \(\overline W=2/\pi\), we get \(\liminf_NS_N\ge\frac2\pi\int_\varepsilon^{1/2}\varphi^{-1}dt\) for every \(\varepsilon>0\), whence \(S_N\to\infty\). Now \(I_{h_N}\ge I_{\eta_N}-I_P\ge2A|\cos2\pi Nx|S_N-L(P)\), so \[ \{x:I_{h_N}(x)\le m\}\subset\Bigl\{x:|\cos2\pi Nx|\le\varepsilon_N\Bigr\},\qquad \varepsilon_N:=\frac{m+L(P)}{2AS_N}\to0 . \] whose measure \(\frac2\pi\arcsin\varepsilon_N\to0\); choosing \(N\) with \(\frac2\pi\arcsin\varepsilon_N<1/k\) gives \(h_N\notin A_{m,k}\), so \(A_{m,k}\) is nowhere dense. By Baire, some \(f\in C(\mathbb{T})\) lies outside \(\bigcup_{m,k}A_{m,k}\); then \(\{I_f<\infty\}=\bigcup_m\{I_f\le m\}\) is null, i.e. \(I_f=\infty\) a.e., which is (a).
For (b) repeat with \(J\) for \(I\), only the perturbation estimate changing: with \(\eta_N(x)=A\sin2\pi Nx\), \(\theta:=2\pi Nx\), \[ |\eta_N(x+t)-\eta_N(x)|=A\bigl|\sin(\theta+2\pi Nt)-\sin\theta\bigr|=A\,W_\theta(Nt),\qquad W_\theta(u):=|\sin(\theta+2\pi u)-\sin\theta| , \] a family that is \(1\)-periodic, bounded by \(2\) and equi-Lipschitz, with mean \[ \overline W_\theta=\frac{1}{2\pi}\int_0^{2\pi}|\sin v-\sin\theta|\,dv\ \ge\ \frac{1}{2\pi}\int_0^{2\pi}|\sin v|\,dv=\frac2\pi, \] since \(a\mapsto\int_0^{2\pi}|\sin v-a|dv\) is convex with derivative \(\lambda\{\sin v<a\}-\lambda\{\sin v>a\}\), vanishing at the median \(a=0\). By the Lemma on \([\varepsilon,1/2]\) with \(\psi=1/\varphi\), uniformly in \(\theta\), \[ \liminf_{N\to\infty}\ \inf_{x}J_{\eta_N}(x)\ \ge\ \frac{2A}{\pi}\int_\varepsilon^{1/2}\frac{dt}{\varphi(t)} \] for every \(\varepsilon>0\), and the right-hand side tends to \(\infty\) as \(\varepsilon\to0\). Hence \(\inf_xJ_{\eta_N}(x)\to\infty\). Consequently, for \(h_N=P+\eta_N\) we get \(J_{h_N}(x)\ge J_{\eta_N}(x)-L(P)>m\) for all \(x\) once \(N\) is large, so the corresponding sets \[ A^{\prime}_{m,k}:=\{g\in C(\mathbb{T}):\lambda(\{x:J_g(x)\le m\})\ge1/k\} \] (closed by the same argument as before) have empty interior. The Baire theorem again yields a continuous \(1\)-periodic \(g\) with \(J_g(x)=\infty\) for a.e. \(x\), i.e. \[ \int_0^1\frac{|g(x+t)-g(x)|}{\varphi(t)}\,dt=\infty\quad\text{for a.e. }x . \]
It remains to exclude convergence of the improper integral \(\lim_{\varepsilon\downarrow0}\int_\varepsilon^1[g(x+t)-g(x)]\varphi(t)^{-1}dt\), which \(J_g=\infty\) alone does not preclude. Set
\begin{equation*} \begin{aligned} K^{(\varepsilon)}_g(x)&:=\int_\varepsilon^1\frac{g(x+t)-g(x)}{\varphi(t)}\,dt,\qquad K^{[\varepsilon]}_g(x):=\sup_{\varepsilon\le\delta<1}\bigl|K^{(\delta)}_g(x)\bigr|,\\ K_g&:=\sup_{0<\varepsilon<1}K^{[\varepsilon]}_g . \end{aligned} \end{equation*}
If the improper integral converges at \(x\), then \(K_g(x)<\infty\). For fixed \(\varepsilon\) the functions \(K^{(\delta)}_g\), \(\delta\ge\varepsilon\), are continuous in \(x\) with the common modulus of continuity \(2\omega_g(\cdot)c_\varepsilon\), where \(c_\varepsilon:=\int_\varepsilon^1\varphi(t)^{-1}dt<\infty\) and \(\omega_g\) is the modulus of continuity of \(g\); hence \(K^{[\varepsilon]}_g\) is continuous, \(|K^{[\varepsilon]}_g-K^{[\varepsilon]}_h|\le2\|g-h\|_\infty c_\varepsilon\), and \(K^{[\varepsilon]}_g\uparrow K_g\) as \(\varepsilon\downarrow0\). Exactly as before, the sets \[ A^{\prime\prime}_{m,k}:=\{g\in C(\mathbb{T}):\ \lambda(\{x:K_g(x)\le m\})\ge1/k\} \] are closed, and have empty interior: given \(f\), \(\delta>0\), take \(P\) and \(A=\delta/4\) as above and apply the Lemma to \(V_\theta(u):=\sin(\theta+2\pi u)-\sin\theta\) (no absolute value), \(1\)-periodic, bounded by \(2\), equi-Lipschitz, with mean \(-\sin\theta\): \[ K^{(\varepsilon)}_{\eta_N}(x)=A\int_\varepsilon^1\frac{V_\theta(Nt)}{\varphi(t)}\,dt=-A\sin\theta\,c_\varepsilon+A\gamma_N(\varepsilon,\theta),\qquad \theta=2\pi Nx, \] where \(\sup_\theta|\gamma_N(\varepsilon,\theta)|\to0\) as \(N\to\infty\) for each fixed \(\varepsilon\). Given \(m\) and \(k\), choose first \(\beta>0\) with \(\frac2\pi\arcsin\beta<1/k\), then \(\varepsilon>0\) with \(A\beta c_\varepsilon>m+L(P)+1\) (possible because \(c_\varepsilon\to\infty\) as \(\varepsilon\downarrow0\), the integral of \(1/\varphi\) being infinite), and finally \(N\) so large that \(A\sup_\theta|\gamma_N(\varepsilon,\theta)|<1\). For every \(x\) with \(|\sin2\pi Nx|\ge\beta\) we then get, with \(h_N=P+\eta_N\),
\begin{equation*} \begin{aligned} K_{h_N}(x)&\ \ge\ \bigl|K^{(\varepsilon)}_{h_N}(x)\bigr| \ \ge\ \bigl|K^{(\varepsilon)}_{\eta_N}(x)\bigr|-\bigl|K^{(\varepsilon)}_{P}(x)\bigr|\\ &\ \ge\ A\beta c_\varepsilon-1-L(P)\ >\ m, \end{aligned} \end{equation*}
since \(|K^{(\varepsilon)}_P(x)|\le J_P(x)\le L(P)\). Hence \(\{K_{h_N}\le m\}\) is contained in \(\{|\sin2\pi Nx|<\beta\}\), whose measure is \(\frac2\pi\arcsin\beta<1/k\), so \(h_N\notin A^{\prime\prime}_{m,k}\) while \(\|h_N-f\|_\infty<\delta\).
Baire applied to all the \(A^{\prime}_{m,k}\) and \(A^{\prime\prime}_{m,k}\) at once gives a continuous \(1\)-periodic \(g\) with \(J_g=K_g=\infty\) a.e.; at such \(x\) the integrand is not Lebesgue integrable and the truncations \(\int_\varepsilon^1\) are unbounded, so the integral diverges in every sense.
Exercises 5.8.114–5.8.120
Construct a continuous function \(f\) such that for every point \(x\) in some everywhere dense set of cardinality of the continuum in \([0,1]\), there is no finite limit
\begin{equation*} \lim_{\varepsilon\to 0}\int_\varepsilon^1 \frac{f(x+t)-f(x-t)}{t}\,dt . \end{equation*}
Take \(f(x)=\int\psi(x-y)\,\mu(dy)\), the potential of a measure \(\mu\) carried by countably many Cantor sets against the odd nondecreasing kernel below; monotonicity of \(\psi\) makes every integrand nonnegative, so divergence comes from monotone convergence with no cancellation. Put
\begin{equation*} \psi(u)=\begin{cases} \dfrac{1}{\log(e/u)}, & 0<u\le 1,\\[2mm] 1, & u>1,\end{cases} \qquad \psi(0)=0,\qquad \psi(-u)=-\psi(u). \end{equation*}
so that \(\psi\) is odd, nondecreasing, continuous and \(|\psi|\le1\). For \(s\in\mathbb{R}\), \(0<\varepsilon<1\) set
\begin{equation*} I_\varepsilon(s)=\int_\varepsilon^1\frac{\psi(s+t)-\psi(s-t)}{t}\,dt,\qquad I_0(s)=\int_0^1\frac{\psi(s+t)-\psi(s-t)}{t}\,dt\in[0,+\infty]. \end{equation*}
The integrand is nonnegative, so \(I_\varepsilon\ge0\) and \(I_\varepsilon\uparrow I_0\) as \(\varepsilon\downarrow0\) by monotone convergence.
Lower bound: let \(0<|s|<1/4\). For \(t\ge2|s|\) we have \(s+t\ge t/2\) and \(s-t\le-t/2\), so
\begin{equation*} \psi(s+t)-\psi(s-t)\ \ge\ \psi(t/2)-\psi(-t/2)=2\psi(t/2). \end{equation*}
Therefore, substituting \(u=t/2\) (which leaves \(dt/t=du/u\) unchanged),
\begin{equation*} I_0(s)\ \ge\ \int_{2|s|}^1\frac{2\psi(t/2)}{t}\,dt\ =\ 2\int_{|s|}^{1/2}\frac{\psi(u)}{u}\,du\ =:\ 2\Lambda(|s|). \end{equation*}
Here \(\Lambda\) is nonincreasing with
\begin{equation*} \Lambda( r)=\int_r^{1/2}\frac{du}{u\log(e/u)}=\log\log(e/r)-\log\log(2e)\ \xrightarrow[r\to 0^+]{}\ +\infty . \end{equation*}
Now fix a nonempty open \((p,q)\subset(0,1)\) and choose \(\ell_0>\ell_1>\cdots\) with
\begin{equation*} \ell_0<\min\{q-p,\ 1/4\},\qquad \ell_{n+1}\le \ell_n/3,\qquad \log\log(e/\ell_n)\ \ge\ 4^n\quad (n\ge 0); \end{equation*}
e.g. \(\ell_n=\min\{3^{-n}\ell_0,\exp(-e^{4^n})\}\) (Check!); \(\ell_0<1/4\) keeps all distances below in the range where the lower bound applies. Build the nested family: \(J_\emptyset\subset(p,q)\) closed of length \(\ell_0\), and \(J_{w0},J_{w1}\subset J_w\) disjoint closed of length \(\ell_{n+1}\) (possible as \(2\ell_{n+1}<\ell_n\)). Set
\begin{equation*} K=\bigcap_{n\ge 0}\ \bigcup_{|w|=n}J_w . \end{equation*}
a nonempty compact subset of \((p,q)\) of cardinality continuum, and let \(\mu\) be the Borel probability measure with \(\mu(J_w)=2^{-|w|}\), the image of the product of the measures \((\delta_0+\delta_1)/2\) on \(\{0,1\}^{\mathbb{N}}\) (Theorem 3.5.1) under the continuous coding map.
Fix \(x\in K\), let \(J^{(n)}\ni x\) be the level-\(n\) interval and \(A_n:=J^{(n)}\setminus J^{(n+1)}\), so the \(A_n\) are disjoint with \(\mu(A_n)=2^{-n-1}\), and \(\mu(\{x\})=0\), so they exhaust \(\mu\) up to a null set. For \(y\in A_n\), \(0<|x-y|\le\ell_n<1/4\), so \(I_0(x-y)\ge2\Lambda(|x-y|)\ge2\Lambda(\ell_n)\) and
\begin{equation*} \int_K I_0(x-y)\,\mu(dy)\ \ge\ \sum_{n\ge 0}\int_{A_n}2\Lambda(|x-y|)\,\mu(dy)\ \ge\ \sum_{n\ge 0}2^{-n-1}\cdot 2\Lambda(\ell_n) =\sum_{n\ge 0}2^{-n}\bigl(\log\log(e/\ell_n)-\log\log(2e)\bigr), \end{equation*}
which diverges, its \(n\)-th term being at least \(2^{-n}(4^n-\log\log(2e))\to+\infty\). Thus
\begin{equation*} \int I_0(x-y)\,\mu(dy)=+\infty\qquad\text{for every }x\in K. \tag{\(*\)} \end{equation*}
Enumerate the rational-endpoint intervals \((p_j,q_j)\subset(0,1)\), take \(K_j,\mu_j\) as above inside each, and put
\begin{equation*} \mu=\sum_{j=1}^\infty 2^{-j}\mu_j , \end{equation*}
a Borel probability measure on \([0,1]\), and \(f(x):=\int\psi(x-y)\,\mu(dy)\), continuous since \(\psi\) is bounded and uniformly continuous (dominated convergence).
Fix \(x\) and \(\varepsilon\in(0,1)\); the integrand \((t,y)\mapsto[\psi(x-y+t)-\psi(x-y-t)]/t\) is nonnegative Borel, so Tonelli gives
\begin{equation*} \int_\varepsilon^1\frac{f(x+t)-f(x-t)}{t}\,dt=\int\!\!\left(\int_\varepsilon^1\frac{\psi(x-y+t)-\psi(x-y-t)}{t}\,dt\right)\mu(dy)=\int I_\varepsilon(x-y)\,\mu(dy). \end{equation*}
As \(\varepsilon\downarrow0\) the integrands increase to \(I_0(x-y)\), so by monotone convergence
\begin{equation*} \lim_{\varepsilon\to 0}\int_\varepsilon^1\frac{f(x+t)-f(x-t)}{t}\,dt=\int I_0(x-y)\,\mu(dy)\in[0,+\infty]. \end{equation*}
Let \(D=\bigcup_{j\ge1}K_j\). If \(x\in K_j\), then, since \(I_0\ge 0\) and \(\mu\ge 2^{-j}\mu_j\),
\begin{equation*} \int I_0(x-y)\,\mu(dy)\ \ge\ 2^{-j}\int I_0(x-y)\,\mu_j(dy)=+\infty \end{equation*}
by \((*)\), so the limit is \(+\infty\) at every \(x\in D\). Finally \(D\) is dense in \([0,1]\), each rational-endpoint interval containing some \(K_j\), and has cardinality continuum, already \(K_1\) doing so.
(Rubel) Let \(f\) be a finite measurable real function on the real line. We consider the following functions with values in \([0,+\infty]\):
\begin{equation*} \varphi(x)=\sup_t |f(x+t)-f(x)|,\qquad \varphi_*(x)=\sup_t|f(x+t)-f(x-t)|, \end{equation*}
\begin{equation*} \Phi(x)=\sup_{t\ne 0}\left|\frac{f(x+t)-f(x)}{t}\right|,\qquad \Phi_*(x)=\sup_{t\ne 0}\left|\frac{f(x+t)-f(x-t)}{t}\right| . \end{equation*}
Show that the functions \(\varphi_*\) and \(\Phi_*\) may not be measurable, although \(\varphi\) and \(\Phi\) are always measurable.
The counterexample is \(f=I_E-I_{E+2}\) for a null set \(E\) with \(E+E\) nonmeasurable (Exercise 1.12.67); \(\varphi\) and \(\Phi\) are measurable because the one-sided suprema collapse to expressions in \(f\) and monotone functions.
(i) \(\varphi\) is measurable. As \(y=x+t\) runs over \(\mathbb{R}\),
\begin{equation*} \varphi(x)=\sup_{y\in\mathbb{R}}|f(y)-f(x)|=\max\bigl(M-f(x),\, f(x)-m\bigr),\qquad M:=\sup_{\mathbb{R}}f\in(-\infty,+\infty],\quad m:=\inf_{\mathbb{R}}f\in[-\infty,+\infty). \end{equation*}
since \(\sup_y|f(y)-c|=\max(\sup_yf(y)-c,\,c-\inf_yf(y))\); \(M,m\) being constants, \(\varphi\) is measurable.
(ii) \(\Phi\) is measurable. With \(y=x+t\),
\begin{equation*} \Phi(x)=\sup_{y\ne x}\left|\frac{f(y)-f(x)}{y-x}\right| =\max\bigl(\Psi_f(x),\,\Psi_{-f}(x)\bigr),\qquad \Psi_g(x):=\sup_{y\ne x}\frac{g(y)-g(x)}{y-x}, \end{equation*}
by \(|a|=\max(a,-a)\), so it suffices that \(\Psi_g\) be measurable for finite measurable \(g\). Fix \(c\) and put \(h(y):=g(y)-cy\); the inequality \((g(y)-g(x))/(y-x)>c\) reads \(h(y)>h(x)\) for \(y>x\) and \(h(y)<h(x)\) for \(y<x\), so
\begin{equation*} \Psi_g(x)>c\iff \Bigl[\exists\,y>x:\ h(y)>h(x)\Bigr]\ \text{ or }\ \Bigl[\exists\,y<x:\ h(y)<h(x)\Bigr] \iff S(x)>h(x)\ \text{ or }\ T(x)<h(x), \end{equation*}
where
\begin{equation*} S(x)=\sup_{y>x}h(y),\qquad T(x)=\inf_{y<x}h(y) \end{equation*}
Both \(S\) and \(T\) are nonincreasing in \(x\), hence Borel, so
\begin{equation*} \{\Psi_g>c\}=\{x:\ S(x)>h(x)\}\cup\{x:\ T(x)<h(x)\} =\bigcup_{q\in\mathbb{Q}}\Bigl(\{S>q\}\cap\{h<q\}\Bigr)\ \cup\ \bigcup_{q\in\mathbb{Q}}\Bigl(\{T<q\}\cap\{h>q\}\Bigr) \end{equation*}
is measurable (a rational separates two extended reals exactly when one exceeds the other). As \(c\) was arbitrary, \(\Psi_g\), hence \(\Phi\), is measurable.
(iii) With \(u=x+t\), \(v=x-t\), so \(u+v=2x\) and \(t=(u-v)/2\),
\begin{equation*} \varphi_*(x)=\sup\{|f(u)-f(v)|:\ u+v=2x\},\qquad \Phi_*(x)=\sup\left\{\frac{2|f(u)-f(v)|}{|u-v|}:\ u+v=2x,\ u\ne v\right\}. \end{equation*}
so both are governed by the sum set of \(\{f\ne0\}\). By Exercise 1.12.67 there is a bounded null \(E_0\) with \(E_0+E_0\) nonmeasurable; affine images \(aE_0+b\) (\(a\ne0\)) are null with \((aE_0+b)+(aE_0+b)=a(E_0+E_0)+2b\) nonmeasurable, so for suitable \(a,b\) we get
\begin{equation*} E\subset\Bigl[\tfrac12-\delta,\ \tfrac12+\delta\Bigr],\qquad \delta:=\tfrac{1}{100}, \end{equation*}
with \(\lambda(E)=0\) and \(E+E\) nonmeasurable. Define
\begin{equation*} f(x)=1 \text{ if } x\in E,\qquad f(x)=-1 \text{ if } x\in E+2,\qquad f(x)=0 \text{ otherwise}. \end{equation*}
The disjoint sets \(E\), \(E+2\) are null, so \(f=0\) a.e. and \(f\) is measurable by completeness.
Since \(|f|\le1\), \(\varphi_*\le2\), with \(\varphi_*(x)=2\) iff some \(u+v=2x\) has one of \(u,v\) in \(E\) and the other in \(E+2\), i.e. \(2x-2\in E+E\). Hence
\begin{equation*} \{x:\ \varphi_*(x)>3/2\}=\{x:\ \varphi_*(x)=2\}=\tfrac12(E+E)+1 , \end{equation*}
nonmeasurable, so \(\varphi_*\) is not measurable.
For \(\Phi_*\), set \(I=[3/2-\delta,3/2+\delta]\supset\tfrac12(E+E)+1\) (as \(E+E\subset[1-2\delta,1+2\delta]\)). For \(x\in I\) and \(u+v=2x\) with \(f(u)\ne f(v)\), one of \(u,v\), say \(u\), lies in \(E\cup(E+2)\) and
\begin{equation*} |t|=\frac{|u-v|}{2}=|u-x|\ \ge\ 1-2\delta . \end{equation*}
Two cases (otherwise \(f(u)=f(v)\) and the quotient is \(0\)):
(a) exactly one of \(u,v\) is in \(E\cup(E+2)\): then \(|f(u)-f(v)|=1\) and the quotient is \(1/|t|\le1/(1-2\delta)\);
(b) one is in \(E\) and the other in \(E+2\): then \(|f(u)-f(v)|=2\) and \(|t|\le1+2\delta\), so the quotient is \(\ge2/(1+2\delta)\).
Case (b) occurs for some admissible pair iff \(2x-2\in E+E\), and with \(\delta=1/100\), \(1/(1-2\delta)=50/49<3/2<100/51=2/(1+2\delta)\). So for \(x\in I\),
\begin{equation*} \Phi_*(x)>\tfrac32\iff 2x-2\in E+E , \end{equation*}
so that
\begin{equation*} \{x:\ \Phi_*(x)>3/2\}\cap I=\tfrac12(E+E)+1 , \end{equation*}
whose right-hand side is nonmeasurable; hence \(\Phi_*\) is not measurable either.
(N.N. Lusin, D.E. Menchoff) Let \(E\subset\mathbb{R}^n\) be a set of finite measure and let \(K\subset E\) be a compact set such that \(E\) has density \(1\) at every point of \(K\). Prove that there exists a compact set \(P\) without isolated points such that \(K\subset P\subset E\) and \(P\) has density \(1\) at every point of \(K\).
Take \(P=K\cup\bigcup_jP_j\), where the \(P_j\) are perfect kernels of compact subsets of \(E\) filling the Whitney cubes of \(\mathbb{R}^n\setminus K\) to within \(\varepsilon_j\). Assume \(K\ne\emptyset\) (else \(P=\emptyset\)), and recall that \(A\) has density \(1\) at \(x\) iff \(\lambda(B(x,r)\setminus A)=o(r^n)\).
Let \(\mathcal{W}=\{Q_j\}\) be the dyadic cubes maximal with respect to \(\mathrm{diam}\,Q\le\mathrm{dist}(Q,K)\). Every \(x\in\mathbb{R}^n\setminus K\) lies in one: with \(d=\mathrm{dist}(x,K)>0\) the property holds for the generation-\(k\) cube through \(x\) once \(2\sqrt n\,2^{-k}\le d\) and fails for large cubes, and the cubes through \(x\) form a chain, so a largest such cube exists and is maximal. No \(Q_j\) meets \(K\), and distinct maximal cubes are disjoint, so \(\mathcal{W}\) partitions \(\mathbb{R}^n\setminus K\), with
\begin{equation*} \mathrm{diam}\,Q_j\le \mathrm{dist}(Q_j,K)\le 4\,\mathrm{diam}\,Q_j , \end{equation*}
the right inequality since the parent fails the property: \(\mathrm{dist}(Q_j,K)\le\mathrm{dist}(\widehat{Q_j},K)+\mathrm{diam}\,\widehat{Q_j}<2\,\mathrm{diam}\,\widehat{Q_j}=4\,\mathrm{diam}\,Q_j\).
Put \(\varepsilon_j:=\min\{\mathrm{diam}\,Q_j,1\}\lambda(Q_j)\), and \(P_j:=\emptyset\) when \(\mathrm{dist}(Q_j,K)\ge1\). Otherwise inner regularity gives a compact \(F_j\subset E\cap Q_j\) with
\begin{equation*} \lambda\bigl((E\cap Q_j)\setminus F_j\bigr)<\varepsilon_j , \end{equation*}
and \(P_j\) the perfect kernel in the Cantor–Bendixson decomposition \(F_j=P_j\cup C_j\) (\(P_j\) closed without isolated points, \(C_j\) countable). Since \(\lambda(C_j)=0\),
\begin{equation*} \lambda\bigl((E\cap Q_j)\setminus P_j\bigr)<\varepsilon_j\qquad\text{for all such } j. \tag{1} \end{equation*}
Then \(K\subset P\subset E\).
Compactness. \(P\) is bounded, since \(P_j\ne\emptyset\) forces \(\mathrm{dist}(Q_j,K)<1\) and \(\mathrm{diam}\,Q_j\le1\), putting \(Q_j\) in the \(2\)-neighbourhood of the bounded \(K\). For closedness, let \(p\) be a limit point of \(P\) with \(p\notin K\), \(d:=\mathrm{dist}(p,K)>0\), and let \(z\in Q_j\cap B(p,d/2)\). Then \(\mathrm{dist}(z,K)\ge d-|z-p|>d/2\), while \(\mathrm{dist}(z,K)\le \mathrm{dist}(Q_j,K)+\mathrm{diam}\,Q_j\le 2\,\mathrm{dist}(Q_j,K)\); hence \(\mathrm{dist}(Q_j,K)>d/4\) and therefore \(\mathrm{diam}\,Q_j\ge \mathrm{dist}(Q_j,K)/4>d/16\). On the other hand \(\mathrm{diam}\,Q_j\le \mathrm{dist}(Q_j,K)\le \mathrm{dist}(z,K)\le |z-p|+d<3d/2\), so \(Q_j\subset B(p,2d)\). Being disjoint, of volume \(\ge(d/(16\sqrt n))^n\), and contained in \(B(p,2d)\), only finitely many such \(Q_j\) exist; as \(K\cap B(p,d/2)=\emptyset\), \(P\cap B(p,d/2)\) lies in a finite union of compact \(P_j\), so \(p\in P\).
Density. Fix \(x\in K\) and \(0<r<1\). Since \(K\subset P\) and the \(Q_j\) cover \(\mathbb{R}^n\setminus K\),
\begin{equation*} B(x,r)\setminus P\ \subset\ \bigcup_{j:\,Q_j\cap B(x,r)\ne\emptyset}\bigl(Q_j\setminus P_j\bigr) \ \subset\ \bigcup_{j:\,Q_j\cap B(x,r)\ne\emptyset}\Bigl[(Q_j\setminus E)\cup\bigl((E\cap Q_j)\setminus P_j\bigr)\Bigr]. \end{equation*}
If \(Q_j\) meets \(B(x,r)\) then \(\mathrm{dist}(Q_j,K)<r<1\), so (1) applies and \(\mathrm{diam}\,Q_j<r\), giving \(Q_j\subset B(x,2r)\). Disjointness and (1) then yield
\begin{equation*} \lambda\bigl(B(x,r)\setminus P\bigr)\ \le\ \sum_{j:\,Q_j\subset B(x,2r)}\lambda(Q_j\setminus E)\ +\ \sum_{j:\,Q_j\subset B(x,2r)}\varepsilon_j \ \le\ \lambda\bigl(B(x,2r)\setminus E\bigr)\ +\ r\,\lambda\bigl(B(x,2r)\bigr), \end{equation*}
using \(\varepsilon_j\le\mathrm{diam}(Q_j)\lambda(Q_j)\le r\lambda(Q_j)\). The first term is \(o(r^n)\) because \(E\) has density \(1\) at \(x\), and the second is \(2^n\omega_nr^{n+1}=o(r^n)\); so \(P\) has density \(1\) at \(x\).
No isolated points: each \(P_j\) has none, and for \(x\in K\) the density just proved gives \(\lambda(P\cap B(x,r))>0\) for small \(r\), so \(B(x,r)\) meets \(P\setminus\{x\}\).
Let \(f\) be a function of bounded variation on \([a,b]\) such that
\begin{equation*} V(f,[a,b])=\int_a^b |f^{\prime}(t)|\,dt . \end{equation*}
Prove that \(f\) is absolutely continuous.
The hypothesis forces \(V(f,[a,x])=\int_a^x|f^{\prime}|\,dt\) for every \(x\), and an indefinite integral is absolutely continuous. Recall \(f^{\prime}\) exists a.e. (Theorem 5.2.6) and is measurable.
Lemma. \(\int_c^d|f^{\prime}(t)|\,dt\le V(f,[c,d])\) for every \([c,d]\subset[a,b]\). Indeed, \(W(x):=V(f,[c,x])\) is nondecreasing (Proposition 5.2.2(i)) and by additivity (5.2.2), for \(c\le x<y\le d\),
\begin{equation*} |f(y)-f(x)|\le V(f,[x,y])=W(y)-W(x), \end{equation*}
so dividing by \(y-x\) gives \(|f^{\prime}|\le W^{\prime}\) at every point where both exist, i.e. a.e. (Theorem 5.2.6 for \(f\) and \(W\)); by Corollary 5.2.7, \(W^{\prime}\) is integrable with \(\int_c^dW^{\prime}\le W(d)-W( c)=V(f,[c,d])\). \(\square\)
Fix \(c\in[a,b]\). By the Lemma,
\begin{equation*} \int_a^c|f^{\prime}(t)|\,dt\le V(f,[a,c]),\qquad \int_c^b|f^{\prime}(t)|\,dt\le V(f,[c,b]), \end{equation*}
all four quantities being finite. Adding, and using additivity of the variation (Proposition 5.2.2(iii)) and the hypothesis,
\begin{equation*} \int_a^b|f^{\prime}(t)|\,dt=\int_a^c|f^{\prime}|\,dt+\int_c^b|f^{\prime}|\,dt\le V(f,[a,c])+V(f,[c,b])=V(f,[a,b])=\int_a^b|f^{\prime}(t)|\,dt , \end{equation*}
so both inequalities, being between finite numbers, are equalities; taking \(c=x\) gives \(V(x):=V(f,[a,x])=\int_a^x|f^{\prime}|\,dt\) for every \(x\), which is absolutely continuous as an indefinite integral of the integrable \(|f^{\prime}|\) (Theorem 5.3.6). Given \(\varepsilon>0\), take \(\delta>0\) from the absolute continuity of \(V\); then for disjoint \((a_i,b_i)\) with \(\sum_i(b_i-a_i)<\delta\),
\begin{equation*} \sum_{i=1}^m|f(b_i)-f(a_i)|\le \sum_{i=1}^m V(f,[a_i,b_i])=\sum_{i=1}^m\bigl(V(b_i)-V(a_i)\bigr)<\varepsilon . \end{equation*}
(i) (Lusin) Prove that there exists no continuous function \(f\) on \([0,1]\) such that \(f^{\prime}(x)=+\infty\) on a set of positive measure.
(ii) Deduce from Theorem 5.8.12 that there is no function with the property mentioned in (i).
(i) \(A:=\{x\in[0,1]:f^{\prime}(x)=+\infty\}\) has outer measure zero for an arbitrary real \(f\), continuity playing no role; here \(f^{\prime}(x)=+\infty\) means \((f(y)-f(x))/(y-x)\to+\infty\) as \(y\to x\). Suppose \(\lambda^*(A)>0\) and, for \(k\in\mathbb{N}\), set
\begin{equation*} S_k=\Bigl\{x\in[0,1]:\ \frac{f(y)-f(x)}{y-x}\ge 1\ \text{ for all } y\in[0,1] \text{ with } 0<|y-x|\le 1/k\Bigr\}. \end{equation*}
Each \(x\in A\) lies in \(S_k\) for large \(k\), so \(A=\bigcup_k(A\cap S_k)\) increasingly; outer measure is continuous along increasing sequences of arbitrary sets (pass to measurable hulls \(H_k\), replaced by \(\bigcap_{m\ge k}H_m\)), so \(\lambda^*(A\cap S_k)>0\) for some \(k\). Splitting \([0,1]\) into finitely many intervals shorter than \(1/k\), subadditivity yields such an interval \(J\) with
\begin{equation*} F:=A\cap S_k\cap J,\qquad \lambda^*(F)>0 . \end{equation*}
Then \(f\) is increasing on \(F\): for \(x<x^{\prime}\) in \(F\) we have \(0<x^{\prime}-x<1/k\), so the defining property of \(S_k\) at \(x\) with \(y=x^{\prime}\) gives
\begin{equation*} f(x^{\prime})-f(x)\ge x^{\prime}-x>0. \tag{2} \end{equation*}
Let \(\alpha=\inf F<\beta=\sup F\) (strict since \(\lambda^*(F)>0\)) and \(F_0:=F\cap(\alpha,\beta)\), so \(\lambda^*(F_0)>0\). Define
\begin{equation*} h(x)=\sup\{f(t):\ t\in F,\ t\le x\},\qquad x\in(\alpha,\beta). \end{equation*}
The set on the right is nonempty (\(\alpha<x\)) and \(f\) is bounded on it by \(f(t^{\prime})\) for any \(t^{\prime}\in F\cap(x,\beta)\), using (2); so \(h\) is real-valued nondecreasing on \((\alpha,\beta)\) and, again by (2), \(h=f\) on \(F_0\).
Let \(N_2\subset(\alpha,\beta)\) be the null set where the nondecreasing \(h\) has no finite derivative (Corollary 5.2.7 on compact subintervals) and \(N_1\) the at most countable set of isolated points of \(F\). Since \(\lambda^*(F_0)>0\), pick
\begin{equation*} x\in F_0\setminus(N_1\cup N_2). \end{equation*}
Being non-isolated in \(F\), it admits \(y_m\in F_0\), \(y_m\ne x\), \(y_m\to x\) (discard finitely many), and there \(f=h\), so
\begin{equation*} \frac{f(y_m)-f(x)}{y_m-x}=\frac{h(y_m)-h(x)}{y_m-x}\ \longrightarrow\ h^{\prime}(x)<+\infty . \end{equation*}
contradicting \(x\in A\). Hence \(\lambda^*(A)=0\), and in particular no continuous \(f\) has \(f^{\prime}=+\infty\) on a set of positive measure.
(ii) At \(x\in A\) all four Dini derivates equal \(+\infty\):
\begin{equation*} \overline{D}^+f(x)=\underline{D}^+f(x)=\overline{D}^-f(x)=\underline{D}^-f(x)=+\infty . \end{equation*}
None of the four alternatives of Theorem 5.8.12 is then possible: (a) needs \(f^{\prime}(x)\) finite, (b) and (c) leave two derivates finite, and (d) needs \(\underline{D}^{\pm}f(x)=-\infty\). Since the alternatives hold a.e., \(A\) lies in a null set, so \(\lambda(A)=0\) by completeness.
(i) (Lusin) Prove that there exists a continuous function \(f\) on \([0,1]\) such that \(f^{\prime}(x)\) exists a.e. and \(f^{\prime}(x)>1\) a.e., but in no interval is \(f\) increasing.
(ii) (Zahorski) Suppose that a set \(E\subset\mathbb{R}^1\) is the countable union of compact sets and every point of \(E\) is its density point. Prove that there exists an approximately continuous function \(\varphi\) such that \(0<\varphi(x)\le 1\) if \(x\in E\) and \(\varphi(x)=0\) if \(x\notin E\).
(iii) Show that there exists an everywhere differentiable function \(f\) on the real line such that \(f^{\prime}\) is discontinuous almost everywhere.
(iv) Prove that there exists a differentiable function \(f\) on \([0,1]\) with a bounded derivative such that on no interval is \(f\) monotone.
(i) Take \(f\) with \(f^{\prime}=g\) a.e. for a finite measurable \(g>1\) integrable on no interval. With \(\{q_n\}\) enumerating the rationals of \([0,1]\), put
\begin{equation*} G(x)=\sum_{n=1}^\infty \frac{4^{-n}}{|x-q_n|}\qquad (x\ne q_n\ \text{for all }n), \end{equation*}
measurable with values in \([0,+\infty]\). The set
\begin{equation*} E_n=\{x:\ 4^{-n}|x-q_n|^{-1}>2^{-n}\}=\{x:\ |x-q_n|<2^{-n}\} \end{equation*}
has measure \(2^{1-n}\), summable, so Borel–Cantelli puts a.e. \(x\) in only finitely many \(E_n\); there \(4^{-n}|x-q_n|^{-1}\le2^{-n}\) eventually and the series converges, i.e. \(G<\infty\) a.e. Set \(g:=1+G\) where \(G<\infty\) and \(g:=2\) elsewhere, so \(g>1\) is finite and measurable. Every nondegenerate \(I\subset[0,1]\) contains some \(q_n\) in its interior, so
\begin{equation*} \int_I g\,dx\ \ge\ 4^{-n}\int_I\frac{dx}{|x-q_n|}=+\infty . \end{equation*}
so \(g\) is integrable on no interval. By Theorem 5.1.4 there is a continuous \(f\) on \([0,1]\), differentiable a.e., with \(f^{\prime}=g>1\) a.e. Were \(f\) nondecreasing on some \([c,d]\), Corollary 5.2.7 would make \(f^{\prime}=g\) integrable there, a contradiction.
(ii) We use the criterion: \(\varphi\) is approximately continuous at \(x\) iff \(\{|\varphi-\varphi(x)|<\varepsilon\}\) has density \(1\) at \(x\) for every \(\varepsilon>0\). (Sufficiency: with \(A_n=\{|\varphi-\varphi(x)|<1/n\}\) and \(r_n\downarrow0\) chosen so that \(\lambda(B(x,r)\setminus A_n)<n^{-1}\lambda(B(x,r))\) for \(r\le r_n\), the set \(\bigcup_n(A_n\cap\{r_{n+1}<|y-x|\le r_n\})\cup\{x\}\) has density \(1\) at \(x\) and carries \(\varphi\) continuously.) Write \(A\prec B\) when \(A\subset B\) and \(B\) has density \(1\) at every point of \(A\); nested instances compose.
Lemma 1. If \(A\) is compact, \(B\) bounded measurable and \(A\prec B\), there is a compact \(C\) with \(A\prec C\prec B\). Proof. Let \(B_1\) be the set of density points of \(B\) lying in \(B\), so \(\lambda(B\setminus B_1)=0\) (Lebesgue) and \(A\subset B_1\), with \(B_1\) of finite measure having density \(1\) at every point of \(A\). Exercise 5.8.116 applied to \(B_1\supset A\) gives a compact \(C\), \(A\subset C\subset B_1\), with \(A\prec C\); and \(C\subset B_1\) says exactly \(C\prec B\). \(\square\)
Lemma 2. If \(K\) is compact, \(E\) measurable and \(K\prec E\), there is an everywhere approximately continuous \(u:\mathbb{R}\to[0,1]\) with \(u=1\) on \(K\) and \(u=0\) off \(E\).
Proof. We may assume \(E\) bounded, replacing it by \(E\cap U\) for a bounded open \(U\supset K\). With \(D\) the dyadic rationals of \([0,1]\), construct compact \(P_r\), \(r\in D\), with
\begin{equation*} P_1=K,\qquad P_s\prec P_r\ \text{ whenever } r<s,\qquad P_0\subset E . \end{equation*}
Lemma 1 with \(A=K\), \(B=E\) gives \(P_0\) with \(K\prec P_0\prec E\); at stage \(n\), for a new \(m=(2k+1)2^{-n-1}\) between \(r=k2^{-n}\) and \(s=(k+1)2^{-n}\), Lemma 1 applied to \(P_s\prec P_r\) gives a compact \(P_m\) with \(P_s\prec P_m\prec P_r\), and composition restores \(P_s\prec P_r\) for all dyadic \(r<s\) so far. Define
\begin{equation*} u(x)=\sup\{r\in D:\ x\in P_r\},\qquad u(x)=0 \text{ if } x\notin P_r \text{ for all } r . \end{equation*}
The \(P_r\) decrease in \(r\), so \(0\le u\le1\), \(u=1\) on \(P_1=K\) and \(u=0\) off \(P_0\subset E\). Fix \(x\), put \(c=u(x)\), \(\varepsilon>0\).
(a) If \(c<1\), take \(s\in D\) with \(c<s<c+\varepsilon\): then \(x\notin P_s\), which is closed, so \(u\le s<c+\varepsilon\) on some \(B(x,\delta)\) (as \(u(y)>s\) would force \(y\in P_t\subset P_s\) for some \(t>s\)).
(b) If \(c>0\), take \(r^{\prime}\in D\) with \(c-\varepsilon<r^{\prime}<c\): since \(c\) is a supremum, \(x\in P_t\) for some \(t>r^{\prime}\), so \(P_t\prec P_{r^{\prime}}\) makes \(P_{r^{\prime}}\) of density \(1\) at \(x\), and \(u\ge r^{\prime}>c-\varepsilon\) there.
So \(\{|u-u(x)|<\varepsilon\}\) contains \(P_{r^{\prime}}\cap B(x,\delta)\) (or a ball, or \(P_{r^{\prime}}\), in the degenerate cases \(c=1\), \(c=0\)), a set of density \(1\) at \(x\). \(\square\)
Now let \(E=\bigcup_mK_m\) with \(K_m\) compact and every point of \(E\) a density point of \(E\), so \(K_m\prec E\); Lemma 2 gives approximately continuous \(u_m:\mathbb{R}\to[0,1]\), \(u_m=1\) on \(K_m\), \(u_m=0\) off \(E\). Put
\begin{equation*} \varphi=\sum_{m=1}^\infty 2^{-m}u_m . \end{equation*}
The series converges uniformly and \(0\le\varphi\le1\). Approximate continuity survives finite sums (two sets of density \(1\) at \(x\) intersect in one) and uniform limits (if \(\sup|\varphi-\varphi_j|<\varepsilon/3\) then \(\{|\varphi_j-\varphi_j(x)|<\varepsilon/3\}\subset\{|\varphi-\varphi(x)|<\varepsilon\}\)), so \(\varphi\) is approximately continuous everywhere. It vanishes off \(E\), and \(x\in K_m\) gives \(\varphi(x)\ge2^{-m}>0\).
Lemma 3. If \(|\varphi|\le M\) is measurable, \(F(x)=\int_0^x\varphi\,dt\), and \(\varphi\) is approximately continuous at \(x\), then \(F^{\prime}(x)=\varphi(x)\). Proof. With \(A=\{|\varphi-\varphi(x)|<\varepsilon\}\) of density \(1\) at \(x\), for \(h>0\),
\begin{equation*} \left|\frac{F(x+h)-F(x)}{h}-\varphi(x)\right| =\frac1h\left|\int_x^{x+h}\bigl(\varphi(t)-\varphi(x)\bigr)dt\right| \le \varepsilon+\frac{2M\,\lambda\bigl([x,x+h]\setminus A\bigr)}{h}, \end{equation*}
the last fraction tending to \(0\) since \(A\) has density \(1\) at \(x\) (similarly for \(h<0\)); as \(\varepsilon\) was arbitrary, \(F^{\prime}(x)=\varphi(x)\). \(\square\)
(iii) Take \(f(x)=\int_0^x\varphi\), for \(\varphi\) from (ii) applied to a \(\sigma\)-compact \(E\) with dense null complement. For \(k\in\mathbb{Z}\), \(j\in\mathbb{N}\) let \(C_{k,j}\subset[k,k+1]\) be nowhere dense compact (a fat Cantor set, Example 1.7.6) with \(\lambda([k,k+1]\setminus C_{k,j})<2^{-j}\), and put
\begin{equation*} E=\bigcup_{k\in\mathbb{Z}}\bigcup_{j\in\mathbb{N}}C_{k,j}. \end{equation*}
Then \(\lambda(\mathbb{R}\setminus E)=0\), so \(E\) has density \(1\) at every point; and \(E\) is meagre, so its complement is dense (Baire). By (ii) there is an approximately continuous \(\varphi\), \(0<\varphi\le1\) on \(E\) and \(\varphi=0\) off \(E\); it is bounded and measurable, so \(f(x)=\int_0^x\varphi\) is everywhere differentiable with \(f^{\prime}=\varphi\) (Lemma 3). At \(x\in E\), \(f^{\prime}(x)>0\) while \(f^{\prime}=0\) at points \(y_n\to x\) outside \(E\), so \(f^{\prime}\) is discontinuous on \(E\), i.e. a.e.
(iv) Take \(f(x)=\int_0^x(\varphi_+-\varphi_-)\) for disjoint dense \(\sigma\)-compact self-dense-point sets \(E^{\pm}\). Enumerate the rational intervals \(\{I_n\}\) and choose inductively disjoint nowhere dense compact \(C_n,D_n\subset I_n\) of positive measure, disjoint from all \(C_m,D_m\), \(m<n\): the closed nowhere dense \(\bigcup_{m<n}(C_m\cup D_m)\) misses some subinterval of \(I_n\), in which two disjoint fat Cantor sets fit. Put
\begin{equation*} B^+=\bigcup_n C_n,\qquad B^-=\bigcup_n D_n . \end{equation*}
disjoint \(\sigma\)-compact sets of positive measure in every interval.
Every bounded measurable \(B\) of positive measure contains a \(\sigma\)-compact \(S\) of positive measure all of whose points are density points of \(S\): pick a density point \(x_0\in B\) (Lebesgue), set \(A_0=\{x_0\}\prec B\), and iterate Lemma 1 to get compact \(A_0\subset A_1\subset\cdots\subset B\) with \(A_k\prec A_{k+1}\prec B\). Then \(S:=\bigcup_kA_k\) has density \(1\) at each of its points (via \(A_{k+1}\subset S\)), and \(\lambda(S)\ge\lambda(A_1)>0\), \(A_1\) having density \(1\) at \(x_0\).
Applying this to \(B^{\pm}\cap I_n\) gives \(\sigma\)-compact \(S^{\pm}_n\subset B^{\pm}\cap I_n\) of that kind. Put
\begin{equation*} E^+=\bigcup_n S^+_n\subset B^+,\qquad E^-=\bigcup_n S^-_n\subset B^- . \end{equation*}
Both are \(\sigma\)-compact, disjoint, dense (each meets every \(I_n\)), and each point of \(E^{\pm}\) is a density point of it, being one of some \(S^{\pm}_n\). By (ii) there are approximately continuous \(\varphi_{\pm}:\mathbb{R}\to[0,1]\) positive exactly on \(E^{\pm}\). Set
\begin{equation*} \varphi=\varphi_+-\varphi_-,\qquad f(x)=\int_0^x\varphi(t)\,dt,\quad x\in[0,1]. \end{equation*}
Then \(|\varphi|\le1\) is approximately continuous everywhere, so \(f\) is differentiable on \([0,1]\) with the bounded derivative \(f^{\prime}=\varphi\) (Lemma 3), and \(\varphi>0\) on \(E^+\), \(\varphi<0\) on \(E^-\) by disjointness. Monotonicity of \(f\) on a nondegenerate \(I\subset[0,1]\) would force \(f^{\prime}\ge0\) (resp. \(\le0\)) on \(I\), contradicted by the points of the dense \(E^-\) (resp. \(E^+\)) in \(I\).
(Hahn, Lusin) Construct two different continuous functions \(F\) and \(G\) on \([0,1]\) such that \(F(0)=G(0)=0\) and \(F^{\prime}(x)=G^{\prime}(x)\) at every point \(x\in[0,1]\), where infinite values of the derivatives are allowed.
Take \(G(x)=\int_0^x w\,dt\) and \(F=G+\theta\), where \(\theta\) is the Cantor function and \(w\ge0\) is integrable but blows up near the Cantor set \(C\) fast enough to force \(G^{\prime}=+\infty\) on \(C\).
Construction of \(w\). Let \(d(t)=\mathrm{dist}(t,C)\), so \(\{d=0\}=C\), and put \(m(\varepsilon)=\lambda\{d<\varepsilon\}\); since \(\{d<\varepsilon\}\downarrow C\) and \(\lambda( C)=0\), we have \(m(\varepsilon)\to0\), so we may choose \(\varepsilon_1>\varepsilon_2>\cdots\to0\) with \(m(\varepsilon_k)\le 2^{-k}/(k+1)\). Let \(\Psi:(0,\infty)\to[1,\infty)\) be continuous and nonincreasing with \(\Psi\equiv1\) on \([\varepsilon_1,\infty)\), \(\Psi(\varepsilon_k)=k\), linear in between; then \(\Psi( r)\to+\infty\) as \(r\to0^+\). Set \(w:=\Psi\circ d\) off \(C\) and \(w:=0\) on \(C\). On \(\{\varepsilon_{k+1}\le d<\varepsilon_k\}\) we have \(w\le\Psi(\varepsilon_{k+1})=k+1\), whence
\begin{equation*} \int_0^1 w\,dt\le 1+\sum_{k\ge1}(k+1)m(\varepsilon_k) \le 1+\sum_{k\ge1}2^{-k}<\infty . \end{equation*}
So \(G\) is absolutely continuous, nondecreasing, \(G(0)=0\) (Theorem 5.3.6).
(i) \(x\notin C\): as \(C\) is closed, \(w=\Psi\circ d\) is continuous near \(x\), so \(G^{\prime}(x)=w(x)<\infty\).
(ii) \(x\in C\): for \(0<|y-x|<\delta\), every \(t\) between \(x\) and \(y\) has \(d(t)\le|t-x|<\delta\), so \(w(t)\ge\Psi(\delta)\) off the null set \(C\), and
\begin{equation*} \frac{G(y)-G(x)}{y-x}=\frac{1}{y-x}\int_x^y w\,dt \ \ge\ \Psi(\delta)\xrightarrow[\delta\to0^+]{}+\infty, \end{equation*}
i.e. \(G^{\prime}(x)=+\infty\).
Now \(\theta\) is continuous and nondecreasing with \(\theta(0)=0\), \(\theta(1)=1\), and is constant on every interval contiguous to \(C\), so \(\theta^{\prime}=0\) off \(C\). Hence \(F\) is continuous, \(F(0)=0\), \(F\ne G\) (since \(F(1)-G(1)=1\)), and \(F^{\prime}(x)=G^{\prime}(x)+0\) for \(x\notin C\), where \(G^{\prime}(x)\) is finite. For \(x\in C\), both \(G\) and \(\theta\) being nondecreasing, the two difference quotients at \(x\) are nonnegative, so
\begin{equation*} \frac{F(y)-F(x)}{y-x}\ \ge\ \frac{G(y)-G(x)}{y-x}\ \xrightarrow[y\to x]{}\ +\infty , \end{equation*}
giving \(F^{\prime}(x)=+\infty=G^{\prime}(x)\).
Exercises 5.8.121–5.8.127
(Tolstoff [952]) Let \(D\) be a bounded region in \(\mathbb{R}^2\) whose boundary \(\partial D\) is a simple piece-wise smooth curve and let \(\varphi\) be a mapping that is continuously differentiable in a neighborhood of the closure of \(D\) and maps \(\partial D\) one-to-one to a contour \(\Gamma\) bounding a region \(G\). Prove that for every bounded measurable function \(f\) on \(G\) one has the equality
\begin{equation*} \int_G f\,dx = k\int_D f\bigl(\varphi(y)\bigr)\det\varphi^{\prime}(y)\,dy , \end{equation*}
where \(k\) is the sign of the integral of \(\det\varphi^{\prime}(y)\) over \(D\) (here \(k\) is automatically nonzero).
The identity is Green’s formula pulled back by \(\varphi\). Write \(\lambda_2\) for planar Lebesgue measure and extend \(f\) by \(0\) off \(G\).
Let \(g\in C^1(\mathbb{R}^2)\) have bounded derivatives and let \(\gamma:[0,1]\to\partial D\) be a positively oriented piecewise smooth simple parametrization. Then
\begin{equation*} \int_D(\partial_{z_1}g)\bigl(\varphi(y)\bigr)\det\varphi^{\prime}(y)\,dy =\int_{\varphi\circ\gamma}g\,dz_2 . \tag{1} \end{equation*}
For \(\varphi\in C^2\) this is Stokes’ formula on \(D\) (whose boundary is a simple piecewise smooth curve) applied to \(\varphi^{*}(g\,dz_2)=(g\circ\varphi)\,d\varphi_2\): equality of the mixed second derivatives of \(\varphi_2\) cancels the two \(\partial_{z_2}g\) terms of \(d(\varphi^{*}\omega)\), leaving \((\partial_{z_1}g)(\varphi)\det\varphi^{\prime}\,dy_1\wedge dy_2\) (Check!). For \(\varphi\) merely \(C^1\), mollify: on a neighbourhood of \(\overline D\) one has \(\varphi_\varepsilon\to\varphi\) and \(\varphi_\varepsilon^{\prime}\to\varphi^{\prime}\) uniformly, so the left side of (1) converges (\(\lambda_2(D)<\infty\)) and so does the right, since on each smooth piece of \(\gamma\) the integrand \(\langle\nabla\varphi_{\varepsilon,2}(\gamma),\gamma^{\prime}\rangle\) converges uniformly with \(|\gamma^{\prime}|\) bounded.
Since \(\varphi|_{\partial D}\) is injective, \(\varphi\circ\gamma\) runs once around the Jordan curve \(\Gamma=\varphi(\partial D)\), rectifiable because \(\varphi\) is Lipschitz on \(\overline D\); let \(k_0\in\{-1,1\}\) be its orientation relative to \(\Gamma^{+}\). Green’s formula for a rectifiable Jordan curve, \(\int_{\Gamma^{+}}g\,dz_2=\int_G\partial_{z_1}g\,dz\), turns (1) into
\begin{equation*} \int_D(\partial_{z_1}g)\bigl(\varphi(y)\bigr)\det\varphi^{\prime}(y)\,dy =k_0\int_G\partial_{z_1}g(z)\,dz . \tag{2} \end{equation*}
Given \(h\in C_c^\infty(\mathbb{R}^2)\), apply (2) to \(g(z)=\int_{-\infty}^{z_1}h(s,z_2)\,ds\), which is smooth with bounded derivatives and \(\partial_{z_1}g=h\):
\begin{equation*} \int_D h\bigl(\varphi(y)\bigr)\det\varphi^{\prime}(y)\,dy=k_0\int_G h(z)\,dz . \tag{3} \end{equation*}
Choosing \(h\) smooth with \(0\le h\le1\) and \(h\equiv1\) on the compact set \(\varphi(\overline D)\cup\overline G\) gives \(\int_D\det\varphi^{\prime}\,dy=k_0\lambda_2(G)\), and \(\lambda_2(G)>0\) since \(G\) is nonempty and open; hence \(\int_D\det\varphi^{\prime}\,dy\ne0\) and \(k=k_0\).
Now let \(\nu:=k\,(\det\varphi^{\prime}\cdot\lambda_2|_D)\circ\varphi^{-1}\), a finite signed Borel measure because \(\det\varphi^{\prime}\) is bounded on \(\overline D\) and \(\lambda_2(D)<\infty\); by the image-measure formula (Theorem 3.6.1),
\begin{equation*} \int_{\mathbb{R}^2}u\,d\nu=k\int_D u\bigl(\varphi(y)\bigr)\det\varphi^{\prime}(y)\,dy \end{equation*}
for bounded Borel \(u\). Multiplying (3) by \(k\) shows \(\int u\,d\nu=\int u\,d(\lambda_2|_G)\) for \(u\in C_c^\infty\), hence for \(u\in C_c\) by uniform approximation, hence on indicators of bounded open sets (monotone approximation from inside, dominated convergence on each part of the Jordan decompositions), hence on all Borel sets by uniqueness of extension for finite measures. Thus \(\nu=\lambda_2|_G\), which is the asserted identity for bounded Borel \(f\).
For bounded Lebesgue measurable \(f\) choose a bounded Borel \(f_0\) with \(\{f\ne f_0\}\subset Z\), \(Z\) Borel and \(\lambda_2(Z)=0\). The left-hand sides agree, and the right-hand sides agree because extending \(\varphi|_{\overline D}\) to a Lipschitz map of \(\mathbb{R}^2\) (possible as \(\varphi\) is \(C^1\) near the compact \(\overline D\)) and applying the area formula, Theorem 5.8.29(i) with \(n=k=2\), yields
\begin{equation*} \int_{D\cap\varphi^{-1}(Z)}|\det\varphi^{\prime}(y)|\,dy =\int_Z\mathrm{Card}\bigl(D\cap\varphi^{-1}(z)\bigr)\,dz=0 . \end{equation*}
(\(\circ\)) (i) Suppose that a function \(f\) is integrable on \([0,1]\) and a function \(\varphi:[0,1]\to[0,1]\) is continuously differentiable. Is it true that the function \(f\bigl(\varphi(x)\bigr)\varphi^{\prime}(x)\) is integrable?
(ii) Let a function \(f\) be integrable on \([a,b]\), let a function \(\varphi:[c,d]\to[a,b]\) be absolutely continuous, and let \(\varphi([c,d])\subset[a,b]\). Suppose, in addition, that the function \(f\bigl(\varphi(x)\bigr)\varphi^{\prime}(x)\) is integrable on \([c,d]\). Prove the equality
\begin{equation*} \int_{\varphi( c)}^{\varphi(d)}f(x)\,dx=\int_c^d f\bigl(\varphi(y)\bigr)\varphi^{\prime}(y)\,dy . \end{equation*}
(i) No. Part (ii) is true in general and is proved first, since the counterexample for (i) rests on it.
(ii) Being absolutely continuous, \(\varphi\) has \(\varphi^{\prime}\in L^1[c,d]\) and \(\varphi(t)-\varphi(s)=\int_s^t\varphi^{\prime}\,du\) (Theorem 5.3.6).
Continuous \(f\): with \(F(x)=\int_a^x f\,dt\), the function \(F\) is \(C^1\), hence Lipschitz with constant \(L=\sup|f|\), so \(F\circ\varphi\) is absolutely continuous, because \(\sum_i|F(\varphi(t_i))-F(\varphi(s_i))|\le L\sum_i|\varphi(t_i)-\varphi(s_i)|\). At a.e. \(y\) the derivative \(\varphi^{\prime}(y)\) exists and \(F\) is differentiable at \(\varphi(y)\), so the chain rule gives \((F\circ\varphi)^{\prime}(y)=f(\varphi(y))\varphi^{\prime}(y)\), and Newton–Leibniz for absolutely continuous functions yields
\begin{equation*} \int_c^d f\bigl(\varphi(y)\bigr)\varphi^{\prime}(y)\,dy =F\bigl(\varphi(d)\bigr)-F\bigl(\varphi( c)\bigr) =\int_{\varphi( c)}^{\varphi(d)}f(x)\,dx . \tag{1} \end{equation*}
Bounded Borel \(f\): let \(\mu:=\varphi^{\prime}\cdot\lambda\) on \([c,d]\), a finite signed Borel measure, so that by Theorem 3.6.1
\begin{equation*} \int_{[a,b]}u\,d(\mu\circ\varphi^{-1}) =\int_c^d u\bigl(\varphi(y)\bigr)\varphi^{\prime}(y)\,dy \end{equation*}
for bounded Borel \(u\), and let \(\sigma\) be \(\pm\lambda\) restricted to the interval with endpoints \(\varphi( c),\varphi(d)\), signed so that \(\int u\,d\sigma=\int_{\varphi( c)}^{\varphi(d)}u\,dx\). By the continuous case these two finite signed Borel measures agree on \(C[a,b]\), hence on indicators of open subintervals (monotone approximation from below and dominated convergence applied to each part of the Jordan decompositions), hence everywhere, the intervals forming a \(\pi\)-system generating the Borel \(\sigma\)-algebra. So (1) holds for bounded Borel \(f\).
Bounded measurable \(f\): choose bounded Borel \(f_0\) with \(f=f_0\) off a Borel set \(E\) of measure zero; the left sides agree. Put \(S:=\{y:\varphi^{\prime}(y)\text{ exists, is finite and nonzero}\}\). By Lemma 5.8.13, \(\lambda(\varphi^{-1}(E)\cap S)=0\), so \(f(\varphi)\varphi^{\prime}=f_0(\varphi)\varphi^{\prime}\) a.e. on \(S\), while for a.e. \(y\notin S\) one has \(\varphi^{\prime}(y)=0\) and both products vanish; so the right sides agree too.
General \(f\): modifying \(f\) on a null set (which by the previous paragraph changes neither side) we may assume \(f\) finite everywhere, and set \(f_N:=\max(-N,\min(f,N))\). Each \(f_N\) is bounded measurable, so (1) holds for \(f_N\); as \(N\to\infty\) we have \(|f_N|\le|f|\in L^1[a,b]\) and \(|f_N(\varphi)\varphi^{\prime}|\le|f(\varphi)\varphi^{\prime}|\in L^1[c,d]\) by hypothesis, so dominated convergence on both sides gives (1) for \(f\).
(i) The printed hint prescribes peak heights \(\varphi(c_n)=n^{-2}\) on \([(n+1)^{-1},n^{-1}]\), which makes the slopes of order \(n^{-2}/\bigl(n(n+1)\bigr)^{-1}\to1\) while \(\varphi^{\prime}(0)=0\), so that \(\varphi\) is differentiable but not \(C^1\); we damp the heights by a logarithm. Take \(f(x)=x^{-1/2}\) on \((0,1]\), \(f(0)=0\), so \(f\in L^1[0,1]\). Fix \(\theta\in C^\infty[0,1]\) with \(0\le\theta\le1\), \(\theta(0)=\theta(1)=0\), \(\theta(1/2)=1\), strictly increasing then strictly decreasing, and with all derivatives vanishing at \(0\) and \(1\) (build it from \(e^{-1/t}\)); let \(M=\sup|\theta^{\prime}|\). For \(n\ge2\) put \(I_n=[1/(n+1),1/n]\), \(\ell_n=1/\bigl(n(n+1)\bigr)\), \(c_n\) the midpoint of \(I_n\), \(a_n=1/\bigl(n^2\log(n+1)\bigr)\), and
\begin{equation*} \varphi(x)=a_n\,\theta\Bigl(\frac{x-1/(n+1)}{\ell_n}\Bigr)\ \ (x\in I_n),\qquad \varphi=0\ \text{ on }\ \{0\}\cup[1/2,1]. \end{equation*}
Then \(0\le\varphi\le a_2<1\), so \(\varphi\) maps \([0,1]\) into \([0,1]\), and \(\varphi\) is \(C^\infty\) on \((0,1]\), all one-sided derivatives at the junctions \(1/n\) vanishing. At \(0\), for \(x\in I_n\),
\begin{equation*} \begin{aligned} \Bigl|\frac{\varphi(x)-\varphi(0)}{x}\Bigr| &\le\frac{a_n}{1/(n+1)}=\frac{n+1}{n^2\log(n+1)}\to0,\\ \sup_{I_n}|\varphi^{\prime}|&\le\frac{M a_n}{\ell_n}=\frac{M(n+1)}{n\log(n+1)}\to0, \end{aligned} \end{equation*}
so \(\varphi^{\prime}(0)=0\) and \(\varphi^{\prime}(x)\to0\) as \(x\to0^+\), i.e. \(\varphi\in C^1[0,1]\). On \(J_n:=[1/(n+1),c_n]\) the function \(\varphi\) increases from \(0\) to \(a_n\), so (ii) applied on \(J_n\) to the bounded truncations \(f_N=\min(f,N)\), followed by monotone convergence (all integrands are nonnegative there, \(\varphi^{\prime}\ge0\)), gives \(\int_{J_n}f(\varphi)\varphi^{\prime}\,dx=\int_0^{a_n}u^{-1/2}\,du=2\sqrt{a_n}\). Hence
\begin{equation*} \int_0^1\bigl|f(\varphi(x))\varphi^{\prime}(x)\bigr|\,dx \ \ge\ \sum_{n\ge2}2\sqrt{a_n} =\sum_{n\ge2}\frac{2}{n\sqrt{\log(n+1)}}=+\infty \end{equation*}
by comparison with the divergent series \(\sum 1/(n\log n)\).
(i) Let \(E\subset\mathbb{R}^1\) be a set of positive Lebesgue measure. Prove that the set \(\mathbb{R}^1\setminus\bigcup_{n=1}^\infty(E+r_n)\), where \(\{r_n\}=\mathbb{Q}\), has measure zero.
(ii) Let \(A\subset\mathbb{R}^1\) be a set of positive outer measure and let \(B\) be an everywhere dense set in \(\mathbb{R}^1\). Prove that for every interval \(I\) one has \(\lambda^*\bigl((A+B)\cap I\bigr)=\lambda(I)\), where \(\lambda\) is Lebesgue measure.
(iii) Suppose we are given two sets \(A,B\subset\mathbb{R}^1\) of positive outer measure. Prove that there exists an interval \(I\) such that \(\lambda^*\bigl((A+B)\cap I\bigr)=\lambda(I)\).
(iv) Construct two sets \(A\) and \(B\) of positive outer measure on the real line such that \(A+B\) contains no open interval.
(v) Suppose we are given two sets \(A\) and \(B\) of positive outer measure on the real line such that at least one of them is measurable. Prove that \(A+B\) contains some open interval.
All five parts run on measurable hulls and two lemmas. A measurable hull of a set \(S\) with \(\lambda^*(S)<\infty\) is a measurable \(H\supset S\) (a \(G_\delta\) will do) with \(\lambda(H)=\lambda^*(S)\); it satisfies
\begin{equation*} \lambda^*(S\cap M)=\lambda(H\cap M)\qquad\text{for every measurable }M. \tag{H} \end{equation*}
Here \(\le\) is clear, and if it were strict for some \(M\), a hull \(H^{\prime}\subset H\cap M\) of \(S\cap M\) would make \(H^{\prime}\cup(H\setminus M)\) a measurable superset of \(S\) of measure \(<\lambda(H)=\lambda^*(S)\). Intersecting with a large interval, any set of positive outer measure may be assumed to have finite positive outer measure.
Lemma A. If \(S,T\) are measurable, \(x_0\) is a density point of \(S\) and \(y_0\) one of \(T\), then \(\lambda((S+c)\cap T)>0\) whenever \(|c-(y_0-x_0)|<\delta\), for some \(\delta>0\). Indeed, by the Lebesgue density theorem pick \(s>0\) with \(\lambda(S\cap(x_0-r,x_0+r))>\frac34\cdot2r\) and likewise for \(T\) at \(y_0\), for all \(0<r\le s\); put \(\delta:=s/100\), \(u:=x_0+c\), \(r:=0.99s\). Then
\begin{equation*} \lambda\bigl((S+c)\cap(u-r,u+r)\bigr) =\lambda\bigl(S\cap(x_0-r,x_0+r)\bigr)>1.5r=1.485s, \end{equation*}
while \((u-r,u+r)\subset(y_0-s,y_0+s)\) and \(\lambda((y_0-s,y_0+s)\setminus T)<0.5s\); were \((S+c)\cap T\) null, this would force \(1.485s\le0.5s\).
Lemma B. For measurable \(S,T\) of finite positive measure the function \(g(x):=\lambda(S\cap(x-T))=(I_S*I_T)(x)\) is continuous with \(\int g\,dx=\lambda(S)\lambda(T)>0\), so \(\{g>0\}\) is open and nonempty: continuity follows from \(|g(x)-g(x^{\prime})|\le\|I_{-T}(\cdot+x-x^{\prime})-I_{-T}\|_{L^1}\) and the continuity of translation in \(L^1\), the integral identity from Fubini’s theorem.
(i) Put \(A:=\bigcup_n(E+r_n)\) and \(B:=\mathbb{R}\setminus A\), both measurable, so that \((E+q)\cap B=\emptyset\) for every rational \(q\). If \(\lambda(B)>0\), shrink \(E\) and \(B\) to finite positive measure and take density points \(x_0\) of \(E\), \(y_0\) of \(B\): Lemma A gives \(\lambda((E+c)\cap B)>0\) for all \(c\) near \(y_0-x_0\), and a rational such \(c\) contradicts the previous sentence. Hence \(\lambda(B)=0\).
(ii) Only \(\ge\) needs proof, and only for bounded \(I\), an unbounded interval being an increasing union of bounded ones. Assume \(0<\lambda^*(A)<\infty\) with hull \(H\), and suppose \(\lambda^*((A+B)\cap I)<\lambda(I)\) for some bounded \(I\). Let \(M\subset I\) be a hull of \((A+B)\cap I\); then \(P:=I\setminus M\) has \(\lambda(P)>0\) and (H) gives \(\lambda^*((A+B)\cap P)=\lambda(M\cap P)=0\). So for each \(b\in B\), the set \((A+b)\cap P\) being null for outer measure, translation invariance of \(\lambda^*\) and (H) yield
\begin{equation*} \lambda\bigl((H+b)\cap P\bigr)=\lambda\bigl(H\cap(P-b)\bigr) =\lambda^*\bigl(A\cap(P-b)\bigr)=0 . \tag{2} \end{equation*}
But with \(x_0\) a density point of \(H\) and \(y_0\) one of \(P\), Lemma A gives \(\delta>0\) with \(\lambda((H+c)\cap P)>0\) for \(|c-(y_0-x_0)|<\delta\), and density of \(B\) lets us take \(c\in B\), contradicting (2).
(iii) Assume \(0<\lambda^*(A),\lambda^*(B)<\infty\), with hulls \(H_A,H_B\). By Lemma B the function \(h:=I_{H_A}*I_{H_B}\) is continuous with \(\int h>0\), so \(W:=\{h>0\}\) is open and nonempty; take a bounded interval \(I\subset W\). If \(\lambda^*((A+B)\cap I)<\lambda(I)\), then exactly as in (ii) there is a measurable \(P\subset I\) with \(\lambda(P)>0\) and \(\lambda^*((A+B)\cap P)=0\), whence, using \((A+b)\cap P\subset(A+B)\cap P\) and (H),
\begin{equation*} g(b):=\lambda\bigl((H_A+b)\cap P\bigr) =\lambda^*\bigl(A\cap(P-b)\bigr)=0\qquad(b\in B). \tag{3} \end{equation*}
Here \(g=I_{-H_A}*I_P\) is continuous by Lemma B, and by Fubini’s theorem
\begin{equation*} \int_{H_B}g(b)\,db=\int_P(I_{H_A}*I_{H_B})(x)\,dx=\int_P h(x)\,dx>0, \end{equation*}
since \(h>0\) on \(I\supset P\) and \(\lambda(P)>0\). Hence \(\lambda(H_B\cap U)>0\) for the open set \(U:=\{g>0\}\), so (H) gives \(\lambda^*(B\cap U)=\lambda(H_B\cap U)>0\), and any \(b\in B\cap U\) contradicts (3).
(iv) Take \(A=B=E\) with \((E+E)\cap\mathbb{Q}=\emptyset\) and \(\lambda^*(E)>0\); every interval contains a rational, so \(E+E\) contains no interval. There are \(\mathfrak{c}\) closed sets of positive measure, each of cardinality \(\mathfrak{c}\) (an uncountable closed set contains a perfect set); enumerate them as \(\{F_\alpha\}_{\alpha<\mathfrak{c}}\) and recursively choose \(x_\alpha\in F_\alpha\) outside
\begin{equation*} \{q-x_\beta:\ \beta<\alpha,\ q\in\mathbb{Q}\}\cup\{q/2:\ q\in\mathbb{Q}\}, \end{equation*}
a set of cardinality \(|\alpha|\cdot\aleph_0+\aleph_0<\mathfrak{c}\). Put \(E:=\{x_\alpha:\alpha<\mathfrak{c}\}\). Then \(2x_\alpha\notin\mathbb{Q}\) and \(x_\alpha+x_\beta\notin\mathbb{Q}\) for \(\beta<\alpha\), so \((E+E)\cap\mathbb{Q}=\emptyset\). And \(\lambda^*(E)>0\): otherwise \(E\subset N\) with \(N\) Borel and null, while by inner regularity \([0,1]\setminus N\) contains a closed set \(F=F_\alpha\) of positive measure, and \(x_\alpha\in E\cap F\) contradicts \(E\subset N\).
(v) Let \(A\) be measurable, so \(\lambda(A)=\lambda^*(A)>0\). Pick measurable \(A_0\subset A\) with \(0<\lambda(A_0)<\infty\), put \(B_0:=B\cap[-N,N]\) with \(0<\lambda^*(B_0)<\infty\), and let \(H\) be a hull of \(B_0\). By Lemma B with \(S=H\), \(T=A_0\), the set \(W:=\{g>0\}\), \(g:=I_H*I_{A_0}\), is open and nonempty, and for \(x\in W\) property (H) applied to the measurable set \(x-A_0\) gives
\begin{equation*} \lambda^*\bigl(B_0\cap(x-A_0)\bigr)=\lambda\bigl(H\cap(x-A_0)\bigr)=g(x)>0 , \end{equation*}
so some \(b\in B_0\cap(x-A_0)\) exists and \(x=(x-b)+b\in A+B\). Thus \(W\subset A+B\), and \(W\) contains an open interval. The case of measurable \(B\) is symmetric.
Let \(E\subset\mathbb{R}^1\) and \(x\in E\). Denote by \(\lambda(E,x,x+h)\) the length of the maximal open interval in \((x,x+h)\) that contains no points of \(E\) (if \(h<0\), then we consider the interval \((x-|h|,x)\)). Let
\begin{equation*} p(E,x)=\limsup_{h\to0}\ \lambda(E,x,x+h)/|h| . \end{equation*}
The set \(E\) is said to be porous at \(x\) if \(p(E,x)<1\). If \(E\) is porous at every point, then we call \(E\) a porous set. Finally, a countable union of porous sets is called \(\sigma\)-porous (this concept is due to E.P. Dolzhenko [230]).
(i) Prove that every porous set has Lebesgue measure zero and is nowhere dense.
(ii) Construct a compact set of measure zero that is not \(\sigma\)-porous.
(iii) Construct a Borel probability measure on the real line that is singular with respect to Lebesgue measure, but vanishes on every \(\sigma\)-porous set.
(iv) Construct a compact set \(K\) on the real line such that every Borel measure on \(K\) is concentrated on a \(\sigma\)-porous set.
Porous sets are null and nowhere dense; items (ii)–(iv) are the theorems of Tkadlec [949], Humke and Preiss [447] and Zajicek [1046] cited in the book’s hint, and we prove here what the tools of this chapter give.
The printed condition “\(E\) is porous at \(x\) if \(p(E,x)<1\)” is a misprint for \(p(E,x)>0\), since otherwise \(E=\mathbb{R}\) would be porous at each of its points and (i) would fail. Note the monotonicity \(P\subset Q\Rightarrow p(P,x)\ge p(Q,x)\): every subset of a porous set is porous.
(i) Nowhere density: if \(\overline{E}\supset(a,b)\), pick \(x\in E\cap(a,b)\); for \(|h|\) small enough every nonempty open subinterval of \((x,x+h)\) meets \(E\), so \(\lambda(E,x,x+h)=0\) and \(p(E,x)=0\), contradicting porosity at \(x\).
Measure zero, in the stronger form \(\lambda^*(E)=0\): as subsets of porous sets are porous and \(E=\bigcup_N(E\cap[-N,N])\), we may assume \(\lambda^*(E)<\infty\). Suppose \(\lambda^*(E)>0\) and let \(H\) be a measurable hull of \(E\) (Exercise 5.8.123), so \(\lambda(H)=\lambda^*(E)>0\) and \(\lambda^*(E\cap M)=\lambda(H\cap M)\) for measurable \(M\). The set \(D\) of density points of \(H\) has \(\lambda(H\setminus D)=0\) by the Lebesgue density theorem, so \(\lambda^*(E\cap D)=\lambda(H\cap D)=\lambda(H)>0\) and there is \(x\in E\cap D\), for which
\begin{equation*} \lim_{r\to0}\frac{\lambda^*\bigl(E\cap(x-r,x+r)\bigr)}{2r} =\lim_{r\to0}\frac{\lambda\bigl(H\cap(x-r,x+r)\bigr)}{2r}=1 . \tag{1} \end{equation*}
But \(c:=p(E,x)>0\) yields \(h_j\to0\), \(h_j\ne0\), with \(\lambda(E,x,x+h_j)>\frac c2|h_j|\), hence open intervals \(J_j\subset(x-r_j,x+r_j)\), \(r_j:=|h_j|\), with \(J_j\cap E=\emptyset\) and \(|J_j|>\frac c2 r_j\); then
\begin{equation*} \lambda^*\bigl(E\cap(x-r_j,x+r_j)\bigr)\le2r_j-|J_j|<2r_j\Bigl(1-\frac c4\Bigr), \end{equation*}
contradicting (1). So \(\lambda^*(E)=0\).
Reduction for (ii)–(iv). Call \(Q\) \(c\)-porous if \(p(Q,x)>c\) for all \(x\in Q\). Since \(P=\bigcup_m P_m\) with \(P_m:=\{x\in P:p(P,x)>1/m\}\) and \(p(P_m,x)\ge p(P,x)>1/m\) there, a set is \(\sigma\)-porous precisely when it is a countable union of \(c_n\)-porous sets, and by monotonicity the pieces may be intersected with any set under consideration.
(ii) Tkadlec [949]. The obvious Baire scheme – a compact null set \(F\) is a Baire space, hence not a countable union of sets nowhere dense in \(F\) – collapses, because porous subsets of \(F\) can be dense in \(F\). Indeed, for infinite compact null \(F\) the countable set \(D\) of endpoints of the complementary intervals is dense in \(F\); each \(x\in D\) is an endpoint of a complementary interval of some length \(g\), so \(\lambda(F,x,x+h)=h\) for \(0<h\le g\) and \(p(F,x)=1\); and an open interval disjoint from \(D\) is disjoint from \(F\) (a point of \(U\cap F\) would be a limit of points of \(D\), which then lie in \(U\)), so \(p(D,x)=p(F,x)=1\). Thus \(D\) is a dense \(1/2\)-porous subset of \(F\). The actual construction uses the abstract porosity machinery of Zajicek [1046].
(iii) Humke and Preiss [447]. Write \(B(x,r):=[x-r,x+r]\).
Density lemma: for a Borel probability measure \(\mu\) and a Borel set \(S\), for \(\mu\)-a.e. \(x\in S\)
\begin{equation*} \lim_{r\to0}\frac{\mu\bigl(S\cap B(x,r)\bigr)}{\mu\bigl(B(x,r)\bigr)}=1 . \tag{2} \end{equation*}
Indeed \(\mathbb{R}\setminus\mathrm{supp}\,\mu\) is a countable union of intervals of measure zero, so \(\mu\)-null, and \(\mu(B(x,r))>0\) for \(x\in\mathrm{supp}\,\mu\); putting \(\nu(A):=\mu(A\setminus S)\) and \(A_c:=\{x\in S\cap\mathrm{supp}\,\mu:\ \overline{D}_\mu\nu(x)\ge c\}\), Lemma 5.8.7(ii) gives \(\nu^*(A_c)\ge c\,\mu^*(A_c)\), while \(A_c\subset S\) and \(\nu(S)=0\) force \(\nu^*(A_c)=0\), hence \(\mu^*(A_c)=0\); take \(c=1/k\), \(k\in\mathbb{N}\).
Mechanism lemma: suppose that for every \(c\in(0,1)\) there are \(\eta( c),r( c)>0\) with
\begin{equation*} \mu(J)\ \ge\ \eta( c)\,\mu\bigl(B(x,r)\bigr) \tag{3} \end{equation*}
whenever \(x\in\mathrm{supp}\,\mu\), \(0<r<r( c)\) and \(J\subset B(x,r)\) is an open interval with \(|J|\ge cr\); then every \(\sigma\)-porous set is \(\mu\)-null. By the reduction and subadditivity of \(\mu^*\) it suffices to treat a \(c\)-porous \(Q\), which by monotonicity and \(\mu(\mathbb{R}\setminus\mathrm{supp}\,\mu)=0\) may be assumed inside \(\mathrm{supp}\,\mu\). If \(\mu^*(Q)>0\), choose Borel \(S\supset Q\) with \(\mu(S)=\mu^*(Q)\); then \(\mu^*(Q\cap A)=\mu(S\cap A)\) for all Borel \(A\), since a strict inequality would give \(\mu(S)=\mu(S\cap A)+\mu(S\setminus A)>\mu^*(Q\cap A)+\mu^*(Q\setminus A)\ge\mu^*(Q)=\mu(S)\). The points of \(S\) where (2) fails lie in a Borel set \(N\) with \(\mu(N)=0\), and \(\mu^*(Q\setminus N)=\mu(S\setminus N)=\mu(S)>0\), so some \(x\in Q\) satisfies (2). Porosity at \(x\) supplies \(r_j\to0\) and open intervals \(J_j\subset B_j:=B(x,r_j)\) with \(J_j\cap Q=\emptyset\) and \(|J_j|>c\,r_j\), so (3) applies for large \(j\) and, as \(Q\cap B_j\subset B_j\setminus J_j\),
\begin{equation*} \mu(S\cap B_j)=\mu^*(Q\cap B_j)\le\mu(B_j)-\mu(J_j)\le\bigl(1-\eta( c)\bigr)\mu(B_j), \end{equation*}
with \(\mu(B_j)>0\), contradicting (2). Hence \(\mu^*(Q)=0\). A measure singular with respect to Lebesgue measure and satisfying (3) for every \(c\in(0,1)\) is constructed in [447], which completes (iii).
(iv) Zajicek [1046]: there is a compact \(K\subset\mathbb{R}\) such that every Borel measure on \(K\) is concentrated on a \(\sigma\)-porous set; in particular a measure as in (iii) carried by \(K\) would satisfy \(\mu(S)=0=\mu(K\setminus S)\), so \(\mu=0\).
(A fully rigorous proof of the remaining step is beyond the scope of this page; see the reference given in the book.)
Prove that every set of positive Lebesgue measure in \(\mathbb{R}^2\) contains the vertices of some equilateral triangle.
Put a density point of \(E\) at the origin: two polar directions \(\pi/3\) apart then share a radius \(t\) with \(E\), and those two points together with the origin are the vertices.
The book’s hint carries three misprints: the rotation is by \(\pi/3\), not \(\pi/6\) (two rays at angle \(\pi/3\) with equal radii \(t\) give third side \(2t\sin(\pi/6)=t\)); the printed chain \(10\pi/11\ge\lambda(\Phi)/2+2(2\pi-\lambda(\Phi))/5\) must run the other way; and “there exists an angle \(\varphi\in E\)” must read \(\varphi\in\Phi\).
By the Lebesgue density theorem \(E\) has a density point, and translating it to the origin and rescaling (both operations preserve equilateral triangles and the ratio below) we may assume \(0\in E\) and, with \(U:=B(0,1)\),
\begin{equation*} \lambda_2(E\cap U)>\frac{10}{11}\,\lambda_2(U)=\frac{10\pi}{11}. \tag{1} \end{equation*}
Polar coordinates \(T(t,\phi)=(t\cos\phi,t\sin\phi)\) are a diffeomorphism of \((0,\infty)\times(0,2\pi)\) onto the plane minus a null ray, so for a.e. \(\phi\) the section \(E_\phi:=\{t\in(0,1]:T(t,\phi)\in E\}\) is measurable, \(F(\phi):=\int_{E_\phi}t\,dt\) is measurable with \(0\le F(\phi)\le\int_0^1t\,dt=\frac12\), and the polar change of variables together with Fubini’s theorem gives \(\lambda_2(E\cap U)=\int_0^{2\pi}F(\phi)\,d\phi\). Put \(\Phi:=\{\phi\in[0,2\pi): F(\phi)\ge\frac25\}\) and split that integral over \(\Phi\) and its complement:
\begin{equation*} \frac{10\pi}{11}<\frac12\lambda(\Phi)+\frac25\bigl(2\pi-\lambda(\Phi)\bigr) =\frac{4\pi}{5}+\frac{1}{10}\lambda(\Phi), \end{equation*}
so that \(\lambda(\Phi)>10\bigl(\frac{10\pi}{11}-\frac{4\pi}{5}\bigr)=\frac{12\pi}{11}>\pi\). On the circle group \(\mathbb{T}=[0,2\pi)\) Lebesgue measure is rotation invariant, so \(\lambda(\Phi-\pi/3)=\lambda(\Phi)>\pi\) and
\begin{equation*} \lambda(\Phi)+\lambda(\Phi-\pi/3)>2\pi\ \ge\ \lambda\bigl(\Phi\cup(\Phi-\pi/3)\bigr), \end{equation*}
whence \(\lambda(\Phi\cap(\Phi-\pi/3))>0\) and there is \(\phi\in\Phi\) with \(\psi:=\phi+\pi/3\pmod{2\pi}\) again in \(\Phi\). If \(\lambda(E_\phi\cap E_\psi)\) were \(0\), then
\begin{equation*} \frac45\le\int_{E_\phi}t\,dt+\int_{E_\psi}t\,dt =\int_{E_\phi\cup E_\psi}t\,dt\le\int_0^1t\,dt=\frac12 , \end{equation*}
which is false; so some \(t>0\) lies in \(E_\phi\cap E_\psi\). Then \(O=(0,0)\), \(P_1=T(t,\phi)\) and \(P_2=T(t,\psi)\) all lie in \(E\), and \(|OP_1|=|OP_2|=t\), \(|P_1P_2|=2t\sin\frac{\pi}{6}=t\).
(\(\circ\)) (Fischer [299]) Let a function \(F\) be continuous on \([0,1]\) and \(F(0)=0\). Prove that \(F\) is the indefinite integral of a function in \(L^2[0,1]\) precisely when the sequence of functions \(n\bigl(F(x+n^{-1})-F(x)\bigr)\) is fundamental in \(L^2[0,1]\), where for \(x>1\) we set \(F(x)=F(1)\).
With \(f_n(x):=n\bigl(F(x+\frac1n)-F(x)\bigr)\), each \(f_n\) is continuous on \([0,1]\), and the two implications are as follows.
Necessity. Let \(F(x)=\int_0^x f\,dt\) with \(f\in L^2[0,1]\), extended by \(0\) to \(\mathbb{R}\); the convention \(F\equiv F(1)\) on \((1,\infty)\) keeps \(F(x)=\int_0^x f\,dt\) for all \(x\ge0\), so the substitution \(t=x+s/n\) gives
\begin{equation*} f_n(x)=n\int_x^{x+1/n}f(t)\,dt=\int_0^1 f\Bigl(x+\frac{s}{n}\Bigr)ds , \end{equation*}
whence, by Minkowski’s inequality for integrals and with \(\tau_hf=f(\cdot+h)\),
\begin{equation*} \|f_n-f\|_{L^2[0,1]}\le\int_0^1\|f(\cdot+s/n)-f\|_{L^2(\mathbb{R})}\,ds \le\sup_{0\le h\le1/n}\|\tau_hf-f\|_{L^2(\mathbb{R})} . \end{equation*}
The right-hand side tends to \(0\) by the continuity of translation in \(L^2(\mathbb{R})\), so \(f_n\to f\) in \(L^2[0,1]\) and \((f_n)\) is fundamental.
Sufficiency. If \((f_n)\) is fundamental, then \(f_n\to f\) in \(L^2[0,1]\) for some \(f\in L^2[0,1]\) by the Riesz–Fischer theorem (Theorem 4.1.4), and since \([0,1]\) has finite measure the Cauchy–Bunyakovskii inequality gives
\begin{equation*} \Bigl|\int_0^t f_n\,dx-\int_0^t f\,dx\Bigr|\le\|f_n-f\|_{L^2[0,1]}\longrightarrow0 \qquad(t\in[0,1]). \end{equation*}
On the other hand the substitution \(y=x+1/n\), legitimate since \(F\) is continuous on \([0,\infty)\), yields
\begin{equation*} \int_0^t f_n(x)\,dx=n\int_{1/n}^{t+1/n}F(y)\,dy-n\int_0^{t}F(x)\,dx =n\int_t^{t+1/n}F-n\int_0^{1/n}F , \end{equation*}
and these are averages of the continuous function \(F\) over intervals shrinking to \(t\) and to \(0\), hence tend to \(F(t)\) and to \(F(0)=0\). Therefore \(F(t)=\int_0^t f(x)\,dx\) on \([0,1]\) with \(f\in L^2[0,1]\).
(Denjoy [214]) Let a function \(f\) be differentiable on \((0,1)\) and let \(\alpha\) and \(\beta\) be such that the set \(\{x:\ \alpha<f^{\prime}(x)<\beta\}\) is nonempty. Prove that this set has positive measure.
This is the Denjoy–Clarkson theorem: the Borel set \(E:=\{x\in(0,1):\alpha<f^{\prime}(x)<\beta\}\) (Borel because \(f^{\prime}\) is of Baire class \(1\) by fact (3) below) cannot be nonempty and null.
Reduction. Fix \(x_0\in E\) and put \(c:=f^{\prime}(x_0)\), \(u(x):=f(x)-cx\), \(\varepsilon:=\min(c-\alpha,\beta-c)>0\). Then \(u^{\prime}=f^{\prime}-c\) and
\begin{equation*} N:=\{x\in(0,1):|u^{\prime}(x)|<\varepsilon\} =\{c-\varepsilon<f^{\prime}<c+\varepsilon\}\subset E \end{equation*}
contains \(x_0\); if \(\lambda(E)=0\) then \(\lambda(N)=0\). So it suffices to show: \(u\) differentiable on \((0,1)\) with \(\lambda(N)=0\) forces \(N=\emptyset\).
Four facts: (1) Darboux’s theorem, so \(u^{\prime}\) maps intervals onto intervals; (2) the mean value theorem; (3) \(u^{\prime}\) is of Baire class \(1\), since on each \([a,b]\subset(0,1)\) it is the pointwise limit of the continuous functions \(x\mapsto n(u(x+1/n)-u(x))\) for large \(n\) and Baire class \(1\) is a local property, whence by the Baire–Osgood theorem \(u^{\prime}|_K\) has a dense set of points of continuity relative to any nonempty \(K\subset(0,1)\) closed in \((0,1)\), such a \(K\) being completely metrizable; (4) Corollary 5.2.7: \(u\) nondecreasing on \([a,b]\) gives \(\int_a^bu^{\prime}\,dx\le u(b)-u(a)\).
Gaps. If \((p,q)\subset(0,1)\) misses \(N\), then \(|u^{\prime}|\ge\varepsilon\) there, so by (1)
\begin{equation*} \text{either}\quad u^{\prime}\ge\varepsilon\ \text{on }(p,q)\quad\text{or}\quad u^{\prime}\le-\varepsilon\ \text{on }(p,q); \tag{G} \end{equation*}
in the first case (2) gives \(u(s)-u(p)\ge\varepsilon(s-p)\) for \(s\in(p,q)\), so letting \(s\downarrow p\) and \(s\uparrow q\) yields \(u^{\prime}(p),u^{\prime}(q)\ge\varepsilon\) (for those endpoints lying in \((0,1)\)), and in the second case \(u^{\prime}(p),u^{\prime}(q)\le-\varepsilon\). Either way the endpoints of an interval free of \(N\) are not in \(N\). (G1)
Assume \(N\ne\emptyset\) and put \(K:=\overline{N}\cap(0,1)\), nonempty and closed in \((0,1)\), with \(N\) dense in \(K\); by (3) choose \(t\in K\) at which \(u^{\prime}|_K\) is continuous relative to \(K\).
(i) \(|u^{\prime}(t)|=\varepsilon\). Taking \(x_k\in N\), \(x_k\to t\), relative continuity gives \(|u^{\prime}(t)|=\lim_k|u^{\prime}(x_k)|\le\varepsilon\). If \(|u^{\prime}(t)|<\varepsilon\), relative continuity supplies an open interval \(V\ni t\), \(V\subset(0,1)\), with \(K\cap V\subset N\); then \(\lambda(K\cap V)=0<\lambda(V)\), so there is \(z\in V\setminus K\), say \(z>t\). Put \(p:=\sup(K\cap[t,z])\); then \(p\in K\cap V\) and \((p,z]\cap K=\emptyset\), so \((p,z)\) is free of \(N\) and (G1) gives \(p\notin N\), contradicting \(K\cap V\subset N\).
Replacing \(u\) by \(-u\), which leaves \(N\) unchanged, we may assume \(u^{\prime}(t)=\varepsilon\).
(ii) \(u^{\prime}>0\) on a neighbourhood of \(t\). Choose \(V=(t-\eta,t+\eta)\subset(0,1)\) with \(u^{\prime}>\varepsilon/2\) on \(K\cap V\). Let \(s\in V\setminus K\) and let \((p,q)\) be the component of \((0,1)\setminus K\) containing \(s\); it misses \(N\), and since \(t\in K\) either \(q\le t\) or \(p\ge t\). If \(q\le t\), then \(q>s>t-\eta\) and \(q<1\), so \(q\in K\cap V\) and \(u^{\prime}(q)>\varepsilon/2>0\), which excludes the second alternative of (G); hence \(u^{\prime}(s)\ge\varepsilon\). If \(p\ge t\), symmetrically \(p\in K\cap V\), \(u^{\prime}(p)>0\) and again \(u^{\prime}(s)\ge\varepsilon\). Thus \(u^{\prime}>0\) on all of \(V\).
(iii) By (ii) and (2), \(u\) is increasing on \(V\), and \(u^{\prime}\ge\varepsilon\) a.e. on \(V\), since off the null set \(N\) one has \(|u^{\prime}|\ge\varepsilon\) together with \(u^{\prime}>0\). So (4) applied on \([a,b]\subset V\) gives \(u(b)-u(a)\ge\int_a^bu^{\prime}\,dx\ge\varepsilon(b-a)\), whence
\begin{equation*} u^{\prime}(x)=\lim_{b\downarrow x}\frac{u(b)-u(x)}{b-x} \ \ge\ \varepsilon\qquad(x\in V), \end{equation*}
i.e. \(N\cap V=\emptyset\), impossible for a neighbourhood of \(t\in\overline{N}\). Hence \(N=\emptyset\), contradicting \(x_0\in N\); therefore \(\lambda(E)>0\).
Exercises 5.8.128–5.8.134
Give an example of a measurable function on \([0,1]\) that has the Darboux property, i.e., on every interval \([a,b] \subset [0,1]\) it assumes all the values between \(f(a)\) and \(f(b)\), but does not have the Denjoy property from the previous exercise, i.e., there exist \(c\) and \(d\) such that the set \(\{x\colon c < f(x) < d\}\) is nonempty and has measure zero.
Take \(f=\psi_k\) on \(C_k\) and \(f\equiv1\) off \(D:=\bigcup_kC_k\), where the \(C_k\) are disjoint affine copies of the Cantor set, one inside each rational subinterval of \([0,1]\), and \(\psi_k:C_k\to[0,1]\) are bijections.
Construction. Enumerate as \(\{J_k\}\) the intervals with rational endpoints in \([0,1]\) and choose \(C_k\) inductively: \(F=C_1\cup\dots\cup C_{k-1}\) is compact and null, hence nowhere dense, so \(J_k\setminus F\) contains a nondegenerate closed interval \([\alpha,\beta]\), and \(C_k:=g_k( C)\) for an affine bijection \(g_k\) of \([0,1]\) onto \([\alpha,\beta]\). The \(C_k\) are then pairwise disjoint compact null sets of cardinality \(\mathfrak{c}\), so the \(\psi_k\) exist and \(\lambda(D)=0\).
Measurability: \(f=1\) off the null set \(D\), so for every \(t\) the set \(\{f>t\}\) differs from \(\emptyset\) or \([0,1]\) by a subset of \(D\), measurable by completeness of Lebesgue measure.
Darboux property: given \(a<b\), pick rationals \(a<p<q<b\) and \(k\) with \(J_k=(p,q)\); then \(C_k\subset(a,b)\) and \(f(C_k)=\psi_k(C_k)=[0,1]\), so \(f\) assumes on \((a,b)\) every value of \([0,1]\), in particular every value between \(f(a)\) and \(f(b)\).
Failure of the Denjoy property: with \(c=1/4\), \(d=1/2\), the set \(\{1/4<f<1/2\}\) is contained in \(D\), hence null, and contains \(\psi_1^{-1}(1/3)\), hence is nonempty.
(Davies [207]) Let a function \(f\) on \([0,1]^2\) be approximately continuous in every variable separately. (i) Prove that \(f\) is Lebesgue measurable. (ii) Prove that \(f\) even belongs to the second Baire class.
Both assertions come from averaging in the second variable; replacing \(f\) by \(\arctan f\) (composition with \(\tan\) restores \(f\) and changes neither measurability nor the Baire class) we may assume \(|f|\le M\).
(P1) A bounded approximately continuous \(g\) on an interval is measurable, by Exercise 5.8.92 (continuity in the density topology), and every point is a Lebesgue point of \(g\) by Exercise 5.8.90(i), so \(g(y)=\lim_{h\to0^+}\frac{1}{2h}\int_{y-h}^{y+h}g(t)\,dt\) at every \(y\).
(P2) Such a \(g\) is of the first Baire class: by (P1) it is the pointwise limit of \(g_n(y)=\frac n2\int_{y-1/n}^{y+1/n}g(t)\,dt\), each Lipschitz with constant \(nM\).
(P3) A uniform limit \(h\) of Baire-\(1\) functions is Baire-\(1\): with \(|h_k-h|\le2^{-k}\) put \(d_1=h_1\), \(d_k=h_k-h_{k-1}\), so \(|d_k|\le3\cdot2^{-k}\) and \(h=\sum_kd_k\); take continuous \(d_{k,j}\to d_k\) pointwise, truncated at \(\pm3\cdot2^{-k}\) (which preserves the convergence), and then \(H_j=\sum_kd_{k,j}\) is continuous and \(H_j\to h\) pointwise by domination.
(i) Every section \(f(\cdot,y)\) and \(f(x,\cdot)\) is bounded and approximately continuous, hence measurable by (P1), and (P1) applied to \(f(x,\cdot)\), with \(f\) extended by \(0\) off the square, gives
\begin{equation*} f(x,y)=\lim_{n\to\infty}A_nf(x,y),\qquad A_nf(x,y):=\frac n2\int_{y-1/n}^{y+1/n}f(x,t)\,dt . \end{equation*}
Since \(|A_nf(x,y)-A_nf(x,y^{\prime})|\le nM|y-y^{\prime}|\) uniformly in \(x\), measurability of each \(x\mapsto A_nf(x,y)\) makes \(A_nf\) a Caratheodory function, hence jointly measurable: the jointly measurable functions \(\sum_j\tau_{j,m}(y)A_nf(x,j/m)\), with \(\tau_{j,m}\) the continuous tent partition of unity on the grid \(\{j/m\}\), converge to \(A_nf\) uniformly, the error being at most \(nM/m\). Thus (i) reduces to
(*) for every interval \(I\) the function \(g_I(x)=\int_If(x,t)\,dt\) is Lebesgue measurable,
which is the hard core of Davies [207] and is taken from there. The natural attempt at () – deduce approximate continuity of \(g_I\) at \(x_0\) from \(|g_I(x)-g_I(x_0)|\le\Delta(x):=\int_I|f(x,t)-f(x_0,t)|\,dt\) by exchanging the order of integration, the inner average over \(x\) tending to \(0\) for each fixed \(t\) by (P1) – presupposes the joint measurability being proved, and the corresponding inequality between iterated upper integrals is false: under the continuum hypothesis Sierpinski’s set \(S\subset[0,1]^2\) has countable horizontal and co-countable vertical sections, so \(F=\mathbb{1}_S\) has \(\int_IF(x,t)\,dt=1\) for every \(x\) but \(∫^{}F(x,t)\,dx=0\) for every \(t\).
(ii) Granted (i), Fubini’s theorem is available and that computation is legitimate:
\begin{equation*} \frac{1}{2h}\int_{x_0-h}^{x_0+h}\Delta(x)\,dx =\int_I\Bigl[\frac{1}{2h}\int_{x_0-h}^{x_0+h}|f(x,t)-f(x_0,t)|\,dx\Bigr]dt\to0 \end{equation*}
as \(h\to0^+\), since the inner averages tend to \(0\) for each fixed \(t\) by (P1) and are bounded by \(2M\) on the finite interval \(I\). So \(x_0\) is a Lebesgue point of the bounded measurable \(g_I\) with value \(g_I(x_0)\), whence \(g_I\) is approximately continuous by Exercise 5.8.90(i) and Baire-\(1\) by (P2). Therefore, for fixed \(n\) and \(m\), the finite sum \(\sum_j\tau_{j,m}(y)A_nf(x,j/m)\) is Baire-\(1\) on the square, each summand being a product of a Baire-\(1\) function of \(x\) and a continuous function of \(y\); it converges to \(A_nf\) uniformly as \(m\to\infty\), so \(A_nf\) is Baire-\(1\) by (P3). Finally \(f=\lim_nA_nf\) pointwise, so \(f\) is of the second Baire class.
(A fully rigorous proof of the remaining step is beyond the scope of this page; see the reference given in the book.)
(\(\circ\)) Prove the following Chebyshev inequality for monotone functions: if \(\varphi\) and \(\psi\) are nondecreasing finite functions on \([0,1]\) and \(\varrho\) is a probability density on \([0,1]\), then \[ \int_0^1 \varphi(x)\psi(x)\varrho(x)\,dx \ \ge \ \int_0^1 \varphi(x)\varrho(x)\,dx \int_0^1 \psi(x)\varrho(x)\,dx . \] If \(\varphi\) is an increasing function and \(\psi\) is decreasing, then the opposite inequality is true.
Everything follows from the sign of \[ H(x,y):=\bigl(\varphi(x)-\varphi(y)\bigr)\bigl(\psi(x)-\psi(y)\bigr)\varrho(x)\varrho(y) . \] A finite monotone function on \([0,1]\) is Borel and bounded, so \(H\) is measurable on the square with \(|H|\le4\|\varphi\|_\infty\|\psi\|_\infty\varrho(x)\varrho(y)\), and \(\varrho(x)\varrho(y)\) has integral \(1\) over the square by Tonelli’s theorem; hence \(H\) and each of the four terms of its expansion are integrable.
(i) \(\varphi,\psi\) both nondecreasing: then \(\varphi(x)-\varphi(y)\) and \(\psi(x)-\psi(y)\) have the same sign for all \(x,y\) and \(\varrho\ge0\), so \(H\ge0\), and expanding by Fubini’s theorem with \(\int_0^1\varrho\,dx=1\),
\begin{equation*} 0\le\int_0^1\!\!\int_0^1H\,dx\,dy =2\int_0^1\varphi\psi\varrho\,dx-2\int_0^1\varphi\varrho\,dx\int_0^1\psi\varrho\,dx . \end{equation*}
(ii) \(\varphi\) nondecreasing, \(\psi\) nonincreasing: the two differences have opposite signs, so \(H\le0\) and the same identity gives \(\int_0^1\varphi\psi\varrho\,dx\le\int_0^1\varphi\varrho\,dx\int_0^1\psi\varrho\,dx\).
Let \(f\colon [0,\infty) \to [0,\infty)\) be a locally integrable function.
(i) Assume \[ \int_E f(x)\,dx \le \sqrt{\lambda(E)} \] for every bounded measurable set \(E\). Prove that \[ \int_0^\infty \frac{f(x)}{1+x}\,dx \ \le \ \int_0^\infty \frac{\sqrt{x}}{(1+x)^2}\,dx \ < \ \frac{1}{2} . \]
(ii) Assume that \[ \int_0^T f(x)\,dx \le T \quad \text{ for all } T. \] Show that the function \[ \frac{f(x)}{1+x^2} \] is integrable.
Both parts follow from one integration by parts. With \(F(x)=\int_0^xf(y)\,dy\), finite by local integrability and nondecreasing, Tonelli’s theorem applied to \(f(x)(-\Phi^{\prime}(s))\mathbb{1}_{\{0\le x\le s\le t\}}\) gives, for every nonincreasing \(\Phi\in C^1[0,\infty)\),
\begin{equation*} \int_0^t f(x)\Phi(x)\,dx=\Phi(t)F(t)+\int_0^t(-\Phi^{\prime}(s))F(s)\,ds . \tag{*} \end{equation*}
(i) Taking \(E=[0,x]\) in the hypothesis gives \(F(x)\le\sqrt{x}\), so \((*)\) with \(\Phi(x)=(1+x)^{-1}\), \(-\Phi^{\prime}(s)=(1+s)^{-2}\ge0\), yields
\begin{equation*} \int_0^t\frac{f(x)}{1+x}\,dx \le\frac{\sqrt{t}}{1+t}+\int_0^t\frac{\sqrt{s}}{(1+s)^2}\,ds , \end{equation*}
and letting \(t\to+\infty\), where \(\sqrt{t}/(1+t)\to0\) and both integrals increase to their improper values by monotone convergence, gives the asserted inequality; in particular \(f(x)/(1+x)\) is integrable. Substituting \(x=u^2\),
\begin{equation*} \int_0^\infty\frac{\sqrt{x}}{(1+x)^2}\,dx=\int_0^\infty\frac{2u^2}{(1+u^2)^2}\,du =2\Bigl(\frac{\pi}{2}-\frac{\pi}{4}\Bigr)=\frac{\pi}{2}, \end{equation*}
since \(\int_0^\infty(1+u^2)^{-2}du=\int_0^{\pi/2}\cos^2\theta\,d\theta=\pi/4\) via \(u=\tan\theta\); the printed final bound \(<1/2\) is a misprint, the correct value of that integral being \(\pi/2\).
(ii) Here \(F(T)\le T\), and \((*)\) with \(\Phi(x)=(1+x^2)^{-1}\), \(-\Phi^{\prime}(s)=2s(1+s^2)^{-2}\ge0\), gives
\begin{equation*} \int_0^t\frac{f(x)}{1+x^2}\,dx\le\frac{t}{1+t^2}+\int_0^t\frac{2s^2}{(1+s^2)^2}\,ds , \end{equation*}
so letting \(t\to+\infty\) and using the computation above we get \(\int_0^\infty f(x)(1+x^2)^{-1}dx\le\pi/2<\infty\), i.e. \(f(x)/(1+x^2)\) is integrable.
(Gordon [374]) In analogy with the definitions in \(\S 5.7\) we shall consider tagged partitions \(P = \{(x_i, E_i)\}\) of the interval \([a,b]\) into finitely many pairwise disjoint measurable sets \(E_i\) with \(x_i \in E_i\). Such a partition \(P\) is said to be subordinate to a positive function \(\delta\) if \(E_i \subset \big(x_i - \delta(x_i),\, x_i + \delta(x_i)\big)\) for all \(i\). Prove that:
(i) a function \(f\) on \([a,b]\) is Riemann integrable precisely when there exists a number \(R\) with the following property: for every \(\varepsilon > 0\), there exists a number \(\delta > 0\) such that \(\big|\sum_{i=1}^n f(x_i)\lambda(E_i) - R\big| < \varepsilon\) for every tagged partition of the interval into measurable sets \(E_i\) subordinate to \(\delta\);
(ii) a function \(f\) on \([a,b]\) is Lebesgue integrable precisely when there exists a number \(L\) with the following property: for every \(\varepsilon > 0\), there exists a positive function \(\delta(\,\cdot\,)\) such that \(\big|\sum_{i=1}^n f(x_i)\lambda(E_i) - L\big| < \varepsilon\) for every tagged partition of the interval into measurable sets \(E_i\) subordinate to the function \(\delta\).
In (i) \(R\) is the Riemann integral and in (ii) \(L\) the Lebesgue integral; write \(S(f,P)=\sum_{i\le n}f(x_i)\lambda(E_i)\).
Subordinate measurable partitions exist for every gauge \(\delta\), and every sufficiently fine Riemann sum is one of the sums \(S(f,P)\): by Cousin’s lemma take a tagged partition of \([a,b]\) into nonoverlapping closed intervals \([t_{j-1},t_j]\) with tags \(y_j\in[t_{j-1},t_j]\) subordinate to \(\delta/2\); merge adjacent intervals with a common tag (such chains have length at most two, so the merged intervals are subordinate to \(\delta\) and the sum is unchanged), after which all tags are distinct; then let \(E_j\) consist of \((t_{j-1},t_j)\) together with those endpoints that are its tag or are unclaimed by the neighbour. The \(E_j\) are disjoint measurable sets with union \([a,b]\), \(y_j\in E_j\), \(\lambda(E_j)=t_j-t_{j-1}\) and \(E_j\subset[t_{j-1},t_j]\).
(i) If \(R\) has the stated property, then by the preceding paragraph every Riemann sum of mesh \(<\delta/2\) equals some \(S(f,P)\) with \(P\) subordinate to the constant gauge \(\delta\), hence differs from \(R\) by less than \(\varepsilon\); so \(f\) is Riemann integrable with integral \(R\).
Conversely let \(f\) be Riemann integrable, \(|f|\le M\), \(R=\int_a^bf\,dx\), which is also its Lebesgue integral. Given \(\varepsilon>0\) put \(\eta=\varepsilon/(2(b-a))\) and \(\sigma=\varepsilon/(16M+1)\). The set \(D_\eta=\{x:\mathrm{osc}(f,x)\ge\eta\}\) is compact and null by the Lebesgue criterion, so there is an open \(U=\bigcup_{j\le N}(a_j,b_j)\supset D_\eta\) with \(\lambda(U)<\sigma\). Every point of the compact set \(K=[a,b]\setminus U\) has oscillation \(<\eta\), so covering \(K\) by finitely many halved oscillation intervals gives \(\delta_1>0\) with \(|f(y)-f(x)|\le\eta\) whenever \(x\in K\) and \(|y-x|<\delta_1\); put \(\delta=\min\{\delta_1,\sigma/(2N)\}\). For \(P\) subordinate to \(\delta\), using \(\sum_i\int_{E_i}f\,d\lambda=R\),
\begin{equation*} |S(f,P)-R|\le\sum_i\int_{E_i}|f(x_i)-f(y)|\,dy \le\eta(b-a)+2M\lambda(U^\delta)<\frac{\varepsilon}{2}+4M\sigma<\varepsilon , \end{equation*}
since cells with tags in \(K\) contribute at most \(\eta\sum_i\lambda(E_i)\), while the disjoint cells with tags in \(U\) lie in the \(\delta\)-neighbourhood \(U^\delta\), of measure \(\le\sigma+2N\delta\le2\sigma\).
(ii) Necessity, with \(L=\int_a^bf\,d\lambda\). (1) Indicators. For \(f=\mathbb{1}_A\) and \(\varepsilon>0\) choose by regularity an open \(O\supset A\) and a closed \(C\subset A\) with \(\lambda(O\setminus A)<\varepsilon\), \(\lambda(A\setminus C)<\varepsilon\), and let \(\delta(x)\) be such that \((x-\delta(x),x+\delta(x))\subset O\) for \(x\in A\) and \((x-\delta(x),x+\delta(x))\cap C=\emptyset\) for \(x\notin A\). For any partial subordinate system \(\{(x_i,E_i)\}_{i\in I}\) with union \(W\), the cells tagged in \(A\) lie in \(O\) and those tagged off \(A\) miss \(C\), whence
\begin{equation*} \sum_{x_i\in A}\lambda(E_i)\le\lambda(W\cap A)+\lambda(O\setminus A),\qquad \lambda(A\cap W)\le\sum_{x_i\in A}\lambda(E_i)+\lambda(A\setminus C), \end{equation*}
that is, \(\bigl|\sum_{i\in I}\mathbb{1}_A(x_i)\lambda(E_i)-\int_W\mathbb{1}_A\,d\lambda\bigr|\le\varepsilon\).
(2) Simple functions. For \(\varphi=\sum_{j\le m}c_j\mathbb{1}_{A_j}\) apply (1) to each \(A_j\) with tolerance \(\varepsilon/(m(1+|c_j|))\) and let \(\delta_\varphi\) be the minimum of the gauges; by linearity, for every partial subordinate system with union \(W\),
\begin{equation*} \Bigl|\sum_{i\in I}\varphi(x_i)\lambda(E_i)-\int_W\varphi\,d\lambda\Bigr| \le\varepsilon . \tag{**} \end{equation*}
(3) General \(f\in L^1[a,b]\). Choose simple \(\varphi_k\to f\) pointwise everywhere with \(|\varphi_k|\le|f|\), so \(\int_a^b|f-\varphi_k|\,d\lambda\to0\) by dominated convergence, and after passing to a subsequence \(\int_a^b|f-\varphi_k|\,d\lambda<\varepsilon2^{-k}\). Let \(k(x)\) be the least \(k\) with \(|f(x)-\varphi_k(x)|<\varepsilon/(b-a)\), let \(B_k=\{k(\cdot)=k\}\), and set \(\delta(x):=\delta_{k(x)}(x)\), where \(\delta_k\) is the gauge of (2) for \(\varphi_k\) and tolerance \(\varepsilon2^{-k}\). If \(P\) is subordinate to \(\delta\), the cells with \(x_i\in B_k\) form a partial \(\delta_k\)-system with union \(W_k\), the \(W_k\) are disjoint with union \([a,b]\) so \(\sum_k\int_{W_k}f\,d\lambda=L\), and \(|\int_{W_k}(\varphi_k-f)\,d\lambda|<\varepsilon2^{-k}\); hence by \((**)\)
\begin{equation*} |S(f,P)-L|\le\frac{\varepsilon}{b-a}\sum_i\lambda(E_i) +\sum_k\varepsilon2^{-k}+\sum_k\varepsilon2^{-k} \le3\varepsilon . \end{equation*}
(ii) Sufficiency. Any free tagged partition \(\widehat P=\{(t_j,I_j)\}_{j\le N}\) subordinate to \(\delta\) in the sense of Definition 5.7.1 converts into a subordinate measurable partition with the same sum. Merge the intervals sharing a tag: for the distinct tags \(t^{(1)},\dots,t^{(m)}\) put \(C_k=\bigcup_{t_j=t^{(k)}}I_j\), which lies in \((t^{(k)}-\delta(t^{(k)}),t^{(k)}+\delta(t^{(k)}))\) and, the \(I_j\) being nonoverlapping, has \(\lambda(C_k)=\sum_{t_j=t^{(k)}}\lambda(I_j)\), so that \(\sum_kf(t^{(k)})\lambda(C_k)=I(f,\widehat P)\). Assign each of the finitely many points lying in several \(C_k\) to just one of them, getting disjoint \(D_k\subset C_k\) with union \([a,b]\), and put \(E_k:=(D_k\setminus T)\cup\{t^{(k)}\}\) with \(T=\{t^{(1)},\dots,t^{(m)}\}\). The \(E_k\) are disjoint measurable sets with union \([a,b]\), \(t^{(k)}\in E_k\), \(E_k\) lies in the window of \(t^{(k)}\), and \(\lambda(E_k)=\lambda(C_k)\), so \(S(f,\{(t^{(k)},E_k)\})=I(f,\widehat P)\). Hence the hypothesis gives \(|I(f,\widehat P)-L|<\varepsilon\) for every free tagged partition subordinate to \(\delta\), i.e. \(f\) is McShane integrable with integral \(L\) (Definition 5.7.3(ii)), and by Theorem 5.7.14 it is Lebesgue integrable with \(\int_a^bf\,d\lambda=L\).
Given a function \(f\) on \([a,b]\), its Banach indicatrix \(N_f \colon \mathbb{R}^1 \to [0,+\infty]\) is defined as follows: \(N_f(y)\) is the cardinality of the set \(f^{-1}(y)\).
(i) Prove that the indicatrix of a continuous function is measurable as a mapping with values in \([0,+\infty]\).
(ii) (Banach [50]) Prove that a continuous function \(f\) is of bounded variation precisely when the function \(N_f\) is integrable. In addition, one has \[ \int_{-\infty}^{+\infty} N_f(y)\,dy = V(f,[a,b]). \tag{5.8.19} \] In particular, \(N_f(y) < \infty\) a.e.
(iii) (H. Kestelman) Prove that for a general function \(f\) of bounded variation, the difference between the left and right sides of (5.8.19) equals the sum of the absolute values of all jumps of \(f\).
(iv) Prove that if a function \(f\) is continuous and \(N_f(y) < \infty\) for all \(y\), then \(f\) is differentiable almost everywhere.
(i) Partition \([a,b]\) into the dyadic cells \(I_{n,1}=[a,a+h_n]\), \(I_{n,k}=(a+(k-1)h_n,a+kh_n]\) for \(k=2,\dots,2^n\), \(h_n=2^{-n}(b-a)\), and put \(g_n:=\sum_{k\le2^n}\mathbb{1}_{f(I_{n,k})}\). Each \(f(I_{n,k})\) is Borel, since \(f(I_{n,1})\) is compact and \(f((c,d])=\bigcup_{m\ge m_0}f([c+1/m,d])\) is a countable union of compacta, so \(g_n\) is Borel; \(g_n\le g_{n+1}\), each \(I_{n,k}\) being the disjoint union of two cells of the next generation and \(f(A\cup B)=f(A)\cup f(B)\); and \(g_n\uparrow N_f\), because \(g_n(y)\) is the number of cells meeting the set \(f^{-1}(y)\), hence at most \(N_f(y)\), while any \(m\) distinct points of \(f^{-1}(y)\) lie in \(m\) distinct cells as soon as \(h_n\) is below their minimal distance. So \(N_f\) is Borel with values in \([0,+\infty]\).
(ii) By monotone convergence, and since continuity of \(f\) makes \(f(I_{n,k})\) an interval with \(\lambda(f(I_{n,k}))=\sup_{I_{n,k}}f-\inf_{I_{n,k}}f=\mathrm{osc}(f,\overline{I_{n,k}})\),
\begin{equation*} \int_{-\infty}^{+\infty}N_f(y)\,dy =\lim_n\sum_{k\le2^n}\lambda\bigl(f(I_{n,k})\bigr)=\lim_n\Omega_n, \qquad\Omega_n:=\sum_k\mathrm{osc}\bigl(f,\overline{I_{n,k}}\bigr), \end{equation*}
where \((\Omega_n)\) is nondecreasing. Each oscillation over a closed interval is attained, so \(\mathrm{osc}(f,\overline{I_{n,k}})\le V(f,\overline{I_{n,k}})\) and hence \(\Omega_n\le V(f,[a,b])\) by additivity of variation. Conversely, given a partition \(a=t_0<\dots<t_m=b\) and \(\varepsilon>0\), uniform continuity provides \(n\) with \(|f(u)-f(v)|<\varepsilon/(2m)\) for \(|u-v|\le h_n\); replacing each \(t_j\) by the nearest point of the dyadic grid changes \(\sum_j|f(t_j)-f(t_{j-1})|\) by at most \(\varepsilon\), and the resulting sum, being over a subset of the grid, is at most \(\Omega_n\). Thus \(V(f,[a,b])\le\lim_n\Omega_n\), which gives (5.8.19); in particular \(N_f\) is integrable exactly when \(f\) is of bounded variation, and then \(N_f<\infty\) a.e.
(iii) Let \(V(x)=V(f,[a,x])\), \(L=V(b)\), \(\Gamma=V([a,b])\), and set \(\widetilde f(V(x)):=f(x)\). This is well defined, since \(V(x)=V(x^{\prime})\) with \(x<x^{\prime}\) forces \(V(f,[x,x^{\prime}])=0\) and \(f\) constant there, and it is \(1\)-Lipschitz on \(\Gamma\) because \(|f(y)-f(x)|\le V(y)-V(x)\). The set \([0,L]\setminus\Gamma\) is a countable disjoint union of intervals, one per one-sided jump of \(V\), of length exactly the absolute value of the corresponding jump of \(f\); extending \(\widetilde f\) affinely across them keeps it \(1\)-Lipschitz, hence continuous, and \(V(\widetilde f,[0,L])=L=V(f,[a,b])\) (the bound \(\le\) by the Lipschitz constant, \(\ge\) because \(\widetilde f\circ V=f\), so every variation sum for \(f\) is one for \(\widetilde f\)). Let \(Q\) be the countable set consisting of the values of \(f\) on nondegenerate intervals of constancy and of all one-sided limits \(f(x_0\pm)\). For \(y\notin Q\) the map \(V\) is injective on \(f^{-1}(y)\), so \(\widetilde f^{-1}(y)\cap\Gamma\) has exactly \(N_f(y)\) points, while the affine piece over the gap of a jump from \(\alpha_j\) to \(\beta_j\) contributes one point precisely when \(y\) lies strictly between \(\alpha_j\) and \(\beta_j\):
\begin{equation*} N_{\widetilde f}(y)=N_f(y) +\sum_j\mathbb{1}_{(\alpha_j\wedge\beta_j,\ \alpha_j\vee\beta_j)}(y), \qquad y\notin Q . \end{equation*}
Integrating, and using (ii) for the continuous function \(\widetilde f\) and monotone convergence for the countable sum,
\begin{equation*} V(f,[a,b])=\int_{-\infty}^{+\infty}N_{\widetilde f}(y)\,dy =\int_{-\infty}^{+\infty}N_f(y)\,dy+\sum_j|\beta_j-\alpha_j| , \end{equation*}
the sum of the absolute values of all jumps of \(f\).
(iv) Let \(N_f(y)<\infty\) for every \(y\). For \(x\in(a,b)\) the finite set \(f^{-1}(f(x))\) leaves \(x\) isolated in it, so there is \(\rho(x)>0\) with \(f\ne f(x)\) for \(0<|t-x|\le\rho(x)\), and by the intermediate value theorem \(f-f(x)\) has a constant sign on each side of \(x\). Thus \(x\) is a strict local minimum, a strict local maximum, an ascending point (\(f<f(x)\) on the left, \(f>f(x)\) on the right) or a descending point.
Strict extrema form a countable set: a strict minimum \(x\) determines rationals \(p<x<q\) with \((p,q)\subset(x-\rho(x),x+\rho(x))\), and \(x\mapsto(p,q)\) is injective, since two distinct minima with the same pair would satisfy \(f(x)>f(x^{\prime})\) and \(f(x^{\prime})>f(x)\); maxima are symmetric.
The ascending points are \(\bigcup_{n,k}S_{n,k}\), where \(S_{n,k}\) is the set of \(x\in J_{n,k}:=[a+(k-1)/(2n),a+k/(2n)]\cap[a,b]\) with \(f<f(x)\) on \((x-\frac1n,x)\cap[a,b]\) and \(f>f(x)\) on \((x,x+\frac1n)\cap[a,b]\); indeed the defining conditions of an ascending point hold with \(1/n\le\rho(x)\). Descending points are ascending points of \(-f\), which is continuous with \(N_{-f}(y)=N_f(-y)<\infty\) and the same points of differentiability, so it suffices to treat a fixed \(S=S_{n,k}\) with at least two points.
Two properties of \(\overline S\subset J_{n,k}\): (a) \(f\) is strictly increasing on \(S\), two of its points being at distance \(\le1/(2n)<1/n\), hence nondecreasing on \(\overline S\) by passage to the limit; (b) for \(x\in\overline S\) one has \(f\le f(x)\) on \((x-\frac1n,x)\cap[a,b]\) and \(f\ge f(x)\) on \((x,x+\frac1n)\cap[a,b]\), as is seen by taking \(x_j\in S\), \(x_j\to x\), and using the defining inequalities at \(x_j\).
Let \(c=\inf\overline S\), \(d=\sup\overline S\), so \(0<d-c\le1/(2n)\), and put \(g(t):=\sup\{f(s):s\in\overline S\cap[c,t]\}\), which by (a) is nondecreasing with \(g=f\) on \(\overline S\); by Lebesgue’s theorem \(g^{\prime}\ge0\) exists a.e. With \(\alpha(t)=\sup(\overline S\cap[c,t])\) and \(\beta(t)=\inf(\overline S\cap[t,d])\), both in \(\overline S\) and within \(d-c<1/n\) of \(t\), property (b) gives the sandwich
\begin{equation*} g\bigl(\alpha(t)\bigr)=f\bigl(\alpha(t)\bigr)\le f(t) \le f\bigl(\beta(t)\bigr)=g\bigl(\beta(t)\bigr), \qquad t\in[c,d]. \end{equation*}
Fix \(x\in\overline S\cap(c,d)\) that is a density point of \(\overline S\) (Theorem 5.6.2) and at which \(g^{\prime}(x)\) exists, and let \(\varepsilon\in(0,1/2)\). One-sided densities also tend to \(1\), so there is \(h_0\in(0,\min(d-x,x-c))\) with \(\lambda(\overline S\cap(x,x+h))>(1-\varepsilon)h\) and \(\lambda(\overline S\cap(x-h,x))>(1-\varepsilon)h\) for \(0<h\le h_0\). Let \(0<t-x<(1-\varepsilon)^2h_0\). All points of \(\overline S\cap(x,t]\) are \(\le\alpha(t)\), so \(\alpha(t)-x>(1-\varepsilon)(t-x)\); and \((t,\beta(t))\cap\overline S=\emptyset\), so with \(h=\min(\beta(t)-x,h_0)\) we get \(\overline S\cap(x,x+h)\subset(x,t]\) and \((1-\varepsilon)h<t-x\), which rules out \(h=h_0\) and gives \(\beta(t)-x<(t-x)/(1-\varepsilon)\). Since the difference quotients of \(g\) are nonnegative, the sandwich yields
\begin{equation*} (1-\varepsilon)\frac{g(\alpha(t))-g(x)}{\alpha(t)-x}\le\frac{f(t)-f(x)}{t-x} \le\frac{1}{1-\varepsilon}\cdot\frac{g(\beta(t))-g(x)}{\beta(t)-x} , \end{equation*}
and letting \(t\downarrow x\), where \(\alpha(t),\beta(t)\downarrow x\), and then \(\varepsilon\downarrow0\), we get \(\lim_{t\downarrow x}(f(t)-f(x))/(t-x)=g^{\prime}(x)\). The left-hand limit is symmetric: now \(x-\beta(t)>(1-\varepsilon)(x-t)\) and \(x-\alpha(t)<(x-t)/(1-\varepsilon)\), and dividing the sandwich by the negative number \(t-x\) reverses it. Hence \(f^{\prime}=g^{\prime}\) a.e. on \(\overline S\supset S\).
Since \([a,b]\) is the union of \(\{a,b\}\), the countable set of strict extrema, the countably many \(S_{n,k}\) and their counterparts for \(-f\), the function \(f\) is differentiable almost everywhere.
(i) Let \(f\) be a continuous function on \([a,b]\) and let \(E\) be a Borel set in \([a,b]\). Show that the function \(y \mapsto N_f(E,y)\) from \(\mathbb{R}^1\) to \([0,+\infty]\) which to every \(y\) puts into correspondence the cardinality of the set \(E \cap f^{-1}(y)\) is Borel measurable.
(ii) Let \(f\) be a continuous function of bounded variation on \([a,b]\) and let \(V(x) = V(f,[a,x])\). Prove that for every Borel set \(B\) in \([a,b]\) one has \[ \lambda\big(V(B)\big) = \int_{-\infty}^{+\infty} N_f(B,y)\,dy . \] Deduce that if \(E \subset [a,b]\) is such that \(\lambda\big(f(E)\big) = 0\), then \(\lambda\big(V(E)\big) = 0\).
(i) The function \(N_f(E,\cdot)\) is Lebesgue measurable for every Borel \(E\), and Borel measurable on the set where the full indicatrix is finite; Borel measurability everywhere cannot be asserted, since already \(\{y:N_f(E,y)\ge1\}=f(E)\) is in general only Suslin.
With \(I_{m,k}\) the dyadic partition of Exercise 5.8.133(i), the functions \(g^E_m:=\sum_{k\le2^m}\mathbb{1}_{f(E\cap I_{m,k})}\) increase pointwise to \(N_f(E,\cdot)\) by the counting argument there. Each \(E\cap I_{m,k}\) is Borel, hence Suslin (Theorem 1.10.4), so its image under a continuous extension of \(f\) to \(\mathbb{R}\) (Tietze) is Suslin by Proposition 1.10.8 and Lebesgue measurable by Theorem 1.10.5; hence so is \(N_f(E,\cdot)\).
For the Borel statement put \(Z:=\{y:N_f([a,b],y)<\infty\}\), a Borel set by 5.8.133(i), and let \(\mathcal{A}\) be the class of Borel \(E\subset[a,b]\) for which \(N_f(E,\cdot)|_Z\) is Borel. Every closed subinterval lies in \(\mathcal{A}\), since \(N_f([c,d],\cdot)\) is the indicatrix of \(f|_{[c,d]}\); if \(A\subset B\) both lie in \(\mathcal{A}\), then on \(Z\) all counts are finite and \(N_f(B\setminus A,\cdot)=N_f(B,\cdot)-N_f(A,\cdot)\) there, so \(B\setminus A\in\mathcal{A}\) (this subtraction, and hence the hint’s induction, is legitimate only on \(Z\)); and \(E_n\uparrow E\) gives \(N_f(E_n,y)\uparrow N_f(E,y)\) by continuity from below of the counting measure of \(f^{-1}(y)\). Thus \(\mathcal{A}\) is a \(\sigma\)-additive class containing the closed subintervals, which form a \(\pi\)-system generating \(\mathcal{B}([a,b])\), so \(\mathcal{A}\supset\mathcal{B}([a,b])\) by Theorem 1.9.3(ii). For \(f\) of bounded variation \(\lambda(\mathbb{R}\setminus Z)=0\) by 5.8.133(ii), so \(N_f(E,\cdot)\) is then Borel up to a null set.
(ii) Both sides are finite Borel measures of total mass \(L:=V(b)\) agreeing on closed intervals.
The measure \(\nu(B):=\int_{\mathbb{R}}N_f(B,y)\,dy\), whose integrand is measurable by (i), is countably additive by monotone convergence, \(N_f(\cdot,y)\) being a measure for each \(y\), and \(\nu([a,b])=V(f,[a,b])=L\) by 5.8.133(ii).
The set function \(\mu(B):=\lambda(V(B))\) is defined because \(V(B)\subset[0,L]\) is Suslin (Proposition 1.10.8), hence measurable. For disjoint \(A,B\) the set \(V(A)\cap V(B)\) is at most countable: \(V(s)=V(t)\) with \(s<t\) forces \(V\) constant on \([s,t]\), and distinct such common values correspond to disjoint nondegenerate intervals of constancy of \(V\). Hence the images of a disjoint sequence are disjoint up to a null set and \(\mu\) is countably additive, with \(\mu([a,b])=\lambda([0,L])=L\), since \(V\) is continuous (\(f\) being continuous) and nondecreasing from \(0\) to \(L\).
On a closed interval, continuity and monotonicity of \(V\) give \(V([c,d])=[V( c),V(d)]\), so
\begin{equation*} \mu([c,d])=V(d)-V( c)=V\bigl(f,[c,d]\bigr) =\int_{\mathbb{R}}N_f([c,d],y)\,dy=\nu([c,d]) \end{equation*}
by 5.8.133(ii) applied to \(f\) on \([c,d]\). Since the closed subintervals form a \(\pi\)-system generating \(\mathcal{B}([a,b])\) and \(\mu,\nu\) are finite with equal total mass, Lemma 1.9.4 gives \(\mu=\nu\) on \(\mathcal{B}([a,b])\).
Deduction: if \(\lambda(f(E))=0\), choose a Borel \(S\supset f(E)\) with \(\lambda(S)=0\); then \(E\subset f^{-1}(S)\), a Borel set, and \(N_f(f^{-1}(S),y)\) equals \(N_f([a,b],y)\) for \(y\in S\) and \(0\) for \(y\notin S\), so the identity just proved gives
\begin{equation*} \lambda\Bigl(V\bigl(f^{-1}(S)\bigr)\Bigr)=\int_S N_f([a,b],y)\,dy=0 . \end{equation*}
As \(V(E)\subset V(f^{-1}(S))\), completeness of Lebesgue measure yields \(\lambda(V(E))=0\).
Exercises 5.8.135–5.8.141
(\(\circ\)) Let \(\mu\) be a measure on a space \(X\) and let a function \(f\) on \(X\times[a,b]\) be such that the functions \(x\mapsto f(x,t)\) are integrable and the functions \(t\mapsto f(x,t)\) are absolutely continuous. Suppose that the function \(\partial f/\partial t\) is integrable with respect to \(\mu\otimes\lambda\), where \(\lambda\) is Lebesgue measure. Prove that the function
\begin{equation*} t\mapsto \int_X f(x,t)\,\mu(dx) \end{equation*}
is absolutely continuous and
\begin{equation*} \frac{d}{dt}\int_X f(x,t)\,\mu(dx)=\int_X \frac{\partial f(x,t)}{\partial t}\,\mu(dx)\quad\text{a.e.} \end{equation*}
\(F(t):=\int_Xf(x,t)\,\mu(dx)\) is absolutely continuous, with \(F^{\prime}(t)=\int_X\partial f(x,t)/\partial t\,\mu(dx)\) a.e.; the two displayed formulas print \(f(t,x)\), a misprint for \(f(x,t)\).
Joint measurability. With \(t_{n,k}=a+k(b-a)2^{-n}\), the functions
\begin{equation*} f_n(x,t)=f(x,a)I_{\{a\}}(t)+\sum_{k<2^n}f(x,t_{n,k+1})I_{(t_{n,k},t_{n,k+1}]}(t) \end{equation*}
are \(\mathcal{A}\otimes\mathcal{B}([a,b])\)-measurable, being finite sums of products of measurable functions of \(x\) and Borel functions of \(t\), and \(f_n\to f\) at every point because each \(t\mapsto f(x,t)\) is continuous; so \(f\) is jointly measurable (Theorem 2.1.5).
A measurable version of the derivative. Extend \(f\) by \(f(x,\pi(t))\), \(\pi=\min(\max(\cdot,a),b)\), still jointly measurable and continuous in \(t\). Each difference quotient \(Q_h=(f(\cdot,\cdot+h)-f)/h\) is then measurable, and by continuity in \(t\) the upper and lower limits of \(Q_h\) as \(h\to0\) may be computed along rational \(h\) alone, hence are measurable; so the set \(M\subset X\times(a,b)\) where they agree and are finite is measurable, it is exactly where \(\partial f/\partial t\) exists finitely, and \(g:=u\,I_M\) with \(u=\limsup_{h\to0}Q_h\) is a measurable function equal to \(\partial f/\partial t\) on \(M\). For each \(x\) the absolutely continuous \(t\mapsto f(x,t)\) is differentiable a.e. (Theorem 5.3.6 with Theorem 5.4.2), so the complement \(N\) of \(M\) has all sections \(N_x\) null, whence \((\mu\otimes\lambda)(N)=\int_X\lambda(N_x)\,\mu(dx)=0\) by Theorems 3.4.5 and 3.4.4, and \(N^t\) is \(\mu\)-null for a.e. \(t\). Thus \(g\) is a version of \(\partial f/\partial t\), the hypothesis says \(g\in L^1(\mu\otimes\lambda)\), and
\begin{equation*} \int_X\frac{\partial f(x,t)}{\partial t}\,\mu(dx) =\int_X g(x,t)\,\mu(dx)\qquad\text{for a.e. }t. \tag{0} \end{equation*}
Fibrewise Newton–Leibniz. For fixed \(x\), Theorem 5.3.6 gives \(h_x\in L^1[a,b]\) with \(f(x,t)=f(x,a)+\int_a^th_x(s)\,ds\), and \(h_x(s)=\partial f(x,s)/\partial s=g(x,s)\) for a.e. \(s\) by Theorem 5.4.2; hence
\begin{equation*} f(x,t)=f(x,a)+\int_a^t g(x,s)\,ds\qquad (x\in X,\ t\in[a,b]). \tag{1} \end{equation*}
Since \(g\in L^1(\mu\otimes\lambda)\) and both measures are \(\sigma\)-finite, Fubini’s theorem applied to \(g\,I_{X\times[a,t]}\) shows that \(G(s):=\int_Xg(x,s)\,\mu(dx)\) is defined for a.e. \(s\), lies in \(L^1[a,b]\), and
\begin{equation*} \int_X\Bigl(\int_a^t g(x,s)\,ds\Bigr)\mu(dx)=\int_a^t G(s)\,ds . \end{equation*}
Integrating (1) in \(x\), where \(x\mapsto f(x,t)\) and \(x\mapsto f(x,a)\) are \(\mu\)-integrable by hypothesis, gives \(F(t)=F(a)+\int_a^tG(s)\,ds\) for all \(t\). By Theorem 5.3.6 this says \(F\) is absolutely continuous, and by Theorem 5.4.2 and (0), \(F^{\prime}(t)=G(t)=\int_X\partial f(x,t)/\partial t\,\mu(dx)\) for a.e. \(t\).
(Tolstoff [951]) (i) Let \(\varphi\) be a positive monotone function on \((0,1]\) with \(\lim_{h\to 0}\varphi(h)=0\). Prove that for every \(\alpha\in(0,1)\), there exists a perfect nowhere dense set \(P\subset[0,1]\) of Lebesgue measure \(\alpha\) such that for a.e. \(x\in P\), there exists a number \(\delta(x)>0\) for which one has \(\lambda\bigl((x,x+h)\setminus P\bigr)<\varphi(|h|)|h|\) whenever \(|h|<\delta(x)\).
(ii) Prove that for every \(\alpha\in(0,1)\), there exists a perfect nowhere dense set \(P\subset[0,1]\) of Lebesgue measure \(\alpha\) such that for sufficiently small \(|h|\) one has
\begin{equation*} \lambda\bigl((x,x+h)\setminus P\bigr)>\varphi(|h|)|h| \quad\text{for all } x\in P . \end{equation*}
(iii) Let \(P\) be a perfect set in \([0,1]\), \([0,1]\setminus P=\bigcup_{n=1}^{\infty}U_n\), where the \(U_n\)’s are disjoint intervals. Suppose that \(\lambda(U_n)\le q^n\), where \(0<q<1\). Show that for every \(\alpha>0\), for a.e. \(x\) there exists \(\delta(x)>0\) such that \(\lambda\bigl((x,x+h)\setminus P\bigr)<|h|^{\alpha}\) whenever \(|h|<\delta(x)\).
(iv) Show that there exist two mutually complementary measurable sets \(A\) and \(B\) in \([0,1]\) such that \(\lambda\bigl(A\cap(x,x+h)\bigr)>\varphi(|h|)|h|\) for a.e. \(x\in B\) whenever \(|h|<\delta(x)\) and \(\lambda\bigl(B\cap(x,x+h)\bigr)>\varphi(|h|)|h|\) for a.e. \(x\in A\) whenever \(|h|<\delta(x)\).
All four parts are Cantor-type constructions plus the Borel–Cantelli lemma. Note first that \(\varphi\) is nondecreasing: a nonincreasing \(\varphi\) would satisfy \(\varphi\ge\varphi(1)>0\), contradicting \(\varphi(h)\to0\). For \(h<0\) we read \((x,x+h)\) as \((x+h,x)\).
(i) Put \(\eta_n:=4^{-n}\), \(c_n:=\varphi(\eta_{n+1})\eta_{n+1}>0\), \(s_0:=1-\alpha\) and \(s_n:=\min\{s_{n-1}/2,\ 2^{n-1}c_n\}\), so that \(s_n\downarrow0\), \(s_n\le2^{-n}s_0\) and \(2^{-n}s_n\le c_n/2<c_n\). Let \(F_0=[0,1]\) and let \(F_n\) arise from \(F_{n-1}\) by deleting from each of its \(2^{n-1}\) intervals the concentric open interval of length \(\ell_n:=2^{1-n}(s_{n-1}-s_n)\); this is possible since \(\ell_n<2^{1-n}(\alpha+s_{n-1})=\delta_{n-1}\), and \(F_n\) is the union of \(2^n\) congruent closed intervals of length \(\delta_n=2^{-n}(\alpha+s_n)\), because \(\lambda(F_n)=1-\sum_{k\le n}2^{k-1}\ell_k=\alpha+s_n\). Then \(P:=\bigcap_nF_n\) is compact with \(\lambda(P)=\alpha\), has empty interior since \(\delta_n\to0\), and has no isolated points since both halves of the rank-\(n\) interval through a point of \(P\) meet \(P\); so \(P\) is perfect and nowhere dense of measure \(\alpha\).
Let \(B_n:=\{x\in P:\mathrm{dist}(x,\mathbb{R}\setminus F_n)<\eta_n\}\), covered by the \(\eta_n\)-neighbourhoods of the \(2\cdot2^n\) endpoints of \(F_n\), so \(\lambda(B_n)\le2\cdot2^{-n}\) is summable and by Borel–Cantelli a.e. \(x\in P\) lies in only finitely many \(B_n\). For such an \(x\) take \(N\) with \(x\notin B_n\) for all \(n\ge N\) and put \(\delta(x):=\eta_N\). Given \(0<|h|<\delta(x)\), choose \(n\ge N\) with \(\eta_{n+1}\le|h|<\eta_n\); every point between \(x\) and \(x+h\) is within \(|h|<\eta_n\le\mathrm{dist}(x,\mathbb{R}\setminus F_n)\) of \(x\), so that interval lies in the single rank-\(n\) interval \(I\ni x\), and \(\lambda(I\cap P)=2^{-n}\alpha\) by congruence. Hence
\begin{equation*} \lambda\bigl((x,x+h)\setminus P\bigr)\le\lambda(I\setminus P)=2^{-n}s_n<c_n =\varphi(\eta_{n+1})\eta_{n+1}\le\varphi(|h|)|h| . \end{equation*}
(ii) Now each interval must be split into many pieces, so that the ranks shrink as fast as we please while the removed proportions stay summable. Put \(\theta_n:=1-\alpha^{2^{-n}}\), so \(\prod_n(1-\theta_n)=\alpha\). With \(F_0=[0,1]\), \(\delta_0=1\), split each rank-\((n-1)\) interval into \(k_n\ge2\) closed intervals of length \(\delta_n=(1-\theta_n)\delta_{n-1}/k_n\) separated by \(k_n-1\) gaps of length \(g_n=\theta_n\delta_{n-1}/(k_n-1)\), so that exactly \(\theta_n\delta_{n-1}\) is removed from each; choose \(k_n\) so large that
\begin{equation*} 2\delta_n\le1,\qquad g_n<g_{n-1},\qquad 2\varphi(2\delta_n)\le\theta_{n+1}, \tag{2} \end{equation*}
possible since \(\delta_n,g_n\to0\) as \(k_n\to\infty\) and \(\varphi(s)\to0\) as \(s\downarrow0\). As in (i), \(P=\bigcap_nF_n\) is perfect and nowhere dense with \(\lambda(P)=\prod_n(1-\theta_n)=\alpha\).
Covering estimate: if \(b\) is the right endpoint of a rank-\(n\) interval and \(J=[b,b+L]\), then \(\lambda(J\setminus P)\ge g_nL/(g_n+\delta_n)\). Indeed, to the right of \(b\) the line decomposes into consecutive blocks \(G_i\cup I_i\) with \(|G_i|\ge g_n\) (a gap between consecutive rank-\(n\) intervals has generation \(m\le n\), hence length \(g_m\ge g_n\) by (2), and \((1,+\infty)\) counts as a last gap) and \(|I_i|=\delta_n\). Taking \(j\) maximal with \(T=\sum_{i\le j}|G_i\cup I_i|\le L\) and \(R=L-T\), one finds in either case (\(|G_{j+1}|\ge R\) or \(|G_{j+1}|<R\), the latter giving \(R<|G_{j+1}|+\delta_n\) by maximality) an integer \(p\ge0\) with \(\lambda(J\setminus P)\ge\max\{pg_n,\ L-p\delta_n\}\), and a maximum dominates the convex combination \((\delta_n\cdot pg_n+g_n(L-p\delta_n))/(\delta_n+g_n)=g_nL/(g_n+\delta_n)\).
Now put \(\delta:=2\delta_1\), let \(x\in P\) and \(0<|h|<\delta\), and pick \(n\ge2\) with \(2\delta_n\le|h|<2\delta_{n-1}\). If \([a,b]\) is the rank-\(n\) interval containing \(x\), then \(b-a=\delta_n\le|h|/2\), so for \(h>0\) the interval \((x,x+h)\) contains \((b,b+h/2)\) and the estimate applies; for \(h<0\) it contains \((a-|h|/2,a)\) and the estimate applies after the reflection \(y\mapsto1-y\), which maps each \(F_n\) onto itself, the splitting being symmetric. Since \(\delta_{n-1}=k_n\delta_n+(k_n-1)g_n>(k_n-1)(g_n+\delta_n)\),
\begin{equation*} \theta_n=\frac{(k_n-1)g_n}{\delta_{n-1}}<\frac{g_n}{g_n+\delta_n}, \end{equation*}
so by (2) at index \(n-1\), monotonicity of \(\varphi\) and \(|h|<2\delta_{n-1}\le1\),
\begin{equation*} \lambda\bigl((x,x+h)\setminus P\bigr)>\theta_n\frac{|h|}{2} \ \ge\ \varphi(2\delta_{n-1})|h| \ \ge\ \varphi(|h|)|h| . \end{equation*}
(iii) The assertion is meant for a.e. \(x\in P\): for \(x\notin P\) and \(\alpha>1\) it fails, small intervals about \(x\) lying inside some \(U_n\). Put \(\beta:=1/(2\alpha)\) and \(d_n:=q^{\beta n}\), summable, and \(E_n:=\{x\in P:\mathrm{dist}(x,U_n)<d_n\}\), contained in the two intervals of length \(d_n\) adjoining \(U_n\), so \(\lambda(E_n)\le2d_n\). By Borel–Cantelli, and since the endpoints of the \(U_n\) form a null set, almost every \(x\in P\) lies in \((0,1)\), is no endpoint of any \(U_n\), and admits \(n_0\) with \(\mathrm{dist}(x,U_n)\ge d_n\) for \(n>n_0\). For such \(x\) put
\begin{equation*} \delta(x):=\min\Bigl\{x,\ 1-x,\ (1-q)^{1/\alpha}, \ \min_{n\le n_0}\mathrm{dist}(x,U_n)\Bigr\}>0 . \end{equation*}
For \(0<|h|<\delta(x)\) a term \(\lambda(U_n\cap(x,x+h))\) can be nonzero only if \(\mathrm{dist}(x,U_n)<|h|\), which is impossible for \(n\le n_0\) and forces \(q^{\beta n}<|h|\), i.e. \(n>\tau:=\log(1/|h|)/(\beta\log(1/q))\), for \(n>n_0\). Hence, by \(\lambda(U_n)\le q^n\),
\begin{equation*} \lambda\bigl((x,x+h)\setminus P\bigr)\le\sum_{n>\tau}q^n\le\frac{q^\tau}{1-q} =\frac{|h|^{2\alpha}}{1-q}<|h|^{\alpha}, \end{equation*}
the last step because \(|h|<(1-q)^{1/\alpha}\).
(iv) Build \(A\) and \(B\) by switching labels at rapidly decreasing scales. Put \(\mu_j:=2^{-j-1}\), \(\ell_0:=1\), \(\ell_j:=\ell_{j-1}/K_j\) with \(K_j\) a multiple of \(2^{j+1}\) chosen so large that
\begin{equation*} 4\ell_j\le1\qquad\text{and}\qquad\varphi(4\ell_j)<2^{-j-5}, \tag{4} \end{equation*}
and \(M_j:=\mu_jK_j\in\mathbb{N}\), \(1\le M_j<K_j\). Let \(\mathcal{D}_j\) partition \([0,1)\) into blocks of length \(\ell_j\). Label \([0,1)\) by \(A\); if a rank-\((j-1)\) block carries a label, give its first \(M_j\) children the opposite label and its remaining \(K_j-M_j\) children the same label. Then \(\lambda\{L_j\ne L_{j-1}\}=M_j/K_j=\mu_j\) is summable, so by Borel–Cantelli the label \(L_j(x)\) stabilizes for a.e. \(x\). Let \(A\) be the set of points stabilizing at \(A\), together with the null non-stabilizing set and \(\{1\}\), and \(B:=[0,1]\setminus A\); both are measurable, each \(\{L_j=A\}\) being a finite union of intervals.
Two facts: a rank-\(j\) block has children of rank \(j+1\) labelled \(A\) of total length at least \(\mu_{j+1}\ell_j\) (namely \(M_{j+1}\ell_{j+1}=\mu_{j+1}\ell_j\) or \((1-\mu_{j+1})\ell_j\), according to its own label), and likewise for \(B\); and a rank-\(m\) block \(J\) labelled \(A\) has \(\lambda(A\cap J)\ge\ell_m\prod_{j>m}(1-\mu_j)\ge\ell_m/2\), since points of \(J\) never switching after rank \(m\) lie in \(A\), and likewise for \(B\).
Let \(x\in(0,1)\) and \(0<|h|<\delta(x):=\min\{4\ell_1,x,1-x\}\), and pick \(n\ge2\) with \(4\ell_n\le|h|<4\ell_{n-1}\). The rank-\(n\) blocks entirely contained in \((x,x+h)\) number at least \(|h|/\ell_n-2\ge|h|/(2\ell_n)\), using \(|h|\ge4\ell_n\), and each such block \(J\) satisfies \(\lambda(A\cap J)\ge\frac12\mu_{n+1}\ell_n\) by the two facts; summing over these disjoint blocks,
\begin{equation*} \lambda\bigl(A\cap(x,x+h)\bigr) \ \ge\ \frac{|h|}{2\ell_n}\cdot\frac{\mu_{n+1}\ell_n}{2}=2^{-n-4}|h| . \end{equation*}
By (4) at \(j=n-1\) and monotonicity, \(\varphi(|h|)\le\varphi(4\ell_{n-1})<2^{-n-4}\), so \(\lambda(A\cap(x,x+h))>\varphi(|h|)|h|\); interchanging \(A\) and \(B\) gives the companion inequality.
(Bary [66, Appendix, 13]) Let \(E \subset [0,1]\) be a measurable set of positive measure and let \(E_0\) be the set of density points of \(E\). Prove that for every \(x \in E_0\) and every number \(\alpha\) there are numbers \(\lambda_n\) such that \(\lambda_n = \alpha n^{-1} + o(1/n)\), \(x + \lambda_n \in E_0\), \(x - \lambda_n \in E_0\) for all \(n\).
Since \(\lambda(E\triangle E_0)=0\), the point \(x\) is a density point of \(E_0\) itself, so the set of radii \(t\) with \(x\pm t\in E_0\) has density \(1\) at \(0\), and the \(\lambda_n\) are chosen from it.
By Theorem 5.6.2 applied to the integrable function \(I_E\), at almost every \(y\) the density of \(E\) equals \(I_E(y)\); hence \(\lambda(E\setminus E_0)=0=\lambda(E_0\setminus E)\), so \(E_0\) is measurable by completeness and \(E,E_0\) have exactly the same density points. Thus \(x\) is a density point of \(E_0\), and splitting a symmetric interval into halves,
\begin{equation*} \lambda\bigl((x,x+r]\setminus E_0\bigr)=o( r),\qquad \lambda\bigl([x-r,x)\setminus E_0\bigr)=o( r),\qquad r\to0^+ . \end{equation*}
Put \(S:=\{t>0:\ x+t\in E_0\ \text{and}\ x-t\in E_0\}\), measurable as the intersection of two affine images of \(E_0\); since \((0,r]\setminus S\) is contained in the union of the two exceptional sets above, \(\lambda((0,r]\setminus S)=o( r)\).
Let \(\alpha>0\) and \(\varepsilon\in(0,1)\), and put \(r:=(1+\varepsilon)\alpha/n\). If \(S\) missed \(J_{n,\varepsilon}:=[(1-\varepsilon)\alpha/n,\ (1+\varepsilon)\alpha/n]\subset(0,r]\), then
\begin{equation*} \lambda\bigl((0,r]\setminus S\bigr)\ge\lambda(J_{n,\varepsilon}) =\frac{2\varepsilon\alpha}{n} =\frac{2\varepsilon}{1+\varepsilon}\,r , \end{equation*}
impossible for large \(n\), as \(r\to0\). So there is \(N(\varepsilon)\) with \(S\cap J_{n,\varepsilon}\ne\emptyset\) for all \(n\ge N(\varepsilon)\). Choose \(N_1<N_2<\cdots\) with \(N_k\ge N(1/k)\), set \(\lambda_n:=0\) for \(n<N_1\), and for \(N_k\le n<N_{k+1}\) pick \(\lambda_n\in S\cap J_{n,1/k}\), so that \(|\lambda_n-\alpha/n|\le\alpha/(kn)\). Then \(x\pm\lambda_n\in E_0\) for every \(n\) (for \(n<N_1\) because \(x\in E_0\)), and \(n|\lambda_n-\alpha/n|\le\alpha/k\) for all \(n\ge N_k\), i.e. \(\lambda_n=\alpha n^{-1}+o(1/n)\).
For \(\alpha<0\) take \(\lambda_n:=-\mu_n\) with \(\mu_n\) the sequence just built for \(|\alpha|\), the two membership conditions being symmetric under \(\lambda_n\mapsto-\lambda_n\); for \(\alpha=0\) take \(\lambda_n:=0\).
(Brodskii [130]) Suppose we are given a continuously differentiable function \(f\) on the plane such that its partial derivatives at the point \((x_0, y_0)\) do not vanish. Suppose that \(x_0\) is a density point of a measurable set \(A \subset \mathbb{R}^1\) and \(y_0\) is a density point of a measurable set \(B \subset \mathbb{R}^1\). Prove that in some neighborhood of \(f(x_0, y_0)\), every point has the form \(f(x,y)\) with \(x \in A\), \(y \in B\).
Translating coordinates we may assume \(x_0=y_0=0\), and we write \(z_0=f(0,0)\), \(a=\partial_xf(0,0)\ne0\), \(b=\partial_yf(0,0)\ne0\), \(c=a/b\); the level curve \(\{f=z\}\) is a bi-Lipschitz graph over the \(x\)-axis, which transports the density of \(B\) at \(0\) to the axis, where it must meet \(A\).
Lipschitz lemma: an \(L\)-Lipschitz \(\psi\) on \(E\subset\mathbb{R}\) satisfies \(\lambda^*(\psi(E))\le L\lambda^*(E)\) – cover \(E\) by intervals \(I_k\) with \(\sum_k|I_k|\le\lambda^*(E)+\varepsilon\) and note that \(\psi(E\cap I_k)\) has diameter at most \(L|I_k|\). Using outer measure here dispenses with measurability questions for the sets pulled back below.
Since \(\partial_yf(0,0)=b\ne0\), the implicit function theorem gives \(\alpha,\beta,\delta_1>0\) and a \(C^1\) map \(\varphi:(-\alpha,\alpha)\times(z_0-\delta_1,z_0+\delta_1)\to(-\beta,\beta)\) with \(\varphi(0,z_0)=0\) and \(\{y\in(-\beta,\beta):f(x,y)=z\}=\{\varphi(x,z)\}\), where \(\partial_x\varphi=-\partial_xf/\partial_yf\) is continuous and equals \(-c\) at \((0,z_0)\). Shrinking \(\alpha\) and \(\delta_1\) so that \(|\partial_x\varphi+c|\le|c|/3\) throughout, the mean value theorem gives, for \(\varphi_z:=\varphi(\cdot,z)\),
\begin{equation*} \tfrac23|c|\,|x-x^{\prime}|\le\bigl|\varphi_z(x)-\varphi_z(x^{\prime})\bigr| \le\tfrac43|c|\,|x-x^{\prime}|, \qquad|x|,|x^{\prime}|<\alpha, \tag{1} \end{equation*}
so \(\varphi_z\) is injective with inverse \(\psi\) Lipschitz of constant \(3/(2|c|)\).
Since \(g(x):=f(x,0)\) has \(g(0)=z_0\) and \(g^{\prime}(0)=a\ne0\), the inverse function theorem gives \(\delta_2\in(0,\delta_1]\) and a continuous \(z\mapsto x_z\in(-\alpha,\alpha)\) with \(f(x_z,0)=z\) and \(x_{z_0}=0\), so \(x_z\to0\) as \(z\to z_0\); by the uniqueness in the implicit function theorem, \(\varphi_z(x_z)=0\).
Now fix \(r>0\) with \(2r<\alpha\), \(\frac43|c|r<\beta\), \(\lambda([-s,s]\setminus A)\le\frac18\cdot2s\) for \(0<s\le2r\) and \(\lambda([-s,s]\setminus B)\le\frac18\cdot2s\) for \(0<s\le\frac43|c|r\) – possible since \(0\) is a density point of \(A\) and of \(B\) – and then \(\delta\in(0,\delta_2]\) with \(|x_z|\le r\) whenever \(|z-z_0|<\delta\).
Let \(|z-z_0|<\delta\), put \(I:=[x_z-r,x_z+r]\subset[-2r,2r]\subset(-\alpha,\alpha)\) and \(C_z:=\{x\in I:\varphi_z(x)\in B\}\). Then \(\lambda^*(I\setminus A)\le\lambda([-2r,2r]\setminus A)\le r/2\). Also \(|\varphi_z(x)|=|\varphi_z(x)-\varphi_z(x_z)|\le\frac43|c|r=:R<\beta\) for \(x\in I\) by (1), so \(\varphi_z(I)\subset[-R,R]\), while \(I\setminus C_z=\psi(\varphi_z(I)\setminus B)\); hence the lemma gives
\begin{equation*} \lambda^*(I\setminus C_z)\le\frac{3}{2|c|}\lambda\bigl([-R,R]\setminus B\bigr) \le\frac{3}{2|c|}\cdot\frac18\cdot\frac83|c|r=\frac r2 . \end{equation*}
Were \(A\cap C_z=\emptyset\), then \(I\subset(I\setminus A)\cup(I\setminus C_z)\) would force \(2r=\lambda(I)\le r/2+r/2=r\). So there is \(x\in A\cap C_z\), and \(y:=\varphi_z(x)\in B\) satisfies \(f(x,y)=z\); undoing the translation, every point of \((f(x_0,y_0)-\delta,f(x_0,y_0)+\delta)\) has this form.
Prove that a necessary and sufficient condition that two sets on the real line are metrically separated in the sense of Exercise 1.12.160 is that at almost all points of one set the density of the other set is zero.
For arbitrary \(A,B\subset\mathbb{R}\) the following are equivalent: (a) \(A\) and \(B\) are metrically separated; (b) there are measurable \(A_0\supset A\), \(B_0\supset B\) with \(\lambda(A_0\cap B_0)=0\); (c) the density of \(B\), taken in the outer sense \(\lambda^{*}(B\cap[x-r,x+r])/2r\to0\), is zero at almost all points of \(A\), i.e. at all points of \(A\) outside a set of outer measure zero. Since (b) is symmetric, (c) holds with \(A\) and \(B\) interchanged as well. No measurability of \(A\) or \(B\) is assumed.
Every \(S\subset\mathbb{R}\) has a measurable envelope \(S^{*}\supset S\) with
\begin{equation*} \lambda^{*}(S\cap E)=\lambda(S^{*}\cap E)\qquad\text{for every measurable }E \tag{1} \end{equation*}
(Section 1.12(iv), (1.12.3)–(1.12.4) and Proposition 1.12.12; for the \(\sigma\)-finite \(\lambda\) take \(S^{*}=\bigcup_{n}H_n\), where \(H_n\subset[n,n+1)\) is the \(G_\delta\) hull \(\bigcap_kU_k\) of \(S\cap[n,n+1)\) with \(\lambda(U_k)\le\lambda^{*}(S\cap[n,n+1))+1/k\) furnished by outer regularity, and sum over the Caratheodory-measurable intervals \([n,n+1)\)). Taking \(E=[x-r,x+r]\) in (1), the outer density of \(S\) at \(x\) is the density of the measurable set \(S^{*}\) at \(x\).
(a) implies (b): choose open \(U_n\supset A\), \(V_n\supset B\) with \(\lambda(U_n\cap V_n)<1/n\) and put \(A_0:=\bigcap_nU_n\), \(B_0:=\bigcap_nV_n\), so \(A_0\cap B_0\subset U_n\cap V_n\) for every \(n\).
(b) implies (a): given \(\varepsilon>0\), outer regularity gives open \(U\supset A_0\), \(V\supset B_0\) with \(\lambda(U\setminus A_0),\lambda(V\setminus B_0)<\varepsilon/2\); a point of \(U\cap V\) outside \(U\setminus A_0\) and \(V\setminus B_0\) lies in \(A_0\cap B_0\), so \(U\cap V\subset(U\setminus A_0)\cup(V\setminus B_0)\cup(A_0\cap B_0)\) and \(\lambda(U\cap V)<\varepsilon\).
(b) implies (c): put \(C:=B_0\setminus A_0\), so that \(C\cap A_0=\emptyset\) and \(B_0\setminus C=A_0\cap B_0\) is null. By Theorem 5.6.2 applied to the indicator of \(\mathbb{R}\setminus C\), almost every point of \(\mathbb{R}\setminus C\) is a density point of \(\mathbb{R}\setminus C\), i.e. has density of \(C\) equal to \(0\); let \(N\) be the null exceptional set. For \(x\in A\setminus N\subset(\mathbb{R}\setminus C)\setminus N\) the density of \(C\) at \(x\) is \(0\), hence so is that of \(B_0\), since \(\lambda(B_0\cap I)\le\lambda(C\cap I)+\lambda(B_0\setminus C)=\lambda(C\cap I)\) for every interval \(I\), hence so is the outer density of \(B\subset B_0\); and \(\lambda^{*}(A\cap N)=0\).
(c) implies (b): let \(B^{*}\) be an envelope of \(B\) and \(D\) the set of its density points, so \(\lambda^{*}(B^{*}\setminus D)=0\) by the density theorem. By (1) the hypothesis says that \(B^{*}\) has density \(0\) at every \(x\in A\setminus N\), where \(\lambda^{*}(N)=0\); such \(x\) are not density points of \(B^{*}\), so \((A\setminus N)\cap D=\emptyset\) and
\begin{equation*} A\cap B^{*}\subset N\cup\bigl((A\setminus N)\cap B^{*}\bigr) \subset N\cup(B^{*}\setminus D), \end{equation*}
giving \(\lambda^{*}(A\cap B^{*})=0\). With \(A^{*}\) an envelope of \(A\), formula (1) for \(S=A\) and \(E=B^{*}\) yields \(\lambda(A^{*}\cap B^{*})=\lambda^{*}(A\cap B^{*})=0\), which is (b).
Let \(f = (f_1, \ldots, f_n) \colon \mathbb{R}^n \to \mathbb{R}^n\), where \(f_i \in W^{1,1}(\mathbb{R}^n)\). Set
\begin{equation*} \Omega = \bigl\{ \det(\partial_{x_j} f_i)_{i,j \le n} \neq 0 \bigr\} \end{equation*}
and denote by \(\lambda|_{\Omega}\) the restriction of Lebesgue measure to \(\Omega\). Show that the measure \(\lambda|_{\Omega} \circ f^{-1}\) is absolutely continuous.
It suffices to prove that \(\lambda(B)=0\) implies \(\lambda(\Omega\cap f^{-1}(B))=0\). Fix Borel representatives of the \(f_i\) and of \(\partial_{x_j}f_i\in L^1\), so \(f\) is Borel and \(\Omega\) is Borel; another choice alters \(\Omega\cap f^{-1}(B)\) only within a null set, so the measure is unaffected.
The \(C^1\) case. If \(g\in C^1(\mathbb{R}^n,\mathbb{R}^n)\), \(U:=\{\det Dg\ne0\}\) and \(\lambda^{*}(B)=0\), then \(\lambda^{*}(U\cap g^{-1}(B))=0\). Indeed \(U\) is open, and by the inverse function theorem every point of \(U\) has an open ball \(V\subset U\) on which \(g\) is injective with \(C^1\) inverse \(h:W\to V\), \(W=g(V)\) open; by Lindeloef’s theorem countably many such balls \(V_j\) cover \(U\), and \(V_j\cap g^{-1}(B)=h_j(B\cap W_j)\). Exhausting \(W_j\) by closed balls \(Q_{j,k}\), on each of which \(\|Dh_j\|\) is bounded, the mean value inequality on the convex \(Q_{j,k}\) makes \(h_j\) Lipschitz there, and a Lipschitz map carries null sets to null sets (the covering computation in the proof of Lemma 3.6.3). Summing over \(k\) and \(j\) gives the claim.
Approximation. Since \(\partial_{x_j}f_i\in L^1\), we have \(W^{1,1}(\mathbb{R}^n)\subset BV(\mathbb{R}^n)\), so Theorem 5.8.27 yields for each \(m\) and each \(i\le n\) a function \(g^{(m)}_i\in C^1(\mathbb{R}^n)\) with \(\lambda\{g^{(m)}_i\ne f_i\}\le4^{-m}/n\) and \(\|f_i-g^{(m)}_i\|_{W^{1,1}}\le4^{-m}/n\); classical and generalized derivatives of \(g^{(m)}_i\) agree a.e. Put \(E_m:=\{g^{(m)}=f\}\), so \(\lambda(\mathbb{R}^n\setminus E_m)\le4^{-m}\), and, by Chebyshev’s inequality,
\begin{equation*} D_m:=\Bigl\{\max_{i,j}\bigl|\partial_{x_j}f_i-\partial_{x_j}g^{(m)}_i\bigr| >2^{-m}\Bigr\}, \qquad\lambda(D_m)\le n\,2^{-m}. \end{equation*}
The sets \(R_m:=(\mathbb{R}^n\setminus E_m)\cup D_m\) have summable measures, so by Borel–Cantelli the Borel set \(G\) of points lying in only finitely many \(R_m\) has full measure. For \(x\in G\) and all large \(m\) we get \(f(x)=g^{(m)}(x)\) and \(Dg^{(m)}(x)\to Df(x)\) entrywise, hence \(\det Dg^{(m)}(x)\to\det Df(x)\), the determinant being a polynomial in the entries.
Conclusion. Put \(H_m:=\{x\in\Omega\cap G\cap E_m:\det Dg^{(m)}(x)\ne0\}\). Then \(\Omega\cap G=\bigcup_mH_m\), since for \(x\in\Omega\cap G\) both \(x\in E_m\) and \(\det Dg^{(m)}(x)\ne0\) hold for all large \(m\). For Borel \(B\) with \(\lambda(B)=0\), on \(H_m\) we have \(f=g^{(m)}\), so \(H_m\cap f^{-1}(B)\subset U_m\cap(g^{(m)})^{-1}(B)\) with \(U_m=\{\det Dg^{(m)}\ne0\}\), a null set by the \(C^1\) case; hence \(\lambda(\Omega\cap G\cap f^{-1}(B))=0\) and, \(G\) having full measure, \(\lambda(\Omega\cap f^{-1}(B))=0\).
For an arbitrary Lebesgue null set \(N\) take a Borel null \(B\supset N\); then \(\Omega\cap f^{-1}(N)\) is null by completeness. For Lebesgue measurable \(A\) write \(A=B_0\cup N_0\) with \(B_0\) Borel and \(N_0\) null (Corollary 1.5.8 applied on each cube \([-k,k]^n\)); then \(\Omega\cap f^{-1}(A)\) is measurable. So \(\lambda|_\Omega\circ f^{-1}\) is defined on the whole Lebesgue \(\sigma\)-algebra and vanishes on null sets, i.e. \(\lambda|_\Omega\circ f^{-1}\ll\lambda\).
Show that the assertion of the previous exercise (Exercise 5.8.140) remains true for any measurable functions \(f_i\) provided that \(\Omega\) is the set of points where the approximate partial derivatives \(\mathrm{ap}\,\partial_{x_j} f_i\) exist and \(\det(\mathrm{ap}\,\partial_{x_j} f_i)_{i,j\le n} \neq 0\).
The proof of Exercise 5.8.140 goes through with Whitney’s Theorem 5.8.14 in place of Theorem 5.8.27: we show that \(\lambda(B)=0\) implies \(\nu(B)=\lambda(\Omega\cap f^{-1}(B))=0\) for \(\nu:=\lambda|_\Omega\circ f^{-1}\), which is the meaning of absolute continuity here, \(\nu\) need not even be \(\sigma\)-finite (for \(n=1\) and \(f(x)=x-k\) on \([k,k+1)\) one has \(\Omega=\mathbb{R}\setminus\mathbb{Z}\) and \(\nu(B)=+\infty\) for every \(B\subset[0,1)\) of positive measure).
Measurability of \(\Omega\). Fix \(i,j\) and put \(q(x,t)=(f_i(x+te_j)-f_i(x))/t\), measurable on \(\mathbb{R}^n\times(\mathbb{R}\setminus\{0\})\) because \((x,t)\mapsto(x+te_j,t)\) is a linear automorphism of determinant \(1\). Writing \(q=q_0\) off a null set with \(q_0\) Borel and applying Fubini to a Borel null hull, for \(x\) outside a null set \(N\) the section \(q(x,\cdot)\) is measurable and agrees a.e. with \(q_0(x,\cdot)\); hence \(x\mapsto D_c(x,r):=\lambda_1\{t:0<|t|\le r,\ q(x,t)>c\}\) is measurable by Proposition 3.3.2(ii). As \(D_c(x,\cdot)\) is nondecreasing, the set \(\{q(x,\cdot)>c\}\) has density \(0\) at \(0\) exactly when \(2^kD_c(x,2^{-k})\to0\), so
\begin{equation*} E_c=\bigcap_{m}\bigcup_{K}\bigcap_{k\ge K} \bigl\{x\notin N:\ D_c(x,2^{-k})\le2^{-k}/m\bigr\} \end{equation*}
is measurable, and so are \(\overline{d}=\inf\{c\in\mathbb{Q}:x\in E_c\}\) and the companion \(\underline{d}\) built from the sets \(E^{\prime}_c\) with \(\{q<c\}\), since \(\{\overline{d}<s\}=\bigcup_{c<s}E_c\) and \(\{\underline{d}>s\}=\bigcup_{c>s}E^{\prime}_c\) by monotonicity in \(c\). For \(x\notin N\) the approximate derivative exists and equals \(L\in\mathbb{R}\) precisely when \(\overline{d}(x)=\underline{d}(x)=L\): if \(\mathrm{ap}\,\partial_{x_j}f_i(x)=L\), then \(\{q(x,\cdot)>c\}\) has density \(0\) for rational \(c>L\) and density \(1\) for rational \(c<L\); conversely, choosing rationals \(c\in(L,L+\delta)\), \(c^{\prime}\in(L-\delta,L)\) with \(x\in E_c\cap E^{\prime}_{c^{\prime}}\), the set \(\{|q(x,\cdot)-L|\ge\delta\}\subset\{q>c\}\cup\{q<c^{\prime}\}\) has density \(0\). So each \(\mathrm{ap}\,\partial_{x_j}f_i\) is a measurable function on a measurable set, \(\Omega\) is measurable, and, \(\lambda(N)=0\), we may assume \(N\cap\Omega=\emptyset\).
Whitney pieces. Fix \(\varepsilon>0\). All approximate partial derivatives exist at every point of \(\Omega\), so condition (ii) of Theorem 5.8.14 holds on \(\Omega\) and its implication (iii) gives, for each \(i\le n\), a closed \(F_i\subset\Omega\) and \(g_i\in C^1(\mathbb{R}^n)\) with \(\lambda(\Omega\setminus F_i)<\varepsilon/n\) and \(f_i=g_i\) on \(F_i\). Put \(F:=\bigcap_iF_i\), closed, with \(\lambda(\Omega\setminus F)<\varepsilon\) and \(f=g:=(g_1,\dots,g_n)\) on \(F\).
Derivatives on \(F\). Let \(F^j\) be the set of \(x\in F\) that are one-dimensional density points of \(F\) in the direction \(e_j\); \(F^j\) is measurable (Proposition 3.3.2(ii) again, \(F\) being closed, with the limit taken along \(r=2^{-k}\)), and \(\lambda(F\setminus F^j)=0\) by the one-dimensional density theorem on each line parallel to \(e_j\) together with Fubini’s theorem. For \(x\in F^j\) the function \(h=f_i-g_i\) vanishes on \(F\), so its difference quotients along \(e_j\) vanish on a set of density \(1\) at \(0\), giving \(\mathrm{ap}\,\partial_{x_j}h(x)=0\); approximate limits being additive and extending ordinary ones, \(\mathrm{ap}\,\partial_{x_j}f_i(x)=\partial_{x_j}g_i(x)\). Hence on \(F^{*}:=\bigcap_jF^j\), of full measure in \(F\),
\begin{equation*} J_g(x)=\det\bigl(\mathrm{ap}\,\partial_{x_j}f_i(x)\bigr)_{i,j\le n}\ne0 . \end{equation*}
Change of variables. \(U:=\{J_g\ne0\}\) is open and contains \(F^{*}\); by the inverse function theorem and Lindeloef’s theorem, \(U=\bigcup_mU_m\) with \(g\) injective on each open ball \(U_m\). For Borel \(B\) with \(\lambda(B)=0\) put \(A_m:=U_m\cap g^{-1}(B)\); Theorem 3.7.1 applied on \(U_m\) to \(I_B\) gives
\begin{equation*} \int_{A_m}|J_g(x)|\,dx=\int_{g(A_m)}I_B(y)\,dy\le\lambda(B)=0 , \end{equation*}
and \(|J_g|>0\) on \(U_m\) forces \(\lambda(A_m)=0\). So \(\lambda(U\cap g^{-1}(B))=0\), and since \(f=g\) on \(F\) with \(\lambda(F\setminus F^{*})=0\), also \(\lambda(F\cap f^{-1}(B))=0\).
Taking \(\varepsilon=1/k\) gives closed \(F_{(k)}\subset\Omega\) with \(\lambda(\Omega\setminus F_{(k)})<1/k\) and \(\lambda(F_{(k)}\cap f^{-1}(B))=0\); with \(G=\bigcup_kF_{(k)}\) we get \(\lambda(\Omega\setminus G)=0\) and hence \(\nu(B)=0\).
Exercise 5.8.142
(Bogachev, Kolesnikov [107]) Let \(U\) be an open ball in \(\mathbb{R}^d\) and let \(F\colon U\to\mathbb{R}^d\) be an integrable mapping such that its derivative \(DF\) in the sense of generalized functions is a bounded measure with values in the space of nonnegative symmetric matrices. Let \(D_{ac}F\) be the operator-valued density of the absolutely continuous component of \(DF\) and let \(\Omega:=\{x\colon \det D_{ac}F(x)>0\}\). Prove that the measure \(\lambda|_{\Omega}\circ F^{-1}\), where \(\lambda\) is Lebesgue measure, is absolutely continuous.
Write \(\mu:=DF=(\mu_{ij})\) with Lebesgue decomposition \(\mu=G\lambda+\mu_s\), \(G=D_{ac}F\); we must show \(\lambda(\Omega\cap F^{-1}(A))=0\) whenever \(A\) is Borel with \(\lambda(A)=0\). Modifying \(F\) on a null set changes nothing, \(\Omega\) being defined through \(\mu\).
Both parts of \(\mu\) are symmetric and nonnegative: \(\mu_{ij}=\mu_{ji}\) gives symmetry, and for each \(v\) the measure \(\mu_v:=\sum_{i,j}v_iv_j\mu_{ij}=\langle\mu(\cdot)v,v\rangle\) is nonnegative, hence so are both parts of its Lebesgue decomposition \(\langle Gv,v\rangle\lambda+\langle\mu_sv,v\rangle\), the decomposition being unique and linear; running \(v\) through a countable dense set gives \(G\ge0\) a.e. We also use that \(A\ge B\ge0\) symmetric implies \(\det A\ge\det B\) (for \(B>0\), \(\det A/\det B=\det(B^{-1/2}AB^{-1/2})\ge1\); in general apply this to \(A+\eta I\ge B+\eta I\) and let \(\eta\to0\)).
A convex potential. With \(\rho_\varepsilon\) a smooth mollifier, \(F_\varepsilon:=F*\rho_\varepsilon\) on \(U^{\varepsilon}\) has \(DF_\varepsilon=\mu*\rho_\varepsilon\), symmetric and nonnegative since \(\langle DF_\varepsilon(x)v,v\rangle=\int\rho_\varepsilon(x-y)\,\mu_v(dy)\ge0\). A smooth field with symmetric Jacobian on a ball is a gradient: \(\varphi_\varepsilon(x):=\int_0^1\langle F_\varepsilon(x_0+t(x-x_0)),x-x_0\rangle\,dt\) satisfies \(\nabla\varphi_\varepsilon=F_\varepsilon\) and \(D^2\varphi_\varepsilon=DF_\varepsilon\ge0\), so \(\varphi_\varepsilon\) is convex; normalize \(\int_{B_1}\varphi_\varepsilon\,dx=0\) for a fixed closed ball \(B_1\subset U\). The Poincare inequality on a ball \(V\supset B_1\), \(\overline V\subset U\), applied to \(u=\varphi_\varepsilon-\varphi_{\varepsilon^{\prime}}\) together with \(|\langle u\rangle_V|\le\lambda(B_1)^{-1}\|u-\langle u\rangle_V\|_{L^1(V)}\), gives \(\|\varphi_\varepsilon-\varphi_{\varepsilon^{\prime}}\|_{L^1(V)}\le C\|F_\varepsilon-F_{\varepsilon^{\prime}}\|_{L^1(V)}\); since \(F_\varepsilon\to F\) in \(L^1(V)\), the \(\varphi_\varepsilon\) converge in \(L^1_{loc}(U)\) to some \(\varphi_0\), and passing to the limit in \(\int\varphi_\varepsilon\partial_{x_i}\psi\,dx=-\int(F_i)_\varepsilon\psi\,dx\) shows the generalized gradient of \(\varphi_0\) is \(F\).
Convex functions with locally bounded \(L^1\) norms are locally uniformly bounded and equi-Lipschitz: Jensen bounds \(u\) above by its local mean, reflection through a point where \(|u|\) is below its mean bounds it below, and a convex function bounded by \(C\) on \(B(z,2r)\) is \(2C/r\)-Lipschitz on \(B(z,r)\). So by Arzela–Ascoli \(\varphi_\varepsilon\to\varphi\) locally uniformly, with \(\varphi\) convex and \(\varphi=\varphi_0\) a.e. Being locally Lipschitz, \(\varphi\) is differentiable on a set \(D\) of full measure (Rademacher) and its generalized partial derivatives are the pointwise ones (integrate by parts along lines and use Fubini), so \(\nabla\varphi=F\) a.e. and we take \(F=\nabla\varphi\) on \(D\); for \(x\in D\) the subdifferential is \(\partial\varphi(x)=\{\nabla\varphi(x)\}\).
Monge–Ampere inequality: for every compact \(K\subset U\),
\begin{equation*} \lambda\bigl(\partial\varphi(K)\bigr)\ \ge\ \int_K\det G(x)\,dx . \tag{2} \end{equation*}
Here \(\partial\varphi(K)\) is compact, being bounded by the local Lipschitz constant of \(\varphi\) and closed by passage to the limit in the subgradient inequality.
(a) For every open \(W\supset\partial\varphi(K)\) one has \(\nabla\varphi_\varepsilon(K)\subset W\) for all small \(\varepsilon\): otherwise \(y_j=\nabla\varphi_{\varepsilon_j}(x_j)\notin W\) with \(x_j\in K\), and by equi-Lipschitzness we may assume \(x_j\to x\in K\), \(y_j\to y\notin W\); passing to the limit in \(\varphi_{\varepsilon_j}(z)\ge\varphi_{\varepsilon_j}(x_j)+\langle y_j,z-x_j\rangle\) gives \(y\in\partial\varphi(x)\subset W\). With outer regularity, \(\limsup_{\varepsilon\to0}\lambda(\nabla\varphi_\varepsilon(K))\le\lambda(\partial\varphi(K))\).
(b) On \(P:=\{x\in K:\det D^2\varphi_\varepsilon(x)>0\}\) the map \(\nabla\varphi_\varepsilon\) is injective: equal gradients \(y\) at \(x\ne x^{\prime}\) make the convex \(\varphi_\varepsilon-\langle y,\cdot\rangle\) constant on \([x,x^{\prime}]\), so \(\langle D^2\varphi_\varepsilon(x)u,u\rangle=0\) for \(u=(x^{\prime}-x)/|x^{\prime}-x|\), and \(D^2\varphi_\varepsilon(x)\ge0\) then forces \(\det D^2\varphi_\varepsilon(x)=0\). Extending \(\nabla\varphi_\varepsilon\) from a closed ball \(K^{\prime}\subset U^{\varepsilon}\) with \(K\) in its interior to a Lipschitz map \(g\) of \(\mathbb{R}^d\), the area formula (Theorem 5.8.29(i) with \(n=k=d\), \(A=P\)) gives
\begin{equation*} \int_P\det D^2\varphi_\varepsilon\,dx =\int_{\mathbb{R}^d}\mathrm{Card}\bigl(P\cap g^{-1}(y)\bigr)dy =\lambda\bigl(\nabla\varphi_\varepsilon(P)\bigr), \end{equation*}
whence \(\int_K\det D^2\varphi_\varepsilon\,dx\le\lambda(\nabla\varphi_\varepsilon(K))\), the integrand vanishing on \(K\setminus P\).
(c) Since \(\mu_s\ge0\), we have \(D^2\varphi_\varepsilon=\mu*\rho_\varepsilon\ge G*\rho_\varepsilon\ge0\), so \(\det D^2\varphi_\varepsilon\ge\det(G*\rho_\varepsilon)\) by the monotonicity of the determinant; and \(G*\rho_\varepsilon\to G\) in \(L^1(K)\), so along a subsequence realizing the lower limit and converging a.e. Fatou’s theorem gives \(\liminf_\varepsilon\int_K\det(G*\rho_\varepsilon)\,dx\ge\int_K\det G\,dx\). Chaining (a), (b), (c) yields (2).
Conclusion. Let \(\lambda(A)=0\) and \(S:=\Omega\cap D\cap F^{-1}(A)\), which is measurable because the entries of \(G\) are Radon–Nikodym densities, hence measurable (Theorem 5.8.8 applied to the Jordan parts of each \(\mu_{ij}\)). If \(\lambda(S)>0\), inner regularity gives a compact \(K\subset S\) with \(\lambda(K)>0\); then \(\partial\varphi(K)=F(K)\subset A\), so the left-hand side of (2) vanishes, while \(\det G>0\) on \(K\subset\Omega\) gives \(\int_K\det G\,dx>0\), a contradiction. Hence \(\lambda(S)=0\), and since \(\lambda(U\setminus D)=0\) we get \(\lambda(\Omega\cap F^{-1}(A))=0\).
Backlinks (2)
1. Measure Theory (Bogachev) /words/library/books/measure_theory_bogachev/
V.I. Bogachev, Measure Theory, Springer, 2007. Two volumes.
Volume I covers constructions and extensions of measures, the Lebesgue integral, operations on measures and functions, the spaces \(L^p\) and spaces of measures, and connections between the integral and derivative. Volume II covers measures on topological spaces, weak convergence, transformations of measures and isomorphisms, conditional measures, and ergodic theory.
Solutions to the Volume I exercises live at Solutions to Bogachev’s Measure Theory, Volume 1.
2. Books /words/library/books/
Here are the books that I have taken the time to create metadata and/or notes for.
Comments