Law of Large Numbers
Weak Law of Large Numbers
Let \((X_n)_{n\ge 1}\) be i.i.d. random variables with finite mean \(m\), and set
The weak law of large numbers says that
Equivalently, for every \(\varepsilon>0\),
Two sequences of random variables \((X_n)\) and \((Y_n)\) are called equivalent if
By the Borel-Cantelli lemma, equivalent sequences differ only finitely many times almost surely.
Lemma. If \((X_n)\) and \((Y_n)\) are equivalent, then
Thus equivalent truncations may be used to prove limit theorems.
Strong Law of Large Numbers
Let \((X_n)_{n\ge 1}\) be i.i.d. and set \(S_n=X_1+\cdots+X_n\). Then we have:
Truncation Method
For \(A>0\), define
Truncation separates the proof into two parts:
-
Control the large jumps:
\[ \sum_n \mathbb{P}(|X_n|>A_n)<\infty. \] -
Prove convergence for bounded variables, usually by variance estimates or Kolmogorov inequalities.
This is the standard route from bounded laws to integrable laws.
Three-Series Theorem
Let \((X_n)\) be independent real random variables and fix \(A>0\). Put
Then \(\sum_n X_n\) converges almost surely if and only if the following three series converge:
The value of \(A\) is not essential; changing \(A\) gives an equivalent criterion.
Proof. We first establish the maximal inequalities used in the proof.
Lemma 1: Let \(Z_1,\ldots,Z_n\) be independent random variables satisfying
and put \(R_k=\sum_{j=1}^k Z_j\). Then, for every \(\varepsilon>0\),
For \(1\le k\le n\), define the first-exit events
The events \(E_k\) are pairwise disjoint and
Since \(E_k\) and \(R_k\) depend only on \(Z_1,\ldots,Z_k\), while \(R_n-R_k\) is independent of them and has mean \(0\),
Summing over \(k\),
which proves the inequality.
As a consequence, if \((Z_n)\) are independent and centered and
then \(\sum_n Z_n\) converges almost surely. Indeed, choose \(N_r\uparrow\infty\) such that
Kolmogorov's inequality gives
The right-hand side is summable, so the Borel--Cantelli lemma shows that the partial sums of \(\sum_n Z_n\) satisfy the Cauchy criterion almost surely.
Lemma 2: Let \(Z_1,\ldots,Z_n\) be independent centered random variables satisfying \(|Z_j|\le C\) a.s., and put \(R_k=\sum_{j=1}^k Z_j\). Then
This centered form is sufficient for the proof below and is slightly sharper than the corresponding estimate stated in the slides.
Let
and let \(E_k\) be the event that \(R_k\) is the first partial sum to leave \([-B,B]\). On \(E_k\),
As in the proof of Lemma 1,
Also, \(|R_n|\le B\) on \(E\). Therefore,
Hence
which gives the desired inequality.
Lemma 3: Let \((U_n)\) be independent and suppose that \(|U_n|\le A\) a.s. If
converges almost surely, then
and the numerical series
converges.
Let \((U_n')\) be an independent copy of \((U_n)\), defined on a product probability space, and set
Then the \(Z_n\)'s are independent and centered,
and \(\sum_n Z_n\) converges almost surely.
Consequently,
Thus, for all sufficiently large \(m\) and every \(n>m\),
Applying Lemma 2 to \(Z_{m+1},\ldots,Z_n\),
Hence
Letting \(n\to\infty\),
and therefore
By Lemma 1,
converges almost surely. Since \(\sum_nU_n\) also converges almost surely, their difference
converges. Hence \(\sum_n\mathbb E[U_n]\) converges.
"\(\Longleftarrow\)"
Assume that
By the first Borel--Cantelli lemma,
so \(X_n=Y_n\) eventually almost surely.
Set
Then the \(Z_n\)'s are independent and centered, and
By Lemma 1, \(\sum_nZ_n\) converges almost surely. Since \(\sum_n\mathbb E[Y_n]\) converges,
converges almost surely. Since \(X_n=Y_n\) eventually,
also converges almost surely.
"\(\Longrightarrow\)"
Assume that
converges almost surely. Then \(X_n\to0\) almost surely, so
The events \(\{|X_n|>A\}\) are independent. By the second Borel--Cantelli lemma,
Moreover, \(X_n=Y_n\) eventually almost surely, so \(\sum_nY_n\) converges almost surely. Since \(|Y_n|\le A\), Lemma 3 gives
and
converges. Thus all three series converge.
Therefore,
Finally, if the criterion holds for one value of \(A>0\), the series \(\sum_nX_n\) converges almost surely; the necessity part then shows that the criterion holds for every \(A>0\).
Maximal Growth
For i.i.d. nonnegative random variables, the behavior of
is controlled by the tail of \(X_1\). In many applications, the largest summand explains why a law of large numbers fails when \(\mathbb{E}|X_1|=\infty\).