Wiener measure and Donsker’s theorem

Jordan Bell
September 4, 2015

1 Relatively compact sets of Borel probability measures on C[0,1]

Let E=C⁢[0,1], let ℬE be the Borel σ-algebra of E, and let 𝒫E be the collection of Borel probability measures on E. We assign 𝒫 the narrow topology, the coarsest topology on 𝒫E such that for each F∈Cb⁢(E) the map μ↦∫EF⁢𝑑μ is continuous.

For f∈E and δ>0 we define

ωf⁢(δ)=sups,t∈[0,1],|s-t|≤δ⁡|f⁢(s)-f⁢(t)|.

For f∈E, ωf⁢(δ)↓0 as δ↓0, and for δ>0, f↦ωf⁢(δ) is continuous. We shall use the following characterization of a relatively compact subset A of E, which is proved using the Arzelà-Ascoli theorem.

Lemma 1.

Let A be a subset of E. A¯ is compact if and only if

supf∈A⁡|f⁢(0)|<∞

and

supf∈A⁡ωf⁢(δ)↓0,δ↓0.

We shall use Prokhorov’s theorem:11 1 K. R. Parthasarathy, Probability Measures on Metric Spaces, p. 47, Chapter II, Theorem 6.7. for X a Polish space and for Γ⊂𝒫X, Γ¯ is compact if and only if for each ϵ>0 there is a compact subset Kϵ of X such that μ⁢(Kϵ)≥1-ϵ for all μ∈Γ. Namely, a subset of 𝒫X is relatively compact if and only if it is tight. We use Prokhorov’s theorem to prove a characterization of relatively compact subsets of 𝒫E, which we then use to prove the characterization in Theorem 3.22 2 K. R. Parthasarathy, Probability Measures on Metric Spaces, p. 213, Chapter VII, Lemma 2.2.

Lemma 2.

Let Γ be a subset of 𝒫E. Γ¯ is compact if and only if for each ϵ>0 there is some Mϵ<∞ and a function δ↦ωϵ⁢(δ) satisfying ωϵ⁢(δ)↓0 as δ↓0 and such that for all μ∈Γ,

μ⁢(Aϵ)≥1-ϵ2,μ⁢(Bϵ)≥1-ϵ2,

where

Aϵ={f∈E:|f⁢(0)|≤Mϵ},Bϵ={f∈E:ωf⁢(δ)≤ωϵ⁢(δ) for all δ>0}.
Proof.

Suppose that Γ satisfies the above conditions. Because f↦|f⁢(0)| is continuous, Aϵ is closed. For δ>0, suppose that fn is a sequence in Bϵ tending to some f∈E. Because g↦ωg⁢(δ) is continuous, ωfn⁢(δ)→ωf⁢(δ), and because ωfn⁢(δ)≤ωϵ⁢(δ) for each n, we get ωf⁢(δ)≤ωϵ⁢(δ) and hence f∈Bϵ, showing that Bϵ is closed. Therefore Kϵ=Aϵ∩Bϵ is closed, i.e. Kϵ=Kϵ¯. The set Kϵ satisfies

supf∈Kϵ⁡|f⁢(0)|≤Mϵ

and

lim supδ↓0⁡supf∈Kϵ⁡ωf⁢(δ)≤lim supδ↓0⁡ωϵ⁢(δ)=0,

thus by Lemma 1, Kϵ is compact. For μ∈Γ,

μ⁢(Kϵ)≥1-ϵ2,

and because Kϵ is compact, this means that Γ is tight, so by Prokhorov’s theorem, Γ is relatively compact.

Now suppose that Γ is relatively compact and let ϵ>0. By Prokhorov’s theorem, there is a compact set Kϵ in E such that μ⁢(Kϵ)≥1-ϵ2 for all μ∈Γ. Define

Mϵ=supf∈Kϵ⁡|f⁢(0)|,ωϵ⁢(δ)=supf∈Kϵ⁡ωf⁢(δ),δ>0.

Because Kϵ is compact, by Lemma 1 we get that Mϵ<∞ and ωϵ⁢(δ)↓0 as δ↓0. For μ∈Γ,

μ⁢(Aϵ)≥μ⁢(Kϵ)≥1-ϵ2,μ⁢(Bϵ)≥μ⁢(Kϵ)≥1-ϵ2,

showing that Γ satisfies the conditions of the theorem. ∎

We now prove the characterization of relatively compact subsets of 𝒫E that we shall use in our proof of Donsker’s theorem.33 3 K. R. Parthasarathy, Probability Measures on Metric Spaces, p. 214, Chapter VII, Theorem 2.2.

Theorem 3 (Relatively compact sets in 𝒫).

Let Γ be a subset of 𝒫E. Γ¯ is compact if and only if the following conditions are satisfied:

  1. 1.

    For each ϵ>0 there is some Mϵ<∞ such that

    μ(f:|f(0)|≤Mϵ)≥1-ϵ2,μ∈Γ.
  2. 2.

    For each ϵ>0 and δ>0 there is some η=η⁢(ϵ,δ)>0 such that

    μ(f:ωf(η)≤δ)≥1-ϵ2,μ∈Γ.
Proof.

Suppose that Γ¯ is compact and let ϵ>0. By Lemma 2, there is some Mϵ<∞ and a function η↦ωϵ⁢(η) satisfying ωϵ⁢(η)↓0 as η↓0 and

μ⁢(Aϵ)≥1-ϵ2,μ⁢(Bϵ)≥1-ϵ2,μ∈Γ.

For δ>0, there is some η=η⁢(ϵ,δ) with ωϵ⁢(η)≤δ. Then for μ∈Γ,

μ(f:ωf(η)≤δ)≥μ(f:ωf(η)≤ωϵ(η))≥μ(Bϵ)≥1-ϵ2.

Now suppose that the conditions of the theorem hold. For each ϵ>0 and n≥1 there is some ηϵ,n>0 such that

μ⁢(Fϵ,n)≥1-ϵ2n+1,μ∈Γ,

where

Fϵ,n={f:ωf⁢(ηϵ,n)≤1n}.

Let

Kϵ={f:|f⁢(0)|≤Mϵ}∩⋂n=1∞Fϵ,n,

for which

μ(Kϵ)≥μ(f:|f(0)|≤Mϵ)≥1-ϵ2,μ∈Γ.

For f∈Kϵ, then for each n≥1 we have f∈Fϵ,n, which means that ωf⁢(ηϵ,n)≤1n, and therefore

supf∈Kϵ⁡ωf⁢(ηϵ,n)≤1n.

Thus for n≥1, if 0<η≤ηϵ,n then

supf∈Kϵ⁡ωf⁢(η)≤1n,

which shows supf∈Kϵ⁡ωf⁢(η)↓0 as η↓0. Then because

supf∈Kϵ⁡|f⁢(0)|≤Mϵ,

applying Lemma 1 we get that Kϵ¯ is compact. The map f↦ωf⁢(ηϵ,n) is continuous, so the set Fϵ,n is closed, and therefore the set Kϵ is closed. Because Kϵ is compact and μ⁢(Kϵ)≥1-ϵ2 for all μ∈Γ, it follows from by Prokhorov’s theorem that Γ is relatively compact. ∎

2 Wiener measure

For t1,…,td∈[0,1], t1<⋯<td, define πt1,…,td:E→ℝd by

πt1,…,td⁢(f)=(f⁢(t1),…,f⁢(td)),f∈E,

which is continuous. We state the following results, which we will use later.44 4 K. R. Parthasarathy, Probability Measures on Metric Spaces, p. 212, Chapter VII, Theorem 2.1.

Theorem 4 (The Borel σ-algebra of E).

ℬE is equal to the σ-algebra generated by {πt:t∈[0,1]}.

Two elements μ and ν of 𝒫E are equal if and only if for any d and any t1<⋯<td, the pushforward measures

μt1,…,td=(πt1,…,td)*⁢μ,νt1,…,td=(πt1,…,td)*⁢ν

are equal.

Let (ξt)t∈[0,1] be a stochastic process with state space ℝ and sample space (Ω,ℱ,P). For t1<⋯<td, let ξt1,…,td=ξt1⊗⋯⊗ξtd and let Pt1,…,td=(ξt1,…,td)*⁢P: for B∈ℬℝd,

Pt1,…,td(B)=((ξt1,…,td)*P)(B)=P(ξt1,…,td-1(B))=P((ξt1,…,ξtd)∈B).

Pt1,…,td is a Borel probability measure on ℝd and is called a finite-dimensional distribution of the stochastic process.

The Kolmogorov continuity theorem55 5 K. R. Parthasarathy, Probability Measures on Metric Spaces, p. 216, Chapter VII, Theorem 3.1 tells us that if there are α,β,K>0 such that for all s,t∈[0,1],

E⁢|ξt-ξs|α≤K⁢|t-s|1+β,

then there is a unique μ∈𝒫E such that for all k and for all t1<⋯<td,

μt1,…,td=Pt1,…,td.

We now define and prove the existence of Wiener measure.66 6 K. R. Parthasarathy, Probability Measures on Metric Spaces, p. 218, Chapter VII, Theorem 3.2.

Theorem 5 (Wiener measure).

There is a unique Borel probability measure W on E satisfying:

  1. 1.

    W(f∈E:f(0)=0)=1.

  2. 2.

    For 0≤t0<t1<⋯<td≤1 the random variables

    πt1-πt0,πt2-πt1,πt3-πt2,πtd-πtd-1

    are independent (E,ℬE,W)→(ℝ,ℬℝ).

  3. 3.

    If 0≤s<t≤1, the random variable πt-πs:(E,ℬE,W)→(ℝ,ℬℝ) is normal with mean 0 and variance t-s.

Proof.

There is a stochastic process (ξt)t∈[0,1] with state space ℝ and some sample space (Ω,ℱ,P), such that (i) P(ξ0=0)=1, (ii) (ξt)t∈[0,1] has independent increments, and (iii) for s<t, ξt-ξs is a normal random variable with mean 0 and variance t-s. (Namely, Brownian motion with starting point 0.) Because ξt-ξs has mean 0 and variance t-s, we calculate (cf. Isserlis’s theorem)

E⁢|ξt-ξs|4=3⁢|t-s|2.

Thus using the Kolmogorov continuity theorem with α=4, β=1, K=3, there is a unique W∈𝒫E such that for all t1<⋯<td,

Wt1,…,td=Pt1,…,td,

i.e. for B∈ℬℝd,

W(πt1⊗⋯⊗πtd∈B)=P(ξt1⊗⋯⊗ξtd∈B).

For t1<⋯<td and B∈ℬℝd, with T:ℝd→ℝd defined by T⁢(x1,…,xd)=(x1,x2-x1,…,xd-xd-1),

W(πt1⊗(πt2-πt1)⊗⋯⊗(πtd-πtd-1)∈B)=W(T∘(πt1⊗πt2⊗⋯⊗πtd)∈B)=W(πt1⊗πt2⊗⋯⊗πtd∈T-1(B))=P(ξt1⊗ξt2⊗⋯⊗ξtd∈T-1(B))=P(T∘(ξt1⊗ξt2⊗⋯⊗ξtd)∈B)=P(ξt1⊗(ξt2-ξt1)⊗⋯⊗(ξtd-ξtd-1)∈B).

Hence, because ξt1,ξt2-ξt1,…,ξtd-ξtd-1 are independent,

(πt1⊗(πt2-πt1)⊗⋯⊗(πtd-πtd-1))*⁢W=(ξt1⊗(ξt2-ξt1)⊗⋯⊗(ξtd-ξtd-1))*⁢P=(ξt1)*⁢P⊗(ξt2-ξt1)*⁢P⊗⋯⊗(ξtd-ξtd-1)*⁢P=(πt1)*⁢W⊗(πt2-πt1)*⁢W⊗⋯⊗(πtd-πtd-1)*⁢W,

which means that the random variables πt1,πt2-πt1,…,πtd-πtd-1 are independent.

If s<t and B1,B2∈ℬℝ, and for T:ℝ2→ℝ2 defined by T⁢(x,y)=(x,y-x),

W((πs,πt-πs)∈(B1,B2)) =W(T∘(πs,πt)∈(B1,B2))
=P((ξs,ξt)∈T-1(B1,B2))
=P((ξs,ξt-ξs)∈(B1,B2)),

which implies that (πt-πs)*⁢W=(ξt-ξs)*⁢P, and because ξt-ξs is a normal random variable with mean 0 and variance t-s, so is πt-πs.

Finally,

W(f:f(0)=0)=W(π0=0)=P(ξ0=0)=1.

∎

(E,ℬE,W) is a probability space, and the stochastic process (πt)t∈[0,1] is a Brownian motion.

3 Interpolation and continuous stochastic processes

Let (ξt)t∈[0,1] be a continuous stochastic process with state space ℝ and sample space (Ω,ℱ,P). To say that the stochastic process is continuous means that for each ω∈Ω the map t↦ξt⁢(ω) is continuous [0,1]→ℝ. Define ξ:Ω→E by

ξ(ω)=(t↦ξt(ω)),ω∈Ω.

For t∈[0,1] and B a Borel set in ℝ,

ξ-1⁢πt-1⁢B={ω∈Ω:ξt⁢(ω)∈B}=ξt-1⁢B,

and because ξt:(Ω,ℱ)→(ℝ,ℬℝ) is measurable this belongs to ℱ. But by Theorem 4, ℬE is generated by the collection {πt-1⁢B:t∈[0,1],B∈ℬℝ}. Now, for f:X→Y and for a nonempty collection ℱ of subsets of Y,77 7 Charalambos D. Aliprantis and Kim C. Border, Infinite Dimensional Analysis: A Hitchhiker’s Guide, third ed., p. 140, Lemma 4.23.

σ⁢(f-1⁢(ℱ))=f-1⁢(σ⁢(ℱ)).

Therefore ξ-1⁢(ℬE)⊂ℱ, which means that ξ:(Ω,ℱ)→(E,ℬE) is measurable. This means that a continuous stochastic proess with index set [0,1] induces a random variable with state space E. Then the pushforward measure of P by ξ is a Borel probability measure on E. We shall end up constructing a sequence of pushforward measures from a sequence of continuous stochastic processes, that converge in 𝒫E to Wiener measure W.

Let (Xn)n≥1 be a sequence of independent identically distributed random variables on a sample space (Ω,ℱ,P) with E⁢(Xn)=0 and V⁢(Xn)=1, and let S0=0 and

Sk=∑i=1kXi.

Then E⁢(Sk)=0 and V⁢(Sk)=k. For t≥0 let

Yt=S[t]+(t-[t])⁢X[t]+1.

Thus, for k≥0 and k≤t≤k+1,

Yt =Sk+(t-k)⁢Xk+1
=Sk+(t-k)⁢(Sk+1-Sk)
=(1-t+k)⁢Sk+(t-k)⁢Sk+1.

For each ω∈Ω, the map t↦Yt⁢(ω) is piecewise linear, equal to Sk⁢(ω) when t=k, and in particular it is continuous. For n≥1, define

Xt(n)=n-1/2⁢Yn⁢t=n-1/2⁢S[n⁢t]+n-1/2⁢(n⁢t-[n⁢t])⁢X[n⁢t]+1,t∈[0,1]. (1)

For 0≤k≤n,

Xk/n(n)=n-1/2⁢Sk.

For each n≥1, (Xt(n))t∈[0,1] is a continuous stochastic process on the sample space (Ω,ℱ,P), and we denote by Pn∈𝒫E the pushforward measure of P by X(n).

4 Donsker’s theorem

Lemma 6.

If Zn and Un are random variables with state space ℝd such that Zn→Z in distribution and Un→0 in distribution, then Zn+Un→0 in distribution.

If Zn are random variables with state space ℝ that converge in distribution to some random variable Z and cn are real numbers that converge to some real number c, then cn⁢Zn→c⁢Z in distribution.

For σ≥0, let νσ2 be the Gaussian measure on ℝ with mean 0 and variance σ2. The characteristic function of νσ2 is, for σ>0,

ν~σ2⁢(ξ)=∫ℝei⁢ξ⁢x⁢𝑑νσ2⁢(x)=∫ℝei⁢ξ⁢x⁢1σ⁢2⁢π⁢e-x22⁢σ2⁢𝑑x=e-12⁢σ2⁢ξ2,

and ν~0⁢(ξ)=1. One checks that c*⁢ν1=νc2 for c≥0.

In following theorem and in what follows, X(n) is the piecewise linear stochastic process defined in (1). We prove that a sequence of finite-dimensional distributions converge to a Gaussian measure.88 8 Bert Fristedt and Lawrence Gray, A Modern Approach to Probability Theory, p. 368, §19.1, Lemma 1.

Theorem 7.

For 0≤t0<t1<t1<⋯<td≤1, the random vectors

(Xt1(n)-Xt0(n),…,Xtd(n)-Xtd-1(n)),(Ω,ℱ,P)→(ℝd,ℬℝd),

converge in distribution to νt1-t0⊗⋯⊗νtd-td-1 as n→∞.

Proof.

For 0<j≤d and n≥1 let

rj,n=[n⁢tj]n,Uj,n=Xtj(n)-Xrj,n(n),

and for 0≤j<d and n≥1 let

sj,n=⌈n⁢tj⌉n,Vj,n=Xsj,n(n)-Xtj(n),

with which

(Xt1(n)-Xt0(n),…,Xtd(n)-Xtd-1(n)) =(Xr1,n(n)-Xs0,n(n),…,Xrd,n(n)-Xsd-1,n(n))
+(U1,n,…,Ud,n)+(V0,n,…,Vd-1,n).

Because E⁢(Xt(n))=0,

E⁢(Uj,n)=0,E⁢(Vj,n)=0.

Furthermore,

V⁢(Uj,n)=V⁢(Xtj(n)-Xrj,n(n))=n-1⁢V⁢(S[n⁢tj]+(n⁢tj-[n⁢tj])⁢X[n⁢tj]+1-S[n⁢rj,n]-(n⁢rj,n-[n⁢rj,n])⁢X[n⁢rj,n]+1)=n-1⁢V⁢(S[n⁢tj]+(n⁢tj-[n⁢tj])⁢X[n⁢tj]+1-S[n⁢tj]-([n⁢tj]-[n⁢tj])⁢X[n⁢rj,n]+1)=n-1⁢(n⁢tj-[n⁢tj])2⁢V⁢(X[n⁢tj]+1)=n-1⁢(n⁢tj-[n⁢tj])2,

and because 0≤n⁢tj-[n⁢tj]<1 this tends to 0 as n→∞. Likewise, V⁢(Vj,n)→0 as n→∞.

For 1≤j≤d,

Xrj,n(n)-Xsj-1,n(n) =n-1/2⁢S[n⁢rj,n]+n-1/2⁢(n⁢rj,n-[n⁢rj,n])⁢X[n⁢rj,n]+1
-n-1/2⁢S[n⁢sj-1,n]-n-1/2⁢(n⁢sj-1,n-[n⁢sj-1,n])⁢X[n⁢sj-1,n]+1
=n-1/2⁢S[n⁢tj]-n-1/2⁢S⌈n⁢tj-1⌉
=n-1/2⁢([n⁢tj]-⌈n⁢tj-1⌉-1)1/2([n⁢tj]-⌈n⁢tj-1⌉-1)1/2⁢∑i=⌈n⁢tj-1⌉+1[n⁢tj]Xi.

By the central limit theorem,

([n⁢tj]-⌈n⁢tj-1⌉-1)1/2⁢∑i=⌈n⁢tj-1⌉+1[n⁢tj]Xi→ν1

in distribution as n→∞. But

n-1/2⁢([n⁢tj]-⌈n⁢tj-1⌉-1)1/2→(tj-tj-1)1/2

as n→∞, and (tj-tj-1)*1/2⁢ν1=νtj-tj-1, so by Lemma 6,

Xrj,n(n)-Xsj-1,n(n)→νtj-tj-1

in distribution as n→∞.

For sufficiently large n, depending on t0,…,td,

t0≤s0,n<r1,n≤t1≤s1,n<r2,n≤⋯≤td-1≤sd-1,n<rd,n≤td.

Check that (U1,n,…,Ud,n)→0 in probability and that (V0,n,…,Vd-1,n)→0 in probability, and hence these random vectors converge to 0 in distribution as n→∞. The random variables Xr1,n(n)-Xs0,n(n),…,Xrd,n(n)-Xsd-1,n(n) are independent, and therefore their joint distribution is equal to the product of their distributions. Now, if μn=μn1⊗⋯⊗μnd and μnj→μj as n→∞, 1≤j≤d, then for ξ∈ℝd,

μ~n⁢(ξ) =μ~n1⁢(ξ1)⁢⋯⁢μ~nd⁢(ξd)
→μ~1⁢(ξ1)⁢⋯⁢μ~d⁢(ξd)
=(μ1⊗⋯⊗μd)~⁢(ξ)

as n→∞, and therefore by Lévy’s continuity theorem, μn→μ1⊗⋯⊗μd as n→∞. This means that the joint distribution of Xr1,n(n)-Xs0,n(n),…,Xrd,n(n)-Xsd-1,n(n) converges to

νt1-t0⊗⋯⊗νtd-td-1

as n→∞. Because (U1,n,…,Ud,n)→0 in distribution as n→∞ and (V0,n,…,Vd-1,n)→0 in distribution as n→∞, applying Lemma 6 we get that

(Xt1(n)-Xt0(n),…,Xtd(n)-Xtd-1(n))→νt1-t0⊗⋯⊗νtd-td-1

in distribution as n→∞, completing the proof. ∎

Let t0=0 and let 0<t1<⋯<td≤1. As X0(n)=0, the above lemma tells us that

(Xt1(n),Xt2(n)-Xt1(n),…,Xtd(n)-Xtd-1(n))→νt1⊗νt2-t1⊗⋯⊗νtd-td-1

in distribution as n→∞. Define g:ℝd→ℝd by

g⁢(x1,x2,…,xd)=(x1,x1+x2,…,x1+x2+⋯+xd).

The function g is continuous and satisfies

g∘(Xt1(n)-Xt0(n),…,Xtd(n)-Xtd-1(n))=(Xt1(n),Xt2(n),…,Xtd(n)).

Then by the continuous mapping theorem,

(Xt1(n),Xt2(n),…,Xtd(n))→g*⁢(νt1⊗νt2-t1⊗⋯⊗νtd-td-1) (2)

in distribution as n→∞.99 9 Allan Gut, Probability: A Graduate Course, second ed., p. 245, Chapter 5, Theorem 10.4.

We prove a result that we use to prove the next lemma, and that lemma is used in the proof of Donsker’s theorem.1010 10 Ioannis Karatzas and Steven E. Shreve, Brownian Motion and Stochastic Calculus, second ed., p. 68, Lemma 4.18.

Lemma 8.

For ϵ>0,

limδ↓0lim supn→∞1δP(max1≤j≤[n⁢δ]+1|Sj|>ϵn1/2)=0.
Proof.

For each δ>0, by the central limit theorem,

([n⁢δ]+1)-1/2⁢S[n⁢δ]+1→Z

in distribution as n→∞, where Z*⁢P=ν1. Because ([n⁢δ]+1)1/2(n⁢δ)1/2→1 as n→∞, by Lemma 6 we then get that

(n⁢δ)-1/2⁢S[n⁢δ]+1→Z

in distribution as n→∞. Now let λ>0, and there is a sequence ϕk in Cb⁢(ℝ) such that ϕk↓1(-∞,-λ]∪[λ,∞)=χλ pointwise as k→∞. For each k, writing X=S[n⁢δ]+1, using the change of variables formula,

P(|X|≥λ(nδ)1/2) =∫Ωχλ⁢(n⁢δ)1/2⁢(X⁢(ω))⁢𝑑P⁢(ω)
=∫Ωχλ⁢((n⁢δ)-1/2⁢X⁢(ω))⁢𝑑P⁢(ω)
≤∫Ωϕk⁢((n⁢δ)-1/2⁢X⁢(ω))⁢𝑑P⁢(ω)
=E⁢(ϕk⁢((n⁢δ)-1/2⁢X)).

Therefore, by the continuous mapping theorem,

lim supn→∞P(|S[n⁢δ]+1|≥λ(nδ)1/2) ≤limn→∞⁡E⁢(ϕk⁢((n⁢δ)-1/2⁢S[n⁢δ]+1))
=E⁢(ϕk∘Z).

Because ϕk↓χλ pointwise as k→∞, using the monotone convergence theorem and then using Chebyshev’s inequality,

E(ϕk∘Z)→E(χλ∘Z)=P(|Z|≥λ)≤λ-3E|Z|3.

We have established that for each λ>0,

lim supn→∞P(|S[n⁢δ]+1|≥λ(nδ)1/2)≤λ-3E|Z|3. (3)

Define

τ=min⁡{j≥1:|Sj|>n1/2⁢ϵ}.

For 0<δ<ϵ2/2, it is a fact that

P(max0≤j≤[n⁢δ]+1|Sj|>n1/2ϵ)≤P(|S[n⁢δ]+1|≥n1/2(ϵ-(2δ)1/2))+∑j=1[n⁢δ]P(|S[n⁢δ]+1|<n1/2(ϵ-(2δ)1/2)|τ=j)P(τ=j).

If τ⁢(ω)=j and |S[n⁢δ]+1⁢(ω)|<n1/2⁢(ϵ-(2⁢δ)1/2) then

|Sj⁢(ω)-S[n⁢δ]+1⁢(ω)|≥|Sj⁢(ω)|-|S[n⁢δ]+1⁢(ω)|>n1/2⁢ϵ-n1/2⁢(ϵ-(2⁢δ)1/2)=(2⁢n⁢δ)1/2.

But by Chebyshev’s inequality and the fact that the random variables X1,X2,… are independent with mean 0 and variance 1,

P(|Sj-S[n⁢δ]+1|>(2nδ)1/2)≤12⁢n⁢δE((Sj-S[n⁢δ]+1)2)=12⁢n⁢δ([nδ]-j)≤12,

so

P(|S[n⁢δ]+1(ω)|<n1/2(ϵ-(2δ)1/2)|τ=j)≤12.

Therefore,

P(max0≤j≤[n⁢δ]+1|Sj|>n1/2ϵ)≤P(|S[n⁢δ]+1|≥n1/2(ϵ-(2δ)1/2))+∑j=1[n⁢δ]12⋅P(τ=j)=P(|S[n⁢δ]+1|≥n1/2(ϵ-(2δ)1/2))+12P(τ≤[nδ])=P(|S[n⁢δ]+1|≥n1/2(ϵ-(2δ)1/2))+12P(max0≤j≤[n⁢δ]+1|Sj|>n1/2ϵ),

so

P(max0≤j≤[n⁢δ]+1|Sj|>n1/2ϵ)≤2P(|S[n⁢δ]+1|≥n1/2(ϵ-(2δ)1/2)).

Now using (3) with λ=(ϵ-(2⁢δ)1/2)⁢δ-1/2,

lim supn→∞P(|S[n⁢δ]+1|≥(ϵ-(2δ)1/2)δ-1/2(nδ)1/2)≤(ϵ-(2δ)1/2)-3δ3/2E|Z|3,

hence

lim supn→∞P(max0≤j≤[n⁢δ]+1|Sj|>n1/2ϵ)≤2(ϵ-(2δ)1/2)-3δ3/2E|Z|3.

Dividing both sides by δ and then taking δ↓0 we obtain the claim. ∎

We prove one more result that we use to prove Donsker’s theorem.1111 11 Ioannis Karatzas and Steven E. Shreve, Brownian Motion and Stochastic Calculus, second ed., p. 69, Lemma 4.19.

Lemma 9.

For T>0 and ϵ>0,

limδ↓0lim supn→∞P(max0≤k≤[n⁢T]+1max1≤j≤[n⁢δ]+1|Sj+k-Sk|>n1/2ϵ)=0.
Proof.

For 0<δ≤T, let m=⌈T/δ⌉, so T/m<δ≤T/(m-1). Then

limn→∞⁡[n⁢T]+1[n⁢δ]+1=Tδ<m,

so for all n≥nδ it is the case that [n⁢T]+1<([n⁢δ]+1)⁢m. Suppose that ω∈Ω is such that there are 1≤j≤[n⁢δ]+1 and 0≤k≤[n⁢T]+1 satisfying

|Sj+k⁢(ω)-Sk⁢(ω)|>n1/2⁢ϵ,

and then let p=[k/([n⁢δ]+1)], which satisfies 0≤p≤m-1 and

([n⁢δ]+1)⁢p≤k<([n⁢δ]+1)⁢(p+1).

Because 1≤j≤[n⁢δ]+1, either

([n⁢δ]+1)⁢p<k+j≤([n⁢δ]+1)⁢(p+1)

or

([n⁢δ]+1)⁢(p+1)<k+j<([n⁢δ]+1)⁢(p+2).

We separate the first case into the cases

|Sk⁢(ω)-S([n⁢δ]+1)⁢p⁢(ω)|>12⁢n1/2⁢ϵ

and

|Sj+k⁢(ω)-S([n⁢δ]+1)⁢p⁢(ω)|>12⁢n1/2⁢ϵ,

and we separate the second case into the cases

|Sk-S([n⁢δ]+1)⁢p⁢(ω)|>13⁢n1/2⁢ϵ,

and

|S([n⁢δ]+1)⁢p⁢(ω)-S([n⁢δ]+1)⁢(p+1)⁢(ω)|>13⁢n1/2⁢ϵ,

and

|S([n⁢δ]+1)⁢(p+1)⁢(ω)-S([n+δ]+1)⁢(p+2)⁢(ω)|>13⁢n1/2⁢ϵ.

It follows that1212 12 This should be worked out more carefully. In Karatzas and Shreve, there is m+1 where I have m.

{max1≤j≤[n⁢δ]+1⁡max0≤k≤[n⁢T]+1⁡|Sj+k-Sk|>n1/2⁢ϵ}⊂⋃p=0m-1{max1≤j≤[n⁢δ]+1|Sj+([n⁢δ]+1)⁢p-S([n⁢δ]+1)⁢p|>13n1/2ϵ}.

For 0≤p≤m-1,

P(max1≤j≤[n⁢δ]+1|Sj+([n⁢δ]+1)⁢p-S([n⁢δ]+1)⁢p|>13n1/2ϵ)≤P(max1≤j≤[n⁢δ]+1|Sj|>13n1/2ϵ),

so

P{max1≤j≤[n⁢δ]+1max0≤k≤[n⁢T]+1|Sj+k-Sk|>n1/2ϵ}≤∑p=0m-1P(max1≤j≤[n⁢δ]+1|Sj|>13n1/2ϵ)=mP(max1≤j≤[n⁢δ]+1|Sj|>13n1/2ϵ).

Lemma 8 tells us

limδ↓0lim supn→∞1δP(max1≤j≤[n⁢δ]+1|Sj|>13n1/2ϵ)=0,

and because m≤Tδ+1=T+δδ,

limδ↓0lim supn→∞P{max1≤j≤[n⁢δ]+1max0≤k≤[n⁢T]+1|Sj+k-Sk|>n1/2ϵ}=0,

proving the claim. ∎

In the following, Pn∈𝒫E denotes the pushforward measure of P by X(n), for X(n) defined in (1). We now prove Donsker’s theorem.1313 13 Ioannis Karatzas and Steven E. Shreve, Brownian Motion and Stochastic Calculus, second ed., p. 70, Theorem 4.20.

Theorem 10 (Donsker’s theorem).

Pn→W.

Proof.

We shall use Theorem 3 to prove that Γ={Pn:n≥1} is relatively compact in 𝒫E. For n≥1,

Pn(f∈E:|f(0)|=0)=P(ω∈Ω:|X0(n)(ω)|=0)=1,

thus the first condition of Theorem 3 is satisfied with Mϵ=0. For the second condition of Theorem 3 to be satisfied it suffices that for each ϵ>0,

limδ↓0lim supn→∞P(sup0≤s,t≤1,|s-t|≤δ|X(n)(s)-X(n)(t)|>ϵ)=0.

Now,

P(sup0≤s,t≤1,|s-t|≤δ|Xs(n)-Xt(n)|>ϵ)=P(sup0≤s,t≤n,|s-t|≤n⁢δ|Ys-Yt|>n1/2ϵ).

Also,

sup0≤s,t≤n,|s-t|≤n⁢δ⁡|Ys-Yt| ≤sup0≤s,t≤n,|s-t|≤n⁢δ⁡|Y-s-Yt|
≤max1≤j≤[n⁢δ]+1⁡max0≤k≤n+1⁡|Sj+k-Sk|,

so applying Lemma 9,

limδ↓0lim supn→∞P(sup0≤s,t≤1,|s-t|≤δ|Xs(n)-Xt(n)|>ϵ)≤limδ↓0lim supn→∞P(max1≤j≤[n⁢δ]+1max0≤k≤n+1|Sj+k-Sk|>n1/2ϵ)→0,

from which we get that Γ is tight in 𝒫E. ∎