The Kolmogorov continuity theorem, Hölder continuity, and the Kolmogorov-Chentsov theorem

Jordan Bell
June 11, 2015

1 Modifications

Let (Ω,ℱ,P) be a probability space, let I be a nonempty set, and let (E,ℰ) be a measurable space. A stochastic process with index set I and state space E is a family (Xt)t∈I of random variables Xt:(Ω,ℱ)→(E,ℰ). If X and Y are stochastic processes, we say that X is a modification of Y if for each t∈I,

P⁢{ω∈Ω:Xt⁢(ω)=Yt⁢(ω)}=1.
Lemma 1.

If X is a modification of Y, then X and Y have the same finite-dimensional distributions.

Proof.

For t1,…,tn∈I, let Ai∈ℰ for each 1≤i≤n, and let

A=⋂i=1nXti-1⁢(Ai)∈ℱ,B=⋂i=1nYti-1⁢(Ai)∈ℱ.

If ω∈A∖B then there is some i for which ω∉Yti-1⁢(Ai), and ω∈Xti-1⁢(Ai) so Xti⁢(ω)≠Yti⁢(ω). Therefore

A⁢△⁢B⊂⋃i=1n{ω∈Ω:Xti⁢(ω)≠Yti⁢(ω)}.

Because X is a modification of Y, the right-hand side is a union of finitely many P-null sets, hence is itself a P-null set. A and B each belong to ℱ, so P⁢(A⁢△⁢B)=0.11 1 We have not assumed that (Ω,ℱ,P) is a complete measure space, so we must verify that a set is measurable before speaking about its measure. Because P⁢(A⁢△⁢B)=0, P⁢(A)=P⁢(B), i.e.

P(Xt1∈A1,…,Xtn∈An)=P(Yt1∈A1,…,Ytn∈An).

This implies that22 2 http://individual.utoronto.ca/jordanbell/notes/finitedimdistributions.pdf

P*⁢(Xt1⊗⋯⊗Xtn)=P*⁢(Yt1⊗⋯⊗Ytn),

namely, X and Y have the same finite-dimensional distributions. ∎

2 Continuous modifications

Let E be a Polish space with Borel σ-algebra ℰ. A stochastic process (Xt)t∈ℝ≥0 is called continuous if for each ω∈Ω, the path t↦Xt⁢(ω) is continuous ℝ≥0→E.

A dyadic rational is an element of

D=⋃i=0∞2-i⁢ℤ.

The Kolmogorov continuity theorem gives conditions under which a stochastic process whose state space is a Polish space has a continuous modification.33 3 Heinz Bauer, Probability Theory, p. 335, Theorem 39.3. It was only after working through the proof given by Bauer that I realized that the statement is true when the state space is a Polish space rather than merely ℝd. In the proof I do not use that |⋅| is a norm on ℝd, and only use that d⁢(x,y)=|x-y| is a metric on ℝd, so it is straightforward to rewrite the proof. This is like the Sobolev lemma,44 4 Walter Rudin, Functional Analysis, second ed., p. 202, Theorem 7.25. which states that if f∈Hs⁢(ℝd) and s>k+d2, then there is some ϕ∈Ck⁢(ℝd) such that f=ϕ almost everywhere. It does not make sense to say that an element of a Sobolev space is itself Ck, because elements of Sobolev spaces are equivalence classes of functions, but it does make sense to say that there is a Ck version of this element.

Theorem 2 (Kolmogorov continuity theorem).

Suppose that (Ω,F,P,(Xt)t∈R≥0) is a stochastic process with state space Rd. If there are α,β,c>0 such that

E⁢(|Xt-Xs|α)≤c⁢|t-s|1+β,s,t∈ℝ≥0, (1)

then the stochastic process has a continuous modification that itself satisfies (1).

Proof.

Let 0<γ<βα and let

δ=β-α⁢γ>0.

For m≥1, let Sm be the set of all pairs (s,t) with

s,t∈{j⁢2-m:0≤j≤2m},

and |s-t|=2-m. There are 2⋅2m such pairs, i.e. |Sm|=2⋅2m. Let

Am=⋃(s,t)∈Sm{|Xs-Xt|≥2-γ⁢m}∈ℱ.

For (s,t)∈Sm, using Chebyshev’s inequality and (1) we get

P(|Xt-Xs|≥2-γ⁢m) ≤(2γ⁢m)α⁢E⁢(|Xt-Xs|α)
≤2α⁢γ⁢m⋅c⁢|t-s|1+β
=c⁢2α⁢γ⁢m⁢2-m⁢(1+β)
<c⁢2-m-δ⁢m.

Hence

P(Am)≤∑(s,t)∈SmP{|Xs-Xt|≥2-γ⁢m}<∑(s,t)∈Smc2-m-δ⁢m=2c⋅2-δ⁢m.

Because ∑mP⁢(Am)<∞, the Borel-Cantelli lemma tells us that

P⁢(⋂n=1∞⋃m=n∞Am)=P⁢(N0)=0,

where for each ω∈Ω∖N0 there is some m0⁢(ω) such that ω∉Am when m≥m0⁢(ω). That is, for ω∈Ω∖N0 there is some m0⁢(ω) such that

|Xt⁢(ω)-Xs⁢(ω)|<2-γ⁢m,m≥m0⁢(ω),(s,t)∈Sm. (2)

Now let ω∈Ω∖N0 and let s,t∈[0,1] be dyadic rationals satisfying

0<|s-t|≤2-m0⁢(ω).

Let m=m⁢(s,t) be the greatest integer such that |s-t|≤2-m:

2-m-1<|s-t|≤2-m, (3)

which implies that m≥m0⁢(ω). There are some i0,j0∈{0,1,2,3,…,2m} such that

s0=i0⁢2-m≤s<(i0+1)⁢2-m,t0=j0⁢2-m≤t<(j0+1)⁢2-m.

As 0≤s-s0<2-m and 0≤t-t0<2-m, there are sequences σj,τj∈{0,1}, j>m, each of which have cofinitely many zero entries, such that

s=s0+∑j>mσj⁢2-j,t=t0+∑j>mτj⁢2-j.

Because 0≤s-s0<2-m and ≤t-t0<2-m,

2-m>|(s-s0)-(t-t0)|=|(s-t)-(s0-t0)|≥|s0-t0|-|s-t|,

and with (3),

|s0-t0|<2-m+|s-t|≤2-m+2-m=2-m+1.

Thus |i0-j0|<2, so |i0-j0|∈{0,1} and so either s0=t0 or (s0,t0)∈Sm. In the first case, |Xt0⁢(ω)-Xs0⁢(ω)|=0. In the second case, since m≥m0⁢(ω), by (2) we have

|Xt0⁢(ω)-Xs0⁢(ω)|<2-γ⁢m. (4)

Define by induction

sn=sn-1+σm+n⁢2-(m+n),n≥1,

i.e.

sn=s0+∑m<j≤m+nσj⁢2-j.

For each n≥1, sn-sn-1∈{0,2-(m+n)}, so either sn=sn-1 or (sn-1,sn)∈Sm+n, and because m+n≥m+1>m0⁢(ω), applying (2) yields

|Xsn⁢(ω)-Xsn-1⁢(ω)|<2-γ⁢(m+n).

Because the sequence σj is eventually equal to 0, the sequence sn is eventually equal to s. Thus

∑n=1∞(Xsn⁢(ω)-Xsn-1⁢(ω))=Xs⁢(ω)-Xs0⁢(ω),

whence

|Xs⁢(ω)-Xs0⁢(ω)|≤∑n=1∞|Xsn⁢(ω)-Xsn-1⁢(ω)|<∑n=1∞2-γ⁢(m+n)=2-γ⁢(m+1)1-2-γ.

By the same reasoning we get

|Xt⁢(ω)-Xt0⁢(ω)|<2-γ⁢(m+1)1-2-γ.

Using these and (4) yields

|Xt⁢(ω)-Xs⁢(ω)| ≤|Xt⁢(ω)-Xt0⁢(ω)|+|Xt0⁢(ω)-Xs0⁢(ω)|+|Xs⁢(ω)-Xs0⁢(ω)|
<2-γ⁢(m+1)1-2-γ+2-γ⁢m+2-γ⁢(m+1)1-2-γ
=C⋅2-γ⁢(m+1),

for C=2γ+21-2-γ. By (3), 2-(m+1)<|t-s|, hence

|Xt⁢(ω)-Xs⁢(ω)|≤C⁢|t-s|γ. (5)

This is true for all dyadic rationals s,t∈[0,1] with |s-t|≤2-m0⁢(ω); when |s-t|=0 it is immediate.

For k≥1, let Xtk=Xk+t, which satisfies (1). By what we have worked out, there is a P-null set N1′∈ℱ such that for each ω∈Ω∖N1′ there is some m1′⁢(ω) such that m≥m1′⁢(ω) and (s,t)∈Sm imply that |Xt1⁢(ω)-Xs1⁢(ω)|<2-γ⁢m. Let N1=N0∪N1′, which is P-null, and for ω∈Ω∖N1 let m1⁢(ω)=max⁡{m0⁢(ω),m1′⁢(ω)}. For s,t∈D∩[0,1] with |s-t|≤2-m1⁢(ω), what we have worked out yields

|Xt⁢(ω)-Xs⁢(ω)|≤C⁢|t-s|γ,|Xt1⁢(ω)-Xs1⁢(ω)|≤C⁢|t-s|γ.

By induction, we get that for each k≥1 there are P-null sets N0⊂N1⊂⋯⊂Nk and for each ω∈Ω∖Nk there is some mk⁢(ω) such that for s,t∈D∩[0,1] with |s-t|≤2-mk⁢(ω),

|Xt⁢(ω)-Xs⁢(ω)| ≤C⁢|t-s|γ
|Xt1⁢(ω)-Xs1⁢(ω)| ≤C⁢|t-s|γ
…
|Xtk⁢(ω)-Xsk⁢(ω)| ≤C⁢|t-s|γ.

Let

Nγ=⋃k≥1Nk,

which is an increasing sequence of sets whose union is P-null. For ω∈Ω∖Nγ, there is a nondecreasing sequence mk⁢(ω) such that when 0≤j≤k and s,t∈D∩[j,j+1] with |s-t|≤2-mk⁢(ω), it is the case that |Xt⁢(ω)-Xs⁢(ω)|≤C⁢|t-s|γ. For s,t∈D∩[0,k+1] with |s-t|≤2-mk⁢(ω), because |s-t|≤12, either there is some 0≤j≤k for which s,t∈[j,j+1] or there is some 1≤j≤k for which, say, s<j<t. In the first case, |Xt⁢(ω)-Xs⁢(ω)|≤C⁢|t-s|γ. In the second case, because |j-s|<|t-s|≤2-mk⁢(ω) and |t-j|<|t-s|≤2-mk⁢(ω), we get, because s,j∈D∩[j-1,j] and j,t∈D∩[j,j+1],

|Xt⁢(ω)-Xs⁢(ω)| ≤|Xt⁢(ω)-Xj⁢(ω)|+|Xj⁢(ω)-Xs⁢(ω)|
≤C⁢|t-j|γ+C⁢|j-s|γ
<2⁢C⁢|t-s|γ.

Thus for

Cγ=2⁢C=2γ+1+41-2-γ,

we have established that for ω∈Ω∖Nγ, k≥1, and s,t∈D∩[0,k+1] satisfying |t-s|≤2-mk⁢(ω), it is the case that

|Xt⁢(ω)-Xs⁢(ω)|≤Cγ⁢|t-s|γ. (6)

This implies that for each ω∈Ω∖Nγ and for k≥1, the mapping t↦Xt⁢(ω) is uniformly continuous on D∩[0,k+1]. For t∈ℝ≥0 and ω∈Ω∖Nγ, define

Yt⁢(ω)=lims∈Ds→t⁡Xs⁢(ω). (7)

For each k≥0, because t↦Xt⁢(ω) is uniformly continuous D∩[0,k+1]→ℝd, where D∩[0,k+1] is dense in [0,k+1] and ℝd is a complete metric space, the map t↦Yt⁢(ω) is uniformly continuous [0,k+1]→ℝd.55 5 Charalambos D. Aliprantis and Kim C. Border, Infinite Dimensional Analysis: A Hitchhiker’s Guide, third ed., p. 77, Lemma 3.11. Then t↦Yt⁢(ω) is continuous ℝ≥0→ℝd. For ω∈Nγ, we define

Yt⁢(ω)=0,t∈ℝ≥0.

Then for each ω∈Ω, t↦Yt⁢(ω) is continuous ℝ≥0→ℝd. For t∈ℝ≥0, ω↦Yt⁢(ω) is the pointwise limit of the sequence of mappings ω↦Xs⁢(ω) as s→t, s∈D. For each s∈D, ω↦Xs⁢(ω) is measurable ℱ→ℬℝd, which implies that ω↦Yt⁢(ω) is itself measurable ℱ→ℬℝd.66 6 Charalambos D. Aliprantis and Kim C. Border, Infinite Dimensional Analysis: A Hitchhiker’s Guide, third ed., p. 142, Lemma 4.29. Namely, (Yt)t∈ℝ≥0 is a continuous stochastic process.

We must show that Y is a modification of X. For s∈D, for all ω∈Ω∖Nγ we have Ys⁢(ω)=Xs⁢(ω). For t∈ℝ≥0, there is a sequence sn∈D tending to t, and then for all ω∈Ω∖Nγ by (7) we have Xsn⁢(ω)→Yt⁢(ω). P⁢(Nγ)=0, namely, Xsn converges to Yt almost surely. Because Xsn converges to Yt almost surely and P is a probability measure, Xsn converges in measure to Yt.77 7 Charalambos D. Aliprantis and Kim C. Border, Infinite Dimensional Analysis: A Hitchhiker’s Guide, third ed., p. 479, Theorem 13.37. On the other hand, for η>0, by Chebyshev’s inequality and (1),

P{|Xsn-Xt|≥η}≤η-αE(|Xsn-Xt|α)≤η-α⋅c|sn-t|1+β,

and because this is true for each η>0, this shows that Xsn converges in measure to Xt. Hence, the limits Yt and Xt are equal as equivalence classes of measurable functions Ω→ℝd.88 8 http://individual.utoronto.ca/jordanbell/notes/L0.pdf That is, P{Yt=Xt}=1. This is true for each t∈ℝ≥0, showing that Y is a modification of X, completing the proof. ∎

3 Hölder continuity

Let (X,d) and (Y,ρ) be metric spaces, let 0<γ<1, and let ϕ:X→Y be a function. For x0∈X, we say that ϕ is γ-Hölder continuous at x0 if there is some 0<ϵx0<1 and some Cx0 such that when d⁢(x,x0)<ϵx0,

ρ⁢(ϕ⁢(x),ϕ⁢(x0))≤Cx0⁢d⁢(x,x0)γ.

We say that ϕ is locally γ-Hölder continuous if for each x0∈X there is some 0<ϵx0<1 and some Cx0 such that when d⁢(x,x0)<ϵx0 and d⁢(y,x0)<ϵx0,

ρ⁢(ϕ⁢(x),ϕ⁢(y))≤Cx0⁢d⁢(x,y)γ.

We say that ϕ is uniformly γ-Hölder continuous if there is some C such that for all x,y∈X,

ρ⁢(ϕ⁢(x),ϕ⁢(y))≤C⁢d⁢(x,y)γ.

We establish properties of Hölder continuous functions in the following.99 9 Achim Klenke, Probability Theory: A Comprehensive Course, p. 448, Lemma 21.3.

Lemma 3.

Let V be a nonempty subset of R≥0, let 0<γ<1, and let f:V→Rd be locally γ-Hölder continuous.

  1. 1.

    If 0<γ′<γ then f is locally γ′-Hölder continuous.

  2. 2.

    If V is compact, then f is uniformly γ-Hölder continuous.

  3. 3.

    If V is an interval of length T>0 and there is some ϵ>0 and some C such that for all s,t∈V with |t-s|≤ϵ we have

    |f⁢(t)-f⁢(s)|≤C⁢|t-s|γ, (8)

    then

    |f⁢(t)-f⁢(s)|≤C⁢⌈Tϵ⌉1-γ⁢|t-s|γ,s,t∈V.
Proof.

For t0∈ℝ≥0, there is some 0<ϵt0<1 and some Ct0 such that when |t-t0|<ϵt0,

|f⁢(t)-f⁢(t0)|≤Ct0⁢|t-t0|γ≤Ct0⁢|t-t0|γ′,

showing that f is locally γ′-Hölder continuous.

With the metric inherited from ℝ≥0, V is a compact metric space. For t∈V and ϵ>0, write

Bϵ⁢(t)={v∈V:|v-t|<ϵ},

which is an open subset of V. Because f is locally γ-Hölder continuous, for each t∈V there is some 0<ϵt<1 and some Ct such that for all u,v∈Bϵt⁢(t),

|f⁢(u)-f⁢(v)|≤Ct⁢|u-v|γ. (9)

Write Ut=Bϵt⁢(t). Because t∈Ut, {Ut:t∈V} is an open cover of V, and because V is compact there are t1,…,tn∈V such that 𝔘={Ut1,…,Utn} is an open cover of V. Because V is a compact metric space, there is a Lebesgue number δ>0 of the open cover 𝔘:1010 10 Charalambos D. Aliprantis and Kim C. Border, Infinite Dimensional Analysis: A Hitchhiker’s Guide, third ed., p. 85, Lemma 3.27. for each t∈V, there is some 1≤i≤n such that Bδ⁢(t)⊂Uti. Let

C=max⁡{Ct1,…,Ctn,2⁢∥f∥u⁢δ-γ},

For s,t∈V with |t-s|<δ, i.e. s∈Bδ⁢(t), there is some 1≤i≤n with s,t∈Uti. By (9),

|f⁢(s)-f⁢(t)|≤Cti⁢|s-t|γ≤C⁢|s-t|γ.

On the other hand, for s,t∈V with |t-s|≥δ,

|f⁢(s)-f⁢(t)|≤2⁢∥f∥u≤2⁢∥f∥u⁢(|s-t|δ)γ=2⁢∥f∥u⁢δ-γ⁢|s-t|γ≤C⁢|s-t|γ.

Thus, for all s,t∈V,

|f⁢(s)-f⁢(t)|≤C⁢|s-t|γ,

showing that f is uniformly γ-Hölder continuous.

Let n=⌈Tϵ⌉. For s,t∈V, because V is an interval of length T, |s-t|≤T≤ϵ⁢n, and then applying (8), because |t-s|n≤ϵ,

|f⁢(t)-f⁢(s)| =|∑k=1nf⁢(s+(t-s)⁢kn)-f⁢(s+(t-s)⁢k-1n)|
≤∑k=1n|f⁢(s+(t-s)⁢kn)-f⁢(s+(t-s)⁢k-1n)|
≤∑k=1nC⁢|t-sn|γ
=C⁢n1-γ⁢|t-s|γ.

∎

The following theorem does not speak about a version of a stochastic process. Rather, it shows what can be said about a stochastic process that satisfies (1) when almost all of its sample paths are continuous.1111 11 Heinz Bauer, Probability Theory, p. 338, Theorem 39.4.

Theorem 4.

If a stochastic process (Xt)t∈R≥0 with state space Rd satisfies (1) and for almost every ω∈Ω the map t↦Xt⁢(ω) is continuous R≥0→Rd, then for almost every ω∈Ω, for every 0<γ<βα, the map t↦Xt⁢(ω) is locally γ-Hölder continuous.

Proof.

There is a P-null set N∈ℱ such that for ω∈Ω∖N, the map t↦Xt⁢(ω) is continuous ℝ≥0→ℝd. For each 0<γ<βα, we have established in (6) that there is a P-null set Nγ∈ℱ such that for k≥1 there is some mk⁢(ω) such that when s,t∈D∩[0,k+1] and |t-s|≤2-mk⁢(ω),

|Xt⁢(ω)-Xs⁢(ω)|≤Cγ⁢|t-s|γ, (10)

where Cγ=2γ+1+41-2-γ. Write δ⁢(k,ω)=2-mk⁢(ω), and let Mγ=Nγ∪N. For ω∈Ω∖Mγ, the map t↦Xt⁢(ω) is continuous ℝ≥0→ℝd. For k≥1 and for s,t∈[0,k+1] satisfying |s-t|≤δ⁢(k,ω), say with s≤t, let m=t-s2 and let s≤sn≤t be a sequence of dyadic rationals decreasing to s and let s≤tn≤t be a sequence of dyadic rationals inceasing to t. Then sn,tn∈D∩[0,k+1] and |sn-tn|≤|s-t|≤δ⁢(k,ω), so by (10),

|Xtn⁢(ω)-Xsn⁢(ω)|≤Cγ⁢|tn-sn|γ.

Because ω∈Ω∖N, Xtn⁢(ω)→Xt⁢(ω) and Xsn⁢(ω)→Xs⁢(ω), so

|Xt⁢(ω)-Xs⁢(ω)| ≤|Xt⁢(ω)-Xtn⁢(ω)|+|Xtn⁢(ω)-Xsn⁢(ω)|+|Xs⁢(ω)-Xsn⁢(ω)|
≤|Xt⁢(ω)-Xtn⁢(ω)|+Cγ⁢|tn-sn|γ+|Xs⁢(ω)-Xsn⁢(ω)|
↓Cγ⁢|t-s|γ,

thus

|Xt⁢(ω)-Xs⁢(ω)|≤Cγ⁢|t-s|γ,

showing that for 0<γ<βα and ω∈Ω∖Mγ, the map t↦Xt⁢(ω) is locally γ-Hölder continuous.

Let 0<γn<βα be a sequence increasing to βα and let

M=⋃n≥1Mγn,

which is a P-null set. Let 0<γ<βα and let n be such that γn≥γ. For ω∈Ω∖M, the map t↦Xt⁢(ω) is locally γn-Hölder continuous, and because γ≤γn this implies that the map is locally γ-Hölder continuous, completing the proof. ∎

Bauer attributes the following theorem to Kolgmorov and Chentsov.1212 12 Nikolai Nikolaevich Chentsov, 1930–1993, obituary in Russian Math. Surveys 48 (1993), no. 2, 161–166. It does not merely state that for any 0<γ<βα there is a modification that is locally γ-Hölder continuous, but that there is a modification that for all 0<γ<βα is locally γ-Hölder continuous.1313 13 Heinz Bauer, Probability Theory, p. 339, Corollary 39.5.

Theorem 5 (Kolmogorov-Chentsov theorem).

If a stochastic process (Xt)t∈R≥0 with state space Rd satisfies (1), then X has a modification Y such that for all ω∈Ω and 0<γ<βα, the path t↦Yt⁢(ω) is locally γ-Hölder continuous.

Proof.

Applying the Kolmogorov continuity theorem, there is a continuous modification Z of X that also satisfies (1). By Theorem 4, there is a P-null set M such that for ω∈Ω∖M and 0<γ<βα, the map t↦Zt⁢(ω) is locally γ-Hölder continuous. For t∈ℝ≥0, define

Yt⁢(ω)={Zt⁢(ω)ω∈Ω∖M0ω∈M,

i.e. Yt=1Ω∖M⁢Zt, which is measurable ℱ→ℬℝd, and so (Yt)t∈ℝ≥0 is a stochastic process. For every ω∈Ω and 0<γ<βα, the map t↦Yt⁢(ω) is locally γ-Hölder continuous. For t∈ℝ≥0,

{Xt≠Yt}={Xt≠Yt,Xt=Zt}∪{Xt≠Yt,Xt≠Zt}⊂{Yt≠Zt}∪{Xt≠Zt}.

Because P(Yt≠Zt)=P(M)=0 and P(Xt≠Zt)=0, since Z is a modification of X, we get P(Xt≠Yt)=0, namely, Y is a modification of X. ∎