Markov kernels, convolution semigroups, and projective families of probability measures

Jordan Bell
June 12, 2015

1 Transition kernels

For a measurable space (E,ℰ), we denote by ℰ+ the set of functions E→[0,∞] that are ℰ→ℬ[0,∞] measurable. It can be proved that if I:ℰ+→[0,∞] is a function such that (i) f=0 implies that I⁢(f)=0, (ii) if f,g∈ℰ+ and a,b≥0 then I⁢(a⁢f+b⁢g)=a⁢I⁢(f)+b⁢I⁢(g), and (iii) if fn is a sequence in ℰ+ that increases pointwise to an element f of ℰ+ then I⁢(fn) increases to I⁢(f), then there a unique measure μ on ℰ such that I⁢(f)=μ⁢f for each f∈ℰ+.11 1 Erhan Çinlar, Probability and Stochastics, p. 28, Theorem 4.21.

Let (E,ℰ) and (F,ℱ) be a measurable space. A transition kernel is a function

K:E×ℱ→[0,∞]

such that (i) for each x∈E, the function Kx:ℱ→[0,∞] defined by

B↦K⁢(x,B)

is a measure on ℱ, and (ii) for each B∈ℱ, the map

x↦K⁢(x,B)

is measurable ℰ→ℬ[0,∞].

If μ is a measure on ℰ, define

(K*⁢μ)⁢(B)=∫EK⁢(x,B)⁢𝑑μ⁢(x),B∈ℱ.

If Bn are pairwise disjoint elements of ℱ, then using that B↦K⁢(x,B) is a measure and the monotone convergence theorem,

(K*⁢μ)⁢(⋃nBn) =∫EK⁢(x,⋃nBn)⁢𝑑μ⁢(x)
=∫E∑nK⁢(x,Bn)⁢d⁢μ⁢(x)
=∑n∫EK⁢(x,Bn)⁢𝑑μ⁢(x)
=∑n(K*⁢μ)⁢(Bn),

showing that K*⁢μ is a measure on ℱ.

If f∈ℱ+, define K*⁢f:E→[0,∞] by

(K*⁢f)⁢(x)=∫Ff⁢(y)⁢𝑑Kx⁢(y),x∈E. (1)

For ϕ=∑j=1kbj⁢1Bj with bj≥0 and Bj∈ℱ, because x↦K⁢(x,Bj) is measurable ℰ→ℬ[0,∞] for each j,

(K*⁢ϕ)⁢(x)=∫F∑j=1kbj⁢1Bj⁢(y)⁢d⁢Kx⁢(y)=∑j=1kbj⁢Kx⁢(Bj)=∑j=1bj⁢K⁢(x,Bj),

is measurable ℰ→ℬ[0,∞]. For f∈ℱ+, there is a sequence of simple functions ϕn with 0≤ϕ1≤ϕ2≤⋯ that converges pointwise to f,22 2 Gerald B. Folland, Real Analysis: Modern Techniques and Their Applications, second ed., p. 47, Theorem 2.10. and then by the monotone convergence theorem, for each x∈E we have

(K*⁢ϕn)⁢(x)=∫Fϕn⁢(y)⁢𝑑Kx⁢(y)→∫Ff⁢(y)⁢𝑑Kx⁢(y)=(K*⁢f)⁢(x),

showing K*⁢ϕn converges pointwise to K*⁢f, and because each K*⁢ϕn is measurable ℰ→ℬ[0,∞], K*⁢f is measurable ℰ→ℬ[0,∞].33 3 Gerald B. Folland, Real Analysis: Modern Techniques and Their Applications, second ed., p. 45, Proposition 2.7. Therefore, if f∈ℱ+ then K*⁢f∈ℰ+. In particular, if K is a transition kernel from (E,ℰ) to (F,ℱ),

(K*⁢1B)⁢(x)=∫F1B⁢(y)⁢𝑑Kx⁢(y)=Kx⁢(B)=K⁢(x,B),x∈E,B∈ℱ. (2)

The following gives conditions under which (2) defines a transition kernel.44 4 Heinz Bauer, Probability Theory, p. 308, Lemma 36.2.

Lemma 1.

Suppose that N:ℱ+→ℰ+ satisfies the following properties:

  1. 1.

    N⁢(0)=0.

  2. 2.

    N⁢(a⁢f+b⁢g)=a⁢N⁢(f)+b⁢N⁢(g) for f,g∈ℱ+ and a,b≥0.

  3. 3.

    If fn is a sequence in ℱ+ increasing to f∈ℱ+, then N⁢(fn)↑N⁢(f).

Then

K⁢(x,B)=(N⁢(1B))⁢(x),x∈E,B∈ℱ,

is a transition kernel from (E,ℰ) to (F,ℱ). K is the unique transition kernel satisfying

K*⁢f=N⁢(f),f∈ℱ+.

If K is a transition kernel from (E,ℰ) to (F,ℱ) and L is a transition kernel from (F,ℱ) to (G,𝒢), the function K*∘L*:𝒢+→ℰ+ satisfies (i) (K*∘L*)⁢(0)=K*⁢(0)=0, (ii) if f,g∈𝒢+ and a,b≥0,

(K*∘L*)⁢(a⁢f+b⁢g) =K*⁢(a⁢L*⁢(f)+b⁢L*⁢(g))
=a⁢K*⁢(L*⁢(f))+K*⁢(L*⁢(g))
=a⁢(K*∘L*)⁢(f)+b⁢(K*∘L*)⁢(g),

and (iii) if fn↑f in 𝒢+, then by the monotone convergence theorem, L*⁢(fn)↑L*⁢(f), and then again applying the monotone convergence theorem, K*⁢(L*⁢(fn))↑K*⁢(L*⁢(f)), i.e.

(K*∘L*)⁢(fn)↑(K*∘L*)⁢(f).

Therefore, from Lemma 1 we get that there is a unique transition kernel from (E,ℰ) to (G,𝒢), denoted K⁢L and called the product of K and L, such that

(K⁢L)*⁢f=(K*∘L*)⁢(f),f∈𝒢+.

For f∈𝒢+ and x∈E,

(K⁢L)*⁢(f)⁢(x) =(K*⁢(L*⁢f))⁢(x)
=∫F(L*⁢f)⁢(y)⁢𝑑Kx⁢(y)
=∫F(∫Gf⁢(z)⁢𝑑Ly⁢(z))⁢𝑑Kx⁢(y).

In particular, for C∈𝒢,

(K⁢L)*⁢(1C)⁢(x)=∫FLy⁢(C)⁢𝑑Kx⁢(y)=∫FL⁢(y,C)⁢𝑑Kx⁢(y). (3)

2 Markov kernels

A Markov kernel from (E,ℰ) to (F,ℱ) is a transition kernel K such that for each x∈E, Kx is a probability measure on ℱ. The unit kernel from (E,ℰ) to (E,ℰ) is

I⁢(x,A)=δx⁢(A). (4)

It is apparent that the unit kernel is a Markov kernel.

If K is a Markov kernel from (E,ℰ) to (F,ℱ) and L is a Markov kernel from (F,ℱ) to (G,𝒢), then for x∈E, by (3) we have

(K⁢L)*⁢(1G)⁢(x)=∫F𝑑Kx⁢(y)=Kx⁢(F)=K⁢(x,F)=1,

and thus by (2),

(K⁢L)x⁢(G)=(K⁢L)⁢(x,G)=1,

showing that for each x∈E, (K⁢L)x is a probability measure. Therefore, the product of two Markov kernels is a Markov kernel.

Let (E,ℰ) be a measurable space and let

Bb⁢(ℰ)

be the set of bounded functions E→ℝ that are measurable ℰ→ℬℝ. Bb⁢(ℰ) is a Banach space with the uniform norm

∥f∥u=supx∈E⁡|f⁢(x)|.

For K a Markov kernel from (E,ℰ) to (F,ℱ) and for f∈Bb⁢(ℱ), define K*⁢f:E→ℝ by

(K*⁢f)⁢(x)=∫Ff⁢(y)⁢𝑑Kx⁢(y),x∈E,

for which

|(K*⁢f)⁢(x)|≤∫F|f⁢(y)|⁢𝑑Kx⁢(y)≤∥f∥u⁢Kx⁢(F)=∥f∥u,

showing that ∥K*⁢f∥u≤∥f∥u. Furthermore, there is a sequence of simple functions ϕn∈Bb⁢(ℱ) that converges to f in the norm ∥⋅∥u.55 5 V. I. Bogachev, Measure Theory, p. 108, Lemma 2.1.8. For x∈E, by the dominated convergence theorem we get that

(K*⁢ϕn)⁢(x)=∫Fϕn⁢(y)⁢𝑑Kx⁢(y)→∫Ff⁢(y)⁢𝑑Kx⁢(y)=(K*⁢f)⁢(x).

Each K*⁢ϕn is measurable ℰ→ℬℝ, hence K*⁢f is measurable ℰ→ℬℝ and so belongs to Bb⁢(ℰ).

3 Markov semigroups

Let (E,ℰ) be a measurable space and for each t≥0, let Pt be a Markov kernel from (E,ℰ) to (E,ℰ). We say that the family (Pt)t∈ℝ≥0 is a Markov semigroup if

Ps+t=Ps⁢Pt,s,t∈ℝ≥0.

For x∈E and A∈ℰ and for s,t≥0, by (2) and (3),

(Ps⁢Pt)⁢(x,A)=((Ps⁢Pt)*⁢1A)⁢(x)=∫EPt⁢(y,A)⁢d⁢(Ps)x⁢(y)

Thus

Ps+t⁢(x,A)=∫EPt⁢(y,A)⁢d⁢(Ps)x⁢(y), (5)

called the Chapman-Kolmogorov equation.

4 Infinitely divisible distributions

Let 𝒫⁢(ℝd) be the collection of Borel probability measures on ℝd. For μ∈𝒫⁢(ℝd), its characteristic function μ~:ℝd→ℂ is defined by

μ~⁢(x)=∫ℝdei⁢⟨x,y⟩⁢𝑑μ⁢(y).

μ~ is uniformly continuous on ℝd and |μ~⁢(x)|≤μ~⁢(0)=1 for all x∈ℝd.66 6 Heinz Bauer, Probability Theory, p. 183, Theorem 22.3. For μ1,…,μn∈𝒫⁢(ℝd), let μ be their convolution:

μ=μ1*⋯*μn,

which for A a Borel set in ℝd is defined by

μ⁢(A)=∫(ℝd)n1A⁢(x1+⋯+xn)⁢d⁢(μ1×⋯×μn)⁢(x1,…,xn).

One computes that77 7 Heinz Bauer, Probability Theory, p. 184, Theorem 22.4.

μ~=μ~1⁢⋯⁢μ~n.

An element μ of 𝒫⁢(ℝd) is called infinitely divisible if for each n≥1, there is some μn∈𝒫⁢(ℝd) such that

μ=μn*⋯*μn⏟n. (6)

Thus,

μ~=(μ~n)n. (7)

On the other hand, if μn∈𝒫⁢(ℝd) is such that (7) is true, then because the characteristic function of μn*⋯*μn is (μ~n)n and the characteristic function of μ is μ~ and these are equal, it follows that μn*⋯*μn and μ are equal.

The following theorem is useful for doing calculations with the characteristic function of an infinitely divisible distribution.88 8 Heinz Bauer, Probability Theory, p. 246, Theorem 29.2.

Theorem 2.

Suppose that μ is an infinitely divisible distribution on ℝd. First,

μ~⁢(x)≠0,x∈ℝd.

Second, there is a unqiue continuous function ϕ:ℝd→ℝ satisfying ϕ⁢(0)=0 and

μ~=|μ~|⁢ei⁢ϕ.

Third, for each n≥1, there is a unique μn∈𝒫⁢(ℝd) for which μ=μn*⋯*μn. The characteristic function of this unique μn is

μ~n=|μ~|1n⁢ei⁢ϕn.

A convolution semigroup is a family (μt)t∈ℝ≥0 of elements of 𝒫⁢(ℝd) such that for s,t∈ℝ≥0,

μs+t=μs*μt.

The convolution semigroup is called continuous when t↦μt is continuous ℝ≥0→𝒫⁢(ℝd), where 𝒫⁢(ℝd) has the narrow topology.

The following theorem connects convolution semigroups and infinitely divisible distributions.99 9 Heinz Bauer, Probability Theory, p. 248, Theorem 29.6.

Theorem 3.

If (μt)t∈ℝ≥0 is a convolution semigroup on ℬℝd, then for each t, the measure μt is infinitely divisible.

If μ∈𝒫⁢(ℝd) is infinitely divisible and t0>0, then there is a unique continuous convolution semigroup (μt)t∈ℝ≥0 such that μt0=μ.

It follows from the above theorem that for a convolution semigroup (μt)t∈ℝ≥0 on ℬℝd, μ1 is infinitely divisible and therefore by Theorem 2, μ~1⁢(x)≠0 for all x. But μ0*μ1=μ1, so μ~0⁢μ~1=μ~1, and μ~0⁢(x)=1 for each x. But δ~0⁢(x)=1 for all x, so

μ0=δ0. (8)

5 Translation-invariant semigroups

Let (Pt)t∈ℝ≥0 be a Markov semigroup on (ℝd,ℬℝd). We say that (Pt)t∈ℝ is translation-invariant if for all x,y∈ℝd, A∈ℬℝd, and t∈ℝ≥0,

Pt⁢(x,A)=Pt⁢(x+y,A+y).

In this case, for t≥0 and for A∈ℬℝd, define

μt⁢(A)=Pt⁢(0,A).

Each μt is a probability measure on ℬℝd, and

μt⁢(A-x)=Pt⁢(0,A-x)=Pt⁢(x,(A-x)+x)=Pt⁢(x,A).

Using that the Chapman-Kolmogorov equation (5) and as (Ps)0⁢(B)=Ps⁢(0,B)=μs⁢(B),

μs+t⁢(A) =Ps+t⁢(0,A)
=∫ℝdPt⁢(y,A)⁢d⁢(Ps)0⁢(y)
=∫ℝdμt⁢(A-y)⁢𝑑μs⁢(y)
=(μt*μs)⁢(A),

showing that (μt)t∈ℝ≥0 is a convolution semigroup on ℬℝd.

On the other hand, if (μt)t∈ℝ≥0 is a convolution semigroup of probability measures on ℬℝd, for t≥0, x∈ℝd, and A∈ℬℝd define

Pt⁢(x,A)=μt⁢(A-x).

Let t≥0. For x∈ℝd, the map A↦Pt⁢(x,A)=μt⁢(A-x) is a probability measure on ℬℝd. The map (x,y)↦x+y is continuous ℝd×ℝd→ℝd, and for A∈ℬℝd, the map 1A:ℝd→ℝ is measurable ℬℝd→ℬℝ. Hence, as ℬℝd×ℝd=ℬℝd⊗ℬℝd, the map (x,y)↦1A⁢(x+y) is measurable ℬℝd⊗ℬℝd→ℬℝ. Thus by Fubini’s theorem,

x↦∫ℝd1A⁢(x+y)⁢𝑑μt⁢(y)=∫ℝd1A-x⁢(y)⁢𝑑μt⁢(y)=μt⁢(A-x)

is measurable ℬℝd→ℬℝ. Hence Pt is a Markov kernel, and thus (Pt)t∈ℝ≥0 is a translation-invariant Markov semigroup.

Define S:ℝd→ℝd by S⁢(x)=-x. For μ,ν∈𝒫⁢(ℝd),

S*⁢(μ*ν)⁢(A) =(μ*ν)⁢(-A)
=∫ℝdμ⁢(-A-y)⁢𝑑ν⁢(y)
=∫ℝdμ⁢(-A+y)⁢𝑑ν¯⁢(y)
=∫ℝdμ¯⁢(A-y)⁢𝑑ν¯⁢(y)
=(μ¯*ν¯)⁢(A),

thus

S*⁢(μ*ν)=(S*⁢μ)*(S*⁢ν). (9)

For μ∈𝒫⁢(ℝd), write

μ¯=S*⁢μ∈𝒫⁢(ℝd),

i.e.,

μ¯⁢(A)=μ⁢(S-1⁢(A))=μ⁢(S⁢(A))=μ⁢(-A).

We calculate

(Pt*⁢1A)⁢(x)=Pt⁢(x,A)=μt⁢(A-x)=∫ℝd1A⁢(x+y)⁢𝑑μt⁢(y).

Then if f is a simple function, f=∑kak⁢1Ak,

(Pt*⁢f)⁢(x)=∑kak⁢∫ℝd1Ak⁢(x+y)⁢𝑑μt⁢(y)=∫ℝdf⁢(x+y)⁢𝑑μt⁢(y).

For f∈Bb⁢(ℬℝd), there is a sequence of simple functions fn that converge to f in the uniform norm, and then by the dominated convergence theorem we get

(Pt*⁢f)⁢(x)=∫ℝdf⁢(x+y)⁢𝑑μt⁢(y).

But

∫ℝdf⁢(x+y)⁢𝑑μt⁢(y) =∫ℝdf⁢(x+S⁢(S⁢(y)))⁢𝑑μt⁢(y)
=∫ℝdf⁢(x+S⁢(y))⁢d⁢(S*⁢μt)⁢(y)
=∫ℝdf⁢(x-y)⁢𝑑μ¯t⁢(y)
=(f*μ¯t)⁢(x).

Therefore for t≥0 and f∈Bb⁢(ℬℝd),

Pt*⁢f=f*μ¯t. (10)

For s,t≥0 and f∈Bb⁢(ℬℝd), by (10), the fact that (μt)t∈ℝ≥0 is a convolution semigroup, and (9), we get

Ps+t*⁢f =f*(S*⁢μs+t)
=f*(S*⁢(μs*μt))
=f*((S*⁢μs)*(S*⁢μt))
=(f*(S*⁢μs))*(S*⁢μt)
=(Ps*⁢f)*(S*⁢μt)
=Pt*⁢(Ps*⁢f).

This shows that (Pt)t∈ℝ≥0 is a Markov semigroup. Moreover, by (8) it holds that μ0=δ0, and hence

P0⁢(x,A)=μ0⁢(A-x)=δ0⁢(A-x)=δx⁢(A).

Namely, P0 is the unit kernel (4).

If (μt)t∈ℝ≥0 is a convolution semigroup and some μt has density qt with respect to Lebesgue measure λd on ℝd,

μt=qt⁢λd,

then writing q¯t⁢(x)=qt⁢(-x), for f∈Bb⁢(ℬℝd) by (10) we have

(Pt*⁢f)⁢(x)=(f*μ¯t)⁢(x)=∫ℝdf⁢(x-y)⁢𝑑μ¯t⁢(y)=∫ℝdf⁢(x+y)⁢qt⁢(y)⁢𝑑λd⁢(y)

so

Pt*f=f*q¯t. (11)

6 The Brownian semigroup

For a∈ℝ and σ>0, let γa,σ2 be the Gaussian measure on ℝ, the probability measure on ℝ whose density with respect to Lebesgue measure is

p⁢(x,a,σ2)=12⁢π⁢σ2⁢exp⁡(-(x-a)22⁢σ2).

For σ=0, let

γa,0=δa.

Define for t∈ℝ≥0,

μt=∏k=1dγ0,t,

which is an element of 𝒫⁢(ℝd). For s,t∈ℝ≥0, we calculate

μs*μt=(∏k=1dγ0,s)*(∏k=1dγ0,t)=∏k=1d(γ0,s*γ0,t)=∏k=1dγ0,s+t=μs+t.

Lévy’s continuity theorem states that if νn is a sequence in 𝒫⁢(ℝd) and there is some ϕ:ℝd→ℂ that is continuous at 0 and to which ν~n converges pointwise, then there is some ν∈𝒫⁢(ℝd) such that ϕ=ν~ and such that νn→ν narrowly. But for t∈ℝ≥0 and x∈ℝd, we calculate

μ~t⁢(x)=∫ℝdei⁢⟨x,y⟩⁢𝑑μt⁢(y)=exp⁡(-t⁢|x|22). (12)

Let ϕ⁢(x)=1 for all x, for which δ~0=ϕ. For tn∈ℝ≥0 tending to 0, let νn=μtn. Then by (12), ν~n converges pointwise to ϕ, so by Lévy’s continuity theorem, νn converges narrowly to δ0. Moreover, because ℝd is a Polish space, 𝒫⁢(ℝd) is a Polish space, and in particular is metrizable. It thus follows that μt converges narrowly to δ0 as t→0. It then follows that t↦μt is continuous ℝ≥0→𝒫⁢(ℝd). Summarizing, (μt)t∈ℝ≥0 is a continuous convolution semigroup.

For t>0, μt has density

gt⁢(x)=∏j=1d(2⁢π⁢t)-1/2⁢e-xj22⁢t=(2⁢π⁢t)-d/2⁢e-|x|22⁢t

with respect to Lebesgue measure λd on ℝd. For t≥0, let

Pt⁢(x,A)=μt⁢(A-x).

We have established that (Pt)t∈ℝ≥0 is a translation-invariant Markov semigroup for which P0⁢(x,A)=δx⁢(A). We call (Pt)t∈ℝ≥0 the Brownian semigroup. For t>0 and f∈Bb⁢(ℬℝd), because g¯t=gt we have by (11),

(Pt⁢f)⁢(x)=(f*gt)⁢(x)=(2⁢π⁢t)-d/2⁢∫ℝdf⁢(x-y)⁢e-|y|22⁢t⁢𝑑λd⁢(y).

7 Projective families

For a nonempty set I, let 𝒦⁢(I) denote the family of finite nonempty subsets of I. We speak in this section about projective families of probability measures.

The following theorem shows how to construct a projective family from a Markov semigroup on a measurable space and a probability measure on this measurable space.1010 10 Heinz Bauer, Probability Theory, p. 314, Theorem 36.4.

Theorem 4.

Let I=ℝ≥0, let (E,ℰ) be a measurable space, let (Pt)t∈I be a Markov semigroup on ℰ, and let μ be a probability measure on ℰ. For J∈𝒦⁢(I), with elements t1<⋯<tn, and for A∈ℰJ, let

PJ⁢(A)=∫E∫E⋯⁢∫E⏟n+1⁢1A⁢(x1,…,xn)⁢d⁢(Ptn-tn-1)xn-1⁢(xn)⁢⋯⁢d⁢(Pt1)x0⁢(x1)⁢d⁢μ⁢(x0).

Then (PJ)J∈𝒦⁢(I) is a projective family of probability measures.

Proof.

Let Ak be pairwise disjoint elements of ℰJ, and call their union A. Then 1A=∑k1Ak, and applying the monotone convergence theorem n+1 times,

∫E∫E⋯⁢∫E⏟n+1⁢1A⁢(x1,…,xn)⁢d⁢(Ptn-tn-1)xn-1⁢(xn)⁢⋯⁢d⁢(Pt1)x0⁢(x1)⁢d⁢μ⁢(x0)=∑k∫E∫E⋯⁢∫E⏟n+1⁢1Ak⁢(x1,…,xn)⁢d⁢(Ptn-tn-1)xn-1⁢(xn)⁢⋯⁢d⁢(Pt1)x0⁢(x1)⁢d⁢μ⁢(x0),

i.e.

PJ⁢(A)=∑kPJ⁢(Ak).

Furthermore, because (Pt)x is a probability measure for each t and for each x and μ is a probability measure, we calculate that

PJ⁢(EJ)=1.

Thus, PJ is a probability measure on ℰJ.

To prove that (PJ)J∈𝒦⁢(I) is a projective family, it suffices to prove that when J,K∈𝒦⁢(I), J⊂K, and K∖J is a singleton, then (πK,J)*⁢PK=PJ. Moreover, because (i) the product σ-algebra ℰJ is generated by the collection of cylinder sets, i.e. sets of the form ∏t∈JAt for At∈ℰ, and (ii) the intersection of finitely many cylinder sets is a cylinder sets, it is proved using the monotone class theorem that if two probability measures on ℰJ coincide on the cylinder sets, then they are equal.1111 11 V. I. Bogachev, Measure Theory, volume I, p. 35, Lemma 1.9.4. Let t1<⋯<tn be the elements of J. To prove that (πK,J)*⁢PK and PJ are equal, it suffices to prove that for any A1,…,An∈ℰ,

(πK,J)*⁢PK⁢(∏j=1nAj)=PJ⁢(∏j=1nAj).

Moreover, for A=∏j=1nAj,

1A=1A1⊗⋯⊗1An,

thus

PJ⁢(∏j=1nAj)=∫E∫E⋯⁢∫E⏟n+1⁢1A1⁢(x1)⁢⋯⁢1An⁢(xn)⁢d⁢(Ptn-tn-1)xn-1⁢(xn)⁢⋯⁢d⁢(Pt1)x0⁢(x1)⁢d⁢μ⁢(x0)=∫E∫A1⋯⁢∫And⁢(Ptn-tn-1)xn-1⁢(xn)⁢⋯⁢d⁢(Pt1)x0⁢(x1)⁢𝑑μ⁢(x0).

Let K∖J={t′}. Either t′<t1, or t′>tn, or there is some 1≤j≤n-1 for which tj<t′<tj+1. Take the case t′<t1. Then

πK,J-1⁢(∏j=1nAj)=∏k=0nBk,

where B0=E and Bj=Aj for 1≤j≤n. Then

(πK,J)*⁢PK⁢(∏j=1nAj)=PK⁢(∏k=0nBk)=∫E∫E∫A1⋯⁢∫And⁢(Ptn-tn-1)xn-1⁢(xn)⁢⋯⁢d⁢(Pt1-t′)x′⁢(x1)⁢d⁢(Pt′)x0⁢(x′)⁢𝑑μ⁢(x0)=∫E∫E∫A1f⁢(x1)⁢d⁢(Pt1-t′)x′⁢(x1)⁢d⁢(Pt′)x0⁢(x′)⁢𝑑μ⁢(x0),

for

f⁢(x1)=∫A2⋯⁢∫And⁢(Ptn-tn-1)xn-1⁢(xn)⁢⋯⁢d⁢(Pt2-t1)x1⁢(x2).

By (1) and because (Pt)t∈I is a Markov semigroup,

∫E∫A1f⁢(x1)⁢d⁢(Pt1-t′)x′⁢(x1)⁢d⁢(Pt′)x0⁢(x′)=∫E∫Ef⁢(x1)⁢1A1⁢(x1)⁢d⁢(Pt1-t′)x′⁢(x1)⁢d⁢(Pt′)x0⁢(x′)=∫EPt1-t′*⁢(f⁢1A1)⁢(x′)⁢d⁢(Pt′)x0⁢(x′)=Pt′*⁢(Pt1-t′*⁢(f⁢1A1))⁢(x0)=Pt1⁢(f⁢1A1)⁢(x0)=∫Ef⁢(x1)⁢1A1⁢(x1)⁢d⁢(Pt1)x0⁢(x1)=∫A1f⁢(x1)⁢d⁢(Pt1)x0⁢(x1)=∫A1∫A2⋯⁢∫And⁢(Ptn-tn-1)xn-1⁢(xn)⁢⋯⁢d⁢(Pt2-t1)x1⁢(x2)⁢d⁢(Pt1)x0⁢(x1).

Thus

(πK,J)*⁢PK⁢(∏j=1nAj)=∫E∫A1∫A2⋯⁢∫And⁢(Ptn-tn-1)xn-1⁢(xn)⁢⋯⁢d⁢(Pt2-t1)x1⁢(x2)⁢d⁢(Pt1)x0⁢(x1)⁢𝑑μ⁢(x0)=PJ⁢(∏j=1nAj).

This shows that the claim is true in the case t′<t1. ∎

Thus, if E is a Polish space with Borel σ-algebra ℰ, let I=ℝ≥0, let (Pt)t∈I be a Markov semigroup on ℰ, and let μ be a probability measure on ℰ. The above theorem tells us that (PJ)𝒦⁢(I) is a projective family, and then the Kolmogorov extension theorem tells us that there is a probability measure1212 12 We write Pμ to indicate that this measure involves μ; it also involves the Markov semigroup, which we do not indicate. Pμ on ℰI such that for any J∈𝒦⁢(I), πJ*⁢Pμ=PJμ. This implies that there is a stochastic process (Xt)t∈I whose finite-dimensional distributions are equal to the probability measures PJ defined in Theorem 4 using the Markov semigroup (Pt)t∈I and the probability measure μ.