The heat kernel on ℝn

Jordan Bell
March 28, 2014

1 Notation

For f∈L1⁢(ℝn), we define f^:ℝn→ℂ by

f^⁢(ξ)=(ℱ⁢f)⁢(ξ)=∫ℝnf⁢(x)⁢e-2⁢π⁢i⁢ξ⁢x⁢𝑑x,ξ∈ℝn.

The statement of the Riemann-Lebesgue lemma is that f^∈C0⁢(ℝn).

We denote by 𝒮n the Fréchet space of Schwartz functions ℝn→ℂ.

If α is a multi-index, we define

Dα=D1α1⁢⋯⁢Dnαn,
Dα=i-|α|⁢Dα=(1i⁢D1)α1⁢⋯⁢(1i⁢Dn)αn,

and

Δ=D12+⋯+Dn2.

2 The heat equation

Fix n, and for t>0, x∈ℝn, define

kt⁢(x)=k⁢(t,x)=(4⁢π⁢t)-n/2⁢exp⁡(-|x|24⁢t).

We call k the heat kernel. It is straightforward to check for any t>0 that kt∈𝒮n. The heat kernel satisfies

kt⁢(x)=(t-1/2)n⁢k1⁢(t-1/2⁢x),t>0,x∈ℝn.

For a>0 and f⁢(x)=e-π⁢a⁢|x|2, it is a fact that f^⁢(ξ)=a-n/2⁢e-π⁢|ξ|2/a. Using this, for any t>0 we get

k^t⁢(ξ)=e-4⁢π2⁢|ξ|2⁢t,ξ∈ℝn.

Thus for any t>0,

∫ℝnkt⁢(x)⁢𝑑x=k^t⁢(0)=1.

Then the heat kernel is an approximate identity: if f∈Lp⁢(ℝn), 1≤p<∞, then ∥f*kt-f∥p→0 as t→0, and if f is a function on ℝn that is bounded and continuous, then for every x∈ℝn, f*kt⁢(x)→f⁢(x) as t→0.11 1 k1, and any kt, belong merely to 𝒮n and not to 𝒟⁢(ℝn), which is demanded in the definition of an approximate identity in Rudin’s Functional Analysis, second ed. For each t>0, because kt∈𝒮n we have f*kt∈C∞⁢(ℝn), and Dα⁢(f*kt)=f*Dα⁢kt for any multi-index α.22 2 Gerald B. Folland, Introduction to Partial Differential Equations, second ed., p. 11, Theorem 0.14.

The heat operator is Dt-Δ and the heat equation is (Dt-Δ)⁢u=0. It is straightforward to check that

(Dt-Δ)⁢k⁢(t,x)=0,t>0,x∈ℝn,

that is, the heat kernel is a solution of the heat equation.

To get some practice proving things about solutions of the heat equation, we work out the following theorem from Folland.33 3 Gerald B. Folland, Introduction to Partial Differential Equations, second ed., p. 144, Theorem 4.4. In Folland’s proof it is not apparent how the hypotheses on u and Dx are used, and we make this explicit.

Theorem 1.

Suppose that u:[0,∞)×Rn→C is continuous, that u is C2 on (0,∞)×Rn, that

(Dt-Δ)⁢u⁢(t,x)=0,t>0,x∈ℝn,

and that u⁢(0,x)=0 for x∈Rn. If for every ϵ>0 there is some C such that

|u⁢(t,x)|≤C⁢eϵ⁢|x|2,|Dx⁢u⁢(t,x)|≤C⁢eϵ⁢|x|2,t>0,x∈ℝn,

then u=0.

Proof.

If f and g are C2 functions on some open set in ℝ×ℝn, such as (0,∞)×ℝn, then

g⁢(∂t⁡f-Δ⁢f)+f⁢(∂t⁡g+Δ⁢g) = ∂t⁡(f⁢g)-g⁢∑j=1n∂j2⁡f+f⁢∑j=1n∂j2⁡g
= ∂t⁡(f⁢g)+∑j=1n∂j⁡(f⁢∂j⁡g-g⁢∂j⁡f)
= divt,x⁢F,

where

F=(f⁢g,f⁢∂1⁡g-g⁢∂1⁡f,…,f⁢∂n⁡g-g⁢∂n⁡f).

Take t0>0, x0∈ℝn, and let f⁢(t,x)=u⁢(t,x) and g⁢(t,x)=k⁢(t0-t,x-x0) for t>0, x∈ℝn. Let 0<a<b<t0 and r>0, and define

Ω={(t,x):|x|<r,a<t<b}.

In Ω we check that (∂t-Δ)⁢f=0 and (∂t+Δ)⁢g=0, so by the divergence theorem,

∫∂⁡ΩF⋅ν=∫Ωdivt,x⁢F=∫Ωg⁢(∂t⁡f-Δ⁢f)+f⁢(∂t⁡g+Δ⁢g)=∫Ωg⋅0+f⋅0=0.

On the other hand, as

∂⁡Ω={(b,x):|x|≤r}∪{(a,x):|x|≤r}∪{(t,x):a<t<b,|x|=r},

we have

∫∂⁡ΩF⋅ν = ∫|x|≤rF⁢(b,x)⋅(1,0,…,0)⁢𝑑x+∫|x|≤rF⁢(a,x)⋅(-1,0,…,0)⁢𝑑x
+∫ab∫|x|=rF⁢(t,x)⋅xr⁢𝑑σ⁢(x)⁢tn-1⁢𝑑t
= ∫|x|≤rf⁢(b,x)⁢g⁢(b,x)⁢𝑑x-∫|x|≤rf⁢(a,x)⁢g⁢(a,x)⁢𝑑x
+∫ab∫|x|=r∑j=1n(f⁢∂j⁡g-g⁢∂j⁡f)⁢(t,x)⁢xjr⁢d⁢σ⁢(x)⁢tn-1⁢d⁢t
= ∫|x|≤ru⁢(b,x)⁢k⁢(t0-b,x-x0)⁢𝑑x-∫|x|≤ru⁢(a,x)⁢k⁢(t0-a,x-x0)⁢𝑑x
+∫ab∫|x|=r∑j=1n(u(t,x)∂jk(t0-t,x-x0)
-k(t0-t,x-x0)∂ju(t,x))xjrdσ(x)tn-1dt,

where σ is surface measure on {|x|=r}=rSn-1. As r→∞, the first two terms tend to

∫ℝnu⁢(b,x)⁢kt0-b⁢(x-x0)⁢𝑑x=∫ℝnu⁢(b,x)⁢kt0-b⁢(x0-x)⁢𝑑x=u⁢(b,⋅)*kt0-b⁢(x0)

and

∫ℝnu⁢(a,x)⁢kt0-a⁢(x-x0)⁢𝑑x=∫ℝnu⁢(a,x)⁢kt0-a⁢(x0-x)⁢𝑑x=u⁢(a,⋅)*kt0-a⁢(x0)

respectively. Let ϵ<14⁢(t0-a), and let C be as given in the statement of the theorem. Using ∂j⁡k⁢(t,x)=-xj2⁢t⁢k⁢(t,x), for any r>0 the third term is bounded by

n⁢∫ab∫|x|=r(C⁢eϵ⁢r2⁢|x-x0|2⁢t⁢k⁢(t0-t,x-x0)+k⁢(t0-t,x-x0)⁢C⁢eϵ⁢r2)⁢𝑑σ⁢(x)⁢tn-1⁢𝑑t,

which is bounded by

n⁢∫ab∫|x|=rC⁢eϵ⁢r2⁢(|x0|+r2⁢a+1)⁢(4⁢π⁢(t0-b))-n/2⁢exp⁡(-r24⁢(t0-a))⁢𝑑σ⁢(x)⁢tn-1⁢𝑑t,

and writing η=14⁢(t0-a)-ϵ and ωn=2⁢πn/2Γ⁢(n/2), the surface area of the sphere of radius 1 in ℝn, this is equal to

(b-a)n⁢rn-1⁢ωn⁢C⁢e-η⁢r2⁢(|x0|+r2⁢a+1)⁢(4⁢π⁢(t0-b))-n/2,

which tends to 0 as r→∞. Therefore,

u⁢(b,⋅)*kt0-b⁢(x0)=u⁢(a,⋅)*kt0-a⁢(x0).

One checks that as b→t0, the left-hand side tends to u⁢(t0,x0), and that as a→0, the right-hand side tends to u⁢(0,x0)=0. Therefore,

u⁢(t0,x0)=0.

This is true for any t0>0, x0∈ℝn, and as u:[0,∞)×ℝn→ℂ is continuous, it follows that u is identically 0. ∎

3 Fundamental solutions

We extend k to ℝ×ℝn as

k⁢(t,x)={(4⁢π⁢t)-n/2⁢exp⁡(-|x|24⁢t)t>0,x∈ℝn0t≤0,x∈ℝn.

This function is locally integrable in ℝ×ℝn, so it makes sense to define Λk∈𝒟′⁢(ℝ×ℝn) by

Λk⁢ϕ=∫ℝ∫ℝnϕ⁢(t,x)⁢k⁢(t,x)⁢𝑑x⁢𝑑t,ϕ∈𝒟⁢(ℝ×ℝn).

Suppose that P is a polynomial in n variables:

P⁢(ξ)=∑cα⁢ξα=∑cα⁢ξ1α1⁢⋯⁢ξnαn.

We say that E∈𝒟′⁢(ℝn) is a fundamental solution of the differential operator

P⁢(D)=∑cα⁢Dα=∑cα⁢i-|α|⁢Dα

if P⁢(D)⁢E=δ. If E=Λf for some locally integrable f, Λf⁢ϕ=∫ℝnϕ⁢(x)⁢f⁢(x)⁢𝑑x, we also say that the function f is a fundamental solution of the differential operator P⁢(D). We now prove that the heat kernel extended to ℝ×ℝn in the above way is a fundamental solution of the heat operator.44 4 Gerald B. Folland, Introduction to Partial Differential Equations, second ed., p. 146, Theorem 4.6.

Theorem 2.

Λk is a fundamental solution of Dt-Δ.

Proof.

For ϵ>0, define Kϵ⁢(t,x)=k⁢(t,x) if t>ϵ and Kϵ⁢(t,x)=0 otherwise. For any ϕ∈𝒟⁢(ℝ×ℝn),

|∫ℝ∫ℝn(k⁢(t,x)-Kϵ⁢(t,x))⁢ϕ⁢(t,x)⁢𝑑x⁢𝑑t| = |∫0ϵ∫ℝnk⁢(t,x)⁢ϕ⁢(t,x)⁢𝑑x⁢𝑑t|
≤ ∥ϕ∥∞⁢∫0ϵ∫ℝnk⁢(t,x)⁢𝑑x⁢𝑑t
= ∥ϕ∥∞⁢∫0ϵ𝑑t
= ∥ϕ∥∞⁢ϵ.

This shows that ΛKϵ→Λk in 𝒟′⁢(ℝ×ℝn), with the weak-* topology. It is a fact that for any multi-index, E↦Dα⁢E is continuous 𝒟′⁢(ℝ×ℝn)→𝒟′⁢(ℝ×ℝn), and hence (Dt-Δ)⁢ΛKϵ→(Dt-Δ)⁢Λk in 𝒟′⁢(ℝ×ℝn). Therefore, to prove the theorem it suffices to prove that (Dt-Δ)⁢ΛKϵ→δ (because 𝒟′⁢(ℝ×ℝn) with the weak-* topology is Hausdorff).

Let ϕ∈𝒟⁢(ℝ×ℝn). Doing integration by parts,

(Dt-Δ)⁢ΛKϵ⁢(ϕ) = ΛKϵ⁢((Dt-Δ)⁢ϕ)
= ∫ℝ∫ℝnKϵ⁢(t,x)⁢(Dt⁢ϕ⁢(t,x)-Δ⁢ϕ⁢(t,x))⁢𝑑x⁢t⁢x
= ∫ϵ∞∫ℝnk⁢(t,x)⁢Dt⁢ϕ⁢(t,x)-k⁢(t,x)⁢Δ⁢ϕ⁢(t,x)⁢d⁢x⁢t⁢x
= ∫ℝn(k⁢(ϵ,x)⁢ϕ⁢(ϵ,x)-∫ϵ∞ϕ⁢(t,x)⁢Dt⁢k⁢(t,x)⁢𝑑t)⁢𝑑x
+∫ϵ∞∫ℝnϕ⁢(t,x)⁢Δ⁢k⁢(t,x)⁢𝑑x⁢𝑑t
= ∫ℝnk⁢(ϵ,x)⁢ϕ⁢(ϵ,x)⁢𝑑x
-∫ϵ∞∫ℝnϕ⁢(t,x)⁢(Dt-Δ)⁢k⁢(t,x)⁢𝑑t⁢𝑑x
= ∫ℝnk⁢(ϵ,x)⁢ϕ⁢(ϵ,x)⁢𝑑x.

So, using kt⁢(x)=kt⁢(-x) and writing ϕt⁢(x)=ϕ⁢(t,x),

(Dt-Δ)⁢ΛKϵ⁢(ϕ) = ∫ℝnkϵ⁢(-x)⁢ϕϵ⁢(x)⁢𝑑x
= kϵ*ϕϵ⁢(0)
= kϵ*ϕ0⁢(0)+kϵ*(ϕϵ-ϕ0)⁢(0).

Using the definition of convolution, the second term is bounded by

supx∈ℝn⁡|ϕϵ⁢(x)-ϕ0⁢(x)|⁢∥kϵ∥1=supx∈ℝn⁡|ϕϵ⁢(x)-ϕ0⁢(x)|,

which tends to 0 as ϵ→0. Because k is an approximate identity, kϵ*ϕ0⁢(0)→ϕ0⁢(0) as ϵ→0. That is,

(Dt-Δ)⁢ΛKϵ⁢(ϕ)→ϕ0⁢(0)=δ⁢(ϕ)

as ϵ→0, showing that (Dt-Δ)⁢ΛKϵ→δ in 𝒟′⁢(ℝ×ℝn) and completing the proof. ∎

4 Functions of the Laplacian

This section is my working through of material in Folland.55 5 Gerald B. Folland, Introduction to Partial Differential Equations, second ed., pp. 149–152, §4B. For f∈𝒮n and for any nonnegative integer k, doing integration by parts we get

ℱ⁢((-Δ)k⁢f)⁢(ξ)=∫ℝn((-Δ)k⁢f)⁢(x)⁢e-2⁢π⁢i⁢ξ⁢x⁢𝑑x=(4⁢π2⁢|ξ|2)k⁢(ℱ⁢f)⁢(ξ),ξ∈ℝn.

Suppose that P is a polynomial in one variable: P⁢(x)=∑ck⁢xk. Then, writing P⁢(-Δ)=∑ck⁢(-Δ)k, we have

ℱ⁢(P⁢(-Δ)⁢f)⁢(ξ) = ∑ck⁢ℱ⁢((-Δ)k⁢f)⁢(ξ)
= ∑ck⁢(4⁢π2⁢|ξ|2)k⁢(ℱ⁢f)⁢(ξ)
= (ℱ⁢f)⁢(ξ)⁢P⁢(4⁢π2⁢|ξ|2).

We remind ourselves that tempered distributions are elements of 𝒮n′, i.e. continuous linear maps 𝒮n→ℂ. The Fourier transform of a tempered distribution Λ is defined by Λ^⁢f=(ℱ⁢Λ)⁢f=Λ⁢f^, f∈𝒮n. It is a fact that the Fourier transform is an isomorphism of locally convex spaces 𝒮n′→𝒮n′.66 6 Walter Rudin, Functional Analysis, second ed., p. 192, Theorem 7.15.

Suppose that ψ:(0,∞)→ℂ is a function such that

Λ⁢f=∫ℝnf⁢(ξ)⁢ψ⁢(4⁢π2⁢|ξ|2)⁢𝑑ξ,f∈𝒮n,

is a tempered distribution. We define ψ⁢(-Δ):𝒮n→𝒮n′ by

ψ⁢(-Δ)⁢f=ℱ-1⁢(f^⁢Λ),f∈𝒮n.

Define fˇ⁢(x)=f⁢(-x); this is not the inverse Fourier transform of f, which we denote by ℱ-1. As well, write τx⁢f⁢(y)=f⁢(y-x). For u∈𝒮n′ and ϕ∈𝒮n, we define the convolution u*ϕ:ℝn→ℂ by

(u*ϕ)⁢(x)=u⁢(τx⁢ϕˇ),x∈ℝn.

One proves that u*ϕ∈C∞⁢(ℝn), that

Dα⁢(u*ϕ)=(Dα⁢u)*ϕ=u*(Dα⁢ϕ)

for any multi-index, that u*ϕ is a tempered distribution, that ℱ⁢(u*ϕ)=ϕ^⁢u^, and that u^*ϕ^=ℱ⁢(ϕ⁢u).77 7 Walter Rudin, Functional Analysis, second ed., p. 195, Theorem 7.19.

We can also write ψ⁢(-Δ) in the following way. There is a unique κψ∈𝒮n′ such that

ℱ⁢κψ=Λ.

For f∈𝒮n, we have ℱ⁢(κψ*f)=f^⁢κ^ψ=f^⁢Λ, but, using the definition of ψ⁢(-Δ) we also have ℱ⁢(ψ⁢(-Δ)⁢f)=ℱ⁢ℱ-1⁢(f^⁢Λ)=f^⁢Λ, so

κψ*f=ψ⁢(-Δ)⁢f.

Moreover, κψ*f∈C∞⁢(ℝn); this shows that ψ⁢(-Δ)⁢f can be interpreted as a tempered distribution or as a function. We call κψ the convolution kernel of ψ⁢(-Δ).

For a fixed t>0, define ψ⁢(s)=e-t⁢s. Then Λ:𝒮n→ℂ defined by

Λ⁢f=∫ℝnf⁢(ξ)⁢ψ⁢(4⁢π2⁢|ξ|2)⁢𝑑ξ=∫ℝnf⁢(ξ)⁢exp⁡(-4⁢π2⁢|ξ|2⁢t)⁢𝑑ξ=∫ℝnf⁢(ξ)⁢k^t⁢(ξ)⁢𝑑ξ

is a tempered distribution. Using the Plancherel theorem, we have

Λ⁢f=∫ℝnf^⁢(ξ)⁢kt⁢(ξ)⁢𝑑ξ.

With κψ∈𝒮n′ such that ℱ⁢κψ=Λ, we have

Λ⁢f=(ℱ⁢κψ)⁢(f)=κψ⁢(f^).

Because f↦f^ is a bijection 𝒮n→𝒮n, this shows that for any f∈𝒮n we have

κψ⁢(f)=∫ℝnf⁢(ξ)⁢kt⁢(ξ)⁢𝑑ξ.

Hence,

et⁢Δ⁢f=κψ*f=kt*f,t>0,f∈𝒮n. (1)

Suppose that ϕ:(0,∞)→ℂ and ω:(0,∞)→(0,∞) are functions and that

ψ⁢(s)=∫0∞ϕ⁢(τ)⁢e-s⁢ω⁢(τ)⁢𝑑τ,s>0.

Manipulating symbols suggests that it may be true that

ψ⁢(-Δ)=∫0∞ϕ⁢(τ)⁢eω⁢(τ)⁢Δ⁢𝑑τ,

and then, for f∈𝒮n,

ψ⁢(-Δ)⁢f=∫0∞ϕ⁢(τ)⁢eω⁢(τ)⁢Δ⁢f⁢𝑑τ=∫0∞ϕ⁢(τ)⁢(kω⁢(τ)*f)⁢𝑑τ,

and hence

κψ⁢(x)=∫0∞ϕ⁢(τ)⁢kω⁢(τ)⁢(x)⁢𝑑τ,x∈ℝn. (2)

Take ψ⁢(s)=s-β with 0<Re⁢β<n2. Because Re⁢β<n2, one checks that

Λ⁢f=∫ℝnf⁢(ξ)⁢(4⁢π2⁢|ξ|2)-β⁢𝑑ξ

is a tempered distribution. As Re⁢β>0, we have

s-β=1Γ⁢(β)⁢∫0∞τβ-1⁢e-s⁢τ⁢𝑑τ

and writing ϕ⁢(τ)=τβ-1Γ⁢(β) and ω⁢(τ)=τ, we suspect from (2) that the convolution kernel of (-Δ)-β is

κψ⁢(x)=∫0∞τβ-1Γ⁢(β)⁢kτ⁢(x)⁢𝑑τ,

which one calculates is equal to

Γ⁢(n2-β)Γ⁢(β)⁢4β⁢πn/2⁢|x|n-2⁢β. (3)

What we have written so far does not prove that this is the convolution kernel of (-Δ)-β because it used (2), but it is straightforward to calculate that indeed the convolution kernel of (-Δ)-β is (3). This calculation is explained in an exercise in Folland.88 8 Gerald B. Folland, Introduction to Partial Differential Equations, second ed., p. 154, Exercise 1.

Taking α=2⁢β and defining

Rα⁢(x)=Γ⁢(n-α2)Γ⁢(α2)⁢2α⁢πn/2⁢|x|n-α,0<Re⁢α<n,x∈ℝn,

we call Rα the Riesz potential of order α. Taking as granted that (3) is the convolution kernel of (-Δ)-β, we have

(-Δ)-α/2⁢f=Rα*f,f∈𝒮n.

Then, if n>2 and α=2 satisfies 0<Re⁢α<n, we work out that

R2⁢(x)=1(n-2)⁢ωn⁢|x|n-2,

where ωn=2⁢πn/2Γ⁢(n/2), and hence

(-Δ)-1⁢f=R2*f,f∈𝒮n,

and applying -Δ we obtain

f=-Δ⁢(R2*f)=(-Δ⁢R2)*f,

hence -Δ⁢R2=δ. That is, R2 is the fundamental solution for -Δ.

Suppose that Re⁢β>0. Then, using the definition of Γ⁢(β) as an integral, with ψ⁢(s)=(1+s)-β, we have

ψ⁢(s)=1Γ⁢(β)⁢∫0∞τβ-1⁢e-(1+s)⁢τ⁢𝑑τ,s>0.

Manipulating symbols suggests that

ψ⁢(-Δ)=1Γ⁢(β)⁢∫0∞τβ-1⁢e-τ⁢eτ⁢Δ⁢𝑑τ,

and using (1), assuming the above is true we would have for all f∈𝒮n,

ψ⁢(-Δ)⁢f=1Γ⁢(β)⁢∫0∞τβ-1⁢e-τ⁢eτ⁢Δ⁢f⁢𝑑τ=1Γ⁢(β)⁢∫0∞τβ-1⁢e-τ⁢(kτ*f)⁢𝑑τ,

whose convolution kernel is

1Γ⁢(β)⁢∫0∞τβ-1⁢e-τ⁢kτ⁢𝑑τ.

We write α=2⁢β and define, for Re⁢α>0,

Bα⁢(x)=1Γ⁢(α2)⁢(4⁢π)n/2⁢∫0∞τα-n2-1⁢e-τ-|x|24⁢τ⁢𝑑τ,x≠0.

We call Bα the Bessel potential of order α. It is straightforward to show, and shown in Folland, that ∥Bα∥1<∞, so Bα∈L1⁢(ℝn). Therefore we can take the Fourier transform of Bα, and one calculates that it is

B^α⁢(ξ)=(1+4⁢π2⁢|ξ|2)-α/2,ξ∈ℝn,

and then

ψ⁢(-Δ)=(1-Δ)-α/2⁢f=Bα*f,f∈𝒮n.

5 Gaussian measure

If μ is a measure on ℝn and f:ℝn→ℂ is a function such that for every x∈ℝn the integral ∫ℝnf⁢(x-y)⁢𝑑μ⁢(y) converges, we define the convolution μ*f:ℝn→ℂ by

(μ*f)⁢(x)=μ⁢(τx⁢fˇ)=∫ℝn(τx⁢fˇ)⁢(y)⁢𝑑μ⁢(y)=∫ℝnfˇ⁢(y-x)⁢𝑑μ⁢(y)=∫ℝnf⁢(x-y)⁢𝑑μ⁢(y).

Let νt be the measure on ℝn with density kt. We call νt Gaussian measure. It satisfies

νt*f⁢(x)=∫ℝnf⁢(x-y)⁢𝑑νt⁢(y)=∫ℝnf⁢(x-y)⁢kt⁢(y)⁢𝑑y=f*kt⁢(x),x∈ℝn.