The Berry-Esseen theorem

Jordan Bell
June 3, 2015

1 Cumulative distribution functions

For a random variable X:(Ω,ℱ,P)→ℝ, we define its cumulative distribution function FX:R→[0,1] by

FX(x)=P(X≤x)=∫{X≤x}dP=∫t≤xd(X*P)(t)=(X*P)((-∞,x]).

A distribution function is a function F:ℝ→[0,1] such that (i) F⁢(-∞)=limx→-∞⁡F⁢(x)=0, (ii) F⁢(∞)=limx→∞⁡F⁢(x)=1, (iii) F is nondecreasing, (iv) F is right-continuous: for each x∈ℝ,

F⁢(x+)=limt↓x⁡F⁢(t)=F⁢(x).

It is a fact that the cumulative distribution function of a random variable is a distribution function and that for any distribution function F there is a random variable X for which F=FX.

Let γ1 be the standard Gaussian measure on R: γ1 has density

p⁢(t,0,1)=12⁢π⁢e-t2/2

with respect to Lebesgue measure on ℝ. Let Φ be the cumulative distribution function of γ1:

Φ⁢(x)=γ1⁢((-∞,x])=∫-∞x𝑑γ1⁢(t)=∫-∞x12⁢π⁢e-t2/2⁢𝑑t.

We first prove the following lemma about distribution functions.11 1 Kai Lai Chung, A Course in Probability Theory, third ed., p. 236, Lemma 1; cf. Allan Gut, Probability: A Graduate Course, second ed., p. 358, Lemma 6.1.

Lemma 1.

Suppose that F is a distribution function, that G:ℝ→ℝ satisfies

G⁢(-∞)=limx→-∞⁡G⁢(x)=0,G⁢(∞)=limx→∞⁡G⁢(x)=1,

and that G is differentiable and its derivative satisfies

M=supx∈ℝ⁡|G′⁢(x)|<∞. (1)

Writing

Δ=12⁢M⁢supx∈ℝ⁡|F⁢(x)-G⁢(x)|,

there is some a∈ℝ such that for all T>0,

2⁢M⁢T⁢Δ⁢(3⁢∫0T⁢Δ1-cos⁡xx2⁢𝑑x-π)≤|∫ℝ1-cos⁡T⁢xx2⁢(F⁢(x+a)-G⁢(x+a))⁢𝑑x|.
Proof.

Because G⁢(-∞)=0 and G⁢(∞)=1, there is some compact interval K such that -1<G⁢(x)<2 for x∈ℝ∖K. Then, because G is continuous it is bounded on K, showing that G is bounded on ℝ, and because M>0 we get Δ<∞.

Write H=F-G. Because H⁢(∞)=0 and H⁢(-∞)=0, there is a compact interval K for which

2⁢M⁢Δ=supx∈ℝ⁡|H⁢(x)|=supx∈K⁡|H⁢(x)|.

By the Bolzano-Weierstrass theorem, either there is a sequence un∈K increasing to some u∈K such that |H⁢(un)|↑2⁢M⁢Δ or there is a sequence un∈K decreasing to some u∈K such that |H⁢(un)|↑2⁢M⁢Δ.22 2 In the proof in Chung there are merely two cases but it is not explained why those are exhaustive. In the first case, either there is a subsequence vn of un such that |H⁢(vn)|=H⁢(vn) or there is a subsequence vn of un such that |H⁢(vn)|=-H⁢(vn). In the first subcase we get H⁢(u-)=2⁢M⁢Δ, thus

F⁢(u-)-G⁢(u)=2⁢M⁢Δ. (2)

In the second subcase we get H⁢(u-)=-2⁢M⁢Δ, thus

F⁢(u-)-G⁢(u)=-2⁢M⁢Δ. (3)

In the second case, either there is a subsequence vn of un such that |H⁢(vn)|=H⁢(vn) or there is a subsequence vn of un such that |H(vn)=-H(vn). In the first subcase we get H⁢(u+)=2⁢M⁢Δ, thus

F⁢(u)-G⁢(u)=2⁢M⁢Δ. (4)

In the second subcase we get H⁢(u+)=2⁢M⁢Δ, thus

F⁢(u)-G⁢(u)=-2⁢M⁢Δ. (5)

We now deal with the subcase (3). Let a=u-Δ. For |x|<Δ, by (1) we have

|G⁢(x+a)-G⁢(u)|=|∫uu+x-ΔG′⁢(y)⁢𝑑y|≤|x-Δ|⁢M=(Δ-x)⁢M,

whence

G⁢(x+a)≥G⁢(u)+(x-Δ)⁢M.

Because x+a=x+u-Δ<u and as F is nondecreasing and using (3),

F⁢(x+a)-G⁢(x+a) ≤F⁢(u-)-G⁢(x+a)
≤F⁢(u-)-(G⁢(u)+(x-Δ)⁢M)
=-M⁢(x+Δ).

Then, because x↦1-cos⁡T⁢xx2⁢x is an odd function,

∫-ΔΔ1-cos⁡T⁢xx2⁢(F⁢(x+a)-G⁢(x+a))⁢𝑑x ≤-M⁢∫-ΔΔ1-cos⁡T⁢xx2⁢(x+Δ)⁢𝑑x
=-2⁢M⁢Δ⁢∫0Δ1-cos⁡T⁢xx2⁢𝑑x.

On the other hand,

|∫(-∞,-Δ)∪(Δ,∞)1-cos⁡T⁢xx2⁢(F⁢(x+a)-G⁢(x+a))⁢𝑑x|≤2⁢M⁢Δ⁢∫(-∞,-Δ)∪(Δ,∞)1-cos⁡T⁢xx2⁢𝑑x=4⁢M⁢Δ⁢∫Δ∞1-cos⁡T⁢xx2⁢𝑑x.

Thus

∫ℝ1-cos⁡T⁢xx2⁢(F⁢(x+a)-G⁢(x+a))⁢𝑑x≤-2⁢M⁢Δ⁢∫0Δ1-cos⁡T⁢xx2⁢𝑑x+4⁢M⁢Δ⁢∫Δ∞1-cos⁡T⁢xx2⁢𝑑x=2⁢M⁢Δ⁢(-3⁢∫0Δ1-cos⁡T⁢xx2⁢𝑑x+2⁢∫0∞1-cos⁡T⁢xx2⁢𝑑x)=2⁢M⁢Δ⁢(-3⁢∫0Δ1-cos⁡T⁢xx2⁢𝑑x+2⋅π⁢T2)=2⁢M⁢T⁢Δ⁢(-3⁢∫0T⁢Δ1-cos⁡xx2⁢𝑑x+π),

which yields the claim of the lemma, for the subcase (3). ∎

We now prove a lemma that gives an inequality for characteristic functions.33 3 Kai Lai Chung, A Course in Probability Theory, third ed., p. 237, Lemma 2; Zhengyan Lin and Zhidong Bai, Probability Inequalities, p. 29, Theorem 4.1.a. We remark that because F is a distribution function, it makes sense to speak about the measure induced by F, and because G is of bounded variation and is continuous, its variation function VG is continuous and the functions VG-G and VG are nondecreasing, and it thus makes sense to speak about the signed measure induced by G=VG-(VG-G), which is equal to the difference of the measures induced by VG and VG-G.

Lemma 2.

Suppose that F is a distribution function, that G:ℝ→ℝ satisfies

G⁢(-∞)=limx→-∞⁡G⁢(x)=0,G⁢(∞)=limx→∞⁡G⁢(x)=1,

that G is differentiable and of bounded variation and that its derivative satisfies

M=supx∈ℝ⁡|G′⁢(x)|<∞,

and that

∫ℝ|F-G|⁢𝑑x<∞.

Write

Δ=12⁢M⁢supx∈ℝ⁡|F⁢(x)-G⁢(x)|

and

f⁢(t)=∫ℝei⁢t⁢x⁢𝑑F⁢(x),g⁢(t)=∫ℝei⁢t⁢x⁢𝑑G⁢(x).

Then for all T>0,

Δ≤1π⁢M⁢∫0T|f⁢(t)-g⁢(t)|t⁢𝑑t+12π⁢T.
Proof.

For any t∈ℝ, because (F-G)⁢(-∞)=0 and (F-G)⁢(∞)=0 and because ∫ℝ|F-G|⁢𝑑x<∞, integrating by parts gives

f⁢(t)-g⁢(t) =∫ℝei⁢t⁢x⁢𝑑F⁢(x)-∫ℝei⁢t⁢x⁢𝑑G⁢(x)
=∫ℝei⁢t⁢x⁢d⁢(F-G)⁢(x)
=-i⁢t⁢∫ℝ(F-G)⁢(x)⁢ei⁢t⁢x⁢𝑑x.

Take a to be the real number that Lemma 1 yields. As

f⁢(t)-g⁢(t)-i⁢t⁢e-i⁢t⁢a⁢(T-|t|)=(T-|t|)⁢∫ℝ(F⁢(x+a)-G⁢(x+a))⁢ei⁢t⁢x⁢𝑑x,

we obtain, using Fubini’s theorem,

∫-TTf⁢(t)-g⁢(t)-i⁢t⁢e-i⁢t⁢a⁢(T-|t|)⁢𝑑t=∫-TT((T-|t|)⁢∫ℝ(F⁢(x+a)-G⁢(x+a))⁢ei⁢t⁢x⁢𝑑x)⁢𝑑t=∫ℝ(F⁢(x+a)-G⁢(x+a))⁢(∫-TT(T-|t|)⁢ei⁢t⁢x⁢𝑑t)⁢𝑑x=2⁢∫ℝ(F⁢(x+a)-G⁢(x+a))⁢1-cos⁡T⁢xx2⁢𝑑x.

Therefore, because F and G are real valued and thus |f⁢(-t)-g⁢(-t)|=|f⁢(t)-g⁢(t)¯|=|f⁢(t)-g⁢(t)|,

|∫ℝ(F⁢(x+a)-G⁢(x+a))⁢1-cos⁡T⁢xx2⁢𝑑x| ≤12⁢∫-TT|f⁢(t)-g⁢(t)||t|⁢(T-|t|)⁢𝑑t
=∫0T|f⁢(t)-g⁢(t)|t⁢(T-t)⁢𝑑t
≤T⁢∫0T|f⁢(t)-g⁢(t)|t⁢𝑑t.

Using this with Lemma 1,

2⁢M⁢T⁢Δ⁢(3⁢∫0T⁢Δ1-cos⁡xx2⁢𝑑x-π) ≤T⁢∫0T|f⁢(t)-g⁢(t)|t⁢𝑑t.

But

3⁢∫0T⁢Δ1-cos⁡xx2⁢𝑑x-π =3⁢∫0∞1-cos⁡xx2⁢𝑑x-3⁢∫T⁢Δ∞1-cos⁡xx2⁢𝑑x-π
≥3⁢∫0∞1-cos⁡xx2⁢𝑑x-6⁢∫T⁢Δ∞1x2⁢𝑑x-π
=3⋅π2-6T⁢Δ-π
=π2-6T⁢Δ,

with which we have

2⁢M⁢T⁢Δ⁢(3⁢∫0T⁢Δ1-cos⁡xx2⁢𝑑x-π)≥2⁢M⁢T⁢Δ⋅(π2-6T⁢Δ)=M⁢T⁢Δ⁢π-12⁢M,

and hence

M⁢T⁢Δ⁢π-12⁢M≤T⁢∫0T|f⁢(t)-g⁢(t)|t⁢𝑑t,

i.e.

Δ≤12π⁢T+1π⁢M⁢∫0T|f⁢(t)-g⁢(t)|t⁢𝑑t,

proving the claim. ∎

2 Berry-Esseen theorem

Let Xn,j, n≥1, 1≤j≤kn, be L3 random variables, with kn→∞, such that for each n, the random variables Xn,j, 1≤j≤kn, are independent, and such that for all n and j,

E⁢(Xn,j)=0.

Let Fn,j be the cumulative distribution function of Xn,j:

Fn,j(x)=P(Xn,j≤x).

Let fn,j be the characteristic function of Xn,j (equivalently, the characteristic function of Fn,j):

fn,j⁢(t)=∫ℝei⁢t⁢x⁢d⁢(Xn,j*⁢P)⁢(x)=∫ℝei⁢t⁢x⁢𝑑Fn,j⁢(x).

Write, for n≥1,

Sn=∑j=1knXn,j,

and let Fn be the cumulative distribution function of Sn:

Fn(x)=P(Sn≤x)

Also, let fn be the characteristic function of Sn (equivalently, the characteristic function of Fn). Because Xn,j, 1≤j≤kn, are independent, we have Sn*⁢P=(Xn,1*⁢P)*⋯*(Xn,kn*⁢P) and hence

fn⁢(t)=∫ℝei⁢t⁢x⁢d⁢(Sn*⁢P)⁢(x)=∏j=1knfn,j⁢(t).

For n≥1 and 1≤j≤kn, write

σn,j2=E⁢(Xn,j2),sn2=∑j=1knσn,j2

and

γn,j=E⁢(|Xn,j|3),Γn=∑j=1knγn,j.

We further assume that for each n,

sn2=∑j=1knσn,j2=1. (6)

We will use the following inequality which we state separately because it is of general use.

Lemma 3.

For n≥1 and |z|<1,

|log⁡(1+z)-∑m=1n-1(-1)m-1⁢zmm|≤|z|nn⁢(1-|z|).

We now prove an inequality for fn, the characteristic function of Sn.44 4 Kai Lai Chung, A Course in Probability Theory, third ed., p. 239, Lemma 3.

Lemma 4.

For n≥1, if |t|<12⁢Γn1/3 then

|fn⁢(t)-e-t2/2|≤Γn⁢|t|3⁢e-t2/2.
Proof.

For 1≤j≤kn and l≥0 and v∈ℝ,

fn,j(l)⁢(v)=(i)l⁢E⁢(Xn,jl⁢ei⁢v⁢Xn,j).

Thus

fn,j⁢(0)=1,fn,j′⁢(0)=i⁢E⁢(Xn,j)=0,fn,j′′⁢(0)=-E⁢(Xn,j2)=-σn,j2,

and

fn,j′′′⁢(v)=-i⁢E⁢(Xn,j3⁢ei⁢v⁢Xn,j).

Then by Taylor’s theorem, there is some s between 0 and t such that

fn,j⁢(t)=1-σn,j22⁢t2-i⁢E⁢(Xn,j3⁢ei⁢s⁢Xn,j)6⁢t3.

Put

-i⁢E⁢(Xn,j3⁢ei⁢s⁢Xn,j)=θ⁢γn,j,

for which

|θ|=|E⁢(Xn,j3⁢ei⁢s⁢Xn,j)|E⁢(|Xn,j|3)≤1.

Because the L2 norm is upper bounded by the L3 norm and because |t|<12⁢Γn1/3,

|σn,j⁢t|≤|γn,j1/3⁢t|≤|Γn1/3⁢t|<12,

and hence

|fn,j⁢(t)-1| =|-σn,j22⁢t2+θ⁢γn,j⁢t36|
≤12⁢|σn,j⁢t|2+γn,j48⁢Γn
<18+148
<14.

Lemma 3 and the inequality |a+b|2≤2⁢(|a|2+|b|2) then tell us that

|log⁡fn,j⁢(t)-(fn,j⁢(t)-1)| ≤|fn,j⁢(t)-1|22⁢(1-|fn,j⁢(t)-1|)
<23⁢|fn,j⁢(t)-1|2
=23⁢|-σn,j22⁢t2+θ⁢γn,j6⁢t3|2
≤43⁢(σn,j44⁢t4+|θ|2⁢γn,j236⁢t6).

Because σn,j≤γn,j1/3 and |θ|≤1,

|log⁡fn,j⁢(t)-(fn,j⁢(t)-1)| ≤43⁢(σn,j⁢γn,j4⁢t4+γn,j236⁢t6)
=43⁢(|σn,j⁢t|4+|γn,j1/3⁢t|336)⁢γn,j⁢|t|3
≤43⁢(12⋅4+18⋅36)⁢γn,j⁢|t|3
=37216⁢γn,j⁢|t|3
<15⁢γn,j⁢|t|3.

Combining this with fn,j⁢(t)=1-σn,j22⁢t2+θ⁢γn,j6⁢t3,

|log⁡fn,j⁢(t)+σn,j22⁢t2|≤|θ⁢γn,j6⁢t3|+15⁢γn,j⁢|t|3≤16⁢γn,j⁢|t|3+15⁢γn,j⁢|t|3≤12⁢γn,j⁢|t|3.

Because this is true for each 1≤j≤kn and because, according to (6), ∑j=1knσn,j2=1,

|log⁡fn⁢(t)+t22|≤|t|⁢322⁢∑j=1knγn,j=|t|32⁢Γn.

For any z∈ℂ it is true that |ez-1|≤|z|⁢e|z|, so the above yields

|fn⁢(t)⁢et2/2-1| =|exp⁡(log⁡(fn⁢(t)⁢et2/2))-1|
≤|log⁡(fn⁢(t)⁢et2/2)|⁢exp⁡(|log⁡(fn⁢(t)⁢et2/2)|)
=|log⁡fn⁢(t)+t22|⁢exp⁡(|log⁡(fn⁢(t)⁢et2/2)|)
≤|t|32⁢Γn⁢exp⁡(|t|32⁢Γn).

But |t|3<18⁢Γn, so

|fn⁢(t)⁢et2/2-1|≤|t|32⁢Γn⁢e1/16≤|t|3⁢Γn,

which completes the proof. ∎

The next lemma gives a different bound on the characteristic function of Sn.55 5 Kai Lai Chung, A Course in Probability Theory, third ed., p. 240, Lemma 4.

Lemma 5.

For n≥1, if |t|<14⁢Γn then

|fn⁢(t)|≤e-t2/3.
Proof.

First, for a distribution function F with characteristic function f,

|f⁢(t)|2 =f⁢(t)⁢f⁢(t)¯
=∫ℝei⁢t⁢x⁢𝑑F⁢(x)⋅∫ℝe-i⁢t⁢x⁢𝑑F⁢(y)
=∫ℝ(∫ℝei⁢t⁢(x-y)⁢𝑑F⁢(x))⁢𝑑F⁢(y)
=∫ℝ(∫ℝcos⁡t⁢(x-y)+i⁢sin⁡t⁢(x-y)⁢d⁢F⁢(x))⁢𝑑F⁢(y).

Because |f⁢(t)|2 is real it follows that

|f⁢(t)|2=∫ℝ(∫ℝcos⁡t⁢(x-y)⁢𝑑F⁢(x))⁢𝑑F⁢(y).

Using

|cos⁡u-(1-u22)|≤|u|36,|a+b|p≤2p-1⁢(|a|p+|b|p),

we have

|cos⁡t⁢(x-y)-(1-(t⁢(x-y))22)|≤23⁢(|t⁢x|3+|t⁢y|3)=2⁢|t|33⁢(|x|3+|y|3)

and then

|f⁢(t)|2 ≤∫ℝ(∫ℝ1-(t⁢(x-y))22+2⁢|t|33⁢(|x|3+|y|3)⁢d⁢F⁢(x))⁢𝑑F⁢(y).

Using this for fn,j, and using that E⁢(Xn,j)=0,

|fn,j⁢(t)|2 ≤∫ℝ(∫ℝ1-(t⁢(x-y))22+2⁢|t|33⁢(|x|3+|y|3)⁢d⁢Fn,j⁢(x))⁢𝑑Fn,j⁢(y)
=∫ℝ1-t2⁢σn,j22-t2⁢y22+2⁢|t|3⁢γn,j3+2⁢|t|3⁢|y|33⁢d⁢F⁢(y)
=1-t2⁢σn,j22-t2⁢σn,j22+2⁢|t|3⁢γn,j3+2⁢|t|3⁢γn,j3
=1-t2⁢σn,j2+4⁢|t|3⁢γn,j3.

Because 1+u≤eu for all u∈ℝ,

|fn,j⁢(t)|2≤exp⁡(-t2⁢σn,j2+4⁢|t|3⁢γn,j3).

Then, by (6),

|fn⁢(t)|2 =∏j=1kn|fn,j⁢(t)|2
≤∏j=1knexp⁡(-t2⁢σn,j2+4⁢|t|3⁢γn,j3)
=exp⁡(-t2⁢∑j=1knσn,j2+4⁢|t|33⁢∑j=1knγn,j)
=exp⁡(-t2+4⁢|t|33⁢Γn).

As |t|<14⁢Γn,

|fn⁢(t)|≤exp⁡(-t22+2⁢|t|33⁢Γn)≤exp⁡(-t22+2⁢|t|212)=e-t2/3,

proving the claim. ∎

We now combine Lemma 4 and Lemma 5.66 6 Kai Lai Chung, A Course in Probability Theory, third ed., p. 240, Lemma 5.

Lemma 6.

For n≥1, if |t|<14⁢Γn then

|fn⁢(t)-e-t2/2|≤16⁢Γn⁢|t|3⁢e-t2/3.
Proof.

Either |t|<12⁢Γn1/3 or 12⁢Γn1/3≤|t|<14⁢Γn. In the first case, Lemma 4 tells us

|fn⁢(t)-e-t2/2|≤Γn⁢|t|3⁢e-t2/2≤Γn⁢|t|3⁢e-t2/3≤16⁢Γn⁢|t|3⁢e-t2/3.

In the second case, Lemma 5 tells us

|fn⁢(t)|≤e-t2/3,

and so, as in this case we have 1≤8⁢Γn⁢|t|3,

|fn⁢(t)-e-t2/2|≤|fn⁢(t)|+e-t2/2≤e-t2/3+e-t2/2≤2⁢e-t2/3≤16⁢Γn⁢|t|3⁢e-t2/3,

showing that the claim is true in both cases. ∎

We finally prove the Berry-Esseen theorem.77 7 Kai Lai Chung, A Course in Probability Theory, third ed., p. 235, Theorem 7.4.1; cf. Allan Gut, Probability: A Graduate Course, second ed., p. 356, Theorem 6.2; John E. Kolassa, Series Approximation Methods in Statistics, p. 25, Theorem 2.6.1; Alexandr A. Borovkov, Probability Theory, p. 659, Theorem A5.1; Ivan Nourdin and Giovanni Peccati, Normal Approximations with Malliavin Calculus: From Stein’s Method to Universality, p. 71, Theorem 3.7.1.

Theorem 7 (Berry-Esseen theorem).

There is some A0<36 such that for each n≥1,

supx∈ℝ⁡|Fn⁢(x)-Φ⁢(x)|≤A0⁢Γn.
Proof.

Let Z be a random variable with Z*⁢P=γ1, i.e. whose cumulative distribution function is Φ. By (6) and because Xn,j, 1≤j≤kn, are independent and satisfy E⁢(Xn,j)=0,

E⁢(Sn2)=∑j=1knE⁢(Xn,j2)=∑j=1knσn,j2=1.

If x<0 then by Chebyshev’s inequality

Fn(x)=P(Sn≤x)=P(-Sn≥-x)≤1x2E(|Sn|2)=1x2

and

Φ(x)=P(Z≤x)=P(-Z≥-x)≤1x2E(|Z|2)=1x2.

If x>0 then also by Chebyshev’s inequality

1-Fn(x)=1-P(Sn≤x)=P(Sn>x)≤1x2

and

1-Φ(x)=1-P(Z≤x)=P(Z>x)≤1x2.

Therefore, because Fn and Φ are nonnegative and 1-Fn and 1-Φ are nonnegative, for all x∈ℝ we have

|Fn⁢(x)-Φ⁢(x)|≤1x2.

Then, because |Fn|≤1 and |Φ|≤1,

∫ℝ|Fn⁢(x)-Φ⁢(x)|⁢𝑑x≤∫|x|≤12⁢𝑑x+∫|x|>11x2⁢𝑑x=6<∞.

Φ′⁢(x)=12⁢π⁢e-x2/2≤12⁢π. We apply Lemma 2 with F=Fn, G=Φ, and M=12⁢π, and because the characteristic function of Φ is ϕ⁢(t)=e-t2/2, we obtain for T=14⁢Γn,

supx∈ℝ⁡|Fn⁢(x)-Φ⁢(x)| ≤2π⁢∫014⁢Γn|fn⁢(t)-ϕ⁢(t)|t⁢𝑑t+96⁢M⁢Γnπ
=2π⁢∫014⁢Γn|fn⁢(t)-e-t2/2|t+96⁢Γnπ⁢2⁢π.

Then applying Lemma 6,

supx∈ℝ⁡|Fn⁢(x)-Φ⁢(x)| ≤2π⁢∫014⁢Γn16⁢Γn⁢t3⁢e-t2/3t⁢𝑑t+96⁢Γnπ⁢2⁢π
=Γn⁢(32π⁢∫014⁢Γnt2⁢e-t2/3⁢𝑑t+96π⁢2⁢π).

This proves the claim with

A0=32π⁢∫0∞t2⁢e-t2/3⁢𝑑t+96π⁢2⁢π=32π⋅3⁢3⁢π4+96π⁢2⁢π=35.64⁢….

∎