The singular value decomposition of compact operators on Hilbert spaces

Jordan Bell
April 3, 2014

1 Preliminaries

The purpose of these notes is to present material about compact operators on Hilbert spaces that is special to Hilbert spaces, rather than what applies to all Banach spaces. We use statements about compact operators on Banach spaces without proof. For instance, any compact operator from one Banach space to another has separable image; the set of compact operators from one Banach space to another is a closed subspace of the set of all bounded linear operators; a Banach space is reflexive if and only if the closed unit ball is weakly compact (Kakutani’s theorem); etc. We do however state precisely each result that we are using for Banach spaces and show that its hypotheses are satisfied.

Let ℕ be the set of positive integers. We say that a set is countable if it is bijective with a subset of ℕ. In this note I do not presume unless I say so that any set is countable or that any Hilbert space is separable. A neighborhood of a point in a topological space is a set that contains an open set that contains the point; one reason why it can be handy to speak about neighborhoods of a point rather than just open sets that contain the point is that the set of all neighborhoods of a point is a filter, whereas it is unlikely that the set of all open sets that contain a point is a filter. If z∈ℂ, we denote z*=z¯.

2 Bounded linear operators

An advantage of working with normed spaces rather than merely topological vector spaces is that continuous linear maps between normed spaces have a simple characterization. If X and Y are normed spaces and T:X→Y is linear, the operator norm of T is

∥T∥=sup∥x∥≤1⁡∥T⁢x∥.

If ∥T∥<∞, then we say that T is bounded.

Theorem 1.

If X and Y are normed spaces, a linear map T:X→Y is continuous if and only if it is bounded.

Proof.

Suppose that T is continuous. In particular T is continuous at 0, so there is some δ>0 such that if ∥x∥≤δ then ∥T⁢x∥=∥T⁢x-T⁢0∥≤1. If x≠0 then, as T is linear,

∥T⁢x∥=1δ⁢∥x∥⁢∥T⁢(δ∥x∥⁢x)∥≤1δ⁢∥x∥.

Thus ∥T∥≤1δ<∞, so T is bounded.

Suppose that T is bounded. Let x0∈X, and let ϵ>0. If ∥x-x0∥≤ϵ∥T∥, then

∥T⁢x-T⁢x0∥=∥T⁢(x-x0)∥≤∥T∥⁢∥x-x0∥≤∥T∥⋅ϵ∥T∥=ϵ.

Hence T is continuous at x0, and so T is continuous. ∎

If X and Y are normed spaces, we denote by ℬ⁢(X,Y) the set of bounded linear maps X→Y. It is straightforward to check that ℬ⁢(X,Y) is a normed space with the operator norm. One proves that if Y is a Banach space, then ℬ⁢(X,Y) is a Banach space,11 1 Walter Rudin, Functional Analysis, second ed., p. 92, Theorem 4.1. and if X is a Banach space one then checks that ℬ⁢(X)=ℬ⁢(X,X) is a Banach algebra. If X is a normed space, we define X*=ℬ⁢(X,ℂ), which is a Banach space, called the dual space of X.

If X and Y are normed spaces and T:X→Y is linear, we say that T has finite rank if

rank⁢T=dim⁡T⁢(X)

is finite. If X is infinite dimensional and Y≠{0}, let ℰ be a Hamel basis for X, let {en:n∈ℕ} be a countable subset of ℰ, and let y∈Y be nonzero. If we define T:X→Y by T⁢en=n⁢∥en∥⁢y and T⁢e=0 if e∈ℰ∖{en:n∈ℕ}, then T is a linear map with finite rank yet T is unbounded. Thus a finite rank linear map is not necessarily bounded. We denote by ℬ00⁢(X,Y) the set of bounded finite rank linear maps X→Y, and check that ℬ00⁢(X,Y) is a vector space. If X is a Banach space, one checks that ℬ00⁢(X) is an ideal of the algebra ℬ⁢(X) (if we either pre- or postcompose a linear map with a finite rank linear map, the image will be finite dimensional).

If X and Y are Banach spaces, we say that T:X→Y is compact if the image of any bounded set under T is precompact (has compact closure). One checks that if a linear map is compact then it is bounded (unlike a finite rank linear map, which is not necessarily bounded). There are several ways to state that a linear map is compact that one proves are equivalent: T is compact if and only if the image of the closed unit ball is precompact; T is compact if and only if the image of the open unit ball is precompact; T is compact if and only if the image under it of any bounded sequence has a convergent subsequence. In a complete metric space, the Heine-Borel theorem asserts that a set is precompact if and only if it is totally bounded (for any ϵ>0, the set can be covered by a finite number of balls of radius ϵ). We denote by ℬ0⁢(X,Y) the set of compact linear maps X→Y, and it is straightforward to check that this is a vector space. One proves that if an operator is in the closure of the compact operators then the image of the closed unit ball under it is totally bounded, and from this it follows that ℬ0⁢(X,Y) is a closed subspace of ℬ⁢(X,Y). ℬ0⁢(X) is an ideal of the algebra ℬ⁢(X): if K∈ℬ0⁢(X) and T∈ℬ⁢(X), one checks that T⁢K∈ℬ0⁢(X) and K⁢T∈ℬ0⁢(X).

Let X and Y be Banach spaces. Using the fact that a bounded set in a finite dimensional normed vector space is precompact, we can prove that a bounded finite rank operator is compact: ℬ00⁢(X,Y)⊆ℬ0⁢(X,Y). Also, it doesn’t take long to prove that the image of a compact operator is separable: if T∈ℬ0⁢(X,Y) then T⁢(X) has a countable dense subset. (We can prove this using the fact that a compact metric space is separable.)

If H is a Hilbert space and Si,i∈I are subsets of H, we define ⋁i∈ISi to be the closure of the span of ⋃i∈ISi. We say that ℰ is an orthonormal basis for H if ⟨e,f⟩=δe,f and H=⋁ℰ.

A sesquilinear form on H is a function f:H×H→ℂ that is linear in its first argument and that satisfies f⁢(x,y)=f⁢(y,x)*. If f is a sesquilinear form on H, we say that f is bounded if

sup{|f(x,y)|:∥x∥,∥y∥≤1}<∞.

The Riesz representation theorem22 2 Walter Rudin, Functional Analysis, second ed., p. 310, Theorem 12.8. states that if f is a bounded sesquilinear form on H, then there is a unique B∈ℬ⁢(H) such that

f⁢(x,y)=⟨x,B⁢y⟩,x,y∈H,

and ∥B∥=sup{|f(x,y)|:∥x∥,∥y∥≤1}. It follows from the Riesz representation that if A∈ℬ⁢(H), then there is a unique A*∈ℬ⁢(H) such that

⟨A⁢x,y⟩=⟨x,A*⁢y⟩,x,y∈H,

and ∥A*∥=∥A∥. ℬ⁢(H) is a C*-algebra: if A,B∈ℬ⁢(H) and λ∈ℂ then A**=A, (A+B)*=A*+B*, (A⁢B)*=B*⁢A*, (λ⁢A)*=λ*⁢A*, and ∥A*⁢A∥=∥A∥2.

We say that A∈ℬ⁢(H) is normal if A*⁢A=A⁢A*, and self-adjoint if A*=A. One proves using the parallelogram law that A∈ℬ⁢(H) is self-adjoint if and only if ⟨A⁢x,x⟩∈ℝ for all x∈H. If A∈ℬ⁢(H) is self-adjoint, we say that A is positive if ⟨A⁢x,x⟩≥0 for all x∈H.

3 Spectrum in Banach spaces

If X and Y are Banach spaces and T∈ℬ⁢(X,Y) is a bijection, then its inverse function T-1:Y→X is linear, since the inverse of a linear bijection is itself linear. Because T is a surjective bounded linear map, by the open mapping theorem it is an open map: if U is an open subset of X then T⁢(U) is an open subset of Y, and it follows that T-1∈ℬ⁢(Y,X). That is, if a bounded linear operator from one Banach space to another is bijective then its inverse function is also a bounded linear operator.

If X is a Banach space and T∈ℬ⁢(X), the spectrum σ⁢(T) of T is the set of those λ∈ℂ such that the map T-λ⁢idX:X→X is not a bijection. One proves that σ⁢(T) is nonempty (the proof uses Liouville’s theorem, which states that a bounded entire function is constant). One also proves that if λ∈σ⁢(T) then |λ|≤∥T∥. We define the spectral radius of T to be

r⁢(T)=supλ∈σ⁢(T)⁡|λ|,

and so r⁢(T)≤∥T∥. Because ℬ0⁢(X) is an ideal in the algebra ℬ⁢(X), if T∈ℬ0⁢(X) is invertible then idX is compact. One checks that if idX is compact then X is finite dimensional (a locally compact topological vector space is finite dimensional), and therefore, if X is an infinite dimensional Banach space and T∈ℬ0⁢(X), then 0∈σ⁢(T).

The resolvent set of T is ρ⁢(T)=ℂ∖σ⁢(T). One proves that ρ⁢(T) is open,33 3 Gert K. Pedersen, Analysis Now, revised printing, p. 131, Theorem 4.1.13. from which it then follows that σ⁢(T) is a compact set. For λ∈ρ⁢(T), we define

R⁢(λ,T)=(T-λ⁢idX)-1∈ℬ⁢(X),

called the resolvent of T.

If X is a Banach space and T∈ℬ0⁢(X), then the point spectrum σpoint⁢(T) of T is the set of those λ∈ℂ such that T-λ⁢idX is not injective. In other words, to say that λ∈σpoint⁢(T) is to say that

dim⁡ker⁡(T-λ⁢idX)>0.

If λ∈σpoint⁢(T), we say that λ is an eigenvalue of T, and call dim⁡ker⁡(T-λ⁢idX) its geometric multiplicity.44 4 There is also a notion of algebraic multiplicity of an eigenvalue: the algebraic multiplicity of λ is defined to be supn∈ℕ⁡dim⁡ker⁡((A-λ⁢idH)n). For self-adjoint operators this is equal to the geometric multiplicity of λ, while for a operator that is not self-adjoint the algebraic multiplicity of an eingevalue may be greater than its geometric multiplicity. Nonzero elements of ker⁡((A-λ⁢idH)n) are called generalized eigenvectors or root vectors. See I. C. Gohberg and M. G. Krein, Introduction to the Theory of Linear Nonselfadjoint Operators in Hilbert Space. It is a fact that each nonzero eigenvalue of T has finite geometric multiplicity, and it is also a fact that if T∈ℬ0⁢(X), then σpoint⁢(T) is a bounded countable set and that if σpoint⁢(T) has a limit point that limit point is 0. The Fredholm alternative tells us that

σ⁢(T)⊆σpoint⁢(T)∪{0}.

If X is infinite dimensional then σ⁢(T)=σpoint⁢(T)∪{0}, and σpoint⁢(T) might or might not include 0. If T∈ℬ00⁢(X), check that σpoint⁢(T) is a finite set.

If H is a Hilbert space, using the fact that a bounded linear operator T∈ℬ⁢(H) is invertible T if and only if both T⁢T* and T*⁢T are bounded below (S is bounded below if there is some c>0 such that ∥S⁢x∥≥c⁢∥x∥ for all x∈X), one can prove that the spectrum of a bounded self-adjoint operator is a set of real numbers, and the spectrum of a bounded positive operator is a set of nonnegative real numbers.

4 Numerical radius

If H is a Hilbert space and A∈ℬ⁢(H), the numerical range of A is the set {⟨A⁢x,x⟩:∥x∥=1}.55 5 The Toeplitz-Hausdorff theorem states that the numerical range of any bounded linear operator is a convex set. See Paul R. Halmos, A Hilbert Space Problem Book, Problem 166. The closure of the numerical range contains the spectrum of A.66 6 Paul R. Halmos, A Hilbert Space Problem Book, Problem 169. If A is normal, then the closure of its numerical range is the convex hull of the spectrum: Problem 171. The numerical radius w⁢(A) of A is the supremum of the numerical range of A: w⁢(A)=sup∥x∥=1⁡|⟨A⁢x,x⟩|. If A∈ℬ⁢(H) is self-adjoint, one can prove that77 7 John B. Conway, A Course in Functional Analysis, second ed., p. 34, Proposition 2.13.

w⁢(A)=∥A∥.

The following theorem asserts that a compact self-adjoint operator A has an eigenvalue whose absolute value is equal to the norm of the operator. Thus in particular, the spectral radius of a compact self-adjoint operator is equal to its numerical radius. Since a self-adjoint operator has real spectrum, to say that |λ|=∥A∥ is to say that either λ=∥A∥ or λ=-∥A∥. A compact operator on a Hilbert space can have empty point spectrum (e.g. the Volterra operator on L2⁢([0,1])) and a bounded self-adjoint operator can have empty point spectrum (e.g. the multiplication operator T⁢ϕ⁢(t)=t⁢ϕ⁢(t) on L2⁢([0,1])), but this theorem shows that if an operator is compact and self-adjoint then its point spectrum is nonempty.

Theorem 2.

If A∈B⁢(H) is compact and self-adjoint then at least one of -∥A∥,∥A∥ is an eigenvalue of A.

Proof.

Because A is self-adjoint, ∥A∥=w⁢(A)=sup∥x∥=1⁡|⟨A⁢x,x⟩|. Also, as A is self-adjoint, ⟨A⁢x,x⟩ is a real number, and thus either ∥A∥=sup∥x∥=1⁡⟨A⁢x,x⟩ or ∥A∥=-inf∥x∥⁡⟨A⁢x,x⟩. In the first case, due to ∥A∥ being a supremum there is a sequence xn, all with norm 1, such that ⟨A⁢xn,xn⟩→∥A∥. Using that A is compact, there is a subsequence xa⁢(n) such that A⁢xa⁢(n) converges to some x, and ∥x∥=1 because each xn has norm 1. Using A=A*,

⟨A⁢xn-∥A∥⁢xn,A⁢xn-∥A∥⁢xn⟩ = ⟨A⁢xn,A⁢xn⟩-⟨A⁢xn,∥A∥⁢xn⟩
-⟨∥A∥⁢xn,A⁢xn⟩+⟨∥A∥⁢xn,∥A∥⁢xn⟩
= ∥A⁢xn∥2-2⁢∥A∥⁢⟨A⁢xn,xn⟩+∥A∥2⁢∥xn∥2
≤ ∥A∥2⁢∥xn∥2-2⁢∥A∥⁢⟨A⁢xn,xn⟩+∥A∥2⁢∥xn∥2
= 2⁢∥A∥2⁢∥xn∥2-2⁢∥A∥⁢⟨A⁢xn,xn⟩
→ 2⁢∥A∥2⁢∥x∥2-2⁢∥A∥⁢∥A∥
= 0.

Therefore, as n→∞, the sequence A⁢xn-∥A∥⁢xn tends to 0, so A⁢x-∥A∥⁢x=0, i.e. x is an eigenvector for the eigenvalue ∥A∥. If ∥A∥=-inf∥x∥=1⁡⟨A⁢x,x⟩ the argument goes the same. ∎

5 Polar decomposition

If H is a Hilbert space and P∈ℬ⁢(H) is positive, there is a unique positive element of ℬ⁢(H), denoted P1/2, satisfying (P1/2)2=P, which we call the positive square root of P.88 8 Gert K. Pedersen, Analysis Now, revised printing, p. 92, Proposition 3.2.11. If A∈ℬ⁢(H) one checks that A*⁢A is positive, and hence A*⁢A has a positive square root, which we denote by |A| and call the absolute value of A. One proves that |A| is the unique positive operator in ℬ⁢(H) satisfying

∥A⁢x∥=∥|A|⁢x∥,x∈H.

An element U of ℬ⁢(H) is said to be a partial isometry if there is a closed subspace X of H such that the restriction of U to X is an isometry X→U⁢(X) and ker⁡U=X⟂. One proves that U*⁢U is the orthogonal projection of H onto X. It can be proved that if A∈ℬ⁢(H) then there is a unique partial isometry U satisfying both ker⁡U=ker⁡A and A=U⁢|A|.99 9 Gert K. Pedersen, Analysis Now, revised printing, p. 96, Theorem 3.2.17. This is called the polar decomposition of A. The polar decomposition satisfies

U*⁢U⁢|A|=|A|,U*⁢A=|A|,U⁢U*⁢A=A.

6 Spectral theorem

If e,f∈H, we define e⊗f:H→H by e⊗f⁢(h)=⟨h,f⟩⁢e. e⊗f is linear, and

∥e⊗f⁢(h)∥=∥⟨h,f⟩⁢e∥=|⟨h,f⟩|⁢∥e∥≤∥h∥⁢∥f∥⁢∥e∥,

so ∥e⊗f∥≤∥e∥⁢∥f∥. Depending on whether f=0 the image of e⊗f is {0} or the span of e, and in either case e⊗f∈ℬ00⁢(H). If either of e or f is 0 then e⊗f has rank 0, and otherwise e⊗f has rank 1, and it is an orthogonal projection precisely when f is a multiple of e.

If ℰ is an orthonormal set in a Hilbert space H, then ℰ is an orthonormal basis for H if and only if the unordered sum

∑e∈ℰe⊗e

converges strongly to idH.1010 10 John B. Conway, A Course in Functional Analysis, second ed., p. 16, Theorem 4.13.

Let’s summarize what we have stated so far about the spectrum and point spectrum of a compact self-adjoint operator on a Hilbert space.

Theorem 3 (Spectrum of compact self-adjoint operators).

If H is a Hilbert space and A∈B0⁢(H) is self-adjoint, then:

  • •

    σ⁢(A) is a nonempty compact subset of ℝ.

  • •

    If H is infinite dimensional, then 0∈σ⁢(A).

  • •

    σ⁢(A)⊆σpoint⁢(A)∪{0}.

  • •

    σpoint⁢(A) is countable.

  • •

    If λ∈ℝ is a limit point of σpoint⁢(A), then λ=0.

  • •

    At least one of ∥A∥,-∥A∥ is an element of σpoint⁢(A).

  • •

    Each nonzero eigenvalue of A has finite geometric multiplicity: If λ∈σpoint⁢(A) and λ≠0, then dim⁡ker⁡(A-λ⁢idH)<∞.

  • •

    If A∈ℬ00⁢(H), then σpoint⁢(A) is a finite set.

We say that A∈ℬ⁢(H) is diagonalizable if there is an orthonormal basis ℰ for H and a bounded set {λe∈ℂ:e∈ℰ} such that the unordered sum

∑e∈ℰλe⁢e⊗e

converges strongly to A.

The following is the spectral theorem for normal compact operators.1111 11 Gert K. Pedersen, Analysis Now, revised printing, p. 108, Theorem 3.3.8.

Theorem 4 (Spectral theorem).

If A∈B0⁢(H) is normal, then A is diagonalizable.

The last assertion of Theorem 3 is that a bounded self-adjoint finite rank operator on a Hilbert space has finitely many elements in its point spectrum. Using the spectral theorem, we get that if A∈ℬ0⁢(H) is self-adjoint and σpoint⁢(A) is finite, then A is finite rank. In the notation we introduce in the following definition, ν⁢(A)<∞ precisely when A has finite rank.

Definition 5.

If A∈ℬ0⁢(H) is self-adjoint, define 0≤ν⁢(A)≤∞ to be the sum of the geometric multiplicities of the nonzero eigenvalues of A:

ν⁢(A)=∑λ∈σpoint⁢(A)∖{0}dim⁡ker⁡(A-λ⁢idH).

Define

(λn(A):n∈ℕ)∈ℝℕ

to be the sequence whose first term is the element of σpoint⁢(A)∖{0} with largest absolute value repeated as many times as its geometric multiplicity. If λ,-λ are both nonzero elements of σpoint⁢(A), we put the positive one first. We repeat this for the remaining elements of σpoint⁢(A)∖{0}. If ν⁢(A)<∞, we define λn⁢(A)=0 for n>ν⁢(A).

Using the spectral theorem and the notation in the above definition we get the following.

Theorem 6.

If A∈B0⁢(H) is self-adjoint, then there is an orthonormal set {en:n∈N} in H such that

∑n∈ℕλn⁢(A)⁢en⊗en

converges strongly to A.

If A∈ℬ0⁢(H), then its absolute value |A| is a positive compact operator, and λn⁢(|A|)≥0 for all n∈ℕ.

Definition 7.

If A∈ℬ0⁢(H) and λ is an eigenvalue of |A|, we call λ a singular value of A, and we define

σn⁢(A)=λn⁢(|A|),n∈ℕ.

Because the absolute value of the absolute value of an operator is the absolute value of the operator, if A∈ℬ0⁢(H) and n∈ℕ then σn⁢(|A|)=σn⁢(A). If A∈ℬ0⁢(H) and λ≠0, one proves that λ is an eigenvalue of A⁢A* if and only if λ is an eigenvalue of A*⁢A and that they have the same geometric multiplicity. From this we get that σn⁢(A)=σn⁢(A*) for all n∈ℕ.

7 Finite rank operators

Theorem 8 (Singular value decomposition).

If H is a Hilbert space and A∈B00⁢(H) has rank⁢A=N, then there is an orthonormal set {en:1≤n≤N} and an orthonormal set {fn:1≤n≤N} such that

A=∑n=1Nσn⁢(A)⁢en⊗fn,A⁢h=∑n=1Nσn⁢(A)⁢⟨h,fn⟩⁢en.
Proof.

|A| is a positive operator with rank⁢|A|=N, and according to Theorem 6, there is an orthonormal set {fn:n∈ℕ} in H such that

|A|=∑n∈ℕλn⁢(|A|)⁢fn⊗fn=∑n∈ℕσn⁢(A)⁢fn⊗fn.

Using the polar decomposition A=U⁢|A|,

A=U⁢|A|=∑n=1Nσn⁢(A)⁢(U⁢fn)⊗fn.

Define en=U⁢fn. As U*⁢U⁢|A|=|A| and as |A|⁢fmσm⁢(A)=fm,

⟨en,em⟩ = ⟨U⁢fn,U⁢fm⟩
= ⟨fn,U*⁢U⁢fm⟩
= ⟨fn,U*⁢U|A|fmσm⁢(A)⟩
= ⟨fn,|A|⁢fmσm⁢(A)⟩
= ⟨fn,fm⟩
= δn,m,

showing that {en:1≤n≤N} is an orthonormal set. ∎

If e,f,x,y∈H, then

⟨e⊗f⁢(x),y⟩=⟨⟨x,f⟩⁢e,y⟩=⟨x,f⟩⁢⟨e,y⟩=⟨y,e⟩*⁢⟨x,f⟩=⟨x,⟨y,e⟩⁢f⟩=⟨x,f⊗e⁢(y)⟩,

so (e⊗f)*=f⊗e.

Theorem 9.

If A∈B00⁢(H) then A*∈B00⁢(H).

Proof.

Let A have the singular value decomposition

A=∑n=1Nσn⁢(A)⁢en⊗fn.

Taking the adjoint, and because σn∈ℝ,

A*=∑n=1Nσn⁢(A)⁢fn⊗en.

A* is a sum of finite rank operators and is therefore itself a finite rank operator. ∎

8 Compact operators

If X and Y are Banach spaces, ℬ00⁢(X,Y)⊆ℬ0⁢(X,Y). But if H is a Hilbert space we can say much more: ℬ00⁢(H) is a dense subset of ℬ0⁢(H). In other words, any compact operator on a Hilbert space can be approximated by a sequence of bounded finite rank operators.1212 12 John B. Conway, A Course in Functional Analysis, second ed., p. 41, Theorem 4.4. As the adjoint An* of each of these finite rank operators An is itself a bounded finite rank operator,

∥A*-An*∥=∥(A-An)*∥=∥A-An∥→0,

so An*→A*. Because each bounded finite rank operator is compact and ℬ0⁢(H) is closed, this establishes that A*∈ℬ0⁢(H). (In fact, it is true that the adjoint of a compact linear operator between Banach spaces is itself compact, but there we don’t have the tool of showing that the adjoint is the limit of the adjoints of finite rank operators.)

If H is a Hilbert space, the weak topology is the topology on H such that a net xα converges to x weakly if for all h∈H the net ⟨xα,h⟩ converges to ⟨x,h⟩ in ℂ. Let 𝔅 be the closed unit ball in H, and let it be a topological space with the subspace topology inherited from H with the weak topology. Thus, a net xα∈𝔅 converges to x∈𝔅 if and only if for all h∈H the net ⟨xα,h⟩ converges to ⟨x,h⟩.

Theorem 10.

If H is a Hilbert space, A∈B⁢(H), and B is the closed unit ball in H with the subspace topology inherited from H with the weak topology, then A is compact if and only if A|B:B→H is continuous.

Proof.

Suppose that A is compact and let xα be a net in 𝔅 that converges weakly to some x∈𝔅. If ϵ>0, then there is some B∈ℬ00⁢(H) with ∥A-B∥<ϵ. Let B have the singular value decomposition

B=∑n=1Nσn⁢(B)⁢en⊗fn.

We have, using that the en are orthonormal,

∥B⁢xα-B⁢x∥2 = ∥∑n=1Nσn⁢(B)⁢⟨xα,fn⟩⁢en-∑n=1Nσn⁢(B)⁢⟨x,fn⟩⁢en∥2
= ∥∑n=1Nσn⁢(B)⁢⟨xα-x,fn⟩⁢en∥2
= ∑n=1Nσn⁢(B)2⁢|⟨xα-x,fn⟩|2.

Eventually this is <ϵ3, and for such α,

∥A⁢xα-A⁢x∥ ≤ ∥A⁢xα-B⁢xα∥+∥B⁢xα-B⁢x∥+∥B⁢x-A⁢x∥
≤ ∥A-B∥⁢∥xα∥+∥B⁢xα-B⁢x∥+∥B-A∥⁢∥x∥
≤ ∥A-B∥+∥B⁢xα-B⁢x∥+∥B-A∥
< ϵ3+ϵ3+ϵ3
= ϵ.

We have shown that A⁢xα→A⁢x in the no rm of H, and this shows that A|𝔅:𝔅→H is continuous.

Suppose that A|𝔅:𝔅→H is continuous. Kakutani’s theorem states that a Banach space is reflexive if and only if the closed unit ball is weakly compact. A Hilbert space is reflexive, hence 𝔅, the closed unit ball with the weak topology, is a compact topological space.1313 13 cf. Paul R. Halmos, A Hilbert Space Problem Book, Problem 17. Since A|𝔅:𝔅→H is continuous and 𝔅 is compact, the image A⁢(𝔅) is compact (the image of a compact set under a continuous map is a compact set). We have shown that the image of the closed unit ball is a compact subset of H, and this shows that A is compact; in fact, to have shown that A is compact we merely needed to show that the image of the closed unit ball is precompact, and H is a Hausdorff space so a compact set is precompact. ∎

A compact linear operator on an infinite dimensional Hilbert space H is not invertible, lest idH be compact. However, operators of the form A-λ⁢idH may indeed be invertible.1414 14 Ward Cheney, Analysis for Applied Mathematics, p. 94, Theorem 2.

Theorem 11.

If A∈B0⁢(H) is a normal operator with diagonalization

A=∑n=1∞λn⁢en⊗en

and 0≠λ∉σpoint⁢(A), then A-λ⁢idH is invertible and

(A-λ⁢idH)-1=-1λ+1λ⁢∑n=1∞λnλn-λ⁢en⊗en,

where the series converges in the strong operator topology.

Proof.

As λn→0 we have α=supn⁡|λn|<∞, and as λ≠0 we have β=infn⁡|λn-λ|>0. Define

TN=-1λ+1λ⁢∑n=1Nλnλn-λ⁢en⊗en∈ℬ00⁢(H),

and if N>M, then, for any h∈H,

∥TN⁢h-TM⁢h∥2 = 1|λ|2⁢∥∑n=M+1Nλnλn-λ⁢⟨h,en⟩⁢en∥2
= 1|λ|2⁢∑n=M+1N|λn|2|λn-λ|2⁢|⟨h,en⟩|2
≤ 1|λ|2⁢∑n=M+1Nα2⁢|⟨h,en⟩|2β2
= α2|λ|2⁢β2⁢∑n=M+1N|⟨h,en⟩|2.

By Bessel’s inequality, ∑n=1∞|⟨h,en⟩|2≤∥h∥2, hence ∑n=N∞|⟨h,en⟩|2→0 as N→∞; this N depends on h, and this is why the claim is stated merely for the strong operator topology and not the norm topology. We have shown that TN⁢h is a Cauchy sequence in H and hence TN⁢h converges. We define B⁢h to be this limit. For h∈H,

∥1λ⁢∑n=1∞λnλn-λ⁢⟨h,en⟩⁢en∥2 = 1|λ|2⁢∑n=1∞|λn|2|λn-λ|2⁢|⟨h,en⟩|2
≤ α2|λ|2⁢β2⁢∑n=1∞|⟨h,en⟩|2
≤ α2|λ|2⁢β2⁢∥h∥2,

whence

∥B⁢h∥ ≤ ∥-1λ⁢h∥+∥1λ⁢∑n=1∞λnλn-λ⁢⟨h,en⟩⁢en∥
≤ 1|λ|⁢∥h∥+α|λ|⁢β⁢∥h∥,

showing that ∥B∥≤1|λ|+α|λ|⁢β. It is straightforward to check that B is linear, thus B∈ℬ⁢(H). (Thus B is a strong limit of finite rank operators. But if H is infinite dimensional then B is in fact not the norm limit of the sequence: for if it were it would be compact, and we will show that B is invertible, which would tell us that idH is compact, contradicting H being infinite dimensional.)

For h∈H,

(A-λ⁢idH)⁢B⁢h = -1λ⁢(A⁢h-λ⁢h)+1λ⁢∑n=1∞λnλn-λ⁢⟨h,en⟩⁢(A⁢en-λ⁢en)
= h-1λ⁢A⁢h+1λ⁢∑n=1∞λnλn-λ⁢⟨h,en⟩⁢(λn⁢en-λ⁢en)
= h-1λ⁢A⁢h+1λ⁢∑n=1∞λn⁢⟨h,en⟩⁢en
= h,

where the final equality is because the series is the diagonalization of A. On the other hand,

B⁢(A-λ⁢idH)⁢h = -1λ⁢(A-λ⁢idH)⁢h+1λ⁢∑n=1∞λnλn-λ⁢⟨(A-λ⁢idH)⁢h,en⟩⁢en
= h-1λ⁢A⁢h+1λ⁢∑n=1∞λnλn-λ⁢⟨A⁢h-λ⁢h,en⟩⁢en
= h-1λ⁢∑n=1∞λn⁢⟨h,en⟩⁢en+1λ⁢∑n=1∞λnλn-λ⁢⟨A⁢h,en⟩⁢en
-∑n=1∞λnλn-λ⁢⟨h,en⟩⁢en
= h-1λ⁢∑n=1∞λn⁢⟨h,en⟩⁢en+1λ⁢∑n=1∞λnλn-λ⁢λn⁢⟨h,en⟩⁢en
-∑n=1∞λnλn-λ⁢⟨h,en⟩⁢en
= h+1λ⁢∑n=1∞-λn⁢(λn-λ)+λn2-λn⁢λλn-λ⁢⟨h,en⟩⁢en
= h,

showing that B=(A-λ⁢idH)-1. ∎

We can start with a function and ask what kind of series it can be expanded into, or we can start with a series and ask what kind of function it defines. The following theorem does the latter. It shows that if en and fn are each orthonormal sequences and λn is a sequence of complex numbers whose limit of 0, then the series

∑n=1∞λn⁢en⊗fn

converges and is an element of ℬ0⁢(H).

Theorem 12.

If H is a Hilbert space, {en:n∈N} is an orthonormal set, {fn:n∈N} is an orthonormal set, and λn∈C is a sequence tending to 0, then the sequence

AN=∑n=1Nλn⁢en⊗fn∈ℬ00⁢(H).

converges to an element of B0⁢(H).

Proof.

Let ϵ>0 and let N0 be such that if n≥N0 then |λn|<ϵ. If N>M≥N0 and h∈H, then, as the en are orthonormal,

∥(AN-AM)⁢h∥2= ∥∑n=M+1Nλn⁢en⊗fn⁢(h)∥2
= ∥∑M+1Nλn⁢⟨h,fn⟩⁢en∥2
= ∑n=M+1N∥λn⁢⟨h,fn⟩⁢en∥2
= ∑n=M+1N|λn|2⁢|⟨h,fn⟩|2
< ϵ2⁢∑n=M+1N|⟨h,fn⟩|2.

By Bessel’s inequality, ∑n=M+1N|⟨h,fn⟩|2≤∥h∥2, and hence

∥(AN-AM)⁢h∥<ϵ⁢∥h∥.

As this holds for all h∈H,

∥AN-AM∥≤ϵ,

showing that AN is a Cauchy sequence, which therefore converges in ℬ⁢(H). As each term in the sequence is finite rank and so compact, the limit is a compact operator. ∎

Continuing the analogy we used with the above theorem, now we start with a function and ask what kind of series it can be expanded into. This is called the singular value decomposition of a compact operator. Helemskii calls the series in the following theorem the Schmidt series of the operator.1515 15 A. Ya. Helemskii, Lectures and Exercises on Functional Analysis, p. 215, Theorem 1. We have already presented the singular value decomposition for finite rank operators in Theorem 8.

Theorem 13 (Singular value decomposition).

If H is a Hilbert space and

A∈ℬ0⁢(H)∖ℬ00⁢(H),

then there is an orthonormal set {en:n∈N} and an orthonormal set {fn:n∈N} such that AN→A, where

AN=∑n=1Nσn⁢(A)⁢en⊗fn.
Proof.

As |A| is self-adjoint and compact, by Theorem 6 there is an orthonormal set {fn:n∈ℕ} such that

|A|=∑n=1∞λn⁢(|A|)⁢fn⊗fn=∑n=1∞σn⁢(A)⁢fn⊗fn.

That is, with |A|N∈ℬ00⁢(H) defined by

|A|N=∑n=1Nσn⁢(A)⁢fn⊗fn,

we have |A|N→|A|.

Let A=U⁢|A| be the polar decomposition of A, and define en=U⁢fn. As U*⁢U⁢|A|=|A| and as σm⁢(A)>0 (because A is not finite rank), we have |A|⁢fmσm⁢(A)=fm, and hence

⟨en,em⟩ = ⟨U⁢fn,U⁢fm⟩
= ⟨fn,U*⁢U⁢fm⟩
= ⟨fn,U*⁢U|A|fmσm⁢(A)⟩
= ⟨fn,|A|⁢fmσm⁢(A)⟩
= ⟨fn,fm⟩
= δn,m,

showing that {en:n∈ℕ} is an orthonormal set. Define

AN=∑n=1Nσn⁢(A)⁢en⊗fn,

and we have AN=U⁢|A|N. As A=U⁢|A| and AN=U⁢|A|N, we get

∥A-AN∥=∥U⁢|A|-U⁢|A|N∥≤∥U∥⁢∥|A|-|A|N∥=∥|A|-|A|N∥→0,

showing that AN→A. ∎

9 Courant min-max theorem

Theorem 14 (Courant min-max theorem).

Let H be an infinite dimensional Hilbert space and let A∈B0⁢(H) be a positive operator. If k∈N then

maxdim⁡S=k⁡minx∈S,∥x∥=1⁡⟨A⁢x,x⟩=λk⁢(A)=σk⁢(A)

and

mindim⁡S=k-1⁡maxx∈S⟂,∥x∥=1⁡⟨A⁢x,x⟩=λk⁢(A)=σk⁢(A).
Proof.

|A| is compact and positive, so according to Theorem 6 there is an orthonormal set {en:n∈ℕ} such that

A=∑n=1∞σn⁢(A)⁢en⊗en.

For k∈ℕ, let Sk=⋁n=k∞{en}. Sk⟂=⋁k=1n-1{en}, so Sk has codimension k-1. (The codimension of a closed subspace of a Hilbert space is the dimension of its orthogonal complement.) If S is a k dimensional subspace of H, then there is some x∈Sk∩S with ∥x∥=1. This is because if V is a closed subspace with codimension k-1 of a Hilbert space and W is a k dimensional subspace of the Hilbert space, then their intersection is a subspace of nonzero dimension. As x∈Sk, there are αn∈ℂ, n≥k, with

x=∑n=k∞αn⁢en,∥x∥2=∑n=k∞|αn|2.

As the sequence σn⁢(A) is nonincreasing,

⟨A⁢x,x⟩ = ⟨∑n=1∞σn⁢(A)⁢⟨x,en⟩⁢en,x⟩
= ⟨∑n=1∞σn⁢(A)⁢⟨∑m=k∞αm⁢em,en⟩⁢en,∑m=k∞αm⁢em⟩
= ⟨∑n=1∞σn⁢(A)⁢∑m=k∞αm⁢δm,n⁢en,∑m=k∞αm⁢em⟩
= ⟨∑n=1∞σn⁢(A)⁢αn⁢χ≥k⁢(n)⁢en,∑m=k∞αm⁢em⟩
= ⟨∑n=k∞σn⁢(A)⁢αn⁢en,∑m=k∞αm⁢em⟩
= ∑n=k∞σn⁢(A)⁢|αn|2
≤ σk⁢(A)⁢∑n=k∞|αn|2
= σk⁢(A),

where we write

χ≥k⁢(n)={1n≥k0n<k.

This shows that if dim⁡S=k then

infx∈S,∥x∥=1⁡⟨A⁢x,x⟩≤σk⁢(A).

Let M=infx∈S,∥x∥=1⁡⟨A⁢x,x⟩, and let xn∈S, ∥xn∥=1, with ⟨A⁢xn,xn⟩→M. As S is a finite dimensional Hilbert space, the unit sphere in it is compact, so there a a subsequence xa⁢(n) that converges to some z∈S, ∥z∥=1. We have

|⟨A⁢z,z⟩-⟨A⁢xn,xn⟩| ≤ |⟨Az,z⟩-⟨Axn,z⟩|+⟨Axn,z⟩-⟨Axn,xn⟩|
= |⟨A(z-xn,z⟩|+|⟨Axn,z-xn⟩|
≤ ∥A∥⁢∥z-xn∥⁢∥z∥+∥A∥⁢∥xn∥⁢∥z-xn∥
= 2⁢∥A∥⁢∥z-xn∥.

As xa⁢(n)→z, we get ⟨A⁢xa⁢(n),xa⁢(n)⟩→⟨A⁢z,z⟩. As ⟨A⁢xn,A⁢xn⟩→M, we get

⟨A⁢z,z⟩=M=infx∈S,∥x∥=1⁡⟨A⁢x,x⟩.

As z∈S and ∥z∥=1, we have in fact

minx∈S,∥x∥=1⁡⟨A⁢x,x⟩=infx∈S,∥x∥=1⁡⟨A⁢x,x⟩≤σk⁢(A).

This is true for any k dimensional subspace of H, so

supdim⁡S=k⁡minx∈S,∥x∥=1⁡⟨A⁢x,x⟩≤σk⁢(A).

If S=Sk+1⟂ then ek∈S, ∥ek∥=1, and

⟨A⁢ek,ek⟩=⟨σk⁢(A)⁢ek,ek⟩=σk⁢(A),

so in fact

maxdim⁡S=k⁡minx∈S,∥x∥=1⁡⟨A⁢x,x⟩=σk⁢(A),

which is the first of the two formulas that we want to prove.

For k≥1, let Sk=⋁n=1k{en}. If S is a k-1 dimensional subspace of H, then S⟂ is a closed subspace with codimension k-1, so the intersection of Sk and S⟂ has nonzero dimension, and so there is some x∈Sk∩S⟂ with ∥x∥=1. As x∈Sk there are α1,…,αk with x=∑n=1kαn⁢en, giving

⟨A⁢x,x⟩ = ⟨∑n=1∞σn⁢(A)⁢⟨x,en⟩⁢en,x⟩
= ⟨∑n=1∞σn⁢(A)⁢⟨∑m=1kαm⁢em,en⟩⁢en,∑m=1kαm⁢em⟩
= ⟨∑n=1∞σn⁢(A)⁢∑m=1kαm⁢δm,n⁢en,∑m=1kαm⁢em⟩
= ⟨∑n=1∞σn⁢(A)⁢αn⁢χ≤k⁢(n)⁢en,∑m=1kαm⁢em⟩
= ⟨∑n=1kσn⁢(A)⁢αn⁢en,∑m=1kαm⁢em⟩
= ∑n=1kσn⁢(A)⁢|αn|2
≥ σk⁢(A)⁢∑n=1k|αn|2
= σk⁢(A).

This shows that

supx∈S⟂,∥x∥=1⁡⟨A⁢x,x⟩≥σk⁢(A).

Define M=supx∈S⟂,∥x∥=1. Because M is a supremum, there is a sequence xn on the unit sphere in Sk-1 such that ⟨A⁢xn,xn⟩→M. The unit sphere in Sk-1 is compact because Sk-1 is finite dimensional, so this sequence has a convergent subsequence xa⁢(n)→z. As

|⟨A⁢z,z⟩-⟨A⁢xn,xn⟩|≤2⁢∥A∥⁢∥z-xn∥

and xa⁢(n)→z, we get

⟨A⁢z,z⟩=M=supx∈S⟂,∥x∥=1⁡⟨A⁢x,x⟩,

whence

maxx∈S⟂,∥x∥=1⁡⟨A⁢x,x⟩=supx∈S⟂,∥x∥=1⁡⟨A⁢x,x⟩≥σk⁢(A).

As this is true for any k-1 dimensional subspace S,

infdim⁡S=k-1⁡maxx∈S⟂,∥x∥=1⁡⟨A⁢x,x⟩≥σk⁢(A).

But for S=Sk-1 we have ek∈S⟂, ∥ek∥=1, and

⟨A⁢ek,ek⟩=⟨σk⁢(A)⁢ek,ek⟩=σk⁢(A),

which implies that

mindim⁡S=k-1⁡maxx∈S⟂,∥x∥=1⁡⟨A⁢x,x⟩=σk⁢(A).

∎

If A∈ℬ⁢(H) is compact, then the eigenvalues of |A| are equal to the singular values of |A|. Therefore the Courant min-max theorem gives expressions for the singular values of a compact linear operator on a Hilbert space, whether or not the operator is itself self-adjoint.

Allahverdiev’s theorem1616 16 I. C. Gohberg and M. G. Krein, Introduction to the Theory of Linear Nonselfadjoint Operators in Hilbert Space, p. 28, Theorem 2.1; cf. J. R. Retherford, Hilbert Space: Compact Operators and the Trace Theorem, p. 75 and p. 106. gives an expression for the singular values of a compact operator that does not involve orthonormal sets, unlike Courant’s min-max theorem. Thus this formula makes sense for a compact operator from one Banach space to another.

Theorem 15 (Allahverdiev’s theorem).

Let H be a Hilbert space and let Fn⁢(H) be the set of bounded finite rank operators of rank ≤n. If A∈B0⁢(H) and n∈N, then

σn⁢(A)=infT∈ℱn-1⁡∥A-T∥.

10 Schatten class operators

If 1≤p<∞ and A∈ℬ0⁢(H), we define

∥A∥p=(∑n∈ℕσn⁢(A)p)1/p,

and define ℬp⁢(H) to be those A∈ℬ0⁢(H) with ∥A∥p<∞. In other words, an element of ℬp⁢(H) is a compact operator whose sequence of singular values is an element of ℓp. We call an element of ℬp⁢(H) a Schatten class operator. We call elements of ℬ1⁢(H) trace class operators and elements of ℬ2⁢(H) Hilbert-Schmidt operators.

If A∈ℬ0⁢(H) is positive, then, according to Theorem 6, there is an orthonormal set {en:n∈ℕ} such that

A=∑n∈ℕλn⁢(A)⁢en⊗en,

where the series converges in the strong operator topology. As the en are orthonormal, we have

Ap=∑n∈ℕλn⁢(A)p⁢en⊗en,

which is itself a positive compact operator, and thus σn⁢(Ap)=σn⁢(A)p for n∈ℕ. Therefore, if A is a positive compact operator, then ∥A∥p=∥Ap∥11/p.

If A∈ℬ0⁢(H) and n∈ℕ, then σn⁢(|A|)=σn⁢(A) and σn⁢(A*)=σn⁢(A). Hence, if 1≤p<∞ then

∥|A|∥p=∥A∥p,∥A*∥p=∥A∥p.

As |A| is compact and self-adjoint, it has an eigenvalue with absolute value ∥A∥, from which it follows that if 1≤p<∞ then ∥A∥≤∥A∥p.

Theorem 16.

If A∈B1⁢(H), B∈B⁢(H), and k∈N, then

σk⁢(B⁢A)≤∥B∥⁢σk⁢(A).
Proof.

For all x∈H,

⟨(B⁢A)*⁢B⁢A⁢x,x⟩ = ⟨B⁢A⁢x,B⁢A⁢x⟩
= ∥B⁢A⁢x∥2
≤ ∥B∥2⁢∥A⁢x∥2
= ∥B∥2⁢⟨A⁢x,A⁢x⟩
= ∥B∥2⁢⟨A*⁢A⁢x,x⟩.

Applying the Courant min-max theorem to the positive operators (B⁢A)*⁢B⁢A and A*⁢A, if k∈ℕ then

σk⁢((B⁢A)*⁢B⁢A) = maxdim⁡S=k⁡minx∈S,∥x∥=1⁡⟨(B⁢A)*⁢B⁢A⁢x,x⟩
≤ ∥B∥2⁢maxdim⁡S=k⁡minx∈S,∥x∥=1⁡⟨A*⁢A⁢x,x⟩
= ∥B∥2⁢σk⁢(A*⁢A).

But

σk⁢((B⁢A)*⁢B⁢A)=σk⁢(|B⁢A|2)=σk⁢(|B⁢A|)2=σk⁢(B⁢A)2

and

σk⁢(A*⁢A)=σk⁢(|A|2)=σk⁢(|A|)2=σk⁢(A)2,

so taking the square root,

σk⁢(B⁢A)≤∥B∥⁢σk⁢(A).

∎

Using Theorem 16, if 1≤p<∞ then

∥B⁢A∥p=(∑n∈ℕσn⁢(B⁢A)p)1/p≤(∑n∈ℕ∥B∥p⁢σk⁢(A)p)1/p=∥B∥⁢∥A∥p.

The following theorem states that the Schatten class operators are Banach spaces.1717 17 Gert K. Pedersen, Analysis Now, revised printing, p. 124, E 3.4.4

Theorem 17.

If 1≤p<∞, then Bp⁢(H) is a Banach space with the norm ∥⋅∥p.

11 Weyl’s inequality

Weyl’s inequality relates the eigenvalues of a self-adjoint compact operator with its singular values.1818 18 Peter D. Lax, Functional Analysis, p. 336, chapter 30, Lemma 7. We use the notation from Definition 5. For N>ν⁢(A) the left hand side is equal to 0 so the inequality is certainly true then.

Theorem 18 (Weyl’s inequality).

If A∈B0⁢(H) is self-adjoint and N≤ν⁢(A), then

∏n=1N|λn⁢(A)|≤∏n=1Nσn⁢(A).
Proof.

Let

EN=⋁n=1Nker⁡(A-λn⁢(A)⁢idH),

which is finite dimensional. Check that EN is an invariant subspace of A, and let AN:EN→EN be the restriction of A to EN. AN is a positive operator. As EN is spanned by eigenvectors for nonzero eigenvalues of A it follows that ker⁡AN={0}, and as EN is finite dimensional, we get that AN is invertible. If AN has polar decomposition AN=UN⁢|AN|, then UN is invertible; if a partial isometry is invertible then it is unitary, so UN is unitary, and therefore the eigenvalues of UN all have absolute value 1. As the determinant of a linear operator on a finite dimensional vector space is the product of its eigenvalues counting algebraic multiplicity,

det⁡|AN|=1|det⁡UN|⁢|det⁡AN|=|det⁡AN|=∏n=1N|λn⁢(A)|. (1)

Let PN be the orthogonal projection onto EN. If v∈EN, then A⁢PN⁢v=AN⁢v, and if v∈EN⟂ then A⁢PN⁢v=A⁢(0)=0. We get that

|A⁢PN|⁢v={|AN|⁢vv∈EN0v∈EN⟂,

and it follows that if 1≤n≤N then σn⁢(AN)=σn⁢(A⁢PN). Using Theorem 16 we get

σn⁢(A⁢PN)≤∥P∥⁢σn⁢(A)≤σn⁢(A);

the second inequality is an equality unless PN=0. We have shown that if 1≤n≤N then σn⁢(AN)≤σn⁢(A), and combining this with (1) gives us

∏n=1N|λn⁢(A)|=det⁡|AN|=∏n=1Nσn⁢(AN)≤∏n=1Nσn⁢(A).

∎

Theorem 19.

If 0<p<∞, A∈B0⁢(H) is self-adjoint, and N∈N, then

∑n=1N|λn⁢(A)|p≤∑n=1Nσn⁢(A)p.
Proof.

Schur’s majorization inequality1919 19 Peter D. Lax, Functional Analysis, p. 337, chapter 30, Lemma 8; cf. J. Michael Steele, The Cauchy-Schwarz Master Class, p. 201, Problem 13.4. states that if a1≥a2≥⋯ and b1≥b2≥⋯ are nonincreasing sequences of real numbers satisfying, for each N∈ℕ,

∑n=1Nan≤∑n=1Nbn,

and ϕ:ℝ→ℝ is a convex function with limx→-∞⁡ϕ⁢(x)=0, then for every N∈ℕ,

∑n=1Nϕ⁢(an)≤∑n=1Nϕ⁢(bn).

With the hypotheses of Theorem 18, for 1≤n≤ν⁢(A), define an=log⁡|λn⁢(A)| and bn=log⁡σn⁢(A) and let ϕ⁢(x)=ep⁢x. By Theorem 18 these satisfy the conditions of Schur’s majorization inequality, which then gives us for 1≤N≤ν⁢(A) that

∑n=1N|λn⁢(A)|p≤∑n=1Nσn⁢(A)p.

If n>ν⁢(A) then λn⁢(A)=0. ∎

12 Rayleigh quotients for self-adjoint operators

If A∈ℬ⁢(H) is self-adjoint, we define the Rayleigh quotient of A by

f⁢(x)=⟨A⁢x,x⟩⟨x,x⟩,x∈H,x≠0,f:H∖{0}→ℝ.

Let X and Y be normed spaces, U an open subset of X, and f:U→Y a function. If x∈U and there is some T∈ℬ⁢(X,Y) such that

limh→0⁡∥f⁢(x+h)-f⁢(x)-T⁢h∥∥h∥=0, (2)

then f is said to be Fréchet differentiable at x, and T is called the Fréchet derivative of f at x;2020 20 Ward Cheney, Analysis for Applied Mathematics, p. 149. it does not take long to prove that if T1,T2∈ℬ⁢(X,Y) both satisfy (2) then T1=T2. We denote the Fréchet derivative of f at x by (D⁢f)⁢x. D⁢f is a map from the set of all points at which f is Fréchet differentiable to ℬ⁢(X,Y).

To say that x is a stationary point of f is to say that f is Fréchet differentiable at x and that the Fréchet derivative of f at x is the zero map. One proves that if T1,T2 are Fréchet derivatives of f at x then T1=T2, and thus speak about the Fréchet derivative of f at x

Theorem 20.

If A∈B⁢(H) is self-adjoint, then each eigenvector of A is a stationary point of the Rayleigh quotient of A.

Proof.

If λ is an eigenvalue of A then, as A is self-adjoint, λ∈ℝ. Let v≠0 satisfy A⁢v=λ⁢v. We have

f⁢(v)=⟨A⁢v,v⟩⟨v,v⟩=⟨λ⁢v,v⟩⟨v,v⟩=λ.

For h≠0, using that A is self-adjoint and that λ∈ℝ,

|f⁢(v+h)-f⁢(v)-0⁢v|∥h∥ = 1∥h∥⋅|⟨A⁢(v+h),v+h⟩⟨v+h,v+h⟩-λ|
= 1∥h∥⁢∥v+h∥2⁢|⟨A⁢(v+h),v+h⟩-λ⁢⟨v+h,v+h⟩|
= 1∥h∥⁢∥v+h∥2|⟨A⁢v,v⟩+⟨A⁢v,h⟩+⟨A⁢h,v⟩+⟨A⁢h,h⟩
-λ⟨v,v⟩-λ⟨v,h⟩-λ⟨h,v⟩-λ⟨h,h⟩|
= 1∥h∥⁢∥v+h∥2|⟨λ⁢v,v⟩+⟨λ⁢v,h⟩+⟨h,λ⁢v⟩+⟨A⁢h,h⟩
-λ⟨v,v⟩-λ⟨v,h⟩-λ⟨h,v⟩-λ⟨h,h⟩|
= 1∥h∥⁢∥v+h∥2⁢|⟨A⁢h,h⟩-λ⁢⟨h,h⟩|
= 1∥h∥⁢∥v+h∥2⁢|⟨A⁢h-λ⁢h,h⟩|.

Therefore

|f⁢(v+h)-f⁢(v)-0⁢v|∥h∥ ≤ ∥A⁢h-λ⁢h∥⁢∥h∥∥h∥⁢∥v+h∥2
= ∥(A-λ⁢idH)⁢h∥∥v+h∥2
= ∥A-λ⁢idH∥⁢∥h∥∥v+h∥2.

As h→0 the right-hand side tends to 0 (one of the terms tends to 0, one doesn’t depend on h, and the denominator is bounded below in terms just of v for sufficiently small h), showing that 0 is the Fréchet derivative of f at v. ∎