In another post I had to find the eigenvalues and eigenvectors of the following matrix,
Eq.1
where is an N-dimensional vector with elements all 1, whilst
is the
identity matrix.
In that original post I said that diagonalizing this matrix (finding its eigenvalues and eigenvectors) was easy. A colleague asked me why I thought it was easy. My answer, “well because it is the only matrix form whose eigen-decomposition I explicitly remember”. And there’s a reason why I remember how to decompose by hand these types of matrices,
- It is easy to do.
- These types of matrices are very common.
Okay, I know that reason 1 is a circular argument, but it’s true. These types of matrices are easy to diagonalize and you’ll agree when you learn the trick to calculate their eigenvalues and eigenvectors. And like all the posts in my Data Science Notes series, I’m basing this post on experience. It is nearly 29 years since I first learnt this trick, and I have used it many time since. The reason I find it so useful is because of reason 2 above.
Reason 2 is equally important. As a Data Scientist you will encounter this form of matrix often, and so there is an advantage to knowing how to calculate the eigenvectors and eigenvalues of a matrix such as in Eq.1, and you learn a lot about eigenvectors and eigenvalues in doing so; you learn what they are, what they represent, how to manipulate them, and how to use them to do useful things. We can view the matrix
in Eq.1 as a low-rank perturbation of the identity matrix, and manipulating matrices by adding a low-rank perturbation is a common trick in the mathematical and statistical sciences, most recently used in LoRA (low-rank adaptation) algorithms for fine-tuning Large Language Models (LLMs). So getting comfortable with the mathematics of eigen-decompositions of low-rank perturbations will pay off.
Let’s begin. To start we’ll need to do a quick recap of what an eigenvector is and what an eigenvalue is.
Recap of what eigenvectors and eigenvalues are
An eigenvector of a square matrix
is a vector that satisfies the following equation,
Eq.2
The scalar quantity is the eigenvalue corresponding to the eigenvector
. Given a vector
that satisfies Eq.2, then the vector
also satisfies Eq.2 for any choice of the scalar
. This means we have the freedom to set the norm of the vector
since we can effectively absorb the effect of any choice of norm into the effective value of
. So without loss of generality we can set
.
An matrix will have
different eigenvectors, some of which may have the same eigenvalue. So overall we have a set
of eigenvectors with corresponding eigenvalues
. Eigenvectors corresponding to eigenvalues that have different values are orthogonal to each other. Eigenvectors corresponding to the same eigenvalue need not be orthogonal to each other, but we can always construct an orthonormal set that spans the same subspace, so without loss of generality we can take the set of eigenvectors
to form an orthonormal set.
The eigen-decomposition of a matrix
We can see from Eq.2 that the effect of the matrix on an eigenvector
is simply to multiply it by a scale factor
. Consequently, the eigenvectors of a matrix are very special. In fact they allow us to characterize or represent the matrix
in the following way,
Eq.3
And once we have all the eigenvectors and corresponding eigenvalues
of a matrix
we can do useful things such as calculate its inverse or calculate its determinant, which occur in multiple places in multi-variate statistics, machine learning, and GenAI. I also needed them for the calculations in my original blog-post . The inverse and determinant of
are given by the usual formulae,
Eq.4
Eq.5
Identifying our eigenvectors
Now, our goal is the work out the eigenvalues and eigenvectors of the matrix,
Eq.6
In Eq.6 I have generalized from the matrix form in Eq.1 a little here, so we have vectors
, which form an orthonormal set, i.e. they are all of unit length and orthogonal to each other.
Let’s write the right-hand-side of Eq. 6 as a sum of two matrices, i.e. we write, in Eq.6 as,
Eq. 7,
with the matrix defined as,
Eq.8
This leads us to our first trick. If we have a vector which is an eigenvector of
with eigenvalue
, then we can easily work out the effect of
on
because the identity matrix
leaves
unchanged. So overall we have,
Eq. 9
Eq.9 tells us that eigenvectors of are eigenvectors of
. We only need to find eigenvectors and eigenvalues of
.
Now we introduce our second trick. We look at the definition of in Eq. 8 and compare it to the eigen-decomposition in Eq.3. They are of the same form. Matrix
is already in the form an eigen-decomposition. The vectors
are the eigenvectors
. And the corresponding eigenvalues are
.
Ok, but we only have vectors
and we need
eigenvectors. Easy, consider any vector,
that is orthogonal to all the vectors
. The effect of
on that vector
is as follows,
Eq.10
So any vector orthogonal to all the vectors will be an eigenvector of
with an eigenvalue of zero.
Now we can combine our first and second tricks together the finally work out the eigenvectors and eigenvalues of . These are,
Eigenvalue = , Eigenvector =
Eigenvalue = , Eigenvector =
……
Eigenvalue = , Eigenvector =
Eigenvalue = , Eigenvectors = Any orthonormal set that spans the subspace orthogonal to the vectors
.
Putting it all together
Let’s put this into action. Let’s go back to our original matrix in Eq.1 at the start of this post. That matrix was of the form,
Eq.11
So comparing Eq.11 to Eq.6 we have ,
,
, and
. So our eigenvectors and corresponding eigenvalues are,
Eigenvalue 1 = , Eigenvector 1 =
Eigenvalues 2 to =
, Eigenvectors 2 to
= an orthonormal set orthogonal to
.
We can now easily calculate the determinant of our matrix in Eq.1 and we get,
Eq.12
We can also easily calculate the inverse of the matrix in Eq.1. In this case we get,
Eq.13
In Eq.13 the vectors form an orthonormal set orthogonal to
.
Now, we can use one final little trick to simplify the right-hand-side of Eq.13. If we recall that the effect of the identity matrix on any vector
from an orthonormal set
is to leave those vectors unchanged, then this means we can write the identity matrix
in terms of an eigen-decomposition, and that decomposition is as follows,
Eq.14
Now Eq.14 is almost what we have in the first term on the right-hand-side of Eq.13. Our sum in Eq.13 runs from to
, whilst the sum in Eq.14 runs from
to
. If only we had another vector
that was orthogonal to
etc.? But we do! We have
. In fact, that is how we defined
etc. – they are all orthogonal to
. This means if we set
, we can re-write the sum in the right-hand-side of Eq.13 as,
Eq.15
where the vectors etc. form an orthonormal set in
. This means (by comparing to Eq.14) we can also write the last expression in Eq.15 as,
Eq.16
Combining this with the final term on the right-hand-side of Eq.13 we get our final simplified expression for , which is,
Eq.17
That’s it. If you want a challenge to test your understanding, then using the ideas and tricks we introduced above see if you can construct a closed form expression for the inverse of the matrix in Eq.6, not just the special case of the matrix in Eq.1.
Other low-rank perturbation tricks
In my original post I wanted to find the eigenvectors and eigenvalues of the matrix in Eq.1. I’ve also illustrated that once we have closed-form expressions for the eigenvectors and eigenvalues we can calculate the determinant and inverse of in Eq.1 very easily. But it is also worth pointing out that there are a couple of other tricks for evaluating the determinant and inverse of matrices of the form in Eq.1 and Eq.6, namely,
- Sylvester’s determinant theorem. The Wikipedia link here is for the Weinstein–Aronszajn identity as it turns out that the deteminant theorem has been mis-attributed to Sylvester – I was taught it as Sylvester’s theorem in university. This allows us to calculate the determinant of a low-rank pertubation in terms of a determinant of a smaller matrix.
- The Woodbury formula. This allows us to express the inverse of a low-rank perturbation in terms of matrix multiplications involving the inverse of the unperturbed matrix. In fact the result in Eq.17 can also be obtained by using the Shermann-Morrison formula, which is a special case of the Woodbury formula.
Sylvester’s theorem and the Woodbury formula only give us information about the determinant and inverse of the low-rank perturbation, so if we explicitly want information about the eigenvectors and eigenvalues we will probably have to use the tricks I have outlined in this post. However, the advantage of the Sylvester and Woodbury formulae is that they can be used when we have a base-case (the unperturbed matrix) that isn’t necessarily the identity matrix, i.e. they are more widely applicable than the tricks I’ve outlined in this post. What the Sylvester and Woodbury formulae also illustrate is that we shouldn’t be frightened of working with low-rank perturbations of base-case matrices. Once you learn a few tricks working with them becomes very easy.
What have we learnt?
- How to calculate analytically (by hand) the eigenvectors and eigenvalues of matrices of the form in Eq.6.
- But more importantly, we’ve learnt along the way that eigenvectors and eigenvalues are not that scary. They are easy to manipulate and work with once you get comfortable with a few concepts that you have probably already learnt at college or at university. Matrices of the form in Eq.6 are the simplest, realistic matrices which bring those concepts to life. These matrices are a great learning ground.
- Even better, matrices of the type in Eq.6 occur a lot in Data Science, Machine Learning, and AI. In fact, it is probably one of Data Science’s secrets; 95% or more of Data Scientists, myself included, stick to using matrices of this form or a low-rank perturbation of some other known base-case matrix when modelling a real problem. And why not!? When modelling a real problem we start with the base-case we understand and change it as little as possible, both for model complexity reasons (number of extra parameters etc.) and for reasons of computational complexity (can we easily compute and understand the properties of the new matrix). So if you get comfortable working with eigenvectors and eigenvalues of low-rank perturbations of simple base-case matrices, you’ve cracked it for 95% of Data Science use-cases.
© 2026 David Hoyle. All Rights Reserved