An equation showing the mathematical form of the log sum exp calculation.

Data Science Notes: 2. Log-Sum-Exp

On January 1, 2026January 31, 2026 By dchoyleIn Algorithms, Coding, Data Science, Data Science Notes, Numerical Analysis, pythonLeave a comment

Summary

Summing up many probabilities that are on very different scales often involves calculation of quantities of the form $\log\left ( \sum_{k} \exp\left( a_{k}\right )\right )$ . This calculation is called log-sum-exp.
Calcuating log-sum-exp the naive way can lead to numerical instabilities. The solution to this numerical problem is the “log-sum-exp” trick.
The scipy.special.logsumexp function provides a very useful implementation of the log-sum-exp trick.
The log-sum-exp function also has uses in machine learning, as it is a smooth, differentiable approximation to the ${\rm max}$ function.

Introduction

This is the second in my series of Data Science Notes series. The first on Bland-Altman plots can be found here. This post is on a very simple numerical trick that ensures accuracy when adding lots of probability contributions together. The trick is so simple that implementations of it exist in standard Python packages, so you only need to call the appropriate function. However, you still need to understand why you can’t just naively code-up the calculation yourself, and why you need to use the numerical trick. As with the Bland-Altman plots, this is something I’ve had to explain to another Data Scientist in the last year.

The log-sum-exp trick

Sometimes you’ll need to calculate a sum of the form, $\sum_{k} \exp\left ( a_{k}\right )$ , where you have values for the $a_{k}$ . Really? Will you? Yes, it will probably be calculating a log-likelihood, or a contribution to a log-likelihood, so the actual calculation you want to do is of the form,

$\log\left ( \sum_{k} \exp \left ( a_{k}\right )\right )$

These sorts of calculations arise where you have log-likelihood or log-probability values $a_{k}$ for individual parts of an overall likelihood calculation. If you come from a physics calculation you’ll also recognise the expression above as the calculation of a log-partition function.

So what we need to do is exponentiate the $a_{k}$ values, sum them, and then take the log at the end. Hence the expression “log-sum-exp”. But why a blogpost on “log-sum-exp”? Surely, it’s an easy calculation. It’s just np.log(np.sum(np.exp(a))) , right ? Not quite.

It depends on the relative values of the $a_{k}$ . Do the naïve calculation np.log(np.sum(np.exp(a))) and it can be horribly inaccurate. Why? Because of overflow and underflow errors.

If we have two values $a_{1}$ and $a_{2}$ and $a_{1}$ is much bigger than $a_{2}$ , when we add $\exp(a_{1})$ to $\exp(a_{2})$ we are using floating point arithmetic to try and add a very large number to a much smaller number. Most likely we will get an overflow error. If would be much better if we’d started with $\exp(a_{1})$ and try to add $\exp(a_{2})$ to it. In fact, we could pre-compute $a_{2} - a_{1}$ , which would be very negative and from this we could easily infer that adding $\exp(a_{2})$ to $\exp(a_{1})$ would make very little difference. In fact, we could just approximate $\exp(a_{1}) + \exp(a_{2})$ by $\exp(a_{1})$ .

But how negative does $a_{2} - a_{1}$ have to be before we ignore the addition of $\exp(a_{2})$ ? We can set some pre-specified threshold, chosen to avoid overflow or underflow errors. From this, we can see how to construct a little Python function that takes an array of values $a_{1}, a_{2},\ldots, a_{N}$ and computes an accurate approximation to $\sum_{k=1}^{N}\exp(a_{k})$ without encountering underflow or overflow errors.

In fact we can go further and approximate the whole sum by first of all identifying the maximum value in an array $a = [a_{1}, a_{2}, \ldots, a_{N}]$ . Let’s say, without loss of generality, the maximum value is $a_{1}$ . We could ensure this by first sorting the array, but it isn’t necessary actually do this to get the log-sum-exp trick to work. We can then subtract $a_{1}$ from all the other values of the array, and we get the result,

$\log \left ( \sum_{k=1}^{N} \exp \left ( a_{k} \right )\right )\;=\; a_{1} + \log \left ( 1\;+\;\sum_{k=2}^{N} \exp \left ( a_{k} - a_{1} \right )\right )$

The values $a_{k} - a_{1}$ are all negative for $k \ge 2$ , so we can easily approximate the logarithm on the right-hand side of the equation by a suitable expansion of $\log (1 + x)$ . This is the “log-sum-exp” trick.

The great news is that this “log-sum-exp” calculation is so common in different scientific fields that there are already Python functions written to do this for us. There is a very convenient “log-sum-exp” function in the SciPy package, which I’ll demonstrate in a moment.

The log-sum-exp function

The sharp-eyed amongst you may have noticed that the last formula above gives us a way of providing upper and lower bounds for the ${\rm max}$ function. We can simply re-arrange the last equation to get,

${\rm max} \left ( a_{1}, a_{2},\ldots, a_{N} \right ) \;\le\; \log \left ( \sum_{k=1}^{N}\exp \left ( a_{k}\right )\right )$

The logarithm calculation on the right-hand side of the inequality above is what we call the log-sum-exp function (lse for short). So we have,

${\rm max} \left ( a_{1}, a_{2},\ldots, a_{N} \right ) \;\le\; {\rm lse}\left ( a_{1}, a_{2}, \ldots, a_{N}\right )$

This gives us an upper bound for the ${\rm max}$ function. Since $a_{k} \le {\rm max}\left ( a_{1}, a_{2},\ldots,a_{N}\right )$ , it is also relatively easy to show that,

${\rm lse} \left ( a_{1}, a_{2}, \ldots, a_{N}\right )\;\le\; {\rm max}\left ( a_{1}, a_{2},\ldots, a_{N} \right )\;+\;\log N$

and so we have have a lower bound for the ${\rm max}$ function. So the log-sum-exp function allows us to compute lower and upper bounds for the maximum of an array of real values, and it can provide an approximation of the maximum function. The advantage is that the log-sum-exp function is smooth and differentiable. In contrast, the maximum function itself is not smooth nor differentiable everywhere, and so is less convenient to work with mathematically. For this reason the log-sum-exp is often called the “real-soft-max” function because it is a “soft” version of the maximum function. It is often used in machine learning settings to replace a maximum calculation.

Calculating log-sum-exp in Python

So how do we calculate the log-sum-exp function in Python. As I said, we can use the SciPy implementation which is in scipy.special. All we need to do is pass an array-like set of values $a_{k}$ . I’ve given a simple example below,

			
# import the packages and functions we need
import numpy as np
from scipy.special import logsumexp
# create the array of a_k values 
a = np.array([70.0, 68.9, 20.3, 72.9, 40.0])
# Calculate log-sum-exp using the scipy function 
lse = logsumexp(a)
# look at the result
print(lse)

		

This will give the result 72.9707742189605

The example above and several more can be found in the Jupyter notebook DataScienceNotes2_LogSumExp.ipynb in the GitHub repository https://github.com/dchoyle/datascience_notes

The great thing about the SciPy implementation of log-sum-exp is that it allows us to include signed scale factors, i.e. we can compute,

$\log \left ( \sum_{k=1}^{N} b_{k}\exp\left ( a_{k}\right ) \right )$

where the values $b_{k}$ are allowed to be negative. This means, that when we are using the SciPy log-sum-exp function to perform the log-sum-exp trick, we can actually use it to calculate numerically stable estimates of sums of the form,

$\log \left ( \exp\left ( a_{1}\right ) \;-\; \exp\left ( a_{2}\right ) \;-\;\exp\left ( a_{3}\right )\;+\;\exp\left ( a_{4}\right )\;+\ldots\; + \exp\left ( a_{N}\right )\right )$ .

Here’a small code snippet illustrating the use of the scipy.special.logsumexp with signed contributions,

			
# Create the array of the a_k values
a = np.array([10.0, 9.99999, 1.2])
b = np.array([1.0, -1.0, 1.0])
# Use the scipy.special log-sum-exp function
lse = logsumexp(a=a, b=b)
# Look at the result
print(lse)

		

This will give the result 1.2642342014146895.

If you look at the output of the example above you’ll see that the final result is much closer to the value of the last array element $a_{3} = 1.2$ . This is because the first two contributions, $\exp(a_{1})$ and $\exp(a_{2})$ almost cancel each other out because the contribution $\exp(a_{2})$ is pre-fixed by a factor of -1. What we get left with is something close to $\log(\exp(a_{3}))\;=\; a_{3}$ .

There is also a small subtlety in using the SciPy logsumexp function with signed contributions. If the substraction of some terms had led to an overall negative result, scipy.special.logsumexp will rerturn NaN as the result. In order to get it to always return a result for us, we have to tell it to return the sign of the final summation as well, by setting the return_sign argument of the function to True. Again, you can find the code example above and others in the notebook DataScienceNotes2_LogSumExp.ipynb in the GitHub repository https://github.com/dchoyle/datascience_notes.

When you are having to combine lots of different probabilities, that are on very different scales, and you need to subtract some of them and add others, the SciPy log-sum-exp function is very very useful.

Extreme Components Analysis

On June 18, 2025August 4, 2025 By dchoyleIn UncategorizedLeave a comment

TL;DR: Both Principal Components Analysis (PCA) and Minor Components Analysis (MCA) can be used for dimensionality reduction, identifying low-dimensional subspaces of interest as those which have the greatest variation in the original data (PCA), or those which have the least variation in the origina data MCA). As real data will contain both directions of unusually high variance and directions of unusually low variance, using just PCA or just MCA will lead to biased estimates of the low-dimensional subspace. The 2003 NeurIPs paper from Welling et al unifies PCA and MCA into a single probabilistic model XCA (Extreme Components Analysis). This post explains the XCA paper of Welling et al and demonstrates the XCA algorithm using simulated data. Code for the demonstration is available from https://github.com/dchoyle/xca_post

A deadline

This post arose because of a deadline I have to meet. I don’t know when the deadline is, I just know there is a deadline. Okay, it is a self-imposed deadline, but it will start to become embarrassing if I don’t hit it.

I was chatting with a connection, Will Faithfull, at a PyData Manchester Leaders meeting almost a year ago. I mentioned that one of my areas of expertise was Principal Components Analysis (PCA), or more specifically, the use of Random Matrix Theory to study the behaviour of PCA when applied to high-dimensional data.

A recap of PCA

In PCA we are trying to approximate a d-dimensional dataset by a reduced number of dimensions $k < d$ . Obviously we want to retain as much of the structure and variation of the original data, so we choose our k-dimensional subspace such that the variance of the original data in the subspace is as high as possible. Given a mean-centered data matrix $\underline{\underline{X}}$ consisting of $N$ observations, we can calculate the sample covariance matrix $\hat{\underline{\underline{C}}}$ as ¹,

$\hat{\underline{\underline{C}}} = \frac{1}{N-1} \underline{\underline{X}}^{\top} \underline{\underline{X}}$

Once we have the (symmetric) matrix $\hat{\underline{\underline{C}}}$ we can easily compute its eigenvectors $\underline{v}_{i}, i=1,\ldots, d$ , and their corresponding eigenvalues $\lambda_{i}$ .

The optimal $k$ -dimensional PCA subspace is then spanned by the $k$ eigenvectors of $\hat{\underline{\underline{C}}}$ that correspond to the $k$ largest eigenvalues of $\hat{\underline{\underline{C}}}$ . These eigenvectors are the directions of greatest variance in the original data. Alternatively, one can just do a Singular Value Decomposition (SVD) of the original data matrix $\underline{\underline{X}}$ , and work with the singular values of $\underline{\underline{X}}$ instead of the eigenvalues of $\hat{\underline{\underline{C}}}$ .

That is a heuristic derivation/justification of PCA (minus the detailed maths) that goes back to Harold Hotelling in 1933². There is a probabilistic model-based derivation due to Tipping and Bishop (1999), which we will return to later.

MCA

Will responded that as part of his PhD, he’d worked on a problem where he was more interested in the directions in the dataset along which the variation is least. The problem Will was working on was “unsupervised change detection in multivariate streaming data”. The solution Will developed was a modular one, chaining together several univariate change detection methods each of which monitored a single feature of the input space. This was combined with a MCA feature extraction and selection pre-processing step. The solution was tested against a problem of unsupervised endogenous eye blink detection.

The idea behind Will’s use of MCA was that for the streaming data he was interested in it was likely that the inter-class variances of various features were likely to be much smaller than intra-class variances, and so any principal components were likely to be dominated by what the classes had in common rather than what had changed, so the directions of greatest variance weren’t very useful for his change detection algorithm.

I’ve put a link here to Will’s PhD in case you are interested in the details of the problem and solution – yes, Will I have read your PhD.

Directions of least variance in a dataset can be found from the same eigen-decomposition of the sample covariance matrix and by selecting the components with the smallest non-zero eigenvalues. Unsurprisingly, focusing on directions of least variance in a dataset is called Minor Components Analysis (MCA)^3,4. Where we have the least variation in the data the data is effectively constrained so, MCA is good for identifying/modelling invariants or constraints within a dataset.

At this point in the conversation, I recalled the last time I’d thought about MCA. That was when an academic colleague and I had a paper accepted at the NeurIPs conference in 2003. Our paper was on kernel PCA applied to high-dimensional data, in particular the eigenvalue distributions that result. As I was moving job and house at the time I was unable to go the conference, so my co-author, Magnus Rattray (now Director of the Institute for Data Science and Artificial Intelligence at the University of Manchester), went instead. On returning, Magnus told me of an interesting conversation he’d had at the conference with Max Welling about our paper. Max also had a paper at the conference, on XCA – Extreme Components Analysis. Max and his collaborators had managed to unify PCA and MCA into a single framework.

I mentioned the XCA paper to Will at the PyData Manchester Leaders meeting and said I’d write something up explaining XCA. It would also give me an excuse to revisit something that I hadn’t looked at since 2003. That conversation with Will was nearly a year ago. Another PyData Manchester Leaders meeting came and went and another will be coming around sometime soon. To avoid having to give a lame apology I thought it was about time I wrote this post.

XCA

Welling et al rightly point out that if we are modelling a dataset as lying in some reduced dimensionality subspace then we consider the data as being a combination of variation and constraint. We have variation of the data within a subspace and a constraint that the data does not fall outside the subspace. So we can model the same dataset focusing either on the variation (PCA) or on the constraints (MCA).

Note that in my blog post I have used a different, more commonly used notation. for the number of features and the number of components, than that used in the Welling et al paper. The mapping between the two notations is given below,

Number of features: My notation = $d$ , Welling et al notation = $D$
Number of components: My notation = $k$ , Welling et al notation = $d$

Probabilistic PCA and MCA

PCA and MCA both have probabilistic formulations, PPCA and PMCA⁵ respectively. Welling et al state that, “probabilistic PCA can be interpreted as a low variance data cloud which has been stretched in certain directions. Probabilistic MCA on the other hand can be thought of as a large variance data cloud which has been pushed inward in certain directions.” In both probabilistic models a $d$ -dimensional datapoint $\underline{x}$ is considered as coming from a zero-mean multivariate Gaussian distribution. In PCA the covariance matrix of the Gaussian is modelled as,

$\underline{\underline{C}}_{PCA} = \sigma^{2}_{0}\underline{\underline{I}}_{d} + \underline{\underline{A}}\,\underline{\underline{A}}^{\top}$

The matrix $\underline{\underline{A}}$ is $k \times d$ and its columns are the principal components that span the low dimensional subspace we are trying to model.

In MCA the covariance matrix is modelled as,

$\underline{\underline{C}}_{MCA}^{-1} = \sigma^{-2}_{0}\underline{\underline{I}}_{d} + \underline{\underline{W}}^{\top}\,\underline{\underline{W}}$

The matrix $\underline{\underline{W}}$ is $d \times (d-k)$ and its rows are the minor components that define the $d-k$ subspace where we want as little variation as possible.

Since in real data we probably have both exceptional directions whose variance is greater than the bulk and exceptional directions whose variance is less than the bulk, both PCA and MCA would lead to biased estimates for these datasets. The problem is that if we use PCA we lump the low variation eigenvalues (minor components) in with our estimate of the isotropic noise, thereby underestimating the true noise variance and consequently biasing our estimate of the large variation PC subspace. Likewise, if we use MCA, we lump all the large variation eigenvalues (principal components) into our estimate of the noise and overestimate the true noise variance, thereby biasing our estimate of the low variation MC subspace.

Probabilistic XCA

In XCA we don’t have that problem. In XCA we include both large variation and small variation directions in our reduced dimensionality subspace. In fact we just have a set of orthogonal directions $\underline{a}_{i}\;,\;i=1,\ldots,k$ that span a low-dimensional subspace and again form the columns of a matrix $\underline{\underline{A}}$ . These are our directions of interest in the data. Some of them, say, $k_{PC}$ , have unusually large variance, some of them, say $k_{MC}$ , have unusually small variance. The overall number of extreme components (XC) is $k = k_{PC} + k_{MC}$ .

As with probabilistic PCA, we then add on top an isotropic noise component to the overall covariance matrix. However, the clever trick used by Welling et al was that they realized that adding noise always increases variances, and so adding noise to all features will make the minor components undetectable as the minor components have, by definition, variances below that of the bulk noise. To circumvent this, Welling et al only added noise to the subspace orthogonal to the subspace spanned by the vectors $\underline{a}_{i}$ . They do this by introducing a projection operator ${\cal{P}}_{A}^{\perp} = \underline{\underline{I}}_{d} - \underline{\underline{A}} \left ( \underline{\underline{A}}^{\top}\,\underline{\underline{A}}\right )^{-1}\underline{\underline{A}}^{\top}$ . Again we model the data as coming from a zero-mean multivariate Gaussian, but for XCA the final covariance matrix is then of the form,

$\underline{\underline{C}}_{XCA} = \sigma^{2}_{0} {\cal{P}}_{A}^{\perp} + \underline{\underline{A}}\,\underline{\underline{A}}^{\top}$

and the XCA model is,

$\underline{x} \sim \underline{\underline{A}}\,\underline{y} + {\cal{P}}_{A}^{\perp} \underline{n}\;\;\;,\;\; \underline{y} \sim {\cal{N}}\left ( \underline{0}, \underline{\underline{I}}_{k}\right )\;\;\;,\;\;\underline{n} \sim {\cal {N}}\left ( \underline{0}, \sigma^{2}_{0}\underline{\underline{I}}_{d}\right )$

We can also start from the MCA side, by defining a projection operator ${\cal{P}}_{W}^{\perp} = \underline{\underline{I}}_{d} - \underline{\underline{W}}^{\top} \left ( \underline{\underline{W}}\,\underline{\underline{W}}^{\top}\right )^{-1}\underline{\underline{W}}$ , where the rows of the $k\times d$ matrix $\underline{\underline{W}}$ span the $k$ dimensional XC subspace we wish to identify. From this MCA-based approach Welling et al derive the probabilistic model for XCA as zero-mean multivariate Gaussian distribution with inverse covariance,

$\underline{\underline{C}}_{XCA}^{-1} = \frac{1}{\sigma^{2}_{0}} {\cal{P}}_{W}^{\perp} + \underline{\underline{W}}^{\top}\,\underline{\underline{W}}$

The two probabilistic forms of XCA are equivalent and so one finds that the matrices $\underline{\underline{A}}$ and $\underline{\underline{W}}$ are related via $\underline{\underline{A}} = \underline{\underline{W}}^{\top}\left ( \underline{\underline{W}}\,\underline{\underline{W}}^{\top} \right )^{-1}$

If we also look at the two ways in which Welling et al derived a probabilistic model for XCA, we can see that they are very similar to the formulations of PPCA and PMCA respectively, just with the replacement of ${\cal{P}}_{A}^{\perp}$ for $\underline{\underline{I}}_{d}$ in the PPCA formulation, and the replacement of ${\cal{P}}_{W}^{\perp}$ for $\underline{\underline{I}}_{d}$ in the PMCA formulation. So with just a redefinition of how we add the noise in the probabilistic model, Welling et al derived a single probabilistic model that unifies PCA and MCA.

Note that we are now defining the minor components subspace as directions of unusually low variance, so we only need a few dimensions, i.e. $k_{MC} \ll d$ , whilst previously when we defined the minor components subspace as the subspace where we wanted to constrain the data away from, we needed $k_{MC} = d - k_{PC}$ directions. The probabilistic formulation of XCA is a very natural and efficient way to express MCA.

Maximum Likelihood solution for XCA

The model likelihood is easily written down and the maximum likelihood solution identified. As one might anticipate the maximum-likelihood estimates for the vectors $\underline{a}_{i}$ are just eigenvectors of $\underline{\underline{\hat{C}}}$ , but we need to work out which ones. We can use the likelihood value at the maximum likelihood solution to do that for us.

Let’s say we want to retain $k=6$ extreme components overall, and we’ll use ${\cal{C}}$ to denote the corresponding set of eigenvalues of $\hat{\underline{\underline{C}}}$ that are retained. The maximum likelihood value for $k$ extreme components (XC) is given by,

$\log L_{ML} = - \frac{Nd}{2}\log \left ( 2\pi e\right )\;-\;\frac{N}{2}\sum_{i\in {\cal{C}}}\lambda_{i}\;-\;\frac{N(d-k)}{2}\log \left ( \frac{1}{d-k}\left [ {\rm tr}\hat{\underline{\underline{C}}} - \sum_{i\in {\cal{C}}} \lambda_{i}\right ]\right )$

All we need to do is evaluate the above equation for all possible subsets ${\cal{C}}$ of size $k$ selected from the $d$ eigenvalues $\lambda_{i}\, , i=1,\ldots,d$ of $\hat{\underline{\underline{C}}}$ . Superificially, this looks like a nasty combinatorial optimization problem, of exponential complexity. But as Welling et al point out, we know from a result proved in the PPCA paper of Tipping and Bishop that in the maximum likelihood solution the non-extreme components have eigenvalues that form a contiguous group, and so the optimal choice of subset ${\cal{C}}$ reduces to determining where to split the ordered eigenvalue spectrum of $\hat{\underline{\underline{C}}}$ . Since we have $k = k_{PC} + k_{MC}$ that reduces to simply determining the optimal number of the largest $\lambda_{i}$ to retain. That makes the optimization problem linear in $k$ .

For example, in our hypothetical example we have said we want $k = 6$ , but that could be a 3 PCs + 3 MCs split, or a 2 PCs + 4MCs split, and so on. To determine which we simply compute the maxium likelihood value for all the possible values of $k_{PC}$ from $k_{PC} = 0$ to $k_{PC} = k$ , each time keeping the largest $k_{PC}$ values of $\lambda_{i}$ and the smallest $k - k_{PC}$ values of $\lambda_{i}$ in our set ${\cal{C}}$ .

Some of the terms in $\log L_{ML}$ don’t change as we vary $k_{PC}$ and can be dropped. Welling et al introduce a quantity ${\cal{K}}$ defined by,

${\cal{K}}\;=\sum_{i\in {\cal{C}}}\lambda_{i}\; + \;(d-k)\log \left ( {\rm tr}\hat{\underline{\underline{C}}} - \sum_{i\in {\cal{C}}} \lambda_{i}\right )$

${\cal{K}}$ is the negative of $\log L_{ML}$ , up to an irrelevant constant and scale. If we then compute ${\cal{K}}\left ( k_{PC}\right )$ for all values of $k_{PC} = 0$ to $k_{PC} = k$ and select the minimum, we can determine the optimal split of $k = k_{PC} + k_{MC}$ .

PCA and MCA as special cases

Potentially, we could find that $k_{PC} = k$ , in which case all the selected extreme components would correspond to principal components, and so the XCA algorithm becomes equivalent to PCA. Likewise, we could get $k_{PC} = 0$ , in which case all the selected extreme components would correspond to minor components and the XCA algorithm becomes equivalent to MCA. So XCA contains pure PCA and pure MCA as special cases. But when do these special cases arise? Obviously, it will depend upon the precise values of the sample covariance eigenvalues $\lambda_{i}$ , or rather the shape of the eigen-spectrum, but Welling et al also give some insight here, namely,

A log-convex sample covariance eigen-spectrum will give PCA
A log-concave sample covariance eigen-spectrum will give MCA
A sample covariance eigen-spectrum that is neither log-convex nor log-concave will yield both principal components and minor components

In layman’s terms, if the plot of the (sorted) eigenvalues on a log scale only bends upwards (has positive second derivative) then XCA will give just principal components, whilst if the plot of the (sorted) eigenvalues on a log scale only bends downwards (has negative second derivative) then we’ll get just minor components. If the plot of the log-eigenvalues has places where the second derivative is positive and places where it is negative, then XCA will yield a mixture of principal and minor components.

Experimental demonstration

To illustrate the XCA theory I produced a Jupyter notebook that generates simulated data containing a known number of principal components and a known number of minor components. The simulated data is drawn from a zero-mean multivariate Gaussian distribution with population covariance matrix $\underline{\underline{C}}$ whose eigenvalues $\Lambda_{i}$ have been set to the following values,

$\begin{array}{cclcl} \Lambda_{i} & = & 3\sigma^{2}\left ( 1 + \frac{i}{3k_{PC}}\right ) & , & i=1,\ldots, k_{PC}\\ \Lambda_{i} & = &\sigma^{2} & , & i=k_{PC}+1,\ldots, d - k_{MC}\\ \Lambda_{i} & = & 0.1\sigma^{2}\left ( 1 + \frac{i}{k_{MC}} \right ) & , & i=d - k_{MC} + 1,\ldots, d\end{array}$

The first $k_{PC}$ eigenvalues represent principal components, as their variance is considerable higher than the ‘noise’ eigenvalues, which are represented by eigenvalues $i=k_{PC}+1$ to $i=d - k_{MC}$ . The last $k_{MC}$ eigenvalues represent minor components, as their variance is considerably lower than the ‘noise’ eigenvalues. Note, I have scaled both the PC and MC population eigenvalues by the ‘noise’ variance $\sigma^{2}$ , so that $\sigma^{2}$ just sets an arbitrary (user-chosen) scale for all the variances. I have chosen a large value of $\sigma^{2}=20$ , so that when I plot the minor component eigenvalues of $\hat{\underline{\underline{C}}}$ I can easily distinguish them from zero (without having to plot on a logarithmic scale).

We would expect the eigenvalues of the sample covariance matrix to follow a similar pattern to the eigenvalues of the population covariance matrix that we used to generate the data, i.e. we expect a small group of noticeably low-valued eigenvalues, a small group of noticeably high-valued eigenvalues, and the bulk (majority) of the eigenvalues to form a continuous spectrum of values.

I generated a dataset consisting of $N=2000$ datapoints, each with $d=200$ features. For this dataset I chose $k_{PC} = k_{MC} = 10$ . From the simulated data I computed the sample covariance matrix $\hat{\underline{\underline{C}}}$ and then calculated the eigenvalues of $\hat{\underline{\underline{C}}}$ .

The left-hand plot below shows all the eigenvalues (sorted from lowest to highest), and I have also zoomed in on just the smallest values (middle plot) and just the largest values (right-hand plot). We can clearly see that the sample covariance eigenvalues follow the pattern we expected.

All the code (with explanations) for my calculations are in a Jupyter notebook and freely available from the github repository https://github.com/dchoyle/xca_post

From the left-hand plot we can see that there are places where the sample covariance eigenspectrum bends upwards and places where it bends downwards, indicating that we would expect the XCA algorithm to retain both principal and minor components. In fact, we can clearly see from the middle and right-hand plots the distinct group of minor component eigenvalues and the distinct group of principal component eigenvalues, and how these correspond to the distinct groups of extreme components in the population covariance eigenvalues. However, it would be interesting to see how the XCA algorithm performs in selecting a value for the number of principal components.

For the eigenvalues above I have calculated the minimum value of ${\cal{K}}$ for a total of $k=20$ extreme components. The minimum value of ${\cal{K}}$ occurs at $k_{PC} = 10$ , indicating that the method of Welling et al estimates that there are 10 principal components in this dataset, and by definition there are then $k - k_{PC} = 20 - 10 = 10$ minor components in the dataset. In this case, the XCA algorithm has identified the dimensionalities of the PC and MC subspaces exactly.

Final comments

That is the post done. I can look Will in the eye when we meet for a beer at the next PyData Manchester Leaders meeting. The post ended up being longer (and more fun) than I expected, and you may have got the impression from my post that the Welling et al paper has completely solved the problem of selecting interesting low-dimensional subspaces in Gaussian distributed data. Note quite true. There are still challenges with XCA, as there are with PCA. For example, we have not said how we choose the total number of extreme components $k$ . That is a whole other model selection problem and one that is particularly interesting for PCA when we have high-dimensional data. This is one of my research areas – see for example my JMLR paper Hoyle2008.

Another challenge that is particularly relevant for high-dimensional data is the question of whether we will see distinct groups of principal and minor component sample covariance eigenvalues at all, even when we have distinct groups of population covariance eigenvalues. I chose very carefully the settings used to generate the simulated data in the example above. I ensured that we had many more samples than features, i.e. $N \gg d$ , and that the extreme component population covariance eigenvalues were distinctly different from the ‘noise’ population eigenvalues. This ensured that the sample covariance eigenvalues separated into three clearly visible groups.

However, in PCA when we have $N < d$ and/or weak signal strengths for the extreme components of the population covariance, then the extreme component sample covariance eigenvalues may not be separated from the bulk of the other eigenvalues. As we increase the ratio $\alpha = N/d$ we observe a series of phase transitions at which each of the extreme components becomes detectable – again this is another of my areas of research expertise [HoyleRattray2003, HoyleRattray2004, HoyleRattray2007]

Footnotes

I have used the usual divide by $N-1$ Bessel correction in the definition of the sample covariance. This is because I have assumed any data matrix will have been explicitly mean-centered. In many of the analyses of PCA the starting assumption is that the data is drawn from a mean-zero distribution, so that the sample mean of any feature is zero only under expectation, not as a constraint. Consequently, most formal analysis of PCA will define the sample covariance matrix with a $1/N$ factor. Since I have to deal with real data, I will never presume the data been drawn from population distribution that has zero-mean and so to model the data with a zero-mean distribution I will explicitly mean-center the data. Therefore, I use the $1/(N-1)$ definition of the sample covariance. Strictly speaking, that means the various theories and analyses I discuss later in the post are not applicable to the data I’ll work with. It is possible to modify the various analyses to explicitly take into account the mean-centering step, but it is tedious to do so. In practice, (for large $N$ ) the difference is largely inconsequential, and formulae derived from analysis of zero-mean distributed data can be accurate for mean-centered data, so we’ll stick with using the $1/(N-1)$ definition for $\hat{\underline{\underline{C}}}$ .
Hotelling, H. “Analysis of a complex of statistical variables into principal components”. Journal of Educational Psychology, 24:417-441 and also 24:498–520, 1933.
https://dx.doi.org/10.1037/h0071325
Oja, E. “Principal components, minor components, and linear neural networks”. Neural Networks, 5(6):927-935, 1992. https://doi.org/10.1016/S0893-6080(05)80089-9
Luo, F.-L., Unbehauen, R. and Cichocki, A. “A Minor Component Analysis Algorithm”. Neural Networks, 10(2):291-297, 1997. https://doi.org/10.1016/S0893-6080(96)00063-9
See for example, Williams, C.K.I. and Agakov, F.V. “Products of gaussians and probabilistic minor components analysis”. Neural Computation, 14(5):1169-1182, 2002. https://doi.org/10.1162/089976602753633439

You’re going to need a bigger algorithm – Amdahl’s Law and your responsibilities as a Data Scientist

On February 2, 2024December 17, 2024 By dchoyleIn Algorithms, Data Science, Mathematical Analysis, Scientific computationLeave a comment

You have some prototype Data Science code based on an algorithm you have designed. The code needs to be productionized, and so sped up to meet the specified production run-times. If you stick to your existing technology stack, unless the runtimes of your prototype code are within a factor of 1000 of your target production runtimes, you’ll need a bigger, better algorithm. There is a limit to what speed up your technology stack can achieve. Why is this? Read on and I’ll explain. And I’ll explain what you can do if you need more than a 1000-fold speed up of your prototype.

Speeding up your code with your current tech stack

There are two ways in which you can speed up your prototype code,

Improve the efficiency of the language constructs used, e.g. in Python replacing for loops with list comprehensions or maps, refactoring subsections of the code etc.
Horizontal scaling of your current hardware, e.g. adding more nodes to a compute cluster, adding more executors to the pool in a Spark cluster.

Point 2 assumes that your calculation is compute bound and not memory bound, but we’ll stick with that assumption for this article. We also exclude the possibility that the productionization team can invent or buy a new technology that is sufficiently different or better than your current tech stack – it would be an unfair ask of the ML engineers to have to invent a whole new technology just to compensate for your poor prototype. They may be able to to, but we are talking solely about using your current tech stack and we assume that it does have some capacity to be horizontally scaled.

So what speed ups can we expect from points 1 and 2 above? Point 1 is always possible. There are always opportunities for improving code efficiency that you or another person will spot when looking at the code for a second time. A more experienced programmer reviewing the code can definitely help. But let’s assume that you’re a reasonably experienced Data Scientist yourself. It is unlikely that your code is so bad that a review by someone else would speed it up by more than a factor of 10 or so.

So if the most we expect from code efficiency improvements is a factor 10 speed up, what speed up can we additionally get from horizontal scaling of your existing tech stack? A factor of 100 at most. Where does this limit of 100 come from? Amdahl’s law.

Amdahl’s law

Amdahl’s law is a great little law. Its origins are in High Performance Computing (HPC), but it has a very intuitive basis and so is widely applicable. Because of that it is worth explaining in detail.

Imagine we have a task that currently takes time T to run. Part of that task can be divided up and performed by separate workers or resources such as compute nodes. Let’s use P to denote the fraction of the task that can be divided up. We choose the symbol P because this part of the overall task can be parallelized. The fraction that can’t be divided up we denote by S, because it is the non-parallelizable or serial part of the task. The serial part of the task represents things like unavoidable overhead and operations in manipulating input and output data-structures and so on.

Obviously, since we’re talking about fractions of the overall runtime T, the fractions P and S must sum to 1, i.e.

The parallelizable part of the task takes time TP to run, whilst the serial part takes time TS to run.

What happens if we do parallelize that parallelizable component P? We’ll parallelize it using N workers or executors. When N=1, the parallelizable part took time TP to run, so with N workers it should (in an ideal world) take time TP/N to run. Now our overall run time, as a function of N is,

This is Amdahl’s law¹. It looks simple but let’s unpack it in more detail. We can write the speed up factor in going from T(N=1) to T(N) as,

The figure below shows plots of the speed-up factor against N, for different values of S.

From the plot in the figure, you can see that the speed up factor initially looks close to linear in N and then saturates. The speed up at saturation depends on the size of the serial component S. There is clearly a limit to the amount of speed up we can achieve. When N is large, we can approximate the speed up factor in Eq.3 as,

From Eq.4 (or from Eq.3) we can see the limiting speed up factor is 1/S. The mathematical approximation in Eq.4 hides the intuition behind the result. The intuition is this; if the total runtime is,

then at some point we will have made N big enough that P/N is smaller than S. This means we have reduced the runtime of the parallelizable part to below that of the serial part. The largest contribution to the overall runtime is now the serial part, not the parallelizable part. Increasing N further won’t change this. We have hit a point of rapidly diminishing returns. And by definition we can’t reduce S by any horizontal scaling. This means that when P/N becomes comparable to S, there is little point in increasing N further and we have effectively reached the saturation speed up.

How small is S?

This is the million-dollar question, as the size of S determines the limiting speed up factor we can achieve through horizontal scaling. A larger value of S means a smaller speed up factor limit. And here’s the depressing part – you’ll be very lucky to get S close to 1%, which would give you a speed up factor limit of 100.

A real-world example

To explain why S = 0.01 is around the lowest serial fraction you’ll observe in a real calculation, I’ll give you a real example. I first came across Amdahl’s law in 2007/2008, whilst working on a genomics project, processing very high-dimensional data sets². The calculations I was doing were statistical hypothesis tests run multiple times.

This is an example of an “embarrassingly parallel” calculation since it just involves splitting up a dataframe into subsets of rows and sending the subsets to the worker nodes of the cluster. There is no sophistication to how the calculation is parallelized, it is almost embarrassing to do – hence the term “embarrassingly parallel”.

The dataframe I had was already sorted in the appropriate order, so parallelization consisted of taking a small number of rows off the top of dataframe and sending to a worker node and repeating. Mathematically, on paper, we had S=0. Timings of actual calculations with different numbers of compute nodes and fitting an Amdahl’s law curve to those timings revealed we had something between S=0.01 and S=0.05.

A value of S=0.01 gaves us a maximum speed up factor of 100 from horizontal scaling. And this was for a problem that on paper had S=0. In reality, there is always some code overhead in manipulating the data. A more realistic limit on S for an average complexity piece of Data Science code would be S=0.05 or S=0.1, meaning we should expect limits on the speed up factor of between 10 and 20.

What to do?

Disappointing isn’t it!? Horizontal scaling will speed up our calculation by at most a factor of 100, and more likely only a factor of 10-20. What does it mean for productionizing our prototype code? If we also include the improvements in the code efficiency, the most we’re likely to be able to speed up our prototype code by is a factor of 1000 overall. It means that as a Data Scientist you have a responsibility to ensure the runtime of your initial prototype is within a factor of 1000 of the production runtime requirements.

If a speed up of 1000 isn’t enough to hit the production run-time requirements, what can we do? Don’t despair. You have several options. Firstly, you can always change the technology underpinning your tech stack. Despite what I said at the beginning of this post, if you are repeatedly finding that horizontal scaling of your current tech stack does not give you the speed-up you require, then there may be a case for either vertical scaling the runtime performance of each worker node or using a superior tech stack if one exists.

If improvement by vertical scaling of individual compute nodes is not possible, then there are still things you can do to mitigate the situation. Put the coffee on, sharpen your pencil, and start work on designing a faster algorithm. There are two approaches you can use here,

Reduce the performance requirements: This could be lowering the accuracy through approximations that are simpler and quicker to calculate. For example, if your code involves significant matrix inversion operations you may be able to approximate a matrix by its diagonal and explicitly hard code the calculation of its inverse rather than performing expensive numerical inversion of the full matrix.
Construct a better algorithm: There are no easy recipes here. You can get some hints on where to focus your effort and attention by identifying the runtime bottlenecks in your initial prototype. This can be done using code profiling tools. Once a bottleneck has been identified, you can then progress by simplifying the problem and constructing a toy problem that has the same mathematical characteristics as the original bottleneck. By speeding up the toy problem you will learn a lot. You can then apply those learnings, even if only approximately, to the original bottleneck problem.

When I first stumbled across Amdahl’s law, I mentioned it to a colleague working on the same project as I was. They were a full-stack software developer and immediately, said “oh, you mean Amdahl’s law about limits on the speed you can write to disk?”. It turns out there is another Amdahl’s Law, often called “Amdahl’s Second Law”, or “Amdahl’s Other Law”, or “Amdahl’s Lesser Law”, or “Amdahl’s Rule-Of-Thumb”. See this blog post, for example, for more details on Amdahl’s Second Law.
Hoyle et. al, “Shared Genomics: High Performance Computing for distributed insights in genomic medical research”, Studies in Health Technology & Informatics 147:232-241, 2009.

How many iterations are needed for the bisection algorithm?

On August 9, 2022September 24, 2022 By dchoyleIn Algorithms, Data Science, Mathematical Analysis, Scientific computationLeave a comment

<TL;DR>

The bisection algorithm is a very simple algorithm for finding the root of a 1-D function.
Working out the number of iterations of the algorithm required to determine the root location within a specified tolerance can be determined from a very simple little hack, which I explain here.
Things get more interesting when we consider variants of the bisection algorithm, where we cut an interval into unequal portions.

</TL;DR>

A little while ago a colleague mentioned that they were repeatedly using an off-the-shelf bisection algorithm to find the root of a function. The algorithm required the user to specify the number of iterations to run the bisection for. Since my colleague was running the algorithm repeatedly they wanted to set the number of iterations efficiently and also to achieve a guaranteed level of accuracy, but they didn’t know how to do this.

I mentioned that it was very simple to do this and it was a couple of lines of arithmetic in a little hack that I’d used many times. Then I realised that the hack was obvious and known to me because I was old – I’d been doing this sort of thing for years. My colleague hadn’t. So I thought the hack would be a good subject for a short blog post.

The idea behind a bisection algorithm is simple and illustrated in Figure 1 below.

How the bisection algorithm works — Figure 1: Schematic of how the bisection algorithm works

At each iteration we determine whether the root is to the right of the current mid-point, in the right-hand interval, or to the left of the current mid-point, in the left-hand interval. In either case, the range within which we locate the root halves. We have gone from knowing it was in the interval $[x_{lower}, x_{upper}]$ , which has width $x_{upper}-x_{lower}$ , to knowing it is in an interval of width $\frac{1}{2}(x_{upper}-x_{lower})$ . So with every iteration we reduce our uncertainty of where the root is located by half. After $N$ iterations we have reduced our initial uncertainty by $(1/2)^{N}$ . Given our initial uncertainty is determined by the initial bracketing of the root, i.e. an interval of width $(x_{upper}^{(initial)}-x_{lower}^{(initial)})$ , we can now work out that after $N$ iterations we have narrowed down the root to an interval of width ${\rm initial\;width} \times \left ( \frac{1}{2}\right ) ^{N}$ . Now if we want to locate the root to within a tolerance ${\rm tol}$ , we just have to keep iterating until the uncertainty reaches ${\rm tol}$ . That is, we run for $N$ iterations where $N$ satisfies,

$\displaystyle N\;=\; -\frac{\ln({\rm initial\;width/tol})}{\ln\left (\frac{1}{2} \right )}$

Strictly speaking we need to run for $\lceil N \rceil$ iterations. Usually I will add on a few extra iterations, e.g. 3 to 5, as an engineering safety factor.

As a means of easily and quickly determining the number of iterations to run a bisection algorithm the calculation above is simple, easy to understand and a great little hack to remember.

Is bisection optimal?

The bisection algorithm works by dividing into two our current estimate of the interval in which the root lies. Dividing the interval in two is efficient. It is like we are playing the childhood game “Guess Who”, where we ask questions about the characters’ features in order to eliminate them.

Asking about a feature that approximately half the remaining characters possess is the most efficient – it has a reasonable probability of applying to the target character and eliminates half of the remaining characters. If we have single question, with a binary outcome and a probability $p$ of one of those outcomes, then the question that has $p = \frac{1}{2}$ maximizes the expected information (the entropy), $p\ln (p)\;+\; (1-p)\ln(1-p)$ .

Dividing the interval unequally

When we first played “Guess Who” as kids we learnt that asking questions with a much lower probability $p$ of being correct didn’t win the game. Is the same true for our root finding algorithm? If instead we divide each interval into unequal portions is the root finding less efficient than when we bisect the interval?

Let’s repeat the derivation but with a different cut-point e.g. 25% along the current interval bracketing the root. In general we can test whether the root is to the left of right of a point that is a proportion $\phi$ along the current interval, meaning the cut-point is $x_{lower} + \phi (x_{upper}-x_{lower})$ . At each iteration we don’t know in advance which side of the cut-point the root lies until we test for it, so in trying to determine in advance the number of iterations we need to run, we have to assume the worst case scenario and assume that the root is still in the larger of the two intervals. The reduction in uncertainty is then, ${\rm max}\{\phi, 1-\phi\}$ . Repeating the derivation we find that we have to run at least,

$\displaystyle N_{Worst\;Case}\;=\;\ -\frac{\ln({\rm initial\;width/tol})}{\ln\left ({\rm max}\{\phi, 1 - \phi \right \})}$

iterations to be guaranteed that we have located the root to within $tol$ .

Now to determine the cut-point $\phi$ that minimizes the upper bound on number of iterations required, we simply differentiate the expression above with respect to $\phi$ . Doing so we find,

$\displaystyle \frac{\partial N_{Worst\;Case}}{\partial \phi} \;=\; -\frac{\ln({\rm initial\;width/tol})}{ (1-\phi) \left ( \ln (1 - \phi) \right )^{2}} \;\;,\;\; \phi < \frac{1}{2}$

and

$\displaystyle \frac{\partial N_{Worst\;Case}}{\partial \phi} \;=\; \frac{\ln({\rm initial\;width/tol})}{\phi \left ( \ln (\phi) \right)^{2}} \;\;,\;\; \phi > \frac{1}{2}$

The minimum of $N_{Worst\;Case}$ is at $\phi =\frac{1}{2}$ , although $\phi=\frac{1}{2}$ is not a stationary point of the upper bound $N_{Worst\;Case}$ , as $N_{Worst\;Case}$ has a discontinuous gradient there.

That is the behaviour of the worst-case scenario. A similar analysis can be applied to the best-case scenario – we simply replace $max$ with $min$ in all the above formula. That is, in the best-case scenario the number of iterations required is given by,

$\displaystyle N_{Best\;Case}\;=\;-\frac{\ln({\rm initial\;width/tol})}{\ln\left ({\rm min}\{\phi, 1 - \phi \right \})}$

Here, the maximum of the best-case number of iterations occurs when $\phi = \frac{1}{2}$ .

That’s the worst-case and best-case scenarios, but how many iterations do we expect to use on average? Let’s look at the expected reduction in uncertainty in the root location after $N$ iterations. In a single iteration a root that is randomly located within our interval will lie, with probability $\phi$ , in segement to the left of our cut-point and leads to a reduction in the uncertainty by a factor of $\phi$ . Similarly, we get a reduction in uncertainty of $1-\phi$ with probability $1-\phi$ if our randomly located root is to the right of the cut-point. So after $N$ iterations the expected reduction in uncertainty is,

$\displaystyle {\rm Expected\;reduction}\;=\;\left ( \phi^{2}\;+\;(1-\phi)^{2}\right )^{N}$

Using this as an approximation to determine the typical number of iterations, we get,

$\displaystyle N_{Expected\;Reduction}\;=\;-\frac{\ln({\rm initial\;width/tol})}{\ln\left ( \phi^{2} + (1-\phi)^{2} \right )}$

This still isn’t the expected number of iterations, but to see how it compares Figure 2 belows shows simulation estimates of $\mathbb{E}\left ( N \right )$ plotted against $\phi$ when the root is random and uniformly distributed within the original interval.

The number of iterations needed for the bisection algorithm — Number of iterations required for the different root finding methods.

For Figure 2 we have set $w = ({\rm initial\;width/tol}) = 0.01$ . Also plotted in Figure 2 are our three theoretical estimates, $\lceil N_{Worst\;Case}\rceil, \lceil N_{Best\;Case}\rceil, \lceil N_{Expected\;Reduction}\rceil$ . The stepped structure in these 3 integer quantities is clearly apparent, as is how many more iterations are required under the worst case method when $\phi \neq \frac{1}{2}$ .

The expected number of iterations required, $\mathbb{E}( N )$ , actually shows a rich structure that isn’t clear unless you zoom in. Some aspects of that structure were unexpected, but requires some more involved mathematics to understand. I may save that for a follow-up post at a later date.

Hoyle Analytics

Tag: Algorithms

Data Science Notes: 2. Log-Sum-Exp

Summary

Introduction

The log-sum-exp trick

The log-sum-exp function

Calculating log-sum-exp in Python

Extreme Components Analysis

A deadline

A recap of PCA

MCA

XCA

Probabilistic PCA and MCA

Probabilistic XCA

Maximum Likelihood solution for XCA

PCA and MCA as special cases

Experimental demonstration

Final comments

Footnotes

You’re going to need a bigger algorithm – Amdahl’s Law and your responsibilities as a Data Scientist

Speeding up your code with your current tech stack

Amdahl’s law

How small is S?

A real-world example

What to do?

How many iterations are needed for the bisection algorithm?

<TL;DR>

</TL;DR>

Is bisection optimal?

Dividing the interval unequally

Summary

Introduction

The log-sum-exp trick

The log-sum-exp function

Calculating log-sum-exp in Python

Share this:

A deadline

A recap of PCA

MCA

XCA

Probabilistic PCA and MCA

Probabilistic XCA

Maximum Likelihood solution for XCA

PCA and MCA as special cases

Experimental demonstration

Final comments

Footnotes

Share this:

Speeding up your code with your current tech stack

Amdahl’s law

How small is S?

A real-world example

What to do?

Share this:

<TL;DR>

</TL;DR>

Is bisection optimal?

Dividing the interval unequally

Share this: