Online and Batch-Incremental Estimation of Covariance Matrices and Means in Python


Estimating the mean and covariance matrix of a dataset is a cornerstone of multivariate statistics and machine learning. While batch (offline) methods are straightforward when all data is available at once, many modern applications require online or streaming estimation — where data arrives sequentially, potentially at very high rates, and storing all past samples is infeasible.

In this Jupyter Notebook below, we explore fully online and batch-incremental estimators for the mean and covariance matrix, including their inverses. We look at how these algorithms work, why they are useful, and how they can adapt to non-stationary data through a forgetting factor. Importantly, these approaches allow us to update estimates efficiently without recomputing matrix inverses from scratch.

The revised companion uses seeded NumPy experiments to check weighted online and batch estimates against direct calculations. Inverse updates begin after sufficient rank is available. The examples distinguish forgetting per batch from forgetting per observation, and compare finite-sample memory with its limiting value. The notebook is the runnable companion to the seven-part series Online Estimation, which derives all of these updates from Appendix B.2 of my PhD thesis.

Related posts: The covariance matrix estimated here is a key ingredient for the Mahalanobis distance — which connects to the Chi-square distribution for anomaly detection thresholds. For Python implementations and benchmarks, see Implementing the Mahalanobis Distance in Python.




Enjoy Reading This Article?

Here are some more articles you might like to read next: