Mathematics of Machine Learning
A course in 10 lectures, from optimization to generative models and optimal transport. Follow the mathematical ideas, then explore them through notebooks and worked examples.
Before you begin. Familiarity with linear algebra, multivariable calculus, and basic probability will help. The lectures emphasize the main concepts and methods; the linked books provide fuller proofs and background.
Each lecture gathers its topics, practical materials, and suggested reading. Transcripts are available for lectures 1–9. Lecture 10 is supported by the OT4ML book, slides, and notebooks.
Lecture 01
Smooth optimization
Read the transcript PDFTopics
- Introduction and motivation
- Gradients, Jacobians, Hessians
- Gradient descent and acceleration
- Stochastic Gradient Descent (SGD)
Notebooks & materials
Further reading
- Convex Optimization, by Boyd and Vandenberghe
- Introduction to Nonlinear Optimization: Theory, Algorithms, and Applications, by Amir Beck
Lecture 02
From smooth to nonsmooth optimization
Read the transcript PDFTopics
- Proofs of gradient descent and acceleration
- Linear models and regularization
- Ridge versus Lasso
- ISTA Algorithm
Notebooks & materials
- Notebook on Linear Regression (specifically the Lasso part)
- Notebook on Interior Point Methods
- Course notes: Mathematical Foundations of Data Sciences
Further reading
- Course Notes on Convexity by Vincent Duval
- Introduction to Nonlinear Optimization: Theory, Algorithms, and Applications, by Amir Beck
Lecture 03
Lasso and compressed sensing
Read the transcript PDFTopics
- Examples of non-smooth functionals (Lasso, TV regularization, constraints)
- Subgradient and proximal operators
- Forward-backward splitting, connection with FISTA
- ADMM, Douglas-Rachford (DR), Primal-Dual
- Compressive sensing theory
Notebooks & materials
- Course notes: Mathematical Foundations of Data Sciences,
- Notebook on Douglas-Rachford Proximal Method
- Proximal Operators Repository (including Python code)
- Non-Smooth Optimization Slides
- Compressed Sensing Slides
Further reading
- A Mathematical Introduction to Compressive Sensing by Simon Foucart and Holger Rauhut (advanced)
- Convex Optimization, by Boyd and Vandenberghe
- Proximal Algorithms, by N. Parikh and S. Boyd
Lecture 04
Kernels and neural-network architectures
Read the transcript PDFTopics
- Transition from ridge regression to kernels
- Multilayer Perceptron (MLP)
- Convolutional Neural Networks (CNN)
- ResNet architecture
- Transformer models
Notebooks & materials
Further reading
- The Elements of Statistical Learning, by Jerome H. Friedman, Robert Tibshirani, and Trevor Hastie
- Machine Learning: A Probabilistic Perspective, by Kevin Patrick Murphy (covers the theory of ML)
Lecture 05
Deep learning: theory and computation
Read the transcript PDFTopics
- Review of MLP and its variants (CNN, ResNet)
- Theoretical framework of two-layer MLPs
- Gradient and Jacobians in neural networks
- Introduction to backpropagation
Notebooks & materials
Further reading
Lecture 06
Differentiable programming
Read the transcript PDFTopics
- Recap on Gradient and Jacobian
- Forward and reverse mode automatic differentiation
- Introduction to PyTorch
- The adjoint method in computational mathematics
Notebooks & materials
Further reading
Lecture 07
Sampling and diffusion models
Read the transcript PDFTopics
- Refresher on Stochastic Gradient Descent (SGD)
- Introduction to Langevin dynamics
- Overview of diffusion models
Notebooks & materials
- Numerical tour on diffusion models
- Course notes on Diffusion Models
Lecture 08
Language models and generative AI
Read the transcript PDFTopics
- Overview of different generative model concepts
- Introduction to generative models (VAE, GANs, U-Net, diffusion)
- Self-supervised learning and next-token prediction
- Tokenizers
- Transformer architectures, FlashAttention
- State space models
Notebooks & materials
Further reading
Lecture 09
Generative models
Read the transcript PDFTopics
- Understanding generative models as density fitting techniques.
- Basics of Maximum Likelihood Estimation and f-divergences.
- Gaussian mixtures and the Expectation-Maximization algorithm.
- Variational Autoencoders (VAE).
- Introduction to Normalizing Flows.
- Generative Adversarial Networks (GANs), Wasserstein GANs (WGANs).
- Diffusion Models.
Lecture 10
Optimal transport
Read OT4MLTopics
- Introduction to Monge and Kantorovich formulations.
- The Sinkhorn algorithm.
- Training of generative models.
- Duality and Wasserstein GANs.
Further reading
- Computational Optimal Transport, by Gabriel Peyré and Marco Cuturi
- Optimal Transport for Applied Mathematicians, by Filippo Santambrogio (advanced)
- Python POT (Python Optimal Transport) toolbox