Machine Learning Theory
Syllabus, Master's level, 1TD179
- Code
- 1TD179
- Education cycle
- Second cycle
- Main field(s) of study and in-depth level
- Computer Science A1F, Data Science A1F, Mathematics A1F
- Grading system
- Pass with distinction (5), Pass with credit (4), Pass (3), Fail (U)
- Finalised by
- The Faculty Board of Science and Technology, 5 February 2026
- Responsible department
- Department of Information Technology
Entry requirements
120 credits in engineering/science including Computer Programming I, Probability and Statistics, Linear Algebra II, Several Variable Calculus/Analysis in Several Variables. Participation in Statistical Machine Learning. Proficiency in English equivalent to the Swedish upper secondary course English 6.
Learning outcomes
On completion of the course, the student should be able to:
- Explain and compare classical (PAC/VC) versus algorithm-dependent (stability, PAC-Bayes, information-theoretic) frameworks for generalization, and specify when each is applicable.
- Derive and manipulate PAC-Bayes bounds (KL-based bounds, Gibbs classifiers), and explain how these can be optimized to produce non-vacuous bounds for deep networks.
- Formally link optimization geometry to generalization: define sharpness/flatness, derive intuitions for why flat minima generalize, and implement/analyze Sharpness-Aware Minimization (SAM).
- Analyze modern phenomena such as double descent and benign overfitting; explain the roles of overparameterization and data covariance structure.
- Utilize kernel/mean-field limits (Neural Tangent Kernel (NTK), infinite width regimes) to reason about the training dynamics of wide networks.
- Conduct a small research study or a rigorous empirical project that synthesizes theory and experiment, including evaluating modern theoretical papers.
Content
A rigorous, concept-driven treatment of modern statistical learning theory, transitioning from classical PAC/VC and uniform convergence/Rademacher tools to algorithm-dependent perspectives. The course covers:
- Algorithmic stability and its relation to generalization.
- PAC-Bayes theory.
- The geometry of loss landscapes (sharpness/flatness, Hessian-based approximations) and Sharpness-Aware Minimization (SAM).
- Overparameterization phenomena, including the Neural Tangent Kernel (NTK) and implicit bias in gradient methods.
- Double descent and benign overfitting in linear models.
- Information-theoretic and compression perspectives linking the various frameworks.
The module concludes with a discussion on open problems and student project presentations that integrate theory and reproducible experiments.
Instruction
Seminars and lectures.
Assessment
Written exam (3 credits). Oral and written presentation of assignments and project work (2 credits).
If there are special reasons for doing so, an examiner may make an exception from the method of assessment indicated and allow a student to be assessed by another method. An example of special reasons might be a certificate regarding special pedagogical support from the disability coordinator of the university.
Reading list
No reading list found.