Fields Academy Shared Graduate Course: Mathematics for AI Safety
Location: Fields Institute, Room 309, Stewart Library
Description
Course Code at the University of Toronto: MAT1510
Registration Deadline: September 20, 2026
Instructor: Professor Yevgeny Liokumovich, University of Toronto
Course Dates: September 14 - December 9, 2026
Mid-Semester Break: October 12-16, 2026
Lecture Times: Mondays | 10:00 AM - 12:00 PM (ET); Wednesdays | 10:00 AM - 11:00 AM (ET)
Office Hours: Fridays | 10:00 AM - 11:00 AM (ET)
Registration Fee:
- Students from our Principal Sponsoring & Affiliate Universities: Free
- Other Students: CAD$500
Capacity Limit: 35 students (auditing is allowed for this course; please indicate in your registration if you are auditing)
Format:
- In-Person: Room 309, Fields Institute
- Online: via Zoom
Course Description
AI safety is a rapidly evolving multidisciplinary field concerned with understanding and mitigating risks posed by increasingly capable AI systems. A central challenge is to reason not only about systems that exist today, but also about risks and failure modes that may emerge as AI systems become more capable and autonomous in the near future.
This course will focus on mathematical questions motivated by AI safety and alignment. We will draw on techniques from statistics and probability, learning theory, geometry, optimization, game theory, and logic. Topics will include capability scaling and forecasting, the theory of deep learning, the geometry and interpretability of neural representations, scalable oversight, multi-agent interactions, and formal approaches to safety.
Many foundational questions in the field remain open, making it unusually accessible to new contributions. The student should aim to develop a viable research project with a novel contribution to the field by the end of the course.
Week-by-Week Topics (the exact contents may change as the field is rapidly developing)
- Week 1: AI capability trajectories
- Risks posed by AI depend on its capabilities. We will discuss neural network scaling laws and their theoretical explanations. AI capability evaluations and current modeling techniques of future AI capabilities. Automation of AI R&D and feedback loops.
- Week 2: Theory of Deep Learning
- Learning theories. SGD vs Bayes. Generalization properties of neural networks. Neural architectures and large-width limits. Training dynamics and phase transitions: grokking, in-context learning.
- Week 3: LLM training stages
- Next-token pretraining. Supervised fine-tuning and instruction tuning. Reward modeling, RLHF and RLAIF. Direct preference optimization and reinforcement learning for reasoning.
- Week 4: Jailbreaking and adversarial attacks
- Jailbreaks and adversarial prompting. Prompt injection and attacks on LLM agents. Red-teaming and automated adversarial attacks. Robustness of safety fine-tuning and defenses against jailbreaks.
- Week 5: Geometry of neural activations
- Linear representation hypothesis, computation in superposition, linear probes, SAEs, feature manifolds.
- Week 6: Applications of interpretability to control and alignment
- Applications of SAEs and linear probes. J-space and logit lens. Activation steering. Persona alignment and monitoring of safety-relevant representations.
- Week 7: Game theory of multi-agent interactions
- Cooperation and competition between AI agents. Repeated games, bargaining and coordination. Multi-agent learning. Commitment, collusion and open-source game theory.
- Week 8: Scalable oversight
- How can weaker evaluators supervise stronger AI systems? Safety via debate. Weak-to-strong generalization. Decomposition and AI-assisted evaluation.
- Week 9: Safety guarantees
- Formal verification methods in AI safety. Robustness certificates and provable guarantees for neural networks. Runtime monitoring and statistical versus worst-case guarantees.
- Week 10: Agency and decision theory
- Mathematical models of agency and goal-directed behavior. Utility theory, sequential decision-making and causal models of incentives. Instrumental goals, power-seeking and corrigibility.
- Week 11: Project presentations
- Week 12: Project presentations
Suggested Readings:
- Iliad Intensive Curriculum
- Introduction to AI Safety, Ethics, and Society, by Dan Hendrycks
Lecture notes from similar courses:
- AI Alignment — Roger Grosse, University of Toronto
- Mathematics for AI Safety — Lionel Levine, Cornell University
- AI Safety — Boaz Barak, Harvard University
Course expectations: The course will be graded on the final project. The course has no formal prerequisites, but students are encouraged to consult the syllabus for some optional pre-reading. If space allows, we will welcome auditors.
Schedule
| 10:00 to 12:00 |
Lecture 01 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 11:00 |
Lecture 02 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 12:00 |
Lecture 03 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 11:00 |
Lecture 04 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 12:00 |
Lecture 05 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 11:00 |
Lecture 06 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto Location:Online |
| 10:00 to 12:00 |
Lecture 07 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 11:00 |
Lecture 08 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 12:00 |
Lecture 09 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 11:00 |
Lecture 10 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 12:00 |
Lecture 11 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 11:00 |
Lecture 12 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto Location:Fields Institute, Room 210 |
| 10:00 to 12:00 |
Lecture 13 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 11:00 |
Lecture 14 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 12:00 |
Lecture 15 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 11:00 |
Lecture 16 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 12:00 |
Lecture 17 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 11:00 |
Lecture 18 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 12:00 |
Lecture 19 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 11:00 |
Lecture 20 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 12:00 |
Lecture 21 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 11:00 |
Lecture 22 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 12:00 |
Lecture 23 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |
| 10:00 to 11:00 |
Lecture 24 | Mathematics for AI Safety
Yevgeny Liokumovich, University of Toronto |


