20 Further Reading
This book is deliberately short: it aims to give you the machinery for deriving an objective, not to catalogue every objective in the literature. These are the books and papers worth going to next.
20.1 General textbooks covering this material in more depth
- Christopher Bishop, Pattern Recognition and Machine Learning (Bishop 2006) — Chapter 1 (probability and decision theory) and Chapter 4 (discriminative classification) cover much of Parts I–V of this book with more statistical machinery.
- Kevin Murphy, Probabilistic Machine Learning: An Introduction (Murphy 2022) and Advanced Topics (Murphy 2023) — the modern, exhaustive reference for everything probabilistic in this book, including proper scoring rules and calibration.
- Trevor Hastie, Robert Tibshirani, and Jerome Friedman, The Elements of Statistical Learning (Hastie et al. 2009) — the standard reference for regularization, model selection, and the statistical view of loss minimization.
- Ian Goodfellow, Yoshua Bengio, and Aaron Courville, Deep Learning (Goodfellow et al. 2016) — Chapter 5 covers maximum likelihood estimation as the basis for essentially all deep learning objectives.
20.2 Information theory
- Thomas Cover and Joy Thomas, Elements of Information Theory (Cover and Thomas 2006) — the canonical treatment of entropy, cross-entropy, and KL divergence, including the inequalities this book only states.
- Claude Shannon, “A Mathematical Theory of Communication” (Shannon 1948) — where entropy, as used everywhere in this book, originates.
20.3 Statistical decision theory and scoring rules
- Abraham Wald, “Statistical Decision Functions” (Wald 1949) and James Berger, Statistical Decision Theory and Bayesian Analysis (Berger 1985) — the formal decision-theoretic framework behind Chapter 8.
- Tilmann Gneiting and Adrian Raftery, “Strictly Proper Scoring Rules, Prediction, and Estimation” (Gneiting and Raftery 2007) — the modern reference on proper scoring rules, well beyond what this book’s single chapter covers.
20.4 Statistical learning theory and surrogate losses
- Vladimir Vapnik, Statistical Learning Theory (Vapnik 1998) — the origin of the margin-based, empirical-risk-minimization view this book leans on throughout Parts V–VI.
- Peter Bartlett, Michael Jordan, and Jon McAuliffe, “Convexity, Classification, and Risk Bounds” (Bartlett et al. 2006) — the precise theory of when a convex surrogate is statistically consistent for 0–1 loss, referenced in Chapter 9.
20.5 Learning to rank
- Christopher Burges, “From RankNet to LambdaRank to LambdaMART: An Overview” (Burges 2010) — a first-hand account of the reasoning behind Chapters 11 and 13, by the author of both methods.
- Tie-Yan Liu’s Learning to Rank for Information Retrieval survey and the original ListNet/ListMLE papers (Cao et al. 2007; Xia et al. 2008) for the full pointwise/pairwise/listwise landscape this book only samples.
20.6 Contrastive and representation learning
- Aäron van den Oord, Yazhe Li, and Oriol Vinyals, “Representation Learning with Contrastive Predictive Coding” (Oord et al. 2018) — the InfoNCE objective used throughout modern self-supervised learning.
- Ting Chen et al., “A Simple Framework for Contrastive Learning of Visual Representations” (SimCLR) (Chen et al. 2020) for a concrete, influential application of the ideas in Chapter 14.
20.7 Language models and preference optimization
- Ashish Vaswani et al., “Attention Is All You Need” (Vaswani et al. 2017) — the architecture whose training objective is the subject of Chapter 15.
- Rafael Rafailov et al., “Direct Preference Optimization: Your Language Model is Secretly a Reward Model” (Rafailov et al. 2023) — the paper underlying Chapter 16; read it directly for the full derivation this book compresses.
- Paul Christiano et al., “Deep Reinforcement Learning from Human Preferences” (Christiano et al. 2017) and Long Ouyang et al., “Training Language Models to Follow Instructions with Human Feedback” (Ouyang et al. 2022) for the RLHF lineage DPO grew out of.