PageSourceSearch

https://fabsilvestri.github.io/AIML-Course/

html fabsilvestri.github.io collected 2026-10-03 09:10:48 UTC 68,976 bytes, 1,415 lines download raw bytes

1<!DOCTYPE html>
2<html lang="en">
3<head>
4<meta charset="utf-8">
5<meta name="viewport" content="width=device-width, initial-scale=1">
6  <link rel="icon" href="assets/img/favicon.svg" type="image/svg+xml">
7  <link rel="icon" href="assets/img/favicon-32.png" sizes="32x32" type="image/png">
8  <link rel="apple-touch-icon" href="assets/img/favicon-180.png">
9  <meta name="theme-color" content="#0b3d62">
10<title>Applications of Machine Learning — BSc Mathematics of Artificial Intelligence</title>
11<meta name="description" content="Applicazioni Informatiche del Machine Learning — a 48-hour applied course: twenty-four lectures on how to build a machine learning system and how to know whether it works.">
12<meta name="robots" content="noindex, nofollow">
13<link rel="stylesheet" href="lib/katex/dist/katex.min.css">
14<link rel="stylesheet" href="assets/css/site.css">
15</head>
16<body>
17
18<a class="skip-link" href="#about">Skip to content</a>
19
20<nav class="site-nav" aria-label="Sections of this page">
21  <div class="wrap">
22    <a class="brand" href="#top">Applications of ML</a>
23    <a class="nav-link" href="#about">Principle</a>
24    <a class="nav-link" href="#method">Method</a>
25    <a class="nav-link" href="#prerequisites">Prerequisites</a>
26    <a class="nav-link" href="#calendar">Calendar</a>
27    <a class="nav-link" href="#lectures">Lectures</a>
28    <a class="nav-link" href="#derivations">Mathematics</a>
29    <a class="nav-link" href="#assessment">Assessment</a>
30    <a class="nav-link" href="#textbook">Textbook</a>
31    <a class="nav-link" href="#practicalities">Practicalities</a>
32    <!-- BEGIN EXERCISES_NAV -->
33    <!-- Held back by EXERCISE_BOOK_PUBLIC in tools/make_site.py -->    <!-- END EXERCISES_NAV -->
34  </div>
35</nav>
36
37<header class="hero" id="top">
38  <div class="wrap">
39    <p class="kicker">Sapienza Università di Roma &middot; BSc Mathematics of Artificial Intelligence</p>
40    <h1>Applications of<br>Machine Learning</h1>
41    <p class="it">Applicazioni Informatiche del Machine Learning</p>
42    <p class="tagline">
43      How to build a machine learning system, and how to know whether it
44      works.
45    </p>
46    <dl class="facts">
47      <div>
48        <dt>Load</dt>
49        <dd>48 academic hours &middot; 24 lectures</dd>
50      </div>
51      <div>
52        <dt>Year</dt>
53        <dd>Third year, BSc</dd>
54      </div>
55      <div>
56        <dt>When</dt>
57        <dd>Tue &amp; Wed &middot; 12:00&ndash;14:00 &middot; from 6 Oct 2026</dd>
58      </div>
59      <div>
60        <dt>Where</dt>
61        <dd><a href="https://www.mat.uniroma1.it/">Aula C &middot; Dip. di Matematica</a></dd>
62      </div>
63      <div>
64        <dt>Teaching language</dt>
65        <dd>Italian &middot; materials in English</dd>
66      </div>
67      <div>
68        <dt>Lecturer</dt>
69        <dd>Fabrizio Silvestri</dd>
70      </div>
71      <div>
72        <dt>Google Classroom</dt>
73        <dd><a href="https://classroom.google.com/c/MjU0MDA3MTk5MjJa">Join the class</a></dd>
74      </div>
75    </dl>
76  </div>
77</header>
78
79<!-- ===================================================================== -->
80
81<section id="about">
82  <div class="wrap">
83    <h2>The organising principle</h2>
84    <p class="lede">One topic per lecture, and every lecture stands on its own.</p>
85
86    <div class="panel">
87      <strong>Each lecture is the mathematics, the method, and a notebook that
88      implements it.</strong> If you miss a lecture, you read its slides and its
89      notebook and you are caught up &mdash; you never have to reconstruct
90      anything from the lecture before it.
91    </div>
92
93    <p>Ninety minutes, in four blocks.</p>
94
95    <div class="grid">
96      <div class="card">
97        <h4>The mathematics <span class="muted">&middot; 20 min</span></h4>
98        <p>The one object the method rests on, derived rather than quoted
99        &mdash; because the method does not make sense without it, and because
100        the derivations are what you still have in five years when the libraries
101        have changed.</p>
102      </div>
103      <div class="card">
104        <h4>The method <span class="muted">&middot; 40 min</span></h4>
105        <p>
105How it works, what its hyperparameters do, the conditions under which
106        it fails, and a worked example with real numbers from real data &mdash;
107        every figure on every slide reproduced by a script in this repository.</p>
108      </div>
109      <div class="card">
110        <h4>Further ground <span class="muted">&middot; 15 min</span></h4>
111        <p>The variants, when to prefer each, and what practitioners actually
112        reach for &mdash; which is not always what the textbook presents first.</p>
113      </div>
114      <div class="card">
115        <h4>The notebook <span class="muted">&middot; 5 min</span></h4>
116        <p>What is in it, what to run, and what to change to see something move.
117        You run it yourself afterwards; the lecture is not a live coding
118        session.</p>
119      </div>
120    </div>
121
122    <p style="margin-top:1.5rem">Ten minutes at the top of each lecture place it
123    in the arc: what we can already do, what we cannot, and which of the two
124    today&rsquo;s method changes.</p>
125  </div>
126</section>
127
128<!-- ===================================================================== -->
129
130<section id="method">
131  <div class="wrap">
132    <h2>Working method</h2>
133    <p class="lede">You will write machine learning code with an assistant
134    &mdash; if not in this course, then in the thesis after it and in the job
135    after that. So the course is explicit about how.</p>
136
137    <div class="tablewrap">
138      <table>
139        <thead>
140          <tr><th>#</th><th>Step</th><th>Who</th></tr>
141        </thead>
142        <tbody>
143          <tr><td class="num">1</td><td><strong>Specify</strong> &mdash; the input, the output, the constraint, the check</td><td>you</td></tr>
144          <tr><td class="num">2</td><td><strong>Generate</strong> &mdash; and then stop, before running anything</td><td>the assistant</td></tr>
145          <tr><td class="num">3</td><td><strong>Read</strong> &mdash; as a reviewer, not as an author</td><td>you</td></tr>
146          <tr><td class="num">4</td><td><strong>Test</strong> &mdash; against a case whose answer you already know</td><td>you</td></tr>
147          <tr><td class="num">5</td><td><strong>Verify</strong> &mdash; that the number means what it appears to mean</td><td>you</td></tr>
148        </tbody>
149      </table>
150    </div>
151
152    <p>Steps 3&ndash;5 are the course. Step 2 is the part that is free.</p>
153
154    <div class="panel">
155      <strong>Where you meet this loop here.</strong> The notebooks in this
156      course are written for you and they are correct &mdash; nothing in them is
157      wrong on purpose. Every code cell is preceded by the specification that
158      would produce it: the box is step&nbsp;1, the cell below it is what
159      step&nbsp;2 returned, and its <em>check</em> line is step&nbsp;4 written
160      down in advance. Read the box, answer the check in your head, then run the
161      cell.
162    </div>
163
164    <h3>Why the emphasis on verification</h3>
165    <p>Machine learning fails <em>silently</em>. A bug in a training loop does not
166    crash &mdash; it returns a plausible number:</p>
167    <ul class="plain">
168      <li>A scaler fitted before the split leaks the test set into training. Great score, broken system.</li>
169      <li>A missing <code>model.eval()</code> leaves dropout active during evaluation.</li>
170      <li>A missing <code>optimizer.zero_grad()</code> accumulates gradients with no error.</li>
171      <li>A metric averaged per batch rather than over the set is wrong by construction.</li>
172      <li>Accuracy on an imbalanced set reports 90% for a model that predicts one class.</li>
173    </ul>
174    <p>Every one of these runs. Every one produces a number. Each is taught in
175    the lecture where it belongs &mdash; leakage in Lecture&nbsp;2, imbalance in
176    Lecture&nbsp;3, and all three PyTorch failures &mdash; the optimiser pair and
177    batch averaging &mdash; in Lecture&nbsp;10, as a property of the method
178    rather than as a trap.</p>
179
180    <div class="panel panel-accent">
181      <strong>The notebooks are ours.</strong> The textbook supplies the
182      syllabus, not the code. No notebook from the author, or from any other
183      third party, is used in this course. Every one is written for it, is
184      complete and correct, and carries above each cell the specification that
185      would produce it.
186    </div>
187
188    <h3>Four rules</h3>
189    <ol>
190      <li><strong>Never keep code you cannot explain line by line.</strong>
191          Part B of the paper puts a cell in front of you and asks what would
192          have to be true for its number to be trusted.</li>
193      <li><strong>Every number needs a baseline.</strong>
193 A metric with nothing to compare it to is decoration.</li>
194      <li><strong>Every model needs an ablation.</strong> Remove a component; show the number moves.</li>
195      <li><strong>Report the failure cases.</strong> A system with no known failure mode has not been tested.</li>
196    </ol>
197  </div>
198</section>
199
200<!-- ===================================================================== -->
201
202<section id="prerequisites">
203  <div class="wrap">
204    <h2>Prerequisites</h2>
205    <p class="lede">Short version: the mathematics of a third-year BSc in
206    Mathematics of Artificial Intelligence, and enough Python to read code
207    critically. If you are missing something, none of it takes long to fix
208    &mdash; and everything below is a pointer, not a reading list to complete
209    before the first lecture.</p>
210
211    <div class="panel">
212      <strong>A note on scope.</strong> Everything <em>taught</em> in this
213      course comes from Chapters 1&ndash;16 of the textbook, or &mdash; for
214      Lectures 19&ndash;22 &mdash; from the lecture notes. The material on this
215      page is the opposite: it is what the course assumes you already have, so
216      the resources here necessarily point outside both. None of it is
217      examinable in its own right.
218    </div>
219
220    <h3>Test yourself in ten minutes</h3>
221    <p>These are not warm-up exercises. Each one is a thing you will actually be
222    asked to do, in the lecture where it appears. Try them before deciding you
223    need to revise anything.</p>
224
225    <ol class="selfcheck">
226      <li>
227        <p class="q">$\mathbf{X}$ is $m \times n$ and $\boldsymbol\theta$ is
228        $n \times 1$. What is the shape of $\mathbf{X}\boldsymbol\theta$, and why
229        is $\mathbf{X}^{\mathsf T}\mathbf{X}$ square?</p>
230        <p class="where">Lecture 2, first ten minutes.</p>
231      </li>
232      <li>
233        <p class="q">Differentiate $\lVert\mathbf{X}\boldsymbol\theta - y\rVert^2$
234        with respect to $\boldsymbol\theta$. Then say what has to be true of the
235        Hessian for the stationary point to be a minimum.</p>
236        <p class="where">Lecture 2. This is the derivation, not a preliminary to it.</p>
237      </li>
238      <li>
239        <p class="q">When is $\mathbf{X}^{\mathsf T}\mathbf{X}$ <em>not</em>
240        invertible? Answer in terms of the columns of $\mathbf{X}$.</p>
241        <p class="where">Lecture 2, and again in Lecture 5 when ridge repairs it.</p>
242      </li>
243      <li>
244        <p class="q">$A$ and $B$ each have variance $\sigma^2$ and correlation
245        $\rho$. What is $\operatorname{Var}\!\left(\tfrac{A+B}{2}\right)$?</p>
246        <p class="where">Lecture 7. This single calculation explains bagging,
247        random forests and extra-trees.</p>
248      </li>
249      <li>
250        <p class="q">What does a 95% confidence interval mean &mdash; and what
251        does it <em>not</em> mean?</p>
252        <p class="where">Lecture 2, on the final test score.</p>
253      </li>
254      <li>
255        <p class="q"><code>a</code> has shape <code>(100, 3)</code>,
256        <code>b</code> has shape <code>(3,)</code>. What does <code>a - b</code>
257        compute, and what shape comes out? Now <code>b</code> has shape
258        <code>(100,)</code> &mdash; what happens?</p>
259        <p class="where">Every lecture. Silent broadcasting is silent.</p>
260      </li>
261      <li>
262        <p class="q">Given a DataFrame <code>df</code>, select the rows where
263        <code>df["x"] &gt; 5</code>, keeping only columns <code>"a"</code> and
264        <code>"b"</code>. Then count how many values in <code>"a"</code> are
265        missing.</p>
266        <p class="where">Lecture 1, in the notebook.</p>
267      </li>
268      <li>
269        <p class="q">Read this traceback out loud and say which line of
270        <em>your</em> code caused it, and why:<br>
271        <code>ValueError: Input X contains NaN.</code></p>
272        <p class="where">Lecture 2, the first time preprocessing meets a
273        missing value.</p>
274      </li>
275    </ol>
276
277    <div class="panel panel-accent">
278      <strong>Could you do six or more?</strong> You are ready; skip the rest of
279      this page. <strong>Fewer?</strong> Find the matching row below. Nothing
280      here needs more than a few evenings, and none of it needs to be finished
281      before Lecture&nbsp;1 &mdash; the items are listed in the order the course
282      first needs them.
283    </div>
284
285    <h3>Filling the gaps</h3>
286
287    <h4 class="gap-head">Mathematics</h4>
288    <div class="tablewrap">
289      <table>
290        <thead>
291          <tr><th>What you need</th><th>First needed</th><th>If you are missing it</th></tr>
292        </thead>
293        <tbody>
294          <tr>
295            <td><strong>Linear algebra</strong><br>
296                <span class="muted">matrix products, transpose, inverse, rank,
297                column space, orthogonality, positive semi-definiteness, SVD</span></td>
298            <td class="num">L2</td>
299            <td>
300              <a href="https://www.3blue1brown.com/topics/linear-algebra">3Blue1Brown, <em>Essence of Linear Algebra</em></a>
301              for the geometric intuition — about three hours, and it is the
302              intuition the Lecture&nbsp;2 derivation assumes.<br>
303              For depth: <a href="https://ocw.mit.edu/courses/18-06-linear-algebra-spring-2010/">MIT 18.06 (Strang)</a>.
304            </td>
305          </tr>
306          <tr>
307            <td><strong>Multivariable calculus</strong><br>
308                <span class="muted">partial derivatives, gradients, the chain
309                rule, Hessians, convexity</span></td>
310            <td class="num">L2</td>
311            <td>
312              <a href="https://www.3blue1brown.com/topics/calculus">3Blue1Brown, <em>Essence of Calculus</em></a>,
313              then chapter 5 of
314              <a href="https://mml-book.github.io/"><em>Mathematics for Machine Learning</em></a>
315              (free PDF) for vector calculus in exactly the notation we use.
316            </td>
317          </tr>
318          <tr>
319            <td><strong>Probability and statistics</strong><br>
320                <span class="muted">expectation, variance, covariance,
321                correlation, independence, sampling, confidence intervals</span></td>
322            <td class="num">L2, L5, L7</td>
323            <td>
324              <a href="https://seeing-theory.brown.edu/">Seeing Theory</a> (Brown)
325              is a visual refresher in an afternoon.<br>
326              For depth: <a href="https://ocw.mit.edu/courses/18-05-introduction-to-probability-and-statistics-spring-2022/">MIT 18.05</a>,
327              or chapter 6 of <em>Mathematics for Machine Learning</em>.
328            </td>
329          </tr>
330        </tbody>
331      </table>
332    </div>
333    <p class="small muted">If you are on this degree programme you almost
334    certainly have all three. The self-check above is a faster way to confirm
335    that than reading the list.</p>
336
337    <h4 class="gap-head">Programming</h4>
338    <div class="tablewrap">
339      <table>
340        <thead>
341          <tr><th>What you need</th><th>First needed</th><th>If you are missing it</th></tr>
342        </thead>
343        <tbody>
344          <tr>
345            <td><strong>Python</strong><br>
346                <span class="muted">functions, lists and dicts, comprehensions,
347                imports, reading a traceback</span></td>
348            <td class="num">L1</td>
349            <td><a href="https://docs.python.org/3/tutorial/">The official Python tutorial</a>,
350                sections 3&ndash;6. A weekend from a standing start; an evening if
351                you know another language.</td>
352          </tr>
353          <tr>
354            <td><strong>NumPy</strong><br>
355                <span class="muted">arrays, shapes, indexing, axes, and
356                <em>broadcasting</em></span></td>
357            <td class="num">L1</td>
358            <td><a href="https://numpy.org/doc/stable/user/absolute_beginners.html">NumPy: the absolute basics</a>,
359                then <a href="https://numpy.org/doc/stable/user/basics.broadcasting.html">the broadcasting rules</a>
360                — read that second page twice. It is the single most common source
361                of code that runs and is wrong.</td>
362          </tr>
363          <tr>
364            <td><strong>pandas</strong><br>
365                <span class="muted">DataFrame, Series, selection, missing values</span></td>
366            <td class="num">L1</td>
367            <td><a href="https://pandas.pydata.org/docs/user_guide/10min.html">10 minutes to pandas</a>.
368                Optimistically named, but an hour genuinely does it.</td>
369          </tr>
370          <tr>
371            <td><strong>matplotlib</strong><br>
372                <span class="muted">enough to draw a histogram and a scatter plot</span></td>
373            <td class="num">L1</td>
374            <td><a href="https://matplotlib.org/stable/tutorials/pyplot.html">
374The pyplot tutorial</a>.
375                Twenty minutes; you will not need more than this in the whole course.</td>
376          </tr>
377          <tr>
378            <td><strong>Google Colab</strong><br>
379                <span class="muted">running cells, restarting the runtime,
380                reading a traceback</span></td>
381            <td class="num">L1</td>
382            <td><a href="https://colab.research.google.com/notebooks/intro.ipynb">Colab&rsquo;s own introduction</a>.
383                Ten minutes. Bring a Google account to the first lecture.</td>
384          </tr>
385          <tr>
386            <td><strong>scikit-learn</strong><br>
387                <span class="muted">the estimator API — <code>fit</code>,
388                <code>transform</code>, <code>predict</code></span></td>
389            <td class="num">L2</td>
390            <td class="taught">Helpful but <strong>not required</strong> — taught
391                from scratch in Lecture&nbsp;2. Lecture&nbsp;1 fits nothing at
392                all: looking properly at data before modelling it is not a
393                preliminary, it is what decides whether the model can work.
394                If you are curious:
395                <a href="https://scikit-learn.org/stable/getting_started.html">Getting started</a>.</td>
396          </tr>
397          <tr>
398            <td><strong>PyTorch</strong></td>
399            <td class="num">L10</td>
400            <td class="taught"><strong>Not required.</strong> Taught from first
401                principles in Lecture&nbsp;10 — tensors, autograd, and the
402                training loop written out in full before any of it is hidden
403                behind a helper. If you insist:
404                <a href="https://docs.pytorch.org/tutorials/beginner/basics/intro.html">the official basics</a>.</td>
405          </tr>
406        </tbody>
407      </table>
408    </div>
409
410    <h3>What you also need on the day</h3>
411    <ul class="plain">
412      <li>A laptop, and a <strong>Google account</strong> for Colab. Nothing is
413          installed locally; nothing depends on your operating system.</li>
414      <li>Access to an <strong>AI coding assistant</strong>. Which one is up to
415          you. The course does not require it &mdash; the notebooks are already
416          written &mdash; but Lecture&nbsp;1 sets out how to use one, and you
417          will need that long after this course.</li>
418      <li>Nothing else. <strong>Every notebook in the course runs on a free
419          CPU runtime</strong>, including the transformer and multimodal ones:
420          where a lecture&rsquo;s deck reports a result from a larger run, the
421          notebook says so and reproduces the ordering at a smaller scale.</li>
422    </ul>
423
424    <h3>What you explicitly do <em>not</em> need</h3>
425    <div class="grid">
426      <div class="card">
427        <h4>Prior machine learning</h4>
428        <p>Helpful, not assumed. The course starts from a problem and a dataset
429        on the first day and builds everything from there.</p>
430      </div>
431      <div class="card">
432        <h4>Deep learning experience</h4>
433        <p>Parts II&ndash;VI assume nothing beyond what Part&nbsp;I
434        established. Neural networks arrive in Lecture&nbsp;9, once the classical
435        models have been covered properly.</p>
436      </div>
437      <div class="card">
438        <h4>Software engineering</h4>
439        <p>
439No production systems, no deployment infrastructure, no build tooling.
440        Notebooks throughout.</p>
441      </div>
442      <div class="card">
443        <h4>Fluent typing</h4>
444        <p>The notebooks are written for you. What is examined is whether you can
445        <em>read</em> them &mdash; say what a cell does, what it measures, and
446        what would break if an argument changed.</p>
447      </div>
448    </div>
449  </div>
450</section>
451
452<!-- ===================================================================== -->
453
454<!-- timetable -->
455<!-- The one place in the course that names days. Everything else refers to
456     lectures by number, and tools/check_decks.py enforces that; this fence
457     lifts the ban for the section whose whole subject is the timetable. -->
458<section id="calendar">
459  <div class="wrap">
460    <h2>Calendar</h2>
461    <p class="lede">Twenty-four lectures on Tuesdays and Wednesdays, from
462    6 October 2026 to 12 January 2027.</p>
463
464    <dl class="facts-row">
465      <div>
466        <dt>Days</dt>
467        <dd>Tuesday and Wednesday</dd>
468      </div>
469      <div>
470        <dt>Hours</dt>
471        <dd>12:00&ndash;14:00</dd>
472      </div>
473      <div>
474        <dt>Room</dt>
475        <dd>Aula C &middot;
476            <a href="https://www.mat.uniroma1.it/">Dipartimento di Matematica
477            &ldquo;Guido Castelnuovo&rdquo;</a></dd>
478      </div>
479    </dl>
480
481    <p class="cal-next" hidden></p>
482
483    <div class="tablewrap">
484      <table class="calendar">
485        <caption>Every lecture in date order, and the days inside the term that
486        carry none.</caption>
487        <tbody>
488      <!-- BEGIN CALENDAR -->
489      <tr class="cal-month">
490        <th colspan="3" scope="colgroup">October 2026</th>
491      </tr>
492      <tr data-date="2026-10-06">
493        <td class="cal-when"><time datetime="2026-10-06">Tue 6</time></td>
494        <td class="cal-n">01</td>
495        <td class="cal-what">What machine learning is, and how we will work</td>
496      </tr>
497      <tr data-date="2026-10-07">
498        <td class="cal-when"><time datetime="2026-10-07">Wed 7</time></td>
499        <td class="cal-n">02</td>
500        <td class="cal-what">The end-to-end project</td>
501      </tr>
502      <tr data-date="2026-10-13">
503        <td class="cal-when"><time datetime="2026-10-13">Tue 13</time></td>
504        <td class="cal-n">03</td>
505        <td class="cal-what">Classification and its metrics</td>
506      </tr>
507      <tr data-date="2026-10-14">
508        <td class="cal-when"><time datetime="2026-10-14">Wed 14</time></td>
509        <td class="cal-n">04</td>
510        <td class="cal-what">Training models</td>
511      </tr>
512      <tr data-date="2026-10-20">
513        <td class="cal-when"><time datetime="2026-10-20">Tue 20</time></td>
514        <td class="cal-n">05</td>
515        <td class="cal-what">Regularisation and the bias–variance trade-off</td>
516      </tr>
517      <tr data-date="2026-10-21">
518        <td class="cal-when"><time datetime="2026-10-21">Wed 21</time></td>
519        <td class="cal-n">06</td>
520        <td class="cal-what">Decision trees</td>
521      </tr>
522      <tr data-date="2026-10-27">
523        <td class="cal-when"><time datetime="2026-10-27">Tue 27</time></td>
524        <td class="cal-n">07</td>
525        <td class="cal-what">Ensembles and random forests</td>
526      </tr>
527      <tr data-date="2026-10-28">
528        <td class="cal-when"><time datetime="2026-10-28">Wed 28</time></td>
529        <td class="cal-n">08</td>
530        <td class="cal-what">Dimensionality reduction and unsupervised learning</td>
531      </tr>
532      <tr class="cal-month">
533        <th colspan="3" scope="colgroup">November 2026</th>
534      </tr>
535      <tr data-date="2026-11-03">
536        <td class="cal-when"><time datetime="2026-11-03">Tue 3</time></td>
537        <td class="cal-n">09</td>
538        <td class="cal-what">Neural networks, from the perceptron up</td>
539      </tr>
540      <tr data-date="2026-11-04">
541        <td class="cal-when"><time datetime="2026-11-04">Wed 4</time></td>
542        <td class="cal-n">10</td>
543        <td class="cal-what">PyTorch</td>
544      </tr>
545      <tr data-date="2026-11-10">
546        <td class="cal-when"><time datetime="2026-11-10">Tue 10</time></td>
547        <td class="cal-n">11</td>
548        <td class="cal-what">Training deep networks</td>
549      </tr>
550      <tr data-date="2026-11-11">
551        <td class="cal-when"><time datetime="2026-11-11">Wed 11</time></td>
552        <td class="cal-n">12</td>
553        <td class="cal-what">Convolutional networks</td>
554      </tr>
555      <tr data-date="2026-11-17">
556        <td class="cal-when"><time datetime="2026-11-17">Tue 17</time></td>
557        <td class="cal-n">13</td>
558        <td class="cal-what">Transfer learning</td>
559      </tr>
560      <tr data-date="2026-11-18">
561        <td class="cal-when"><time datetime="2026-11-18">Wed 18</time></td>
562        <td class="cal-n">14</td>
563        <td class="cal-what">Detection and segmentation</td>
564      </tr>
565      <tr data-date="2026-11-24">
566        <td class="cal-when"><time datetime="2026-11-24">Tue 24</time></td>
567        <td class="cal-n">15</td>
568        <td class="cal-what">Time series</td>
569      </tr>
570      <tr data-date="2026-11-25">
571        <td class="cal-when"><time datetime="2026-11-25">Wed 25</time></td>
572        <td class="cal-n">16</td>
573        <td class="cal-what">Recurrent networks</td>
574      </tr>
575      <tr class="cal-month">
576        <th colspan="3" scope="colgroup">December 2026</th>
577      </tr>
578      <tr data-date="2026-12-01">
579        <td class="cal-when"><time datetime="2026-12-01">Tue 1</time></td>
580        <td class="cal-n">17</td>
581        <td class="cal-what">Text</td>
582      </tr>
583      <tr data-date="2026-12-02">
584        <td class="cal-when"><time datetime="2026-12-02">Wed 2</time></td>
585        <td class="cal-n">18</td>
586        <td class="cal-what">Attention and transformers</td>
587      </tr>
588      <tr class="cal-off">
589        <td class="cal-when"><time datetime="2026-12-08">Tue 8</time></td>
590        <td class="cal-n">&mdash;</td>
591        <td class="cal-what">Immacolata &mdash; the university is closed</td>
592      </tr>
593      <tr data-date="2026-12-09">
594        <td class="cal-when"><time datetime="2026-12-09">Wed 9</time></td>
595        <td class="cal-n">19</td>
596        <td class="cal-what">Information retrieval: the lexical foundation</td>
597      </tr>
598      <tr data-date="2026-12-15">
599        <td class="cal-when"><time datetime="2026-12-15">Tue 15</time></td>
600        <td class="cal-n">20</td>
601        <td class="cal-what">Information retrieval: dense retrieval</td>
602      </tr>
603      <tr data-date="2026-12-16">
604        <td class="cal-when"><time datetime="2026-12-16">Wed 16</time></td>
605        <td class="cal-n">21</td>
606        <td class="cal-what">Recommender systems: from ratings to factors</td>
607      </tr>
608      <tr data-date="2026-12-22">
609        <td class="cal-when"><time datetime="2026-12-22">Tue 22</time></td>
610        <td class="cal-n">22</td>
611        <td class="cal-what">Recommender systems: neural, and evaluated honestly</td>
612      </tr>
613      <tr data-date="2026-12-23">
614        <td class="cal-when"><time datetime="2026-12-23">Wed 23</time></td>
615        <td class="cal-n">23</td>
616        <td class="cal-what">Vision transformers and multimodal retrieval</td>
617      </tr>
618      <tr class="cal-off">
619        <td class="cal-when">24 Dec &ndash; 6 Jan</td>
620        <td class="cal-n">&mdash;</td>
621        <td class="cal-what">Christmas break</td>
622      </tr>
623      <tr class="cal-month">
624        <th colspan="3" scope="colgroup">January 2027</th>
625      </tr>
626      <tr data-date="2027-01-12">
627        <td class="cal-when"><time datetime="2027-01-12">Tue 12</time></td>
628        <td class="cal-n">24</td>
629        <td class="cal-what">Generation, retrieval-augmented systems, and where this leaves you</td>
630      </tr>      <!-- END CALENDAR -->
631        </tbody>
632      </table>
633    </div>
634
635    <p class="panel"><strong>The material is not published in advance.</strong>
636    A lecture&rsquo;s slides, notebook and notes appear on this page at
637    <strong>11:30</strong> on the day that lecture is taught, half an hour
638    before it begins, and stay there for the rest of the course.</p>
639  </div>
640</section>
641<!-- /timetable -->
642
643<!-- ===================================================================== -->
644
645<section id="lectures">
646  <div class="wrap">
647    <h2>Lectures</h2>
648    <p class="lede">Twenty-four lectures in six parts. Each lecture&rsquo;s
649    slides, notebook and notes appear on its own card at 11:30 on the day it is
650    taught.</p>
651
652    <noscript>
653      <p class="panel"><strong>This page needs JavaScript to show the
654      material.</strong> A lecture&rsquo;s slides, notebook and notes are added
655      to its card by a script, on the day of that lecture. With JavaScript
656      turned off the cards stay shut.</p>
657    </noscript>
658
659    <!-- BEGIN LECTURES -->
660    <div class="part-head">
661      <h3>Part I &mdash; Tabular data and classical models</h3>
662      <span class="part-meta">Lectures 1–8 &middot; Chapters 1–8 &middot; runs on CPU</span>
663    </div>
664    <ol class="lectures">
665      <li class="lecture" data-n="01" data-reveal="2026-10-06T11:30:00+02:00" data-material="slides,pdf,notebook,notes">
666        <span class="n">01</span>
667        <div class="body">
668          <p class="t">What machine learning is, and how we will work</p>
669          <p class="meta">California housing <span class="badge badge-ch">Ch 1–2</span></p>
670          <p class="when"><time datetime="2026-10-06">Tue 6 Oct 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
671        </div>
672        <div class="links">
673          <span class="btn-locked">Opens Tue 6 Oct, 11:30</span>
674        </div>
675      </li>
676      <li class="lecture" data-n="02" data-reveal="2026-10-07T11:30:00+02:00" data-material="slides,pdf,notebook,notes">
677        <span class="n">02</span>
678        <div class="body">
679          <p class="t">The end-to-end project</p>
680          <p class="meta">California housing <span class="badge badge-ch">Ch 2</span></p>
681          <p class="when"><time datetime="2026-10-07">Wed 7 Oct 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
682          <p class="thread">Derivation &middot; Least squares and the normal equation</p>
683        </div>
684        <div class="links">
685          <span class="btn-locked">Opens Wed 7 Oct, 11:30</span>
686        </div>
687      </li>
688      <li class="lecture" data-n="03" data-reveal="2026-10-13T11:30:00+02:00" data-material="slides,pdf,notebook,notes">
689        <span class="n">03</span>
690        <div class="body">
691          <p class="t">Classification and its metrics</p>
692          <p class="meta">MNIST <span class="badge badge-ch">Ch 3</span></p>
693          <p class="when"><time datetime="2026-10-13">Tue 13 Oct 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
694          <p class="thread">Derivation &middot; Imbalance, and the non-monotonicity of precision</p>
695        </div>
696        <div class="links">
697          <span class="btn-locked">Opens Tue 13 Oct, 11:30</span>
698        </div>
699      </li>
700      <li class="lecture" data-n="04" data-reveal="2026-10-14T11:30:00+02:00" data-material="slides,pdf,notebook,notes">
701        <span class="n">04</span>
702        <div class="body">
703          <p class="t">Training models</p>
704          <p class="meta">Titanic <span class="badge badge-ch">Ch 4</span></p>
705          <p class="when"><time datetime="2026-10-14">Wed 14 Oct 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
706          <p class="thread">Derivation &middot; Gradient descent</p>
707        </div>
708        <div class="links">
709          <span class="btn-locked">Opens Wed 14 Oct, 11:30</span>
710        </div>
711      </li>
712      <li class="lecture" data-n="05" data-reveal="2026-10-20T11:30:00+02:00" data-material="slides,pdf,notebook,notes">
713        <span class="n">05</span>
714        <div class="body">
715          <p class="t">Regularisation and the bias–variance trade-off</p>
716          <p class="meta">Titanic <span class="badge badge-ch">Ch 4</span></p>
717          <p class="when"><time datetime="2026-10-20">Tue 20 Oct 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
718          <p class="thread">Derivation &middot; The bias–variance decomposition</p>
719        </div>
720        <div class="links">
721          <span class="btn-locked">Opens Tue 20 Oct, 11:30</span>
722        </div>
723      </li>
724      <li class="lecture" data-n="06" data-reveal="2026-10-21T11:30:00+02:00" data-material="slides,pdf,notebook,notes">
725        <span class="n">06</span>
726        <div class="body">
727          <p class="t">Decision trees</p>
728          <p class="meta">CoverType <span class="badge badge-ch">Ch 5</span></p>
729          <p class="when"><time datetime="2026-10-21">Wed 21 Oct 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
730          <p class="thread">Derivation &middot; Impurity: Gini and entropy</p>
731        </div>
732        <div class="links">
733          <span class="btn-locked">Opens Wed 21 Oct, 11:30</span>
734        </div>
735      </li>
736      <li class="lecture" data-n="07" data-reveal="2026-10-27T11:30:00+01:00" data-material="slides,pdf,notebook,notes">
737        <span class="n">07</span>
738        <div class="body">
739          <p class="t">Ensembles and random forests</p>
740          <p class="meta">CoverType <span class="badge badge-ch">Ch 6</span></p>
741          <p class="when"><time datetime="2026-10-27">Tue 27 Oct 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
742          <p class="thread">Derivation &middot; The variance of an average of correlated predictors</p>
743        </div>
744        <div class="links">
745          <span class="btn-locked">Opens Tue 27 Oct, 11:30</span>
746        </div>
747      </li>
748      <li class="lecture" data-n="08" data-reveal="2026-10-28T11:30:00+01:00" data-material="slides,pdf,notebook,notes">
749        <span class="n">08</span>
750        <div class="body">
751          <p class="t">Dimensionality reduction and unsupervised learning</p>
752          <p class="meta">Olivetti faces <span class="badge badge-ch">Ch 7–8</span></p>
753          <p class="when"><time datetime="2026-10-28">Wed 28 Oct 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
754          <p class="thread">Derivation &middot; PCA via the SVD; Johnson–Lindenstrauss</p>
755        </div>
756        <div class="links">
757          <span class="btn-locked">Opens Wed 28 Oct, 11:30</span>
758        </div>
759      </li>
760    </ol>
761    <div class="part-head">
762      <h3>Part II &mdash; Neural networks</h3>
763      <span class="part-meta">Lectures 9–11 &middot; Chapters 9–11 &middot; runs on CPU</span>
764    </div>
765    <ol class="lectures">
766      <li class="lecture" data-n="09" data-reveal="2026-11-03T11:30:00+01:00" data-material="slides,pdf,notebook,notes">
767        <span class="n">09</span>
768        <div class="body">
769          <p class="t">Neural networks, from the perceptron up</p>
770          <p class="meta">Fashion-MNIST <span class="badge badge-ch">Ch 9</span></p>
771          <p class="when"><time datetime="2026-11-03">Tue 3 Nov 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
772          <p class="thread">Derivation &middot; What a layer computes</p>
773        </div>
774        <div class="links">
775          <span class="btn-locked">Opens Tue 3 Nov, 11:30</span>
776        </div>
777      </li>
778      <li class="lecture" data-n="10" data-reveal="2026-11-04T11:30:00+01:00" data-material="slides,pdf,notebook,notes">
779        <span class="n">10</span>
780        <div class="body">
781          <p class="t">PyTorch</p>
782          <p class="meta">Fashion-MNIST <span class="badge badge-ch">Ch 10</span></p>
783          <p class="when"><time datetime="2026-11-04">Wed 4 Nov 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
784          <p class="thread">Derivation &middot; Backpropagation as reverse-mode automatic differentiation</p>
785        </div>
786        <div class="links">
787          <span class="btn-locked">Opens Wed 4 Nov, 11:30</span>
788        </div>
789      </li>
790      <li class="lecture" data-n="11" data-reveal="2026-11-10T11:30:00+01:00" data-material="slides,pdf,notebook,notes">
791        <span class="n">11</span>
792        <div class="body">
793          <p class="t">Training deep networks</p>
794          <p class="meta">CIFAR-10 <span class="badge badge-ch">Ch 11</span></p>
795          <p class="when"><time datetime="2026-11-10">Tue 10 Nov 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
796          <p class="thread">Derivation &middot; Variance propagation and weight initialisation</p>
797        </div>
798        <div class="links">
799          <span class="btn-locked">Opens Tue 10 Nov, 11:30</span>
800        </div>
801      </li>
802    </ol>
803    <div class="part-head">
804      <h3>Part III &mdash; Computer vision</h3>
805      <span class="part-meta">Lectures 12–14 &middot; Chapter 12 &middot; runs on CPU</span>
806    </div>
807    <ol class="lectures">
808      <li class="lecture" data-n="12" data-reveal="2026-11-11T11:30:00+01:00" data-material="slides,pdf,notebook,notes">
809        <span class="n">12</span>
810        <div class="body">
811          <p class="t">Convolutional networks</p>
812          <p class="meta">Flowers102 <span class="badge badge-ch">Ch 12</span></p>
813          <p class="when"><time datetime="2026-11-11">Wed 11 Nov 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
814          <p class="thread">Derivation &middot; Weight sharing, equivariance and memory</p>
815        </div>
816        <div class="links">
817          <span class="btn-locked">Opens Wed 11 Nov, 11:30</span>
818        </div>
819      </li>
820      <li class="lecture" data-n="13" data-reveal="2026-11-17T11:30:00+01:00" data-material="slides,pdf,notebook,notes">
821        <span class="n">13</span>
822        <div class="body">
823          <p class="t">Transfer learning</p>
824          <p class="meta">Flowers102 <span class="badge badge-ch">Ch 12</span></p>
825          <p class="when"><time datetime="2026-11-17">Tue 17 Nov 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
826        </div>
827        <div class="links">
828          <span class="btn-locked">Opens Tue 17 Nov, 11:30</span>
829        </div>
830      </li>
831      <li class="lecture" data-n="14" data-reveal="2026-11-18T11:30:00+01:00" data-material="slides,pdf,notebook,notes">
832        <span class="n">14</span>
833        <div class="body">
834          <p class="t">Detection and segmentation</p>
835          <p class="meta">COCO <span class="badge badge-ch">Ch 12</span></p>
836          <p class="when"><time datetime="2026-11-18">Wed 18 Nov 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
837          <p class="thread">Derivation &middot; IoU’s vanishing gradient; mAP</p>
838        </div>
839        <div class="links">
840          <span class="btn-locked">Opens Wed 18 Nov, 11:30</span>
841        </div>
842      </li>
843    </ol>
844    <div class="part-head">
845      <h3>Part IV &mdash; Sequences and language</h3>
846      <span class="part-meta">Lectures 15–18 &middot; Chapters 13–15 &middot; runs on CPU</span>
847    </div>
848    <ol class="lectures">
849      <li class="lecture" data-n="15" data-reveal="2026-11-24T11:30:00+01:00" data-material="slides,pdf,notebook,notes">
850        <span class="n">15</span>
851        <div class="body">
852          <p class="t">Time series</p>
853          <p class="meta">Chicago transit ridership <span class="badge badge-ch">Ch 13</span></p>
854          <p class="when"><time datetime="2026-11-24">Tue 24 Nov 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
855          <p class="thread">Derivation &middot; Stationarity, differencing and autocorrelation</p>
856        </div>
857        <div class="links">
858          <span class="btn-locked">Opens Tue 24 Nov, 11:30</span>
859        </div>
860      </li>
861      <li class="lecture" data-n="16" data-reveal="2026-11-25T11:30:00+01:00" data-material="slides,pdf,notebook,notes">
862        <span class="n">16</span>
863        <div class="body">
864          <p class="t">Recurrent networks</p>
865          <p class="meta">Chicago transit ridership <span class="badge badge-ch">Ch 13</span></p>
866          <p class="when"><time datetime="2026-11-25">Wed 25 Nov 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
867        </div>
868        <div class="links">
869          <span class="btn-locked">Opens Wed 25 Nov, 11:30</span>
870        </div>
871      </li>
872      <li class="lecture" data-n="17" data-reveal="2026-12-01T11:30:00+01:00" data-material="slides,pdf,notebook,notes">
873        <span class="n">17</span>
874        <div class="body">
875          <p class="t">Text</p>
876          <p class="meta">IMDb <span class="badge badge-ch">Ch 14</span></p>
877          <p class="when"><time datetime="2026-12-01">Tue 1 Dec 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
878          <p class="thread">Derivation &middot; Softmax, cross-entropy and logits</p>
879        </div>
880        <div class="links">
881          <span class="btn-locked">Opens Tue 1 Dec, 11:30</span>
882        </div>
883      </li>
884      <li class="lecture" data-n="18" data-reveal="2026-12-02T11:30:00+01:00" data-material="slides,pdf,notebook,notes">
885        <span class="n">18</span>
886        <div class="body">
887          <p class="t">Attention and transformers</p>
888          <p class="meta">IMDb <span class="badge badge-ch">Ch 14–15</span></p>
889          <p class="when"><time datetime="2026-12-02">Wed 2 Dec 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
890          <p class="thread">Derivation &middot; Scaled dot-product attention</p>
891        </div>
892        <div class="links">
893          <span class="btn-locked">Opens Wed 2 Dec, 11:30</span>
894        </div>
895      </li>
896    </ol>
897    <div class="part-head">
898      <h3>Part V &mdash; Information retrieval and recommender systems</h3>
899      <span class="part-meta">Lectures 19–22 &middot; Outside the book &middot; examinable</span>
900    </div>
901    <ol class="lectures">
902      <li class="lecture" data-n="19" data-reveal="2026-12-09T11:30:00+01:00" data-material="slides,pdf,notebook,notes">
903        <span class="n">19</span>
904        <div class="body">
905          <p class="t">Information retrieval: the lexical foundation</p>
906          <p class="meta">SciFact (BEIR) <span class="badge badge-math">Outside the book</span></p>
907          <p class="when"><time datetime="2026-12-09">Wed 9 Dec 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
908          <p class="thread">Derivation &middot; Evaluating a ranking: MRR, AP, NDCG</p>
909        </div>
910        <div class="links">
911          <span class="btn-locked">Opens Wed 9 Dec, 11:30</span>
912        </div>
913      </li>
914      <li class="lecture" data-n="20" data-reveal="2026-12-15T11:30:00+01:00" data-material="slides,pdf,notebook,notes">
915        <span class="n">20</span>
916        <div class="body">
917          <p class="t">Information retrieval: dense retrieval</p>
918          <p class="meta">SciFact (BEIR) <span class="badge badge-math">Outside the book</span></p>
919          <p class="when"><time datetime="2026-12-15">Tue 15 Dec 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
920        </div>
921        <div class="links">
922          <span class="btn-locked">Opens Tue 15 Dec, 11:30</span>
923        </div>
924      </li>
925      <li class="lecture" data-n="21" data-reveal="2026-12-16T11:30:00+01:00" data-material="slides,pdf,notebook,notes">
926        <span class="n">21</span>
927        <div class="body">
928          <p class="t">Recommender systems: from ratings to factors</p>
929          <p class="meta">MovieLens <span class="badge badge-math">Outside the book</span></p>
930          <p class="when"><time datetime="2026-12-16">Wed 16 Dec 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
931          <p class="thread">Derivation &middot; Matrix factorisation, and its relation to the SVD</p>
932        </div>
933        <div class="links">
934          <span class="btn-locked">Opens Wed 16 Dec, 11:30</span>
935        </div>
936      </li>
937      <li class="lecture" data-n="22" data-reveal="2026-12-22T11:30:00+01:00" data-material="slides,pdf,notebook,notes">
938        <span class="n">22</span>
939        <div class="body">
940          <p class="t">Recommender systems: neural, and evaluated honestly</p>
941          <p class="meta">MovieLens <span class="badge badge-math">Outside the book</span></p>
942          <p class="when"><time datetime="2026-12-22">Tue 22 Dec 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
943        </div>
944        <div class="links">
945          <span class="btn-locked">Opens Tue 22 Dec, 11:30</span>
946        </div>
947      </li>
948    </ol>
949    <div class="part-head">
950      <h3>Part VI &mdash; Multimodal models, and closing the course</h3>
951      <span class="part-meta">Lectures 23–24 &middot; Chapters 15–16 &middot; runs on CPU</span>
952    </div>
953    <ol class="lectures">
954      <li class="lecture" data-n="23" data-reveal="2026-12-23T11:30:00+01:00" data-material="slides,pdf,notebook,notes">
955        <span class="n">23</span>
956        <div class="body">
957          <p class="t">Vision transformers and multimodal retrieval</p>
958          <p class="meta">COCO <span class="badge badge-ch">Ch 15–16</span></p>
959          <p class="when"><time datetime="2026-12-23">Wed 23 Dec 2026</time> &middot; 12:00–14:00 &middot; Aula C</p>
960          <p class="thread">Derivation &middot; The contrastive objective and its temperature</p>
961        </div>
962        <div class="links">
963          <span class="btn-locked">Opens Wed 23 Dec, 11:30</span>
964        </div>
965      </li>
966      <li class="lecture" data-n="24" data-reveal="2027-01-12T11:30:00+01:00" data-material="slides,pdf,notebook,notes">
967        <span class="n">24</span>
968        <div class="body">
969          <p class="t">Generation, retrieval-augmented systems, and where this leaves you</p>
970          <p class="meta">COCO and the Part V corpora <span class="badge badge-ch">Ch 15–16</span></p>
971          <p class="when"><time datetime="2027-01-12">Tue 12 Jan 2027</time> &middot; 12:00–14:00 &middot; Aula C</p>
972        </div>
973        <div class="links">
974          <span class="btn-locked">Opens Tue 12 Jan, 11:30</span>
975        </div>
976      </li>
977    </ol>    <!-- END LECTURES -->
978  </div>
979</section>
980
981<!-- ===================================================================== -->
982
983<section id="derivations">
984  <div class="wrap">
985    <h2>The mathematics</h2>
986    <p class="lede">Each lecture derives the one object its method rests on. Not
987    a parallel theory course: every derivation below does visible work on the
988    method taught beside it, and they are 40% of the written paper.</p>
989
990    <ol class="threads">
991    <!-- BEGIN DERIVATIONS -->
992      <li>Least squares and the normal equation <span class="where">&middot; Lecture 2</span></li>
993      <li>Imbalance, and the non-monotonicity of precision <span class="where">&middot; Lecture 3</span></li>
994      <li>Gradient descent <span class="where">&middot; Lecture 4</span></li>
995      <li>The bias–variance decomposition <span class="where">&middot; Lecture 5</span></li>
996      <li>Impurity: Gini and entropy <span class="where">&middot; Lecture 6</span></li>
997      <li>The variance of an average of correlated predictors <span class="where">&middot; Lecture 7</span></li>
998      <li>PCA via the SVD; Johnson–Lindenstrauss <span class="where">&middot; Lecture 8</span></li>
999      <li>What a layer computes <span class="where">&middot; Lecture 9</span></li>
1000      <li>Backpropagation as reverse-mode automatic differentiation <span class="where">&middot; Lecture 10</span></li>
1001      <li>Variance propagation and weight initialisation <span class="where">&middot; Lecture 11</span></li>
1002      <li>Weight sharing, equivariance and memory <span class="where">&middot; Lecture 12</span></li>
1003      <li>IoU’s vanishing gradient; mAP <span class="where">&middot; Lecture 14</span></li>
1004      <li>Stationarity, differencing and autocorrelation <span class="where">&middot; Lecture 15</span></li>
1005      <li>Softmax, cross-entropy and logits <span class="where">&middot; Lecture 17</span></li>
1006      <li>Scaled dot-product attention <span class="where">&middot; Lecture 18</span></li>
1007      <li>Evaluating a ranking: MRR, AP, NDCG <span class="where">&middot; Lecture 19</span></li>
1008      <li>Matrix factorisation, and its relation to the SVD <span class="where">&middot; Lecture 21</span></li>
1009      <li>The contrastive objective and its temperature <span class="where">&middot;
1009 Lecture 23</span></li>    <!-- END DERIVATIONS -->
1010    </ol>
1011
1012    <div class="panel">
1013      The derivations are cross-referential and the order matters. Lecture&nbsp;5
1014      completes Lecture&nbsp;2; Lecture&nbsp;14 uses Lecture&nbsp;3;
1015      Lecture&nbsp;21 and Lecture&nbsp;23 both use Lecture&nbsp;8; Lecture&nbsp;20
1016      and Lecture&nbsp;22 are the same architecture twice; Lecture&nbsp;23 is
1017      evaluated with Lecture&nbsp;19&rsquo;s metrics.
1018    </div>
1019  </div>
1020</section>
1021
1022<!-- ===================================================================== -->
1023
1024<section id="assessment">
1025  <div class="wrap">
1026    <h2>Assessment</h2>
1027    <p class="lede">The written examination carries the mark. The oral is
1028    optional, short, and can move that mark in either direction.</p>
1029
1030    <div class="grid">
1031      <div class="card">
1032        <h4>Written examination</h4>
1033        <p>Two hours, closed book, no formula sheet. Three parts: the
1034        derivations, choosing a method for a stated situation, and reading
1035        results &mdash; as technical exercises, closed choices with their
1036        reason, and short open answers.</p>
1037        <p class="muted">Marked out of 30, pass at 18. From the written
1038        alone, the mark recorded is capped at <strong>27</strong>.</p>
1039      </div>
1040      <div class="card">
1041        <h4>Oral examination &mdash; optional</h4>
1042        <p>Seven to ten minutes. <strong>Three questions, on any topic from
1043        the course.</strong></p>
1044        <p class="muted">No notes, no computer. Open to anyone who passed the
1045        written.</p>
1046      </div>
1047    </div>
1048
1049    <div class="panel">
1050      <span class="panel-title-inline">How the mark is made</span>
1051      <strong>The oral moves your written mark, up or down, and the result is
1052      final.</strong> <strong>28, 29 and 30 exist only through the oral</strong>,
1053      and lode is a separate decision, from 30.<br><br>
1054      Sitting it is <strong>your choice</strong>, made when you see your written
1055      mark &mdash; and <strong>binding once registered</strong>. Both are taken
1056      in the same session: a written pass does <strong>not</strong> carry
1057      forward.
1058    </div>
1059
1060    <div class="panel panel-accent">
1061      <strong>Why it is built this way.</strong> A grade you can improve at no
1062      risk is one everybody attempts, which turns the oral into a queue rather
1063      than an examination. Making it a real decision &mdash; three marks up,
1064      three marks down &mdash; means the people who sit it are the people with
1065      something to show. Nothing about it is hidden: the questions are published,
1066      the arithmetic is published, and neither changes after you have decided.
1067    </div>
1068
1069    <h3>Can a course this applied really have a written examination?</h3>
1070    <p>Yes &mdash; and this course is better suited to one than most applied
1071    courses, for a specific reason.</p>
1072    <div class="panel">
1073      The competence being built here is not typing. It is
1074      <strong>judgement about whether a result can be trusted</strong>, and that
1075      judgement is exercised by reading: reading code you did not write, reading a
1076      metric, reading a curve. Reading is paper-native. A written examination
1077      tests it <em>more</em> directly than a lab does &mdash; at a keyboard a
1078      student can arrive at the right answer by running things until they look
1079      right, whereas on paper they have to actually know.
1080    </div>
1081    <p>Three rules keep it honest.</p>
1082    <ol>
1083      <li><strong>No question tests API recall.</strong> Asking for the
1084          arguments of <code>train_test_split</code> would test what a
1085          docstring is for. Syntax errors in handwritten code cost nothing;
1086          logic errors cost everything.</li>
1087      <li><strong>No question is answerable by reciting a definition.</strong>
1088          Every one requires applying it to a stated situation.</li>
1089      <li><strong>The paper states what it needs.</strong> There is no formula
1090          sheet, and none is required: any definition or standard result a
1091          question depends on is printed in the question itself. What you are
1092          expected to supply is the reasoning, never the recall.</li>
1093    </ol>
1094    <p class="small muted">
1094This is why closed book costs you nothing here. Each
1095    of the eighteen derivations is one you have <em>performed</em> &mdash; and a
1096    result you can rebuild in three lines is not a result you need to have
1097    memorised.</p>
1098    <p>What a written examination cannot see is whether the reasoning on the
1099    page is yours. That is what the oral is for &mdash; and it is why it is
1100    short, and why it can go to any part of the course.</p>
1101
1102    <h3>The written examination</h3>
1103    <p>The three parts below are weighted <em>within the paper</em>. The paper
1104    carries your mark: up to 27 on its own, and up to 30 if you sit the oral.</p>
1105
1106    <p>Every paper mixes three <strong>forms</strong> of question, and the parts
1107    below say what each one is <em>about</em> rather than what shape it takes:
1108    <strong>technical exercises</strong> worked on the page &mdash; a
1109    derivation, a gradient, an arithmetic check on a stated table;
1110    <strong>closed questions</strong>, where you choose between two or three
1111    stated options; and <strong>open questions</strong> answered in a few
1112    sentences. A closed choice is rarely enough on its own: where a question
1113    asks for the reason as well as the choice, both are needed for full
1114    marks.</p>
1115    <div class="tablewrap">
1116      <table>
1117        <thead>
1118          <tr><th>Part</th><th>What it asks</th><th class="num">Weight<br><span class="th-sub">of the paper</span></th><th class="num">Time</th></tr>
1119        </thead>
1120        <tbody>
1121          <tr>
1122            <td><strong>A &middot; The derivations</strong></td>
1123            <td>Derive, state, apply. The eighteen objects developed in the
1124                lectures — the normal equation, the bias&ndash;variance
1125                decomposition, variance reduction by averaging, reverse-mode
1126                differentiation, cross-entropy, NDCG, and the rest.</td>
1127            <td class="num">40%</td>
1128            <td class="num">~48 min</td>
1129          </tr>
1130          <tr>
1131            <td><strong>B &middot; Choosing a method</strong></td>
1132            <td>A situation in four lines: this data, this constraint, this
1133                requirement. Which method, why that one rather than the obvious
1134                alternative, and what evidence would make you change your mind.
1135                <span class="muted">Some questions supply a short code excerpt and
1136                ask what it would take for its reported number to be
1137                trustworthy.</span></td>
1138            <td class="num">35%</td>
1139            <td class="num">~42 min</td>
1140          </tr>
1141          <tr>
1142            <td><strong>C &middot; Reading results</strong></td>
1143            <td>Plots and tables — learning curves, per-fold scores, a confusion
1144                matrix, a precision/recall curve, a ranking. Say what they show,
1145                what they do not show, and which way the number moves if you
1146                change the stated thing.</td>
1147            <td class="num">25%</td>
1148            <td class="num">~30 min</td>
1149          </tr>
1150        </tbody>
1151      </table>
1152    </div>
1153
1154    <div class="panel">
1155      <strong>Every exercise in the course, with its solution.</strong> Each
1156      deck ends with five questions in this style, answered on the following
1157      lecture&rsquo;s deck. All 120 exercises are collected in one place:
1158      <!-- BEGIN EXERCISES_PANEL -->
1159      published once the course has run.      <!-- END EXERCISES_PANEL -->
1160    </div>
1161
1162    <h3>Three specimen questions</h3>
1163    <p>One from each part, at the intended level and scale &mdash; each is
1164    calibrated to the time budget above.</p>
1165    <p class="small muted"><strong>Illustration only.</strong> These three show
1166    the kind of thing each part asks and roughly how much of it. They are not a
1167    template, not a syllabus, and not a promise: the questions on the paper you
1168    sit may differ in form, in topic, and in how they are put.</p>
1169
1170    <div class="specimen">
1171      <h4>Part A &mdash; derive, then apply</h4>
1172      <p>Two regressors each have variance $\sigma^2$ and are correlated with
1173      coefficient $\rho$.</p>
1174      <ol class="tight">
1175        <li>Show that the variance of their average is
1176            $\frac{\sigma^{2}(1+\rho)}{2}$.</li>
1177        <li>Bagging and a random forest differ in one respect. State it, and say
1178            which term in your expression it attacks.</li>
1179        <li>An ensemble of one hundred <em>identical</em> trees has $\rho = 1$.
1180            What does your expression predict, and is that the right answer?</li>
1181      </ol>
1182    </div>
1183
1184    <div class="specimen">
1185      <h4>Part B &mdash; choosing a method</h4>
1186      <p>A colleague has <strong>4,000 labelled rows</strong> and
1187      <strong>sixty numeric features</strong>, many of them near-duplicates of
1188      one another. Every prediction has to be justified to a regulator.</p>
1189      <ol class="tight">
1190        <li>Ordinary least squares, ridge, lasso, or a random forest &mdash;
1191            which do you fit first, and what specifically about the situation
1192            decides it?</li>
1193        <li>Name the one piece of evidence that would make you abandon that
1194            choice for one of the others.</li>
1195      </ol>
1196      <p>They send you this program, and the number it printed.</p>
1197<pre><code>best = None
1198for a in [0.01, 0.1, 1, 10, 100]:
1199    m = Ridge(alpha=a).fit(X_train, y_train)
1200    s = root_mean_squared_error(y_test, m.predict(X_test))
1201    if best is None or s &lt; best[1]:
1202        best = (a, s)
1203
1204print(f"best alpha={best[0]}, test RMSE={best[1]:,.0f}")</code></pre>
1205      <ol class="tight" start="3">
1206        <li>This program never creates a validation set. <strong>Which object is
1207            doing that job?</strong></li>
1208        <li>As an estimate of the model&rsquo;s error on new data, the printed
1209            RMSE is <strong>(i)</strong> too optimistic, <strong>(ii)</strong> too
1210            pessimistic, or <strong>(iii)</strong> neither. Choose one and give the
1211            reason in a sentence.</li>
1212        <li>Rewrite it so the printed number is honest. Two statements is
1213            enough.</li>
1214      </ol>
1215      <p class="small muted">Every part has a determinate answer &mdash; ridge,
1216      because the near-duplicate features make $\mathbf{X}^{\mathsf T}\mathbf{X}$
1217      ill-conditioned and the regulator rules out the forest; <em>the test
1218      set</em>; <em>(i)</em>; a cross-validated search on the training data,
1219      scored once on the test set. Nothing here rewards a well-phrased opinion,
1220      and none of it can be answered without knowing what a validation set is
1221      for.</p>
1222    </div>
1223
1224    <div class="specimen">
1225      <h4>Part C &mdash; reading results</h4>
1226      <p>Two teams report the following for the same dataset.</p>
1227      <div class="tablewrap">
1228        <table>
1229          <thead><tr><th>Team</th><th class="num">Training RMSE</th><th class="num">10-fold CV RMSE</th></tr></thead>
1230          <tbody>
1231            <tr><td>A</td><td class="num">$0</td><td class="num">$68,574</td></tr>
1232            <tr><td>B</td><td class="num">$68,233</td><td class="num">$68,282</td></tr>
1233          </tbody>
1234        </table>
1235      </div>
1236      <ol class="tight">
1237        <li>Name each pathology.</li>
1238        <li>Which model is more useful? Explain why that is not the same question
1239            as which is better fitted.</li>
1240        <li>One of these is helped by collecting more training data and the other
1241            is not. Which, and why?</li>
1242      </ol>
1243    </div>
1244
1245    <h3>The oral examination</h3>
1246    <p>Optional. Three questions, answered aloud, on any topic the course
1247    covered &mdash; there is no separate syllabus for it, and no list of
1248    questions to learn.</p>
1249
1250    <div class="cols-oral">
1251      <div class="card">
1252        <h4>What happens</h4>
1253        <ul class="plain">
1254          <li>You decide <strong>after</strong> seeing your written mark. Until
1255              then there is nothing to opt into.</li>
1256          <li>Three questions, on any topic from the course.</li>
1257          <li>Seven to ten minutes, no notes, no computer.</li>
1258        </ul>
1259      </div>
1260      <div class="card">
1261        <h4>What it is worth</h4>
1262        <ul class="plain">
1263          <li><strong>Up.</strong> The only route to a mark above 27, and to
1264              lode.</li>
1265          <li><strong>Down.</strong> The same conversation can lower it. This is
1266              not a formality.</li>
1267          <li><strong>Worth sitting from below the cap too</strong> &mdash; the
1268              movement is measured from <em>your</em> mark, wherever it is.</li>
1269        </ul>
1270      </div>
1271    </div>
1272    <p class="small muted">Nothing is collected and nothing is graded as an
1273    artefact. The notebooks are examinable as <em>experience</em>: you are
1274    expected to be able to account for what they do and why, without the code in
1275    front of you.</p>
1276    <div class="panel panel-accent">
1277      <strong>Rule one is checked twice.</strong> Never keep code you cannot
1278      explain. Part B of the paper puts a cell in front of every candidate; the
1279      oral puts a question in front of the ones who choose it. Both apply to
1280      code an assistant wrote for you exactly as they apply to code you typed.
1281    </div>
1282
1283    <p style="margin-top:1.75rem">Note what is absent: there are no marks for a
1284    notebook that runs. <strong>Evidence earns marks.</strong></p>
1285  </div>
1286</section>
1287
1288<!-- ===================================================================== -->
1289
1290<section id="textbook">
1291  <div class="wrap">
1292    <h2>Textbook and scope</h2>
1293    <p>Aurélien Géron, <em>Hands-On Machine Learning with Scikit-Learn and
1294    PyTorch</em>, O&rsquo;Reilly, 2025 &mdash; <strong>Chapters 1&ndash;16</strong>,
1295    covering Lectures 1&ndash;18 and 23&ndash;24.</p>
1296
1297    <p><strong>Lectures 19&ndash;22</strong> &mdash; information retrieval and
1298    recommender systems &mdash; sit outside the book and are taught from the
1299    lecture notes. They are examinable on the same terms as everything else,
1300    and for those four lectures <strong>the notes below are the primary
1301    source</strong> &mdash; not a supplement to a chapter, because there is no
1302    chapter.</p>
1303
1304    <p class="small">Written notes exist for <strong>every</strong> lecture
1305    &mdash; the <em>Notes (PDF)</em> button that appears on each card above on
1306    the day of its lecture. For the other twenty they set out the
1307    lecture&rsquo;s argument at length beside the chapter it is taught from;
1308    the four listed here are the ones with no chapter behind them.</p>
1309
1310    <div class="tablewrap">
1311      <table>
1312        <thead><tr><th>Extended lecture notes</th><th>Dataset</th></tr></thead>
1313        <tbody>
1314          <tr><td>19 &middot; Information retrieval: the lexical foundation</td>
1315              <td>SciFact (BEIR)</td></tr>
1316          <tr><td>20 &middot; Information retrieval: dense retrieval</td>
1317              <td>SciFact (BEIR)</td></tr>
1318          <tr><td>21 &middot; Recommender systems: from ratings to factors</td>
1319              <td>
1319MovieLens 1M</td></tr>
1320          <tr><td>22 &middot; Recommender systems: neural, and evaluated honestly</td>
1321              <td>MovieLens 1M</td></tr>
1322        </tbody>
1323      </table>
1324    </div>
1325
1326    <div class="panel">
1327      <strong>Scope discipline.</strong> The examinable surface is Chapters
1328      1&ndash;16 plus the notes for Lectures 19&ndash;22. Every section of every
1329      lecture is marked <em>examinable</em>, <em>not examinable &mdash;
1330      engineering</em>, or <em>beyond the syllabus, for context</em>, so you
1331      never have to guess.
1332    </div>
1333
1334    <h3>Chapter coverage</h3>
1335    <div class="tablewrap">
1336      <table>
1337        <thead><tr><th>Chapter</th><th>Lectures</th></tr></thead>
1338        <tbody>
1339          <tr><td>1 &middot; The machine learning landscape</td><td class="num">1</td></tr>
1340          <tr><td>2 &middot; End-to-end project</td><td class="num">1, 2</td></tr>
1341          <tr><td>3 &middot; Classification</td><td class="num">3</td></tr>
1342          <tr><td>4 &middot; Training models</td><td class="num">4, 5</td></tr>
1343          <tr><td>5 &middot; Decision trees</td><td class="num">6</td></tr>
1344          <tr><td>6 &middot; Ensembles and random forests</td><td class="num">7</td></tr>
1345          <tr><td>7 &middot; Dimensionality reduction</td><td class="num">8</td></tr>
1346          <tr><td>8 &middot; Unsupervised learning</td><td class="num">8</td></tr>
1347          <tr><td>9 &middot; Introduction to artificial neural networks</td><td class="num">9</td></tr>
1348          <tr><td>10 &middot; Building networks with PyTorch</td><td class="num">10</td></tr>
1349          <tr><td>11 &middot; Training deep networks</td><td class="num">11</td></tr>
1350          <tr><td>12 &middot; Deep computer vision</td><td class="num">12, 13, 14</td></tr>
1351          <tr><td>13 &middot; Sequences</td><td class="num">15, 16</td></tr>
1352          <tr><td>14 &middot; NLP with RNNs and attention</td><td class="num">17, 18</td></tr>
1353          <tr><td>15 &middot; Transformers</td><td class="num">18, 23, 24</td></tr>
1354          <tr><td>16 &middot; Vision and multimodal transformers</td><td class="num">23, 24</td></tr>
1355          <tr><td><em>Lecture notes</em> &middot; Information retrieval</td><td class="num">19, 20</td></tr>
1356          <tr><td><em>Lecture notes</em> &middot; Recommender systems</td><td class="num">21, 22</td></tr>
1357        </tbody>
1358      </table>
1359    </div>
1360  </div>
1361</section>
1362
1363<!-- ===================================================================== -->
1364
1365<section id="practicalities">
1366  <div class="wrap">
1367    <h2>Practicalities</h2>
1368    <ul class="plain">
1369      <li>Announcements go through
1370          <a href="https://classroom.google.com/c/MjU0MDA3MTk5MjJa">the course&rsquo;s Google Classroom</a>; join it
1371          with your Sapienza account.</li>
1372      <li>Every lecture has a Colab notebook, linked from its own slides and
1373          from the list above from the morning of that lecture.</li>
1374      <li>Every notebook runs on a free CPU runtime, and says at the top
1375          roughly how long it takes.</li>
1376      <li>Every deck has a <strong>PDF</strong> button beside it, one page per
1377          slide. Printing a deck from the browser gives the same thing.</li>
1378      <li>In a deck: <code>M</code> opens the menu, <code>S</code> opens the
1379          speaker notes, <code>C</code> toggles the chalkboard, <code>F</code>
1380          goes full screen.</li>
1381      <li>Lectures refer to one another by number, so the material does not
1382          depend on the timetable.</li>
1383      <li>Nothing in any notebook is wrong on purpose.</li>
1384    </ul>
1385  </div>
1386</section>
1387
1388<footer>
1389  <div class="wrap">
1390    <p class="draft">Work in progress &mdash; not yet announced to students</p>
1391    <p><strong>Applicazioni Informatiche del Machine Learning</strong> &middot;
1392       BSc Mathematics of Artificial Intelligence &middot; Sapienza Università di Roma</p>
1393    <p>Fabrizio Silvestri &middot;
1394       <a href="https://fabsilvestri.github.io/">fabsilvestri.github.io</a></p>
1395    <p>Slides built with <a href="https://revealjs.com/">reveal.js</a> and
1396       <a href="https://katex.org/">KaTeX</a>
1396, both vendored locally so the decks
1397       work without a network connection.</p>
1398  </div>
1399</footer>
1400
1401<!-- KaTeX, vendored locally so the page needs no network beyond this host -->
1402<script defer src="assets/js/site-nav.js"></script>
vendor: 1 bytes, line 1402
1402
1403<script defer src="assets/js/reveal-material.js"></script>
vendor: 1 bytes, line 1403
1403
1404<script defer src="lib/katex/dist/katex.min.js"></script>
vendor: 1 bytes, line 1404
1404
1405<script defer src="lib/katex/dist/contrib/auto-render.min.js"
1406        onload="renderMathInElement(document.body, {
1407          delimiters: [
1408            {left: '$$', right: '$$', display: true},
1409            {left: '$',  right: '$',  display: false}
1410          ],
1411          throwOnError: false
1412        });"></script>
1412
1413
1414</body>
1415</html>

Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.