1<!DOCTYPE html> 2<html lang="en"> 3<head> 4<meta charset="utf-8"> 5<meta name="viewport" content="width=device-width, initial-scale=1"> 6 <link rel="icon" href="assets/img/favicon.svg" type="image/svg+xml"> 7 <link rel="icon" href="assets/img/favicon-32.png" sizes="32x32" type="image/png"> 8 <link rel="apple-touch-icon" href="assets/img/favicon-180.png"> 9 <meta name="theme-color" content="#0b3d62"> 10<title>Applications of Machine Learning â BSc Mathematics of Artificial Intelligence</title> 11<meta name="description" content="Applicazioni Informatiche del Machine Learning â a 48-hour applied course: twenty-four lectures on how to build a machine learning system and how to know whether it works."> 12<meta name="robots" content="noindex, nofollow"> 13<link rel="stylesheet" href="lib/katex/dist/katex.min.css"> 14<link rel="stylesheet" href="assets/css/site.css"> 15</head> 16<body> 17 18<a class="skip-link" href="#about">Skip to content</a> 19 20<nav class="site-nav" aria-label="Sections of this page"> 21 <div class="wrap"> 22 <a class="brand" href="#top">Applications of ML</a> 23 <a class="nav-link" href="#about">Principle</a> 24 <a class="nav-link" href="#method">Method</a> 25 <a class="nav-link" href="#prerequisites">Prerequisites</a> 26 <a class="nav-link" href="#calendar">Calendar</a> 27 <a class="nav-link" href="#lectures">Lectures</a> 28 <a class="nav-link" href="#derivations">Mathematics</a> 29 <a class="nav-link" href="#assessment">Assessment</a> 30 <a class="nav-link" href="#textbook">Textbook</a> 31 <a class="nav-link" href="#practicalities">Practicalities</a> 32 <!-- BEGIN EXERCISES_NAV --> 33 <!-- Held back by EXERCISE_BOOK_PUBLIC in tools/make_site.py --> <!-- END EXERCISES_NAV --> 34 </div> 35</nav> 36 37<header class="hero" id="top"> 38 <div class="wrap"> 39 <p class="kicker">Sapienza Università di Roma · BSc Mathematics of Artificial Intelligence</p> 40 <h1>Applications of<br>Machine Learning</h1> 41 <p class="it">Applicazioni Informatiche del Machine Learning</p> 42 <p class="tagline"> 43 How to build a machine learning system, and how to know whether it 44 works. 45 </p> 46 <dl class="facts"> 47 <div> 48 <dt>Load</dt> 49 <dd>48 academic hours · 24 lectures</dd> 50 </div> 51 <div> 52 <dt>Year</dt> 53 <dd>Third year, BSc</dd> 54 </div> 55 <div> 56 <dt>When</dt> 57 <dd>Tue & Wed · 12:00–14:00 · from 6 Oct 2026</dd> 58 </div> 59 <div> 60 <dt>Where</dt> 61 <dd><a href="https://www.mat.uniroma1.it/">Aula C · Dip. di Matematica</a></dd> 62 </div> 63 <div> 64 <dt>Teaching language</dt> 65 <dd>Italian · materials in English</dd> 66 </div> 67 <div> 68 <dt>Lecturer</dt> 69 <dd>Fabrizio Silvestri</dd> 70 </div> 71 <div> 72 <dt>Google Classroom</dt> 73 <dd><a href="https://classroom.google.com/c/MjU0MDA3MTk5MjJa">Join the class</a></dd> 74 </div> 75 </dl> 76 </div> 77</header> 78 79<!-- ===================================================================== --> 80 81<section id="about"> 82 <div class="wrap"> 83 <h2>The organising principle</h2> 84 <p class="lede">One topic per lecture, and every lecture stands on its own.</p> 85 86 <div class="panel"> 87 <strong>Each lecture is the mathematics, the method, and a notebook that 88 implements it.</strong> If you miss a lecture, you read its slides and its 89 notebook and you are caught up — you never have to reconstruct 90 anything from the lecture before it. 91 </div> 92 93 <p>Ninety minutes, in four blocks.</p> 94 95 <div class="grid"> 96 <div class="card"> 97 <h4>The mathematics <span class="muted">· 20 min</span></h4> 98 <p>The one object the method rests on, derived rather than quoted 99 — because the method does not make sense without it, and because 100 the derivations are what you still have in five years when the libraries 101 have changed.</p> 102 </div> 103 <div class="card"> 104 <h4>The method <span class="muted">· 40 min</span></h4> 105 <p>
105How it works, what its hyperparameters do, the conditions under which 106 it fails, and a worked example with real numbers from real data — 107 every figure on every slide reproduced by a script in this repository.</p> 108 </div> 109 <div class="card"> 110 <h4>Further ground <span class="muted">· 15 min</span></h4> 111 <p>The variants, when to prefer each, and what practitioners actually 112 reach for — which is not always what the textbook presents first.</p> 113 </div> 114 <div class="card"> 115 <h4>The notebook <span class="muted">· 5 min</span></h4> 116 <p>What is in it, what to run, and what to change to see something move. 117 You run it yourself afterwards; the lecture is not a live coding 118 session.</p> 119 </div> 120 </div> 121 122 <p style="margin-top:1.5rem">Ten minutes at the top of each lecture place it 123 in the arc: what we can already do, what we cannot, and which of the two 124 today’s method changes.</p> 125 </div> 126</section> 127 128<!-- ===================================================================== --> 129 130<section id="method"> 131 <div class="wrap"> 132 <h2>Working method</h2> 133 <p class="lede">You will write machine learning code with an assistant 134 — if not in this course, then in the thesis after it and in the job 135 after that. So the course is explicit about how.</p> 136 137 <div class="tablewrap"> 138 <table> 139 <thead> 140 <tr><th>#</th><th>Step</th><th>Who</th></tr> 141 </thead> 142 <tbody> 143 <tr><td class="num">1</td><td><strong>Specify</strong> — the input, the output, the constraint, the check</td><td>you</td></tr> 144 <tr><td class="num">2</td><td><strong>Generate</strong> — and then stop, before running anything</td><td>the assistant</td></tr> 145 <tr><td class="num">3</td><td><strong>Read</strong> — as a reviewer, not as an author</td><td>you</td></tr> 146 <tr><td class="num">4</td><td><strong>Test</strong> — against a case whose answer you already know</td><td>you</td></tr> 147 <tr><td class="num">5</td><td><strong>Verify</strong> — that the number means what it appears to mean</td><td>you</td></tr> 148 </tbody> 149 </table> 150 </div> 151 152 <p>Steps 3–5 are the course. Step 2 is the part that is free.</p> 153 154 <div class="panel"> 155 <strong>Where you meet this loop here.</strong> The notebooks in this 156 course are written for you and they are correct — nothing in them is 157 wrong on purpose. Every code cell is preceded by the specification that 158 would produce it: the box is step 1, the cell below it is what 159 step 2 returned, and its <em>check</em> line is step 4 written 160 down in advance. Read the box, answer the check in your head, then run the 161 cell. 162 </div> 163 164 <h3>Why the emphasis on verification</h3> 165 <p>Machine learning fails <em>silently</em>. A bug in a training loop does not 166 crash — it returns a plausible number:</p> 167 <ul class="plain"> 168 <li>A scaler fitted before the split leaks the test set into training. Great score, broken system.</li> 169 <li>A missing <code>model.eval()</code> leaves dropout active during evaluation.</li> 170 <li>A missing <code>optimizer.zero_grad()</code> accumulates gradients with no error.</li> 171 <li>A metric averaged per batch rather than over the set is wrong by construction.</li> 172 <li>Accuracy on an imbalanced set reports 90% for a model that predicts one class.</li> 173 </ul> 174 <p>Every one of these runs. Every one produces a number. Each is taught in 175 the lecture where it belongs — leakage in Lecture 2, imbalance in 176 Lecture 3, and all three PyTorch failures — the optimiser pair and 177 batch averaging — in Lecture 10, as a property of the method 178 rather than as a trap.</p> 179 180 <div class="panel panel-accent"> 181 <strong>The notebooks are ours.</strong> The textbook supplies the 182 syllabus, not the code. No notebook from the author, or from any other 183 third party, is used in this course. Every one is written for it, is 184 complete and correct, and carries above each cell the specification that 185 would produce it. 186 </div> 187 188 <h3>Four rules</h3> 189 <ol> 190 <li><strong>Never keep code you cannot explain line by line.</strong> 191 Part B of the paper puts a cell in front of you and asks what would 192 have to be true for its number to be trusted.</li> 193 <li><strong>Every number needs a baseline.</strong>
193 A metric with nothing to compare it to is decoration.</li> 194 <li><strong>Every model needs an ablation.</strong> Remove a component; show the number moves.</li> 195 <li><strong>Report the failure cases.</strong> A system with no known failure mode has not been tested.</li> 196 </ol> 197 </div> 198</section> 199 200<!-- ===================================================================== --> 201 202<section id="prerequisites"> 203 <div class="wrap"> 204 <h2>Prerequisites</h2> 205 <p class="lede">Short version: the mathematics of a third-year BSc in 206 Mathematics of Artificial Intelligence, and enough Python to read code 207 critically. If you are missing something, none of it takes long to fix 208 — and everything below is a pointer, not a reading list to complete 209 before the first lecture.</p> 210 211 <div class="panel"> 212 <strong>A note on scope.</strong> Everything <em>taught</em> in this 213 course comes from Chapters 1–16 of the textbook, or — for 214 Lectures 19–22 — from the lecture notes. The material on this 215 page is the opposite: it is what the course assumes you already have, so 216 the resources here necessarily point outside both. None of it is 217 examinable in its own right. 218 </div> 219 220 <h3>Test yourself in ten minutes</h3> 221 <p>These are not warm-up exercises. Each one is a thing you will actually be 222 asked to do, in the lecture where it appears. Try them before deciding you 223 need to revise anything.</p> 224 225 <ol class="selfcheck"> 226 <li> 227 <p class="q">$\mathbf{X}$ is $m \times n$ and $\boldsymbol\theta$ is 228 $n \times 1$. What is the shape of $\mathbf{X}\boldsymbol\theta$, and why 229 is $\mathbf{X}^{\mathsf T}\mathbf{X}$ square?</p> 230 <p class="where">Lecture 2, first ten minutes.</p> 231 </li> 232 <li> 233 <p class="q">Differentiate $\lVert\mathbf{X}\boldsymbol\theta - y\rVert^2$ 234 with respect to $\boldsymbol\theta$. Then say what has to be true of the 235 Hessian for the stationary point to be a minimum.</p> 236 <p class="where">Lecture 2. This is the derivation, not a preliminary to it.</p> 237 </li> 238 <li> 239 <p class="q">When is $\mathbf{X}^{\mathsf T}\mathbf{X}$ <em>not</em> 240 invertible? Answer in terms of the columns of $\mathbf{X}$.</p> 241 <p class="where">Lecture 2, and again in Lecture 5 when ridge repairs it.</p> 242 </li> 243 <li> 244 <p class="q">$A$ and $B$ each have variance $\sigma^2$ and correlation 245 $\rho$. What is $\operatorname{Var}\!\left(\tfrac{A+B}{2}\right)$?</p> 246 <p class="where">Lecture 7. This single calculation explains bagging, 247 random forests and extra-trees.</p> 248 </li> 249 <li> 250 <p class="q">What does a 95% confidence interval mean — and what 251 does it <em>not</em> mean?</p> 252 <p class="where">Lecture 2, on the final test score.</p> 253 </li> 254 <li> 255 <p class="q"><code>a</code> has shape <code>(100, 3)</code>, 256 <code>b</code> has shape <code>(3,)</code>. What does <code>a - b</code> 257 compute, and what shape comes out? Now <code>b</code> has shape 258 <code>(100,)</code> — what happens?</p> 259 <p class="where">Every lecture. Silent broadcasting is silent.</p> 260 </li> 261 <li> 262 <p class="q">Given a DataFrame <code>df</code>, select the rows where 263 <code>df["x"] > 5</code>, keeping only columns <code>"a"</code> and 264 <code>"b"</code>. Then count how many values in <code>"a"</code> are 265 missing.</p> 266 <p class="where">Lecture 1, in the notebook.</p> 267 </li> 268 <li> 269 <p class="q">Read this traceback out loud and say which line of 270 <em>your</em> code caused it, and why:<br> 271 <code>ValueError: Input X contains NaN.</code></p> 272 <p class="where">Lecture 2, the first time preprocessing meets a 273 missing value.</p> 274 </li> 275 </ol> 276 277 <div class="panel panel-accent"> 278 <strong>Could you do six or more?</strong> You are ready; skip the rest of 279 this page. <strong>Fewer?</strong> Find the matching row below. Nothing 280 here needs more than a few evenings, and none of it needs to be finished 281 before Lecture 1 — the items are listed in the order the course 282 first needs them. 283 </div> 284 285 <h3>Filling the gaps</h3> 286 287 <h4 class="gap-head">Mathematics</h4> 288 <div class="tablewrap"> 289 <table> 290 <thead> 291 <tr><th>What you need</th><th>First needed</th><th>If you are missing it</th></tr> 292 </thead> 293 <tbody> 294 <tr> 295 <td><strong>Linear algebra</strong><br>
296 <span class="muted">matrix products, transpose, inverse, rank, 297 column space, orthogonality, positive semi-definiteness, SVD</span></td> 298 <td class="num">L2</td> 299 <td> 300 <a href="https://www.3blue1brown.com/topics/linear-algebra">3Blue1Brown, <em>Essence of Linear Algebra</em></a> 301 for the geometric intuition â about three hours, and it is the 302 intuition the Lecture 2 derivation assumes.<br> 303 For depth: <a href="https://ocw.mit.edu/courses/18-06-linear-algebra-spring-2010/">MIT 18.06 (Strang)</a>. 304 </td> 305 </tr> 306 <tr> 307 <td><strong>Multivariable calculus</strong><br> 308 <span class="muted">partial derivatives, gradients, the chain 309 rule, Hessians, convexity</span></td> 310 <td class="num">L2</td> 311 <td> 312 <a href="https://www.3blue1brown.com/topics/calculus">3Blue1Brown, <em>Essence of Calculus</em></a>, 313 then chapter 5 of 314 <a href="https://mml-book.github.io/"><em>Mathematics for Machine Learning</em></a> 315 (free PDF) for vector calculus in exactly the notation we use. 316 </td> 317 </tr> 318 <tr> 319 <td><strong>Probability and statistics</strong><br> 320 <span class="muted">expectation, variance, covariance, 321 correlation, independence, sampling, confidence intervals</span></td> 322 <td class="num">L2, L5, L7</td> 323 <td> 324 <a href="https://seeing-theory.brown.edu/">Seeing Theory</a> (Brown) 325 is a visual refresher in an afternoon.<br> 326 For depth: <a href="https://ocw.mit.edu/courses/18-05-introduction-to-probability-and-statistics-spring-2022/">MIT 18.05</a>, 327 or chapter 6 of <em>Mathematics for Machine Learning</em>. 328 </td> 329 </tr> 330 </tbody> 331 </table> 332 </div> 333 <p class="small muted">If you are on this degree programme you almost 334 certainly have all three. The self-check above is a faster way to confirm 335 that than reading the list.</p> 336 337 <h4 class="gap-head">Programming</h4> 338 <div class="tablewrap"> 339 <table> 340 <thead> 341 <tr><th>What you need</th><th>First needed</th><th>If you are missing it</th></tr> 342 </thead> 343 <tbody> 344 <tr> 345 <td><strong>Python</strong><br> 346 <span class="muted">functions, lists and dicts, comprehensions, 347 imports, reading a traceback</span></td> 348 <td class="num">L1</td> 349 <td><a href="https://docs.python.org/3/tutorial/">The official Python tutorial</a>, 350 sections 3–6. A weekend from a standing start; an evening if 351 you know another language.</td> 352 </tr> 353 <tr> 354 <td><strong>NumPy</strong><br> 355 <span class="muted">arrays, shapes, indexing, axes, and 356 <em>broadcasting</em></span></td> 357 <td class="num">L1</td> 358 <td><a href="https://numpy.org/doc/stable/user/absolute_beginners.html">NumPy: the absolute basics</a>, 359 then <a href="https://numpy.org/doc/stable/user/basics.broadcasting.html">the broadcasting rules</a> 360 â read that second page twice. It is the single most common source 361 of code that runs and is wrong.</td> 362 </tr> 363 <tr> 364 <td><strong>pandas</strong><br> 365 <span class="muted">DataFrame, Series, selection, missing values</span></td> 366 <td class="num">L1</td> 367 <td><a href="https://pandas.pydata.org/docs/user_guide/10min.html">10 minutes to pandas</a>. 368 Optimistically named, but an hour genuinely does it.</td> 369 </tr> 370 <tr> 371 <td><strong>matplotlib</strong><br> 372 <span class="muted">enough to draw a histogram and a scatter plot</span></td> 373 <td class="num">L1</td> 374 <td><a href="https://matplotlib.org/stable/tutorials/pyplot.html">
374The pyplot tutorial</a>. 375 Twenty minutes; you will not need more than this in the whole course.</td> 376 </tr> 377 <tr> 378 <td><strong>Google Colab</strong><br> 379 <span class="muted">running cells, restarting the runtime, 380 reading a traceback</span></td> 381 <td class="num">L1</td> 382 <td><a href="https://colab.research.google.com/notebooks/intro.ipynb">Colab’s own introduction</a>. 383 Ten minutes. Bring a Google account to the first lecture.</td> 384 </tr> 385 <tr> 386 <td><strong>scikit-learn</strong><br> 387 <span class="muted">the estimator API â <code>fit</code>, 388 <code>transform</code>, <code>predict</code></span></td> 389 <td class="num">L2</td> 390 <td class="taught">Helpful but <strong>not required</strong> â taught 391 from scratch in Lecture 2. Lecture 1 fits nothing at 392 all: looking properly at data before modelling it is not a
393 preliminary, it is what decides whether the model can work. 394 If you are curious: 395 <a href="https://scikit-learn.org/stable/getting_started.html">Getting started</a>.</td> 396 </tr> 397 <tr> 398 <td><strong>PyTorch</strong></td> 399 <td class="num">L10</td> 400 <td class="taught"><strong>Not required.</strong> Taught from first 401 principles in Lecture 10 â tensors, autograd, and the 402 training loop written out in full before any of it is hidden 403 behind a helper. If you insist: 404 <a href="https://docs.pytorch.org/tutorials/beginner/basics/intro.html">the official basics</a>.</td> 405 </tr> 406 </tbody> 407 </table> 408 </div> 409 410 <h3>What you also need on the day</h3> 411 <ul class="plain"> 412 <li>A laptop, and a <strong>Google account</strong> for Colab. Nothing is 413 installed locally; nothing depends on your operating system.</li> 414 <li>Access to an <strong>AI coding assistant</strong>. Which one is up to 415 you. The course does not require it — the notebooks are already 416 written — but Lecture 1 sets out how to use one, and you 417 will need that long after this course.</li> 418 <li>Nothing else. <strong>Every notebook in the course runs on a free 419 CPU runtime</strong>, including the transformer and multimodal ones: 420 where a lecture’s deck reports a result from a larger run, the 421 notebook says so and reproduces the ordering at a smaller scale.</li> 422 </ul> 423 424 <h3>What you explicitly do <em>not</em> need</h3> 425 <div class="grid"> 426 <div class="card"> 427 <h4>Prior machine learning</h4> 428 <p>Helpful, not assumed. The course starts from a problem and a dataset 429 on the first day and builds everything from there.</p> 430 </div> 431 <div class="card"> 432 <h4>Deep learning experience</h4> 433 <p>Parts II–VI assume nothing beyond what Part I 434 established. Neural networks arrive in Lecture 9, once the classical 435 models have been covered properly.</p> 436 </div> 437 <div class="card"> 438 <h4>Software engineering</h4> 439 <p>
439No production systems, no deployment infrastructure, no build tooling. 440 Notebooks throughout.</p> 441 </div> 442 <div class="card"> 443 <h4>Fluent typing</h4> 444 <p>The notebooks are written for you. What is examined is whether you can 445 <em>read</em> them — say what a cell does, what it measures, and 446 what would break if an argument changed.</p> 447 </div> 448 </div> 449 </div> 450</section> 451 452<!-- ===================================================================== --> 453 454<!-- timetable --> 455<!-- The one place in the course that names days. Everything else refers to 456 lectures by number, and tools/check_decks.py enforces that; this fence 457 lifts the ban for the section whose whole subject is the timetable. --> 458<section id="calendar"> 459 <div class="wrap"> 460 <h2>Calendar</h2> 461 <p class="lede">Twenty-four lectures on Tuesdays and Wednesdays, from 462 6 October 2026 to 12 January 2027.</p> 463 464 <dl class="facts-row"> 465 <div> 466 <dt>Days</dt> 467 <dd>Tuesday and Wednesday</dd> 468 </div> 469 <div> 470 <dt>Hours</dt> 471 <dd>12:00–14:00</dd> 472 </div> 473 <div> 474 <dt>Room</dt> 475 <dd>Aula C · 476 <a href="https://www.mat.uniroma1.it/">Dipartimento di Matematica 477 “Guido Castelnuovo”</a></dd> 478 </div> 479 </dl> 480 481 <p class="cal-next" hidden></p> 482 483 <div class="tablewrap"> 484 <table class="calendar"> 485 <caption>Every lecture in date order, and the days inside the term that 486 carry none.</caption> 487 <tbody> 488 <!-- BEGIN CALENDAR --> 489 <tr class="cal-month"> 490 <th colspan="3" scope="colgroup">October 2026</th> 491 </tr> 492 <tr data-date="2026-10-06"> 493 <td class="cal-when"><time datetime="2026-10-06">Tue 6</time></td> 494 <td class="cal-n">01</td> 495 <td class="cal-what">What machine learning is, and how we will work</td> 496 </tr> 497 <tr data-date="2026-10-07"> 498 <td class="cal-when"><time datetime="2026-10-07">Wed 7</time></td> 499 <td class="cal-n">02</td> 500 <td class="cal-what">The end-to-end project</td> 501 </tr> 502 <tr data-date="2026-10-13"> 503 <td class="cal-when"><time datetime="2026-10-13">Tue 13</time></td> 504 <td class="cal-n">03</td> 505 <td class="cal-what">Classification and its metrics</td> 506 </tr> 507 <tr data-date="2026-10-14"> 508 <td class="cal-when"><time datetime="2026-10-14">Wed 14</time></td> 509 <td class="cal-n">04</td> 510 <td class="cal-what">Training models</td> 511 </tr> 512 <tr data-date="2026-10-20"> 513 <td class="cal-when"><time datetime="2026-10-20">Tue 20</time></td> 514 <td class="cal-n">05</td> 515 <td class="cal-what">Regularisation and the biasâvariance trade-off</td> 516 </tr> 517 <tr data-date="2026-10-21"> 518 <td class="cal-when"><time datetime="2026-10-21">Wed 21</time></td> 519 <td class="cal-n">06</td> 520 <td class="cal-what">Decision trees</td> 521 </tr> 522 <tr data-date="2026-10-27"> 523 <td class="cal-when"><time datetime="2026-10-27">Tue 27</time></td> 524 <td class="cal-n">07</td> 525 <td class="cal-what">Ensembles and random forests</td> 526 </tr> 527 <tr data-date="2026-10-28"> 528 <td class="cal-when"><time datetime="2026-10-28">Wed 28</time></td> 529 <td class="cal-n">08</td> 530 <td class="cal-what">Dimensionality reduction and unsupervised learning</td> 531 </tr> 532 <tr class="cal-month"> 533 <th colspan="3" scope="colgroup">November 2026</th> 534 </tr> 535 <tr data-date="2026-11-03"> 536 <td class="cal-when"><time datetime="2026-11-03">Tue 3</time></td> 537 <td class="cal-n">09</td> 538 <td class="cal-what">Neural networks, from the perceptron up</td> 539 </tr> 540 <tr data-date="2026-11-04"> 541 <td class="cal-when"><time datetime="2026-11-04">Wed 4</time></td> 542 <td class="cal-n">10</td> 543 <td class="cal-what">PyTorch</td> 544 </tr> 545 <tr data-date="2026-11-10"> 546 <td class="cal-when"><time datetime="2026-11-10">Tue 10</time></td> 547 <td class="cal-n">11</td> 548 <td class="cal-what">Training deep networks</td> 549 </tr> 550 <tr data-date="2026-11-11"> 551 <td class="cal-when"><time datetime="2026-11-11">Wed 11</time></td> 552 <td class="cal-n">12</td> 553 <td class="cal-what">Convolutional networks</td> 554 </tr> 555 <tr data-date="2026-11-17"> 556 <td class="cal-when"><time datetime="2026-11-17">Tue 17</time></td> 557 <td class="cal-n">13</td> 558 <td class="cal-what">Transfer learning</td> 559 </tr> 560 <tr data-date="2026-11-18"> 561 <td class="cal-when"><time datetime="2026-11-18">Wed 18</time></td> 562 <td class="cal-n">14</td> 563 <td class="cal-what">Detection and segmentation</td> 564 </tr> 565 <tr data-date="2026-11-24"> 566 <td class="cal-when"><time datetime="2026-11-24">Tue 24</time></td> 567 <td class="cal-n">15</td> 568 <td class="cal-what">Time series</td> 569 </tr> 570 <tr data-date="2026-11-25"> 571 <td class="cal-when"><time datetime="2026-11-25">Wed 25</time></td> 572 <td class="cal-n">16</td> 573 <td class="cal-what">Recurrent networks</td> 574 </tr> 575 <tr class="cal-month"> 576 <th colspan="3" scope="colgroup">December 2026</th> 577 </tr> 578 <tr data-date="2026-12-01"> 579 <td class="cal-when"><time datetime="2026-12-01">Tue 1</time></td> 580 <td class="cal-n">17</td> 581 <td class="cal-what">Text</td> 582 </tr> 583 <tr data-date="2026-12-02"> 584 <td class="cal-when"><time datetime="2026-12-02">Wed 2</time></td> 585 <td class="cal-n">18</td> 586 <td class="cal-what">Attention and transformers</td> 587 </tr> 588 <tr class="cal-off"> 589 <td class="cal-when"><time datetime="2026-12-08">Tue 8</time></td> 590 <td class="cal-n">—</td> 591 <td class="cal-what">Immacolata — the university is closed</td> 592 </tr> 593 <tr data-date="2026-12-09"> 594 <td class="cal-when"><time datetime="2026-12-09">Wed 9</time></td> 595 <td class="cal-n">19</td> 596 <td class="cal-what">Information retrieval: the lexical foundation</td> 597 </tr> 598 <tr data-date="2026-12-15"> 599 <td class="cal-when"><time datetime="2026-12-15">Tue 15</time></td> 600 <td class="cal-n">20</td> 601 <td class="cal-what">Information retrieval: dense retrieval</td> 602 </tr> 603 <tr data-date="2026-12-16"> 604 <td class="cal-when"><time datetime="2026-12-16">Wed 16</time></td> 605 <td class="cal-n">21</td> 606 <td class="cal-what">Recommender systems: from ratings to factors</td> 607 </tr> 608 <tr data-date="2026-12-22"> 609 <td class="cal-when"><time datetime="2026-12-22">Tue 22</time></td> 610 <td class="cal-n">22</td> 611 <td class="cal-what">Recommender systems: neural, and evaluated honestly</td> 612 </tr> 613 <tr data-date="2026-12-23"> 614 <td class="cal-when"><time datetime="2026-12-23">Wed 23</time></td> 615 <td class="cal-n">23</td> 616 <td class="cal-what">Vision transformers and multimodal retrieval</td> 617 </tr> 618 <tr class="cal-off"> 619 <td class="cal-when">24 Dec – 6 Jan</td> 620 <td class="cal-n">—</td> 621 <td class="cal-what">Christmas break</td> 622 </tr> 623 <tr class="cal-month"> 624 <th colspan="3" scope="colgroup">January 2027</th> 625 </tr> 626 <tr data-date="2027-01-12"> 627 <td class="cal-when"><time datetime="2027-01-12">Tue 12</time></td> 628 <td class="cal-n">24</td> 629 <td class="cal-what">Generation, retrieval-augmented systems, and where this leaves you</td> 630 </tr> <!-- END CALENDAR --> 631 </tbody> 632 </table> 633 </div> 634 635 <p class="panel"><strong>The material is not published in advance.</strong> 636 A lecture’s slides, notebook and notes appear on this page at 637 <strong>11:30</strong> on the day that lecture is taught, half an hour 638 before it begins, and stay there for the rest of the course.</p> 639 </div> 640</section> 641<!-- /timetable --> 642 643<!-- ===================================================================== --> 644 645<section id="lectures"> 646 <div class="wrap"> 647 <h2>Lectures</h2> 648 <p class="lede">Twenty-four lectures in six parts. Each lecture’s 649 slides, notebook and notes appear on its own card at 11:30 on the day it is 650 taught.</p> 651 652 <noscript> 653 <p class="panel"><strong>This page needs JavaScript to show the 654 material.</strong> A lecture’s slides, notebook and notes are added 655 to its card by a script, on the day of that lecture. With JavaScript 656 turned off the cards stay shut.</p> 657 </noscript> 658 659 <!-- BEGIN LECTURES --> 660 <div class="part-head"> 661 <h3>Part I — Tabular data and classical models</h3>
662 <span class="part-meta">Lectures 1â8 · Chapters 1â8 · runs on CPU</span> 663 </div> 664 <ol class="lectures"> 665 <li class="lecture" data-n="01" data-reveal="2026-10-06T11:30:00+02:00" data-material="slides,pdf,notebook,notes"> 666 <span class="n">01</span> 667 <div class="body"> 668 <p class="t">What machine learning is, and how we will work</p> 669 <p class="meta">California housing <span class="badge badge-ch">Ch 1â2</span></p> 670 <p class="when"><time datetime="2026-10-06">Tue 6 Oct 2026</time> · 12:00â14:00 · Aula C</p> 671 </div> 672 <div class="links"> 673 <span class="btn-locked">Opens Tue 6 Oct, 11:30</span> 674 </div> 675 </li> 676 <li class="lecture" data-n="02" data-reveal="2026-10-07T11:30:00+02:00" data-material="slides,pdf,notebook,notes"> 677 <span class="n">02</span> 678 <div class="body"> 679 <p class="t">The end-to-end project</p> 680 <p class="meta">California housing <span class="badge badge-ch">Ch 2</span></p> 681 <p class="when"><time datetime="2026-10-07">Wed 7 Oct 2026</time> · 12:00â14:00 · Aula C</p> 682 <p class="thread">Derivation · Least squares and the normal equation</p> 683 </div> 684 <div class="links"> 685 <span class="btn-locked">Opens Wed 7 Oct, 11:30</span> 686 </div> 687 </li> 688 <li class="lecture" data-n="03" data-reveal="2026-10-13T11:30:00+02:00" data-material="slides,pdf,notebook,notes"> 689 <span class="n">03</span> 690 <div class="body"> 691 <p class="t">Classification and its metrics</p> 692 <p class="meta">MNIST <span class="badge badge-ch">Ch 3</span></p> 693 <p class="when"><time datetime="2026-10-13">Tue 13 Oct 2026</time> · 12:00â14:00 · Aula C</p> 694 <p class="thread">Derivation · Imbalance, and the non-monotonicity of precision</p> 695 </div> 696 <div class="links"> 697 <span class="btn-locked">Opens Tue 13 Oct, 11:30</span> 698 </div> 699 </li> 700 <li class="lecture" data-n="04" data-reveal="2026-10-14T11:30:00+02:00" data-material="slides,pdf,notebook,notes"> 701 <span class="n">04</span> 702 <div class="body"> 703 <p class="t">Training models</p> 704 <p class="meta">Titanic <span class="badge badge-ch">Ch 4</span></p> 705 <p class="when"><time datetime="2026-10-14">Wed 14 Oct 2026</time> · 12:00â14:00 · Aula C</p> 706 <p class="thread">Derivation · Gradient descent</p> 707 </div> 708 <div class="links"> 709 <span class="btn-locked">Opens Wed 14 Oct, 11:30</span> 710 </div> 711 </li> 712 <li class="lecture" data-n="05" data-reveal="2026-10-20T11:30:00+02:00" data-material="slides,pdf,notebook,notes"> 713 <span class="n">05</span> 714 <div class="body"> 715 <p class="t">Regularisation and the biasâvariance trade-off</p> 716 <p class="meta">Titanic <span class="badge badge-ch">Ch 4</span></p> 717 <p class="when"><time datetime="2026-10-20">Tue 20 Oct 2026</time> · 12:00â14:00 · Aula C</p> 718 <p class="thread">Derivation · The biasâvariance decomposition</p> 719 </div> 720 <div class="links"> 721 <span class="btn-locked">Opens Tue 20 Oct, 11:30</span> 722 </div> 723 </li> 724 <li class="lecture" data-n="06" data-reveal="2026-10-21T11:30:00+02:00" data-material="slides,pdf,notebook,notes"> 725 <span class="n">06</span> 726 <div class="body"> 727 <p class="t">Decision trees</p> 728 <p class="meta">CoverType <span class="badge badge-ch">Ch 5</span></p> 729 <p class="when"><time datetime="2026-10-21">Wed 21 Oct 2026</time> · 12:00â14:00 · Aula C</p> 730 <p class="thread">Derivation · Impurity: Gini and entropy</p> 731 </div> 732 <div class="links"> 733 <span class="btn-locked">Opens Wed 21 Oct, 11:30</span> 734 </div> 735 </li> 736 <li class="lecture" data-n="07" data-reveal="2026-10-27T11:30:00+01:00" data-material="slides,pdf,notebook,notes">
737 <span class="n">07</span> 738 <div class="body"> 739 <p class="t">Ensembles and random forests</p> 740 <p class="meta">CoverType <span class="badge badge-ch">Ch 6</span></p> 741 <p class="when"><time datetime="2026-10-27">Tue 27 Oct 2026</time> · 12:00â14:00 · Aula C</p> 742 <p class="thread">Derivation · The variance of an average of correlated predictors</p> 743 </div> 744 <div class="links"> 745 <span class="btn-locked">Opens Tue 27 Oct, 11:30</span> 746 </div> 747 </li> 748 <li class="lecture" data-n="08" data-reveal="2026-10-28T11:30:00+01:00" data-material="slides,pdf,notebook,notes"> 749 <span class="n">08</span> 750 <div class="body"> 751 <p class="t">Dimensionality reduction and unsupervised learning</p> 752 <p class="meta">Olivetti faces <span class="badge badge-ch">Ch 7â8</span></p> 753 <p class="when"><time datetime="2026-10-28">Wed 28 Oct 2026</time> · 12:00â14:00 · Aula C</p> 754 <p class="thread">Derivation · PCA via the SVD; JohnsonâLindenstrauss</p> 755 </div> 756 <div class="links"> 757 <span class="btn-locked">Opens Wed 28 Oct, 11:30</span> 758 </div> 759 </li> 760 </ol> 761 <div class="part-head"> 762 <h3>Part II — Neural networks</h3> 763 <span class="part-meta">Lectures 9â11 · Chapters 9â11 · runs on CPU</span> 764 </div> 765 <ol class="lectures"> 766 <li class="lecture" data-n="09" data-reveal="2026-11-03T11:30:00+01:00" data-material="slides,pdf,notebook,notes"> 767 <span class="n">09</span> 768 <div class="body"> 769 <p class="t">Neural networks, from the perceptron up</p> 770 <p class="meta">Fashion-MNIST <span class="badge badge-ch">Ch 9</span></p> 771 <p class="when"><time datetime="2026-11-03">Tue 3 Nov 2026</time> · 12:00â14:00 · Aula C</p> 772 <p class="thread">Derivation · What a layer computes</p> 773 </div> 774 <div class="links"> 775 <span class="btn-locked">Opens Tue 3 Nov, 11:30</span> 776 </div> 777 </li> 778 <li class="lecture" data-n="10" data-reveal="2026-11-04T11:30:00+01:00" data-material="slides,pdf,notebook,notes"> 779 <span class="n">10</span> 780 <div class="body"> 781 <p class="t">PyTorch</p> 782 <p class="meta">Fashion-MNIST <span class="badge badge-ch">Ch 10</span></p> 783 <p class="when"><time datetime="2026-11-04">Wed 4 Nov 2026</time> · 12:00â14:00 · Aula C</p> 784 <p class="thread">Derivation · Backpropagation as reverse-mode automatic differentiation</p> 785 </div> 786 <div class="links"> 787 <span class="btn-locked">Opens Wed 4 Nov, 11:30</span> 788 </div> 789 </li> 790 <li class="lecture" data-n="11" data-reveal="2026-11-10T11:30:00+01:00" data-material="slides,pdf,notebook,notes"> 791 <span class="n">11</span> 792 <div class="body"> 793 <p class="t">Training deep networks</p> 794 <p class="meta">CIFAR-10 <span class="badge badge-ch">Ch 11</span></p> 795 <p class="when"><time datetime="2026-11-10">Tue 10 Nov 2026</time> · 12:00â14:00 · Aula C</p> 796 <p class="thread">Derivation · Variance propagation and weight initialisation</p> 797 </div> 798 <div class="links"> 799 <span class="btn-locked">Opens Tue 10 Nov, 11:30</span> 800 </div> 801 </li> 802 </ol> 803 <div class="part-head"> 804 <h3>Part III — Computer vision</h3> 805 <span class="part-meta">Lectures 12â14 · Chapter 12 · runs on CPU</span> 806 </div> 807 <ol class="lectures"> 808 <li class="lecture" data-n="12" data-reveal="2026-11-11T11:30:00+01:00" data-material="slides,pdf,notebook,notes"> 809 <span class="n">12</span> 810 <div class="body"> 811 <p class="t">Convolutional networks</p> 812 <p class="meta">Flowers102 <span class="badge badge-ch">Ch 12</span></p> 813 <p class="when"><time datetime="2026-11-11">Wed 11 Nov 2026</time> · 12:00â14:00 · Aula C</p> 814 <p class="thread">Derivation · Weight sharing, equivariance and memory</p> 815 </div> 816 <div class="links">
817 <span class="btn-locked">Opens Wed 11 Nov, 11:30</span> 818 </div> 819 </li> 820 <li class="lecture" data-n="13" data-reveal="2026-11-17T11:30:00+01:00" data-material="slides,pdf,notebook,notes"> 821 <span class="n">13</span> 822 <div class="body"> 823 <p class="t">Transfer learning</p> 824 <p class="meta">Flowers102 <span class="badge badge-ch">Ch 12</span></p> 825 <p class="when"><time datetime="2026-11-17">Tue 17 Nov 2026</time> · 12:00â14:00 · Aula C</p> 826 </div> 827 <div class="links"> 828 <span class="btn-locked">Opens Tue 17 Nov, 11:30</span> 829 </div> 830 </li> 831 <li class="lecture" data-n="14" data-reveal="2026-11-18T11:30:00+01:00" data-material="slides,pdf,notebook,notes"> 832 <span class="n">14</span> 833 <div class="body"> 834 <p class="t">Detection and segmentation</p> 835 <p class="meta">COCO <span class="badge badge-ch">Ch 12</span></p> 836 <p class="when"><time datetime="2026-11-18">Wed 18 Nov 2026</time> · 12:00â14:00 · Aula C</p> 837 <p class="thread">Derivation · IoUâs vanishing gradient; mAP</p> 838 </div> 839 <div class="links"> 840 <span class="btn-locked">Opens Wed 18 Nov, 11:30</span> 841 </div> 842 </li> 843 </ol> 844 <div class="part-head"> 845 <h3>Part IV — Sequences and language</h3> 846 <span class="part-meta">Lectures 15â18 · Chapters 13â15 · runs on CPU</span> 847 </div> 848 <ol class="lectures"> 849 <li class="lecture" data-n="15" data-reveal="2026-11-24T11:30:00+01:00" data-material="slides,pdf,notebook,notes"> 850 <span class="n">15</span> 851 <div class="body"> 852 <p class="t">Time series</p> 853 <p class="meta">Chicago transit ridership <span class="badge badge-ch">Ch 13</span></p> 854 <p class="when"><time datetime="2026-11-24">Tue 24 Nov 2026</time> · 12:00â14:00 · Aula C</p> 855 <p class="thread">Derivation · Stationarity, differencing and autocorrelation</p> 856 </div> 857 <div class="links"> 858 <span class="btn-locked">Opens Tue 24 Nov, 11:30</span> 859 </div> 860 </li> 861 <li class="lecture" data-n="16" data-reveal="2026-11-25T11:30:00+01:00" data-material="slides,pdf,notebook,notes"> 862 <span class="n">16</span> 863 <div class="body"> 864 <p class="t">Recurrent networks</p> 865 <p class="meta">Chicago transit ridership <span class="badge badge-ch">Ch 13</span></p> 866 <p class="when"><time datetime="2026-11-25">Wed 25 Nov 2026</time> · 12:00â14:00 · Aula C</p> 867 </div> 868 <div class="links"> 869 <span class="btn-locked">Opens Wed 25 Nov, 11:30</span> 870 </div> 871 </li> 872 <li class="lecture" data-n="17" data-reveal="2026-12-01T11:30:00+01:00" data-material="slides,pdf,notebook,notes"> 873 <span class="n">17</span> 874 <div class="body"> 875 <p class="t">Text</p> 876 <p class="meta">IMDb <span class="badge badge-ch">Ch 14</span></p> 877 <p class="when"><time datetime="2026-12-01">Tue 1 Dec 2026</time> · 12:00â14:00 · Aula C</p> 878 <p class="thread">Derivation · Softmax, cross-entropy and logits</p> 879 </div> 880 <div class="links"> 881 <span class="btn-locked">Opens Tue 1 Dec, 11:30</span> 882 </div> 883 </li> 884 <li class="lecture" data-n="18" data-reveal="2026-12-02T11:30:00+01:00" data-material="slides,pdf,notebook,notes"> 885 <span class="n">18</span> 886 <div class="body"> 887 <p class="t">Attention and transformers</p> 888 <p class="meta">IMDb <span class="badge badge-ch">Ch 14â15</span></p> 889 <p class="when"><time datetime="2026-12-02">Wed 2 Dec 2026</time> · 12:00â14:00 · Aula C</p> 890 <p class="thread">Derivation · Scaled dot-product attention</p> 891 </div> 892 <div class="links"> 893 <span class="btn-locked">Opens Wed 2 Dec, 11:30</span> 894 </div> 895 </li> 896 </ol> 897 <div class="part-head"> 898 <h3>Part V — Information retrieval and recommender systems</h3>
899 <span class="part-meta">Lectures 19â22 · Outside the book · examinable</span> 900 </div> 901 <ol class="lectures"> 902 <li class="lecture" data-n="19" data-reveal="2026-12-09T11:30:00+01:00" data-material="slides,pdf,notebook,notes"> 903 <span class="n">19</span> 904 <div class="body"> 905 <p class="t">Information retrieval: the lexical foundation</p> 906 <p class="meta">SciFact (BEIR) <span class="badge badge-math">Outside the book</span></p> 907 <p class="when"><time datetime="2026-12-09">Wed 9 Dec 2026</time> · 12:00â14:00 · Aula C</p> 908 <p class="thread">Derivation · Evaluating a ranking: MRR, AP, NDCG</p> 909 </div> 910 <div class="links"> 911 <span class="btn-locked">Opens Wed 9 Dec, 11:30</span> 912 </div> 913 </li> 914 <li class="lecture" data-n="20" data-reveal="2026-12-15T11:30:00+01:00" data-material="slides,pdf,notebook,notes"> 915 <span class="n">20</span> 916 <div class="body"> 917 <p class="t">Information retrieval: dense retrieval</p> 918 <p class="meta">SciFact (BEIR) <span class="badge badge-math">Outside the book</span></p> 919 <p class="when"><time datetime="2026-12-15">Tue 15 Dec 2026</time> · 12:00â14:00 · Aula C</p> 920 </div> 921 <div class="links"> 922 <span class="btn-locked">Opens Tue 15 Dec, 11:30</span> 923 </div> 924 </li> 925 <li class="lecture" data-n="21" data-reveal="2026-12-16T11:30:00+01:00" data-material="slides,pdf,notebook,notes"> 926 <span class="n">21</span> 927 <div class="body"> 928 <p class="t">Recommender systems: from ratings to factors</p> 929 <p class="meta">MovieLens <span class="badge badge-math">Outside the book</span></p> 930 <p class="when"><time datetime="2026-12-16">Wed 16 Dec 2026</time> · 12:00â14:00 · Aula C</p> 931 <p class="thread">Derivation · Matrix factorisation, and its relation to the SVD</p> 932 </div> 933 <div class="links"> 934 <span class="btn-locked">Opens Wed 16 Dec, 11:30</span> 935 </div> 936 </li> 937 <li class="lecture" data-n="22" data-reveal="2026-12-22T11:30:00+01:00" data-material="slides,pdf,notebook,notes"> 938 <span class="n">22</span> 939 <div class="body"> 940 <p class="t">Recommender systems: neural, and evaluated honestly</p> 941 <p class="meta">MovieLens <span class="badge badge-math">Outside the book</span></p> 942 <p class="when"><time datetime="2026-12-22">Tue 22 Dec 2026</time> · 12:00â14:00 · Aula C</p> 943 </div> 944 <div class="links"> 945 <span class="btn-locked">Opens Tue 22 Dec, 11:30</span> 946 </div> 947 </li> 948 </ol> 949 <div class="part-head"> 950 <h3>Part VI — Multimodal models, and closing the course</h3> 951 <span class="part-meta">Lectures 23â24 · Chapters 15â16 · runs on CPU</span> 952 </div> 953 <ol class="lectures"> 954 <li class="lecture" data-n="23" data-reveal="2026-12-23T11:30:00+01:00" data-material="slides,pdf,notebook,notes"> 955 <span class="n">23</span> 956 <div class="body"> 957 <p class="t">Vision transformers and multimodal retrieval</p> 958 <p class="meta">COCO <span class="badge badge-ch">Ch 15â16</span></p> 959 <p class="when"><time datetime="2026-12-23">Wed 23 Dec 2026</time> · 12:00â14:00 · Aula C</p> 960 <p class="thread">Derivation · The contrastive objective and its temperature</p> 961 </div> 962 <div class="links"> 963 <span class="btn-locked">Opens Wed 23 Dec, 11:30</span> 964 </div> 965 </li> 966 <li class="lecture" data-n="24" data-reveal="2027-01-12T11:30:00+01:00" data-material="slides,pdf,notebook,notes"> 967 <span class="n">24</span> 968 <div class="body"> 969 <p class="t">Generation, retrieval-augmented systems, and where this leaves you</p> 970 <p class="meta">COCO and the Part V corpora <span class="badge badge-ch">Ch 15â16</span></p> 971 <p class="when"><time datetime="2027-01-12">Tue 12 Jan 2027</time> · 12:00â14:00 · Aula C</p> 972 </div> 973 <div class="links">
974 <span class="btn-locked">Opens Tue 12 Jan, 11:30</span> 975 </div> 976 </li> 977 </ol> <!-- END LECTURES --> 978 </div> 979</section> 980 981<!-- ===================================================================== --> 982 983<section id="derivations"> 984 <div class="wrap"> 985 <h2>The mathematics</h2> 986 <p class="lede">Each lecture derives the one object its method rests on. Not 987 a parallel theory course: every derivation below does visible work on the 988 method taught beside it, and they are 40% of the written paper.</p> 989 990 <ol class="threads"> 991 <!-- BEGIN DERIVATIONS --> 992 <li>Least squares and the normal equation <span class="where">· Lecture 2</span></li> 993 <li>Imbalance, and the non-monotonicity of precision <span class="where">· Lecture 3</span></li> 994 <li>Gradient descent <span class="where">· Lecture 4</span></li> 995 <li>The biasâvariance decomposition <span class="where">· Lecture 5</span></li> 996 <li>Impurity: Gini and entropy <span class="where">· Lecture 6</span></li> 997 <li>The variance of an average of correlated predictors <span class="where">· Lecture 7</span></li> 998 <li>PCA via the SVD; JohnsonâLindenstrauss <span class="where">· Lecture 8</span></li> 999 <li>What a layer computes <span class="where">· Lecture 9</span></li> 1000 <li>Backpropagation as reverse-mode automatic differentiation <span class="where">· Lecture 10</span></li> 1001 <li>Variance propagation and weight initialisation <span class="where">· Lecture 11</span></li> 1002 <li>Weight sharing, equivariance and memory <span class="where">· Lecture 12</span></li> 1003 <li>IoUâs vanishing gradient; mAP <span class="where">· Lecture 14</span></li> 1004 <li>Stationarity, differencing and autocorrelation <span class="where">· Lecture 15</span></li> 1005 <li>Softmax, cross-entropy and logits <span class="where">· Lecture 17</span></li> 1006 <li>Scaled dot-product attention <span class="where">· Lecture 18</span></li> 1007 <li>Evaluating a ranking: MRR, AP, NDCG <span class="where">· Lecture 19</span></li> 1008 <li>Matrix factorisation, and its relation to the SVD <span class="where">· Lecture 21</span></li> 1009 <li>The contrastive objective and its temperature <span class="where">·
1009 Lecture 23</span></li> <!-- END DERIVATIONS --> 1010 </ol> 1011 1012 <div class="panel"> 1013 The derivations are cross-referential and the order matters. Lecture 5 1014 completes Lecture 2; Lecture 14 uses Lecture 3; 1015 Lecture 21 and Lecture 23 both use Lecture 8; Lecture 20 1016 and Lecture 22 are the same architecture twice; Lecture 23 is 1017 evaluated with Lecture 19’s metrics. 1018 </div> 1019 </div> 1020</section> 1021 1022<!-- ===================================================================== --> 1023 1024<section id="assessment"> 1025 <div class="wrap"> 1026 <h2>Assessment</h2> 1027 <p class="lede">The written examination carries the mark. The oral is 1028 optional, short, and can move that mark in either direction.</p> 1029 1030 <div class="grid"> 1031 <div class="card"> 1032 <h4>Written examination</h4> 1033 <p>Two hours, closed book, no formula sheet. Three parts: the 1034 derivations, choosing a method for a stated situation, and reading 1035 results — as technical exercises, closed choices with their 1036 reason, and short open answers.</p> 1037 <p class="muted">Marked out of 30, pass at 18. From the written 1038 alone, the mark recorded is capped at <strong>27</strong>.</p> 1039 </div> 1040 <div class="card"> 1041 <h4>Oral examination — optional</h4> 1042 <p>Seven to ten minutes. <strong>Three questions, on any topic from 1043 the course.</strong></p> 1044 <p class="muted">No notes, no computer. Open to anyone who passed the 1045 written.</p> 1046 </div> 1047 </div> 1048 1049 <div class="panel"> 1050 <span class="panel-title-inline">How the mark is made</span> 1051 <strong>The oral moves your written mark, up or down, and the result is 1052 final.</strong> <strong>28, 29 and 30 exist only through the oral</strong>, 1053 and lode is a separate decision, from 30.<br><br> 1054 Sitting it is <strong>your choice</strong>, made when you see your written 1055 mark — and <strong>binding once registered</strong>. Both are taken 1056 in the same session: a written pass does <strong>not</strong> carry 1057 forward. 1058 </div> 1059 1060 <div class="panel panel-accent"> 1061 <strong>Why it is built this way.</strong> A grade you can improve at no 1062 risk is one everybody attempts, which turns the oral into a queue rather 1063 than an examination. Making it a real decision — three marks up, 1064 three marks down — means the people who sit it are the people with 1065 something to show. Nothing about it is hidden: the questions are published, 1066 the arithmetic is published, and neither changes after you have decided. 1067 </div> 1068 1069 <h3>Can a course this applied really have a written examination?</h3> 1070 <p>Yes — and this course is better suited to one than most applied 1071 courses, for a specific reason.</p> 1072 <div class="panel"> 1073 The competence being built here is not typing. It is 1074 <strong>judgement about whether a result can be trusted</strong>, and that 1075 judgement is exercised by reading: reading code you did not write, reading a 1076 metric, reading a curve. Reading is paper-native. A written examination 1077 tests it <em>more</em> directly than a lab does — at a keyboard a 1078 student can arrive at the right answer by running things until they look 1079 right, whereas on paper they have to actually know. 1080 </div> 1081 <p>Three rules keep it honest.</p> 1082 <ol> 1083 <li><strong>No question tests API recall.</strong> Asking for the 1084 arguments of <code>train_test_split</code> would test what a 1085 docstring is for. Syntax errors in handwritten code cost nothing; 1086 logic errors cost everything.</li> 1087 <li><strong>No question is answerable by reciting a definition.</strong> 1088 Every one requires applying it to a stated situation.</li> 1089 <li><strong>The paper states what it needs.</strong> There is no formula 1090 sheet, and none is required: any definition or standard result a 1091 question depends on is printed in the question itself. What you are 1092 expected to supply is the reasoning, never the recall.</li> 1093 </ol> 1094 <p class="small muted">
1094This is why closed book costs you nothing here. Each 1095 of the eighteen derivations is one you have <em>performed</em> — and a 1096 result you can rebuild in three lines is not a result you need to have 1097 memorised.</p> 1098 <p>What a written examination cannot see is whether the reasoning on the 1099 page is yours. That is what the oral is for — and it is why it is 1100 short, and why it can go to any part of the course.</p> 1101 1102 <h3>The written examination</h3> 1103 <p>The three parts below are weighted <em>within the paper</em>. The paper 1104 carries your mark: up to 27 on its own, and up to 30 if you sit the oral.</p> 1105 1106 <p>Every paper mixes three <strong>forms</strong> of question, and the parts 1107 below say what each one is <em>about</em> rather than what shape it takes: 1108 <strong>technical exercises</strong> worked on the page — a 1109 derivation, a gradient, an arithmetic check on a stated table; 1110 <strong>closed questions</strong>, where you choose between two or three 1111 stated options; and <strong>open questions</strong> answered in a few 1112 sentences. A closed choice is rarely enough on its own: where a question 1113 asks for the reason as well as the choice, both are needed for full 1114 marks.</p> 1115 <div class="tablewrap"> 1116 <table> 1117 <thead> 1118 <tr><th>Part</th><th>What it asks</th><th class="num">Weight<br><span class="th-sub">of the paper</span></th><th class="num">Time</th></tr> 1119 </thead> 1120 <tbody> 1121 <tr> 1122 <td><strong>A · The derivations</strong></td> 1123 <td>Derive, state, apply. The eighteen objects developed in the 1124 lectures â the normal equation, the bias–variance 1125 decomposition, variance reduction by averaging, reverse-mode 1126 differentiation, cross-entropy, NDCG, and the rest.</td> 1127 <td class="num">40%</td> 1128 <td class="num">~48 min</td> 1129 </tr> 1130 <tr> 1131 <td><strong>B · Choosing a method</strong></td> 1132 <td>A situation in four lines: this data, this constraint, this 1133 requirement. Which method, why that one rather than the obvious 1134 alternative, and what evidence would make you change your mind. 1135 <span class="muted">Some questions supply a short code excerpt and 1136 ask what it would take for its reported number to be 1137 trustworthy.</span></td> 1138 <td class="num">35%</td> 1139 <td class="num">~42 min</td> 1140 </tr> 1141 <tr> 1142 <td><strong>C · Reading results</strong></td> 1143 <td>Plots and tables â learning curves, per-fold scores, a confusion 1144 matrix, a precision/recall curve, a ranking. Say what they show, 1145 what they do not show, and which way the number moves if you 1146 change the stated thing.</td> 1147 <td class="num">25%</td> 1148 <td class="num">~30 min</td> 1149 </tr> 1150 </tbody> 1151 </table> 1152 </div> 1153 1154 <div class="panel"> 1155 <strong>Every exercise in the course, with its solution.</strong> Each 1156 deck ends with five questions in this style, answered on the following 1157 lecture’s deck. All 120 exercises are collected in one place: 1158 <!-- BEGIN EXERCISES_PANEL --> 1159 published once the course has run. <!-- END EXERCISES_PANEL --> 1160 </div> 1161 1162 <h3>Three specimen questions</h3> 1163 <p>One from each part, at the intended level and scale — each is 1164 calibrated to the time budget above.</p> 1165 <p class="small muted"><strong>Illustration only.</strong> These three show 1166 the kind of thing each part asks and roughly how much of it. They are not a 1167 template, not a syllabus, and not a promise: the questions on the paper you 1168 sit may differ in form, in topic, and in how they are put.</p> 1169 1170 <div class="specimen"> 1171 <h4>Part A — derive, then apply</h4> 1172 <p>Two regressors each have variance $\sigma^2$ and are correlated with 1173 coefficient $\rho$.</p> 1174 <ol class="tight"> 1175 <li>Show that the variance of their average is 1176 $\frac{\sigma^{2}(1+\rho)}{2}$.</li> 1177 <li>Bagging and a random forest differ in one respect. State it, and say 1178 which term in your expression it attacks.</li> 1179 <li>An ensemble of one hundred <em>identical</em> trees has $\rho = 1$. 1180 What does your expression predict, and is that the right answer?</li> 1181 </ol> 1182 </div> 1183 1184 <div class="specimen"> 1185 <h4>Part B — choosing a method</h4> 1186 <p>A colleague has <strong>4,000 labelled rows</strong> and 1187 <strong>sixty numeric features</strong>, many of them near-duplicates of 1188 one another. Every prediction has to be justified to a regulator.</p> 1189 <ol class="tight"> 1190 <li>Ordinary least squares, ridge, lasso, or a random forest — 1191 which do you fit first, and what specifically about the situation 1192 decides it?</li> 1193 <li>Name the one piece of evidence that would make you abandon that 1194 choice for one of the others.</li> 1195 </ol> 1196 <p>They send you this program, and the number it printed.</p> 1197<pre><code>best = None 1198for a in [0.01, 0.1, 1, 10, 100]: 1199 m = Ridge(alpha=a).fit(X_train, y_train) 1200 s = root_mean_squared_error(y_test, m.predict(X_test)) 1201 if best is None or s < best[1]: 1202 best = (a, s) 1203 1204print(f"best alpha={best[0]}, test RMSE={best[1]:,.0f}")</code></pre> 1205 <ol class="tight" start="3"> 1206 <li>This program never creates a validation set. <strong>Which object is 1207 doing that job?</strong></li> 1208 <li>As an estimate of the model’s error on new data, the printed 1209 RMSE is <strong>(i)</strong> too optimistic, <strong>(ii)</strong> too 1210 pessimistic, or <strong>(iii)</strong> neither. Choose one and give the 1211 reason in a sentence.</li> 1212 <li>Rewrite it so the printed number is honest. Two statements is 1213 enough.</li> 1214 </ol> 1215 <p class="small muted">Every part has a determinate answer — ridge, 1216 because the near-duplicate features make $\mathbf{X}^{\mathsf T}\mathbf{X}$ 1217 ill-conditioned and the regulator rules out the forest; <em>the test 1218 set</em>; <em>(i)</em>; a cross-validated search on the training data, 1219 scored once on the test set. Nothing here rewards a well-phrased opinion,
1220 and none of it can be answered without knowing what a validation set is 1221 for.</p> 1222 </div> 1223 1224 <div class="specimen"> 1225 <h4>Part C — reading results</h4> 1226 <p>Two teams report the following for the same dataset.</p> 1227 <div class="tablewrap"> 1228 <table> 1229 <thead><tr><th>Team</th><th class="num">Training RMSE</th><th class="num">10-fold CV RMSE</th></tr></thead> 1230 <tbody> 1231 <tr><td>A</td><td class="num">$0</td><td class="num">$68,574</td></tr> 1232 <tr><td>B</td><td class="num">$68,233</td><td class="num">$68,282</td></tr> 1233 </tbody> 1234 </table> 1235 </div> 1236 <ol class="tight"> 1237 <li>Name each pathology.</li> 1238 <li>Which model is more useful? Explain why that is not the same question 1239 as which is better fitted.</li> 1240 <li>One of these is helped by collecting more training data and the other 1241 is not. Which, and why?</li> 1242 </ol> 1243 </div> 1244 1245 <h3>The oral examination</h3> 1246 <p>Optional. Three questions, answered aloud, on any topic the course 1247 covered — there is no separate syllabus for it, and no list of 1248 questions to learn.</p> 1249 1250 <div class="cols-oral"> 1251 <div class="card"> 1252 <h4>What happens</h4> 1253 <ul class="plain"> 1254 <li>You decide <strong>after</strong> seeing your written mark. Until 1255 then there is nothing to opt into.</li> 1256 <li>Three questions, on any topic from the course.</li> 1257 <li>Seven to ten minutes, no notes, no computer.</li> 1258 </ul> 1259 </div> 1260 <div class="card"> 1261 <h4>What it is worth</h4> 1262 <ul class="plain"> 1263 <li><strong>Up.</strong> The only route to a mark above 27, and to 1264 lode.</li> 1265 <li><strong>Down.</strong> The same conversation can lower it. This is 1266 not a formality.</li> 1267 <li><strong>Worth sitting from below the cap too</strong> — the 1268 movement is measured from <em>your</em> mark, wherever it is.</li> 1269 </ul> 1270 </div> 1271 </div> 1272 <p class="small muted">Nothing is collected and nothing is graded as an 1273 artefact. The notebooks are examinable as <em>experience</em>: you are 1274 expected to be able to account for what they do and why, without the code in 1275 front of you.</p> 1276 <div class="panel panel-accent"> 1277 <strong>Rule one is checked twice.</strong> Never keep code you cannot 1278 explain. Part B of the paper puts a cell in front of every candidate; the 1279 oral puts a question in front of the ones who choose it. Both apply to 1280 code an assistant wrote for you exactly as they apply to code you typed. 1281 </div> 1282 1283 <p style="margin-top:1.75rem">Note what is absent: there are no marks for a 1284 notebook that runs. <strong>Evidence earns marks.</strong></p> 1285 </div> 1286</section> 1287 1288<!-- ===================================================================== --> 1289 1290<section id="textbook"> 1291 <div class="wrap"> 1292 <h2>Textbook and scope</h2> 1293 <p>Aurélien Géron, <em>Hands-On Machine Learning with Scikit-Learn and 1294 PyTorch</em>, O’Reilly, 2025 — <strong>Chapters 1–16</strong>, 1295 covering Lectures 1–18 and 23–24.</p> 1296 1297 <p><strong>Lectures 19–22</strong> — information retrieval and 1298 recommender systems — sit outside the book and are taught from the 1299 lecture notes. They are examinable on the same terms as everything else, 1300 and for those four lectures <strong>the notes below are the primary 1301 source</strong> — not a supplement to a chapter, because there is no 1302 chapter.</p> 1303 1304 <p class="small">Written notes exist for <strong>every</strong> lecture 1305 — the <em>Notes (PDF)</em> button that appears on each card above on 1306 the day of its lecture. For the other twenty they set out the 1307 lecture’s argument at length beside the chapter it is taught from; 1308 the four listed here are the ones with no chapter behind them.</p> 1309 1310 <div class="tablewrap"> 1311 <table> 1312 <thead><tr><th>Extended lecture notes</th><th>Dataset</th></tr></thead> 1313 <tbody> 1314 <tr><td>19 · Information retrieval: the lexical foundation</td> 1315 <td>SciFact (BEIR)</td></tr> 1316 <tr><td>20 · Information retrieval: dense retrieval</td> 1317 <td>SciFact (BEIR)</td></tr> 1318 <tr><td>21 · Recommender systems: from ratings to factors</td> 1319 <td>
1319MovieLens 1M</td></tr> 1320 <tr><td>22 · Recommender systems: neural, and evaluated honestly</td> 1321 <td>MovieLens 1M</td></tr> 1322 </tbody> 1323 </table> 1324 </div> 1325 1326 <div class="panel"> 1327 <strong>Scope discipline.</strong> The examinable surface is Chapters 1328 1–16 plus the notes for Lectures 19–22. Every section of every 1329 lecture is marked <em>examinable</em>, <em>not examinable — 1330 engineering</em>, or <em>beyond the syllabus, for context</em>, so you 1331 never have to guess. 1332 </div> 1333 1334 <h3>Chapter coverage</h3> 1335 <div class="tablewrap"> 1336 <table> 1337 <thead><tr><th>Chapter</th><th>Lectures</th></tr></thead> 1338 <tbody> 1339 <tr><td>1 · The machine learning landscape</td><td class="num">1</td></tr> 1340 <tr><td>2 · End-to-end project</td><td class="num">1, 2</td></tr> 1341 <tr><td>3 · Classification</td><td class="num">3</td></tr> 1342 <tr><td>4 · Training models</td><td class="num">4, 5</td></tr> 1343 <tr><td>5 · Decision trees</td><td class="num">6</td></tr> 1344 <tr><td>6 · Ensembles and random forests</td><td class="num">7</td></tr> 1345 <tr><td>7 · Dimensionality reduction</td><td class="num">8</td></tr> 1346 <tr><td>8 · Unsupervised learning</td><td class="num">8</td></tr> 1347 <tr><td>9 · Introduction to artificial neural networks</td><td class="num">9</td></tr> 1348 <tr><td>10 · Building networks with PyTorch</td><td class="num">10</td></tr> 1349 <tr><td>11 · Training deep networks</td><td class="num">11</td></tr> 1350 <tr><td>12 · Deep computer vision</td><td class="num">12, 13, 14</td></tr> 1351 <tr><td>13 · Sequences</td><td class="num">15, 16</td></tr> 1352 <tr><td>14 · NLP with RNNs and attention</td><td class="num">17, 18</td></tr> 1353 <tr><td>15 · Transformers</td><td class="num">18, 23, 24</td></tr> 1354 <tr><td>16 · Vision and multimodal transformers</td><td class="num">23, 24</td></tr> 1355 <tr><td><em>Lecture notes</em> · Information retrieval</td><td class="num">19, 20</td></tr> 1356 <tr><td><em>Lecture notes</em> · Recommender systems</td><td class="num">21, 22</td></tr> 1357 </tbody> 1358 </table> 1359 </div> 1360 </div> 1361</section> 1362 1363<!-- ===================================================================== --> 1364 1365<section id="practicalities"> 1366 <div class="wrap"> 1367 <h2>Practicalities</h2> 1368 <ul class="plain"> 1369 <li>Announcements go through 1370 <a href="https://classroom.google.com/c/MjU0MDA3MTk5MjJa">the course’s Google Classroom</a>; join it 1371 with your Sapienza account.</li> 1372 <li>Every lecture has a Colab notebook, linked from its own slides and 1373 from the list above from the morning of that lecture.</li> 1374 <li>Every notebook runs on a free CPU runtime, and says at the top 1375 roughly how long it takes.</li> 1376 <li>Every deck has a <strong>PDF</strong> button beside it, one page per 1377 slide. Printing a deck from the browser gives the same thing.</li> 1378 <li>In a deck: <code>M</code> opens the menu, <code>S</code> opens the 1379 speaker notes, <code>C</code> toggles the chalkboard, <code>F</code> 1380 goes full screen.</li> 1381 <li>Lectures refer to one another by number, so the material does not 1382 depend on the timetable.</li> 1383 <li>Nothing in any notebook is wrong on purpose.</li> 1384 </ul> 1385 </div> 1386</section> 1387 1388<footer> 1389 <div class="wrap"> 1390 <p class="draft">Work in progress — not yet announced to students</p> 1391 <p><strong>Applicazioni Informatiche del Machine Learning</strong> · 1392 BSc Mathematics of Artificial Intelligence · Sapienza Università di Roma</p> 1393 <p>Fabrizio Silvestri · 1394 <a href="https://fabsilvestri.github.io/">fabsilvestri.github.io</a></p> 1395 <p>Slides built with <a href="https://revealjs.com/">reveal.js</a> and 1396 <a href="https://katex.org/">KaTeX</a>
1396, both vendored locally so the decks 1397 work without a network connection.</p> 1398 </div> 1399</footer> 1400 1401<!-- KaTeX, vendored locally so the page needs no network beyond this host -->
1402<script defer src="assets/js/site-nav.js"></script>
vendor: 1 bytes, line 1402
1402
1403<script defer src="assets/js/reveal-material.js"></script>
vendor: 1 bytes, line 1403
1403
1404<script defer src="lib/katex/dist/katex.min.js"></script>
vendor: 1 bytes, line 1404
1404
1405<script defer src="lib/katex/dist/contrib/auto-render.min.js" 1406 onload="renderMathInElement(document.body, { 1407 delimiters: [ 1408 {left: '$$', right: '$$', display: true}, 1409 {left: '$', right: '$', display: false} 1410 ], 1411 throwOnError: false 1412 });"></script>
1412 1413 1414</body> 1415</html>
Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.