1<!DOCTYPE html> <html lang="en"> <head> <meta http-equiv="Content-Type" content="text/html; charset=UTF-8"> <meta charset="utf-8"> <meta name="viewport" content="width=device-width, initial-scale=1, shrink-to-fit=no"> <meta http-equiv="X-UA-Compatible" content="IE=edge"> <title> Lorenzo Pacchiardi </title> <meta name="author" content="Lorenzo Pacchiardi"> <meta name="description" content="Lorenzo Pacchiardi, Assistant Research Professor at the Leverhulme Centre for the Future of Intelligence, University of Cambridge. "> <meta name="keywords" content="jekyll, jekyll-theme, academic-website, portfolio-website"> <link rel="stylesheet" href="/assets/css/bootstrap.min.css?a4b3f509e79c54a512b890d73235ef04"> <link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/[email protected]/css/mdb.min.css" integrity="sha256-jpjYvU3G3N6nrrBwXJoVEYI/0zw8htfFnhT9ljN3JJw=" crossorigin="anonymous"> <link defer rel="stylesheet" href="https://unpkg.com/[email protected]/dist/bootstrap-table.min.css"> <link rel="stylesheet" href="/assets/css/academicons.min.css?f0b7046b84e425c55f3463ac249818f5"> <link rel="stylesheet" type="text/css" href="https://fonts.googleapis.com/css?family=Roboto:300,400,500,700|Roboto+Slab:100,300,400,500,700|Material+Icons"> <link rel="stylesheet" href="/assets/css/jekyll-pygments-themes-github.css?591dab5a4e56573bf4ef7fd332894c99" media="" id="highlight_theme_light"> <link rel="shortcut icon" href="data:image/svg+xml,<svg%20xmlns=%22http://www.w3.org/2000/svg%22%20viewBox=%220%200%20100%20100%22><text%20y=%22.9em%22%20font-size=%2290%22>%F0%9F%91%A8%E2%80%8D%F0%9F%92%BB</text></svg>"> <link rel="stylesheet" href="/assets/css/main.css?d41d8cd98f00b204e9800998ecf8427e"> <link rel="canonical" href="http://lorenzopacchiardi.me//"> <link rel="stylesheet" href="/assets/css/jekyll-pygments-themes-native.css?5847e5ed4a4568527aa6cfab446049ca" media="none" id="highlight_theme_dark">
1<script src="/assets/js/theme.js?bf50d6d9dd867d3e0f3b0add94449649"></script>
1 </head> <body class="fixed-top-nav "> <header> <nav id="navbar" class="navbar navbar-light navbar-expand-sm fixed-top" role="navigation"> <div class="container"> <div class="navbar-brand social"> <a href="mailto:%6C%70%36%36%36@%63%61%6D.%61%63.%75%6B" title="email"><i class="fa-solid fa-envelope"></i></a> <a href="https://orcid.org/0000-0003-4760-7638" title="ORCID" rel="external nofollow noopener" target="_blank"><i class="ai ai-orcid"></i></a> <a href="https://scholar.google.com/citations?user=9EAb0uEAAAAJ" title="Google Scholar" rel="external nofollow noopener" target="_blank"><i class="ai ai-google-scholar"></i></a> <a href="https://github.com/LoryPack" title="GitHub" rel="external nofollow noopener" target="_blank"><i class="fa-brands fa-github"></i></a> <a href="https://www.linkedin.com/in/lorenzo-pacchiardi" title="LinkedIn" rel="external nofollow noopener" target="_blank"><i class="fa-brands fa-linkedin"></i></a> <a href="https://twitter.com/LPacchiardi" title="X" rel="external nofollow noopener" target="_blank"><i class="fa-brands fa-x-twitter"></i></a> </div> <button class="navbar-toggler collapsed ml-auto" type="button" data-toggle="collapse" data-target="#navbarNav" aria-controls="navbarNav" aria-expanded="false" aria-label="Toggle navigation"> <span class="sr-only">Toggle navigation</span> <span class="icon-bar top-bar"></span> <span class="icon-bar middle-bar"></span> <span class="icon-bar bottom-bar"></span> </button> <div class="collapse navbar-collapse text-right" id="navbarNav"> <ul class="navbar-nav ml-auto flex-nowrap"> <li class="nav-item active"> <a class="nav-link" href="/">about <span class="sr-only">(current)</span> </a> </li> <li class="nav-item "> <a class="nav-link" href="/publications/">publications </a> </li> <li class="nav-item "> <a class="nav-link" href="/talks/">talks </a> </li> <li class="nav-item "> <a class="nav-link" href="/blog/">blog </a> </li> <li class="nav-item "> <a class="nav-link" href="/resources/">resources </a> </li> <li class="toggle-container"> <button id="light-toggle" title="Change theme"> <i class="fa-solid fa-moon"></i> <i class="fa-solid fa-sun"></i> </button> </li> </ul> </div> </div> </nav> <progress id="progress" value="0"> <div class="progress-container"> <span class="progress-bar"></span> </div> </progress> </header> <div class="container mt-5" role="main"> <div class="post"> <header class="post-header"> <h1 class="post-title"> <span class="font-weight-bold">Lorenzo</span> Pacchiardi </h1> <p class="desc"></p> </header> <article> <div class="profile float-right"> <figure> <picture> <source class="responsive-img-srcset" srcset="/assets/img/prof_pic-480.webp 480w,/assets/img/prof_pic-800.webp 800w,/assets/img/prof_pic-1400.webp 1400w," sizes="(min-width: 800px) 231.0px, (min-width: 576px) 30vw, 95vw" type="image/webp"> <img src="/assets/img/prof_pic.jpg?ebed4ec7ffbef62222a1c9bdc058971f" class="img-fluid z-depth-1 rounded" width="100%" height="auto" alt="prof_pic.jpg" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </source></picture> </figure> </div> <div class="clearfix"> <p><em>Assistant Research Professor, University of Cambridge</em></p> <p>I am an Assistant Research Professor at the <a href="http://lcfi.ac.uk/" rel="external nofollow noopener" target="_blank">Leverhulme Centre for the Future of Intelligence</a> at the University of Cambridge. I lead a research project (funded by <a href="https://www.openphilanthropy.org/" rel="external nofollow noopener" target="_blank">Open Philanthropy</a>) on developing a benchmark for measuring the ability of LLMs to perform data science tasks. I am more broadly interested in <a href="https://arxiv.org/abs/2502.15620" rel="external nofollow noopener" target="_blank">AI evaluation</a>, particularly in <a href="https://arxiv.org/abs/2502.14445" rel="external nofollow noopener" target="_blank">predictability</a> and <a href="https://arxiv.org/abs/2503.06378" rel="external nofollow noopener" target="_blank">cognitive evaluation</a>, and I closely collaborate with <a href="http://josephorallo.webs.upv.es/" rel="external nofollow noopener" target="_blank">Prof José Hernández-Orallo</a> and <a href="http://lcfi.ac.uk/people/lucy-cheke/" rel="external nofollow noopener" target="_blank">Prof Lucy Cheke</a>. I contribute to the <a href="https://aievaluation.substack.com/" rel="external nofollow noopener" target="_blank">AI evaluation newsletter</a>. I was faculty on the first edition of the <a href="https://ai-evaluation.org/" rel="external nofollow noopener" target="_blank">International Programme on AI Evaluation</a>, teaching modules on benchmarking and predictive evaluation.</p> <p>I am deeply familiar with EU AI policy (having been involved in several initiatives) and am currently part of the <a href="https://digital-strategy.ec.europa.eu/en/policies/ai-scientific-panel" rel="external nofollow noopener" target="_blank">AI Act scientific panel</a>, a 60-expert body advising the EU AI Office on general-purpose AI model risks, classification, and methodology. I am one of the co-founders of the Italian AI policy think tank <a href="https://www.cepte.it/" rel="external nofollow noopener" target="_blank">CePTE</a>.</p> <p>I am also a board member at <a href="https://www.meridiancambridge.org/" rel="external nofollow noopener" target="_blank">Meridian</a>, which runs AI safety and biosecurity programmes and hosts researchers, students and professionals in high-impact careers in central Cambridge. I also collaborate with <a href="https://www.unjournal.org/" rel="external nofollow noopener" target="_blank">The Unjournal</a> to make impactful research more rigorous, and I co-founded <a href="https://academicjobsitaly.com/" rel="external nofollow noopener" target="_blank">AcademicJobsItaly.com</a> to make the Italian academic job market more accessible.</p> <p>I previously worked on <a href="https://arxiv.org/abs/2309.15840" rel="external nofollow noopener" target="_blank">detecting lying in large language models</a> with <a href="https://owainevans.github.io/" rel="external nofollow noopener" target="_blank">Dr Owain Evans</a> (through the MATS programme) and on <a href="https://artificialintelligenceact.eu/standard-setting/" rel="external nofollow noopener" target="_blank">technical standards for AI</a> for the <a href="https://artificialintelligenceact.eu/" rel="external nofollow noopener" target="_blank">
1EU AI Act</a> at the <a href="https://futureoflife.org/" rel="external nofollow noopener" target="_blank">Future of Life Institute</a>. I have also shortly advised <a href="https://www.rand.org/" rel="external nofollow noopener" target="_blank">RAND</a> on AI evaluation.</p> <p>I obtained a PhD in Statistics and Machine Learning at Oxford, during which I worked on Bayesian simulation-based inference, generative models and probabilistic forecasting (with applications to meteorology). My supervisors were Prof. <a href="https://warwick.ac.uk/fac/sci/statistics/staff/academic-research/dutta/" rel="external nofollow noopener" target="_blank">Ritabrata Dutta</a> (Uni. Warwick) and Prof. <a href="https://www.stats.ox.ac.uk/people/geoff-nicholls" rel="external nofollow noopener" target="_blank">Geoff Nicholls</a> (Uni. Oxford).</p> <p>Before my PhD studies, I obtained a Bachelorâs degree in Physical Engineering from Politecnico di Torino (Italy) and an MSc in Physics of Complex Systems from a joint programme between Politecnico di Torino, <a href="https://www.sissa.it/" rel="external nofollow noopener" target="_blank">SISSA</a> and <a href="https://www.ictp.it/" rel="external nofollow noopener" target="_blank">ICTP</a> (Italy), and Université Paris-Sud (France). I did my MSc thesis at <a href="https://lighton.ai/" rel="external nofollow noopener" target="_blank">LightOn</a>, a machine learning startup in Paris.</p> </div> <h2> <a href="/news/" style="color: inherit">news</a> </h2> <div class="news"> <div class="table-responsive" style="max-height: 60vw"> <table class="table table-sm table-borderless"> <tr> <th scope="row" style="width: 20%">May 19, 2026</th> <td> I co-authored a foundational <a href="https://doi.org/10.5281/zenodo.20344324" rel="external nofollow noopener" target="_blank">preprint</a> on how AI evaluation should adapt for Continual Learning AI. </td> </tr> <tr> <th scope="row" style="width: 20%">Apr 01, 2026</th> <td> Our paper <a href="https://www.nature.com/articles/s41586-026-10303-2" rel="external nofollow noopener" target="_blank">âGeneral Scales Unlock AI Evaluation with Explanatory and Predictive Powerâ</a> has been accepted and published in <a href="https://www.nature.com/" rel="external nofollow noopener" target="_blank"><em>Nature</em></a>! ð </td> </tr> <tr> <th scope="row" style="width: 20%">Mar 01, 2026</th> <td> Our paper <a href="https://openreview.net/forum?id=1QcY6LPcdQ" rel="external nofollow noopener" target="_blank">âNo Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probesâ</a> has been accepted to the <a href="https://openreview.net/group?id=ICLR.cc/2026/Workshop/Trustworthy_AI" rel="external nofollow noopener" target="_blank">ICLR 2026 Trustworthy AI workshop</a>! ð </td> </tr> <tr> <th scope="row" style="width: 20%">May 16, 2025</th> <td> Our <a href="https://ai-evaluation-paradigms.github.io/" rel="external nofollow noopener" target="_blank">survey on AI evaluation</a> was accepted at <a href="https://2025.ijcai.org/" rel="external nofollow noopener" target="_blank">IJCAI 2025 survey track</a> and our <a href="https://predictaboard.github.io/" rel="external nofollow noopener" target="_blank">PredictaBoard</a> was accepted at ACL 2025 Findings. <img class="emoji" title=":tada:" alt=":tada:" src="https://github.githubassets.com/images/icons/emoji/unicode/1f389.png" height="20" width="20"> </td> </tr> <tr> <th scope="row" style="width: 20%">Mar 11, 2025</th> <td> Our new <a href="https://arxiv.org/abs/2503.06378" rel="external nofollow noopener" target="_blank">preprint</a> shows how to extract the most predictive and explanatory power from AI benchmarks by automatically annotating the demands posed by each question. Check it out! </td> </tr> </table> </div> </div> <h2> <a href="/publications/" style="color: inherit">selected publications</a> </h2> <div class="publications"> <ol class="bibliography"> <li> <div class="row"> <div class="col-sm-2 abbr"> <abbr class="badge">Nature</abbr> </div> <div id="zhou2025generalscalesunlockai" class="col-sm-8"> <div class="title">General Scales Unlock AI Evaluation with Explanatory and Predictive Power</div> <div class="author"> Lexin Zhou, <em>Lorenzo Pacchiardi</em>, Fernando MartÃnez-Plumed, Katherine M. Collins, Yael Moros-Daval, Seraphina Zhang, Qinlin Zhao, Yitian Huang, Luning Sun, Jonathan E. Prunty, Zongqian Li, Pablo Sánchez-GarcÃa, Kexin Jiang Chen, Pablo A. M. Casares, Jiyun Zu, John Burden, Behzad Mehrbakhsh, David Stillwell, Manuel Cebrian, Jindong Wang, Peter Henderson, Sherry Tongshuang Wu, Patrick C. Kyllonen, Lucy Cheke, Xing Xie, and José Hernández-Orallo </div> <div class="periodical"> <em>Nature</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a class="abstract btn btn-sm z-depth-0" role="button">Abs</a> <a class="bibtex btn btn-sm z-depth-0" role="button">Bib</a> <a href="https://kinds-of-intelligence-cfi.github.io/ADELE/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">HTML</a> <a href="https://www.nature.com/articles/s41586-026-10303-2.pdf" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://github.com/Kinds-of-Intelligence-CFI/ADeLe-AIEvaluation" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Code</a> </div> <div class="abstract hidden"> <p>Ensuring safe and effective use of artificial intelligence (AI) requires understanding and anticipating its performance on new tasks, from advanced scientific challenges to transformed workplace activities. So far, benchmarking has guided progress in AI but has offered limited explanatory and predictive power for general-purpose AI systems, attributed to limited transferability across specific tasks. Here we introduce general scales for AI evaluation that elicit demand profiles explaining what capabilities common AI benchmarks truly measure, extract ability profiles quantifying the general strengths and limits of AI systems and robustly predict AI performance for new task instances. Our fully automated methodology builds on 18 rubrics, capturing a broad range of cognitive and intellectual demands, which place different task instances on the same general scales, illustrated on 15 large language models (LLMs) and 63 tasks. Both the demand and the ability profiles on these scales bring new insights such as construct val
1idity through benchmark sensitivity and specificity and explain conflicting claims about whether AI has reasoning capabilities. Ultimately, high predictive power at the instance level becomes possible using the general scales, providing superior estimates over strong black-box baseline predictors, especially in out-of-distribution settings (new tasks and benchmarks). The scales, rubrics, battery, techniques and results presented here constitute a solid foundation for a science of AI evaluation, underpinning the reliable deployment of AI in the years ahead.</p> </div> <div class="bibtex hidden"> <figure class="highlight"><pre><code class="language-bibtex" data-lang="bibtex"><span class="nc">@article</span><span class="p">{</span><span class="nl">zhou2025generalscalesunlockai</span><span class="p">,</span> 2 <span class="na">title</span> <span class="p">=</span> <span class="s">{{General Scales Unlock AI Evaluation with Explanatory and Predictive Power}}</span><span class="p">,</span> 3 <span class="na">author</span> <span class="p">=</span> <span class="s">{Zhou, Lexin and Pacchiardi, Lorenzo and MartÃnez-Plumed, Fernando and Collins, Katherine M. and Moros-Daval, Yael and Zhang, Seraphina and Zhao, Qinlin and Huang, Yitian and Sun, Luning and Prunty, Jonathan E. and Li, Zongqian and Sánchez-GarcÃa, Pablo and Chen, Kexin Jiang and Casares, Pablo A. M. and Zu, Jiyun and Burden, John and Mehrbakhsh, Behzad and Stillwell, David and Cebrian, Manuel and Wang, Jindong and Henderson, Peter and Wu, Sherry Tongshuang and Kyllonen, Patrick C. and Cheke, Lucy and Xie, Xing and Hernández-Orallo, José}</span><span class="p">,</span> 4 <span class="na">year</span> <span class="p">=</span> <span class="s">{2026}</span><span class="p">,</span> 5 <span class="na">journal</span> <span class="p">=</span> <span class="s">{Nature}</span><span class="p">,</span> 6 <span class="na">volume</span> <span class="p">=</span> <span class="s">{652}</span><span class="p">,</span> 7 <span class="na">pages</span> <span class="p">=</span> <span class="s">{58-67}</span><span class="p">,</span> 8 <span class="na">url</span> <span class="p">=</span> <span class="s">{https://kinds-of-intelligence-cfi.github.io/ADELE/}</span><span class="p">,</span> 9<span class="p">}</span></code></pre></figure> </div> </div> </div> </li> <li> <div class="row"> <div class="col-sm-2 abbr"> <abbr class="badge">arXiv</abbr> </div> <div id="brundage2026frontieraiauditingrigorous" class="col-sm-8"> <div class="title">Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies</div> <div class="author"> Miles Brundage, Noemi Dreksler, Aidan Homewood, Sean McGregor, Patricia Paskov, Conrad Stosz, Girish Sastry, A. Feder Cooper, George Balston, Steven Adler, Stephen Casper, Markus Anderljung, Grace Werner, Soren Mindermann, Vasilios Mavroudis, Ben Bucknall, Charlotte Stix, Jonas Freund, <em>Lorenzo Pacchiardi</em>, Jose Hernandez-Orallo, Matteo Pistillo, Michael Chen, Chris Painter, Dean W. Ball, Cullen OâKeefe, Gabriel Weil, Ben Harack, Graeme Finley, Ryan Hassan, Scott Emmons, Charles Foster, Anka Reuel, Bri Treece, Yoshua Bengio, Daniel Reti, Rishi Bommasani, Cristian Trout, Ali Shahin Shamsabadi, Rajiv Dattani, Adrian Weller, Robert Trager, Jaime Sevilla, Lauren Wagner, Lisa Soder, Ketan Ramakrishnan, Henry Papadatos, Malcolm Murray, and Ryan Tovcimak </div> <div class="periodical"> 2026 </div> <div class="periodical"> </div> <div class="links"> <a class="abstract btn btn-sm z-depth-0" role="button">Abs</a> <a class="bibtex btn btn-sm z-depth-0" role="button">Bib</a> <a href="https://arxiv.org/abs/2601.11699" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">HTML</a> <a href="https://arxiv.org/pdf/2601.11699.pdf" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> <div class="abstract hidden"> <p>Frontier AI is becoming critical societal infrastructure, but outsiders lack reliable ways to judge whether leading developersâ safety and security claims are accurate and whether their practices meet relevant standards. Compared to other social and technological systems we rely on daily such as consumer products, corporate financial statements, and food supply chains, AI is subject to less rigorous third-party scrutiny along several dimensions. Ambiguity about whether AI systems are trustworthy can discourage deployment in some contexts where the technology could be benef
9icial, and make it more likely when itâs dangerous. Public transparency alone cannot close this gap: many safety- and security-relevant details are legitimately confidential and require expert interpretation. We define frontier AI auditing as rigorous third-party verification of frontier AI developersâ safety and security claims, and evaluation of their systems and practices against relevant standards, based on deep, secure access to non-public information. To make rigor legible and comparable, we introduce AI Assurance Levels (AAL-1 to AAL-4), ranging from time-bounded system audits to continuous, deception-resilient verification.</p> </div> <div class="bibtex hidden"> <figure class="highlight"><pre><code class="language-bibtex" data-lang="bibtex"><span class="nc">@misc</span><span class="p">{</span><span class="nl">brundage2026frontieraiauditingrigorous</span><span class="p">,</span> 10 <span class="na">title</span> <span class="p">=</span> <span class="s">{Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies}</span><span class="p">,</span> 11 <span class="na">author</span> <span class="p">=</span> <span class="s">{Brundage, Miles and Dreksler, Noemi and Homewood, Aidan and McGregor, Sean and Paskov, Patricia and Stosz, Conrad and Sastry, Girish and Cooper, A. Feder and Balston, George and Adler, Steven and Casper, Stephen and Anderljung, Markus and Werner, Grace and Mindermann, Soren and Mavroudis, Vasilios and Bucknall, Ben and Stix, Charlotte and Freund, Jonas and Pacchiardi, Lorenzo and Hernandez-Orallo, Jose and Pistillo, Matteo and Chen, Michael and Painter, Chris and Ball, Dean W. and O'Keefe, Cullen and Weil, Gabriel and Harack, Ben and Finley, Graeme and Hassan, Ryan and Emmons, Scott and Foster, Charles and Reuel, Anka and Treece, Bri and Bengio, Yoshua and Reti, Daniel and Bommasani, Rishi and Trout, Cristian and Shamsabadi, Ali Shahin and Dattani, Rajiv and Weller, Adrian and Trager, Robert and Sevilla, Jaime and Wagner, Lauren and Soder, Lisa and Ramakrishnan, Ketan and Papadatos, Henry and Murray, Malcolm and Tovcimak, Ryan}</span><span class="p">,</span> 12 <span class="na">year</span> <span class="p">=</span> <span class="s">{2026}</span><span class="p">,</span> 13 <span class="na">eprint</span> <span class="p">=</span> <span class="s">{2601.11699}</span><span class="p">,</span> 14 <span class="na">archiveprefix</span> <span class="p">=</span> <span class="s">{arXiv}</span><span class="p">,</span> 15 <span class="na">primaryclass</span> <span class="p">=</span> <span class="s">{cs.CY}</span><span class="p">,</span> 16 <span class="na">url</span> <span class="p">=</span> <span class="s">{https://arxiv.org/abs/2601.11699}</span><span class="p">,</span> 17 <span class="na">publisher</span> <span class="p">=</span> <span class="s">{arXiv}</span><span class="p">,</span> 18<span class="p">}</span></code></pre></figure> </div> </div> </div> </li> <li> <div class="row"> <div class="col-sm-2 abbr"> <abbr class="badge">Zenodo</abbr> </div> <div id="pacchiardi2026continuallearningrequiresevaluating" class="col-sm-8"> <div class="title">Continual Learning Requires Evaluating Trajectories</div> <div class="author"> <em>Lorenzo Pacchiardi</em>, Patricia Paskov, Seán à hÃigeartaigh, Fernando MartıÌnez-Plumed, Katherine M. Collins, Fazl Barez, Jonathan Prunty, Matteo Gabriel Mecattaf, Zafeirios Fountas, Risto Uuk, Sanmi Koyejo, Cozmin Ududec, and José Hernández-Orallo </div> <div class="periodical"> 2026 </div> <div class="periodical"> </div> <div class="links"> <a class="abstract btn btn-sm z-depth-0" role="button">Abs</a> <a class="bibtex btn btn-sm z-depth-0" role="button">Bib</a> <a href="https://cl-eval.github.io/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">HTML</a> <a href="https://zenodo.org/records/20344324/files/paper.pdf" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> <div class="abstract hidden"> <p>AI systems increasingly incorporate continual learning mechanisms allowing their behaviour to adapt after deployment, from (1) in-context learning and (2) memory features already in wide use to (3) post-deployment weight modification under research. We argue that, by treating AI systems as frozen artefacts whose performance and safety are assessed at release, current evaluation practices structurally ignore the behavioural trajectory of a system that continues to learn from experience. Our position is that evaluation of continual learning systems should be centred on behavioural trajectories, with the complementary goals of characterising the landscape of possible behaviours and forecasting how behaviour will evolve from a given set of experiences. This can be operationalised through trajectory elicitation sandboxes and predictive monitors that forecast behavioural evolution, but may face fundamental obstacles analogous to those seen in dynamical systems. These are best addressed by (1) applying trajectory-centred evaluation to todayâs continual learning systems and (2) relying on the resulting evidence to design systems amenable to it, yielding a virtuous cycle in which systems and their evaluations co-evolve.</p> </div> <div class="bibtex hidden"> <figure class="highlight"><pre><code class="language-bibtex" data-lang="bibtex"><span class="nc">@misc</span><span class="p">{</span><span class="nl">pacchiardi2026continuallearningrequiresevaluating</span><span class="p">,</span> 19 <span class="na">title</span> <span class="p">=</span> <span class="s">{Continual Learning Requires Evaluating Trajectories}</span><span class="p">,</span> 20 <span class="na">author</span> <span class="p">=</span> <span class="s">{Pacchiardi, Lorenzo and Paskov, Patricia and h{\'E}igeartaigh, Se{\'a}n {\'O} and Mart{\'\i}nez-Plumed, Fernando and Collins, Katherine M. and Barez, Fazl and Prunty, Jonathan and Mecattaf, Matteo Gabriel and Fountas, Zafeirios and Uuk, Risto and Koyejo, Sanmi and Ududec, Cozmin and Hern{\'a}ndez-Orallo, Jos{\'e}}</span><span class="p">,</span> 21 <span class="na">year</span> <span class="p">=</span> <span class="s">{2026}</span><span class="p">,</span> 22 <span class="na">doi</span> <span class="p">=</span> <span class="s">{10.5281/zenodo.20344324}</span><span class="p">,</span> 23 <span class="na">url</span> <span class="p">=</span> <span class="s">{https://cl-eval.github.io/}</span><span class="p">,</span> 24<span class="p">}</span></code></pre></figure> </div> </div> </div> </li> <li> <div class="row"> <div class="col-sm-2 abbr"> <abbr class="badge">NeurIPS Workshop</abbr> </div> <div id="pacchiardi2025a" class="col-sm-8"> <div class="title">A Framework for the Categorisation of General-Purpose AI Models under the EU AI Act<
24/div> <div class="author"> <em>Lorenzo Pacchiardi</em>, John Burden, Fernando MartıÌnez-Plumed, Jose Hernandez-Orallo, Emilia Gomez, and David Fernández-Llorca </div> <div class="periodical"> <em>In NeurIPS 2025 Workshop on Regulatable ML</em> , 2025 </div> <div class="periodical"> </div> <div class="links"> <a class="abstract btn btn-sm z-depth-0" role="button">Abs</a> <a class="bibtex btn btn-sm z-depth-0" role="button">Bib</a> <a href="https://openreview.net/forum?id=uE33aEsyX1" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">HTML</a> <a href="https://openreview.net/pdf?id=uE33aEsyX1" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> <div class="abstract hidden"> <p>We propose a framework for categorising AI models as General-Purpose AI (GPAI) models, based on their capabilities and generality, as defined in the European Union (EU) AI Act. Our framework breaks down the core components of the GPAI definition into measurable elements, focusing on four primary cognitive domains: Attention and Scan, Comprehension and Compositional Expression, Conceptualisation, Learning and Abstraction, and Quantitative and Logical Reasoning. We suggest using the Annotated Demand Levels (ADeLe) procedure to evaluate AI modelsâ capabilities in these domains, and provide a methodology for combining domain-level scores into a single measure of generality. The framework is illustrated with empirical results from existing models, and policy recommendations are made for selecting thresholds and metrics for GPAI categorisation.</p> </div> <div class="bibtex hidden"> <figure class="highlight"><pre><code class="language-bibtex" data-lang="bibtex"><span class="nc">@inproceedings</span><span class="p">{</span><span class="nl">pacchiardi2025a</span><span class="p">,</span> 25 <span class="na">Workshop</span><span class="err">},</span> 26 <span class="err">title</span> <span class="p">=</span> <span class="s">{A Framework for the Categorisation of General-Purpose {AI} Models under the {EU} {AI} Act}</span><span class="p">,</span> 27 <span class="na">author</span> <span class="p">=</span> <span class="s">{Pacchiardi, Lorenzo and Burden, John and Mart{\'\i}nez-Plumed, Fernando and Hernandez-Orallo, Jose and Gomez, Emilia and Fern{\'a}ndez-Llorca, David}</span><span class="p">,</span> 28 <span class="na">booktitle</span> <span class="p">=</span> <span class="s">{NeurIPS 2025 Workshop on Regulatable ML}</span><span class="p">,</span> 29 <span class="na">year</span> <span class="p">=</span> <span class="s">{2025}</span><span class="p">,</span> 30 <span class="na">url</span> <span class="p">=</span> <span class="s">{https://openreview.net/forum?id=uE33aEsyX1}</span><span class="p">,</span> 31<span class="p">}</span></code></pre></figure> </div> </div> </div> </li> <li> <div class="row"> <div class="col-sm-2 abbr"> <abbr class="badge">IJCAI</abbr> </div> <div id="burden2025paradigmsaievaluationmapping" class="col-sm-8"> <div class="title">Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture</div> <div class="author"> John Burden*, Marko TeÅ¡iÄ*, <em>Lorenzo Pacchiardi*</em>, and José Hernández-Orallo </div> <div class="periodical"> <em>IJCAI 2025 Survey Track</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a class="abstract btn btn-sm z-depth-0" role="button">Abs</a> <a class="bibtex btn btn-sm z-depth-0" role="button">Bib</a> <a href="https://ai-evaluation-paradigms.github.io/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">HTML</a> </div> <div class="abstract hidden"> <p>Research in AI evaluation has grown increasingly complex and multidisciplinary, attracting researchers with diverse backgrounds and objectives. As a result, divergent evaluation paradigms have emerged, often developing in isolation, adopting conflicting terminologies, and overlooking each otherâs contributions. This fragmentation has led to insular research trajectories and communication barriers both among different paradigms and with the general public, contributing to unmet expectations for deployed AI systems. To help bridge this insularity, in this paper we survey recent work in the AI evaluation landscape and identify six main paradigms. We characterise major recent contributions within each paradigm across key dimensions related to their goals, methodologies and research cultures. By clarifying the unique combination of questions and approaches associated with each paradigm, we aim to increase awareness of the breadth of current evaluation approaches and foster cross-pollination between different paradigms. We also identify potential gaps in the field to inspire future research directions.
31</p> </div> <div class="bibtex hidden"> <figure class="highlight"><pre><code class="language-bibtex" data-lang="bibtex"><span class="nc">@article</span><span class="p">{</span><span class="nl">burden2025paradigmsaievaluationmapping</span><span class="p">,</span> 32 <span class="na">title</span> <span class="p">=</span> <span class="s">{Paradigms of {AI} Evaluation: Mapping Goals, Methodologies and Culture}</span><span class="p">,</span> 33 <span class="na">author</span> <span class="p">=</span> <span class="s">{Burden*, John and TeÅ¡iÄ*, Marko and Pacchiardi*, Lorenzo and Hernández-Orallo, José}</span><span class="p">,</span> 34 <span class="na">year</span> <span class="p">=</span> <span class="s">{2025}</span><span class="p">,</span> 35 <span class="na">eprint</span> <span class="p">=</span> <span class="s">{2502.15620}</span><span class="p">,</span> 36 <span class="na">archiveprefix</span> <span class="p">=</span> <span class="s">{arXiv}</span><span class="p">,</span> 37 <span class="na">primaryclass</span> <span class="p">=</span> <span class="s">{cs.AI}</span><span class="p">,</span> 38 <span class="na">journal</span> <span class="p">=</span> <span class="s">{IJCAI 2025 Survey Track}</span><span class="p">,</span> 39 <span class="na">url</span> <span class="p">=</span> <span class="s">{https://arxiv.org/abs/2502.15620}</span><span class="p">,</span> 40<span class="p">}</span></code></pre></figure> </div> </div> </div> </li> <li> <div class="row"> <div class="col-sm-2 abbr"> <abbr class="badge">ICLR 2024</abbr> </div> <div id="pacchiardi2023catch" class="col-sm-8"> <div class="title">How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions</div> <div class="author"> <em>Lorenzo Pacchiardi*</em>, Alex J Chan*, Sören Mindermann, Ilan Moscovitz, Alexa Y Pan, Yarin Gal, Owain Evans, and Jan Brauner </div> <div class="periodical"> <em>The Twelfth International Conference on Learning Representations</em>, 2024 </div> <div class="periodical"> </div> <div class="links"> <a class="abstract btn btn-sm z-depth-0" role="button">Abs</a> <a class="bibtex btn btn-sm z-depth-0" role="button">Bib</a> <a href="https://openreview.net/forum?id=567BjxgaTp" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">HTML</a> <a href="https://github.com/lorypack/llm-liedetector" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Code</a> </div> <div class="abstract hidden"> <p>Large language models (LLMs) can "lie", which we define as outputting false statements despite "knowing" the truth in a demonstrable sense. LLMs might "lie", for example, when instructed to output misinformation. Here, we develop a simple lie detector that requires neither access to the LLMâs activations (black-box) nor ground-truth knowledge of the fact in question. The detector works by asking a predefined set of unrelated follow-up questions after a suspected lie, and feeding the LLMâs yes/no answers into a logistic regression classifier. Despite its simplicity, this lie detector is highly accurate and surprisingly general. When trained on examples from a single setting â prompting GPT-3.5 to lie about factual questions â the detector generalises out-of-distribution to (1) other LLM architectures, (2) LLMs fine-tuned to lie, (3) sycophantic lies, and (4) lies emerging in real-life scenarios such as sales. These results indicate that LLMs have distinctive lie-related behavioural patterns, consistent across architectures and contexts, which could enable general-purpose lie detection. </p> </div> <div class="bibtex hidden"> <figure class="highlight"><pre><code class="language-bibtex" data-lang="bibtex"><span class="nc">@article</span><span class="p">{</span><span class="nl">pacchiardi2023catch</span><span class="p">,</span> 41 <span class="na">title</span> <span class="p">=</span> <span class="s">{How to Catch an {AI} Liar: Lie Detection in Black-Box {LLM}s by Asking Unrelated Questions}</span><span class="p">,</span> 42 <span class="na">author</span> <span class="p">=</span> <span class="s">{Pacchiardi*, Lorenzo and Chan*, Alex J and Mindermann, S{\"o}ren and Moscovitz, Ilan and Pan, Alexa Y and Gal, Yarin and Evans, Owain and Brauner, Jan}</span><span class="p">,</span> 43 <span class="na">
43journal</span> <span class="p">=</span> <span class="s">{The Twelfth International Conference on Learning Representations}</span><span class="p">,</span> 44 <span class="na">year</span> <span class="p">=</span> <span class="s">{2024}</span><span class="p">,</span> 45 <span class="na">url</span> <span class="p">=</span> <span class="s">{https://openreview.net/forum?id=567BjxgaTp}</span><span class="p">,</span> 46<span class="p">}</span></code></pre></figure> </div> </div> </div> </li> </ol> </div> <div class="social"> <div class="contact-icons"> <a href="mailto:%6C%70%36%36%36@%63%61%6D.%61%63.%75%6B" title="email"><i class="fa-solid fa-envelope"></i></a> <a href="https://orcid.org/0000-0003-4760-7638" title="ORCID" rel="external nofollow noopener" target="_blank"><i class="ai ai-orcid"></i></a> <a href="https://scholar.google.com/citations?user=9EAb0uEAAAAJ" title="Google Scholar" rel="external nofollow noopener" target="_blank"><i class="ai ai-google-scholar"></i></a> <a href="https://github.com/LoryPack" title="GitHub" rel="external nofollow noopener" target="_blank"><i class="fa-brands fa-github"></i></a> <a href="https://www.linkedin.com/in/lorenzo-pacchiardi" title="LinkedIn" rel="external nofollow noopener" target="_blank"><i class="fa-brands fa-linkedin"></i></a> <a href="https://twitter.com/LPacchiardi" title="X" rel="external nofollow noopener" target="_blank"><i class="fa-brands fa-x-twitter"></i></a> </div> <div class="contact-note"></div> </div> </article> </div> </div> <footer class="fixed-bottom" role="contentinfo"> <div class="container mt-0"> © Copyright 2026 Lorenzo Pacchiardi. Powered by <a href="https://jekyllrb.com/" target="_blank" rel="external nofollow noopener">Jekyll</a> with <a href="https://github.com/alshedivat/al-folio" rel="external nofollow noopener" target="_blank">al-folio</a> theme. Hosted by <a href="https://pages.github.com/" target="_blank" rel="external nofollow noopener">GitHub Pages</a>. Photos from <a href="https://unsplash.com" target="_blank" rel="external nofollow noopener">Unsplash</a>. Last updated: September 21, 2026. </div> </footer>
46<script src="https://cdn.jsdelivr.net/npm/[email protected]/dist/jquery.min.js" integrity="sha256-/xUj+3OJU5yExlq6GSYGSHk7tPXikynS7ogEvDej/m4=" crossorigin="anonymous"></script>
46
46<script src="/assets/js/bootstrap.bundle.min.js"></script>
46
46<script src="https://cdn.jsdelivr.net/npm/[email protected]/js/mdb.min.js" integrity="sha256-NdbiivsvWt7VYCt6hYNT3h/th9vSTL4EDWeGs5SN3DA=" crossorigin="anonymous"></script>
46
46<script defer src="https://cdn.jsdelivr.net/npm/[email protected]/dist/masonry.pkgd.min.js" integrity="sha256-Nn1q/fx0H7SNLZMQ5Hw5JLaTRZp0yILA/FRexe19VdI=" crossorigin="anonymous"></script>
46
46<script defer src="https://cdn.jsdelivr.net/npm/imagesloaded@4/imagesloaded.pkgd.min.js"></script>
46
46<script defer src="/assets/js/masonry.js" type="text/javascript"></script>
46
46<script defer src="https://cdn.jsdelivr.net/npm/[email protected]/dist/medium-zoom.min.js" integrity="sha256-7PhEpEWEW0XXQ0k6kQrPKwuoIomz8R8IYyuU1Qew4P8=" crossorigin="anonymous"></script>
46
46<script defer src="/assets/js/zoom.js?85ddb88934d28b74e78031fd54cf8308"></script>
46
46<script defer src="https://unpkg.com/[email protected]/dist/bootstrap-table.min.js"></script>
46
46<script src="/assets/js/no_defer.js?2930004b8d7fcd0a8e00fdcfc8fc9f24"></script>
46
46<script defer src="/assets/js/common.js?4a129fbf39254905f505c7246e641eaf"></script>
46
46<script defer src="/assets/js/copy_code.js?7254ae07fe9cc5f3a10843e1c0817c9c" type="text/javascript"></script>
46
46<script async src="https://d1bxh8uas1mnw7.cloudfront.net/assets/embed.js"></script>
46
46<script async src="https://badge.dimensions.ai/badge.js"></script>
46
46<script type="text/javascript">window.MathJax={tex:{tags:"ams"}};</script>
46
46<script defer type="text/javascript" id="MathJax-script" src="https://cdn.jsdelivr.net/npm/[email protected]/es5/tex-mml-chtml.js"></script>
46
46<script defer src="https://polyfill.io/v3/polyfill.min.js?features=es6"></script>
46
46<script async src="https://www.googletagmanager.com/gtag/js?id="></script>
46
46<script>function gtag(){window.dataLayer.push(arguments)}window.dataLayer=window.dataLayer||[],gtag("js",new Date),gtag("config","");</script>
46
46<script type="text/javascript">function progressBarSetup(){"max"in document.createElement("progress")?(initializeProgressElement(),$(document).on("scroll",function(){progressBar.attr({value:getCurrentScrollPosition()})}),$(window).on("resize",initializeProgressElement)):(resizeProgressBar(),$(document).on("scroll",resizeProgressBar),$(window).on("resize",resizeProgressBar))}function getCurrentScrollPosition(){return $(window).scrollTop()}function initializeProgressElement(){let e=$("#navbar").outerHeight(!0);$("body").css({"padding-top":e}),$("progress-container").css({"padding-top":e}),progressBar.css({top:e}),progressBar.attr({max:getDistanceToScroll(),value:getCurrentScrollPosition()})}function getDistanceToScroll(){return $(document).height()-$(window).height()}function resizeProgressBar(){progressBar.css({width:getWidthPercentage()+"%"})}function getWidthPercentage(){return getCurrentScrollPosition()/getDistanceToScroll()*100}const progressBar=$("#progress");window.onload=function(){setTimeout(progressBarSetup,50)};</script>
46 </body> </html>
Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.