1<!DOCTYPE html> 2<html> 3 <head> 4 <meta charset="utf-8"> 5<meta name="viewport" content="width=device-width, initial-scale=1"> 6<meta name="twitter:card" content="summary_large_image"> 7<meta name="twitter:site" content="@janmatthis"> 8<meta name="twitter:creator" content="@janmatthis"> 9<meta name="twitter:title" content="Benchmarking Simulation-Based Inference"> 10<meta name="twitter:description" content="Project website with paper summary, interactive results, and code."> 11<meta name="twitter:image" content="https://sbi-benchmark.github.io/assets/img/social_card.jpg"> 12<title>SBI Benchmark</title> 13<link rel="stylesheet" href="https://sbi-benchmark.github.io/assets/css/bulma.min.css"> 14<link rel="stylesheet" href="https://sbi-benchmark.github.io/assets/css/custom.css"> 15<link rel="stylesheet" href="//cdnjs.cloudflare.com/ajax/libs/highlight.js/10.4.1/styles/github.min.css">
16<script src="//cdnjs.cloudflare.com/ajax/libs/highlight.js/10.4.1/highlight.min.js"></script>
16 17 </head> 18 <body class="has-navbar-fixed-top"> 19 <div id="wrapper"> 20 <nav class="navbar is-fixed-top is-white is-transparent has-shadow"> 21 <div class="navbar-brand"> 22 <a class="navbar-item" href="/"> 23 <!-- <b>SBI Benchmark</b> --> 24 <!-- <img src="img/logo.png" /> --> 25 </a> 26 <div class="navbar-burger burger" data-target="navbar"> 27 <span></span> 28 <span></span> 29 <span></span> 30 </div> 31 </div> 32 33 <div id="navbar" class="navbar-menu"> 34 <div class="navbar-start"> 35 36 37 <a class="navbar-item is-current" href="."> 38 Overview 39 </a> 40 41 42 43 <div class="navbar-item has-dropdown is-hoverable"> 44 <a class="navbar-link not-current" href="."> 45 Interactive Results 46 </a> 47 <div class="navbar-dropdown is-boxed"> 48 49 <a class="navbar-item not-current" href="https://sbi-benchmark-streamlit-metrics-5h5b07.streamlit.app" target="_blank"> 50 Metrics 51 </a> 52 53 <a class="navbar-item not-current" href="https://sbi-benchmark-streamlit-posteriors-3rtkv1.streamlit.app" target="_blank"> 54 Posteriors 55 </a> 56 57 <a class="navbar-item not-current" href="https://sbi-benchmark-streamlit-correlations-y32m8z.streamlit.app" target="_blank"> 58 Correlations 59 </a> 60 61 </div> 62 </div> 63 64 65 66 <a class="navbar-item not-current" href="code/"> 67 Code & Reproducibility 68 </a> 69 70 71 72 <a class="navbar-item not-current" href="contribute/"> 73 Contributions 74 </a> 75 76 77 </div> 78 79 <div class="navbar-end"> 80 <div class="navbar-item"> 81 <div class="field is-grouped"> 82 <p class="control"> 83 <a class="button is-dark" href="https://github.com/sbi-benchmark/sbibm" target="_blank"> 84 <span class="icon"> 85 <img src="https://sbi-benchmark.github.io/img/github_white.svg" /> 86 </span> 87 <span style="padding-left: 6px">Code</span> 88 </a> 89 </p> 90 <p class="control"> 91 <a class="button is-dark" href="https://arxiv.org/abs/2101.04653" target="_blank"> 92 <span class="icon" style="padding: 1px;"> 93 <img src="https://sbi-benchmark.github.io/img/adobeacrobatreader_white.svg" /> 94 </span> 95 <span style="padding-left: 6px"> 96 Paper 97 </span> 98 </a> 99 </p> 100 </div> 101 </div> 102 </div> 103 </div> 104</nav> 105 <section class="section"> 106 <div class="container"> 107 <div class="columns"> 108 <div class="column is-7 is-offset-2"> 109 <div class="content"> 110 <h1 id="benchmarking-simulation-based-inference">Benchmarking Simulation-Based Inference</h1> 111<p>Many domains of science make use of numerical simulators for which (1) it is easy to simulate but (2) the likelihood is unavailable, making it hard to perform statistical identification of model parameters consistent with observed data. <a href="https://www.pnas.org/content/early/2020/05/28/1912789117" target="_blank">Simulation-based inference (SBI)</a> deals with this 'likelihood-free' setting. Although recent advances have led to a large number of SBI algorithms, a public benchmark for such algorithms has been lacking: We set out to fill this gap, carefully select tasks and metrics, and evaluate several canonical algorithms. Through this website you can explore all results of our <a href="http://proceedings.mlr.press/v130/lueckmann21a.html" target="_blank">manuscript</a>, including comparisons on all metrics and plotting of posteriors for each of more than 10 000 runs.</p> 112<p>Keep reading for a brief summary of the manuscript, or jump right into interactive results (through the menu on top). We provide a framework to benchmark your own algorithms, code and results for reproducibility, and invite contributions.</p> 113<h2 id="our-motivation-and-methods">Our Motivation and Methods</h2> 114<p>Open benchmarks can be an important component of transparent and reproducible computational research. However, a benchmark framework for SBI has been lacking, possibly due to the challenging endeavour of designing benchmarking tasks and defining suitable performance metrics.</p> 115<p>We selected a set of initial algorithms representing four distinct approaches to SBI, analyzed multiple performance metrics which have been used in the literature, and implemented ten tasks, including ones popular in the field.</p> 116<p><img alt="" src="/img/algorithms.png" /></p> 117<p><small><em>Overview of algorithms. Classification and schemes following <a href="https://www.pnas.org/content/early/2020/05/28/1912789117" target="_blank">Cranmer et al. (2020)</a>.</em></small></p> 118<p><strong>Algorithms</strong>. We compare algorithms belonging to four distinct approaches to SBI: Classical ABC approaches as well as model-based approaches approximating likelihoods, posteriors, or density ratios. We contrast algorithms that use the prior distribution to propose parameters against ones that sequentially adapt the proposal. Keeping our initial selection of algorithms focused allowed us to carefully consider implementation details and hyperparameters.</p> 119<p><strong>Metrics.</strong> The shortcomings of commonly used metrics (see paper for details) led us to focus on tasks for which a likelihood can be evaluated, which allowed us to calculate reference (âground-truthâ) posteriors. These reference posteriors are made available to allow rapid evaluation of SBI algorithms. While we compare algorithms in terms of classifier 2-sample tests (C2ST) in the paper, the website allows comparisons in terms of all metrics we considered.</p> 120<p><strong>Tasks.</strong> We focused on eight purely statistical problems and two problems relevant in applied domains, with diverse dimensionalities of parameters and data.</p> 121<h2 id="key-findings">Key Findings</h2> 122<p>The full potential of the benchmark will be realized when it is populated with additional community-contributed algorithms and tasks. However, our initial version already provides useful insights:</p> 123<ul> 124<li>the choice of performance metric is critical (commonly used 125 metrics such as the likelihood of true parameters or Maximum Mean Discrepancy with median heuristic have important shortcomings);</li> 126<li>the performance of the algorithms on some tasks leaves substantial room for improvement;</li> 127<li>sequential estimation generally improves sample efficiency;</li> 128<li>for small and moderate simulation budgets, neural-network based approaches outperform classical ABC algorithms, confirming recent progress in the field;</li> 129<li>but that there is no algorithm to rule them all.</li> 130</ul> 131<p>The performance ranking of algorithms is task-dependent, pointing to a need for better guidance or automated procedures for choosing which algorithm to use when. In the manuscript, we included
131some considerations and recommendations for practitioners, based on our current results and understanding, and dedicated a page to discussing various limitations.</p> 132<p><strong>Find out more in our manuscript, <a href="http://proceedings.mlr.press/v130/lueckmann21a.html">available through PMLR</a> or:</strong></p> 133<blockquote> 134<p><a href="https://arxiv.org/abs/2101.04653">arXiv.org/abs/2101.04653</a></p> 135</blockquote> 136<p><em>We believe that the full potential of the benchmark will be revealed as more researchers participate and contribute. In order to facilitate this process, we provide a benchmarking framework which is designed to be highly extensible and easily used. See <a href="code/">Code & Reproducibility</a> for explanations and examples.</em></p> 137<h2 id="3-minute-summary">3-Minute Summary</h2> 138<div id="presentation-embed-38952956"></div>
139<script src='https://slideslive.com/embed_presentation.js'></script>
vendor: 1 bytes, line 139
139
140<script> 141 embed = new SlidesLiveEmbed('presentation-embed-38952956', { 142 presentationId: '38952956', 143 autoPlay: false, 144 verticalEnabled: true, 145 zoomRatio: 0.3 146 }); 147</script>
147 148 149<h2 id="citation">Citation</h2> 150<pre><code class="language-bibtex">@InProceedings{lueckmann2021benchmarking, 151 title = {Benchmarking Simulation-Based Inference}, 152 author = {Lueckmann, Jan-Matthis and Boelts, Jan and Greenberg, David and Goncalves, Pedro and Macke, Jakob}, 153 booktitle = {Proceedings of The 24th International Conference on Artificial Intelligence and Statistics}, 154 pages = {343--351}, 155 year = {2021}, 156 editor = {Banerjee, Arindam and Fukumizu, Kenji}, 157 volume = {130}, 158 series = {Proceedings of Machine Learning Research}, 159 month = {13--15 Apr}, 160 publisher = {PMLR} 161} 162</code></pre> 163<h2 id="support">Support</h2> 164<p>This work was supported by the German Research Foundation (DFG; SFB 1233 PN 276693517, SFB 1089, SPP 2041, Germanyâs Excellence Strategy â EXC number 2064/1 PN 390727645) and the German Federal Ministry of Education and Research (BMBF; project â<a href="https://fit.uni-tuebingen.de/Project/Details?id=9199">ADIMEM</a>â, FKZ 01IS18052 A-D).</p> 165 </div> 166 </div> 167 168 <div class="column is-3 is-hidden-touch"> 169 <aside class="menu has-navbar-fixed-top"> 170 171 172 <ul class="menu-list"> 173 174 <li> 175 176 <ul> 177 178 <li> 179 <a href="#our-motivation-and-methods" class="is-size-7">Our Motivation and Methods</a> 180 181</li> 182 183 <li> 184 <a href="#key-findings" class="is-size-7">Key Findings</a> 185 186</li> 187 188 <li> 189 <a href="#3-minute-summary" class="is-size-7">3-Minute Summary</a> 190 191</li> 192 193 <li> 194 <a href="#citation" class="is-size-7">Citation</a> 195 196</li> 197 198 <li> 199 <a href="#support" class="is-size-7">Support</a> 200 201</li> 202 203 </ul> 204 205</li> 206 207 </ul> 208 209 210</aside> 211 </div> 212 213 </div> 214 </div> 215 </section> 216 </div> 217 <footer class="footer"> 218 <div class="has-text-centered"> 219 <p> 220 If you have questions or comments, please do not hesitate to <a href="mailto:[email protected],[email protected]">contact us</a>. 221 </p> 222 </div> 223</footer> 224 </body> 225
225<script type="text/javascript"> 226 document.addEventListener('DOMContentLoaded', () => { 227 228 // Get all "navbar-burger" elements 229 const $navbarBurgers = Array.prototype.slice.call(document.querySelectorAll('.navbar-burger'), 0); 230 231 // Check if there are any navbar burgers 232 if ($navbarBurgers.length > 0) { 233 234 // Add a click event on each of them 235 $navbarBurgers.forEach( el => { 236 el.addEventListener('click', () => { 237 238 // Get the target from the "data-target" attribute 239 const target = el.dataset.target; 240 const $target = document.getElementById(target); 241 242 // Toggle the "is-active" class on both the "navbar-burger" and the "navbar-menu" 243 el.classList.toggle('is-active'); 244 $target.classList.toggle('is-active'); 245 246 }); 247 }); 248 } 249 250 }); 251</script>
vendor: 1 bytes, line 251
251
252<script>hljs.initHighlightingOnLoad();</script>
252 253</html>
Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.