PageSourceSearch

https://sbi-benchmark.github.io/

html sbi-benchmark.github.io collected 2026-10-03 09:54:12 UTC 12,378 bytes, 253 lines download raw bytes

1<!DOCTYPE html>
2<html>
3  <head>
4    <meta charset="utf-8">
5<meta name="viewport" content="width=device-width, initial-scale=1">
6<meta name="twitter:card" content="summary_large_image">
7<meta name="twitter:site" content="@janmatthis">
8<meta name="twitter:creator" content="@janmatthis">
9<meta name="twitter:title" content="Benchmarking Simulation-Based Inference">
10<meta name="twitter:description" content="Project website with paper summary, interactive results, and code.">
11<meta name="twitter:image" content="https://sbi-benchmark.github.io/assets/img/social_card.jpg">
12<title>SBI Benchmark</title>
13<link rel="stylesheet" href="https://sbi-benchmark.github.io/assets/css/bulma.min.css">
14<link rel="stylesheet" href="https://sbi-benchmark.github.io/assets/css/custom.css">
15<link rel="stylesheet" href="//cdnjs.cloudflare.com/ajax/libs/highlight.js/10.4.1/styles/github.min.css">
16<script src="//cdnjs.cloudflare.com/ajax/libs/highlight.js/10.4.1/highlight.min.js"></script>
16
17  </head>
18  <body class="has-navbar-fixed-top">
19    <div id="wrapper">
20      <nav class="navbar is-fixed-top is-white is-transparent has-shadow">
21  <div class="navbar-brand">
22    <a class="navbar-item" href="/">
23      <!-- <b>SBI Benchmark</b> -->
24      <!-- <img src="img/logo.png" /> -->
25    </a>
26    <div class="navbar-burger burger" data-target="navbar">
27      <span></span>
28      <span></span>
29      <span></span>
30    </div>
31  </div>
32
33  <div id="navbar" class="navbar-menu">
34    <div class="navbar-start">
35      
36        
37          <a class="navbar-item is-current" href=".">
38            Overview
39          </a>
40        
41      
42        
43          <div class="navbar-item has-dropdown is-hoverable">
44              <a class="navbar-link not-current" href=".">
45                Interactive Results
46              </a>
47              <div class="navbar-dropdown is-boxed">
48                
49                  <a class="navbar-item not-current" href="https://sbi-benchmark-streamlit-metrics-5h5b07.streamlit.app" target="_blank">
50                    Metrics
51                  </a>
52                
53                  <a class="navbar-item not-current" href="https://sbi-benchmark-streamlit-posteriors-3rtkv1.streamlit.app" target="_blank">
54                    Posteriors
55                  </a>
56                
57                  <a class="navbar-item not-current" href="https://sbi-benchmark-streamlit-correlations-y32m8z.streamlit.app" target="_blank">
58                    Correlations
59                  </a>
60                
61              </div>
62          </div>
63        
64      
65        
66          <a class="navbar-item not-current" href="code/">
67            Code & Reproducibility
68          </a>
69        
70      
71        
72          <a class="navbar-item not-current" href="contribute/">
73            Contributions
74          </a>
75        
76      
77    </div>
78
79    <div class="navbar-end">
80      <div class="navbar-item">
81        <div class="field is-grouped">
82          <p class="control">
83            <a class="button is-dark" href="https://github.com/sbi-benchmark/sbibm" target="_blank">
84              <span class="icon">
85                <img src="https://sbi-benchmark.github.io/img/github_white.svg" />
86              </span>
87              <span style="padding-left: 6px">Code</span>
88            </a>
89          </p>
90          <p class="control">
91            <a class="button is-dark" href="https://arxiv.org/abs/2101.04653" target="_blank">
92              <span class="icon" style="padding: 1px;">
93                <img src="https://sbi-benchmark.github.io/img/adobeacrobatreader_white.svg" />
94              </span>
95              <span style="padding-left: 6px">
96                Paper
97              </span>
98            </a>
99          </p>
100        </div>
101      </div>
102    </div>
103  </div>
104</nav>
105      <section class="section">
106        <div class="container">
107          <div class="columns">
108            <div class="column is-7 is-offset-2">
109              <div class="content">
110                 <h1 id="benchmarking-simulation-based-inference">Benchmarking Simulation-Based Inference</h1>
111<p>Many domains of science make use of numerical simulators for which (1) it is easy to simulate but (2) the likelihood is unavailable, making it hard to perform statistical identification of model parameters consistent with observed data. <a href="https://www.pnas.org/content/early/2020/05/28/1912789117" target="_blank">Simulation-based inference (SBI)</a> deals with this 'likelihood-free' setting. Although recent advances have led to a large number of SBI algorithms, a public benchmark for such algorithms has been lacking: We set out to fill this gap, carefully select tasks and metrics, and evaluate several canonical algorithms. Through this website you can explore all results of our <a href="http://proceedings.mlr.press/v130/lueckmann21a.html" target="_blank">manuscript</a>, including comparisons on all metrics and plotting of posteriors for each of more than 10 000 runs.</p>
112<p>Keep reading for a brief summary of the manuscript, or jump right into interactive results (through the menu on top). We provide a framework to benchmark your own algorithms, code and results for reproducibility, and invite contributions.</p>
113<h2 id="our-motivation-and-methods">Our Motivation and Methods</h2>
114<p>Open benchmarks can be an important component of transparent and reproducible computational research. However, a benchmark framework for SBI has been lacking, possibly due to the challenging endeavour of designing benchmarking tasks and defining suitable performance metrics.</p>
115<p>We selected a set of initial algorithms representing four distinct approaches to SBI, analyzed multiple performance metrics which have been used in the literature, and implemented ten tasks, including ones popular in the field.</p>
116<p><img alt="" src="/img/algorithms.png" /></p>
117<p><small><em>Overview of algorithms. Classification and schemes following <a href="https://www.pnas.org/content/early/2020/05/28/1912789117" target="_blank">Cranmer et al. (2020)</a>.</em></small></p>
118<p><strong>Algorithms</strong>. We compare algorithms belonging to four distinct approaches to SBI: Classical ABC approaches as well as model-based approaches approximating likelihoods, posteriors, or density ratios. We contrast algorithms that use the prior distribution to propose parameters against ones that sequentially adapt the proposal. Keeping our initial selection of algorithms focused allowed us to carefully consider implementation details and hyperparameters.</p>
119<p><strong>Metrics.</strong> The shortcomings of commonly used metrics (see paper for details) led us to focus on tasks for which a likelihood can be evaluated, which allowed us to calculate reference (‘ground-truth’) posteriors. These reference posteriors are made available to allow rapid evaluation of SBI algorithms. While we compare algorithms in terms of classifier 2-sample tests (C2ST) in the paper, the website allows comparisons in terms of all metrics we considered.</p>
120<p><strong>Tasks.</strong> We focused on eight purely statistical problems and two problems relevant in applied domains, with diverse dimensionalities of parameters and data.</p>
121<h2 id="key-findings">Key Findings</h2>
122<p>The full potential of the benchmark will be realized when it is populated with additional community-contributed algorithms and tasks. However, our initial version already provides useful insights:</p>
123<ul>
124<li>the choice of performance metric is critical (commonly used
125  metrics such as the likelihood of true parameters or Maximum Mean Discrepancy with median heuristic have important shortcomings);</li>
126<li>the performance of the algorithms on some tasks leaves substantial room for improvement;</li>
127<li>sequential estimation generally improves sample efficiency;</li>
128<li>for small and moderate simulation budgets, neural-network based approaches outperform classical ABC algorithms, confirming recent progress in the field;</li>
129<li>but that there is no algorithm to rule them all.</li>
130</ul>
131<p>The performance ranking of algorithms is task-dependent, pointing to a need for better guidance or automated procedures for choosing which algorithm to use when. In the manuscript, we included 
131some considerations and recommendations for practitioners, based on our current results and understanding, and dedicated a page to discussing various limitations.</p>
132<p><strong>Find out more in our manuscript, <a href="http://proceedings.mlr.press/v130/lueckmann21a.html">available through PMLR</a> or:</strong></p>
133<blockquote>
134<p><a href="https://arxiv.org/abs/2101.04653">arXiv.org/abs/2101.04653</a></p>
135</blockquote>
136<p><em>We believe that the full potential of the benchmark will be revealed as more researchers participate and contribute. In order to facilitate this process, we provide a benchmarking framework which is designed to be highly extensible and easily used. See <a href="code/">Code &amp; Reproducibility</a> for explanations and examples.</em></p>
137<h2 id="3-minute-summary">3-Minute Summary</h2>
138<div id="presentation-embed-38952956"></div>
139<script src='https://slideslive.com/embed_presentation.js'></script>
vendor: 1 bytes, line 139
139
140<script>
141    embed = new SlidesLiveEmbed('presentation-embed-38952956', {
142        presentationId: '38952956',
143        autoPlay: false,
144        verticalEnabled: true,
145        zoomRatio: 0.3
146    });
147</script>
147
148
149<h2 id="citation">Citation</h2>
150<pre><code class="language-bibtex">@InProceedings{lueckmann2021benchmarking,
151 title     = {Benchmarking Simulation-Based Inference},
152 author    = {Lueckmann, Jan-Matthis and Boelts, Jan and Greenberg, David and Goncalves, Pedro and Macke, Jakob},
153 booktitle = {Proceedings of The 24th International Conference on Artificial Intelligence and Statistics},
154 pages     = {343--351},
155 year      = {2021},
156 editor    = {Banerjee, Arindam and Fukumizu, Kenji},
157 volume    = {130},
158 series    = {Proceedings of Machine Learning Research},
159 month     = {13--15 Apr},
160 publisher = {PMLR}
161}
162</code></pre>
163<h2 id="support">Support</h2>
164<p>This work was supported by the German Research Foundation (DFG; SFB 1233 PN 276693517, SFB 1089, SPP 2041, Germany’s Excellence Strategy – EXC number 2064/1 PN 390727645) and the German Federal Ministry of Education and Research (BMBF; project ’<a href="https://fit.uni-tuebingen.de/Project/Details?id=9199">ADIMEM</a>’, FKZ 01IS18052 A-D).</p> 
165              </div>
166            </div>
167            
168            <div class="column is-3 is-hidden-touch">
169              <aside class="menu has-navbar-fixed-top">
170  
171    
172      <ul class="menu-list">
173        
174          <li>
175  
176    <ul>
177      
178        <li>
179  <a href="#our-motivation-and-methods" class="is-size-7">Our Motivation and Methods</a>
180  
181</li>
182      
183        <li>
184  <a href="#key-findings" class="is-size-7">Key Findings</a>
185  
186</li>
187      
188        <li>
189  <a href="#3-minute-summary" class="is-size-7">3-Minute Summary</a>
190  
191</li>
192      
193        <li>
194  <a href="#citation" class="is-size-7">Citation</a>
195  
196</li>
197      
198        <li>
199  <a href="#support" class="is-size-7">Support</a>
200  
201</li>
202      
203    </ul>
204  
205</li>
206        
207      </ul>
208    
209  
210</aside>
211            </div>
212            
213          </div>
214        </div>
215      </section>
216    </div>
217    <footer class="footer">
218  <div class="has-text-centered">
219    <p>
220      If you have questions or comments, please do not hesitate to <a href="mailto:[email protected],[email protected]">contact us</a>.
221    </p>
222  </div>
223</footer>
224  </body>
225  
225<script type="text/javascript">
226  document.addEventListener('DOMContentLoaded', () => {
227
228    // Get all "navbar-burger" elements
229    const $navbarBurgers = Array.prototype.slice.call(document.querySelectorAll('.navbar-burger'), 0);
230
231    // Check if there are any navbar burgers
232    if ($navbarBurgers.length > 0) {
233
234      // Add a click event on each of them
235      $navbarBurgers.forEach( el => {
236        el.addEventListener('click', () => {
237
238          // Get the target from the "data-target" attribute
239          const target = el.dataset.target;
240          const $target = document.getElementById(target);
241
242          // Toggle the "is-active" class on both the "navbar-burger" and the "navbar-menu"
243          el.classList.toggle('is-active');
244          $target.classList.toggle('is-active');
245
246        });
247      });
248    }
249
250  });
251</script>
vendor: 1 bytes, line 251
251
252<script>hljs.initHighlightingOnLoad();</script>
252
253</html>

Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.