PageSourceSearch

https://ad-l-jepa.github.io/

html ad-l-jepa.github.io collected 2026-10-03 08:32:28 UTC 24,931 bytes, 534 lines download raw bytes

1<!doctype html>
2<html lang="en">
3
4<head>
5  <meta charset="utf-8">
6  <meta name="viewport" content="width=device-width, initial-scale=1">
7  <title>Self-Supervised Representation Learning with Joint Embedding Predictive Architecture for Automotive LiDAR
8    Object Detection</title>
9  
9<script src="template.v2.js"></script>
9
10  
10<script src="https://d3js.org/d3.v5.min.js"></script>
10
11  
11<script src="https://d3js.org/d3-collection.v1.min.js"></script>
11
12  
12<script src="https://rawgit.com/nstrayer/slid3r/master/dist/slid3r.js"></script>
12
13  
13<script src="cross_fade.js"></script>
13
14  <link rel="stylesheet" href="style.css">
15  <!-- 
15<script>
16    window.MathJax = {
17      tex: {
18        inlineMath: [['$', '$'], ['\\(', '\\)']],
19        displayMath: [['\\[', '\\]'], ['$$', '$$']],
20        packages: { '[+]': ['ams', 'textmacros'] }  // enables \text
21      },
22      svg: { fontCache: 'global' }
23    };
24  </script>
24
25  
25<script defer src="https://cdn.jsdelivr.net/npm/mathjax@3/es5/tex-svg.js"></script>
25 -->
26  <link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/[email protected]/dist/katex.min.css"
27    integrity="sha384-n8MVd4RsNIU0tAv4ct0nTaAbDJwPJzDEaqSD1odI+WdtXRGWt2kTvGFasHpSy3SV" crossorigin="anonymous">
28  
28<script defer src="https://cdn.jsdelivr.net/npm/[email protected]/dist/katex.min.js"
29    integrity="sha384-XjKyOOlGwcjNTAIQHIpgOno0Hl1YQqzUOEleOLALmuqehneUG+vnGctmUb0ZY0l8"
30    crossorigin="anonymous"></script>
30
31  
31<script defer src="https://cdn.jsdelivr.net/npm/[email protected]/dist/contrib/auto-render.min.js"
32    integrity="sha384-+VBxd3r6XgURycqtZ117nYw44OOcIax56Z4dCRWbxyPt0Koah1uHoK0o4+/RRE05"
33    crossorigin="anonymous"></script>
33
34  
34<script>
35    document.addEventListener("DOMContentLoaded", function () {
36      renderMathInElement(document.body, {
37        delimiters: [
38          { left: '$$', right: '$$', display: true },
39          { left: '$', right: '$', display: false },
40          { left: '\\(', right: '\\)', display: false },
41          { left: '\\[', right: '\\]', display: true }
42        ],
43        throwOnError: false
44      });
45    });
46  </script>
46
47</head>
48
49<body>
50  <div class="header-container">
51    <div class="header-content">
52      <h1>Self-Supervised Representation Learning with Joint Embedding Predictive Architecture for Automotive LiDAR
53        Object Detection</h1>
54      <p>AD-L-JEPA: A novel self-supervised pre-training framework with a joint embedding predictive architecture (JEPA)
55        for automotive LiDAR object detection.</p>
56      <div class="button-container">
57        <a href="https://arxiv.org/abs/2501.04969" class="button">Paper</a>
58        <a href="https://github.com/HaoranZhuExplorer/adljepa" class="button">Code</a>
59      </div>
60    </div>
61    <div class="header-image">
62      <img src="images/teaser.png" alt="Teaser Image">
63    </div>
64  </div>
65  <d-article>
66    <div class="byline">
67      <div class="byline-container">
68        <div class="byline-column">
69          <h3>Authors</h3>
70          <p><a href="https://arxiv.org/search/cs?searchtype=author&query=Zhu,+H" class="author-link">Haoran Zhu</a></p>
71          <p><a href="https://arxiv.org/search/cs?searchtype=author&query=Dong,+Z" class="author-link">Zhenyuan Dong</a>
72          </p>
73          <p><a href="https://arxiv.org/search/cs?searchtype=author&query=Topollai,+K" class="author-link">Kristi
74              Topollai</a></p>
75          <p><a href="https://arxiv.org/search/cs?searchtype=author&query=Sha,+B" class="author-link">Beiyao Sha</a></p>
76          <p><a href="https://arxiv.org/search/cs?searchtype=author&query=Choromanska,+A" class="author-link">Anna
77              Choromanska</a></p>
78        </div>
79        <div class="byline-column">
80          <h3>Affiliations</h3>
81          <p><a href="https://cs.nyu.edu/home/index.html" class="affiliation-link">New York University</a></p>
82        </div>
83        <div class="byline-column">
84          <h3>Resources</h3>
85          <p><a href="https://arxiv.org/abs/2501.04969" class="affiliation-link">Paper</a></p>
86          <p><a href="https://github.com/HaoranZhuExplorer/adljepa" class="affiliation-link">Code Repository</a></p>
87        </div>
88      </div>
89    </div>
90    <d-contents>
91      <nav>
92        <h4>Contents</h4>
93        <div><a href="#abstract">Abstract</a></div>
94        <div><a href="#method">Method</a></div>
95        <div><a href="#results">Results</a></div>
96        <div><a href="#conclusion">Conclusion</a></div>
97      </nav>
98    </d-contents>
99    <section id="abstract">
100      <h2>Abstract</h2>
101      <p>
102        Recently, self-supervised representation learning relying on vast amounts of unlabeled data has been explored as
103        a pre-training method for autonomous driving. However, directly applying popular contrastive or generative
104        methods to this problem is insufficient and may even lead to negative transfer. In this paper, we present
105        AD-L-JEPA, a novel self-supervised pre-training framework with a joint embedding predictive architecture (JEPA)
106        for automotive LiDAR object detection. Unlike existing methods, AD-L-JEPA is neither generative nor contrastive.
107        Instead of explicitly generating masked regions, our method predicts Bird's-Eye-View embeddings to capture the
108        diverse nature of driving scenes. Furthermore, our approach eliminates the need to manually form contrastive
109        pairs by employing explicit variance regularization to avoid representation collapse. Experimental results
110        demonstrate consistent improvements on the LiDAR 3D object detection downstream task across the KITTI3D, Waymo,
111        and ONCE datasets, while reducing GPU hours by 1.9x-2.7x and GPU memory by 2.8x-4x compared with the
112        state-of-the-art method Occupancy-MAE. Notably, on the largest ONCE dataset, pre-training on 100K frames yields
113        a 1.61 mAP gain, better than all other methods pre-trained on either 100K or 500K frames, and pre-training on
114        500K frames yields a 2.98 mAP gain, better than all other methods pre-trained on either 500K or 1M frames.
115        AD-L-JEPA constitutes the first JEPA-based pre-training method for autonomous driving. It offers better quality,
116        faster, and more GPU-memory-efficient self-supervised representation learning. The source code of AD-L-JEPA is
117        ready to be released.
118      </p>
119    </section>
120
121    <section id="introduction">
122      <h2>Introduction</h2>
123      <p>
124        Unlike human drivers, current autonomous driving (AD) systems still require large amounts of labeled data for
125        training. This supervised-only paradigm is expensive due to labeling costs and limits the scalability of these
126        systems. Recently, researchers have proposed self-supervised learning (SSL) across camera, LiDAR, and radar
127        modalities to pre-train the network without any labels and then fine-tune it with labeled data to adapt to
128        specific downstream tasks.
129      </p>
130      <p>
131        In SSL, the two most popular learning paradigms are contrastive methods and generative methods. However,
132        directly applying these methods for pre-training in AD is challenging and can even hurt downstream performance.
133        This stems from both the difficulty of defining meaningful contrastive pairs via data augmentation in driving
134        scenarios that contain multiple objects, and the fact that explicit scene generation is time-consuming and
135        insufficient to capture semantic representations of diverse driving scenarios.
136      </p>
137      <p>
138        In this paper, we present <strong>AD-L-JEPA</strong> (Autonomous Driving with LiDAR data via a Joint Embedding
139        Predictive Architecture), a novel self-supervised pre-training framework for automotive LiDAR object detection
140        that, as opposed to existing methods, is neither generative nor contrastive. Our method learns self-supervised
141        representations in Bird's Eye View (BEV) space and predicts embeddings for spatially masked regions. It omits
142        the need to create human-crafted positive/negative pairs, as required by contrastive learning. Furthermore,
143        rather than explicitly reconstructing unknown parts of the data as generative methods do, it predicts BEV
144        embeddings instead.
145      </p>
146      <figure style="margin-top: 20px; margin-bottom: 20px;">
147        <img src="images/ad-l-jepa/intuition.png" alt="Intuition of AD-L-JEPA" style="width: 100%;">
148        <figcaption><strong>Figure 1:</strong> Intuition of AD-L-JEPA. Unlike contrastive methods that require negative
149          pairs or generative methods that reconstruct raw data, AD-L-JEPA predicts latent embeddings of masked regions
150          from visible regions in the BEV space.</figcaption>
151      </figure>
152    </section>
153
154    <section id="method">
155      <h2>Method</h2>
156      <p>
157        The architecture of AD-L-JEPA is shown in Figure 1. The overarching intuition behind our framework is as
158        follows: for the visible parts of the point cloud scene, the network is trained in a self-supervised manner to
159        predict how the invisible parts should appear in the embedding space. This enables the learning of geometrically
160        and semantically reasonable representations, as well as adapting to the high uncertainty nature of the AD scenes
161        by avoiding the explicit reconstruction of the invisible parts of the data.
162      </p>
163      <figure style="margin-top: 20px; margin-bottom: 20px;">
164        <img src="images/ad-l-jepa/architecture.png" alt="AD-L-JEPA Architecture" style="width: 100%;">
165        <figcaption><strong>Figure 2:</strong> Overview of the AD-L-JEPA architecture: We introduce modified BEV-guided
166          masking to mask the input point cloud in both empty and non-empty regions. The network predicts BEV embeddings
167          at masked regions, leveraging variance regularization at non-empty regions following the output of the context
168          encoder and the lightweight spatial predictor. It also employs a moving average update of the target encoder
169          to learn diverse, high-level semantic representations.</figcaption>
170      </figure>
171
172      <h3>Modified BEV-Guided Masking</h3>
173      <p>
174        To learn effective representations in a self-supervised manner, masking is used to create invisible and visible
175        regions. The network is then trained to predict embeddings of the invisible regions based on the visible ones.
176        We have two design recipes for masking in AD scenarios: (1) masks are first created in the BEV embedding space
177        and recursively upsampled to the input point cloud to identify points to be masked; (2) both empty and non-empty
178        areas should be included in the visible and invisible regions created by the masks. These two criteria can be
179        achieved by modifying the BEV-guided masking originally proposed in [Lin et al. 2024].
180      </p>
181      <figure style="margin-top: 20px; margin-bottom: 20px;">
182        <img src="images/ad-l-jepa/masking.png" alt="Modified BEV-Guided Masking" style="width: 100%;">
183        <figcaption><strong>Figure 3:</strong> Comparison of original BEV-guided masking with our modified version that
184          creates masks in both empty and non-empty regions.</figcaption>
185      </figure>
186
187      <h3>Context & Target Encoders</h3>
188      <p>
189        The context encoder $f_\theta$ and target encoder $f_{\bar{\theta}}$ are backbones responsible for extracting
190        context embeddings from the unmasked point cloud and target embeddings from the masked point cloud,
191        respectively. The context encoder will later be used for fine-tuning on the downstream tasks after
192        self-supervised representation learning. It receives input point cloud features and outputs embeddings in a
193        downsampled 3D space. We obtain BEV embeddings by reshaping the 3D embeddings.
194      </p>
195
196      <h3>Predictor</h3>
197      <p>
198        The predictor is a lightweight, three-layer convolutional network $g_\phi$ that predicts target BEV embeddings
199        from visible context BEV embeddings. We denote the predicted embedding, after the $L_2$ normalization is applied
200        to each BEV grid's embedding dimension, as $\boldsymbol{\hat{s}}_c = g_\phi(\boldsymbol{\hat{z}}_c)$.
201      </p>
202
203      <h3>Training Objectives</h3>
204      <p>
205        We pre-train the network in a self-supervised manner with two losses to ensure we learn high-quality,
206        non-collapsed embeddings: a cosine similarity-based embedding prediction loss and a variance regularization
207        loss.
208      </p>
209      <p>
210        <strong>Embedding Prediction Loss:</strong>
211        $$
212        \mathcal{L_{\text{jepa}}} = \frac{\alpha_0}{\sum |P_n|} \sum (1 - \text{sim}(\hat{s}_c, \hat{s}_t)) +
213        \frac{\alpha_1}{\sum |Q_n|} \sum (1 - \text{sim}(\hat{s}_c, \hat{s}_t))
214        $$
215        where $P_n$ and $Q_n$ are subsets of masked empty and non-empty BEV grids, respectively.
216      </p>
217      <p>
218        <strong>Variance Regularization Loss:</strong>
219        $$
220        \mathcal{L_\text{reg}} = \beta_1 \sum v(\boldsymbol{\hat{z}}_c) + \beta_2 \sum v(\boldsymbol{\hat{s}}_c)
221        $$
222        This loss ensures that the average variance across all embedding dimensions is larger than some threshold,
223        preventing representation collapse.
224      </p>
225      <p>
226        The overall self-supervised learning loss is:
227        $$ \mathcal{L} = \lambda_{\text{jepa}} \mathcal{L_\text{jepa}} + \lambda_{\text{reg}} \mathcal{L_\text{reg}} $$
228      </p>
229      <p>
230        The parameters of the target encoder are updated through a moving average of the context encoder's parameters,
231        $\bar{\theta} \leftarrow \eta \bar{\theta} + (1 - \eta) \theta$, to further avoid representation collapse.
232      </p>
233    </section>
234
235    <section id="results">
236      <h2>Results</h2>
237      <p>
238        We evaluate our pre-training method on three datasets of increasing scale: KITTI3D, Waymo, and ONCE. We compare
239        against state-of-the-art self-supervised methods like Occupancy-MAE and ALSO.
240      </p>
241
242      <h3>Pre-training Efficiency</h3>
243      <p>
244        Unlike Occupancy-MAE, which uses computationally expensive dense 3D convolutions to reconstruct invisible
245        regions, AD-L-JEPA employs a joint-embedding predictive architecture at the BEV level and omits those layers.
246        This results in <strong>2.8x–3.4x lower GPU memory usage</strong> and <strong>2.7x fewer GPU hours</strong> for
247        pre-training on the 20% and 100% splits of the Waymo dataset, and <strong>3.1x–4x lower GPU memory
248          usage</strong> and <strong>1.9x fewer GPU hours</strong> for pre-training on the ONCE 100k split.
249      </p>
250
251      <h3>Visual Comparison</h3>
252      <p>
253        We visualize the impact of pre-training label efficiency. AD-L-JEPA consistently outperforms baselines across
254        different label efficiencies.
255      </p>
256      <figure style="margin-top: 20px; margin-bottom: 20px;">
257        <img src="images/ad-l-jepa/visual_comparison.png" alt="Visual Comparison of Label Efficiency"
258          style="width: 100%;">
259        <figcaption><strong>Figure 4:</strong> Visual comparison of label efficiency. AD-L-JEPA demonstrates superior
260          performance even with limited labeled data.</figcaption>
261      </figure>
262
263      <h3>Downstream Fine-tuning Performance</h3>
264
265      <h4>KITTI3D (PV-RCNN)</h4>
266      <div style="overflow-x: auto;">
267        <table class="display-table" style="width: 100%; margin-bottom: 20px;">
268          <thead>
269            <tr>
270              <th>Method</th>
271              <th>Cars</th>
272              <th>Ped.</th>
273              <th>Cycl.</th>
274              <th>Overall</th>
275              <th>Diff.</th>
276            </tr>
277          </thead>
278          <tbody>
279            <tr>
280              <td>No pre-training</td>
281              <td>84.65</td>
282              <td>56.19</td>
283              <td>72.19</td>
284              <td>71.01</td>
285              <td>-</td>
286            </tr>
287            <tr>
288              <td>Occupancy-MAE</td>
289              <td>84.34</td>
290              <td>57.55</td>
291              <td>71.33</td>
292              <td>71.07</td>
293              <td>+0.06</td>
294            </tr>
295            <tr>
296              <td>ALSO</td>
297              <td>84.64</td>
298              <td>57.09</td>
299              <td><strong>73.72</strong></td>
300              <td>71.82</td>
301              <td>+0.81</td>
302            </tr>
303            <tr style="background-color: #f0f8ff;">
304              <td><strong>AD-L-JEPA (ours)</strong></td>
305              <td><strong>85.07</strong></td>
306              <td><strong>59.68</strong></td>
307              <td>73.02</td>
308              <td><strong>72.59</strong></td>
309              <td><strong>+1.58</strong></td>
310            </tr>
311          </tbody>
312        </table>
313      </div>
314
315      <h4>Waymo (CenterPoint, 100% Data)</h4>
316      <div style="overflow-x: auto;">
317        <table class="display-table" style="width: 100%; margin-bottom: 20px;">
318          <thead>
319            <tr>
320              <th>Method</th>
321              <th>Veh.</th>
322              <th>Ped.</th>
323              <th>Cycl.</th>
324              <th>Overall</th>
325              <th>Diff.</th>
326            </tr>
327          </thead>
328          <tbody>
329            <tr>
330              <td>No pre-training</td>
331              <td>63.28</td>
332              <td>63.95</td>
333              <td>66.77</td>
334              <td>64.67</td>
335              <td>-</td>
336            </tr>
337            <tr>
338              <td>Occupancy-MAE</td>
339              <td>63.53</td>
340              <td><strong>64.73</strong></td>
341              <td>67.77</td>
342              <td>65.34</td>
343              <td>+0.67</td>
344            </tr>
345            <tr style="background-color: #f0f8ff;">
346              <td><strong>AD-L-JEPA (ours)</strong></td>
347              <td><strong>63.58</strong></td>
348              <td>64.58</td>
349              <td><strong>68.07</strong></td>
350              <td><strong>65.41</strong></td>
351              <td><strong>+0.74</strong></td>
352            </tr>
353          </tbody>
354        </table>
355      </div>
356
357      <h4>ONCE (SECOND)</h4>
358      <div style="overflow-x: auto;">
359        <table class="display-table" style="width: 100%; margin-bottom: 20px;">
360          <thead>
361            <tr>
362              <th>Method</th>
363              <th>Veh.</th>
364              <th>Ped.</th>
365              <th>Cycl.</th>
366              <th>Overall</th>
367              <th>Diff.</th>
368            </tr>
369          </thead>
370          <tbody>
371            <tr>
372              <td>No pre-training</td>
373              <td>71.19</td>
374              <td>26.44</td>
375              <td>58.04</td>
376              <td>51.89</td>
377              <td>-</td>
378            </tr>
379            <tr>
380              <td>Occupancy-MAE (100k)</td>
381              <td><strong>73.54</strong></td>
382              <td>25.93</td>
383              <td>58.34</td>
384              <td>52.60</td>
385              <td>+0.71</td>
386            </tr>
387            <tr style="background-color: #f0f8ff;">
388              <td><strong>AD-L-JEPA (100k)</strong></td>
389              <td>73.18</td>
390              <td>29.19</td>
391              <td>58.14</td>
392              <td>53.50</td>
393              <td>
393+1.61</td>
394            </tr>
395            <tr style="background-color: #e6f2ff;">
396              <td><strong>AD-L-JEPA (500k)</strong></td>
397              <td>73.25</td>
398              <td>31.91</td>
399              <td><strong>59.47</strong></td>
400              <td><strong>54.87</strong></td>
401              <td><strong>+2.98</strong></td>
402            </tr>
403            <tr style="background-color: #d9ebff;">
404              <td><strong>AD-L-JEPA (1M)</strong></td>
405              <td>73.01</td>
406              <td><strong>31.94</strong></td>
407              <td>59.16</td>
408              <td>54.70</td>
409              <td>+2.81</td>
410            </tr>
411          </tbody>
412        </table>
413      </div>
414
415      <h3>Transfer Learning (Waymo -> KITTI)</h3>
416      <p>
417        We evaluate transfer learning by pre-training on Waymo and fine-tuning on KITTI. AD-L-JEPA consistently
418        outperforms baselines across different label efficiencies.
419      </p>
420      <div style="overflow-x: auto;">
421        <table class="display-table" style="width: 100%; margin-bottom: 20px;">
422          <thead>
423            <tr>
424              <th>Method (100% Labels)</th>
425              <th>Cars</th>
426              <th>Ped.</th>
427              <th>Cycl.</th>
428              <th>Overall</th>
429              <th>Diff.</th>
430            </tr>
431          </thead>
432          <tbody>
433            <tr>
434              <td>No pre-training</td>
435              <td><strong>81.99</strong></td>
436              <td>52.02</td>
437              <td>65.07</td>
438              <td>66.36</td>
439              <td>-</td>
440            </tr>
441            <tr>
442              <td>Occupancy-MAE</td>
443              <td>81.65</td>
444              <td>51.51</td>
445              <td>66.72</td>
446              <td>66.63</td>
447              <td>+0.27</td>
448            </tr>
449            <tr style="background-color: #f0f8ff;">
450              <td><strong>AD-L-JEPA (ours)</strong></td>
451              <td>80.92</td>
452              <td><strong>52.45</strong></td>
453              <td><strong>69.76</strong></td>
454              <td><strong>67.71</strong></td>
455              <td><strong>+1.35</strong></td>
456            </tr>
457          </tbody>
458        </table>
459      </div>
460
461      <h3>Other Evaluations</h3>
462
463      <h4>Occupancy Estimation</h4>
464      <figure style="margin-top: 20px; margin-bottom: 20px;">
465        <img src="images/ad-l-jepa/visual_occupancy.png" alt="Occupancy Estimation" style="width: 100%;">
466        <figcaption><strong>Figure 5:</strong> Masked region occupancy estimation evaluated by comparing BEV embeddings
467          obtained by AD-L-JEPA with the learnable empty token via the cosine similarity. Unmasked regions are ignored
468          and the cosine similarity in this case is represented in white color.</figcaption>
469      </figure>
470
471      <h4>Singular Value Decomposition Analysis</h4>
472      <figure style="margin-top: 20px; margin-bottom: 20px;">
473        <img src="images/ad-l-jepa/svd.png" alt="SVD Analysis" style="width: 100%;">
474        <figcaption><strong>Figure 6:</strong> Sorted normalized singular values and the corresponding cumulative
475          explained variance, obtained by singular value decomposition of pre-trained BEV embeddings. Embeddings are
476          obtained either with AD-L-JEPA or Occupancy-MAE.</figcaption>
477      </figure>
478
479    </section>
480
481    <section id="discussion" style="margin-top:40px;">
482      <h2>Discussion</h2>
483      <p>
484        <strong>Saturation on Large-Scale Data:</strong> Interestingly, AD-L-JEPA pre-trained on 1M frames, although
485        significantly better than other methods, falls slightly behind AD-L-JEPA pre-trained on 500K frames. This small
486        drop aligns with existing literature showing that increasing the number of unlabeled samples consistently boosts
487        performance but saturates at a point. Such saturation can be explained by the data redundancy of highly similar
488        driving scenarios in the 1M frame setting. To validate, we took AD-L-JEPA pretrained on 100K frames and tested
489        it on 16K unseen LiDAR samples from the 500K/1M sets. The 500K set showed a higher average loss (0.44 vs. 0.43),
490        implying richer diversity and stronger fine-tuning transfer.
491      </p>
492      <p>
493        <strong>Future Work:</strong> For future work, we plan to extend AD-L-JEPA to leverage temporal dynamics and to
494        incorporate action-conditioned self-supervised representation learning in AD scenarios.
495      </p>
496    </section>
497
498    </section>
499
500    <section id="conclusion" style="margin-top:40px;">
501      <h2>Conclusion</h2>
502      <p>
503        AD-L-JEPA constitutes the first JEPA-based pre-training method for autonomous driving. It offers better quality,
504        faster, and more GPU-memory-efficient self-supervised representation learning.
505      </p>
506    </section>
507
508  </d-article>
509  <d-appendix>
510    <p>
511      This webpage template is adapted from <a href="https://rae-dit.github.io/">RAE</a>.
512    </p>
513    <h3>BibTeX</h3>
514    <p class="bibtex">
515      @misc{zhu2025selfsupervisedrepresentationlearningjoint,<br>
516      &nbsp;&nbsp;title={Self-Supervised Representation Learning with Joint Embedding Predictive Architecture for
517      Automotive LiDAR Object Detection},<br>
518      &nbsp;&nbsp;author={Haoran Zhu and Zhenyuan Dong and Kristi Topollai and Beiyao Sha and Anna Choromanska},<br>
519      &nbsp;&nbsp;year={2025},<br>
520      &nbsp;&nbsp;eprint={2501.04969},<br>
521      &nbsp;&nbsp;archivePrefix={arXiv},<br>
522      &nbsp;&nbsp;primaryClass={cs.RO}<br>
523      }
524    </p>
525    <d-footnote-list></d-footnote-list>
526    <d-citation-list></d-citation-list>
527  </d-appendix>
528
529  <!-- bibliography will be inlined during Distill pipeline's pre-rendering -->
530  <d-bibliography src="bibliography.bib"></d-bibliography>
531  
531<script src="contents_bar.js"></script>
531 <!-- for scroll/toc -->
532</body>
533
534</html>

Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.