1<!DOCTYPE html> 2<html> 3 4<head> 5 <meta charset="utf-8"> 6 <meta name="description" 7 content="Desktop gaming data effectively pretrains embodied AI: 152Ã compression via OWA Toolkit, YouTube pseudo-labeling with Generalist-IDM, achieving 96.6% on LIBERO manipulation and 83.3% on CANVAS navigation with 1.3K hours of data."> 8 <meta name="keywords" content="Embodied Ai, Vision-Language-Action Models, Inverse Dynamics Models"> 9 <meta name="viewport" content="width=device-width, initial-scale=1"> 10 <title>D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI</title> 11 12 <!-- Google tag (gtag.js) -->
vendor: 68 bytes, lines 12-13
12 13 <script async src="https://www.googletagmanager.com/gtag/js?id=
13G-5J9LZW868J
vendor: 16 bytes, lines 13-14
13"></script> 14
14<script> 15
vendor: 155 bytes, lines 15-19
15window.dataLayer = window.dataLayer || []; 16 function gtag() { dataLayer.push(arguments); } 17 gtag('js', new Date()); 18 19 gtag('config', '
19G-5J9LZW868J
vendor: 8 bytes, lines 19-20
19'); 20
20</script>
20 21 22 <link href="https://fonts.googleapis.com/css?family=Google+Sans|Noto+Sans|Castoro" rel="stylesheet"> 23 24 <link rel="stylesheet" href="../static/css/bulma.min.css"> 25 <link rel="stylesheet" href="../static/css/slick.css"> 26 <link rel="stylesheet" href="../static/css/slick-theme.css"> 27 <link rel="stylesheet" href="../static/css/fontawesome.all.min.css"> 28 <link rel="stylesheet" href="https://cdn.jsdelivr.net/gh/jpswalsh/academicons@1/css/academicons.min.css"> 29 <link rel="stylesheet" href="../static/css/index.css"> 30 <link rel="icon" type="image/png" sizes="32x32" href="../static/images/favicon-32x32.png"> 31 <link rel="icon" type="image/png" sizes="16x16" href="../static/images/favicon-16x16.png"> 32 <link rel="apple-touch-icon" sizes="180x180" href="../static/images/apple-touch-icon.png"> 33 34
34<script src="https://ajax.googleapis.com/ajax/libs/jquery/3.5.1/jquery.min.js"></script>
34 35
35<script defer src="../static/js/fontawesome.all.min.js"></script>
35 36
36<script src="../static/js/slick.min.js"></script>
36 37
37<script src="../static/js/index.js"></script>
37 38
38<script src="../static/js/research-projects.js?v=20260905-corl"></script>
38 39</head> 40 41<body> 42 43 <nav class="navbar" role="navigation" aria-label="main navigation"> 44 <div class="navbar-brand"> 45 <a role="button" class="navbar-burger" aria-label="menu" aria-expanded="false"> 46 <span aria-hidden="true"></span> 47 <span aria-hidden="true"></span> 48 <span aria-hidden="true"></span> 49 </a> 50 </div> 51 <div class="navbar-menu"> 52 <div class="navbar-start" style="flex-grow: 1; justify-content: center;"> 53 <a class="navbar-item" href="../"> 54 <span class="icon"> 55 <i class="fas fa-home"></i> 56 </span> 57 </a> 58 59 <div class="navbar-item has-dropdown is-hoverable"> 60 <a class="navbar-link"> 61 More Research 62 </a> 63 <div class="navbar-dropdown" id="research-dropdown-items"> 64 <!-- Auto-populated by ../static/js/research-projects.js (single source of truth) --> 65 </div> 66 </div> 67 </div> 68 </div> 69 70 </div> 71 </nav> 72 73 <section class="hero"> 74 <div class="hero-body"> 75 <div class="container is-max-desktop"> 76 <div class="columns is-centered"> 77 <div class="column has-text-centered"> 78 <h1 class="title is-2 publication-title">D2E: Scaling Vision-Action Pretraining on Desktop Data 79 for Transfer to Embodied AI</h1> 80 <div class="is-size-5 publication-authors"> 81 <span class="author-block"> 82 <a href="https://suhwanchoi.me/">Suhwan Choi</a><sup>â 1</sup>,</span> 83 </span> 84 <span class="author-block"> 85 <a href="https://lastdefiance20.github.io/">Jaeyoon Jung</a><sup>â 1</sup>,</span> 86 </span> 87 <span class="author-block"> 88 <a href="https://hbseong97.github.io/">Haebin Seong</a><sup>â 1</sup>,</span> 89 </span> 90 <span class="author-block"> 91 <a href="https://minchankim.me/">Minchan Kim</a><sup>1</sup>,</span> 92 </span> 93 <span class="author-block"> 94 <a href="https://scholar.google.com/citations?user=Jh3S9aAAAAAJ">Minyeong Kim</a><sup>2</sup>,</span> 95 </span> 96 <br> 97 <span class="author-block"> 98 <a href="https://scholar.google.com/citations?user=VxekekYAAAAJ">Yongjun Cho</a><sup>1</sup>,</span> 99 </span> 100 <span class="author-block"> 101 Yoonshik Kim<sup>1</sup>,</span> 102 </span> 103 <span class="author-block"> 104 <a href="https://www.linkedin.com/in/yu-been-park-7223a6234/">Yubeen Park</a><sup>1</sup>,</span> 105 </span> 106 <span class="author-block"> 107 <a href="https://yj-yu.github.io/home/">Youngjae Yu</a><sup>â¡3</sup>,</span> 108 </span> 109 <span class="author-block"> 110 <a href="https://scholar.google.com/citations?user=7iaKhrEAAAAJ">Yunsung Lee</a><sup>â¡1</sup>,</span> 111 </span> 112 </div> 113 114 <div class="is-size-6 publication-authors"> 115 <span class="author-block"> 116 <sup>â </sup> Equal contribution, <sup>â¡</sup> Co-corresponding author, <sup>1</sup> 117 MAUM.AI, <sup>2</sup> Stanford University, <sup>3</sup> Seoul National University 118 </div> 119 <br> 120 121 <div class="is-size-5 publication-venue">
122 <span class="venue-block">International Conference on Learning Representations (ICLR) 123 2026</span> 124 <br> 125 </div> 126 <br> 127 128 <div class="column has-text-centered"> 129 <div class="publication-links"> 130 <!-- PDF Link. --> 131 <span class="link-block"> 132 <a href="https://arxiv.org/abs/2510.05684" 133 class="external-link button is-normal is-rounded is-dark"> 134 <span class="icon"> 135 <i class="fas fa-file-pdf"></i> 136 </span> 137 <span>Paper</span> 138 </a> 139 </span> 140 141 <!-- Video Link. --> 142 <!-- <span class="link-block"> 143 <a href="https://youtu.be/oJuU4x02azI" 144 class="external-link button is-normal is-rounded is-dark"> 145 <span class="icon"> 146 <i class="fab fa-youtube"></i> 147 </span> 148 <span>Video</span> 149 </a> 150 </span> --> 151 <!-- Code Link. --> 152 <span class="link-block"> 153 <a href="https://github.com/worv-ai/D2E" 154 class="external-link button is-normal is-rounded is-dark"> 155 <span class="icon"> 156 <i class="fab fa-github-alt"></i> 157 </span> 158 <span>Code</span> 159 </a> 160 </span> 161 <!-- Model Link. --> 162 <span class="link-block"> 163 <a href="https://huggingface.co/open-world-agents/Generalist-IDM-1B" 164 class="external-link button is-normal is-rounded is-dark is-disabled"> 165 <span class="icon"> 166 <i class="fas fa-robot"></i> 167 </span> 168 <span>Model (G-IDM)</span> 169 </a> 170 </span> 171 <!-- Dataset Link. --> 172 <span class="link-block"> 173 <a href="https://huggingface.co/datasets/open-world-agents/D2E-480p" 174 class="external-link button is-normal is-rounded is-dark is-disabled"> 175 <span class="icon"> 176 <i class="fas fa-database"></i> 177 </span> 178 <span>Dataset (480p)</span> 179 </a> 180 </span> 181 <span class="link-block"> 182 <a href="https://huggingface.co/datasets/open-world-agents/D2E-Original" 183 class="external-link button is-normal is-rounded is-dark is-disabled"> 184 <span class="icon"> 185 <i class="fas fa-database"></i> 186 </span> 187 <span>Dataset (FHD/QHD)</span> 188 </a> 189 </span> 190 <!-- Demo Link. --> 191 <!-- <span class="link-block"> 192 <a href="https://huggingface.co/spaces
192/maum-ai/CANVAS-DEMO" 193 class="external-link button is-normal is-rounded is-dark"> 194 <span class="icon"> 195 <i class="fas fa-images"></i> </span> 196 <span>Demo</span> 197 </a> 198 </span> --> 199 </div> 200 201 </div> 202 </div> 203 </div> 204 </div> 205 </div> 206 </section> 207 208 <section class="hero teaser"> 209 <div class="container is-max-desktop is-centered has-text-justified is-size-5"> 210 <div class="hero-body"> 211 <figure id="teaser"> 212 <img src="./static/images/1_teaser.png" alt="d2e teaser" /> 213 </figure> 214 <p> 215 Existing approaches (e.g., DROID) for collecting embodied AI data are expensive, low diversity, 216 and hard to scale. D2E leverages desktop data which is cheap, high diversity, and easy to scale. 217 The OWA Toolkit captures 335 hours of rich desktop demonstrations across 31 games with 218 152Ã compression. The Generalist-IDM uses next-event prediction with temporal offset (NEP-Ï) 219 to achieve OOD generalization, enabling pseudo-labeling of 1K+ hours of YouTube gameplay. 220 Vision-Action Pretraining transfers desktop-pretrained representations to embodied AI, achieving 221 96.6% success on LIBERO manipulation and 83.3% on CANVAS navigation benchmarks which demonstrates 222 desktop-to-robotics transfer. 223 </p> 224 </div> 225 </div> 226 </section> 227 228 <!-- Results Carousel --> 229 <section class="hero is-light is-small"> 230 <div class="hero-body"> 231 <div class="container"> 232 <h1 class="title has-text-centered">Pseudo-Label result on YouTube dataset</h1> 233 <div class="columns is-multiline"> 234 <div class="column is-3"> 235 <video poster="" id="brotato" autoplay controls muted loop playsinline width="100%"> 236 <source src="./static/videos/youtube_brotato.mp4" type="video/mp4"> 237 </video> 238 <p class="is-size-6 mt-2 has-text-centered">Brotato</p> 239 </div> 240 <div class="column is-3"> 241 <video poster="" id="csgo2" autoplay controls muted loop playsinline width="100%"> 242 <source src="./static/videos/youtube_csgo2.mp4" type="video/mp4"> 243 </video> 244 <p class="is-size-6 mt-2 has-text-centered">CSGO2</p> 245 </div> 246 <div class="column is-3"> 247 <video poster="" id="stardew" autoplay controls muted loop playsinline width="100%"> 248 <source src="./static/videos/youtube_stardew.mp4" type="video/mp4"> 249 </video> 250 <p class="is-size-6 mt-2 has-text-centered">Stardew Valley</p> 251 </div> 252 <div class="column is-3"> 253 <video poster="" id="minecraft" autoplay controls muted loop playsinline width="100%"> 254 <source src="./static/videos/youtube_minecraft.mp4" type="video/mp4"> 255 </video> 256 <p class="is-size-6 mt-2 has-text-centered">Minecraft</p> 257 </div> 258 <div class="column is-3"> 259 <video poster="" id="slime" autoplay controls muted loop playsinline width="100%"> 260 <source src="./static/videos/youtube_slime.mp4" type="video/mp4"> 261 </video> 262 <p class="is-size-6 mt-2 has-text-centered">Slime Rancher</p> 263 </div> 264 <div class="column is-3"> 265 <video poster="" id="raft" autoplay controls muted loop playsinline width="100%"> 266 <source src="./static/videos/youtube_raft.mp4" type="video/mp4"> 267 </video> 268 <p class="is-size-6 mt-2 has-text-centered">Raft</p> 269 </div> 270 <div class="column is-3"> 271 <video poster="" id="barony" autoplay controls muted loop playsinline width="100%"> 272 <source src="./static/videos/youtube_barony.mp4" type="video/mp4"> 273 </video> 274 <p class="is-size-6 mt-2 has-text-centered">Barony</p> 275 </div> 276 <div class="column is-3"> 277 <video poster="" id="dinkum" autoplay controls muted loop playsinline width="100%"> 278 <source src="./static/videos/youtube_dinkum.mp4" type="video/mp4"> 279 </video> 280 <p class="is-size-6 mt-2 has-text-centered">Dinkum</p> 281 </div> 282 </div> 283 <div class="content has-text-centered mt-4"> 284 <p class="is-size-6"> 285 Generalist-IDM uses a single model to label actions on video-only 286 data, across 2D/3D games and visual navigation/UI interactions without separate processing. <br> 287 Remarkably, in Counter-Strike videos where spectator mode begins around 10 seconds, 288 it can distinguish between active gameplay and spectator mode by recognizing subtle UI elements, 289 avoiding action predictions during spectator phases. 290 </p> 291 </div> 292 </div> 293 </div> 294 </section> 295 296 <section class="section"> 297 <div class="container is-max-desktop"> 298 <!-- Abstract. --> 299 <div class="columns is-centered has-text-centered"> 300 <div class="column is-four-fifths"> 301 <h2 class="title is-3">Abstract</h2> 302 <div class="content has-text-justified"> 303 <p> 304 Large language models leverage internet-scale text data, yet embodied AI remains constrained 305 by the prohibitive costs of physical trajectory collection. Desktop 306 environments---particularly gaming---offer a compelling alternative: they provide rich 307 sensorimotor interactions at scale while maintaining the structured observation-action 308 coupling essential for embodied learning. We present D2E (Desktop to Embodied AI), a 309 framework that demonstrates desktop interactions can serve as an effective pretraining 310 substrate for robotics embodied AI tasks. Unlike prior work that remained domain-specific 311 (e.g., VPT for Minecraft) or kept data proprietary (e.g., SIMA), D2E establishes a complete 312 pipeline from scalable desktop data collection to verified transfer in embodied domains. Our 313 framework comprises three components: (1) the OWA Toolkit that unifies diverse desktop 314 interactions into a standardized format with 152Ã compression, (2) the Generalist-IDM that 315 achieves strong zero-shot generalization across unseen games through timestamp-based event 316 prediction, enabling internet-scale pseudo-labeling, and (3) VAPT that transfers 317 desktop-pretrained representations to physical manipulation and navigation. Using 1.3K+ 318 hours of data (259 hours of human demonstrations, and 1K+ hours of pseudo-labeled gameplay), 319 we achieve a total of 96.6% success rate on LIBERO manipulation and 83.3% on CANVAS 320 navigation benchmarks. This validates that sensorimotor primitives in digital interactions 321 exhibit sufficient invariance to transfer meaningfully to physical embodied tasks, 322 establishing desktop pretraining as a practical paradigm for robotics. We will make all our 323 work public, including the OWA toolkit, datasets of human-collected and pseudo-labeled, and 324 VAPT-trained models. 325 </p> 326 </div> 327 </div> 328 </div> 329 330 <!--/ Abstract. --> 331 </div> 332 </section> 333 334 <section class="section"> 335 <div class="container is-max-desktop"> 336 <div class="columns is-centered has-text-centered"> 337 <div class="column is-full-width"> 338 <h2 class="title is-3">Generalist Inverse Dynamics Model (G-IDM)</h2> 339 <div class="content has-text-justified"> 340 <figure id="gidm_id"> 341 <img src="./static/images/3_gidm_id.png" alt="gidm_id" /> 342 </figure> 343 <p> 344 The Generalist Inverse Dynamics Model (G-IDM) learns to predict actions from observation 345 transitions across diverse desktop environments. 346 Trained on our multi-domain corpus collected via the OWA Toolkit, G-IDM achieves strong 347 performance across all in-distribution environments, yielding large gains in Pearson 348 correlation (e.g., +39.5 points on Stardew Valley X) and ke
348yboard accuracy 349 (e.g., +57.6 points on Brotato), demonstrating robust generalization over diverse control 350 dynamics. 351 </p> 352 </div> 353 </div> 354 </div> 355 </div> 356 </section> 357 358 <section class="section"> 359 <div class="container is-max-desktop"> 360 <div class="columns is-centered has-text-centered"> 361 <div class="column is-full-width"> 362 <h2 class="title is-3">NEP-τ: Temporal Offset Ablation</h2> 363 <div class="content has-text-justified"> 364 <figure id="offset_ablation"> 365 <img src="./static/images/2_offset.png" alt="temporal_offset_ablation" /> 366 </figure> 367 <p> 368 A key design choice of the Generalist-IDM is <strong>NEP-τ</strong> (Next-Event 369 Prediction with Temporal Offset), which shifts the observation window forward by τ 370 milliseconds to incorporate future visual context when predicting the current action. 371 Without any offset (τ = 0), Pearson correlations collapse near zero and keyboard accuracy 372 drops sharply, confirming that future context is essential for resolving the current action. 373 A small offset (τ = 50 ms) recovers mouse prediction but remains suboptimal for keyboard 374 accuracy. Performance stabilizes at <strong>τ ≥ 100 ms</strong> with only minor variation 375 up to 200 ms, showing that NEP-τ is robust to the exact offset once sufficient future 376 context is provided. We adopt τ = 100 ms as the default in all experiments. 377 </p> 378 </div> 379 </div> 380 </div> 381 </div> 382 </section> 383 384 <!-- G-IDM Video Examples --> 385 <section class="hero is-light is-small"> 386 <div class="hero-body"> 387 <div class="container"> 388 <div class="field"> 389 <div class="control"> 390 <div class="select is-fullwidth"> 391 <select id="gidm-game-selector"> 392 <option value="game2">Brotato (2D)</option> 393 <option value="game1">Minecraft (3D)</option> 394 </select> 395 </div> 396 </div> 397 </div> 398 399 <div id="game1" class="game-videos" style="display: none;"> 400 <div class="columns is-multiline"> 401 <div class="column is-4"> 402 <video poster="" id="minecraft-gt" controls muted loop playsinline width="100%"> 403 <source src="./static/videos/minecraft-gt.mp4" type="video/mp4"> 404 </video> 405 <p class="is-size-6 mt-2 has-text-centered">Ground Truth</p> 406 </div> 407 <div class="column is-4"> 408 <video poster="" id="minecraft-idm" controls muted loop playsinline width="100%"> 409 <source src="./static/videos/minecraft-idm.mp4" type="video/mp4"> 410 </video> 411 <p class="is-size-6 mt-2 has-text-centered">IDM</p> 412 </div> 413 <div class="column is-4"> 414 <video poster="" id="minecraft-gidm" controls muted loop playsinline width="100%"> 415 <source src="./static/videos/minecraft-gidm.mp4" type="video/mp4"> 416 </video> 417 <p class="is-size-6 mt-2 has-text-centered">G-IDM</p> 418 </div> 419 </div> 420 </div> 421 422 <div id="game2" class="game-videos"> 423 <div class="columns is-multiline"> 424 <div class="column is-4"> 425 <video poster="" id="brotato-gt" controls muted loop playsinline width="100%"> 426 <source src="./static/videos/brotato-gt.mp4" type="video/mp4"> 427 </video> 428 <p class="is-size-6 mt-2 has-text-centered">
428Ground Truth</p> 429 </div> 430 <div class="column is-4"> 431 <video poster="" id="brotato-idm" controls muted loop playsinline width="100%"> 432 <source src="./static/videos/brotato-idm.mp4" type="video/mp4"> 433 </video> 434 <p class="is-size-6 mt-2 has-text-centered">IDM</p> 435 </div> 436 <div class="column is-4"> 437 <video poster="" id="brotato-gidm" controls muted loop playsinline width="100%"> 438 <source src="./static/videos/brotato-gidm.mp4" type="video/mp4"> 439 </video> 440 <p class="is-size-6 mt-2 has-text-centered">G-IDM</p> 441 </div> 442 </div> 443 </div> 444 </div> 445 </div> 446 </section> 447 448 <section class="section"> 449 <div class="container is-max-desktop"> 450 <div class="columns is-centered has-text-centered"> 451 <div class="column is-full-width"> 452 <h2 class="title is-3">Out-of-Distribution Generalization</h2> 453 <div class="content has-text-justified"> 454 <figure id="gidm_ood"> 455 <img src="./static/images/4_gidm_ood.png" alt="gidm_ood" /> 456 </figure> 457 <p> 458 We evaluate G-IDM on two unseen games: Battlefield 6 (3D) and Ogu and the Secret Forest 459 (2D). In Battlefield 6, G-IDM achieves <strong>63%</strong> keyboard accuracy, matching 460 or slightly outperforming the Specialist-IDM. When provided with a few-shot prefix, the 461 predicted mouse scale improves significantly, demonstrating in-context adaptation to mouse 462 sensitivity. In Ogu and the Secret Forest, G-IDM more than doubles the Specialist-IDM's 463 performance (from ~12% to nearly 28%), showing substantial gains even under a large domain 464 gap. 465 </p> 466 </div> 467 </div> 468 </div> 469 </div> 470 </section> 471 472 <!-- OOD Results --> 473 <section class="hero is-light is-small"> 474 <div class="hero-body"> 475 <div class="container"> 476 <div class="field"> 477 <div class="control"> 478 <div class="select is-fullwidth"> 479 <select id="ood-game-selector"> 480 <option value="ood1">Battlefield 6 (3D)</option> 481 <option value="ood2">Ogu and the Secret Forest (2D)</option> 482 </select> 483 </div> 484 </div> 485 </div> 486 487 <div id="ood1" class="ood-videos"> 488 <div class="columns is-multiline"> 489 <div class="column is-one-fifth"> 490 <video poster="" id="battlefield-gt" controls muted loop playsinline width="100%"> 491 <source src="./static/videos/battlefield-gt.mp4" type="video/mp4"> 492 </video> 493 <p class="is-size-6 mt-2 has-text-centered">Ground Truth</p> 494 </div> 495 <div class="column is-one-fifth"> 496 <video poster="" id="battlefield-idm-ft" controls muted loop playsinline width="100%"> 497 <source src="./static/videos/battlefield-idm-ft.mp4" type="video/mp4"> 498 </video> 499 <p class="is-size-6 mt-2 has-text-centered">IDM (Fine Tune)</p> 500 </div> 501 <div class="column is-one-fifth"> 502 <video poster="" id="battlefield-gidm-zero" controls muted loop playsinline width="100%"> 503 <source src="./static/videos/battlefield-gidm-zero.mp4" type="video/mp4"> 504 </video> 505 <p class="is-size-6 mt-2 has-text-centered">G-IDM (Zero Shot)</p> 506 </div> 507 <div class="column is-one-fifth"> 508 <video poster="" id="battlefield-gidm-few" controls muted loop playsinline width="100%"> 509 <source src="./static/videos/battlefield-gidm-few.mp4" type="video/mp4"> 510 </video> 511 <p class="is-size-6 mt-2 has-text-centered">G-IDM (Few Shot)</p> 512 </div> 513 <div class="column is-one-fifth"> 514 <video poster="" id="battlefield-gidm-ft" controls muted loop playsinline width="100%"> 515 <source src="./static/videos/battlefield-gidm-ft.mp4" type="video/mp4"> 516 </video> 517 <p class="is-size-6 mt-2 has-text-centered">G-IDM (Fine Tune)</p> 518 </div> 519 </div> 520 <div class="content has-text-justified mt-4"> 521 <p class="is-size-6"> 522 In Battlefield 6, for G-IDM (Zero Shot) we can observe that scale of mouse movement is 523 different with GT, 524 but we can observe that scale of movement remain nearly same for G-IDM (Few Shot). 525 This improvement occurs because providing context examples helps the model calibrate the 526 appropriate movement scale. 527 </p> 528 </div> 529 </div> 530 531 <div id="ood2" class="ood-videos" style="display: none;"> 532 <div class="columns is-multiline"> 533 <div class="column is-one-fifth"> 534 <video poster="" id="oguforest-gt" controls muted loop playsinline width="100%"> 535 <source src="./static/videos/oguforest-gt.mp4" type="video/mp4"> 536 </video> 537 <p class="is-size-6 mt-2 has-text-centered">
537Ground Truth</p> 538 </div> 539 <div class="column is-one-fifth"> 540 <video poster="" id="oguforest-idm-ft" controls muted loop playsinline width="100%"> 541 <source src="./static/videos/oguforest-idm-ft.mp4" type="video/mp4"> 542 </video> 543 <p class="is-size-6 mt-2 has-text-centered">IDM (Fine Tune)</p> 544 </div> 545 <div class="column is-one-fifth"> 546 <video poster="" id="oguforest-gidm-zero" controls muted loop playsinline width="100%"> 547 <source src="./static/videos/oguforest-gidm-zero.mp4" type="video/mp4"> 548 </video> 549 <p class="is-size-6 mt-2 has-text-centered">G-IDM (Zero Shot)</p> 550 </div> 551 <div class="column is-one-fifth"> 552 <video poster="" id="oguforest-gidm-few" controls muted loop playsinline width="100%"> 553 <source src="./static/videos/oguforest-gidm-few.mp4" type="video/mp4"> 554 </video> 555 <p class="is-size-6 mt-2 has-text-centered">G-IDM (Few Shot)</p> 556 </div> 557 <div class="column is-one-fifth"> 558 <video poster="" id="oguforest-gidm-ft" controls muted loop playsinline width="100%"> 559 <source src="./static/videos/oguforest-gidm-ft.mp4" type="video/mp4"> 560 </video> 561 <p class="is-size-6 mt-2 has-text-centered">G-IDM (Fine Tune)</p> 562 </div> 563 </div> 564 </div> 565 </div> 566 </div> 567 </section> 568 569 <section class="section"> 570 <div class="container is-max-desktop"> 571 <div class="columns is-centered has-text-centered"> 572 <div class="column is-full-width"> 573 <h2 class="title is-3">Desktop-to-Embodied Transfer</h2> 574 <div class="content has-text-justified"> 575 <p> 576 To validate the effectiveness of desktop pretraining for embodied AI, we evaluate our 577 approach on three challenging downstream tasks: 578 <strong>LIBERO manipulation</strong>, <strong>CANVAS navigation</strong>, and 579 <strong>SO101 real-world pick-and-place</strong>. These 580 benchmarks represent diverse embodied scenarios 581 requiring different sensorimotor skills - precise object manipulation, spatial navigation, 582 and real-world pick-and-place. 583 Our Vision-Action Pretraining (VAPT) framework transfers desktop-pretrained representations 584 to these physical domains, 585 demonstrating that sensorimotor patterns learned from gaming environments can generalize to 586 real-world robotic tasks. 587 </p> 588 </div> 589 </div> 590 </div> 591 </div> 592 </section> 593 594 <section class="section"> 595 <div class="container is-max-desktop"> 596 <div class="columns is-centered has-text-centered"> 597 <div class="column is-full-width"> 598 <h2 class="title is-3">LIBERO Manipulation Results</h2> 599 <div class="content has-text-justified"> 600 <figure id="libero_results"> 601 <img src="./static/images/5_libero.png" alt="libero_results" /> 602 </figure> 603 <p> 604 VAPT without pseudo-labels achieves <strong>96.6%</strong> total success and 605 <strong>93.6%</strong> on long-horizon tasks, comparable to or surpassing much larger 606 models such as π<sub>0</sub> (3.3B) and OpenVLA (7B). 607 Our 1B-parameter model shows particularly strong advantages on long-horizon tasks that 608 require careful action sequencing. 609 </p> 610 </div> 611 </div> 612 </div> 613 </div> 614 </section> 615 616 <!-- LIBERO Video Examples --> 617 <section class="hero is-light is-small"> 618 <div class="hero-body"> 619 <div class="container"> 620 <div class="field"> 621 <div class="control"> 622 <div class="select is-fullwidth"> 623 <select id="libero-task-selector"> 624 <option value="task1">TASK: Put both the alphabet soup and the cream cheese box in the
625 basket</option> 626 <option value="task2">TASK: Put both the alphabet soup and the tomato sauce in the 627 basket</option> 628 <option value="task3">TASK: Put both moka pots on the stove</option> 629 </select> 630 </div> 631 </div> 632 </div> 633 634 <div id="task1" class="task-videos"> 635 <div class="columns is-multiline"> 636 <div class="column is-4"> 637 <video poster="" id="libero-1-base" controls muted loop playsinline width="100%"> 638 <source src="./static/videos/libero-1-base.mp4" type="video/mp4"> 639 </video> 640 <p class="is-size-6 mt-2 has-text-centered">Baseline (42%)</p> 641 </div> 642 <div class="column is-4"> 643 <video poster="" id="libero-1-ft" controls muted loop playsinline width="100%"> 644 <source src="./static/videos/libero-1-ft.mp4" type="video/mp4"> 645 </video> 646 <p class="is-size-6 mt-2 has-text-centered">+ VAPT w/o pseudo (98%)</p> 647 </div> 648 <div class="column is-4"> 649 <video poster="" id="libero-1-ptft" controls muted loop playsinline width="100%"> 650 <source src="./static/videos/libero-1-ptft.mp4" type="video/mp4"> 651 </video> 652 <p class="is-size-6 mt-2 has-text-centered">+ VAPT w/ pseudo (100%)</p> 653 </div> 654 </div> 655 </div> 656 657 <div id="task2" class="task-videos" style="display: none;"> 658 <div class="columns is-multiline"> 659 <div class="column is-4"> 660 <video poster="" id="libero-2-base" controls muted loop playsinline width="100%"> 661 <source src="./static/videos/libero-2-base.mp4" type="video/mp4"> 662 </video> 663 <p class="is-size-6 mt-2 has-text-centered">Baseline (36%)</p> 664 </div> 665 <div class="column is-4"> 666 <video poster="" id="libero-2-ft" controls muted loop playsinline width="100%"> 667 <source src="./static/videos/libero-2-ft.mp4" type="video/mp4"> 668 </video> 669 <p class="is-size-6 mt-2 has-text-centered">+ VAPT w/o pseudo (88%)</p> 670 </div> 671 <div class="column is-4"> 672 <video poster="" id="libero-2-ptft" controls muted loop playsinline width="100%"> 673 <source src="./static/videos/libero-2-ptft.mp4" type="video/mp4"> 674 </video> 675 <p class="is-size-6 mt-2 has-text-centered">+ VAPT w/ pseudo (90%)</p> 676 </div> 677 </div> 678 </div> 679 680 <div id="task3" class="task-videos" style="display: none;"> 681 <div class="columns is-multiline"> 682 <div class="column is-4"> 683 <video poster="" id="libero-3-base" controls muted loop playsinline width="100%"> 684 <source src="./static/videos/libero-3-base.mp4" type="video/mp4"> 685 </video> 686 <p class="is-size-6 mt-2 has-text-centered">Baseline (10%)</p> 687 </div> 688 <div class="column is-4"> 689 <video poster="" id="libero-3-ft" controls muted loop playsinline width="100%"> 690 <source src="./static/videos/libero-3-ft.mp4" type="video/mp4"> 691 </video> 692 <p class="is-size-6 mt-2 has-text-centered">+ VAPT w/o pseudo (88%)</p> 693 </div> 694 <div class="column is-4"> 695 <video poster="" id="libero-3-ptft" controls muted loop playsinline width="100%"> 696 <source src="./static/videos/libero-3-ptft.mp4" type="video/mp4"> 697 </video> 698 <p class="is-size-6 mt-2 has-text-centered">+ VAPT w/ pseudo (46%)</p> 699 </div> 700 </div> 701 </div> 702 </div> 703 </div> 704 </section> 705 706 <section class="section"> 707 <div class="container is-max-desktop"> 708 <div class="columns is-centered has-text-centered"> 709 <div class="column is-full-width"> 710 <h2 class="title is-3">CANVAS Navigation Results</h2> 711 <div class="content has-text-justified"> 712 <figure id="canvas_results"> 713 <img src="./static/images/7_canvas.png" alt="canvas_results" /> 714 </figure> 715 <p> 716 Adding pseudo-labeled demonstrations increases navigation performance to 717 <strong>83.3%</strong>, an 8-point improvement over the baseline. 718 The benefit is especially large under misleading instructions, as in 719 <em>sim_orchard</em> (86.7% vs. 53.3%) and <em>sim_street_sidewalk</em> 720 (73.3% vs. 40.0%), indicating that pseudo-labeling is particularly useful for navigation 721 tasks where success depends on high-level planning rather than precise low-level control. 722 </p> 723 </div> 724 </div> 725 </div> 726 </div> 727 </section> 728 729 <!-- CANVAS Video Examples --> 730 <section class="hero is-light is-small"> 731 <div class="hero-body"> 732 <div class="container"> 733 <div class="field"> 734 <div class="control"> 735 <div class="select is-fullwidth"> 736 <select id="canvas-env-selector"> 737 <option value="env1">Environment: sim_gallery</option> 738 <option value="env2">Environment: sim_street_sidewalk</option> 739 </select> 740 </div> 741 </div> 742 </div> 743 744 <div id="env1" class="env-videos"> 745 <div class="columns is-multiline"> 746 <div class="column is-6"> 747 <video poster="" id="canvas-1-base" controls muted loop playsinline width="100%"> 748 <source src="./static/videos/3_baseline_fail.mp4" type="video/mp4"> 749 </video> 750 <p class="is-size-6 mt-2 has-text-centered">Baseline (fail)</p> 751 </div> 752 <div class="column is-6"> 753 <video poster="" id="canvas-1-vapt" controls muted loop playsinline width="100%"> 754 <source src="./static/videos/3_vapt_success.mp4" type="video/mp4"> 755 </video> 756 <p class="is-size-6 mt-2 has-text-centered">+ VAPT w/ pseudo (success)</p> 757 </div> 758 </div> 759 </div> 760 761 <div id="env2" class="env-videos" style="display: none;"> 762 <div class="columns is-multiline"> 763 <div class="column is-6"> 764 <video poster="" id="canvas-2-base" controls muted loop playsinline width="100%"> 765 <source src="./static/videos/1_baseline_fail.mp4" type="video/mp4"> 766 </video> 767 <p class="is-size-6 mt-2 has-text-centered">Baseline (fail)</p> 768 </div> 769 <div class="column is-6"> 770 <video poster="" id="canvas-2-vapt" controls muted loop playsinline width="100%"> 771 <source src="./static/videos/1_vapt_success.mp4" type="video/mp4"> 772 </video> 773 <p class="is-size-6 mt-2 has-text-centered">+ VAPT w/ pseudo (success)</p> 774 </div> 775 </div> 776 </div> 777 </div> 778 </div> 779 </section> 780 781 <section class="section"> 782 <div class="container is-max-desktop"> 783 <div class="columns is-centered has-text-centered"> 784 <div class="column is-full-width"> 785 <h2 class="title is-3">Meta-World & SO101 Real-World Results</h2> 786 <div class="content has-text-justified"> 787 <figure id="meta_real_results"> 788 <img src="./static/images/6_meta_real.png" alt="meta_world_so101_results" /> 789 </figure> 790 <p> 791 VAPT consistently outperforms the baseline on Meta-World across all difficulty levels, 792 with gains most pronounced on Hard and Very Hard tasks. 793 We further validate our approach with a real-world pick-and-place experiment using an 794 SO101 robot arm, following the evaluation protocol of SmolVLA (Shukor et al., 2025). The 795 task requires grasping a blue cube and placing it in a white box, with the cube placed at 796 five distinct initial positions. We collect 208 demonstration episodes and evaluate each 797 trained policy over 30 rollouts. The baseline InternVL3-1B achieves a 70% success rate, 798 while both VAPT variants reach <strong>80%</strong>, confirming that VAPT transfers 799 effectively to real-world hardware. 800 </p> 801 </div> 802 </div> 803 </div> 804 </div> 805 </section> 806 807 <section class="hero is-light is-small"> 808 <div class="hero-body"> 809 <div class="container"> 810 <div class="columns is-multiline"> 811 <div class="column is-6"> 812 <video poster="" id="so101-baseline-right" controls muted loop playsinline width="100%"> 813 <source src="./static/videos/so101_baseline_fail_right.mp4" type="video/mp4"> 814 </video> 815 <p class="is-size-6 mt-2 has-text-centered">Baseline (fail) - right view</p> 816 </div> 817 <div class="column is-6"> 818 <video poster="" id="so101-vapt-right" controls muted loop playsinline width="100%"> 819 <source src="./static/videos/so101_ours_success_right.mp4" type="video/mp4"> 820 </video> 821 <p class="is-size-6 mt-2 has-text-centered">+ VAPT (success) - right view</p> 822 </div> 823 <div class="column is-6"> 824 <video poster="" id="so101-baseline-top" controls muted loop playsinline width="100%"> 825 <source src="./static/videos/so101_baseline_fail_top.mp4" type="video/mp4"> 826 </video> 827 <p class="is-size-6 mt-2 has-text-centered">Baseline (fail) - top view</p> 828 </div> 829 <div class="column is-6">
830 <video poster="" id="so101-vapt-top" controls muted loop playsinline width="100%"> 831 <source src="./static/videos/so101_ours_success_top.mp4" type="video/mp4"> 832 </video> 833 <p class="is-size-6 mt-2 has-text-centered">+ VAPT (success) - top view</p> 834 </div> 835 </div> 836 </div> 837 </div> 838 </section> 839 840 <section class="section" id="BibTeX"> 841 <div class="container is-max-desktop content"> 842 <h2 class="title">BibTeX</h2> 843 <pre><code>@inproceedings{choi2026d2e, 844 title={D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI}, 845 author={Choi, Suhwan and Jung, Jaeyoon and Seong, Haebin and Kim, Minchan and Kim, Minyeong and Cho, Yongjun and Kim, Yoonshik and Park, Yu and Yu, Youngjae and Lee, Yunsung}, 846 booktitle={International Conference on Learning Representations}, 847 volume={2026}, 848 pages={46207--46236}, 849 year={2026} 850}</code></pre> 851 </div> 852 </section> 853 <br> 854 <center class="is-size-10"> 855 The website design was based on <a 856 href="https://github.com/general-navigation-models/general-navigation-models.github.io"><span 857 class="dnerf">general-navigation-models</span></a> adapted from <a href="https://nerfies.github.io" 858 class="external-link"><span class="dnerf">Nerfies</span></a>. 859 </center> 860 <br> 861</body> 862 863</html>
Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.