1<!DOCTYPE html> 2<html lang="en"> 3<head> 4 <meta charset="UTF-8"> 5 <title>GHOST: Hierarchical Sub-Goal Policies for Generalizing Robot Manipulation</title> 6 <link rel="icon" href="data:image/svg+xml,<svg xmlns=%22http://www.w3.org/2000/svg%22 viewBox=%220 0 100 100%22><text y=%22.9em%22 font-size=%2290%22>ð»</text></svg>"> 7 <link rel="stylesheet" href="style.css"> 8</head> 9<body> 10 <div class="toc"> 11 <h3>Content</h3> 12 <hr> 13 <ul> 14 <li><a href="#abstract">Abstract</a></li> 15 <li><a href="#approach">Approach</a></li> 16 <li class="toc-subsection"><a href="#data-collection">Data Collection</a></li> 17 <li class="toc-subsection"><a href="#high-level-policy">High-Level Policy</a></li> 18 <li class="toc-subsection"><a href="#low-level-policy">Low-Level Policy</a></li> 19 <li><a href="#interactive-visualization">Interactive Visualization</a></li> 20 <li><a href="#experiments">Experiments</a></li> 21 <li><a href="#bibtex">BibTeX</a></li> 22 </ul> 23 </div> 24 25 <div class="main-content"> 26 <div class="hero-text">ð» GHOST</div> 27 <div class="sub-hero-text">Hierarchical Sub-Goal Policies for Generalizing Robot Manipulation</div> 28 29 <!-- Authors --> 30 <div class="authors"> 31 <a href="https://sriramsk1999.github.io/" target="_blank">Sriram Krishna</a><sup>1</sup>, 32 <a href="https://beisner.me/" target="_blank">Ben Eisner</a><sup>1</sup>, 33 <a href="https://www.linkedin.com/in/haotian-zhan-96935126a/" target="_blank">Haotian Zhan</a><sup>1</sup>, 34 <a href="https://yingyuan0414.github.io/" target="_blank">Ying Yuan</a><sup>1</sup>, 35 <a href="https://haoyuzhen.com/" target="_blank">Haoyu Zhen</a><sup>2</sup>, 36 <a href="https://people.csail.mit.edu/ganchuang/" target="_blank">Chuang Gan</a><sup>2</sup>, <br> 37 <a href="https://shubhtuls.github.io/" target="_blank">Shubham Tulsiani</a><sup>1</sup>, 38 <a href="https://davheld.github.io/" target="_blank">David Held</a><sup>1</sup> 39 <span class="affiliation"><sup>1</sup>Robotics Institute, Carnegie Mellon University <sup>2</sup>UMass Amherst</span> 40 <span class="affiliation" style="color: #555; text-align: center; font-size: 20px; margin-top: 6px;">Robotics: Science and Systems (RSS) 2026</span> 41 </div> 42 <!-- End Authors --> 43 44 <!-- Quick Links --> 45 <div class="quick-links"> 46 <a href="ghost.pdf" target="_blank">[pdf]</a> 47 <a href="https://arxiv.org/abs/2606.10025" target="_blank">[arxiv]</a> 48 <a href="https://github.com/r-pad/ghost" target="_blank">[code]</a> 49 </div> 50 51 <!-- Teaser Figure --> 52 <div style="width: 100%; margin: 20px 0;"> 53 <img src="figs/ghost-teaser.png" alt="GHOST Teaser Figure" style="width: 100%; border-radius: 15px; border: 1.5px solid #000;"> 54 </div> 55 <!-- Caption for Teaser --> 56 <p class="figure-caption"> 57 <b>GHOST</b> learns skills from robot teleoperation data and optionally uses human demonstrations to generalize to novel tasks. We train a hierarchical policy that decouples embodiment-agnostic goal prediction (Ï<sub>hi</sub>) from embodiment-specific action execution (Ï<sub>lo</sub>). By training Ï<sub>hi</sub> on both robot and human data and Ï<sub>lo</sub> purely on robot data, GHOST transfers learned manipulation skills to out-of-distribution tasks with third person human video demonstrations. 58 </p> 59 60 <div class="tagline" id="abstract">Abstract.</div> 61 62 <div class="section"> 63 We present <b>GHOST</b>, a framework for learning visuomotor manipulation policies that <i>generalize</i> beyond the training distribution. GHOST factorizes control into (i) a high-level policy that predicts the next sub-goal as a <i>distribution</i> over 3D end-effector poses from multi-view RGB-D observations, and (ii) a low-level goal-conditioned controller that executes embodiment-specific actions. To condition image-based policies on 3D goals, we introduce a simple spatial interface that projects predicted goals into the image plane and represents them as <i>end-effector heatmaps</i>. Across a suite of manipulation tasks, this hierarchical factorization consistently improves performance and robustness compared to a flat Diffusion Policy. 64 <br><br> 65 Further, we show that this hierarchical interface also makes it easy to incorporate human demonstrations without relying on (noisy) action retargeting. As sub-goals are largely embodiment-agnostic, we train the high-level policy on human video to specify how learned skills should be applied and composed, while keeping the low-level policy trained purely on robot data. This hierarchy enables adaptation to novel objects and task variations using a small number of human demonstrations. 66 </div> 67 68 <div class="tagline" id="approach">
68Approach.</div> 69 70 <div class="section"> 71 The GHOST framework learns a hierarchy over <i>sub-goal end-effector poses</i>. During training, sub-goal boundaries provide supervision for the high-level planner Ï<sub>hi</sub>, while the full robot trajectories provide action supervision for the low-level controller Ï<sub>lo</sub>. At test time, Ï<sub>hi</sub> predicts the next end-effector sub-goal (as a distribution) from workspace observations and language, and Ï<sub>lo</sub> executes goal-conditioned control toward that sub-goal. This separation isolates long-horizon reasoning in Ï<sub>hi</sub> and maintains precise, embodiment-specific control in Ï<sub>lo</sub>. 72 </div> 73 74 <div class="section-subtitle" id="data-collection">1. Data Collection and Processing.</div> 75 76 <div class="section"> 77 We collect two types of demonstrations: robot teleoperation data and human demonstrations. For robot data, we collect demonstrations across multiple task variants, each demonstrating the same skill (e.g., pick-and-place). When possible, sub-goals <i>s</i> are extracted automatically by identifying the timesteps where the gripper state transitions from open → close and vice versa. For human data, we optionally collect demonstrations of the same skill in a novel setting using either new objects or novel applications of the skill. We track 3D hand poses using off-the-shelf hand pose estimators, resolving the weak-perspective scale ambiguity by segmenting the hand with Grounded-SAM and scaling the detection with the observed depth of the hand. Sub-goals are manually annotated at the timesteps where a sub-goal terminates. 78 79 <!-- Demonstration Videos: Robot vs Human --> 80 <div style="width: 100%; margin: 30px 0;"> 81 <div style="display: flex; gap: 20px; justify-content: center; align-items: flex-start; margin-bottom: 10px;"> 82 <div style="width: 48%; text-align: center;"> 83 <video style="width: 100%; border-radius: 15px; border: 1.5px solid #000;" autoplay loop muted playsinline> 84 <source src="vids/onesie_teleop.mp4" type="video/mp4"> 85 </video> 86 </div> 87 <div style="width: 48%; text-align: center;"> 88 <video style="width: 100%; border-radius: 15px; border: 1.5px solid #000;" autoplay loop muted playsinline> 89 <source src="vids/towel_human_demo.mp4" type="video/mp4"> 90 </video> 91 </div> 92 </div> 93 </div> 94 95 <!-- Human Demo Supercut --> 96 <div style="width: 100%; margin: 20px 0; text-align: center;"> 97 <video style="width: 100%; border-radius: 15px; border: 1.5px solid #000;" autoplay loop muted playsinline> 98 <source src="vids/human_demo_supercut.mp4" type="video/mp4"> 99 </video> 100 <p class="figure-caption"> 101 <b>Human demonstration collection.</b> We collect human video demonstrations for out-of-distribution generalization, annotating sub-goal timesteps where key manipulation events occur (e.g., grasping, placing, folding transitions). 102 </p> 103 </div> 104 105 To enable cross-embodiment transfer, we represent the end-effector pose as a sparse set of 4 3D keypoints rather than a position and SO(3) rotation. For robot data, we sample a point cloud from the gripper mesh and select 4 points at the base of the gripper, the tips of the two gripper fingers, and the grasping center. For human data, we extract hand point clouds from the MANO mesh of the hand pose and select 4 corresponding points - the palm, tips of the thumb and index finger, and the grasping center. This unified end-effector representation allows the high-level policy to be trained on both robot and human demonstrations. 106 </div> 107 108 <!-- 3D Viewers for End-Effector Representations --> 109 <div style="width: 100%; margin: 30px 0; display: flex; gap: 20px; justify-content: center; align-items: flex-start;"> 110 <div style="width: 48%; text-align: center;"> 111 <canvas id="aloha-viewer" style="width: 100%; height: 350px; border-radius: 15px; border: 1.5px solid #000; background: linear-gradient(135deg, #e3f2fd 0%, #bbdefb 100%);"></canvas> 112 <p style="font-family: 'Roboto Mono', monospace; font-size: 14px; margin-top: 10px; color: #333;">ALOHA Gripper</p> 113 </div> 114 <div style="width: 48%; text-align: center;">
115 <canvas id="hand-viewer" style="width: 100%; height: 350px; border-radius: 15px; border: 1.5px solid #000; background: linear-gradient(135deg, #f0f4f8 0%, #d9e2ec 100%);"></canvas> 116 <p style="font-family: 'Roboto Mono', monospace; font-size: 14px; margin-top: 10px; color: #333;">MANO Hand</p> 117 </div> 118 </div> 119 <p class="figure-caption"> 120 <b>End-effector representations for robot and human demonstrations.</b> Instead of representing the gripper pose as a translation and SO(3) rotation, we represent the gripper pose as a set of 3D keypoints (red spheres): gripper base, fingertip locations, and grasping center. <i>Left</i>: Parallel-jaw gripper. <i>Right</i>: MANO hand. (Interactive - drag to rotate, scroll to zoom) 121 </p> 122 123 <div class="section-subtitle" id="high-level-policy">2. High-Level Goal Prediction Policy.</div> 124 125 <!-- High-Level Policy Figure --> 126 <div style="width: 100%; margin: 20px 0;"> 127 <img src="figs/dino_3dgp.png" alt="High-Level Goal Prediction Architecture" style="width: 100%; border-radius: 15px; border: 1.5px solid #000;"> 128 </div> 129 <p class="figure-caption"> 130 <b>GHOST High-level sub-goal prediction architecture:</b> RGB-D observations from multiple cameras are processed with a DINOv3 encoder, with patch tokens augmented by 3D coordinates. Additional context (gripper state, language embedding, embodiment name) is encoded via separate MLPs. A decoder-only transformer processes all tokens, with each patch predicting the GMM parameters of a 3D sub-goal distribution over end-effector keypoints. 131 </p> 132 133 <div class="section"> 134 The high-level policy Ï<sub>hi</sub> predicts the end-effector pose at the next sub-goal timestep. We parameterize it as a decoder-only transformer operating on RGB-D observations from multiple cameras. Each RGB-D image is processed independently with a frozen DINOv3 encoder, producing K×C patch tokens (K per camera). Each patch token is augmented with the features of the 3D coordinate [x, y, z] of the patch center, processed with an MLP. Additional tokens encode the current gripper pose and the task language embedding via Flan-T5, and a learnable token specifies whether the demonstration is from a human or robot. We also include register tokens, which have been shown to improve performance on dense prediction tasks. To handle the multimodality inherent in demonstration data, we model the goal distribution as a dense per-patch Gaussian Mixture Model (GMM). Each patch token predicts (1) a mixing weight w<sub>i</sub>, and (2) four 3D residual vectors {δ<sub>i,1</sub>, δ<sub>i,2</sub>, δ<sub>i,3</sub>, δ<sub>i,4</sub>} relative to the patch center p<sub>i</sub>; the global goal distribution is a mixture of K×C isotropic Gaussians (one per patch). 135 </div> 136 137 <!-- GMM Inference Figure --> 138 <div style="width: 100%; margin: 20px 0;"> 139 <video style="width: 100%; border-radius: 15px; border: 1.5px solid #000;" autoplay loop muted playsinline> 140 <source src="vids/gmm_quiver_vis.mp4" type="video/mp4"> 141 </video> 142 </div> 143 <p class="figure-caption"> 144 <b>High-level policy GMM predictions at inference.</b> We visualize the mixture components of the GMM, with the opacity encoding mixing weight w<sub>i</sub> and arrows showing projections of 3D residuals δ<sub>i</sub> from patch centers p<sub>i</sub>. 145 </p> 146 147 <div class="section-subtitle" id="low-level-policy">3. Low-Level Goal-Conditioned Policy.</div> 148 149 <div class="section"> 150 We instantiate Ï<sub>lo</sub> as a Diffusion Policy that generates action chunks by denoising conditioned on observations. Each point in the 3D goal e<sub>s</sub> is projected onto the image plane of each camera using the camera parameters, yielding sparse 2D coordinates p<sub>i</sub>. These are converted to dense <i>end-effector heatmaps</i> H â â<sup>3×h×w</sup>, where each channel <i>c</i> encodes the square-root pixel distance field â(||x<sub>xy</sub> - p<sub>c</sub>||<sub>2</sub> / d<sub>max</sub>) from a keypoint, normalized by the image diagonal d<sub>max</sub>. We use heatmaps rather than single-pixel binary masks because we empirically find them to yield better performance; we hypothesize this is due to the denser supervisory signal. For practical implementation reasons, we select three of the heatmap channels corresponding to non-collinear points in the goal and pass these channels through standard 3-channel image encoder networks before incorporating into the policy network. 151 </div> 152 153 <!-- End-Effector Heatmap Figure --> 154 <div style="width: 100%; margin: 20px 0;"> 155 <img src="figs/eef_heatmap.png" alt="End-Effector Heatmap Generation" style="width: 100%; border-radius: 15px; border: 1.5px solid #000;"> 156 </div> 157 <p class="figure-caption"> 158 <b>End-effector heatmap generation for goal conditioning.</b> Predicted 3D keypoints are projected to 2D coordinates, and then converted to dense distance field heatmaps that encode spatial proximity to each keypoint. 159 </p> 160 161 <!-- Low-Level Architecture Figure --> 162 <div style="width: 100%; margin: 20px 0;"> 163 <img src="figs/low_level_arch.png" alt="Low-Level Goal-Conditioned Policy Architecture" style="width: 100%; border-radius: 15px; border: 1.5px solid #000;"> 164 </div> 165 <p class="figure-caption"> 166 <b>GHOST Low-level goal-conditioned policy architecture.</b> Images from each camera and the projected end-effector heatmap images are processed independently with ResNet encoders and concatenated along with the proprioceptive input into a global conditioning vector for the Diffusion Policy. 167 </p> 168 169 <div class="tagline" id="interactive-visualization">Interactive 4D Visualization.</div> 170 171 <div class="section"> 172 Explore an interactive visualization of a GHOST rollout on the onesie folding task. The viewer contains synchronized feeds from 3 workspace cameras alongside a 3D point cloud reconstruction of the scene, all scrubbable over time. Drag to rotate the 3D view, scroll to zoom. Use the timeline at the bottom to scrub through the episode. (Video downsampled for efficient streaming). 173 </div> 174 175 <div style="width: 85vw; position: relative; margin-left: calc((100% - 85vw) / 2); box-sizing: border-box;"> 176 <iframe src="https://app.rerun.io/version/latest/?url=https://ghost-human-demo.github.io/fold_onesie_gc_20260124_recording.rrd" style="width: 100%; height: 85vh; border: 1.5px solid #000; border-radius: 15px;" allow="fullscreen"></iframe> 177 </div> 178 <p class="figure-caption"> 179 <b>Interactive 4D visualization of a GHOST rollout.</b> 3 synchronized camera views and a 3D point cloud reconstruction, powered by <a href="https://rerun.io" target="_blank" style="color: #555; text-decoration: underline;">Rerun</a>. Drag to rotate the 3D view, scroll to zoom, and use the timeline to scrub through the episode. 180 </p> 181 182 <div class="tagline" id="experiments">Experiments and Results.</div> 183 184 <div class="section"> 185 We evaluate GHOST across multiple challenging real-world manipulation tasks to demonstrate both in-distribution performance improvements and out-of-distribution generalization through human demonstrations. 186 </div> 187 188 <h3 style="font-family: 'Roboto Mono', monospace; font-size: 18px; margin-top: 30px;">Task Overview</h3> 189 <table style="width: 100%; border-collapse: collapse; margin: 20px 0; font-family: 'Roboto Mono', monospace; font-size: 14px;"> 190 <thead> 191 <tr style="background-color: #333; color: white;"> 192 <th style="padding: 12px; text-align: left; border: 1px solid #ddd;">Task</th> 193 <th style="padding: 12px; text-align: left; border: 1px solid #ddd;">Data Source</th> 194 <th style="padding: 12px; text-align: center; border: 1px solid #ddd;"># Demos</th> 195 <th style="padding: 12px; text-align: left; border: 1px solid #ddd;">Generalization Type</th> 196 </tr> 197 </thead> 198 <tbody> 199 <!-- Pick-and-Place --> 200 <tr style="background-color: #f0f0f0;"> 201 <td colspan="4" style="padding: 8px; border: 1px solid #ddd; font-weight: bold; font-style: italic; text-align: center;">Pick-and-Place</td> 202 </tr> 203 <tr style="background-color: rgba(90, 155, 213, 0.2);"> 204 <td style="padding: 10px; border: 1px solid #ddd;">plate-on-table</td> 205 <td style="padding: 10px; border: 1px solid #ddd;">Robot</td> 206 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">20</td> 207 <td style="padding: 10px; border: 1px solid #ddd;">â</td> 208 </tr> 209 <tr style="background-color: rgba(90, 155, 213, 0.2);"> 210 <td style="padding: 10px; border: 1px solid #ddd;">plate-in-bin</td> 211 <td style="padding: 10px; border: 1px solid #ddd;">Robot</td> 212 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">20</td> 213 <td style="padding: 10px; border: 1px solid #ddd;">â</td> 214 </tr> 215 <tr style="background-color: rgba(90, 155, 213, 0.2);"> 216 <td style="padding: 10px; border: 1px solid #ddd;">
216mug-in-bin</td> 217 <td style="padding: 10px; border: 1px solid #ddd;">Robot</td> 218 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">20</td> 219 <td style="padding: 10px; border: 1px solid #ddd;">â</td> 220 </tr> 221 <tr style="background-color: rgba(247, 150, 70, 0.2);"> 222 <td style="padding: 10px; border: 1px solid #ddd;">mug-on-table</td> 223 <td style="padding: 10px; border: 1px solid #ddd;">Human</td> 224 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">20</td> 225 <td style="padding: 10px; border: 1px solid #ddd;">Object combination</td> 226 </tr> 227 <!-- Cloth Folding --> 228 <tr style="background-color: #f0f0f0;"> 229 <td colspan="4" style="padding: 8px; border: 1px solid #ddd; font-weight: bold; font-style: italic; text-align: center;">Cloth Folding</td> 230 </tr> 231 <tr style="background-color: rgba(90, 155, 213, 0.2);"> 232 <td style="padding: 10px; border: 1px solid #ddd;">fold-onesie</td> 233 <td style="padding: 10px; border: 1px solid #ddd;">Robot</td> 234 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">33</td> 235 <td style="padding: 10px; border: 1px solid #ddd;">â</td> 236 </tr> 237 <tr style="background-color: rgba(90, 155, 213, 0.2);"> 238 <td style="padding: 10px; border: 1px solid #ddd;">fold-shirt</td> 239 <td style="padding: 10px; border: 1px solid #ddd;">Robot</td> 240 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">50</td> 241 <td style="padding: 10px; border: 1px solid #ddd;">â</td> 242 </tr> 243 <tr style="background-color: rgba(247, 150, 70, 0.2);"> 244 <td style="padding: 10px; border: 1px solid #ddd;">fold-onesie-ood</td> 245 <td style="padding: 10px; border: 1px solid #ddd;">Human</td> 246 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">17</td> 247 <td style="padding: 10px; border: 1px solid #ddd;">Object instance</td> 248 </tr> 249 <tr style="background-color: rgba(247, 150, 70, 0.2);"> 250 <td style="padding: 10px; border: 1px solid #ddd;">fold-towel</td> 251 <td style="padding: 10px; border: 1px solid #ddd;">Human</td> 252 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">50</td> 253 <td style="padding: 10px; border: 1px solid #ddd;">Object category + Skill composition</td> 254 </tr> 255 <!-- Hammer Pin --> 256 <tr style="background-color: #f0f0f0;"> 257 <td colspan="4" style="padding: 8px; border: 1px solid #ddd; font-weight: bold; font-style: italic; text-align: center;">Hammer Pin</td> 258 </tr> 259 <tr style="background-color: rgba(90, 155, 213, 0.2);"> 260 <td style="padding: 10px; border: 1px solid #ddd;">hammer-pin</td> 261 <td style="padding: 10px; border: 1px solid #ddd;">Human+Robot</td> 262 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">100</td> 263 <td style="padding: 10px; border: 1px solid #ddd;">â</td> 264 </tr> 265 </tbody> 266 </table> 267 268 <!-- Task Demonstrations Placeholder --> 269 <div class="video-gallery-section" id="pick-place-gallery"> 270 <div class="video-gallery-container"> 271 <div class="video-gallery" id="videoGalleryPickPlace"> 272 <!-- Video 1: Plate Table 0 --> 273 <div style="width: 350px; flex-shrink: 0;"> 274 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 275 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;"> 276 <span style="font-weight: bold;">plate-on-table</span> 277 <span style="color: #666;">1/1</span> 278 <span style="font-size: 16px; color: #22c55e;">✓</span> 279 </div> 280 </div> 281 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 282 <source src="vids/plate_table_0.mp4" type="video/mp4"> 283 </video> 284 </div> 285 <!-- Video 2: Plate Table 1 --> 286 <div style="width: 350px; flex-shrink: 0;"> 287 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 288 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;">
289 <span style="font-weight: bold;">plate-on-table</span> 290 <span style="color: #666;">1/1</span> 291 <span style="font-size: 16px; color: #22c55e;">✓</span> 292 </div> 293 </div> 294 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 295 <source src="vids/plate_table_1.mp4" type="video/mp4"> 296 </video> 297 </div> 298 <!-- Video 3: Plate Table 4 --> 299 <div style="width: 350px; flex-shrink: 0;"> 300 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 301 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;"> 302 <span style="font-weight: bold;">plate-on-table</span> 303 <span style="color: #666;">1/1</span> 304 <span style="font-size: 16px; color: #22c55e;">✓</span> 305 </div> 306 </div> 307 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 308 <source src="vids/plate_table_4.mp4" type="video/mp4"> 309 </video> 310 </div> 311 <!-- Video 4: Mug Table 1 --> 312 <div style="width: 350px; flex-shrink: 0;"> 313 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 314 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;"> 315 <span style="font-weight: bold;">mug-on-table</span> 316 <span style="color: #666;">1/1</span> 317 <span style="font-size: 16px; color: #22c55e;">✓</span> 318 </div> 319 </div> 320 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 321 <source src="vids/mug_table_1.mp4" type="video/mp4"> 322 </video> 323 </div> 324 <!-- Video 5: Mug Table 3 --> 325 <div style="width: 350px; flex-shrink: 0;"> 326 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 327 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;"> 328 <span style="font-weight: bold;">mug-on-table</span> 329 <span style="color: #666;">1/1</span> 330 <span style="font-size: 16px; color: #22c55e;">✓</span> 331 </div> 332 </div> 333 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 334 <source src="vids/mug_table_3.mp4" type="video/mp4"> 335 </video> 336 </div> 337 <!-- Video 6: Mug Table 5 --> 338 <div style="width: 350px; flex-shrink: 0;"> 339 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 340 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;"> 341 <span style="font-weight: bold;">mug-on-table</span> 342 <span style="color: #666;">0/1</span> 343 <span style="font-size: 16px; color: #ef4444;">✗</span> 344 </div> 345 </div> 346 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 347 <source src="vids/mug_table_5.mp4" type="video/mp4"> 348 </video> 349 </div> 350 </div> 351 </div> 352 <div class="gallery-caption-container">
353 <div class="gallery-nav-controls"> 354 <button class="gallery-nav left" id="scrollLeftBtnPickPlace"><</button> 355 <button class="gallery-nav right" id="scrollRightBtnPickPlace">></button> 356 </div> 357 <p class="figure-caption gallery-caption"> 358 <b>Pick and Place:</b> GHOST generalizes pick-and-place skills to novel object combinations (mug-on-table) after training on in-distribution tasks (plate-on-table, plate-in-bin, mug-in-bin). We overlay a colormap on the image to visualize the predicted goal. 359 </p> 360 </div> 361 </div> 362 363 <!-- Hammer Pin Gallery --> 364 <div class="video-gallery-section" id="hammer-pin-gallery"> 365 <div class="video-gallery-container"> 366 <div class="video-gallery" id="videoGalleryHammerPin"> 367 <!-- Video 1: Hammer Pin 3 (Success) --> 368 <div style="width: 350px; flex-shrink: 0;"> 369 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 370 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;"> 371 <span style="font-weight: bold;">hammer-pin</span> 372 <span style="color: #666;">1/1</span> 373 <span style="font-size: 16px; color: #22c55e;">✓</span> 374 </div> 375 </div> 376 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 377 <source src="vids/hammer_pin_3.mp4" type="video/mp4"> 378 </video> 379 </div> 380 <!-- Video 2: Hammer Pin 4 (Success) --> 381 <div style="width: 350px; flex-shrink: 0;"> 382 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 383 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;"> 384 <span style="font-weight: bold;">hammer-pin</span> 385 <span style="color: #666;">1/1</span> 386 <span style="font-size: 16px; color: #22c55e;">✓</span> 387 </div> 388 </div> 389 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 390 <source src="vids/hammer_pin_4.mp4" type="video/mp4"> 391 </video> 392 </div> 393 <!-- Video 3: Hammer Pin 5 (Success) --> 394 <div style="width: 350px; flex-shrink: 0;"> 395 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 396 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;"> 397 <span style="font-weight: bold;">hammer-pin</span> 398 <span style="color: #666;">1/1</span> 399 <span style="font-size: 16px; color: #22c55e;">✓</span> 400 </div> 401 </div> 402 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 403 <source src="vids/hammer_pin_5.mp4" type="video/mp4"> 404 </video> 405 </div> 406 <!-- Video 4: Hammer Pin 1 (Failure) --> 407 <div style="width: 350px; flex-shrink: 0;"> 408 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 409 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;"> 410 <span style="font-weight: bold;">hammer-pin</span> 411 <span style="color: #666;">0/1</span>
412 <span style="font-size: 16px; color: #ef4444;">✗</span> 413 </div> 414 </div> 415 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 416 <source src="vids/hammer_pin_1.mp4" type="video/mp4"> 417 </video> 418 </div> 419 <!-- Video 5: Hammer Pin 2 (Failure) --> 420 <div style="width: 350px; flex-shrink: 0;"> 421 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 422 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;"> 423 <span style="font-weight: bold;">hammer-pin</span> 424 <span style="color: #666;">0/1</span> 425 <span style="font-size: 16px; color: #ef4444;">✗</span> 426 </div> 427 </div> 428 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 429 <source src="vids/hammer_pin_2.mp4" type="video/mp4"> 430 </video> 431 </div> 432 </div> 433 </div> 434 <div class="gallery-caption-container"> 435 <div class="gallery-nav-controls"> 436 <button class="gallery-nav left" id="scrollLeftBtnHammerPin"><</button> 437 <button class="gallery-nav right" id="scrollRightBtnHammerPin">></button> 438 </div> 439 <p class="figure-caption gallery-caption"> 440 <b>Hammer Pin:</b> Pick up a hammer and strike the target pin. The task requires precise grasping of the hammer tool and striking the correct pin. We overlay a colormap on the image to visualize the predicted goal. 441 </p> 442 </div> 443 </div> 444 445 <!-- Cloth Folding Placeholder --> 446 <div class="video-gallery-section" id="cloth-folding-gallery"> 447 <div class="video-gallery-container"> 448 <div class="video-gallery" id="videoGalleryClothFolding"> 449 <!-- Video 1: Fold Onesie 0 --> 450 <div style="width: 350px; flex-shrink: 0;"> 451 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 452 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;"> 453 <span style="font-weight: bold;">fold-onesie</span> 454 <span style="color: #666;">5/5</span> 455 <span style="font-size: 16px; color: #22c55e;">✓</span> 456 </div> 457 </div> 458 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 459 <source src="vids/fold_onesie_0.mp4" type="video/mp4"> 460 </video> 461 </div> 462 <!-- Video 2: Fold Onesie 2 --> 463 <div style="width: 350px; flex-shrink: 0;"> 464 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 465 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;"> 466 <span style="font-weight: bold;">fold-onesie</span> 467 <span style="color: #666;">5/5</span> 468 <span style="font-size: 16px; color: #22c55e;">✓</span> 469 </div> 470 </div> 471 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 472 <source src="vids/fold_onesie_2.mp4" type="video/mp4"> 473 </video> 474 </div> 475 <!-- Video 3: Fold Onesie 6 --> 476 <div style="width: 350px; flex-shrink: 0;"> 477 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 478 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;">
479 <span style="font-weight: bold;">fold-onesie</span> 480 <span style="color: #666;">5/5</span> 481 <span style="font-size: 16px; color: #22c55e;">✓</span> 482 </div> 483 </div> 484 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 485 <source src="vids/fold_onesie_6.mp4" type="video/mp4"> 486 </video> 487 </div> 488 <!-- Video 4: OOD Onesie 0 --> 489 <div style="width: 350px; flex-shrink: 0;"> 490 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 491 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;"> 492 <span style="font-weight: bold;">fold-onesie-ood</span> 493 <span style="color: #666;">5/5</span> 494 <span style="font-size: 16px; color: #22c55e;">✓</span> 495 </div> 496 </div> 497 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 498 <source src="vids/ood_onesie_0.mp4" type="video/mp4"> 499 </video> 500 </div> 501 <!-- Video 5: OOD Onesie 1 --> 502 <div style="width: 350px; flex-shrink: 0;"> 503 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 504 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;"> 505 <span style="font-weight: bold;">fold-onesie-ood</span> 506 <span style="color: #666;">4/5</span> 507 <span style="font-size: 16px; color: #ef4444;">✗</span> 508 </div> 509 </div> 510 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 511 <source src="vids/ood_onesie_1.mp4" type="video/mp4"> 512 </video> 513 </div> 514 <!-- Video 6: OOD Onesie 4 --> 515 <div style="width: 350px; flex-shrink: 0;"> 516 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 517 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;"> 518 <span style="font-weight: bold;">fold-onesie-ood</span> 519 <span style="color: #666;">5/5</span> 520 <span style="font-size: 16px; color: #22c55e;">✓</span> 521 </div> 522 </div> 523 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 524 <source src="vids/ood_onesie_4.mp4" type="video/mp4"> 525 </video> 526 </div> 527 <!-- Video 7: Fold Towel 0 --> 528 <div style="width: 350px; flex-shrink: 0;"> 529 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 530 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;"> 531 <span style="font-weight: bold;">fold-towel</span> 532 <span style="color: #666;">0/1</span> 533 <span style="font-size: 16px; color: #ef4444;">✗</span> 534 </div> 535 </div> 536 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 537 <source src="vids/fold_towel_0.mp4" type="video/mp4"> 538 </video> 539 </div> 540 <!-- Video 8: Fold Towel 2 --> 541 <div style="width: 350px; flex-shrink: 0;"> 542 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 543 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;">
544 <span style="font-weight: bold;">fold-towel</span> 545 <span style="color: #666;">0/1</span> 546 <span style="font-size: 16px; color: #ef4444;">✗</span> 547 </div> 548 </div> 549 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 550 <source src="vids/fold_towel_2.mp4" type="video/mp4"> 551 </video> 552 </div> 553 <!-- Video 9: Fold Towel 4 --> 554 <div style="width: 350px; flex-shrink: 0;"> 555 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 556 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;"> 557 <span style="font-weight: bold;">fold-towel</span> 558 <span style="color: #666;">1/1</span> 559 <span style="font-size: 16px; color: #22c55e;">✓</span> 560 </div> 561 </div> 562 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 563 <source src="vids/fold_towel_4.mp4" type="video/mp4"> 564 </video> 565 </div> 566 <!-- Video 10: Fold Towel 8 --> 567 <div style="width: 350px; flex-shrink: 0;"> 568 <div style="background-color: #f5f5f5; padding: 8px 12px; border-radius: 15px 15px 0 0; border: 1.5px solid #000; border-bottom: none;"> 569 <div style="display: flex; justify-content: space-between; align-items: center; font-family: 'Roboto Mono', monospace; font-size: 12px;"> 570 <span style="font-weight: bold;">fold-towel</span> 571 <span style="color: #666;">1/1</span> 572 <span style="font-size: 16px; color: #22c55e;">✓</span> 573 </div> 574 </div> 575 <video style="width: 100%; height: 220px; border-radius: 0 0 15px 15px; border: 1.5px solid #000; border-top: none; object-fit: cover;" autoplay loop muted playsinline> 576 <source src="vids/fold_towel_8.mp4" type="video/mp4"> 577 </video> 578 </div> 579 </div> 580 </div> 581 <div class="gallery-caption-container"> 582 <div class="gallery-nav-controls"> 583 <button class="gallery-nav left" id="scrollLeftBtnClothFolding"><</button> 584 <button class="gallery-nav right" id="scrollRightBtnClothFolding">></button> 585 </div> 586 <p class="figure-caption gallery-caption"> 587 <b>Cloth Folding:</b> After training on folding onesies and shirts (in-distribution), GHOST shows meaningful progress in generalizing to novel object instances (fold-onesie-ood) and novel object categories with skill composition (fold-towel). We overlay a colormap on the image to visualize the predicted goal. 588 </p> 589 </div> 590 </div> 591 592 <!-- Results Tables --> 593 <div class="section"> 594 <h3 style="font-family: 'Roboto Mono', monospace; font-size: 18px; margin-top: 30px;">Pick and Place</h3> 595 <table style="width: 100%; border-collapse: collapse; margin: 20px 0; font-family: 'Roboto Mono', monospace; font-size: 14px;"> 596 <thead> 597 <tr style="background-color: #333; color: white;"> 598 <th style="padding: 12px; text-align: left; border: 1px solid #ddd;">Method</th> 599 <th style="padding: 12px; text-align: center; border: 1px solid #ddd; background-color: rgb(90, 155, 213);"><span style="color: white;">plate-on-table</span></th> 600 <th style="padding: 12px; text-align: center; border: 1px solid #ddd; background-color: rgb(247, 150, 70);"><span style="color: white;">mug-on-table</span></th> 601 </tr> 602 </thead> 603 <tbody> 604 <tr style="background-color: #f9f9f9;"> 605 <td style="padding: 10px; border: 1px solid #ddd;">DP</td> 606 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">80.0 ± 15.0</td> 607 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">13.3 ± 9.2</td> 608 </tr> 609 <tr> 610 <td style="padding: 10px; border: 1px solid #ddd;">MimicPlay</td> 611 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">65.0 ± 15.0</td> 612 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">28.3 ± 12.5</td> 613 </tr> 614 <tr style="background-color: #f9f9f9;"> 615 <td style="padding: 10px; border: 1px solid #ddd;">GHOST (Ours - Robot Only)</td> 616 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; text-decoration: underline;">83.3 ± 10.8</td> 617 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; text-decoration: underline;">55.0 ± 15.0</td> 618 </tr> 619 <tr> 620 <td style="padding: 10px; border: 1px solid #ddd;">GHOST (Ours)</td> 621 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; background-color: #d3d3d3; font-weight: bold;">98.3 ± 2.5</td> 622 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; background-color: #d3d3d3; font-weight: bold;">
62263.3 ± 13.3</td> 623 </tr> 624 </tbody> 625 </table> 626 627 <h3 style="font-family: 'Roboto Mono', monospace; font-size: 18px; margin-top: 40px;">Onesie Folding</h3> 628 <table style="width: 100%; border-collapse: collapse; margin: 20px 0; font-family: 'Roboto Mono', monospace; font-size: 14px;"> 629 <thead> 630 <tr style="background-color: #333; color: white;"> 631 <th style="padding: 12px; text-align: left; border: 1px solid #ddd;">Task</th> 632 <th style="padding: 12px; text-align: left; border: 1px solid #ddd;">Method</th> 633 <th style="padding: 12px; text-align: center; border: 1px solid #ddd;">1-step</th> 634 <th style="padding: 12px; text-align: center; border: 1px solid #ddd;">2-step</th> 635 <th style="padding: 12px; text-align: center; border: 1px solid #ddd;">3-step</th> 636 <th style="padding: 12px; text-align: center; border: 1px solid #ddd;">4-step</th> 637 <th style="padding: 12px; text-align: center; border: 1px solid #ddd;">Final</th> 638 </tr> 639 </thead> 640 <tbody> 641 <!-- fold-onesie (ID) --> 642 <tr> 643 <td rowspan="4" style="padding: 10px; border: 1px solid #ddd; font-weight: bold; background-color: rgba(90, 155, 213, 0.2);">fold-onesie</td> 644 <td style="padding: 10px; border: 1px solid #ddd;">DP</td> 645 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">90.0</td> 646 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; text-decoration: underline;">76.7</td> 647 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">53.3</td> 648 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">40.0</td> 649 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">10.0 ± 11.7</td> 650 </tr> 651 <tr> 652 <td style="padding: 10px; border: 1px solid #ddd;">MimicPlay</td> 653 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; text-decoration: underline;">93.3</td> 654 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; text-decoration: underline;">76.7</td> 655 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">70.0</td> 656 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">70.0</td> 657 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">46.7 ± 16.7</td> 658 </tr> 659 <tr> 660 <td style="padding: 10px; border: 1px solid #ddd;">GHOST (Ours - Robot Only)</td> 661 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; background-color: #d3d3d3; font-weight: bold;">100.0</td> 662 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; background-color: #d3d3d3; font-weight: bold;">100.0</td> 663 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; text-decoration: underline;">96.7</td> 664 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; text-decoration: underline;">86.7</td> 665 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; text-decoration: underline;">80.0 ± 15.0</td> 666 </tr> 667 <tr style="border-bottom: 3px solid #666;"> 668 <td style="padding: 10px; border: 1px solid #ddd; border-bottom: 3px solid #666;">GHOST (Ours)</td> 669 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; background-color: #d3d3d3; font-weight: bold; border-bottom: 3px solid #666;">100.0</td> 670 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; background-color: #d3d3d3; font-weight: bold; border-bottom: 3px solid #666;">100.0</td> 671 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; background-color: #d3d3d3; font-weight: bold; border-bottom: 3px solid #666;">100.0</td> 672 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; background-color: #d3d3d3; font-weight: bold; border-bottom: 3px solid #666;">90.0</td> 673 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; background-color: #d3d3d3; font-weight: bold; border-bottom: 3px solid #666;">
67383.3 ± 13.3</td> 674 </tr> 675 <!-- fold-onesie-ood (OOD) --> 676 <tr> 677 <td rowspan="4" style="padding: 10px; border: 1px solid #ddd; font-weight: bold; background-color: rgba(247, 150, 70, 0.2);">fold-onesie-ood<br><span style="font-weight: normal; font-size: 12px;">(Novel Object Instance)</span></td> 678 <td style="padding: 10px; border: 1px solid #ddd;">DP</td> 679 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; text-decoration: underline;">76.7</td> 680 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; text-decoration: underline;">60.0</td> 681 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">46.7</td> 682 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">26.7</td> 683 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">10.0 ± 10.0</td> 684 </tr> 685 <tr> 686 <td style="padding: 10px; border: 1px solid #ddd;">MimicPlay</td> 687 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">63.3</td> 688 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">36.7</td> 689 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">30.0</td> 690 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">20.0</td> 691 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">0.0 ± 0.0</td> 692 </tr> 693 <tr> 694 <td style="padding: 10px; border: 1px solid #ddd;">GHOST (Ours - Robot Only)</td> 695 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; background-color: #d3d3d3; font-weight: bold;">100.0</td> 696 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; background-color: #d3d3d3; font-weight: bold;">100.0</td> 697 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; text-decoration: underline;">73.3</td> 698 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; text-decoration: underline;">60.0</td> 699 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; text-decoration: underline;">43.3 ± 16.7</td> 700 </tr> 701 <tr style="border-bottom: 3px solid #666;"> 702 <td style="padding: 10px; border: 1px solid #ddd; border-bottom: 3px solid #666;">GHOST (Ours)</td> 703 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; background-color: #d3d3d3; font-weight: bold; border-bottom: 3px solid #666;">100.0</td> 704 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; background-color: #d3d3d3; font-weight: bold; border-bottom: 3px solid #666;">100.0</td> 705 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; background-color: #d3d3d3; font-weight: bold; border-bottom: 3px solid #666;">93.3</td> 706 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; background-color: #d3d3d3; font-weight: bold; border-bottom: 3px solid #666;">86.7</td> 707 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; background-color: #d3d3d3; font-weight: bold; border-bottom: 3px solid #666;">56.7 ± 16.7</td> 708 </tr> 709 </tbody> 710 </table> 711 712 <h3 style="font-family: 'Roboto Mono', monospace; font-size: 18px; margin-top: 40px;">Towel Folding</h3> 713 <table style="width: 100%; border-collapse: collapse; margin: 20px 0; font-family: 'Roboto Mono', monospace; font-size: 14px;"> 714 <thead> 715 <tr style="background-color: #333; color: white;"> 716 <th style="padding: 12px; text-align: left; border: 1px solid #ddd;">Method</th> 717 <th style="padding: 12px; text-align: center; border: 1px solid #ddd; background-color: rgb(247, 150, 70);"><span style="color: white;">fold-towel (Novel Category + Skill Composition)</span></th> 718 </tr> 719 </thead> 720 <tbody> 721 <tr style="background-color: #f9f9f9;"> 722 <td style="padding: 10px; border: 1px solid #ddd;">MimicPlay</td> 723 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">16.7 ± 13.3</td> 724 </tr> 725 <tr> 726 <td style="padding: 10px; border: 1px solid #ddd;">GHOST (Ours)</td> 727 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; background-color: #d3d3d3; font-weight: bold;">36.7 ± 16.7</td> 728 </tr> 729 </tbody> 730 </table> 731 732 <h3 style="font-family: 'Roboto Mono', monospace; font-size: 18px; margin-top: 40px;">Hammer Pin Results</h3> 733 <table style="width: 100%; border-collapse: collapse; margin: 20px 0; font-family: 'Roboto Mono', monospace; font-size: 14px;"> 734 <thead> 735 <tr style="background-color: #333; color: white;"> 736 <th style="padding: 12px; text-align: left; border: 1px solid #ddd;">Method</th> 737 <th style="padding: 12px; text-align: center; border: 1px solid #ddd; background-color: rgb(90, 155, 213);"><span style="color: white;">hammer-pin</span></th> 738 </tr> 739 </thead> 740 <tbody> 741 <tr style="background-color: #f9f9f9;"> 742 <td style="padding: 10px; border: 1px solid #ddd;">DP</td> 743 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">16.7 ± 13.3</td> 744 </tr> 745 <tr> 746 <td style="padding: 10px; border: 1px solid #ddd;">MimicPlay</td> 747 <td style="padding: 10px; text-align: center; border: 1px solid #ddd;">33.3 ± 16.7</td> 748 </tr> 749 <tr style="background-color: #f9f9f9;"> 750 <td style="padding: 10px; border: 1px solid #ddd;">GHOST (Ours - Robot Only)</td> 751 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; text-decoration: underline;">50.0 ± 16.7</td> 752 </tr> 753 <tr> 754 <td style="padding: 10px; border: 1px solid #ddd;">GHOST (Ours)</td> 755 <td style="padding: 10px; text-align: center; border: 1px solid #ddd; background-color: #d3d3d3; font-weight: bold;">70.0 ± 16.7</td> 756 </tr> 757 </tbody> 758 </table> 759 </div> 760 761 <!-- Results Summary --> 762 <div class="section" style="margin-top: 40px;"> 763 <h3 style="font-family: 'Roboto Mono', monospace; font-size: 18px; margin-bottom: 15px;">Key Findings</h3> 764 765 <p style="font-family: 'Roboto Mono', monospace; font-size: 14px; line-height: 1.6; margin-bottom: 15px;"> 766 <b>Do hierarchical policies improve in-distribution performance even without human data?</b><br> 767 Yes. For <span style="color: rgb(90, 155, 213); font-weight: bold;">plate-on-table</span>, nearly all methods saturate in performance, as the task is simple with sufficient training data. However, for long-horizon complex tasks, we see significant increases: on <span style="color: rgb(90, 155, 213); font-weight: bold;">fold-onesie</span>, performance increases from 10% (DP) to 80% (GHOST - Robot Only) final success, showing large benef
767its from hierarchical decomposition. Similarly, on <span style="color: rgb(90, 155, 213); font-weight: bold;">hammer-pin</span>, which requires precise grasping of the hammer tool and striking the correct pin, performance significantly improves from DP (16.7%) to GHOST - Robot Only (50%). 768 </p> 769 770 <p style="font-family: 'Roboto Mono', monospace; font-size: 14px; line-height: 1.6; margin-bottom: 15px;"> 771 <b>Do human demonstrations enable transferring learned skills to novel object instances, categories, and contexts?</b><br> 772 Yes. Human demonstrations unlock meaningful OOD transfer of learned skills to novel objects and skill compositions. GHOST achieves 63.3% success on <span style="color: rgb(247, 150, 70); font-weight: bold;">mug-on-table</span>, a task featuring a combination of objects unseen in robot demonstrations. On <span style="color: rgb(247, 150, 70); font-weight: bold;">fold-onesie-ood</span>, GHOST achieves 56.7% final success vs 43.3% for GHOST - Robot Only and 0% for MimicPlay. On the hardest task of generalizing a policy to a novel object category and skill combination (<span style="color: rgb(247, 150, 70); font-weight: bold;">fold-towel</span>), GHOST achieves 36.7% success as compared to 16.7% with the MimicPlay baseline. 773 </p> 774 </div> 775 776 <!-- BibTeX --> 777 <div class="bibtex-code" id="bibtex"> 778 <div class="bibtex-title">BibTeX</div> 779 <pre><code>@inproceedings{krishna2026ghost, 780 title = {GHOST: Hierarchical Sub-Goal Policies for Generalizing Robot Manipulation}, 781 author = {Krishna, Sriram and Eisner, Ben and Zhan, Haotian and Yuan, Ying and 782 Zhen, Haoyu and Gan, Chuang and Tulsiani, Shubham and Held, David}, 783 booktitle = {Robotics: Science and Systems (RSS)}, 784 year = {2026} 785}</code></pre> 786 </div> 787 788 </div> <!-- End of main-content div --> 789 790 <div class="footer"> 791 Website template modified from <a href="https://videomimic.net/" target="_blank">VideoMimic</a>. 792 </div> 793 794 <!-- JavaScript for Video Gallery Navigation --> 795
795<script> 796 document.addEventListener('DOMContentLoaded', function() { 797 const galleries = [ 798 { 799 galleryInnerId: 'videoGalleryPickPlace', 800 scrollLeftBtnId: 'scrollLeftBtnPickPlace', 801 scrollRightBtnId: 'scrollRightBtnPickPlace' 802 }, 803 { 804 galleryInnerId: 'videoGalleryHammerPin', 805 scrollLeftBtnId: 'scrollLeftBtnHammerPin', 806 scrollRightBtnId: 'scrollRightBtnHammerPin' 807 }, 808 { 809 galleryInnerId: 'videoGalleryClothFolding', 810 scrollLeftBtnId: 'scrollLeftBtnClothFolding', 811 scrollRightBtnId: 'scrollRightBtnClothFolding' 812 } 813 ]; 814 815 galleries.forEach(galleryConfig => { 816 const galleryContainer = document.getElementById(galleryConfig.galleryInnerId); 817 const scrollLeftBtn = document.getElementById(galleryConfig.scrollLeftBtnId); 818 const scrollRightBtn = document.getElementById(galleryConfig.scrollRightBtnId); 819 820 if (galleryContainer && scrollLeftBtn && scrollRightBtn) { 821 const scrollAmount = 365; // Width of placeholder + gap 822 823 scrollLeftBtn.addEventListener('click', () => { 824 galleryContainer.parentElement.scrollBy({ left: -scrollAmount, behavior: 'smooth' }); 825 }); 826 827 scrollRightBtn.addEventListener('click', () => { 828 galleryContainer.parentElement.scrollBy({ left: scrollAmount, behavior: 'smooth' }); 829 }); 830 } 831 }); 832 }); 833 </script>
833 834 835 <!-- Three.js for 3D Viewers --> 836
836<script src="https://cdnjs.cloudflare.com/ajax/libs/three.js/r128/three.min.js"></script>
836 837
837<script src="https://cdn.jsdelivr.net/npm/[email protected]/examples/js/controls/OrbitControls.js"></script>
837 838
838<script src="https://cdn.jsdelivr.net/npm/[email protected]/examples/js/loaders/PLYLoader.js"></script>
838 839 840 <!-- 3D Viewer Script --> 841
841<script> 842 function createViewer(canvasId, meshFiles, cameraPosition) { 843 const canvas = document.getElementById(canvasId); 844 const width = canvas.clientWidth; 845 const height = canvas.clientHeight; 846 847 // Scene setup 848 const scene = new THREE.Scene(); 849 scene.background = null; 850 851 // Camera 852 const camera = new THREE.PerspectiveCamera(45, width / height, 0.1, 1000); 853 camera.position.set(cameraPosition.x, cameraPosition.y, cameraPosition.z); 854 855 // Renderer 856 const renderer = new THREE.WebGLRenderer({ 857 canvas: canvas, 858 alpha: true, 859 antialias: true 860 }); 861 renderer.setSize(width, height); 862 renderer.setPixelRatio(window.devicePixelRatio); 863 864 // Lights 865 const ambientLight = new THREE.AmbientLight(0xffffff, 0.6); 866 scene.add(ambientLight); 867 868 const directionalLight1 = new THREE.DirectionalLight(0xffffff, 0.5); 869 directionalLight1.position.set(5, 5, 5); 870 scene.add(directionalLight1); 871 872 const directionalLight2 = new THREE.DirectionalLight(0xffffff, 0.3); 873 directionalLight2.position.set(-5, -5, -5); 874 scene.add(directionalLight2); 875 876 // Controls 877 const controls = new THREE.OrbitControls(camera, renderer.domElement); 878 controls.enableDamping = true; 879 controls.dampingFactor = 0.05; 880 controls.screenSpacePanning = false; 881 controls.minDistance = 0.1; 882 controls.maxDistance = 10; 883 884 // Load PLY files 885 const loader = new THREE.PLYLoader(); 886 const meshes = []; 887 let loadedCount = 0; 888 889 meshFiles.forEach((file, index) => { 890 loader.load(file.path, function(geometry) { 891 geometry.computeVertexNormals(); 892 893 const material = new THREE.MeshPhongMaterial({ 894 vertexColors: true, 895 specular: 0x111111, 896 shininess: 30 897 }); 898 899 const mesh = new THREE.Mesh(geometry, material); 900 meshes.push(mesh); 901 scene.add(mesh); 902 903 loadedCount++; 904 905 // After all meshes are loaded, center them as a group 906 if (loadedCount === meshFiles.length) { 907 const box = new THREE.Box3(); 908 meshes.forEach(m => box.expandByObject(m)); 909 const center = box.getCenter(new THREE.Vector3()); 910 meshes.forEach(m => m.position.sub(center)); 911 } 912 }, 913 // Progress callback 914 function(xhr) { 915 console.log((xhr.loaded / xhr.total * 100) + '% loaded: ' + file.path); 916 }, 917 // Error callback 918 function(error) { 919 console.error('Error loading ' + file.path + ':', error); 920 }); 921 }); 922 923 // Animation loop 924 function animate() { 925 requestAnimationFrame(animate); 926 controls.update(); 927 renderer.render(scene, camera); 928 } 929 animate(); 930 931 // Handle window resize 932 window.addEventListener('resize', () => { 933 const newWidth = canvas.clientWidth; 934 const newHeight = canvas.clientHeight; 935 camera.aspect = newWidth / newHeight; 936 camera.updateProjectionMatrix(); 937 renderer.setSize(newWidth, newHeight); 938 }); 939 } 940 941 // Initialize viewers when page loads 942 document.addEventListener('DOMContentLoaded', function() { 943 // ALOHA Gripper viewer 944 createViewer('aloha-viewer', [ 945 { path: 'meshes/aloha_gripper.ply' }, 946 { path: 'meshes/aloha_spheres.ply' } 947 ], { x: 0.15, y: 0.1, z: 0.15 }); 948 949 // Hand viewer 950 createViewer('hand-viewer', [ 951 { path: 'meshes/hand_gripper.ply' }, 952 { path: 'meshes/hand_spheres.ply' } 953 ], { x: 0.15, y: 0.1, z: 0.15 }); 954 }); 955 </script>
955 956 957</body> 958</html>
Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.