1<!DOCTYPE html> 2<html> 3<head> 4 <meta charset="utf-8"> 5 <meta name="description" 6 content="UniFunc3D: Unified Active Spatial-Temporal Grounding for 3D Functionality Segmentation"> 7 <meta name="keywords" content="UniFunc3D, 3D functionality segmentation, spatial-temporal grounding, MLLM"> 8 <meta name="viewport" content="width=device-width, initial-scale=1"> 9 <title>UniFunc3D: Unified Active Spatial-Temporal Grounding for 3D Functionality Segmentation</title> 10 11 <link href="https://fonts.googleapis.com/css?family=Google+Sans|Noto+Sans|Castoro" 12 rel="stylesheet"> 13 14 <link rel="stylesheet" href="./static/css/bulma.min.css"> 15 <link rel="stylesheet" href="./static/css/fontawesome.all.min.css"> 16 <link rel="stylesheet" 17 href="https://cdn.jsdelivr.net/gh/jpswalsh/academicons@1/css/academicons.min.css"> 18 <link rel="stylesheet" href="./static/css/index.css"> 19 20
20<script src="https://ajax.googleapis.com/ajax/libs/jquery/3.5.1/jquery.min.js"></script>
20 21
21<script defer src="./static/js/fontawesome.all.min.js"></script>
21 22</head> 23<body> 24 25<nav class="navbar" role="navigation" aria-label="main navigation"> 26 <div class="navbar-brand"> 27 <a role="button" class="navbar-burger" aria-label="menu" aria-expanded="false"> 28 <span aria-hidden="true"></span> 29 <span aria-hidden="true"></span> 30 <span aria-hidden="true"></span> 31 </a> 32 </div> 33 <div class="navbar-menu"> 34 <div class="navbar-start" style="flex-grow: 1; justify-content: center;"> 35 <a class="navbar-item" href="https://jiaying.link"> 36 <span class="icon"> 37 <i class="fas fa-home"></i> 38 </span> 39 </a> 40 41 <div class="navbar-item has-dropdown is-hoverable"> 42 <a class="navbar-link"> 43 More Research 44 </a> 45 <div class="navbar-dropdown"> 46 <a class="navbar-item" href="https://jiaying.link/cvpr2020-pgd/"> 47 PMD 48 </a> 49 <a class="navbar-item" href="https://jiaying.link/cvpr2021-gsd/"> 50 GSD 51 </a> 52 <a class="navbar-item" href="https://jiaying.link/neurips2022-gsds/"> 53 GSD-S 54 </a> 55 <a class="navbar-item" href="https://jiaying.link/cvpr2023-vmd/"> 56 VMD 57 </a> 58 </div> 59 </div> 60 </div> 61 62 </div> 63</nav> 64 65 66<section class="hero"> 67 <div class="hero-body"> 68 <div class="container is-max-desktop"> 69 <div class="columns is-centered"> 70 <div class="column has-text-centered"> 71 <h1 class="title is-1 publication-title">UniFunc3D: Unified Active Spatial-Temporal Grounding for 3D Affordance Segmentation</h1> 72 <div class="is-size-5 publication-authors"> 73 <span class="author-block"> 74 <a href="https://jiaying.link">Jiaying Lin</a>, 75 </span> 76 <span class="author-block"> 77 <a href="https://www.danxurgb.net/">Dan Xu</a> 78 </span> 79 </div> 80 81 <div class="is-size-5 publication-authors"> 82 <span class="author-block"> 83 The Hong Kong University of Science and Technology (HKUST) 84 </span> 85 </div> 86 87 <div class="is-size-5 publication-authors" style="margin-top: 0.5rem;"> 88 <span class="author-block">NeurIPS 2026</span> 89 </div> 90 91 <div class="column has-text-centered"> 92 <div class="publication-links"> 93 <!-- PDF Link. --> 94 <!-- <span class="link-block"> 95 <a href="" 96 class="external-link button is-normal is-rounded is-dark"> 97 <span class="icon"> 98 <i class="fas fa-file-pdf"></i> 99 </span> 100 <span>Paper</span> 101 </a> 102 </span> --> 103 <!-- arXiv Link. --> 104 <span class="link-block"> 105 <a href="https://arxiv.org/abs/2603.23478" 106 class="external-link button is-normal is-rounded is-dark"> 107 <span class="icon"> 108 <i class="ai ai-arxiv"></i> 109 </span> 110 <span>arXiv</span> 111 </a> 112 </span> 113 <!-- Code Link. --> 114 <span class="link-block"> 115 <a href="" 116 class="external-link button is-normal is-rounded is-dark"> 117 <span class="icon"> 118 <i class="fab fa-github"></i> 119 </span> 120 <span>Code (Coming soon)</span> 121 </a> 122 </span> 123 </div> 124 </div> 125 </div> 126 </div> 127 </div> 128 </div> 129</section> 130 131 132<section class="section"> 133 <div class="container is-max-desktop"> 134 <div class="content has-text-centered"> 135 <video id="teaser-video" autoplay muted loop playsinline controls width="100%"> 136 <source src="./static/video/unifunc_demo_mp4.mp4" type="video/mp4"> 137 </video> 138 </div> 139 </div> 140</section> 141 142 143<section class="section"> 144 <div class="container is-max-desktop"> 145 <div class="content has-text-centered"> 146 <img alt="teaser" class="center" src="./static/images/teaser_plug_device.jpg"> 147 </div> 148 <div class="content has-text-justified" style="margin-top: 1rem;"> 149 <b>Overview of UniFunc3D compared to existing fragmented pipelines.</b> 150 (Top) Prior methods like Fun3DU rely on a visually blind text-only LLM for initial task parsing. 151 Coupled with single-scale passive heuristic frame selection, this fragmented approach suffers from three critical failure modes: 152 semantic misinterpretations, spatial-temporal context inconsistencies, and imperceptible small targets. 153 (Bottom) Our proposed UniFunc3D addresses these limitations by utilizing a unified Multimodal Large Language Model (MLLM) as an active observer, 154 consolidating semantic, temporal, and spatial reasoning into a single forward pass. 155 </div> 156 </div> 157</section> 158 159 160<section class="section"> 161 <div class="container is-max-desktop"> 162 <!-- Abstract. --> 163 <div class="columns is-centered has-text-centered"> 164 <div class="column is-four-fifths"> 165 <h2 class="title is-2">Abstract</h2> 166 <div class="content has-text-justified"> 167 Functionality segmentation in 3D scenes requires an agent to ground implicit natural-language instructions into precise masks of fine-grained interactive elements. 168 Existing methods rely on fragmented pipelines that suffer from visual blindness during initial task parsing. 169 We observe that these methods are limited by single-scale, passive and heuristic frame selection. 170 We present UniFunc3D, a unified and training-free framework that treats the multimodal large language model as an active observer. 171 By consolidating semantic, temporal, and spatial reasoning into a single forward pass, UniFunc3D performs joint reasoning to ground task decomposition in direct visual evidence. 172 Our approach introduces active spatial-temporal grounding with a coarse-to-fine strategy. 173 This allows the model to select correct video frames adaptively and focus on high-detail interactive parts while preserving the global context necessary for disambiguation. 174 On SceneFun3D, UniFunc3D achieves state-of-the-art performance, surpassing both training-free and training-based methods by a large margin with a relative <b>59.9% mIoU improvement</b>, without any task-specific training. 175 </div> 176 </div> 177 </div> 178 <!--/ Abstract. --> 179 </div> 180</section> 181 182 183<section class="section"> 184 <div class="container is-max-desktop"> 185 <div class="columns is-centered has-text-centered"> 186 <div class="column is-four-fifths"> 187 <h2 class="title is-2">Method</h2> 188 </div> 189 </div> 190 <div class="hero-body"> 191 <div class="content has-text-justified"> 192 UniFunc3D employs a single unified MLLM with <b>active spatial-temporal grounding</b> across two stages: 193 <br><br> 194 <b>(1) Active spatial-temporal grounding with joint functional object identification.</b> 195 The <i>coarse stage</i> (Round 1) actively surveys low-resolution video frames across multiple sampling iterations and selects the most informative candidate via visual verification. 196 The <i>fine stage</i> (Round 2) processes a dense temporal window at native high resolution, delivering zoom-in capability while preserving global scene context for precise localization. 197 <br><br> 198 <b>(2) Visual mask generation and verification.</b> 199 Predicted affordance points prompt SAM3 for segmentation. Each mask is then verified by the same MLLM through visual overlay inspection before 3D lifting. Verified masks undergo multi-view agreement and 3D lifting to produce the final point cloud mask. 200 </div> 201 202 <div class="content has-text-centered"> 203 <img alt="pipeline" class="center" src="./static/images/pipeline.jpg"> 204 </div>
205 <figcaption> 206 <b>Method overview.</b> The coarse stage actively surveys low-resolution frames with multiple sampling iterations; the fine stage processes a dense temporal window at native high resolution. 207 Visual mask generation and verification uses SAM3 with MLLM-based mask verification, then multi-view 3D lifting to obtain the final 3D masks. 208 </figcaption> 209 </div> 210 </div> 211</section> 212 213 214<section class="section"> 215 <div class="container is-max-desktop"> 216 <div class="columns is-centered has-text-centered"> 217 <div class="column is-four-fifths"> 218 <h2 class="title is-2">Results</h2> 219 </div> 220 </div> 221 <div class="hero-body"> 222 <h3 class="title is-3">State-of-the-art on SceneFun3D</h3> 223 <div class="content has-text-justified"> 224 UniFunc3D-30B achieves the best performance across <i>all</i> methods on all reported metrics on both splits of SceneFun3D. 225 Remarkably, our training-free approach outperforms both training-free and training-based methods, including those using significantly larger models (72B). 226 <ul> 227 <li>Compared to Fun3DU (training-free): <b>+84.9% relative AP<sub>50</sub></b> and <b>+59.9% relative mIoU</b> on split0.</li> 228 <li>Compared to AffordBot-72B (training-based, fine-tuned for 1000 epochs): <b>+49.4% AP<sub>50</sub></b> and <b>+68.5% mIoU</b>.</li> 229 <li>UniFunc3D also achieves a <b>3.2× speedup</b> over Fun3DU (~26 min vs. ~82 min per scene).</li> 230 </ul> 231 </div> 232 233 <div class="content has-text-centered" style="overflow-x: auto;"> 234 <table class="table is-bordered is-striped is-narrow is-hoverable is-fullwidth"> 235 <thead> 236 <tr> 237 <th rowspan="2">Method</th> 238 <th colspan="5" style="text-align:center;">Split0 (30 scenes, val)</th> 239 <th colspan="5" style="text-align:center;">Split1 (200 scenes, train)</th> 240 </tr> 241 <tr> 242 <th>AP<sub>50</sub></th><th>AP<sub>25</sub></th><th>AR<sub>50</sub></th><th>AR<sub>25</sub></th><th>mIoU</th> 243 <th>AP<sub>50</sub></th><th>AP<sub>25</sub></th><th>AR<sub>50</sub></th><th>AR<sub>25</sub></th><th>mIoU</th> 244 </tr> 245 </thead> 246 <tbody> 247 <tr><td colspan="11"><i>Training-based methods:</i></td></tr> 248 <tr> 249 <td>TASA-72B</td> 250 <td>26.9</td><td>28.6</td><td>â</td><td>â</td><td>19.7</td> 251 <td colspan="5" style="text-align:center;"><i>trained on split1</i></td> 252 </tr> 253 <tr> 254 <td>AffordBot-72B</td> 255 <td>20.91</td><td>24.76</td><td>18.99</td><td>22.84</td><td>14.42</td> 256 <td colspan="5" style="text-align:center;"><i>trained on split1</i></td> 257 </tr> 258 <tr><td colspan="11"><i>Training-free methods:</i></td></tr> 259 <tr> 260 <td>Fun3DU-9B</td> 261 <td>16.9</td><td>33.3</td><td>38.2</td><td>46.7</td><td>15.2</td> 262 <td>12.6</td><td>23.1</td><td>32.9</td><td>40.5</td><td>11.5</td> 263 </tr> 264 <tr> 265 <td>UniFunc3D-8B (Ours)</td> 266 <td>23.82</td><td>44.04</td><td>46.07</td><td>55.51</td><td>20.92</td> 267 <td>16.24</td><td>29.02</td><td>38.91</td><td>48.15</td><td>14.23</td> 268 </tr> 269 <tr style="font-weight:bold;"> 270 <td>UniFunc3D-30B (Ours)</td> 271 <td>31.24</td><td>51.01</td><td>46.97</td><td>58.88</td><td>24.30</td> 272 <td>21.32</td><td>35.76</td><td>40.03</td><td>51.00</td><td>17.09</td> 273 </tr> 274 </tbody> 275 </table> 276 </div> 277 </div> 278 279 <!-- Qualitative Results --> 280 <div class="hero-body"> 281 <h3 class="title is-3">Qualitative Results</h3> 282 <div class="content has-text-justified"> 283 Qualitative comparison across five representative queries. Our method clearly outperforms prior methods in spatial disambiguation and handling small interactive objects. 284 For example, for <i>"Open the top left drawer of the cabinet with the beauty products on top"</i>, 285 our method finds the correct <b>top-left knob</b>, while AffordBot finds the wrong top-right knob and Fun3DU mistakenly segments the drawer face. 286 </div> 287 288 <!-- Column header --> 289 <div class="columns is-centered is-vcentered" style="margin-bottom:0; font-weight:bold;"> 290 <div class="column is-2 has-text-centered"></div> 291 <div class="column has-text-centered"><p class="is-size-7">AffordBot</p></div> 292 <div class="column has-text-centered"><p class="is-size-7">Fun3DU</p></div> 293 <div class="column has-text-centered"><p class="is-size-7"><b>Ours</b></p></div> 294 <div class="column has-text-centered"><p class="is-size-7">GT</p></div> 295 </div> 296 297 <!-- q1 --> 298 <div class="columns is-centered is-vcentered" style="margin-bottom:0.4rem;"> 299 <div class="column is-2 has-text-centered"> 300 <p class="is-size-7"><i>"Open the top left drawer of the cabinet with the beauty products on top"</i></p> 301 </div> 302 <div class="column"><img src="./static/images/qual/q1_affordbot.jpg" alt="AffordBot"></div> 303 <div class="column"><img src="./static/images/qual/q1_fun3du.jpg" alt="Fun3DU"></div> 304 <div class="column"><img src="./static/images/qual/q1_ours.jpg" alt="Ours"></div> 305 <div class="column"><img src="./static/images/qual/q1_gt.jpg" alt="GT"></div> 306 </div> 307 308 <!-- q2 --> 309 <div class="columns is-centered is-vcentered" style="margin-bottom:0.4rem;"> 310 <div class="column is-2 has-text-centered"> 311 <p class="is-size-7"><i>"Turn on the ceiling light"</i></p> 312 </div> 313 <div class="column"><img src="./static/images/qual/q2_affordbot.jpg" alt="AffordBot"></div> 314 <div class="column"><img src="./static/images/qual/q2_fun3du.jpg" alt="Fun3DU"></div> 315 <div class="column"><img src="./static/images/qual/q2_ours.jpg" alt="Ours"></div> 316 <div class="column"><img src="./static/images/qual/q2_gt.jpg" alt="GT"></div> 317 </div> 318 319 <!-- q3 --> 320 <div class="columns is-centered is-vcentered" style="margin-bottom:0.4rem;"> 321 <div class="column is-2 has-text-centered"> 322 <p class="is-size-7"><i>"Control the water flow in the bathtub using the drain control dial"</i></p> 323 </div> 324 <div class="column"><img src="./static/images/qual/q3_affordbot.jpg" alt="AffordBot"></div> 325 <div class="column"><img src="./static/images/qual/q3_fun3du.jpg" alt="Fun3DU"></div> 326 <div class="column"><img src="./static/images/qual/q3_ours.jpg" alt="Ours"></div> 327 <div class="column"><img src="./static/images/qual/q3_gt.jpg" alt="GT"></div> 328 </div> 329 330 <!-- q4 --> 331 <div class="columns is-centered is-vcentered" style="margin-bottom:0.4rem;"> 332 <div class="column is-2 has-text-centered"> 333 <p class="is-size-7"><i>"Select a washing program"</i></p> 334 </div> 335 <div class="column"><img src="./static/images/qual/q4_affordbot.jpg" alt="AffordBot"></div> 336 <div class="column"><img src="./static/images/qual/q4_fun3du.jpg" alt="Fun3DU"></div> 337 <div class="column"><img src="./static/images/qual/q4_ours.jpg" alt="Ours"></div> 338 <div class="column"><img src="./static/images/qual/q4_gt.jpg" alt="GT"></div> 339 </div> 340 341 <!-- q5 --> 342 <div class="columns is-centered is-vcentered"> 343 <div class="column is-2 has-text-centered"> 344 <p class="is-size-7"><i>"Flush the toilet"</i></p> 345 </div> 346 <div class="column"><img src="./static/images/qual/q5_affordbot.jpg" alt="AffordBot"></div> 347 <div class="column"><img src="./static/images/qual/q5_fun3du.jpg" alt="Fun3DU"></div> 348 <div class="column"><img src="./static/images/qual/q5_ours.jpg" alt="Ours"></div> 349 <div class="column"><img src="./static/images/qual/q5_gt.jpg" alt="GT"></div> 350 </div> 351 352 </div> 353 <!-- / Qualitative Results --> 354 355 </div> 356</section> 357 358 359<section class="section" id="BibTeX">
360 <div class="container is-max-desktop content"> 361 <h2 class="title">BibTeX</h2> 362 <pre><code>@InProceedings{Lin_UniFunc3D, 363 author = {Lin, Jiaying and Xu, Dan}, 364 title = {UniFunc3D: Unified Active Spatial-Temporal Grounding for 3D Functionality Segmentation}, 365 booktitle = {arXiv}, 366 year = {2026}, 367}</code></pre> 368 </div> 369</section> 370 371 372<footer class="footer"> 373 <div class="container"> 374 <div class="content has-text-centered"> 375 <a class="icon-link" href="https://jiaying.link"> 376 <i class="fas fa-home"></i> 377 </a> 378 </div> 379 <div class="columns is-centered"> 380 <div class="column is-8"> 381 <div class="content"> 382 <p> 383 This website is licensed under a <a rel="license" 384 href="http://creativecommons.org/licenses/by-sa/4.0/">Creative 385 Commons Attribution-ShareAlike 4.0 International License</a>. 386 </p> 387 <p> 388 We borrow this website template from <a 389 href="https://github.com/nerfies/nerfies.github.io">Nerfies</a>. 390 </p> 391 </div> 392 </div> 393 </div> 394 </div> 395</footer> 396 397
398<script type="text/javascript"> 399 var sc_project=12773779; 400 var sc_invisible=1; 401 var sc_security="f1222376"; 402 </script>
402 403
403<script type="text/javascript" 404 src="https://www.statcounter.com/counter/counter.js" 405 async></script>
405 406 <noscript><div class="statcounter"><a title="Web Analytics" 407 href="https://statcounter.com/" target="_blank"><img 408 class="statcounter" 409 src="https://c.statcounter.com/12773779/0/f1222376/1/" 410 alt="Web Analytics" 411 referrerPolicy="no-referrer-when-downgrade"></a></div></noscript> 412<!-- End of Statcounter Code --> 413</body> 414</html>
Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.