PageSourceSearch

https://jiaying.link/unifunc3d/

html jiaying.link collected 2026-09-28 06:49:11 UTC 18,841 bytes, 414 lines download raw bytes

1<!DOCTYPE html>
2<html>
3<head>
4  <meta charset="utf-8">
5  <meta name="description"
6        content="UniFunc3D: Unified Active Spatial-Temporal Grounding for 3D Functionality Segmentation">
7  <meta name="keywords" content="UniFunc3D, 3D functionality segmentation, spatial-temporal grounding, MLLM">
8  <meta name="viewport" content="width=device-width, initial-scale=1">
9  <title>UniFunc3D: Unified Active Spatial-Temporal Grounding for 3D Functionality Segmentation</title>
10
11  <link href="https://fonts.googleapis.com/css?family=Google+Sans|Noto+Sans|Castoro"
12        rel="stylesheet">
13
14  <link rel="stylesheet" href="./static/css/bulma.min.css">
15  <link rel="stylesheet" href="./static/css/fontawesome.all.min.css">
16  <link rel="stylesheet"
17        href="https://cdn.jsdelivr.net/gh/jpswalsh/academicons@1/css/academicons.min.css">
18  <link rel="stylesheet" href="./static/css/index.css">
19
20  
20<script src="https://ajax.googleapis.com/ajax/libs/jquery/3.5.1/jquery.min.js"></script>
20
21  
21<script defer src="./static/js/fontawesome.all.min.js"></script>
21
22</head>
23<body>
24
25<nav class="navbar" role="navigation" aria-label="main navigation">
26  <div class="navbar-brand">
27    <a role="button" class="navbar-burger" aria-label="menu" aria-expanded="false">
28      <span aria-hidden="true"></span>
29      <span aria-hidden="true"></span>
30      <span aria-hidden="true"></span>
31    </a>
32  </div>
33  <div class="navbar-menu">
34    <div class="navbar-start" style="flex-grow: 1; justify-content: center;">
35      <a class="navbar-item" href="https://jiaying.link">
36      <span class="icon">
37          <i class="fas fa-home"></i>
38      </span>
39      </a>
40
41      <div class="navbar-item has-dropdown is-hoverable">
42        <a class="navbar-link">
43          More Research
44        </a>
45        <div class="navbar-dropdown">
46          <a class="navbar-item" href="https://jiaying.link/cvpr2020-pgd/">
47            PMD
48          </a>
49          <a class="navbar-item" href="https://jiaying.link/cvpr2021-gsd/">
50            GSD
51          </a>
52          <a class="navbar-item" href="https://jiaying.link/neurips2022-gsds/">
53            GSD-S
54          </a>
55          <a class="navbar-item" href="https://jiaying.link/cvpr2023-vmd/">
56            VMD
57          </a>
58        </div>
59      </div>
60    </div>
61
62  </div>
63</nav>
64
65
66<section class="hero">
67  <div class="hero-body">
68    <div class="container is-max-desktop">
69      <div class="columns is-centered">
70        <div class="column has-text-centered">
71          <h1 class="title is-1 publication-title">UniFunc3D: Unified Active Spatial-Temporal Grounding for 3D Affordance Segmentation</h1>
72          <div class="is-size-5 publication-authors">
73            <span class="author-block">
74              <a href="https://jiaying.link">Jiaying Lin</a>,
75            </span>
76            <span class="author-block">
77              <a href="https://www.danxurgb.net/">Dan Xu</a>
78            </span>
79          </div>
80
81          <div class="is-size-5 publication-authors">
82            <span class="author-block">
83              The Hong Kong University of Science and Technology (HKUST)
84            </span>
85          </div>
86
87          <div class="is-size-5 publication-authors" style="margin-top: 0.5rem;">
88            <span class="author-block">NeurIPS 2026</span>
89          </div>
90
91          <div class="column has-text-centered">
92            <div class="publication-links">
93              <!-- PDF Link. -->
94              <!-- <span class="link-block">
95                <a href=""
96                   class="external-link button is-normal is-rounded is-dark">
97                  <span class="icon">
98                      <i class="fas fa-file-pdf"></i>
99                  </span>
100                  <span>Paper</span>
101                </a>
102              </span> -->
103              <!-- arXiv Link. -->
104              <span class="link-block">
105                <a href="https://arxiv.org/abs/2603.23478"
106                   class="external-link button is-normal is-rounded is-dark">
107                  <span class="icon">
108                      <i class="ai ai-arxiv"></i>
109                  </span>
110                  <span>arXiv</span>
111                </a>
112              </span>
113              <!-- Code Link. -->
114              <span class="link-block">
115                <a href=""
116                   class="external-link button is-normal is-rounded is-dark">
117                  <span class="icon">
118                      <i class="fab fa-github"></i>
119                  </span>
120                  <span>Code (Coming soon)</span>
121                  </a>
122              </span>
123            </div>
124          </div>
125        </div>
126      </div>
127    </div>
128  </div>
129</section>
130
131
132<section class="section">
133  <div class="container is-max-desktop">
134    <div class="content has-text-centered">
135      <video id="teaser-video" autoplay muted loop playsinline controls width="100%">
136        <source src="./static/video/unifunc_demo_mp4.mp4" type="video/mp4">
137      </video>
138    </div>
139  </div>
140</section>
141
142
143<section class="section">
144  <div class="container is-max-desktop">
145    <div class="content has-text-centered">
146      <img alt="teaser" class="center" src="./static/images/teaser_plug_device.jpg">
147    </div>
148    <div class="content has-text-justified" style="margin-top: 1rem;">
149      <b>Overview of UniFunc3D compared to existing fragmented pipelines.</b>
150      (Top) Prior methods like Fun3DU rely on a visually blind text-only LLM for initial task parsing.
151      Coupled with single-scale passive heuristic frame selection, this fragmented approach suffers from three critical failure modes:
152      semantic misinterpretations, spatial-temporal context inconsistencies, and imperceptible small targets.
153      (Bottom) Our proposed UniFunc3D addresses these limitations by utilizing a unified Multimodal Large Language Model (MLLM) as an active observer,
154      consolidating semantic, temporal, and spatial reasoning into a single forward pass.
155    </div>
156  </div>
157</section>
158
159
160<section class="section">
161  <div class="container is-max-desktop">
162    <!-- Abstract. -->
163    <div class="columns is-centered has-text-centered">
164      <div class="column is-four-fifths">
165        <h2 class="title is-2">Abstract</h2>
166        <div class="content has-text-justified">
167          Functionality segmentation in 3D scenes requires an agent to ground implicit natural-language instructions into precise masks of fine-grained interactive elements.
168          Existing methods rely on fragmented pipelines that suffer from visual blindness during initial task parsing.
169          We observe that these methods are limited by single-scale, passive and heuristic frame selection.
170          We present UniFunc3D, a unified and training-free framework that treats the multimodal large language model as an active observer.
171          By consolidating semantic, temporal, and spatial reasoning into a single forward pass, UniFunc3D performs joint reasoning to ground task decomposition in direct visual evidence.
172          Our approach introduces active spatial-temporal grounding with a coarse-to-fine strategy.
173          This allows the model to select correct video frames adaptively and focus on high-detail interactive parts while preserving the global context necessary for disambiguation.
174          On SceneFun3D, UniFunc3D achieves state-of-the-art performance, surpassing both training-free and training-based methods by a large margin with a relative <b>59.9% mIoU improvement</b>, without any task-specific training.
175        </div>
176      </div>
177    </div>
178    <!--/ Abstract. -->
179  </div>
180</section>
181
182
183<section class="section">
184  <div class="container is-max-desktop">
185    <div class="columns is-centered has-text-centered">
186      <div class="column is-four-fifths">
187        <h2 class="title is-2">Method</h2>
188      </div>
189    </div>
190    <div class="hero-body">
191      <div class="content has-text-justified">
192        UniFunc3D employs a single unified MLLM with <b>active spatial-temporal grounding</b> across two stages:
193        <br><br>
194        <b>(1) Active spatial-temporal grounding with joint functional object identification.</b>
195        The <i>coarse stage</i> (Round 1) actively surveys low-resolution video frames across multiple sampling iterations and selects the most informative candidate via visual verification.
196        The <i>fine stage</i> (Round 2) processes a dense temporal window at native high resolution, delivering zoom-in capability while preserving global scene context for precise localization.
197        <br><br>
198        <b>(2) Visual mask generation and verification.</b>
199        Predicted affordance points prompt SAM3 for segmentation. Each mask is then verified by the same MLLM through visual overlay inspection before 3D lifting. Verified masks undergo multi-view agreement and 3D lifting to produce the final point cloud mask.
200      </div>
201
202      <div class="content has-text-centered">
203        <img alt="pipeline" class="center" src="./static/images/pipeline.jpg">
204      </div>
205      <figcaption>
206        <b>Method overview.</b> The coarse stage actively surveys low-resolution frames with multiple sampling iterations; the fine stage processes a dense temporal window at native high resolution.
207        Visual mask generation and verification uses SAM3 with MLLM-based mask verification, then multi-view 3D lifting to obtain the final 3D masks.
208      </figcaption>
209    </div>
210  </div>
211</section>
212
213
214<section class="section">
215  <div class="container is-max-desktop">
216    <div class="columns is-centered has-text-centered">
217      <div class="column is-four-fifths">
218        <h2 class="title is-2">Results</h2>
219      </div>
220    </div>
221    <div class="hero-body">
222      <h3 class="title is-3">State-of-the-art on SceneFun3D</h3>
223      <div class="content has-text-justified">
224        UniFunc3D-30B achieves the best performance across <i>all</i> methods on all reported metrics on both splits of SceneFun3D.
225        Remarkably, our training-free approach outperforms both training-free and training-based methods, including those using significantly larger models (72B).
226        <ul>
227          <li>Compared to Fun3DU (training-free): <b>+84.9% relative AP<sub>50</sub></b> and <b>+59.9% relative mIoU</b> on split0.</li>
228          <li>Compared to AffordBot-72B (training-based, fine-tuned for 1000 epochs): <b>+49.4% AP<sub>50</sub></b> and <b>+68.5% mIoU</b>.</li>
229          <li>UniFunc3D also achieves a <b>3.2&times; speedup</b> over Fun3DU (~26 min vs. ~82 min per scene).</li>
230        </ul>
231      </div>
232
233      <div class="content has-text-centered" style="overflow-x: auto;">
234        <table class="table is-bordered is-striped is-narrow is-hoverable is-fullwidth">
235          <thead>
236            <tr>
237              <th rowspan="2">Method</th>
238              <th colspan="5" style="text-align:center;">Split0 (30 scenes, val)</th>
239              <th colspan="5" style="text-align:center;">Split1 (200 scenes, train)</th>
240            </tr>
241            <tr>
242              <th>AP<sub>50</sub></th><th>AP<sub>25</sub></th><th>AR<sub>50</sub></th><th>AR<sub>25</sub></th><th>mIoU</th>
243              <th>AP<sub>50</sub></th><th>AP<sub>25</sub></th><th>AR<sub>50</sub></th><th>AR<sub>25</sub></th><th>mIoU</th>
244            </tr>
245          </thead>
246          <tbody>
247            <tr><td colspan="11"><i>Training-based methods:</i></td></tr>
248            <tr>
249              <td>TASA-72B</td>
250              <td>26.9</td><td>28.6</td><td>—</td><td>—</td><td>19.7</td>
251              <td colspan="5" style="text-align:center;"><i>trained on split1</i></td>
252            </tr>
253            <tr>
254              <td>AffordBot-72B</td>
255              <td>20.91</td><td>24.76</td><td>18.99</td><td>22.84</td><td>14.42</td>
256              <td colspan="5" style="text-align:center;"><i>trained on split1</i></td>
257            </tr>
258            <tr><td colspan="11"><i>Training-free methods:</i></td></tr>
259            <tr>
260              <td>Fun3DU-9B</td>
261              <td>16.9</td><td>33.3</td><td>38.2</td><td>46.7</td><td>15.2</td>
262              <td>12.6</td><td>23.1</td><td>32.9</td><td>40.5</td><td>11.5</td>
263            </tr>
264            <tr>
265              <td>UniFunc3D-8B (Ours)</td>
266              <td>23.82</td><td>44.04</td><td>46.07</td><td>55.51</td><td>20.92</td>
267              <td>16.24</td><td>29.02</td><td>38.91</td><td>48.15</td><td>14.23</td>
268            </tr>
269            <tr style="font-weight:bold;">
270              <td>UniFunc3D-30B (Ours)</td>
271              <td>31.24</td><td>51.01</td><td>46.97</td><td>58.88</td><td>24.30</td>
272              <td>21.32</td><td>35.76</td><td>40.03</td><td>51.00</td><td>17.09</td>
273            </tr>
274          </tbody>
275        </table>
276      </div>
277    </div>
278
279    <!-- Qualitative Results -->
280    <div class="hero-body">
281      <h3 class="title is-3">Qualitative Results</h3>
282      <div class="content has-text-justified">
283        Qualitative comparison across five representative queries. Our method clearly outperforms prior methods in spatial disambiguation and handling small interactive objects.
284        For example, for <i>"Open the top left drawer of the cabinet with the beauty products on top"</i>,
285        our method finds the correct <b>top-left knob</b>, while AffordBot finds the wrong top-right knob and Fun3DU mistakenly segments the drawer face.
286      </div>
287
288      <!-- Column header -->
289      <div class="columns is-centered is-vcentered" style="margin-bottom:0; font-weight:bold;">
290        <div class="column is-2 has-text-centered"></div>
291        <div class="column has-text-centered"><p class="is-size-7">AffordBot</p></div>
292        <div class="column has-text-centered"><p class="is-size-7">Fun3DU</p></div>
293        <div class="column has-text-centered"><p class="is-size-7"><b>Ours</b></p></div>
294        <div class="column has-text-centered"><p class="is-size-7">GT</p></div>
295      </div>
296
297      <!-- q1 -->
298      <div class="columns is-centered is-vcentered" style="margin-bottom:0.4rem;">
299        <div class="column is-2 has-text-centered">
300          <p class="is-size-7"><i>"Open the top left drawer of the cabinet with the beauty products on top"</i></p>
301        </div>
302        <div class="column"><img src="./static/images/qual/q1_affordbot.jpg" alt="AffordBot"></div>
303        <div class="column"><img src="./static/images/qual/q1_fun3du.jpg" alt="Fun3DU"></div>
304        <div class="column"><img src="./static/images/qual/q1_ours.jpg" alt="Ours"></div>
305        <div class="column"><img src="./static/images/qual/q1_gt.jpg" alt="GT"></div>
306      </div>
307
308      <!-- q2 -->
309      <div class="columns is-centered is-vcentered" style="margin-bottom:0.4rem;">
310        <div class="column is-2 has-text-centered">
311          <p class="is-size-7"><i>"Turn on the ceiling light"</i></p>
312        </div>
313        <div class="column"><img src="./static/images/qual/q2_affordbot.jpg" alt="AffordBot"></div>
314        <div class="column"><img src="./static/images/qual/q2_fun3du.jpg" alt="Fun3DU"></div>
315        <div class="column"><img src="./static/images/qual/q2_ours.jpg" alt="Ours"></div>
316        <div class="column"><img src="./static/images/qual/q2_gt.jpg" alt="GT"></div>
317      </div>
318
319      <!-- q3 -->
320      <div class="columns is-centered is-vcentered" style="margin-bottom:0.4rem;">
321        <div class="column is-2 has-text-centered">
322          <p class="is-size-7"><i>"Control the water flow in the bathtub using the drain control dial"</i></p>
323        </div>
324        <div class="column"><img src="./static/images/qual/q3_affordbot.jpg" alt="AffordBot"></div>
325        <div class="column"><img src="./static/images/qual/q3_fun3du.jpg" alt="Fun3DU"></div>
326        <div class="column"><img src="./static/images/qual/q3_ours.jpg" alt="Ours"></div>
327        <div class="column"><img src="./static/images/qual/q3_gt.jpg" alt="GT"></div>
328      </div>
329
330      <!-- q4 -->
331      <div class="columns is-centered is-vcentered" style="margin-bottom:0.4rem;">
332        <div class="column is-2 has-text-centered">
333          <p class="is-size-7"><i>"Select a washing program"</i></p>
334        </div>
335        <div class="column"><img src="./static/images/qual/q4_affordbot.jpg" alt="AffordBot"></div>
336        <div class="column"><img src="./static/images/qual/q4_fun3du.jpg" alt="Fun3DU"></div>
337        <div class="column"><img src="./static/images/qual/q4_ours.jpg" alt="Ours"></div>
338        <div class="column"><img src="./static/images/qual/q4_gt.jpg" alt="GT"></div>
339      </div>
340
341      <!-- q5 -->
342      <div class="columns is-centered is-vcentered">
343        <div class="column is-2 has-text-centered">
344          <p class="is-size-7"><i>"Flush the toilet"</i></p>
345        </div>
346        <div class="column"><img src="./static/images/qual/q5_affordbot.jpg" alt="AffordBot"></div>
347        <div class="column"><img src="./static/images/qual/q5_fun3du.jpg" alt="Fun3DU"></div>
348        <div class="column"><img src="./static/images/qual/q5_ours.jpg" alt="Ours"></div>
349        <div class="column"><img src="./static/images/qual/q5_gt.jpg" alt="GT"></div>
350      </div>
351
352    </div>
353    <!-- / Qualitative Results -->
354
355  </div>
356</section>
357
358
359<section class="section" id="BibTeX">
360  <div class="container is-max-desktop content">
361    <h2 class="title">BibTeX</h2>
362    <pre><code>@InProceedings{Lin_UniFunc3D,
363    author    = {Lin, Jiaying and Xu, Dan},
364    title     = {UniFunc3D: Unified Active Spatial-Temporal Grounding for 3D Functionality Segmentation},
365    booktitle = {arXiv},
366    year      = {2026},
367}</code></pre>
368  </div>
369</section>
370
371
372<footer class="footer">
373  <div class="container">
374    <div class="content has-text-centered">
375      <a class="icon-link" href="https://jiaying.link">
376        <i class="fas fa-home"></i>
377      </a>
378    </div>
379    <div class="columns is-centered">
380      <div class="column is-8">
381        <div class="content">
382          <p>
383            This website is licensed under a <a rel="license"
384                                                href="http://creativecommons.org/licenses/by-sa/4.0/">Creative
385            Commons Attribution-ShareAlike 4.0 International License</a>.
386          </p>
387          <p>
388            We borrow this website template from <a
389            href="https://github.com/nerfies/nerfies.github.io">Nerfies</a>.
390          </p>
391        </div>
392      </div>
393    </div>
394  </div>
395</footer>
396
397
398<script type="text/javascript">
399  var sc_project=12773779;
400  var sc_invisible=1;
401  var sc_security="f1222376";
402  </script>
402
403  
403<script type="text/javascript"
404  src="https://www.statcounter.com/counter/counter.js"
405  async></script>
405
406  <noscript><div class="statcounter"><a title="Web Analytics"
407  href="https://statcounter.com/" target="_blank"><img
408  class="statcounter"
409  src="https://c.statcounter.com/12773779/0/f1222376/1/"
410  alt="Web Analytics"
411  referrerPolicy="no-referrer-when-downgrade"></a></div></noscript>
412<!-- End of Statcounter Code -->
413</body>
414</html>

Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.