PageSourceSearch

https://yangcaoai.github.io/SpatialSpeak/

html yangcaoai.github.io collected 2026-10-03 10:12:15 UTC 20,147 bytes, 152 lines download raw bytes

1<!doctype html>
2<html lang="en">
3<head>
4  <meta charset="utf-8">
5  <meta name="viewport" content="width=device-width,initial-scale=1">
6  <meta name="theme-color" content="#f8faf7">
7  <title>SpatialSpeak | Spatial Chain-of-Thought Reasoning</title>
8  <meta name="description" content="SpatialSpeak connects QA-native reconstruction with spatial chain-of-thought reasoning. Explore colored 3D point clouds, recorded reasoning examples, and research results.">
9  <meta property="og:title" content="SpatialSpeak">
10  <meta property="og:description" content="QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning">
11  <meta property="og:type" content="website">
12  <meta property="og:image" content="assets/videos/sample-1.webp?v=e0fa39db7df5">
13  <link rel="icon" href="assets/favicon.svg?v=17b893ce4f2b" type="image/svg+xml">
14  <link rel="stylesheet" href="styles.css?v=61bb3c21332b">
15  
15<script defer src="site-config.js?v=e7ea751c6144"></script>
15
16  
16<script defer src="cloud-annotations.js?v=fb72581adddd"></script>
16
17  
17<script defer src="cloud-viewer.js?v=c9ee09d54e38"></script>
17
18  
18<script defer src="app.js?v=1e4b45937048"></script>
18
19  
19<script defer src="demo.js?v=e9012dc6e60d"></script>
19
20</head>
21<body>
22<a class="skip-link" href="#main">Skip to content</a>
23<header class="site-header">
24  <a class="wordmark" href="#top" aria-label="SpatialSpeak home">SpatialSpeak</a>
25  <nav aria-label="Main navigation"><a href="#demo">Demo</a><a href="#explore">3D Scenes</a><a href="#method">Method</a><a href="#results">Results</a><a href="#citation">Citation</a></nav>
26  <span class="header-note"></span>
27</header>
28<main id="main">
29  <section class="hero section-width" id="top">
30    <p class="eyebrow"><span class="status-dot"></span> MULTI-VIEW SPATIAL INTELLIGENCE</p>
31    <h1>Spatial<span>Speak</span><span class="title-period">.</span></h1>
32    <p class="paper-title">QA-Native Reconstruction with Local and Global Context<br class="desktop-break"> for Spatial Chain-of-Thought Reasoning</p>
33    <p class="authors"><span><a href="https://yangcaoai.github.io/" target="_blank" rel="noopener noreferrer">Yang Cao</a><sup>1</sup></span><span><a href="https://zestfuljx.github.io/" target="_blank" rel="noopener noreferrer">Jiaxin Zhang</a><sup>3</sup></span><span><a href="https://daveredrum.github.io/" target="_blank" rel="noopener noreferrer">Dave Zhenyu Chen</a><sup>2</sup></span><span><a href="https://zhongyingji.github.io/" target="_blank" rel="noopener noreferrer">Yingji Zhong</a><sup>1</sup></span><span><a href="https://gaoruiyuan.com/" target="_blank" rel="noopener noreferrer">Ruiyuan Gao</a><sup>2</sup></span><span><a href="https://racheltechie.github.io/" target="_blank" rel="noopener noreferrer">Lanqing Hong</a><sup>2</sup></span><span><a href="https://www.danxurgb.net/" target="_blank" rel="noopener noreferrer">Dan Xu</a><sup>1,*</sup></span></p>
34    <p class="affiliations"><span><sup>1</sup> Hong Kong University of Science and Technology</span><span><sup>2</sup> Huawei Noah’s Ark Lab</span><span><sup>3</sup> Harbin Institute of Technology</span></p>
35    <div class="resource-links" aria-label="Project resources">
36      <a class="button primary" href="#demo"><span aria-hidden="true">▶</span> Watch the demo</a>
37      <a class="button primary" data-resource="paperUrl" href="https://arxiv.org/pdf/2609.33616" target="_blank" rel="noopener noreferrer">Read the paper ↗</a>
38      <a class="button primary" data-resource="codeUrl" href="https://github.com/yangcaoai/SpatialSpeak-VLM" target="_blank" rel="noopener noreferrer">Code on GitHub ↗</a>
39    </div>
40  </section>
41
42  <section class="demo-section section-width" id="demo" aria-label="SpatialSpeak video demonstration">
43    <div class="demo-video-card">
44      <video id="teaser-video" controls autoplay muted loop playsinline preload="metadata" poster="assets/videos/sample-1.webp?v=e0fa39db7df5" aria-label="Three recorded paper examples: rotating point clouds with model responses appearing line by line">
45        <source src="assets/videos/spatialspeak-demo.mp4?v=ef40b64e46d6" type="video/mp4">
46        <p>Your browser does not support embedded video. <a href="assets/videos/spatialspeak-demo.mp4?v=ef40b64e46d6">Download the demonstration.</a></p>
47      </video>
48    </div>
49    <div class="demo-chapters" aria-label="Video samples">
50      <button data-video-time="0" aria-pressed="true"><span>Sample 1</span> scene0030_00</button>
51      <button data-video-time="14" aria-pressed="false"><span>Sample 2</span> scene0441_00</button>
52      <button data-video-time="28" aria-pressed="false"><span>Sample 3</span> scene0578_00</button>
53    </div>
54    <a class="video-download" href="assets/videos/spatialspeak-demo.mp4?v=ef40b64e46d6" download>Download demo video <span aria-hidden="true">↓</span></a>
55  </section>
56
57  <section class="paper-overview section-width" id="overview" aria-labelledby="overview-title">
58    <div class="section-heading"><h2 id="overview-title">Learning local geometry and global context makes spatial CoT more effective.</h2></div>
59    <figure><button class="figure-button" data-figure="1" aria-label="Enlarge research overview and performance comparison"><img src="assets/figures/figure-1.svg?v=5428e6490016" alt="SpatialSpeak overview with its ReVSI performance advantage: 62.8 compared with 54.1 for the strongest baseline, a gain of 8.7 points" width="534" height="291" loading="lazy"></button>
60    <figcaption><strong>62.8 on ReVSI · +8.7 points.</strong> Reconstruction pretraining connects geometric understanding with spatial reasoning. <span class="figure-actions">
60<button class="text-link" data-figure="1">Enlarge ↗</button> <a class="text-link figure-pdf" href="assets/figures/figure-1.pdf" target="_blank" rel="noopener">View PDF ↗</a></span></figcaption></figure>
61  </section>
62
63  <section class="explore section-width" id="explore" aria-labelledby="explore-title">
64    <div class="section-heading"><div><p class="eyebrow">THE SCENES BEHIND THE PAPER</p><h2 id="explore-title">Explore the 3D Scenes</h2></div><p>Drag to rotate. Scroll or pinch to zoom.<br>Choose a scene to inspect its reconstruction and paper example.</p></div>
65    <div class="scene-tabs" role="tablist" aria-label="Choose a 3D scene">
66      <button id="tab-classroom" role="tab" aria-selected="true" aria-controls="scene-panel" data-scene="classroom"><span>scene0030_00</span></button>
67      <button id="tab-bathroom" role="tab" aria-selected="false" aria-controls="scene-panel" tabindex="-1" data-scene="bathroom"><span>scene0441_00</span></button>
68      <button id="tab-chairs" role="tab" aria-selected="false" aria-controls="scene-panel" tabindex="-1" data-scene="chairs"><span>scene0578_00</span></button>
69      <button id="tab-reconstruction" role="tab" aria-selected="false" aria-controls="scene-panel" tabindex="-1" class="metric-tab-start" data-scene="reconstruction"><span>scene0086_02</span></button>
70      <button id="tab-reconstruction-b" role="tab" aria-selected="false" aria-controls="scene-panel" tabindex="-1" data-scene="reconstruction-b"><span>scene0645_00</span></button>
71    </div>
72    <div class="scene-panel" id="scene-panel" role="tabpanel" aria-labelledby="tab-classroom">
73      <div class="scene-visual">
74        <div class="viewer" id="viewer">
75          <div id="measurement-preview" class="measurement-preview" hidden></div>
76          <div class="viewer-top"><span><i class="status-dot"></i> <span id="scene-kind">Colored reconstruction</span></span><span id="point-count">Loading scene…</span></div>
77          <img id="cloud-fallback" src="assets/frames/image26.webp" alt="A colored point cloud of a classroom with two blackboards" class="cloud-fallback">
78          <canvas id="cloud-canvas" tabindex="0" aria-label="Interactive point cloud. Drag or use arrow keys to rotate. Use plus and minus to zoom, R to reset."></canvas>
79          <div class="viewer-message" id="viewer-message" role="status">Loading 3D scene…</div>
80          <div class="viewer-bottom"><span class="viewer-hint">↔ Drag to rotate <span>· Scroll to zoom</span></span><div class="viewer-controls"><button id="measurement-toggle" hidden aria-pressed="true">Rotate in 3D</button><button id="rotate-button" aria-pressed="true" title="Toggle automatic rotation">Ⅱ <span>Pause rotation</span></button><button id="reset-view" title="Reset view">Reset view</button><button id="fullscreen-button" aria-label="Expand 3D viewer" title="Expand viewer">⤢</button></div></div>
81        </div>
82        <div class="input-views"><div class="views-label"><span>INPUT VIEWS</span><small id="figure-ref"></small></div><div id="input-frames" class="frame-strip"></div></div>
83      </div>
84      <div class="reasoning-panel" aria-live="polite">
85        <p class="eyebrow" id="example-type">OBJECT COUNTING</p>
86        <h3 id="scene-question">How many blackboards are in the scene?</h3>
87        <p class="example-note" id="example-note">A recorded example from the paper, paired with its reconstruction.</p>
88        <div class="response-heading"><span>SpatialSpeak response</span><button id="replay-response" aria-label="Replay the recorded model response">↻ Replay</button></div><ol class="reasoning-steps" id="reasoning-steps"></ol>
89        <div class="answer-strip"><div><span id="answer-label">SpatialSpeak</span><strong id="answer-value">2 <small>blackboards</small></strong></div><div><span id="truth-label">Ground truth</span><strong id="truth-value">2</strong></div></div>
90        <p class="baseline-result" id="baseline-result">Without QA-RP and CoT-VC: <strong>3</strong></p>
91        <button class="text-link" id="view-example-figure">View the paper figure <span aria-hidden="true">↗</span></button>
92      </div>
93    </div>
94    <noscript><p class="no-script-note">Enable JavaScript to rotate the point clouds and switch examples. The paper figures and research summary below remain available.</p></noscript>
95  </section>
96
97
98  <section class="metric-scenes section-width" id="measurements" aria-labelledby="metric-scenes-title">
99    <div class="section-heading"><h2 id="metric-scenes-title">Metric Reconstruction</h2><p>The red arrows show exactly which span is measured.</p></div>
100    <div class="metric-scene-grid">
101      <article class="metric-scene-card"><h3>scene0086_02</h3><div class="measurement-pair"><figure><figcaption>SpatialSpeak</figcaption><img class="paper-measurement" src="assets/figures/scene0086_02-ours.svg?v=eab4e9c64280" alt="SpatialSpeak: red arrows mark the 1.24-meter span" loading="lazy"></figure><figure><figcaption>Ground truth</figcaption><img class="paper-measurement" src="assets/figures/scene0086_02-gt.svg?v=72f4712054ec" alt="Ground truth: the same span is 1.25 meters" loading="lazy"></figure></div><p class="measurement-baselines">CUT3R: 1.38 m · MapAnything: 1.63 m</p><a class="text-link" href="#explore" data-select-scene="reconstruction">Explore this scene ↗</a></article>
102      <article class="metric-scene-card"><h3>scene0645_00</h3><div class="measurement-pair"><figure><figcaption>SpatialSpeak</figcaption><img class="paper-measurement" src="assets/figures/scene0645_00-ours.svg?v=905ba1710ac0" alt="SpatialSpeak: red arrows mark the 2.52-meter bed length" loading="lazy"></figure><figure><figcaption>Ground truth</figcaption><img class="paper-measurement" src="assets/figures/scene0645_00-gt.svg?v=b674e032e48b" alt="Ground truth: the same bed length is 2.54 meters" loading="lazy"></figure></div><p class="measurement-baselines">CUT3R: 2.82 m · MapAnything: 2.90 m</p><a class="text-link" href="#explore" data-select-scene="reconstruction-b">Explore this scene ↗</a></article>
103    </div>
104    <p class="measurement-note">Original measurement annotations from the paper. Switch to the 3D view to inspect the corresponding reconstruction.</p>
105  </section>
106
107  <section class="method-section" id="method" aria-labelledby="method-title">
108    <div class="section-width">
109      <div class="section-heading"><div><p class="eyebrow">TWO STAGES, ONE LANGUAGE INTERFACE</p><h2 id="method-title">Method</h2></div></div>
110      <div class="stage-grid">
111        <article class="stage"><span class="stage-number">01</span><div><p class="stage-kicker">QA-RP · RECONSTRUCTION PRETRAINING</p><p>Marked-point 3D queries teach fine-grained local geometry. Object-center queries teach global scene context across views. Both use text-based question answering.</p></div></article>
112        <article class="stage"><span class="stage-number">02</span><div><p class="stage-kicker">CoT-VC · SPATIAL REASONING</p><p>Spatial chain-of-thought connects estimated geometry to the requested answer. Reliability assessment and visual compensation support refinement when needed.</p></div></article>
113      </div>
114      <figure class="method-figure"><button class="figure-button" data-figure="2" aria-label="Enlarge method overview"><img src="assets/figures/figure-2.svg?v=374b242655a4" alt="The SpatialSpeak framework: QA-RP pretraining combines local point and global object-center queries, followed by CoT-VC finetuning for spatial reasoning" loading="lazy" width="622" height="294"></button><figcaption>QA-native reconstruction pretraining connects local and global scene understanding with spatial chain-of-thought learning. <span class="figure-actions"><button class="text-link" data-figure="2">Enlarge ↗</button> <a class="text-link figure-pdf" href="assets/figures/figure-2.pdf" target="_blank" rel="noopener">View PDF ↗</a></span></figcaption></figure>
115    </div>
116  </section>
117
118  <section class="results section-width" id="results" aria-labelledby="results-title">
119    <div class="section-heading"><div><p class="eyebrow">EXPERIMENTS</p><h2 id="results-title">Experiments</h2></div><p>Results reported in the paper.<br>Higher is better.</p></div>
120    <div class="metrics"><div><strong>
12062.8<span>%</span></strong><p>ReVSI average</p></div><div><strong>+8.7<span>pts</span></strong><p>Over the strongest compared baseline</p></div><div><strong>4<span>B</span></strong><p>VLM backbone parameters</p></div></div>
121    <div class="results-grid">
122      <article class="benchmark-chart"><div class="chart-heading"><h3>ReVSI performance</h3><span>Average score (%)</span></div><div class="bar-chart" id="bar-chart">
123        <div class="bar-row ours"><span>SpatialSpeak-4B</span><div class="bar-track"><i style="width:95.111%"></i></div><strong>62.8</strong></div>
124        <div class="bar-row"><span>SpatialStack-4B</span><div class="bar-track"><i style="width:75.778%"></i></div><strong>54.1</strong></div>
125        <div class="bar-row"><span>GeoThinker-8B</span><div class="bar-track"><i style="width:75.556%"></i></div><strong>54.0</strong></div>
126        <div class="bar-row"><span>VLM-3R-7B</span><div class="bar-track"><i style="width:66.889%"></i></div><strong>50.1</strong></div>
127        <div class="bar-row"><span>Cambrian-S-7B</span><div class="bar-track"><i style="width:64.667%"></i></div><strong>49.1</strong></div>
128        <div class="chart-axis" aria-hidden="true"><span>20</span><span>65</span></div>
129      </div><p class="chart-note">Displayed score range: 20–65%. Selected comparisons from Table 6. SpatialSpeak, SpatialStack, and GeoThinker are evaluated under the same 32-frame setting. VLM-3R and Cambrian-S results are sourced from the ReVSI paper.</p></article>
130      <article class="ablation"><p class="eyebrow">WHY RECONSTRUCTION MATTERS</p><h3>Spatial CoT gains more<br>after QA-RP.</h3><p>The improvement from CoT-VC grows from <strong>2.6</strong> to <strong>6.9 points</strong> when the model first learns reconstruction.</p><div class="gain-row"><span>Without QA-RP</span><div class="gain-track"><i style="width:37.681%"></i></div><strong>+2.6</strong></div><div class="gain-row highlight"><span>With QA-RP</span><div class="gain-track"><i style="width:100%"></i></div><strong>+6.9</strong></div><p class="chart-note">ReVSI ablation, Table 1. Baseline: 52.4. CoT-VC only: 55.0. QA-RP only: 55.9. Full method: 62.8.</p></article>
131    </div>
132    <div class="table-wrap"><table><caption>Spatial reasoning across benchmarks</caption><thead><tr><th scope="col">Benchmark</th><th scope="col">Setting</th><th scope="col">SpatialSpeak-4B</th><th scope="col">Best compared baseline</th></tr></thead><tbody><tr><th scope="row">ReVSI</th><td>32-frame evaluation</td><td class="score">62.8</td><td>54.1 <small>SpatialStack-4B</small></td></tr><tr><th scope="row">VSI-Bench</th><td>Normal training</td><td class="score">63.3</td><td>55.4 <small>Omni-View-7B</small></td></tr><tr><th scope="row">VSI-Bench</th><td>Scaled training</td><td class="score">73.0</td><td>72.6 <small>GeoThinker-8B</small></td></tr><tr><th scope="row">SPAR-Bench</th><td>Overall average</td><td class="score">76.0</td><td>72.0 <small>SpatialStack-4B</small></td></tr></tbody></table></div>
133    <p class="table-note">Scores and comparison settings follow Tables 6–9 of the paper. Normal and scaled training settings use different training data and should be read separately.</p>
134  </section>
135
136
137
138  <section class="abstract-section section-width" aria-labelledby="abstract-title"><h2 id="abstract-title">Abstract</h2><p>Vision-language models (VLMs) can benefit from geometric priors for multi-view spatial reasoning, yet answer-only training does not directly supervise the intermediate geometric estimates and their use in deriving quantitative spatial answers. We hypothesize that spatial chain-of-thought (CoT) supervision becomes more effective when the VLM first jointly learns complementary local geometry and global scene context through multi-view reconstruction.</p><p>We introduce <strong>SpatialSpeak</strong>, a two-stage framework that connects QA-native reconstruction pretraining with spatial CoT learning. In Stage I, QA-Native Reconstruction Pretraining (QA-RP) combines marked-point 3D queries for fine-grained local geometry with object-center queries for global scene context across views. Both tasks are formulated as text-based question answering, allowing geometric estimation and subsequent reasoning to share the same autoregressive output interface. In Stage II, spatial CoT with Visual Compensation (CoT-VC) trains the model to express question-relevant geometric estimates and use them to derive answers, with reliability assessment and visual compensation supporting answer refinement when needed.</p><p>On ReVSI, QA-RP increases the gain from CoT-VC from 2.6 to 6.9 points, and ablations show that both local and global reconstruction supervision are beneficial. SpatialSpeak achieves state-of-the-art results on ReVSI, VSI-Bench, and SPAR-Bench, with a ReVSI score of 62.8 that exceeds the strongest compared baseline by 8.7 points.</p></section>
139
140  <section class="citation-section section-width" id="citation" aria-labelledby="citation-title"><div class="section-heading"><div><p class="eyebrow">REFERENCE</p><h2 id="citation-title">Citation</h2></div><button class="button secondary" id="copy-citation">Copy BibTeX <span aria-hidden="true">⧉</span></button></div><pre><code id="bibtex">@article{cao2026spatialspeak,
141  title = {SpatialSpeak: QA-Native Reconstruction with Local and Global
142           Context for Spatial Chain-of-Thought Reasoning},
143  author = {Cao, Yang and Zhang, Jiaxin and Chen, Dave Zhenyu and
144            Zhong, Yingji and Gao, Ruiyuan and Hong, Lanqing and Xu, Dan},
145  journal = {arXiv preprint arXiv:2609.33616},
146  year = {2026}
147}</code></pre><p class="citation-note">arXiv:2609.33616</p><span id="copy-status" class="visually-hidden" role="status"></span></section>
148</main>
149<footer class="section-width"><a class="wordmark" href="#top">SpatialSpeak<span class="title-period">.</span></a><p>QA-native reconstruction for spatial reasoning.</p></footer>
150<dialog id="figure-dialog" aria-labelledby="dialog-title"><div class="dialog-top"><h2 id="dialog-title">Paper figure</h2><a id="dialog-pdf" class="text-link" href="assets/figures/figure-1.pdf" target="_blank" rel="noopener">View original PDF ↗</a><button id="close-figure" aria-label="Close figure">×</button></div><img id="dialog-image" alt=""><p id="dialog-caption"></p></dialog>
151</body>
152</html>

Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.