1<!doctype html> 2<html lang="en"> 3<head> 4 <meta charset="utf-8"> 5 <meta name="viewport" content="width=device-width,initial-scale=1"> 6 <meta name="theme-color" content="#f8faf7"> 7 <title>SpatialSpeak | Spatial Chain-of-Thought Reasoning</title> 8 <meta name="description" content="SpatialSpeak connects QA-native reconstruction with spatial chain-of-thought reasoning. Explore colored 3D point clouds, recorded reasoning examples, and research results."> 9 <meta property="og:title" content="SpatialSpeak"> 10 <meta property="og:description" content="QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning"> 11 <meta property="og:type" content="website"> 12 <meta property="og:image" content="assets/videos/sample-1.webp?v=e0fa39db7df5"> 13 <link rel="icon" href="assets/favicon.svg?v=17b893ce4f2b" type="image/svg+xml"> 14 <link rel="stylesheet" href="styles.css?v=61bb3c21332b"> 15
15<script defer src="site-config.js?v=e7ea751c6144"></script>
15 16
16<script defer src="cloud-annotations.js?v=fb72581adddd"></script>
16 17
17<script defer src="cloud-viewer.js?v=c9ee09d54e38"></script>
17 18
18<script defer src="app.js?v=1e4b45937048"></script>
18 19
19<script defer src="demo.js?v=e9012dc6e60d"></script>
19 20</head> 21<body> 22<a class="skip-link" href="#main">Skip to content</a> 23<header class="site-header"> 24 <a class="wordmark" href="#top" aria-label="SpatialSpeak home">SpatialSpeak</a> 25 <nav aria-label="Main navigation"><a href="#demo">Demo</a><a href="#explore">3D Scenes</a><a href="#method">Method</a><a href="#results">Results</a><a href="#citation">Citation</a></nav> 26 <span class="header-note"></span> 27</header> 28<main id="main"> 29 <section class="hero section-width" id="top"> 30 <p class="eyebrow"><span class="status-dot"></span> MULTI-VIEW SPATIAL INTELLIGENCE</p> 31 <h1>Spatial<span>Speak</span><span class="title-period">.</span></h1> 32 <p class="paper-title">QA-Native Reconstruction with Local and Global Context<br class="desktop-break"> for Spatial Chain-of-Thought Reasoning</p> 33 <p class="authors"><span><a href="https://yangcaoai.github.io/" target="_blank" rel="noopener noreferrer">Yang Cao</a><sup>1</sup></span><span><a href="https://zestfuljx.github.io/" target="_blank" rel="noopener noreferrer">Jiaxin Zhang</a><sup>3</sup></span><span><a href="https://daveredrum.github.io/" target="_blank" rel="noopener noreferrer">Dave Zhenyu Chen</a><sup>2</sup></span><span><a href="https://zhongyingji.github.io/" target="_blank" rel="noopener noreferrer">Yingji Zhong</a><sup>1</sup></span><span><a href="https://gaoruiyuan.com/" target="_blank" rel="noopener noreferrer">Ruiyuan Gao</a><sup>2</sup></span><span><a href="https://racheltechie.github.io/" target="_blank" rel="noopener noreferrer">Lanqing Hong</a><sup>2</sup></span><span><a href="https://www.danxurgb.net/" target="_blank" rel="noopener noreferrer">Dan Xu</a><sup>1,*</sup></span></p> 34 <p class="affiliations"><span><sup>1</sup> Hong Kong University of Science and Technology</span><span><sup>2</sup> Huawei Noahâs Ark Lab</span><span><sup>3</sup> Harbin Institute of Technology</span></p> 35 <div class="resource-links" aria-label="Project resources"> 36 <a class="button primary" href="#demo"><span aria-hidden="true">â¶</span> Watch the demo</a> 37 <a class="button primary" data-resource="paperUrl" href="https://arxiv.org/pdf/2609.33616" target="_blank" rel="noopener noreferrer">Read the paper â</a> 38 <a class="button primary" data-resource="codeUrl" href="https://github.com/yangcaoai/SpatialSpeak-VLM" target="_blank" rel="noopener noreferrer">Code on GitHub â</a> 39 </div> 40 </section> 41 42 <section class="demo-section section-width" id="demo" aria-label="SpatialSpeak video demonstration"> 43 <div class="demo-video-card"> 44 <video id="teaser-video" controls autoplay muted loop playsinline preload="metadata" poster="assets/videos/sample-1.webp?v=e0fa39db7df5" aria-label="Three recorded paper examples: rotating point clouds with model responses appearing line by line"> 45 <source src="assets/videos/spatialspeak-demo.mp4?v=ef40b64e46d6" type="video/mp4"> 46 <p>Your browser does not support embedded video. <a href="assets/videos/spatialspeak-demo.mp4?v=ef40b64e46d6">Download the demonstration.</a></p> 47 </video> 48 </div> 49 <div class="demo-chapters" aria-label="Video samples"> 50 <button data-video-time="0" aria-pressed="true"><span>Sample 1</span> scene0030_00</button> 51 <button data-video-time="14" aria-pressed="false"><span>Sample 2</span> scene0441_00</button> 52 <button data-video-time="28" aria-pressed="false"><span>Sample 3</span> scene0578_00</button> 53 </div> 54 <a class="video-download" href="assets/videos/spatialspeak-demo.mp4?v=ef40b64e46d6" download>Download demo video <span aria-hidden="true">â</span></a> 55 </section> 56 57 <section class="paper-overview section-width" id="overview" aria-labelledby="overview-title"> 58 <div class="section-heading"><h2 id="overview-title">Learning local geometry and global context makes spatial CoT more effective.</h2></div> 59 <figure><button class="figure-button" data-figure="1" aria-label="Enlarge research overview and performance comparison"><img src="assets/figures/figure-1.svg?v=5428e6490016" alt="SpatialSpeak overview with its ReVSI performance advantage: 62.8 compared with 54.1 for the strongest baseline, a gain of 8.7 points" width="534" height="291" loading="lazy"></button> 60 <figcaption><strong>62.8 on ReVSI · +8.7 points.</strong> Reconstruction pretraining connects geometric understanding with spatial reasoning. <span class="figure-actions">
60<button class="text-link" data-figure="1">Enlarge â</button> <a class="text-link figure-pdf" href="assets/figures/figure-1.pdf" target="_blank" rel="noopener">View PDF â</a></span></figcaption></figure> 61 </section> 62 63 <section class="explore section-width" id="explore" aria-labelledby="explore-title"> 64 <div class="section-heading"><div><p class="eyebrow">THE SCENES BEHIND THE PAPER</p><h2 id="explore-title">Explore the 3D Scenes</h2></div><p>Drag to rotate. Scroll or pinch to zoom.<br>Choose a scene to inspect its reconstruction and paper example.</p></div> 65 <div class="scene-tabs" role="tablist" aria-label="Choose a 3D scene"> 66 <button id="tab-classroom" role="tab" aria-selected="true" aria-controls="scene-panel" data-scene="classroom"><span>scene0030_00</span></button> 67 <button id="tab-bathroom" role="tab" aria-selected="false" aria-controls="scene-panel" tabindex="-1" data-scene="bathroom"><span>scene0441_00</span></button> 68 <button id="tab-chairs" role="tab" aria-selected="false" aria-controls="scene-panel" tabindex="-1" data-scene="chairs"><span>scene0578_00</span></button> 69 <button id="tab-reconstruction" role="tab" aria-selected="false" aria-controls="scene-panel" tabindex="-1" class="metric-tab-start" data-scene="reconstruction"><span>scene0086_02</span></button> 70 <button id="tab-reconstruction-b" role="tab" aria-selected="false" aria-controls="scene-panel" tabindex="-1" data-scene="reconstruction-b"><span>scene0645_00</span></button> 71 </div> 72 <div class="scene-panel" id="scene-panel" role="tabpanel" aria-labelledby="tab-classroom"> 73 <div class="scene-visual"> 74 <div class="viewer" id="viewer"> 75 <div id="measurement-preview" class="measurement-preview" hidden></div> 76 <div class="viewer-top"><span><i class="status-dot"></i> <span id="scene-kind">Colored reconstruction</span></span><span id="point-count">Loading sceneâ¦</span></div> 77 <img id="cloud-fallback" src="assets/frames/image26.webp" alt="A colored point cloud of a classroom with two blackboards" class="cloud-fallback"> 78 <canvas id="cloud-canvas" tabindex="0" aria-label="Interactive point cloud. Drag or use arrow keys to rotate. Use plus and minus to zoom, R to reset."></canvas> 79 <div class="viewer-message" id="viewer-message" role="status">Loading 3D sceneâ¦</div> 80 <div class="viewer-bottom"><span class="viewer-hint">â Drag to rotate <span>· Scroll to zoom</span></span><div class="viewer-controls"><button id="measurement-toggle" hidden aria-pressed="true">Rotate in 3D</button><button id="rotate-button" aria-pressed="true" title="Toggle automatic rotation">â ¡ <span>Pause rotation</span></button><button id="reset-view" title="Reset view">Reset view</button><button id="fullscreen-button" aria-label="Expand 3D viewer" title="Expand viewer">⤢</button></div></div> 81 </div> 82 <div class="input-views"><div class="views-label"><span>INPUT VIEWS</span><small id="figure-ref"></small></div><div id="input-frames" class="frame-strip"></div></div> 83 </div> 84 <div class="reasoning-panel" aria-live="polite"> 85 <p class="eyebrow" id="example-type">OBJECT COUNTING</p> 86 <h3 id="scene-question">How many blackboards are in the scene?</h3> 87 <p class="example-note" id="example-note">A recorded example from the paper, paired with its reconstruction.</p> 88 <div class="response-heading"><span>SpatialSpeak response</span><button id="replay-response" aria-label="Replay the recorded model response">â» Replay</button></div><ol class="reasoning-steps" id="reasoning-steps"></ol> 89 <div class="answer-strip"><div><span id="answer-label">SpatialSpeak</span><strong id="answer-value">2 <small>blackboards</small></strong></div><div><span id="truth-label">Ground truth</span><strong id="truth-value">2</strong></div></div> 90 <p class="baseline-result" id="baseline-result">Without QA-RP and CoT-VC: <strong>3</strong></p> 91 <button class="text-link" id="view-example-figure">View the paper figure <span aria-hidden="true">â</span></button> 92 </div> 93 </div> 94 <noscript><p class="no-script-note">Enable JavaScript to rotate the point clouds and switch examples. The paper figures and research summary below remain available.</p></noscript> 95 </section> 96 97 98 <section class="metric-scenes section-width" id="measurements" aria-labelledby="metric-scenes-title"> 99 <div class="section-heading"><h2 id="metric-scenes-title">Metric Reconstruction</h2><p>The red arrows show exactly which span is measured.</p></div> 100 <div class="metric-scene-grid"> 101 <article class="metric-scene-card"><h3>scene0086_02</h3><div class="measurement-pair"><figure><figcaption>SpatialSpeak</figcaption><img class="paper-measurement" src="assets/figures/scene0086_02-ours.svg?v=eab4e9c64280" alt="SpatialSpeak: red arrows mark the 1.24-meter span" loading="lazy"></figure><figure><figcaption>Ground truth</figcaption><img class="paper-measurement" src="assets/figures/scene0086_02-gt.svg?v=72f4712054ec" alt="Ground truth: the same span is 1.25 meters" loading="lazy"></figure></div><p class="measurement-baselines">CUT3R: 1.38 m · MapAnything: 1.63 m</p><a class="text-link" href="#explore" data-select-scene="reconstruction">Explore this scene â</a></article> 102 <article class="metric-scene-card"><h3>scene0645_00</h3><div class="measurement-pair"><figure><figcaption>SpatialSpeak</figcaption><img class="paper-measurement" src="assets/figures/scene0645_00-ours.svg?v=905ba1710ac0" alt="SpatialSpeak: red arrows mark the 2.52-meter bed length" loading="lazy"></figure><figure><figcaption>Ground truth</figcaption><img class="paper-measurement" src="assets/figures/scene0645_00-gt.svg?v=b674e032e48b" alt="Ground truth: the same bed length is 2.54 meters" loading="lazy"></figure></div><p class="measurement-baselines">CUT3R: 2.82 m · MapAnything: 2.90 m</p><a class="text-link" href="#explore" data-select-scene="reconstruction-b">Explore this scene â</a></article> 103 </div> 104 <p class="measurement-note">Original measurement annotations from the paper. Switch to the 3D view to inspect the corresponding reconstruction.</p> 105 </section> 106 107 <section class="method-section" id="method" aria-labelledby="method-title"> 108 <div class="section-width"> 109 <div class="section-heading"><div><p class="eyebrow">TWO STAGES, ONE LANGUAGE INTERFACE</p><h2 id="method-title">Method</h2></div></div> 110 <div class="stage-grid"> 111 <article class="stage"><span class="stage-number">01</span><div><p class="stage-kicker">QA-RP · RECONSTRUCTION PRETRAINING</p><p>Marked-point 3D queries teach fine-grained local geometry. Object-center queries teach global scene context across views. Both use text-based question answering.</p></div></article> 112 <article class="stage"><span class="stage-number">02</span><div><p class="stage-kicker">CoT-VC · SPATIAL REASONING</p><p>Spatial chain-of-thought connects estimated geometry to the requested answer. Reliability assessment and visual compensation support refinement when needed.</p></div></article> 113 </div> 114 <figure class="method-figure"><button class="figure-button" data-figure="2" aria-label="Enlarge method overview"><img src="assets/figures/figure-2.svg?v=374b242655a4" alt="The SpatialSpeak framework: QA-RP pretraining combines local point and global object-center queries, followed by CoT-VC finetuning for spatial reasoning" loading="lazy" width="622" height="294"></button><figcaption>QA-native reconstruction pretraining connects local and global scene understanding with spatial chain-of-thought learning. <span class="figure-actions"><button class="text-link" data-figure="2">Enlarge â</button> <a class="text-link figure-pdf" href="assets/figures/figure-2.pdf" target="_blank" rel="noopener">View PDF â</a></span></figcaption></figure> 115 </div> 116 </section> 117 118 <section class="results section-width" id="results" aria-labelledby="results-title"> 119 <div class="section-heading"><div><p class="eyebrow">EXPERIMENTS</p><h2 id="results-title">Experiments</h2></div><p>Results reported in the paper.<br>Higher is better.</p></div> 120 <div class="metrics"><div><strong>
12062.8<span>%</span></strong><p>ReVSI average</p></div><div><strong>+8.7<span>pts</span></strong><p>Over the strongest compared baseline</p></div><div><strong>4<span>B</span></strong><p>VLM backbone parameters</p></div></div> 121 <div class="results-grid"> 122 <article class="benchmark-chart"><div class="chart-heading"><h3>ReVSI performance</h3><span>Average score (%)</span></div><div class="bar-chart" id="bar-chart"> 123 <div class="bar-row ours"><span>SpatialSpeak-4B</span><div class="bar-track"><i style="width:95.111%"></i></div><strong>62.8</strong></div> 124 <div class="bar-row"><span>SpatialStack-4B</span><div class="bar-track"><i style="width:75.778%"></i></div><strong>54.1</strong></div> 125 <div class="bar-row"><span>GeoThinker-8B</span><div class="bar-track"><i style="width:75.556%"></i></div><strong>54.0</strong></div> 126 <div class="bar-row"><span>VLM-3R-7B</span><div class="bar-track"><i style="width:66.889%"></i></div><strong>50.1</strong></div> 127 <div class="bar-row"><span>Cambrian-S-7B</span><div class="bar-track"><i style="width:64.667%"></i></div><strong>49.1</strong></div> 128 <div class="chart-axis" aria-hidden="true"><span>20</span><span>65</span></div> 129 </div><p class="chart-note">Displayed score range: 20â65%. Selected comparisons from Table 6. SpatialSpeak, SpatialStack, and GeoThinker are evaluated under the same 32-frame setting. VLM-3R and Cambrian-S results are sourced from the ReVSI paper.</p></article> 130 <article class="ablation"><p class="eyebrow">WHY RECONSTRUCTION MATTERS</p><h3>Spatial CoT gains more<br>after QA-RP.</h3><p>The improvement from CoT-VC grows from <strong>2.6</strong> to <strong>6.9 points</strong> when the model first learns reconstruction.</p><div class="gain-row"><span>Without QA-RP</span><div class="gain-track"><i style="width:37.681%"></i></div><strong>+2.6</strong></div><div class="gain-row highlight"><span>With QA-RP</span><div class="gain-track"><i style="width:100%"></i></div><strong>+6.9</strong></div><p class="chart-note">ReVSI ablation, Table 1. Baseline: 52.4. CoT-VC only: 55.0. QA-RP only: 55.9. Full method: 62.8.</p></article> 131 </div> 132 <div class="table-wrap"><table><caption>Spatial reasoning across benchmarks</caption><thead><tr><th scope="col">Benchmark</th><th scope="col">Setting</th><th scope="col">SpatialSpeak-4B</th><th scope="col">Best compared baseline</th></tr></thead><tbody><tr><th scope="row">ReVSI</th><td>32-frame evaluation</td><td class="score">62.8</td><td>54.1 <small>SpatialStack-4B</small></td></tr><tr><th scope="row">VSI-Bench</th><td>Normal training</td><td class="score">63.3</td><td>55.4 <small>Omni-View-7B</small></td></tr><tr><th scope="row">VSI-Bench</th><td>Scaled training</td><td class="score">73.0</td><td>72.6 <small>GeoThinker-8B</small></td></tr><tr><th scope="row">SPAR-Bench</th><td>Overall average</td><td class="score">76.0</td><td>72.0 <small>SpatialStack-4B</small></td></tr></tbody></table></div> 133 <p class="table-note">Scores and comparison settings follow Tables 6â9 of the paper. Normal and scaled training settings use different training data and should be read separately.</p> 134 </section> 135 136 137 138 <section class="abstract-section section-width" aria-labelledby="abstract-title"><h2 id="abstract-title">Abstract</h2><p>Vision-language models (VLMs) can benefit from geometric priors for multi-view spatial reasoning, yet answer-only training does not directly supervise the intermediate geometric estimates and their use in deriving quantitative spatial answers. We hypothesize that spatial chain-of-thought (CoT) supervision becomes more effective when the VLM first jointly learns complementary local geometry and global scene context through multi-view reconstruction.</p><p>We introduce <strong>SpatialSpeak</strong>, a two-stage framework that connects QA-native reconstruction pretraining with spatial CoT learning. In Stage I, QA-Native Reconstruction Pretraining (QA-RP) combines marked-point 3D queries for fine-grained local geometry with object-center queries for global scene context across views. Both tasks are formulated as text-based question answering, allowing geometric estimation and subsequent reasoning to share the same autoregressive output interface. In Stage II, spatial CoT with Visual Compensation (CoT-VC) trains the model to express question-relevant geometric estimates and use them to derive answers, with reliability assessment and visual compensation supporting answer refinement when needed.</p><p>On ReVSI, QA-RP increases the gain from CoT-VC from 2.6 to 6.9 points, and ablations show that both local and global reconstruction supervision are beneficial. SpatialSpeak achieves state-of-the-art results on ReVSI, VSI-Bench, and SPAR-Bench, with a ReVSI score of 62.8 that exceeds the strongest compared baseline by 8.7 points.</p></section> 139 140 <section class="citation-section section-width" id="citation" aria-labelledby="citation-title"><div class="section-heading"><div><p class="eyebrow">REFERENCE</p><h2 id="citation-title">Citation</h2></div><button class="button secondary" id="copy-citation">Copy BibTeX <span aria-hidden="true">â§</span></button></div><pre><code id="bibtex">@article{cao2026spatialspeak, 141 title = {SpatialSpeak: QA-Native Reconstruction with Local and Global 142 Context for Spatial Chain-of-Thought Reasoning},
143 author = {Cao, Yang and Zhang, Jiaxin and Chen, Dave Zhenyu and 144 Zhong, Yingji and Gao, Ruiyuan and Hong, Lanqing and Xu, Dan}, 145 journal = {arXiv preprint arXiv:2609.33616}, 146 year = {2026} 147}</code></pre><p class="citation-note">arXiv:2609.33616</p><span id="copy-status" class="visually-hidden" role="status"></span></section> 148</main> 149<footer class="section-width"><a class="wordmark" href="#top">SpatialSpeak<span class="title-period">.</span></a><p>QA-native reconstruction for spatial reasoning.</p></footer> 150<dialog id="figure-dialog" aria-labelledby="dialog-title"><div class="dialog-top"><h2 id="dialog-title">Paper figure</h2><a id="dialog-pdf" class="text-link" href="assets/figures/figure-1.pdf" target="_blank" rel="noopener">View original PDF â</a><button id="close-figure" aria-label="Close figure">Ã</button></div><img id="dialog-image" alt=""><p id="dialog-caption"></p></dialog> 151</body> 152</html>
Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.