1<!DOCTYPE html> 2<html xmlns="http://www.w3.org/1999/html"> 3 <head> 4 <meta charset="utf-8" /> 5 <meta 6 name="description" 7 content="S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Model with Spatio-Temporal Visual Representation 8" 9 /> 10 <meta name="google-site-verification" content="ckOPAEVbXcPaPfTw-55IBv8ONk4piVPU6rT_egFFEDc" /> 11 <meta name="keywords" content="Diffusion Model, Dexterous Manipulation, Robot Learning" /> 12 <meta name="viewport" content="width=device-width, initial-scale=1" /> 13 <title> 14 S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Model with Spatio-Temporal Visual Representation. 15 </title> 16 <link rel="preconnect" href="https://fonts.googleapis.com" /> 17 <link rel="preconnect" href="https://fonts.gstatic.com" crossorigin /> 18 <link 19 href="https://fonts.googleapis.com/css2?family=Roboto:wght@300;400;500;700&display=swap" 20 rel="stylesheet" 21 /> 22 <link href="./public/index.css" rel="stylesheet" /> 23 <link href="./public/media.css" rel="stylesheet" /> 24 <link href="./public/sidebars.css" rel="stylesheet" /> 25
25<script src="https://code.jquery.com/jquery-3.3.1.min.js"></script>
25 26
26<script src="./public/js/base.js"></script>
26 27 </head> 28 29 <body> 30 <div class="sidebarsWrapper"> 31 <div class="sidebars"> 32 <a class="barWrapper" clear href="#abstract-a" id="bar2" 33 ><span>Abstract</span> 34 <div class="bar"></div 35 ></a> 36 <a class="barWrapper" clear href="#methods-a" id="bar3" 37 ><span>Methods</span> 38 <div class="bar"></div 39 ></a> 40 <a class="barWrapper" clear href="#results-a" id="bar4" 41 ><span>Results</span> 42 <div class="bar"></div 43 ></a> 44 <a class="barWrapper" clear href="#visualization-a" id="bar5" 45 ><span>Visualizations</span> 46 <div class="bar"></div 47 ></a> 48 <a class="barWrapper" clear href="#citation" id="bar6" 49 ><span>Citation</span> 50 <div class="bar"></div 51 ></a> 52<!-- <a class="barWrapper" clear href="#attn-a" id="bar5"--> 53<!-- ><span>Attention Analysis</span>--> 54<!-- <div class="bar"></div--> 55<!-- ></a>--> 56 </div> 57 </div> 58 <main class="content"> 59 <section class="heading" style="text-align: center!important;"> 60 <h1 class="title"> 61 S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Model with Spatio-Temporal Visual Representation 62 </h1> 63 <section class="authors"> 64 <ul> 65 <li> 66 <span 67 ><a 68 href="https://yichen928.github.io/" 69 rel="noreferrer" 70 target="_blank" 71 >Yichen Xie</a 72 ><sup>1 * †</sup></span 73 > 74 </li> 75 <li> 76 <span 77 ><a 78 href="https://scholar.google.com/citations?user=QW6Ro8IAAAAJ&hl=zh-TW" 79 rel="noreferrer" 80 target="_blank" 81 >Runsheng Xu</a 82 ><sup>2 *</sup></span 83 > 84 </li> 85 <li> 86 <span 87 ><a 88 href="https://scholar.google.com/citations?user=v6o-fksAAAAJ&hl=zh-CN" 89 rel="noreferrer" 90 target="_blank" 91 >Tong He</a 92 ><sup>2</sup></span 93 > 94 </li> 95 <li> 96 <span 97 ><a 98 href="https://jyhjinghwang.github.io/" 99 rel="noreferrer" 100 target="_blank" 101 >Jyh-Jing Hwang</a 102 ><sup>2</sup></span 103 > 104 </li> 105 <li> 106 <span 107 ><a 108 href="https://www.cs.cornell.edu/~katieluo/" 109 rel="noreferrer" 110 target="_blank" 111 >Katie Z Luo</a 112 ><sup>3 †</sup></span 113 > 114 </li> 115 <li> 116 <span 117 ><a 118 href="https://jingweij.github.io/" 119 rel="noreferrer" 120 target="_blank" 121 >Jingwei Ji</a 122 ><sup>2</sup></span 123 > 124 </li> 125 <li> 126 <span 127 ><a 128 href="https://www.cs.cornell.edu/~hubert/" 129 rel="noreferrer" 130 target="_blank" 131 >Hubert Lin</a 132 ><sup>2</sup></span 133 > 134 </li> 135 <li> 136 <span 137 ><a 138 href="http://letianchen.me/" 139 rel="noreferrer" 140 target="_blank" 141 >Letian Chen</a 142 ><sup>4 †</sup></span 143 > 144 </li> 145 <li> 146 <span 147 ><a 148 href="https://luyiren.me/" 149 rel="noreferrer" 150 target="_blank" 151 >Yiren Lu</a 152 ><sup>2</sup></span 153 > 154 </li> 155 <li> 156 <span 157 ><a 158 href="https://scholar.google.com/citations?hl=en&user=tiCAVTQAAAAJ&view_op=list_works&sortby=pubdate" 159 rel="noreferrer" 160 target="_blank" 161 >
161Zhaoqi Leng</a 162 ><sup>2</sup></span 163 > 164 </li> 165 <li> 166 <span 167 ><a 168 href="https://scholar.google.com/citations?user=T04c3fwAAAAJ&hl=en" 169 rel="noreferrer" 170 target="_blank" 171 >Dragomir Anguelov</a 172 ><sup>2</sup></span 173 > 174 </li> 175 <li> 176 <span 177 ><a 178 href="https://scholar.google.com/citations?user=6POeyBoAAAAJ&hl=en" 179 rel="noreferrer" 180 target="_blank" 181 >Mingxing Tan</a 182 ><sup>2</sup></span 183 > 184 </li> 185 </ul> 186 </section> 187 <section class="affiliations"> 188 <ul> 189 <li><sup>1</sup>UC Berkeley,</li> 190 <li><sup>2</sup>Waymo LLC,</li> 191 <li><sup>3</sup>Cornell University,</li> 192 <li><sup>4</sup>Georgia Institute of Technology </li> 193 </ul> 194 </section> 195 <section class="equal"> 196 <p> 197 <sup>*</sup>Equal contribution 198 </p> 199 </section> 200 <section class="intern"> 201 <p> 202 <sup> †</sup>Work done as interns in Waymo 203 </p> 204 </section> 205 <section class="conference"> 206 <h3> 207 CVPR 2025 208 </h3> 209 </section> 210 <section class="logo"> 211 <br> 212 <div style="display: flex; width: 100%; height:auto; margin: auto; gap: 10%; justify-content: center; align-items: center"> 213 <img 214 style="width: 15%; height: auto" 215 src="./public/images/Waymo_logo.png" 216 /> 217 </div> 218 </section> 219 <section class="links"> 220 <ul> 221 <a href="https://arxiv.org/abs/2505.24139" rel="noreferrer" target="_blank"> 222 <li> 223 <span class="icon"> <img src="./public/paper.svg" /> </span 224 ><span>Paper</span> 225 </li> 226 </a> 227 </ul> 228 </section> 229 <a class="anchor" id="abstract-a"></a> 230 <h2>Abstract</h2> 231 <p class="abstract" style="font-family: 'Times New Roman', Arial; text-align: justify"> 232 The latest advancements in multi-modal large language models (MLLMs) have spurred a strong renewed interest in end-to-end motion planning approaches for autonomous driving. Many end-to-end approaches rely on human annotations to learn intermediate perception and prediction tasks, while purely self-supervised approachesâwhich directly learn from sensor inputs to generate planning trajectories without human annotationsâoften underperform the state of the art. 233 We observe a key gap in the input representation space: end-to-end approaches built on MLLMs are often pretrained with reasoning tasks in 2D image space rather than the native 3D space in which autonomous vehicles plan. 234 To this end, we propose S4-Driver, a <u>S</u>calable <u>S</u>elf-<u>S</u>upervised motion planning algorithm with <u>S</u>patio-temporal visual representation, based on the popular PaLI multimodal large language model. S4-Driver uses a novel sparse volume strategy to seamlessly transform the strong visual representation of MLLMs from perspective view to 3D space without the need to finetune the vision encoder. This representation aggregates multi-view and multi-frame visual inputs and enables better prediction of planning trajectories in 3D space. 235 To validate our method, we run experiments on both nuScenes and Waymo Open Motion Dataset (with in-house camera data). 236 Results show that S4-Driver performs favorably against existing supervised multi-task approaches while requiring no human annotations. It also demonstrates great scalability when pretrained on large volumes of unannotated driving logs. 237 </p> 238 </section> 239 240 <section class="methods" style="text-align: center!important;"> 241 <a class="anchor" id="methods-a"></a> 242 <br> 243 <h2>Self-Supervised Learning Framework</h2> 244 <br> 245 <div style="display: flex; margin: auto; width: 95%; height: auto; justify-content: space-between; margin-bottom: -15px; margin-top: -25px"> 246 <img style="width: 40%;height: auto;" src="./public/images/mtl.png"> 247 <img style="width: 40%;height: auto;" src="./public/images/ssl.png"> 248 </div> 249 <div class="col-title" style="display: flex; width: 95%; height: auto; justify-content: space-between; margin: 0; font-family: 'Times New Roman',serif"> 250 <p style="margin-left: 5em">Multi-task learning framework.</p> 251 <p style="margin-right: 2em">Self-supervised planning framework.</p> 252 </div> 253 <h2>Enhancing MLLMs for End-to-End Motion Planning</h2> 254 <br> 255 <div style="display: flex; margin: auto; width: 100%; height: auto; justify-content: space-between"> 256 <img style="width: 100%;height: auto;" src="./public/images/overview.png"> 257 </div> 258 <br> 259 <div style="display: flex; margin: auto; width: 100%; height: auto; justify-content: space-between; font-family: 'Times New Roman',serif"> 260 <img style="width: 60%;height: auto;" src="./public/images/roadmap_new.png"> 261 <p style="width: 50%;height: auto;margin-top: 100px;margin-left: 1em;text-align: left;"><u>Top:<br>Overview of S4-Driver frameworks.</u><br>We enhance the PaLI model for motion planning by incorporating meta-decision, spatio-temporal visual representation, and multi-decoding aggregation.<br><br><br><u>Left:<br>A roadmap for enhancing MLLM for planning.</u><br> We show the performance on Waymo Open Motion Dataset after including each module, while shadow items are not adopted in the subsequent steps.</p> 262 </div> 263 </section> 264 265 <section class="results" style="text-align: center!important;"> 266 <a class="anchor" id="results-a"></a> 267 <h2>Results</h2> 268 <p style="font-family: 'Times New Roman',serif;font-weight: bold">Results on nuScenes Dataset.</p> 269 <div style="display: flex; margin: auto; width: 70%; height: auto; justify-content: space-between; margin-top: -15px"> 270 <img style="width: 100%;height: auto;" src="./public/images/nuscenes.png"> 271 </div> 272 <br> 273 <p style="font-family: 'Times New Roman',serif;font-weight: bold">
273Results on Waymo Open Motion Dataset (with internal camera data).</p> 274 <div style="display: flex; margin: auto; width: 70%; height: auto; justify-content: space-between; margin-top: -15px"> 275 <img style="width: 100%;height: auto;" src="./public/images/womd.png"> 276 </div> 277 <p style="font-family: 'Times New Roman',serif;font-weight: bold">Data Scaling-up with Raw Driving Logs.</p> 278 <div style="display: flex; margin: auto; width: 50%; height: auto; justify-content: space-between; margin-top: -15px"> 279 <img style="width: 100%;height: auto;" src="./public/images/scale-up.png"> 280 </div> 281 282 <a class="anchor" id="visualization-a"></a> 283 <h2>Visualizations</h2> 284 <div class="col-title" style="display: flex; width:100%; height: auto; justify-content: space-between; margin: 0; font-family: 'Times New Roman',serif;font-weight: bold"> 285 <p style="margin-left: 9em">Extreme Weather.</p> 286 <p style="margin-right: 10em">Severe Shadow.</p> 287 </div> 288 <div style="display: flex; margin: auto; width: 100%; height: auto; justify-content: space-between; margin-top: -15px"> 289 <img style="width: 47%;height: auto;" src="./public/images/vis_snow.png"> 290 <img style="width: 47%;height: auto;" src="./public/images/vis_shadow.png"> 291 </div> 292 <p style="font-family: 'Times New Roman',serif;font-weight: bold">Reacting to Traffic Signals.</p> 293 <div style="display: flex; margin: auto; width: 100%; height: auto; justify-content: space-between; margin-top: -15px"> 294 <img style="width: 47%;height: auto;" src="./public/images/vis_red.png"> 295 <img style="width: 47%;height: auto;" src="./public/images/vis_green.png"> 296 </div> 297 <p style="font-family: 'Times New Roman',serif;font-weight: bold">Bad Lighting Condition.</p> 298 <div style="display: flex; margin: auto; width: 100%; height: auto; justify-content: space-between; margin-top: -15px"> 299 <img style="width: 47%;height: auto;" src="./public/images/vis_night1.png"> 300 <img style="width: 47%;height: auto;" src="./public/images/vis_night2.png"> 301 </div> 302 <p style="font-family: 'Times New Roman',serif;font-weight: bold">Turning.</p> 303 <div style="display: flex; margin: auto; width: 100%; height: auto; justify-content: space-between; margin-top: -15px"> 304 <img style="width: 47%;height: auto;" src="./public/images/vis_turn1.png"> 305 <img style="width: 47%;height: auto;" src="./public/images/vis_turn2.png"> 306 </div> 307 <div style="display: flex; margin: auto; width: 100%; height: auto; justify-content: space-between;"> 308 <img style="width: 47%;height: auto;" src="./public/images/vis_turn3.png"> 309 <img style="width: 47%;height: auto;" src="./public/images/vis_turn4.png"> 310 </div> 311 <p style="font-family: 'Times New Roman',serif;font-weight: bold">Keeping the Lane.</p> 312 <div style="display: flex; margin: auto; width: 100%; height: auto; justify-content: space-between; margin-top: -15px"> 313 <img style="width: 47%;height: auto;" src="./public/images/vis_lane1.png"> 314 <img style="width: 47%;height: auto;" src="./public/images/vis_lane2.png"> 315 </div> 316 <div style="display: flex; margin: auto; width: 100%; height: auto; justify-content: space-between;"> 317 <img style="width: 47%;height: auto;" src="./public/images/vis_lane3.png"> 318 <img style="width: 47%;height: auto;" src="./public/images/vis_lane4.png"> 319 </div> 320 </section> 321 322 <a class="anchor" id="citation"></a> 323 <section class="citation" style="text-align: justify;"> 324 <h2>Bibtex</h2> 325 <pre> 326<code>@InProceedings{xie2025s4driver, 327 title={S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Model with Spatio-Temporal Visual Representation}, 328 author={Xie, Yichen and Xu, Runsheng and He, Tong and Hwang, Jyh-Jing and Luo, Katie Z and Ji, Jingwei and Lin, Hubert and Chen, Letian and Lu, Yiren and Leng, Zhaoqi and Anguelov, Dragomir and Tan, Mingxing}, 329 booktitle = {IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, 330 year = {2025}, 331}</code></pre> 332 </section> 333 <br /> 334 </main> 335 </body> 336</html>
Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.