PageSourceSearch

https://s4-driver.github.io/

html s4-driver.github.io collected 2026-10-03 09:53:05 UTC 16,153 bytes, 336 lines download raw bytes

1<!DOCTYPE html>
2<html xmlns="http://www.w3.org/1999/html">
3  <head>
4    <meta charset="utf-8" />
5    <meta
6      name="description"
7      content="S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Model with Spatio-Temporal Visual Representation
8"
9    />
10    <meta name="google-site-verification" content="ckOPAEVbXcPaPfTw-55IBv8ONk4piVPU6rT_egFFEDc" />
11    <meta name="keywords" content="Diffusion Model, Dexterous Manipulation, Robot Learning" />
12    <meta name="viewport" content="width=device-width, initial-scale=1" />
13    <title>
14      S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Model with Spatio-Temporal Visual Representation.
15    </title>
16    <link rel="preconnect" href="https://fonts.googleapis.com" />
17    <link rel="preconnect" href="https://fonts.gstatic.com" crossorigin />
18    <link
19      href="https://fonts.googleapis.com/css2?family=Roboto:wght@300;400;500;700&display=swap"
20      rel="stylesheet"
21    />
22    <link href="./public/index.css" rel="stylesheet" />
23    <link href="./public/media.css" rel="stylesheet" />
24    <link href="./public/sidebars.css" rel="stylesheet" />
25    
25<script src="https://code.jquery.com/jquery-3.3.1.min.js"></script>
25
26    
26<script src="./public/js/base.js"></script>
26
27  </head>
28
29  <body>
30    <div class="sidebarsWrapper">
31      <div class="sidebars">
32        <a class="barWrapper" clear href="#abstract-a" id="bar2"
33          ><span>Abstract</span>
34          <div class="bar"></div
35        ></a>
36        <a class="barWrapper" clear href="#methods-a" id="bar3"
37          ><span>Methods</span>
38          <div class="bar"></div
39        ></a>
40        <a class="barWrapper" clear href="#results-a" id="bar4"
41          ><span>Results</span>
42          <div class="bar"></div
43        ></a>
44        <a class="barWrapper" clear href="#visualization-a" id="bar5"
45          ><span>Visualizations</span>
46          <div class="bar"></div
47        ></a>
48        <a class="barWrapper" clear href="#citation" id="bar6"
49          ><span>Citation</span>
50          <div class="bar"></div
51        ></a>
52<!--        <a class="barWrapper" clear href="#attn-a" id="bar5"-->
53<!--          ><span>Attention Analysis</span>-->
54<!--          <div class="bar"></div-->
55<!--        ></a>-->
56      </div>
57    </div>
58    <main class="content">
59      <section class="heading" style="text-align: center!important;">
60        <h1 class="title">
61          S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Model with Spatio-Temporal Visual Representation
62        </h1>
63        <section class="authors">
64          <ul>
65            <li>
66              <span
67                ><a
68                  href="https://yichen928.github.io/"
69                  rel="noreferrer"
70                  target="_blank"
71              >Yichen Xie</a
72                ><sup>1 * &#8224;</sup></span
73              >
74            </li>
75            <li>
76              <span
77                ><a
78                  href="https://scholar.google.com/citations?user=QW6Ro8IAAAAJ&hl=zh-TW"
79                  rel="noreferrer"
80                  target="_blank"
81              >Runsheng Xu</a
82                ><sup>2 *</sup></span
83              >
84            </li>
85            <li>
86              <span
87                ><a
88                  href="https://scholar.google.com/citations?user=v6o-fksAAAAJ&hl=zh-CN"
89                  rel="noreferrer"
90                  target="_blank"
91              >Tong He</a
92                ><sup>2</sup></span
93              >
94            </li>
95            <li>
96              <span
97                ><a
98                  href="https://jyhjinghwang.github.io/"
99                  rel="noreferrer"
100                  target="_blank"
101                  >Jyh-Jing Hwang</a
102                ><sup>2</sup></span
103              >
104            </li>
105            <li>
106              <span
107                ><a
108                  href="https://www.cs.cornell.edu/~katieluo/"
109                  rel="noreferrer"
110                  target="_blank"
111                  >Katie Z Luo</a
112                ><sup>3 &#8224;</sup></span
113              >
114            </li>
115            <li>
116              <span
117                ><a
118                  href="https://jingweij.github.io/"
119                  rel="noreferrer"
120                  target="_blank"
121                  >Jingwei Ji</a
122                ><sup>2</sup></span
123              >
124            </li>
125            <li>
126              <span
127                ><a
128                  href="https://www.cs.cornell.edu/~hubert/"
129                  rel="noreferrer"
130                  target="_blank"
131                  >Hubert Lin</a
132                ><sup>2</sup></span
133              >
134            </li>
135            <li>
136              <span
137                ><a
138                  href="http://letianchen.me/"
139                  rel="noreferrer"
140                  target="_blank"
141                  >Letian Chen</a
142                ><sup>4 &#8224;</sup></span
143              >
144            </li>
145            <li>
146              <span
147                ><a
148                  href="https://luyiren.me/"
149                  rel="noreferrer"
150                  target="_blank"
151                  >Yiren Lu</a
152                ><sup>2</sup></span
153              >
154            </li>
155            <li>
156              <span
157                ><a
158                  href="https://scholar.google.com/citations?hl=en&user=tiCAVTQAAAAJ&view_op=list_works&sortby=pubdate"
159                  rel="noreferrer"
160                  target="_blank"
161                  >
161Zhaoqi Leng</a
162                ><sup>2</sup></span
163              >
164            </li>
165            <li>
166              <span
167                ><a
168                  href="https://scholar.google.com/citations?user=T04c3fwAAAAJ&hl=en"
169                  rel="noreferrer"
170                  target="_blank"
171                  >Dragomir Anguelov</a
172                ><sup>2</sup></span
173              >
174            </li>
175            <li>
176              <span
177                ><a
178                  href="https://scholar.google.com/citations?user=6POeyBoAAAAJ&hl=en"
179                  rel="noreferrer"
180                  target="_blank"
181                  >Mingxing Tan</a
182                ><sup>2</sup></span
183              >
184            </li>
185          </ul>
186        </section>
187        <section class="affiliations">
188          <ul>
189            <li><sup>1</sup>UC Berkeley,</li>
190            <li><sup>2</sup>Waymo LLC,</li>
191            <li><sup>3</sup>Cornell University,</li>
192            <li><sup>4</sup>Georgia Institute of Technology            </li>
193          </ul>
194        </section>
195        <section class="equal">
196          <p>
197            <sup>*</sup>Equal contribution
198          </p>
199        </section>
200        <section class="intern">
201          <p>
202            <sup> &#8224</sup>Work done as interns in Waymo
203          </p>
204        </section>
205        <section class="conference">
206          <h3>
207          CVPR 2025
208          </h3>
209        </section>
210        <section class="logo">
211          <br>
212          <div style="display: flex; width: 100%; height:auto; margin: auto; gap: 10%; justify-content: center; align-items: center">
213            <img
214            style="width: 15%; height: auto"
215            src="./public/images/Waymo_logo.png"
216            />
217          </div>
218        </section>
219        <section class="links">
220          <ul>
221            <a href="https://arxiv.org/abs/2505.24139" rel="noreferrer" target="_blank">
222              <li>
223                <span class="icon"> <img src="./public/paper.svg" /> </span
224                ><span>Paper</span>
225              </li>
226            </a>
227          </ul>
228        </section>
229        <a class="anchor" id="abstract-a"></a>
230        <h2>Abstract</h2>
231        <p class="abstract" style="font-family: 'Times New Roman', Arial; text-align: justify">
232          The latest advancements in multi-modal large language models (MLLMs) have spurred a strong renewed interest in end-to-end motion planning approaches for autonomous driving. Many end-to-end approaches rely on human annotations to learn intermediate perception and prediction tasks, while purely self-supervised approaches—which directly learn from sensor inputs to generate planning trajectories without human annotations—often underperform the state of the art.
233          We observe a key gap in the input representation space: end-to-end approaches built on MLLMs are often pretrained with reasoning tasks in 2D image space rather than the native 3D space in which autonomous vehicles plan.
234          To this end, we propose S4-Driver, a <u>S</u>calable <u>S</u>elf-<u>S</u>upervised motion planning algorithm with <u>S</u>patio-temporal visual representation,  based on the popular PaLI multimodal large language model. S4-Driver uses a novel sparse volume strategy to seamlessly transform the strong visual representation of MLLMs from perspective view to 3D space without the need to finetune the vision encoder. This representation aggregates multi-view and multi-frame visual inputs and enables better prediction of planning trajectories in 3D space.
235          To validate our method, we run experiments on both nuScenes and Waymo Open Motion Dataset (with in-house camera data).
236          Results show that S4-Driver performs favorably against existing supervised multi-task approaches while requiring no human annotations. It also demonstrates great scalability when pretrained on large volumes of unannotated driving logs.
237        </p>
238      </section>
239
240      <section class="methods" style="text-align: center!important;">
241        <a class="anchor" id="methods-a"></a>
242        <br>
243        <h2>Self-Supervised Learning Framework</h2>
244        <br>
245        <div style="display: flex; margin: auto; width: 95%; height: auto; justify-content: space-between; margin-bottom: -15px; margin-top: -25px">
246          <img style="width: 40%;height: auto;" src="./public/images/mtl.png">
247          <img style="width: 40%;height: auto;" src="./public/images/ssl.png">
248        </div>
249        <div class="col-title" style="display: flex; width: 95%; height: auto; justify-content: space-between; margin: 0; font-family: 'Times New Roman',serif">
250          <p style="margin-left: 5em">Multi-task learning framework.</p>
251          <p style="margin-right: 2em">Self-supervised planning framework.</p>
252        </div>
253        <h2>Enhancing MLLMs for End-to-End Motion Planning</h2>
254        <br>
255        <div style="display: flex; margin: auto; width: 100%; height: auto; justify-content: space-between">
256          <img style="width: 100%;height: auto;" src="./public/images/overview.png">
257        </div>
258        <br>
259        <div style="display: flex; margin: auto; width: 100%; height: auto; justify-content: space-between; font-family: 'Times New Roman',serif">
260          <img style="width: 60%;height: auto;" src="./public/images/roadmap_new.png">
261          <p style="width: 50%;height: auto;margin-top: 100px;margin-left: 1em;text-align: left;"><u>Top:<br>Overview of S4-Driver frameworks.</u><br>We enhance the PaLI model for motion planning by incorporating meta-decision, spatio-temporal visual representation, and multi-decoding aggregation.<br><br><br><u>Left:<br>A roadmap for enhancing MLLM for planning.</u><br> We show the performance on Waymo Open Motion Dataset after including each module, while shadow items are not adopted in the subsequent steps.</p>
262        </div>
263      </section>
264
265      <section class="results" style="text-align: center!important;">
266      <a class="anchor" id="results-a"></a>
267      <h2>Results</h2>
268      <p style="font-family: 'Times New Roman',serif;font-weight: bold">Results on nuScenes Dataset.</p>
269      <div style="display: flex; margin: auto; width: 70%; height: auto; justify-content: space-between; margin-top: -15px">
270        <img style="width: 100%;height: auto;" src="./public/images/nuscenes.png">
271      </div>
272      <br>
273      <p style="font-family: 'Times New Roman',serif;font-weight: bold">
273Results on Waymo Open Motion Dataset (with internal camera data).</p>
274      <div style="display: flex; margin: auto; width: 70%; height: auto; justify-content: space-between; margin-top: -15px">
275        <img style="width: 100%;height: auto;" src="./public/images/womd.png">
276      </div>
277      <p style="font-family: 'Times New Roman',serif;font-weight: bold">Data Scaling-up with Raw Driving Logs.</p>
278      <div style="display: flex; margin: auto; width: 50%; height: auto; justify-content: space-between; margin-top: -15px">
279        <img style="width: 100%;height: auto;" src="./public/images/scale-up.png">
280      </div>
281
282      <a class="anchor" id="visualization-a"></a>
283      <h2>Visualizations</h2>
284      <div class="col-title" style="display: flex; width:100%; height: auto; justify-content: space-between; margin: 0; font-family: 'Times New Roman',serif;font-weight: bold">
285        <p style="margin-left: 9em">Extreme Weather.</p>
286        <p style="margin-right: 10em">Severe Shadow.</p>
287      </div>
288      <div style="display: flex; margin: auto; width: 100%; height: auto; justify-content: space-between; margin-top: -15px">
289        <img style="width: 47%;height: auto;" src="./public/images/vis_snow.png">
290        <img style="width: 47%;height: auto;" src="./public/images/vis_shadow.png">
291      </div>
292      <p style="font-family: 'Times New Roman',serif;font-weight: bold">Reacting to Traffic Signals.</p>
293      <div style="display: flex; margin: auto; width: 100%; height: auto; justify-content: space-between; margin-top: -15px">
294        <img style="width: 47%;height: auto;" src="./public/images/vis_red.png">
295        <img style="width: 47%;height: auto;" src="./public/images/vis_green.png">
296      </div>
297      <p style="font-family: 'Times New Roman',serif;font-weight: bold">Bad Lighting Condition.</p>
298      <div style="display: flex; margin: auto; width: 100%; height: auto; justify-content: space-between; margin-top: -15px">
299        <img style="width: 47%;height: auto;" src="./public/images/vis_night1.png">
300        <img style="width: 47%;height: auto;" src="./public/images/vis_night2.png">
301      </div>
302      <p style="font-family: 'Times New Roman',serif;font-weight: bold">Turning.</p>
303      <div style="display: flex; margin: auto; width: 100%; height: auto; justify-content: space-between; margin-top: -15px">
304        <img style="width: 47%;height: auto;" src="./public/images/vis_turn1.png">
305        <img style="width: 47%;height: auto;" src="./public/images/vis_turn2.png">
306      </div>
307      <div style="display: flex; margin: auto; width: 100%; height: auto; justify-content: space-between;">
308        <img style="width: 47%;height: auto;" src="./public/images/vis_turn3.png">
309        <img style="width: 47%;height: auto;" src="./public/images/vis_turn4.png">
310      </div>
311      <p style="font-family: 'Times New Roman',serif;font-weight: bold">Keeping the Lane.</p>
312      <div style="display: flex; margin: auto; width: 100%; height: auto; justify-content: space-between; margin-top: -15px">
313        <img style="width: 47%;height: auto;" src="./public/images/vis_lane1.png">
314        <img style="width: 47%;height: auto;" src="./public/images/vis_lane2.png">
315      </div>
316      <div style="display: flex; margin: auto; width: 100%; height: auto; justify-content: space-between;">
317        <img style="width: 47%;height: auto;" src="./public/images/vis_lane3.png">
318        <img style="width: 47%;height: auto;" src="./public/images/vis_lane4.png">
319      </div>
320      </section>
321
322      <a class="anchor" id="citation"></a>
323      <section class="citation" style="text-align: justify;">
324        <h2>Bibtex</h2>
325        <pre>
326<code>@InProceedings{xie2025s4driver,
327  title={S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Model with Spatio-Temporal Visual Representation},
328  author={Xie, Yichen and Xu, Runsheng and He, Tong and Hwang, Jyh-Jing and Luo, Katie Z and Ji, Jingwei and Lin, Hubert and Chen, Letian and Lu, Yiren and Leng, Zhaoqi and Anguelov, Dragomir and Tan, Mingxing},
329  booktitle = {IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
330  year      = {2025},
331}</code></pre>
332      </section>
333      <br />
334    </main>
335  </body>
336</html>

Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.