1<!DOCTYPE html> 2<html> 3 4<head> 5 <meta charset="utf-8"> 6 <meta name="description" content="VimTS: A Unified Video and Image Text Spotter for Enhancing the Cross-domain Generalization"> 7 <meta name="keywords" content="cross-domain text spotting"> 8 <meta name="viewport" content="width=device-width, initial-scale=1"> 9 <title>VimTS</title> 10 11 <link rel="stylesheet" href="https://fonts.googleapis.com/css?family=Google+Sans|Noto+Sans|Castoro"> 12 <link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/[email protected]/css/bulma.min.css"> 13 <link rel="stylesheet" href="https://maxcdn.bootstrapcdn.com/bootstrap/4.5.2/css/bootstrap.min.css"> 14 <link rel="stylesheet" href="https://cdn.jsdelivr.net/gh/jpswalsh/academicons@1/css/academicons.min.css"> 15 <link rel="stylesheet" href="https://cdnjs.cloudflare.com/ajax/libs/font-awesome/5.15.1/css/all.min.css"> 16 <link rel="stylesheet" href="./static/css/index.css"> 17 <link rel="icon" href="https://cdn-icons-png.flaticon.com/512/954/954591.png"> 18 <link href="https://fonts.googleapis.com/icon?family=Material+Icons" rel="stylesheet"> 19 20 21
21<script src="https://ajax.googleapis.com/ajax/libs/jquery/3.5.1/jquery.min.js"></script>
21 22
22<script defer src="https://cdnjs.cloudflare.com/ajax/libs/font-awesome/5.15.1/js/all.min.js"></script>
22 23
23<script type="module" src="https://gradio.s3-us-west-2.amazonaws.com/4.16.0/gradio.js"></script>
23 24</head> 25 26<style> 27 .expandable-card .card-text-container { 28 max-height: 200px; 29 overflow-y: hidden; 30 position: relative; 31 } 32 33 .expandable-card.expanded .card-text-container { 34 max-height: none; 35 } 36 37 .expand-btn { 38 position: relative; 39 display: none; 40 background-color: rgba(255, 255, 255, 0.8); 41 /* margin-top: -20px; */ 42 /* justify-content: center; */ 43 color: #510c75; 44 border-color: transparent; 45 } 46 47 .expand-btn:hover { 48 background-color: rgba(200, 200, 200, 0.8); 49 text-decoration: none; 50 border-color: transparent; 51 color: #510c75; 52 } 53 54 .expand-btn:focus { 55 outline: none; 56 text-decoration: none; 57 } 58 59 .expandable-card:not(.expanded) .card-text-container:after { 60 content: ""; 61 position: absolute; 62 bottom: 0; 63 left: 0; 64 width: 100%; 65 height: 90px; 66 background: linear-gradient(rgba(255, 255, 255, 0.2), rgba(255, 255, 255, 1)); 67 } 68 69 .expandable-card:not(.expanded) .expand-btn { 70 margin-top: -40px; 71 } 72 73 .card-body { 74 padding-bottom: 5px; 75 } 76 77 .vertical-flex-layout { 78 justify-content: center; 79 align-items: center; 80 height: 100%; 81 display: flex; 82 flex-direction: column; 83 gap: 5px; 84 } 85 86 .figure-img { 87 max-width: 100%; 88 height: auto; 89 } 90 91 .adjustable-font-size { 92 font-size: calc(0.5rem + 2vw); 93 } 94 95 .chat-history { 96 flex-grow: 1; 97 overflow-y: auto; 98 /* overflow-x: hidden; */ 99 padding: 5px; 100 border-bottom: 1px solid #ccc; 101 margin-bottom: 10px; 102 } 103 104 #gradio pre { 105 background-color: transparent; 106 } 107 108 .video-container { 109 display: flex; 110 justify-content: space-around; 111 } 112 113 .video { 114 flex: 1; 115 width: 400px; /* 设置åºå®å®½åº¦ */ 116 height: 400px; /* 设置åºå®é«åº¦ */ 117 } 118 119 .video video { 120 width: 100%; 121 height: 100%; 122 object-fit: cover; 123 } 124 125</style> 126 127 128 129 130<body> 131 132 <section class="hero"> 133 <div class="hero-body"> 134 <div class="container is-max-desktop"> 135 <div class="columns is-centered"> 136 <div class="column has-text-centered"> 137 <h1 class="title is-1 publication-title">VimTS: <span class="is-size-2"><span class="is-size-1">A</span> <span class="is-size-1"></span>U</span>nified <span class="is-size-1">V</span>ideo <span class="is-size-1">a</span>nd <span class="is-size-1">I</span>mage <span class="is-size-1">T</span>ext</span> <span class="is-size-1">S</span>potter</span> <span class="is-size-1">f</span>or <span class="is-size-1">E</span>nhancing</span> <span class="is-size-1">t</span>he <span class="is-size-1">C</span>ross-domain</span> <span class="is-size-1">G</span>eneralization</span></h1> 138 <div class="is-size-5 publication-authors"> 139 <span class="author-block"> 140 <a style="color:#141111;font-weight:normal;">Yuliang Liu</a>, 141 </span> 142 <span class="author-block"> 143 <a style="color:#141111;font-weight:normal;">Mingxin Huang</a>, 144 </span> 145 <span class="author-block"> 146 <a style="color:#141111;font-weight:normal;">Hao Yan</a>, 147 </span> 148 <span class="author-block"> 149 <a style="color:#141111;font-weight:normal;">Linger Deng</a>, 150 </span> 151 <span class="author-block"> 152 <a style="color:#141111;font-weight:normal;">Weijia Wu</a>, 153 </span> 154 <span class="author-block"> 155 <a style="color:#141111;font-weight:normal;">Hao Lu</a>, 156 </span> 157 <span class="author-block"> 158 <a style="color:#141111;font-weight:normal;">Chunhua Shen</a>, 159 </span> 160 <span class="author-block"> 161 <a style="color:#141111;font-weight:normal;">Lianwen Jin</a>, 162 </span> 163 <span class="author-block"> 164 <a style="color:#141111;font-weight:normal;">Xiang Bai</a> 165 </span> 166 </div> 167 168 <div class="is-size-5 publication-authors">
169 <span class="author-block"><b style="color:#f68946; font-weight:normal">▶ </b> Huazhong University of Science and Technology</b></span> 170 <span class="author-block"><b style="color:#008AD7; font-weight:normal">▶ </b> South China University of Technology</span> 171 <span class="author-block"><b style="color:#F2A900; font-weight:normal">▶ </b> Zhejiang University</span> 172 </div> 173 174 175 176 <!-- <div class="column has-text-centered"> 177 <h3 class="title is-3 publication-title">Improved Baselines with Visual Instruction Fine-tuning</h3> 178 <div class="is-size-5 publication-authors"> 179 <span class="author-block"> 180 <a href="https://hliu.cc/" style="color:#f68946;font-weight:normal;">Haotian Liu<sup>*</sup></a>, 181 </span> 182 <span class="author-block"> 183 <a href="https://chunyuan.li/" style="color:#008AD7;font-weight:normal;">Chunyuan Li<sup>*</sup></a>, 184 </span> 185 <span class="author-block"> 186 <a href="https://yuheng-li.github.io" style="color:#008AD7;font-weight:normal;">Yuheng Li</a>, 187 </span> 188 <span class="author-block"> 189 <a href="https://pages.cs.wisc.edu/~yongjaelee/" style="color:#f68946;font-weight:normal;">Yong Jae 190 Lee</a> 191 </span> 192 </div> 193 194 <div class="is-size-5 publication-authors"> 195 <span class="author-block"><b style="color:#f68946; font-weight:normal">▶ </b> University of 196 Wisconsin-Madison</b></span> 197 <span class="author-block"><b style="color:#008AD7; font-weight:normal">▶ </b> Microsoft Research</span> 198 </div> --> 199 200 <div class="column has-text-centered"> 201 <div class="publication-links"> 202 <span class="link-block"> 203 <a href="https://arxiv.org/abs/2404.19652" target="_blank" 204 class="external-link button is-normal is-rounded is-dark"> 205 <span class="icon"> 206 <i class="ai ai-arxiv"></i> 207 </span> 208 <span>arXiv</span> 209 </a> 210 </span> 211 <span class="link-block"> 212 <a href="https://github.com/Yuliang-Liu/VimTS" target="_blank" 213 class="external-link button is-normal is-rounded is-dark"> 214 <span class="icon"> 215 <i class="fab fa-github"></i> 216 </span> 217 <span>Code</span> 218 </a> 219 </span> 220 </div> 221 </div> 222 </div> 223 </div> 224 </div> 225 </div> 226 </section> 227 228 <section class="hero teaser"> 229 <div class="container is-max-desktop"> 230 <div class="hero-body"> 231 <h4 class="subtitle has-text-centered"> 232 ð¥<span style="color: #ff3860"> 233 VimTS is a unified video and image text spotter for enhancing the cross-domain generalization. It outperforms the state-of-the-art method by an average of 2.6% in six cross-domain benchmarks such as TT-to-IC15, CTW1500-to-TT, and TT-to-CTW1500. For video-level cross-domain adaption, our method even surpasses the previous end-to-end video spotting method in ICDAR2015 video and DSText v2 by an average of 5.5% on the MOTA metric, using only image-level data. 234 </h4> 235 </div> 236 </div> 237 </section> 238 239 240 241<section class="section"> 242 <div class="columns is-centered has-text-centered"> 243 <div class="column is-six-fifths"> 244 <h2 class="title is-3"><img id="painting_icon" width="3%" src="https://cdn-icons-png.flaticon.com/512/5886/5886212.png"> Video</h2> 245 </div> 246 </div> 247 248 249 <div class="video-container"> 250 <div class="video"> 251 <video class="my-video video1" autoplay loop muted> 252 <source src="images/1.mp4" type="video/mp4"> 253 Your browser does not support the video tag. 254 </video> 255 </div> 256 <div class="video"> 257 <video class="my-video video1" autoplay loop muted> 258 <source src="images/2.mp4" type="video/mp4"> 259 Your browser does not support the video tag. 260 </video> 261 </div> 262 <div class="video"> 263 <video class="my-video video1" autoplay loop muted> 264 <source src="images/3.mp4" type="video/mp4"> 265 Your browser does not support the video tag. 266 </video> 267 </div> 268</div> 269 270 271 272 <!-- <div class="video-container"> 273 <video class="my-video video1" autoplay loop muted> 274 <source src="images/1.mp4" type="video/mp4"> 275 Your browser does not support the video tag. 276 </video> 277 </div> 278 279 <div class="video-container"> 280 <video class="my-video video1" autoplay loop muted> 281 <source src="images/2.mp4" type="video/mp4"> 282 Your browser does not support the video tag. 283 </video> 284 </div> 285 286 <div class="video-container"> 287 <video class="my-video video1" autoplay loop muted> 288 <source src="images/3.mp4" type="video/mp4"> 289 Your browser does not support the video tag. 290 </video> 291 </div> --> 292</section> 293 294 295<section class="section"> 296 <!-- Results. --> 297 <div class="columns is-centered has-text-centered"> 298 <div class="column is-six-fifths"> 299 <h2 class="title is-3"><img id="painting_icon" width="3%" src="https://cdn-icons-png.flaticon.com/512/5886/5886212.png"> Framework</h2> 300 </div> 301 </div> 302 <!-- </div> --> 303 <!--/ Results. --> 304<div class="container is-max-desktop"> 305 306 307 <div class="columns is-centered"> 308 <div class="column is-full-width"> 309 <div class="content has-text-justified"> 310 <p> 311 Overall framework of our method. 312 <figure style="text-align: center;"> 313 <img id="teaser" width="100%" src="images/framework.png"> 314 </figure> 315 </p> 316 <p> 317 Overall framework of CoDeF-based synthetic method. 318 <figure style="text-align: center;"> 319 <img id="teaser" width="80%" src="images/flowtext.png"> 320 </figure> 321 </p> 322<!-- CSS Code: Place this code in the document's head (between the 'head' tags) --> 323 </div> 324 </div> 325 </div> 326</section> 327 328<section class="section"> 329 <!-- Results. --> 330 <div class="columns is-centered has-text-centered"> 331 <div class="column is-six-fifths"> 332 <h2 class="title is-3"><img id="painting_icon" width="3%" src="https://cdn-icons-png.flaticon.com/512/5886/5886212.png"> VTD-368k</h2> 333 </div> 334 </div> 335 <!-- </div> --> 336 <!--/ Results. --> 337<div class="container is-max-desktop"> 338 339 <div class="columns is-centered"> 340 <div class="column is-full-width"> 341 <div class="content has-text-justified"> 342 <p> 343 We manually collect and filter text-free, open-source and unrestricted videos from NExT-QA, Charades-Ego, Breakfast, A2D, MPI-Cooking, ActorShift and Hollywood. By utilizing the CoDeF, our synthetic method facilitates the achievement of realistic and stable text flow propagation, significantly reducing the occurrence of distortions. 344 <figure style="text-align: center;"> 345 <img id="teaser" width="100%" src="images/VTD-368K.png"> 346 </figure> 347<!-- CSS Code: Place this code in the document's head (between the 'head' tags) --> 348 </div> 349 </div> 350 </div> 351</section> 352 353<section class="section"> 354 <div class="container is-max-desktop"> 355 <!-- Abstract. --> 356 <div class="columns is-centered has-text-centered"> 357 <div class="column is-four-fifths"> 358 <h2 class="title is-3">Benchmark Experiments</h2> 359 <div class="content has-text-justified"> 360 <p> 361 For image-level cross-domain text spotting, we conduct experiments on six cross-domain scenarios to evaluate VimTS. 362 For video-level cross-domain text spotting, we conduct experiments on two popular video text spotting benchmarks to evaluate VimTS. 363 The results are presented in the following. 364 </p> 365 <p> 366 <div align=center> 367 <img src="images/cd_e2e.png" style="width: 100%" alt> 368 </div> 369 <div align=center> 370 <img src="images/cd_video.png" style="width: 100%" alt> 371 </div> 372 <div align=center> 373 <img src="images/cp_mlmms.png" style="width: 55%" alt> 374 </div> 375 </div> 376 377 378 </p> 379 <ul> 380 <li> 381 <b>Conclusion: It is worth mentioning that our method demonstrates it is viable that still text images can be 382 learned to be well transferred to video text images. Since still images require significantly less annotation 383 effort compared to video image, exploring methods to bridge the domain gaps will be highly valuable. 384 Furthermore, we demonstrate that current Large Multimodal Models still face limitations in cross-domain text 385 spotting. Using fewer parameters and less data to improve the generalization of Large Multimodal Models in 386 text spotting is worth further exploration. 387 </li> 388 </ul> 389 </div> 390 </div> 391 </div> 392</section> 393 394 395 396<section class="section"> 397 <!-- Results. --> 398 <div class="columns is-centered has-text-centered"> 399 <div class="column is-six-fifths"> 400 <h2 class="title is-3"><img id="painting_icon" width="3%" src="https://cdn-icons-png.flaticon.com/512/5886/5886212.png"> Some Visualization</h2> 401 </div> 402 </div> 403 <!-- </div> --> 404 <!--/ Results. --> 405<div class="container is-max-desktop"> 406 407 <div class="columns is-centered"> 408 <div class="column is-full-width"> 409 <div align=center> 410 <img src="images/vis.png" style="width: 70%" alt> 411 </div> 412 </div> 413<!-- CSS Code: Place this code in the document's head (between the 'head' tags) --> 414 </div> 415 </div> 416 </div> 417</section> 418 419<section class="section"> 420 <!-- Results. --> 421 <div class="columns is-centered has-text-centered"> 422 <div class="column is-six-fifths"> 423 <h2 class="title is-3"><img id="painting_icon" width="3%" src="https://cdn-icons-png.flaticon.com/512/5886/5886212.png"> Compared with MLMMS</h2> 424 </div> 425 </div> 426 <!-- </div> --> 427 <!--/ Results. --> 428 <div class="container is-max-desktop"> 429 430 <div class="columns is-centered"> 431 <div class="column is-full-width"> 432 <div align=center> 433 <img src="images/cp_mlmms_imgs_2.png" style="width: 70%" alt> 434 </div> 435 </div> 436 <!-- CSS Code: Place this code in the document's head (between the 'head' tags) --> 437 </div> 438 </div> 439 </div> 440 </section> 441 442 443 <section class="section" id="BibTeX"> 444 <div class="container is-max-desktop content"> 445 <h2 class="title">BibTeX</h2> 446 <pre><code> 447 @misc{liuvimts, 448 author={Liu, Yuliang and Huang, Mingxin and Yan, Hao and Deng, Linger and Wu, Weijia and Lu, Hao and Shen, Chunhua and Jin, Lianwen and Bai, Xiang}, 449 title={VimTS: A Unified Video and Image Text Spotter for Enhancing the Cross-domain Generalization}, 450 publisher={arXiv preprint arXiv:2404.19652}, 451 year={2024}, 452 } 453 </code></pre> 454 </div> 455 </section> 456 457</body> 458 459</html>
Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.