1 2<!doctype html> 3<html lang="en"> 4<head> 5 <meta charset="utf-8"> 6 <meta name="viewport" content="width=device-width, initial-scale=1, shrink-to-fit=no"> 7 <meta name="google-site-verification" content="eGDOZ_6azobM9Vcl7r072IFo1FJ-TfNvGkmz6YbLCLo" /> 8 9 <!-- Preload critical resources --> 10 <link rel="preload" href="https://fonts.cdnfonts.com/css/corbel" as="style"> 11 <link rel="preload" href="html_pages/resources/bootstrap.min.css" as="style"> 12 13 <!-- Font imports --> 14 <link href="https://fonts.cdnfonts.com/css/corbel" rel="stylesheet"> 15 16 <!-- CSS --> 17 <link href="html_pages/resources/bootstrap.min.css" rel="stylesheet"> 18 <link href="html_pages/resources/main.css" rel="stylesheet"> 19 20 <!-- JavaScript --> 21
21<script src="js/jquery-2.1.3.min.js" defer></script>
21 22
22<script src="html_pages/resources/main.js" defer></script>
22 23 24 <title id="title">Video Thinking Test</title> 25</head> 26 27<body data-new-gr-c-s-check-loaded="14.1110.0" data-gr-ext-installed=""> 28 <header> 29 <nav> 30 <a class="h7 pt-10" style="margin: 5px; white-space: nowrap; font-weight: 900; font-size: 1.7rem;" href="#">Video-TT</a> 31 </nav> 32 <nav> 33 <a href="#abstract">Abstract</a> 34 <a href="#video">Demo</a> 35 <a href="#annotation">Annotation</a> 36 <a href="#statistic">Statistic</a> 37 <a href="#performance">Performance</a> 38 <a href="#acknowledgement">Acknowledgement</a> 39 </nav> 40 </header> 41 42 <section class="jumbotron text-center pb-2" id="Video-TT"> 43 <div class="video-background"> 44 <video playsinline="playsinline" autoplay="autoplay" muted="muted" loop="loop"> 45 <source src="assets/header_video.mp4" type="video/mp4"> 46 </video> 47 </div> 48 <div class="container"> 49 <br><br><br> 50 <h1 class="jumbotron-heading" style="font-size: 6rem; font-weight: 900;">Video Thinking Test</h1> 51 <h5 class="pt-1" style="font-size: 2rem; font-weight: normal">A Holistic Benchmark for Advanced Video Reasoning and Understanding </h5> 52 <br> 53 <a href="https://zhangyuanhan-ai.github.io/" target="_blank" rel="noopener noreferrer" style="font-size: 1.2rem; color: white !important;">Yuanhan Zhang*<a>, 54 <a style="font-size: 1.2rem; color: white !important;">Yunice Chew*</a>, 55 <a href="https://scholar.google.com/citations?user=kMui170AAAAJ&hl=zh-CN" target="_blank" rel="noopener noreferrer" style="font-size: 1.2rem; color: white !important;">Yuhao Dong</a>, 56 <a style="font-size: 1.2rem; color: white !important;"> Aria Leo</a>, 57 <a style="font-size: 1.2rem; color: white !important;"> Bo Hu</a>, 58 <a href="https://liuziwei7.github.io/" target="_blank" rel="noopener noreferrer" style="font-size: 1.2rem; color: white !important;">Ziwei Liu</a> 59 <br><br> 60 <a style="font-size: 1.2rem; color: white !important;"><i>* Equal contribution</i></a> 61 <br><br> 62 <a style="font-size: 1.2rem; color: white !important;"><i>ICCV 2025</i></a> 63 <br><br> 64 <a href="https://liuziwei7.github.io/team.html" target="_blank" rel="noopener noreferrer"> 65 <div style="display: inline-block; background: radial-gradient(ellipse at center, rgba(20, 20, 20, 1) 0%, rgba(20, 20, 20, 0) 80%); border-radius: 20px; padding: 10px;"> 66 <img class="paper-btn" src="assets/ntu.png" style="height: 80px; width: auto; display: block;"> 67 </div> 68 </a> 69 <br><br><br> 70 <div style="display: flex; justify-content: center; gap: 15px; margin-top: 10px;"> 71 <!-- Paper Button --> 72 <a href="https://arxiv.org/abs/2507.15028" target="_blank" rel="noopener noreferrer" class="jumbotron-button"> 73 <img src="assets/arxiv.svg" alt="" style="height: 20px; width: auto;"> 74 Paper 75 </a> 76 <!-- Dataset Button --> 77 <a href="https://huggingface.co/datasets/lmms-lab/video-tt" target="_blank" rel="noopener noreferrer" class="jumbotron-button"> 78 <img src="assets/huggingface.png" alt="" style="height: 20px; width: auto;"> 79 Dataset 80 </a> 81 <!-- Example Button --> 82 <a href="more_samples.html" rel="noopener noreferrer" class="jumbotron-button"> 83 <img src="assets/movie.svg" alt="" style="height: 20px; width: auto;"> 84 Examples 85 </a> 86 </div> 87 </div> 88 89 </section> 90 91 92 93 <section id="abstract"> 94 <div class="container"> 95 <h1 class="jumbotron-heading">Abstract</h1> 96 <div style="justify-content: space-around; flex-wrap: wrap; margin-top: 30px"> 97 <div class="image-item"> 98 <img src="assets/teaser.png" class="zoomable-image" style="margin-top: -20px; width: 100%; max-width: 8000px; height: auto;"> 99 </div> 100 </div> 101 102 <p>We introduce <b>the Video Thinking Test (Video-TT)</b>, a benchmark designed to assess if video LLMs can interpret real-world videos as effectively as humans. <b>Video-TT</b> <b>1)</b> differentiates between errors due to inadequate frame sampling and genuine gaps in understanding complex visual narratives, and <b>2)</b> evaluates robustness against natural adversarial questions. <b>Video-TT</b> comprises 1,000 YouTube Shorts videos, each with one open-ended question and four adversarial questions that probe visual and narrative complexity. Our evaluation shows a significant gap between video LLMs and human performance, underscoring the need for benchmarks like <b>Video-TT</b> to advance video understanding.</p> 103 </div> 104 105 </section> 106 107 <section id="video" style="margin-top: 20px;"> 108 <div class="container"> 109 <h1 class="jumbotron-heading">Video Demo</h1> 110 <div class="youtube-container" style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden; max-width: 100%; height: auto;"> 111 <iframe 112 src="https://www.youtube.com/embed/vjL-munUong" 113 title="YouTube video player" 114 frameborder="0" 115 allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" 116 allowfullscreen 117 style="position: absolute; top: 0; left: 0; width: 100%; height: 100%;"> 118 </iframe> 119 </div> 120 </div> 121 122 </section> 123 124 <section id="annotation" style="margin-top: 20px;"> 125 <div class="container"> 126 <h1 class="jumbotron-heading">Data Annotation Pipeline</h1> 127 128 <div style="display: flex; justify-content: center; flex-wrap: wrap; margin-top: 30px;"> 129 <div class="image-item"> 130 <img 131 src="assets/annotation_pipeline.png" 132 alt="Data Annotation Pipeline Diagram" 133 class="zoomable-image" 134 style="margin-top: -20px; width: 100%; height: auto;" 135 > 136 </div> 137 </div> 138 139 <div style="text-align: left; margin-top: 40px;"> 140 <p><strong>The dataset is built through a multi-step annotation and verification pipeline:</strong></p> 141 <ul class="info-list"> 142 <li> 143 <strong style="color: #E97132;">Ensuring Complexity:</strong> (Visual Complexity) Measures how visually challenging a video is, based on factors like unclear or unusual content, fast motion, complex object arrangements, and visual illusions that hinder recognition. (Narrative Complexity) Reflects how cognitively demanding the storyline is, including elements like plot twists, montage-style editing, subtle technical manipulations, and reliance on world knowledge for full comprehension. 144 </li> 145 <li> 146 <strong style="color: #E97132;">Primary Question Annotation:</strong> Annotators select videos and create QA pairs requiring either <em>visual</em> or <em>narrative complexity</em>. A question is retained only if at least one top model (GPT-4o, LLaVA-Video, Qwen2.5-VL) fails to answer it c
146orrectly. 147 </li> 148 <li> 149 <strong style="color: #4EA72E;">Answer & Rationale:</strong> Annotators provide the correct answer, a detailed reasoning process, and critique of incorrect model responses. 150 </li> 151 <li> 152 <strong style="color: #4EA72E;">Sampling Check:</strong> Questions must be answerable from 80 uniformly sampled frames, ensuring reliance on visual rather than auditory cues. 153 </li> 154 <li> 155 <strong style="color: #4E95D9;">Adversarial Question Expansion:</strong> Annotators create four challenging variants per primary question based on model failures, with minimal edits to the original answer and rationale. 156 </li> 157 <li> 158 <strong style="color: #b7b7b7;">Alignment Check:</strong> A consensus-based process among three annotators ensures consistency. Questions without agreement are discarded. 159 </li> 160 </ul> 161 </div> 162 </div> 163 </section> 164 165 <section id="statistic" style="margin-top: 20px;"> 166 <div class="container"> 167 <h1 class="jumbotron-heading">Dataset Statistics</h1> 168 <div style="justify-content: space-around; flex-wrap: wrap; margin-top: 30px"> 169 <div class="image-item"> 170 <img src="assets/statistics.jpg" class="zoomable-image" style="margin-top: -20px; width: 100%; max-width: 8000px; height: auto;"> 171 </div> 172 </div> 173 174 <p>The dataset includes 5,000 question-answer pairs across 1,000 videos. Questions were first grouped by reasoning levelâelement, event, or plotâbased on video content. They were further categorized by the type of inquiry (e.g., Attributes, Localization). When a specific complexity factor appeared frequently within a category (e.g., over 50 instances), it was promoted to a sub-category (e.g., Element AttributesâIllusion). In total, 18 distinct question types are identified.</p> 175 </div> 176 177 </section> 178 179 <section id="performance" style="margin-top: 20px;"> 180 <div class="container text-left"> 181 <h1 class="jumbotron-heading">Performance</h1> 182 <div style="justify-content: space-around; flex-wrap: wrap; margin-top: 30px"> 183 <div class="image-item"> 184 185 <img src="assets/performance.png" class="zoomable-image" style="margin-top: -20px; width: 100%; max-width: 8000px; height: auto;"> 186 </div> 187 188 <p> 189 Among video-language models, performance varies widely. While open-source models like InternVL-2.5-8B perform well on straightforward questions (65.7% on Correctly-Led), they struggle with misleading prompts (24.5% on Wrongly-Led). LLaVA-Video-72B emerges as the strongest open-source model overall. Proprietary models such as GPT-4o and Gemini Pro outperform most open-source models, with GPT-4o showing greater robustness to misleading prompts (67.5% Correctly-Led, 39.8% Wrongly-Led), though still far behind human-level reasoning. Notably, LLaVA-Video-72B approaches GPT-4o's accuracy in multiple-choice settings, but falls short on primary open-ended questionsâhighlighting both a limitation of current open-source systems and a bias in existing benchmarks that overemphasize multiple-choice formats. 190 191 Regarding natural adversarial robustness, the table reveals that humans remain the gold standard with 64.4% accuracy. GPT-4o ranks highest among models at 36.0%, but still significantly lags behind. Open-source models like InternVL-2.5-7B perform poorly (10.9%), with even larger variants offering minimal gains. These results underscore the challenge of building models that can resist adversarial perturbations and the need for more rigorous benchmarks targeting open-ended and robustness-centric tasks. 192 193 </p> 194 </div> 195 </div> 196 </section> 197 198 199 200 <section id="acknowledgement" style="margin-top: 20px;"> 201 <div class="container text-left"> 202 <h1 class="jumbotron-heading">Acknowledgement</h1> 203 <div style="display: flex; justify-content: space-around; flex-wrap: wrap; margin-top: 20px"> 204 <p style="font-weight: 200"> 205 This study is supported by the Ministry of Education, Singapore, under its MOE AcRF Tier 2 (MOE-T2EP20221-0012, MOE-T2EP20223-0002), and under the RIE2020 Industry Alignment Fund â Industry Collaboration Projects (IAF-ICP) Funding Initiative, as well as cash and in-kind contribution from the industry partner(s). 206 Homepage credits: 207 <a href="https://snap-research.github.io/Panda-70M/" target="_blank" rel="noopener noreferrer" style="font-size: 1.2rem; color: white !important;">Panda-70M<a> 208 </p> 209 </div> 210 </div> 211 </section> 212 213 214 215 <!-- The Modal --> 216 <div id="imageModal" class="modal"> 217 <span class="close">×</span> 218 <img class="modal-content" id="modalImage"> 219 </div> 220 221</body> 222</html>
Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.