1<!DOCTYPE html> 2<html> 3<head> 4 <meta charset="utf-8"> 5 <meta name="keywords" content="harmful dataset, harmful content detection, Vision Language Model"> 6 <meta name="viewport" content="width=device-width, initial-scale=1"> 7 <title>T2Vs Meet VLMs: A Scalable Multimodal Dataset for Visual Harmfulness Recognition</title> 8 9 <!-- Global site tag (gtag.js) - Google Analytics -->
vendor: 66 bytes, lines 9-10
9 10 <script async src="https://www.googletagmanager.com/gtag/js?id=
10G-PYVRSFMDRL
vendor: 14 bytes, lines 10-11
10"></script> 11
11<script> 12
vendor: 155 bytes, lines 12-20
12window.dataLayer = window.dataLayer || []; 13 14 function gtag() { 15 dataLayer.push(arguments); 16 } 17 18 gtag('js', new Date()); 19 20 gtag('config', '
20G-PYVRSFMDRL
vendor: 6 bytes, lines 20-21
20'); 21
21</script>
21 22 23 <link href="https://fonts.googleapis.com/css?family=Google+Sans|Noto+Sans|Castoro" 24 rel="stylesheet"> 25 26 <link rel="stylesheet" href="./static/css/bulma.min.css"> 27 <link rel="stylesheet" href="./static/css/bulma-carousel.min.css"> 28 <link rel="stylesheet" href="./static/css/bulma-slider.min.css"> 29 <link rel="stylesheet" href="./static/css/fontawesome.all.min.css"> 30 <link rel="stylesheet" 31 href="https://cdn.jsdelivr.net/gh/jpswalsh/academicons@1/css/academicons.min.css"> 32 <link rel="stylesheet" href="./static/css/index.css"> 33 <link rel="icon" href="./static/images/favicon.svg"> 34 35
35<script src="https://ajax.googleapis.com/ajax/libs/jquery/3.5.1/jquery.min.js"></script>
35 36
36<script defer src="./static/js/fontawesome.all.min.js"></script>
36 37
37<script src="./static/js/bulma-carousel.min.js"></script>
37 38
38<script src="./static/js/bulma-slider.min.js"></script>
38 39
39<script src="./static/js/index.js"></script>
39 40</head> 41<body> 42 43 44<section class="hero"> 45 <div class="hero-body"> 46 <div class="container is-max-desktop"> 47 <div class="columns is-centered"> 48 <div class="column has-text-centered"> 49 <h1 class="title is-1 publication-title">T2Vs Meet VLMs: A Scalable Multimodal Dataset for Visual Harmfulness Recognition</h1> 50 <div class="is-size-5 publication-authors"> 51 <span class="author-block"> 52 <a href="https://github.com/denny3388">Chen Yeh</a><sup>1*</sup>,</span> 53 <span class="author-block"> 54 <a href="https://github.com/thisismingggg">You-Ming Chang</a><sup>1*</sup>,</span> 55 <span class="author-block"> 56 <a href="https://walonchiu.github.io">Wei-Chen Chiu</a><sup>1</sup>, 57 </span> 58 <span class="author-block"> 59 <a href="https://ningyu1991.github.io/">Ning Yu</a><sup>2</sup> 60 </span> 61 </div> 62 <span class="author-block"> 63 (<sup>*</sup>Both authors contribute equally) 64 </span> 65 <div class="is-size-5 publication-authors"> 66 <span class="author-block"><sup>1</sup>National Yang Ming Chiao Tung University,</span> 67 <span class="author-block"><sup>2</sup>Netflix Eyeline Studios</span> 68 </div> 69 <span class="author-block"> 70 🎉 Accepted to <b>NeurIPS'24 Datasets and Benchmarks Track</b> 🎉 71 </span> 72 73 <div class="column has-text-centered"> 74 <div class="publication-links"> 75 <!-- NeurIPS PDF Link. --> 76 <span class="link-block"> 77 <!-- <a href="https://arxiv.org/pdf/2011.12948" class="external-link button is-normal is-rounded is-dark" 78 target="_blank"> --> 79 <a class="external-link button is-normal is-rounded is-dark" 80 target="_blank"> 81 <span class="icon"> 82 <i class="fas fa-file-pdf"></i> 83 </span> 84 <span>Paper</span> 85 </a> 86 </span> 87 <!-- Dataset Link. --> 88 <span class="link-block"> 89 <a href="https://huggingface.co/datasets/denny3388/VHD11K/tree/main" 90 class="external-link button is-normal is-rounded is-dark" 91 target="_blank"> 92 <!-- <span class="icon"> 93 <i class="fa fa-database"></i> 94 </span> --> 95 <span class="icon"> 96 <img class="icon-padding" 97 src="./static/images/hf-logo-pirate.png"> 98 </span> 99 <span>Data (VHD11K)</span> 100 </a> 101 </span> 102 <span class="link-block"> 103 <a href="https://arxiv.org/abs/2409.19734" 104 class="external-link button is-normal is-rounded is-dark" 105 target="_blank"> 106 <span class="icon"> 107 <i class="ai ai-arxiv"></i> 108 </span> 109 <span>arXiv (w/ Appendix)</span> 110 </a> 111 </span> 112 <!-- Video Link. --> 113 <span class="link-block"> 114 <a href="https://recorder-v3.slideslive.com/?share=92650&s=ae1ca94b-6fdf-4ea4-811f-384c8ce114c8" 115 class="external-link button is-normal is-rounded is-dark" 116 target="_blank" > 117 <!-- <a class="external-link button is-normal is-rounded is-dark" 118 target="_blank"> --> 119 <span class="icon"> 120 <i class="fa fa-video"></i> 121 </span> 122 <span>Video (Please use 123 Google Chrome)</span> 124 </a> 125 </span> 126 <!-- Code Link. --> 127 <span class="link-block"> 128 <a href="https://github.com/nctu-eva-lab/VHD11K" 129 class="external-link button is-normal is-rounded is-dark" 130 target="_blank">
131 <span class="icon"> 132 <i class="fab fa-github"></i> 133 </span> 134 <span>Code</span> 135 </a> 136 </span> 137 </div> 138 139 </div> 140 </div> 141 </div> 142 </div> 143 </div> 144</section> 145 146<!-- Brief Intro --> 147<section class="section"> 148 <div class="container is-max-desktop"> 149 <!-- Overview Image. --> 150 <div class="columns is-centered "> 151 <div class="column is-four-fifths has-text-centered interpolation-panel"> 152 <img src="./static/images/overview.png" 153 class="overview-image" 154 alt="Dataset curating process image."/> 155 <p>Overview: dataset curating process. Please note that the white rectangle masks serve as censorship, and are not included as inputs. <i>"A."</i>, <i>"N."</i> and <i>"J."</i> stand for the affirmative debater, the negative debater and the judge respectively.</p> 156 </div> 157 </div> 158 <!--/ Overview Image. --> 159 160 <!-- Overview. --> 161 <div class="columns is-centered has-text-centered"> 162 <div class="column is-four-fifths"> 163 <h2 class="title is-3">💡 Overview</h2> 164 <div class="content has-text-justified"> 165 <p> 166 We propose a comprehensive and extensive harmful dataset, <b>Visual Harmful Dataset 11K (VHD11K)</b>, consisting of <b>10,000 images</b> and <b>1,000 videos</b>, crawled from the Internet and generated by 4 generative models, across a total of <b>10 harmful categories</b> covering a full spectrum of harmful concepts with non-trival definition. 167 </p> 168 <p> 169 We also propose a novel annotation framework by formulating the annotation process as a <b>Multi-agent Visual Question Answering (VQA) Task</b>, having 3 different VLMs "<b>debate</b>" about whether the given image/video is harmful, and incorporating the in-context learning strategy in the debating process. 170 </p> 171 </div> 172 </div> 173 </div> 174 <!--/ Overview. --> 175 176 <!-- Paper video. --> 177 <!-- <div class="columns is-centered has-text-centered"> 178 <div class="column is-four-fifths"> 179 <h2 class="title is-3">📹 Video</h2> 180 <div class="publication-video"> 181 <iframe src="https://slideslive.com/js_embed/presentations/38990022" 182 frameborder="0" allow="autoplay; encrypted-media" allowfullscreen></iframe> 183 </div> 184 </div> 185 </div> --> 186 <!--/ Paper video. --> 187 188 <!-- VHD11K. --> 189 <div class="columns is-centered has-text-centered"> 190 <div class="column is-four-fifths"> 191 <h2 class="title is-3">📚 VHD11K: Our Proposed Multimodal Dataset for Visual Harmfulness Recognition</h2> 192 <div class="content has-text-justified"> 193 <p> 194 The entire dataset is publicly available at <a href="https://eva-lab.synology.me:8001/sharing/2iar2UrZs">here</a>. Under the shared folder, there are: 195 </p> 196 <pre><code>dataset_10000_1000 197|--croissant-vhd11k.json # metadata of VHD11K 198|--harmful_image_10000_ann.json # annotaion file of harmful images of VHD11K 199 (image name, harmful type, arguments, ...) 200|--harmful_images_10000.zip # 10000 harmful images of VHD11K 201|--harmful_video_1000_ann.json # annotaion file of harmful videos of VHD11K 202 (video name, harmful type, arguments, ...) 203|--harmful_videos_1000.zip # 1000 harmful videos of VHD11K 204|--ICL_samples.zip # in-context learning samples used in annoators 205|--ICL_images # in-context learning images 206|--ICL_videos_frames # frames of each in-context learning video</code></pre> 207 </div> 208 </div> 209 </div> 210 <!--/ VHD11K. --> 211 212 <!-- Evaluation. --> 213 <div class="columns is-centered has-text-centered"> 214 <div class="column is-four-fifths"> 215 <h2 class="title is-3">📊 Evaluation</h2> 216 <!-- Evaluation Tables --> 217 <div class="content has-text-justified"> 218 <div class="column is-four has-text-centered"> 219 <img src="./static/images/benchmarking.png" 220 class="interpolation-image" 221 alt="Pretrained models banchmarking table."/> 222 <p>Harmfulness detection accuracies of <b><u>pretrained</u></b> baseline methods.</p> 223 224 <img src="./static/images/VHD11K_vs_SMID.png" 225 class="interpolation-image" 226 style="padding-top: 25px;" 227 alt="VHD11K and SMID comparison table."/> 228 229 <p>Harmfulness detection accuracies of <b><u>prompt-tuned methods on VHD11K and SMID</u></b>
229.</p> 230 </div> 231 232 <!-- Evaluation text --> 233 <div class="column content has-text-justified"> 234 <p> 235 Evaluation and experimental results demonstrate that 236 <ol> 237 <li> 238 the <b>great alignment</b> between the annotation from our novel annotation framework and those from human, ensuring the reliability of VHD11K. 239 </li> 240 <li> 241 our full-spectrum harmful dataset <b>successfully identifies the inability of existing harmful content detection methods</b> to detect extensive harmful contents and improves the performance of existing harmfulness recognition methods. 242 </li> 243 <li> 244 our dataset <b>outperforms the baseline dataset, SMID,</b> as evidenced by the superior improvement in harmfulness recognition methods. 245 </li> 246 </ol> 247 </p> 248 </div> 249 </div> 250 </div> 251 <!--/ Evaluation. --> 252 253 </div> 254 <!-- Evaluation. --> 255 256 <!-- Example. --> 257 <h2 class="title is-3 has-text-centered">🙌 Examples</h2> 258 <div class="columns is-centered"> 259 <div class="column is-four-fifths"> 260 <h2 class="title is-4">Annotating Process</h2> 261 <div class="content has-text-justified"> 262 <p> 263 When inputting a visual content, the two debaters (i.e. affirmative and negative debaters) engage in a two-round debate on whether the given content is harmful. The "judge" then makes the final decision based on the arguments of both sides. For detailed role definitions for each of the three agents, please refer to the appendix of the paper. 264 </p> 265 </div> 266 267 <div class="columns is-vcentered interpolation-panel"> 268 <div class="column has-text-centered"> 269 <img src="./static/images/debate.png" 270 class="image_ICL_samples-image" 271 alt="Image ICL sample examples."/> 272 <p>An example of the debate annotation framework. Please note that the white rectangle masks serve as censorship, and are not included as inputs. <i>"Affirm."</i> and <i>"Neg."</i> stand for the affirmative and negative debaters, respectively.</p> 273 </div> 274 </div> 275 276 <h2 class="title is-4 is-centered">In-context Learning (ICL) Samples</h2> 277 <div class="content is-centered"> 278 <p> 279 Here are the in-context learning samples of the image/video annotator and the corresponding expected responses. 280 The white rectangles simply serve as censorship, and are not included as input. 281 </p> 282 </div> 283 284 <div class="columns is-centered interpolation-panel"> 285 <!-- Image --> 286 <div class="column"> 287 <h2 class="title is-5">Image</h2> 288 <div class=" has-text-centered"> 289 <img src="./static/images/image_ICL_samples_2.png" 290 class="image_ICL_samples-image" 291 alt="Image ICL sample examples."/> 292 </div> 293 </div> 294 <!--/ Image --> 295 296 <!-- Video. --> 297 <div class="column"> 298 <h2 class="title is-5">Video</h2> 299 <div class=" has-text-centered"> 300 <img src="./static/images/video_frame_ICL_samples.png" 301 class="video_frame_ICL_samples-image" 302 alt="Video frame ICL sample examples."/> 303 </div> 304 </div> 305 <!--/ Video. --> 306 </div> 307 </div> 308 </div> 309 <!-- Example. --> 310 311 </div> 312</section> 313 314<section class="section" id="BibTeX"> 315 <div class="container is-max-desktop content"> 316 317 <div class="columns is-centered"> 318 <div class="column is-four-fifths"> 319 320 <!-- Acknowledgement. --> 321 <h2 class="title is-3">💪 Acknowledgement</h2> 322 323 <div class="content has-text-justified"> 324 <p> 325 This project is built upon the the giant sholder of <a href="https://github.com/microsoft/autogen">Autogen</a>. Great thanks to them! 326 </p> 327 </div> 328 <!--/ Acknowledgement. --> 329 330 <h2 class="title is-3">👐 BibTeX</h2> 331 <pre><code>@inproceedings{yeh2024t2vs, 332 author={Chen Yeh and You-Ming Chang and Wei-Chen Chiu and Ning Yu}, 333 booktitle={Advances in Neural Information Processing Systems}, 334 title={T2Vs Meet VLMs: A Scalable Multimodal Dataset for Visual Harmfulness Recognition}, 335 year={2024} 336}</code></pre> 337 </div> 338 </div> 339 </div> 340</section> 341 342 343<footer class="footer"> 344 <div class="container"> 345 <div class="content has-text-centered"> 346 <a class="icon-link" 347 href="./static/videos/nerfies_paper.pdf"> 348 <i class="fas fa-file-pdf"></i> 349 </a> 350 <a class="icon-link" href="https://github.com/keunhong" class="external-link" disabled> 351 <i class="fab fa-github"></i> 352 </a> 353 </div> 354 <div class="columns is-centered"> 355 <div class="column is-8"> 356 <div class="content"> 357 <p> 358 This website is licensed under a <a rel="license" 359 href="http://creativecommons.org/licenses/by-sa/4.0/">Creative 360 Commons Attribution-ShareAlike 4.0 International License</a>. 361 </p> 362 <p> 363 This means you are free to borrow the <a 364 href="https://github.com/nerfies/nerfies.github.io">source code</a> of this website, 365 we just ask that you link back to this page in the footer. 366 Please remember to remove the analytics code included in the header of the website which 367 you do not want on your website. 368 </p> 369 </div> 370 </div> 371 </div> 372 </div> 373</footer> 374 375</body> 376</html>
Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.