PageSourceSearch

https://binzhubz.github.io/

html binzhubz.github.io collected 2026-10-03 08:42:47 UTC 55,820 bytes, 5 lines download raw bytes

1<!DOCTYPE html> <html lang="en"> <head> <meta http-equiv="Content-Type" content="text/html; charset=UTF-8"> <meta charset="utf-8"> <meta name="viewport" content="width=device-width, initial-scale=1, shrink-to-fit=no"> <meta http-equiv="X-UA-Compatible" content="IE=edge"> <title> Bin Zhu </title> <meta name="author" content="Bin Zhu"> <meta name="description" content=""> <meta name="keywords" content="Multimedia, Multimodal Large Language Model, Egocentric Video Understanding, AI for Healthcare"> <link rel="stylesheet" href="/assets/css/bootstrap.min.css?a4b3f509e79c54a512b890d73235ef04"> <link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/[email protected]/css/mdb.min.css" integrity="sha256-jpjYvU3G3N6nrrBwXJoVEYI/0zw8htfFnhT9ljN3JJw=" crossorigin="anonymous"> <link defer rel="stylesheet" href="/assets/css/academicons.min.css?f0b7046b84e425c55f3463ac249818f5"> <link defer rel="stylesheet" href="/assets/css/scholar-icons.css?62b2ac103a88034e6882a5be5f3e2772"> <link defer rel="stylesheet" type="text/css" href="https://fonts.googleapis.com/css?family=Roboto:300,400,500,700|Roboto+Slab:100,300,400,500,700|Material+Icons&amp;display=swap"> <link defer rel="stylesheet" href="/assets/css/jekyll-pygments-themes-github.css?591dab5a4e56573bf4ef7fd332894c99" media="" id="highlight_theme_light"> <link rel="shortcut icon" href="data:image/svg+xml,&lt;svg%20xmlns=%22http://www.w3.org/2000/svg%22%20viewBox=%220%200%20100%20100%22&gt;&lt;text%20y=%22.9em%22%20font-size=%2290%22&gt;%E2%9A%9B%EF%B8%8F&lt;/text&gt;&lt;/svg&gt;"> <link rel="stylesheet" href="/assets/css/main.css?d41d8cd98f00b204e9800998ecf8427e"> <link rel="canonical" href="https://binzhubz.github.io/"> 
1<script src="/assets/js/theme.js?a81d82887dd692e91686b43de4542f18"></script>
1 <link defer rel="stylesheet" href="/assets/css/jekyll-pygments-themes-native.css?5847e5ed4a4568527aa6cfab446049ca" media="none" id="highlight_theme_dark"> 
1<script>
2    initTheme();
3  </script>
3 </head> <body class="fixed-top-nav "> <header> <nav id="navbar" class="navbar navbar-light navbar-expand-sm fixed-top" role="navigation"> <div class="container"> <button class="navbar-toggler collapsed ml-auto" type="button" data-toggle="collapse" data-target="#navbarNav" aria-controls="navbarNav" aria-expanded="false" aria-label="Toggle navigation"> <span class="sr-only">Toggle navigation</span> <span class="icon-bar top-bar"></span> <span class="icon-bar middle-bar"></span> <span class="icon-bar bottom-bar"></span> </button> <div class="collapse navbar-collapse text-right" id="navbarNav"> <ul class="navbar-nav ml-auto flex-nowrap"> <li class="nav-item active"> <a class="nav-link" href="/">about <span class="sr-only">(current)</span> </a> </li> <li class="nav-item "> <a class="nav-link" href="/publications/">publications </a> </li> <li class="nav-item "> <a class="nav-link" href="/team/">team </a> </li> <li class="nav-item "> <a class="nav-link" href="/service/">service </a> </li> <li class="nav-item "> <a class="nav-link" href="/grants&amp;talks/">grants&amp;talks </a> </li> <li class="nav-item "> <a class="nav-link" href="/teaching/">teaching </a> </li> <li class="nav-item"> <button id="search-toggle" title="Search" onclick="openSearchModal()"> <span class="nav-link">ctrl k <i class="ti ti-search"></i></span> </button> </li> <li class="toggle-container"> <button id="light-toggle" title="Change theme"> <i class="ti ti-sun-moon" id="light-toggle-system"></i> <i class="ti ti-moon-filled" id="light-toggle-dark"></i> <i class="ti ti-sun-filled" id="light-toggle-light"></i> </button> </li> </ul> </div> </div> </nav> <progress id="progress" value="0"> <div class="progress-container"> <span class="progress-bar"></span> </div> </progress> </header> <div class="container mt-5" role="main"> <div class="post"> <header class="post-header"> <h1 class="post-title"> Bin Zhu </h1> <p class="desc"></p> </header> <article> <div class="profile float-right"> <figure> <picture> <source class="responsive-img-srcset" srcset="/assets/img/prof_pic-480.webp 480w,/assets/img/prof_pic-800.webp 800w,/assets/img/prof_pic-1400.webp 1400w," type="image/webp" sizes="(min-width: 930px) 270.0px, (min-width: 576px) 30vw, 95vw"> <img src="/assets/img/prof_pic.jpg?bb9369170263d31fecad467fbb82c6eb" class="img-fluid z-depth-1 rounded" width="100%" height="auto" alt="prof_pic.jpg" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </source></picture> </figure> <div class="more-info"> <a href="https://scholar.google.com.hk/citations?user=xn0ZcJQAAAAJ&amp;hl=en" target="_blank" title="Google Scholar" rel="external nofollow noopener"><i class="ai ai-google-scholar"></i> Google Scholar</a> <p> <a href="mailto:[email protected]"><i class="fas fa-envelope"></i> E-mail</a></p> <p> 80 Stamford Road, Singapore 178902 </p> </div> </div> <div class="clearfix"> <p> I am an Assistant Professor of Computer Science and a Lee Kong Chian Fellow at the <a href="https://computing.smu.edu.sg/" rel="external nofollow noopener" target="_blank">School of Computing and Information Systems</a>, <a href="https://www.smu.edu.sg/" rel="external nofollow noopener" target="_blank">Singapore Management University (SMU)</a>. Before joining SMU, I was a Postdoctoral Researcher working with <a href="https://dimadamen.github.io/" rel="external nofollow noopener" target="_blank"> Prof. Dima Damen </a> at the University of Bristol, contributing to the <a href="https://www.robots.ox.ac.uk/~vgg/projects/visualai/" rel="external nofollow noopener" target="_blank">EPSRC Visual AI Program Grant</a> led by <a href="https://www.robots.ox.ac.uk/~az/" rel="external nofollow noopener" target="_blank">Prof. Andrew Zisserman </a>. I earned my Ph.D. degree from <a href="https://www.cs.cityu.edu.hk/" rel="external nofollow noopener" target="_blank">Department of Computer Science</a>, <a href="https://www.cityu.edu.hk" rel="external nofollow noopener" target="_blank">City University of Hong Kong</a> in 2021, under the supervision of <a href="https://faculty.smu.edu.sg/profile/ngo-chong-wah-601" rel="external nofollow noopener" target="_blank">Prof. Chong-Wah Ngo</a> and <a href="https://www.cs.cityu.edu.hk/~wkchan/" rel="external nofollow noopener" target="_blank">Dr. Wing-Kwong Chan</a>. Earlier, I obtained my master and bachelor degrees from <a href="http://www.zju.edu.cn/" rel="external nofollow noopener" target="_blank">Zhejiang University</a> and <a href="http://www.seu.edu.cn/" rel="external nofollow noopener" target="_blank">Southeast University</a> respectively. </p> <p>My research interest lies in Human Centered Multimedia Computing, broadly spanning Multimodal AI, Robotics, and Egocentric Vision. I aim to develop intelligent systems that integrate vision, language, and action to understand and generate multimodal content, reason about complex environments, and interact effectively with humans and the physical world. A major focus of my current research is trustworthy AI, with an emphasis on building robust, reliable, and safe intelligent systems for real-world environments. </p> <p>🔥🔥🔥<span style="color: rgb(255, 0, 0)">Openings:</span> I am actively looking for self-motivated Ph.D. students, CSC visiting students, and (remote) interns. In addition, I am looking for postdoctoral researchers and funded visiting students to work on embodied AI. If you are interested in working with me, please feel free to drop me an email with your CV and other supporting documents (if any). </p> </div> <h2> <a href="/news/" style="color: inherit">news</a> </h2> <div class="news"> <div class="table-responsive" style="max-height: 60vw"> <table class="table table-sm table-borderless"> <tr> <th scope="row" style="width: 20%">Aug 21, 2026</th> <td> Two papers on Cultural Video Generation Benchmarking and Video World Models for Robotic Safety have been accepted to EMNLP 2026, as a Main Conference paper and a Findings paper respectively. Congratulations to Xianjing, Huiqiong and all collaborators! </td> </tr> <tr> <th scope="row" style="width: 20%">Aug 14, 2026</th> <td> I delivered an online talk at Tianjin University on “Beyond Task Success: Towards Trustworthy Embodied Intelligence”. </td> </tr> <tr> <th scope="row" style="width: 20%">Jul 29, 2026</th> <td> I will serve as Senior Program Committee for AAAI 2027. </td> </tr> <tr> <th scope="row" style="width: 20%">Jul 10, 2026</th> <td> Two papers on Multimodal Large Language Model Safety and Spatial Intelligence have been accepted to ACM MM 2026. Congrats to Pengkun, Yian and all collaborators! </td> </tr> <tr> <th scope="row" style="width: 20%">Jul 07, 2026</th> <td> I delivered a talk at <a href="https://teai.fudan.edu.cn/" rel="external nofollow noopener" target="_blank">Institute of Trustworthy Embodied AI, Fudan University</a> on “Beyond Task Success: Towards Trustworthy Embodied Intelligence”. </td> </tr> <tr> <th scope="row" style="width: 20%">Jun 29, 2026</th> <td> One paper on Retrieval Augmented Generation is accepted by TOMM. Congrats to Hailong and all collaborators! </td> </tr> <tr> <th scope="row" style="width: 20%">Jun 19, 2026</th> <td> I am honored to have been awarded the Lee Kong Chian Fellowship! </td> </tr> <tr> <th scope="row" style="width: 20%">Jun 18, 2026</th> <td>
3 Two papers on Vision-Language-Action Safety and Long-tail Learning have been accepted to ECCV 2026. Congrats to Ninghao, Jianggang and all collaborators! </td> </tr> <tr> <th scope="row" style="width: 20%">Jun 08, 2026</th> <td> I was invited to deliver a keynote talk, titled “Beyond Task Success: Towards Trustworthy Foundation Models for Robotics,” at the <a href="https://trustvlm.github.io/TrustVLM-ICMR26/" rel="external nofollow noopener" target="_blank">TrustVLM Workshop at ICMR 2026</a>. </td> </tr> <tr> <th scope="row" style="width: 20%">Apr 23, 2026</th> <td> I was invited to deliver a talk at the <a href="https://singaporevisionday.github.io/svd2026/" rel="external nofollow noopener" target="_blank">Singapore Vision Day 2026</a> on “Towards Trustworthy Embodied Intelligence”. </td> </tr> <tr> <th scope="row" style="width: 20%">Apr 15, 2026</th> <td> Three papers on Multimodal Large Language Models and Vision-Language-Segmentation are accepted by ICMR 2026. Congrats to Yinxuan, Pengkun, Chengxi and all collaborators! </td> </tr> <tr> <th scope="row" style="width: 20%">Apr 07, 2026</th> <td> Two papers on Text-to-Video Generation and Video Large Language Model Safety have been accepted to ACL 2026, as a main conference paper and a Findings paper respectively. Congrats to Xianjing, Pengkun and all collaborators! </td> </tr> <tr> <th scope="row" style="width: 20%">Mar 18, 2026</th> <td> I am invited to serve as Special Session Co-Chair at <a href="https://www.mmm2027.net/organisers" rel="external nofollow noopener" target="_blank">MMM 2027</a>. </td> </tr> <tr> <th scope="row" style="width: 20%">Mar 17, 2026</th> <td> One paper on Text-to-Image Generation is accepted by ICME 2026. Congrats to Xiaoyu and all collaborators! </td> </tr> <tr> <th scope="row" style="width: 20%">Feb 21, 2026</th> <td> One paper on Runtime Monitoring for Robotic Manipulation is accepted by CVPR 2026. Congrats to Shijie and all collaborators! </td> </tr> <tr> <th scope="row" style="width: 20%">Jan 30, 2026</th> <td> Our paper <a href="https://arxiv.org/pdf/2506.09677" rel="external nofollow noopener" target="_blank">“Benchmarking Gaslighting Negation Attacks Against Reasoning Models”</a> has been selected as a Best Paper Candidate at <a href="https://mmm2026.cz" rel="external nofollow noopener" target="_blank">MMM 2026 </a>. </td> </tr> <tr> <th scope="row" style="width: 20%">Jan 18, 2026</th> <td> Three papers on Speech LLM Safety, 3D Hand Motion Generation and Recipe Generation are accepted by ICASSP 2026. Congrats to Jinyang, ChingLam, Guoshan and all collaborators! </td> </tr> <tr> <th scope="row" style="width: 20%">Dec 22, 2025</th> <td> One paper on Vision Language Navigation is accepted by TIP. Congrats to Guangzhao and all collaborators! </td> </tr> <tr> <th scope="row" style="width: 20%">Dec 08, 2025</th> <td> Our project “VISTA: A Value-Informed Safety &amp; Trust Architecture for Autonomous Agents” has been awarded by <a href="https://aisingapore.org/" rel="external nofollow noopener" target="_blank"> AI Singapore (AISG)</a>, with more than SGD$1,000,000 in total funding. I will serve as Co-Principal Investigator. Congrats Zhiguang! </td> </tr> <tr> <th scope="row" style="width: 20%">Nov 29, 2025</th> <td> I will serve as Organization Chair for Pacific Graphics 2026. </td> </tr> <tr> <th scope="row" style="width: 20%">Nov 21, 2025</th> <td> I will serve as Area Chair for ICME 2026. </td> </tr> <tr> <th scope="row" style="width: 20%">Nov 20, 2025</th> <td> I am invited to serve as Workshop Chair at <a href="https://icmr2026.org/organization.html" rel="external nofollow noopener" target="_blank">ACM ICMR 2026 </a>. </td> </tr> <tr> <th scope="row" style="width: 20%">Nov 08, 2025</th> <td> One paper on Reinforcement Learning for Robotic Manipulation is accepted as an Oral presentation by AAAI 2026. Congrats to Jiarui and all collaborators! </td> </tr> <tr> <th scope="row" style="width: 20%">Oct 30, 2025</th> <td> Two papers on Large Reasoning Model Safety and Multimodal Large Language Model are accepted by MMM 2026. </td> </tr> <tr> <th scope="row" style="width: 20%">Sep 08, 2025</th> <td> I was invited to deliver a talk at the <a href="https://www.nextcenter.org/" rel="external nofollow noopener" target="_blank">National University of Singapore NExT Research Centre</a> on “Food Computing from an Egocentric Video Perspective”. </td> </tr> <tr> <th scope="row" style="width: 20%">Sep 02, 2025</th> <td> I was honored to serve on the Board of Examiners for Ph.D. Candidate <a href="https://gabrielegoletto.github.io/" rel="external nofollow noopener" target="_blank">Gabriele Goletto</a>, whose thesis focused on Egocentric Vision. Congratulations to Dr. Goletto on a successful defense! 🎓 </td> </tr> <tr> <th scope="row" style="width: 20%">Aug 26, 2025</th> <td> One paper on Large Lithium-ion Battery Model is accepted by <a href="https://www.nature.com/articles/s41467-025-63678-7" rel="external nofollow noopener" target="_blank">Nature Communications</a> as co-corresponding author. </td> </tr> <tr> <th scope="row" style="width: 20%">Aug 22, 2025</th> <td> I am awarded a <a href="https://www.moe.gov.sg/" rel="external nofollow noopener" target="_blank">Singapore Ministry of Education (MOE) Academic Research Fund (AcRF) Tier 2 grant</a> as Principal Investigator for my project “Self-Adaptive Planning with Environmental Awareness for Embodied Agents”, with total funding of SGD$959,166. </td> </tr> <tr> <th scope="row" style="width: 20%">Jul 05, 2025</th> <td> One paper on Multimodal Large Language Model is accepted by ACM MM 2025. </td> </tr> <tr> <th scope="row" style="width: 20%">Jun 26, 2025</th> <td> One paper on Multimodal Large Language Model is accepted by ICCV 2025. </td> </tr> <tr> <th scope="row" style="width: 20%">Jun 19, 2025</th> <td> One paper on Recipe Progress Tracking in Non-Visual Cooking are accepted by ASSETS 2025. </td> </tr> <tr> <th scope="row" style="width: 20%">Jun 09, 2025</th> <td> One paper on Cooking Procedural Image Generation is accepted by ACM TOMM. </td> </tr> <tr> <th scope="row" style="width: 20%">Jun 06, 2025</th> <td> I will serve as the <a href="https://sites.google.com/view/www-acmicmr-org" rel="external nofollow noopener" target="_blank">Program Co-Chair for ACM ICMR 2027</a>, which will be held in Singapore! </td> </tr> <tr> <th scope="row" style="width: 20%">Apr 23, 2025</th> <td> One paper on Nutrition Estimation is accepted by ICMR 2025. </td> </tr> <tr> <th scope="row" style="width: 20%">Mar 21, 2025</th> <td> One paper on Ingredient Recognition is accepted by ICME 2025 (oral). </td> </tr> <tr> <th scope="row" style="width: 20%">Mar 08, 2025</th> <td> One paper on Egocentric Video Understanding is accepted by CVPR 2025. </td> </tr> <tr> <th scope="row" style="width: 20%">Feb 23, 2025</th> <td> One paper on Recipe Following in Cooking Video is accepted by CHI (LBW) 2025. </td> </tr> <tr> <th scope="row" style="width: 20%">Jan 16, 2025</th> <td> One paper on Large Multimodal Model in Food Domain is accepted by IEEE TMM. </td> </tr> <tr> <th scope="row" style="width: 20%">Dec 09, 2024</th> <td> Two papers on Grasp Generation and Text-to-Hand-Image Generation are accepted by AAAI 2025. </td> </tr> <tr> <th scope="row" style="width: 20%">Dec 05, 2024</th> <td> We are excited to announce our <a href="https://2025.ieeeicme.org/ss10-multimedia-for-cooking-and-eating-activities/" rel="external nofollow noopener" target="_blank">Special Session on Multimedia for Cooking and Eating Activities </
3a> at ICME 2025. We warmly invite you to submit your papers! </td> </tr> <tr> <th scope="row" style="width: 20%">Oct 29, 2024</th> <td> One paper on Recipe Generation is accepted by WACV 2025. </td> </tr> <tr> <th scope="row" style="width: 20%">Jul 15, 2024</th> <td> One paper on Cross-modal Recipe Retrieval is accepted by ECCV 2024. </td> </tr> <tr> <th scope="row" style="width: 20%">Jul 15, 2024</th> <td> One paper on Time-series Weight Prediction is accepted by ACM MM 2024 and is further selected as an Oral presentation (3.97%). </td> </tr> <tr> <th scope="row" style="width: 20%">Jun 15, 2024</th> <td> One paper on Text-driven Video Prediction is accepted by ACM TOMM. </td> </tr> <tr> <th scope="row" style="width: 20%">Feb 15, 2024</th> <td> Two papers on Unsupervised Video Hashing and Generalizable Food Recognition are accepted by IEEE TMM. </td> </tr> <tr> <th scope="row" style="width: 20%">Jan 01, 2024</th> <td> I joined <a href="https://www.smu.edu.sg/" rel="external nofollow noopener" target="_blank"> Singapore Management University </a> as an Assistant Professor of Computer Science. <img class="emoji" title=":sparkles:" alt=":sparkles:" src="https://github.githubassets.com/images/icons/emoji/unicode/2728.png" height="20" width="20"> <img class="emoji" title=":smile:" alt=":smile:" src="https://github.githubassets.com/images/icons/emoji/unicode/1f604.png" height="20" width="20"> </td> </tr> </table> </div> </div> <h2> <a href="/publications/" style="color: inherit">selected publications</a> </h2> <div class="publications"> <ol class="bibliography"> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">arXiv</abbr> <figure> <picture> <img src="/assets/img/publication_preview/LIBEROVPro.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="LIBEROVPro.png" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="li2026LIBERO-VPro" class="col-sm-8"> <div class="title">LIBERO-VPro: Benchmarking Closed-Loop Visual Robustness of Robotic Foundation Models</div> <div class="author"> Huiqiong Li, Zhiting Mei, Anirudha Majumdar, Jingjing Chen, Yu-Gang Jiang, and <em>Bin Zhu</em> </div> <div class="periodical"> <em>arXiv preprint arXiv:2609.24350</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2609.24350" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://huiqiongli.github.io/LIBERO-VPro/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">arXiv</abbr> <figure> <picture> <img src="/assets/img/publication_preview/MoWAM.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="MoWAM.png" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="wang2026mowam" class="col-sm-8"> <div class="title">MoWAM: Explicit Future Motion Prediction for Efficient World Action Models</div> <div class="author"> Jiayu Wang, <em>Bin Zhu</em>, Yue Yu, and Jingjing Chen </div> <div class="periodical"> <em>arXiv preprint arXiv:2609.20709</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2609.20709" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="#" class="btn btn-sm z-depth-0" role="button">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">EMNLP Findings</abbr> <figure> <picture> <img src="/assets/img/publication_preview/RoboTrust.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="RoboTrust.png" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="li2026robotrustbench" class="col-sm-8"> <div class="title">RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation</div> <div class="author"> Huiqiong Li, Jiayu Wang, Zhiting Mei, Anirudha Majumdar, Jingjing Chen, and <em>Bin Zhu</em> </div> <div class="periodical"> <em>In Findings of the Conference on Empirical Methods in Natural Language Processing</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2606.01600" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://huiqiongli.github.io/RoboTrustBench/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">EMNLP Main</abbr> <figure> <picture> <img src="/assets/img/publication_preview/CultureVidBench.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="CultureVidBench.png" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="han2026culturevidbench" class="col-sm-8"> <div class="title">CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation</div> <div class="author"> Xianjing Han, Yuhan Su, Yang Deng, Dong Ma, Wee Peng Tay, and <em>Bin Zhu</em> </div> <div class="periodical"> <em>In Proceedings of the Conference on Empirical Methods in Natural Language Processing</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2608.01942" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://hanxjing.github.io/CultureVidBench/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">MM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/ACMMM2026-gaslighting.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="ACMMM2026-gaslighting.png" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="gaslightingMM2026" class="col-sm-8"> <div class="title">Semantic-Structural Decoupling: Disentangling Semantic Attention from Structural Bias in the Attention Manifold</div> <div class="author"> Pengkun Jiao, <em>Bin Zhu</em>, Jingjing Chen, and Yu-Gang Jiang </div> <div class="periodical"> <em>In Proceedings of the 34th ACM International Conference on Multimedia</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2607.24017" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://github.com/pengkun-jiao/SPAR" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Code</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">MM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/ACMMM2026-SpatialImaginer.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="ACMMM2026-SpatialImaginer.png" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="li2026spatialimaginer" class="col-sm-8"> <div class="title">SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning</div> <div class="author"> Yian Li, Yang Jiao, <em>Bin Zhu</em>, Tianwen Qian, Shaoxiang Chen, Jingjing Chen, and Yu-Gang Jiang </div> <div class="periodical"> <em>In Proceedings of the 34th ACM International Conference on Multimedia</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2604.17385" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="#" class="btn btn-sm z-depth-0" role="button">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ECCV</abbr> <figure> <picture> <img src="/assets/img/publication_preview/IGAR.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="IGAR.png" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="IGAR" class="col-sm-8"> <div class="title">Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration</div> <div class="author"> Ninghao Zhang, <em>Bin Zhu</em>, Shijie Zhou, and Jingjing Chen </div> <div class="periodical"> <em>In European Conference on Computer Vision</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2603.06001" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://ray-nh.github.io/igar/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ECCV</abbr> <figure> <picture> <img src="/assets/img/publication_preview/VICAL.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="VICAL.png" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="VICAL" class="col-sm-8"> <div class="title">VICAL: Vicinal Consistency Alignment for Long-Tailed Visual Recognition</div> <div class="author"> Jianggang Zhu, Zheng Wang, <em>Bin Zhu</em>, Yi-Ping Phoebe Chen, and Jingjing Chen </div> <div class="periodical"> <em>In European Conference on Computer Vision</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="/assets/pdf/" class="btn btn-sm z-depth-0" role="button">PDF</a> <a href="#" class="btn btn-sm z-depth-0" role="button">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">CVPR Main</abbr> <figure> <picture> <img src="/assets/img/publication_preview/CVPR2026.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="CVPR2026.png" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="zhouCVPR" class="col-sm-8"> <div class="title">RC-NF: Robot-Conditioned Normalizing Flow for Real-Time Anomaly Detection in Robotic Manipulation</div> <div class="author"> Shijie Zhou, <em>Bin Zhu</em>, Jiarui Yang, Xiangyu Zhao, Jingjing Chen, and Yu-Gang Jiang </div> <div class="periodical"> <em>In Proceedings of the IEEE/CVF Computer Vision and Pattern Recognition Conference</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2603.11106" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://heikaishuizz.github.io/RC-NF/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ACL Main</abbr> <figure> <picture> <img src="/assets/img/publication_preview/OSCBench.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="OSCBench.png" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="OSCBench" class="col-sm-8"> <div class="title">OSCBench: Benchmarking Object State Change in Text-to-Video Generation</div> <div class="author"> Xianjing Han, <em>Bin Zhu</em>, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, and Jingjing Chen </div> <div class="periodical"> <em>In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2603.11698" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://hanxjing.github.io/OSCBench/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ACL Findings</abbr> <figure> <picture> <img src="/assets/img/publication_preview/VideoLLM-ACL26.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="VideoLLM-ACL26.png" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="VideoLLM-Gaslighting" class="col-sm-8"> <div class="title">Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models</div> <div class="author"> Ziyao Tang, Pengkun Jiao, <em>Bin Zhu</em>, Huiyan Qi, Jingjing Chen, and Yu-Gang Jiang </div> <div class="periodical"> <em>In Findings of the 64th Annual Meeting of the Association for Computational Linguistics</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2604.17873" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://pengkun-jiao.github.io/GasVideo-1000/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">AAAI Oral</abbr> <figure> <picture> <img src="/assets/img/publication_preview/AAAI2026.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="AAAI2026.png" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="yang2025actor" class="col-sm-8"> <div class="title">Actor-Critic for Continuous Action Chunks: A Reinforcement Learning Framework for Long-Horizon Robotic Manipulation with Sparse Reward</div> <div class="author"> Jiarui Yang, <em>Bin Zhu</em>, Jingjing Chen, and Yu-Gang Jiang </div> <div class="periodical"> <em>In Proceedings of the AAAI Conference on Artificial Intelligence</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2508.11143" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://github.com/flyfaerss/ac3" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Code</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">TIP</abbr> <figure> <picture> <img src="/assets/img/publication_preview/VLN-TIP2026.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="VLN-TIP2026.png" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="VLN-TIP2026" class="col-sm-8"> <div class="title">ThinkMatter: Panoramic-Aware Instructional Semantics for Monocular Vision-and-Language Navigation</div> <div class="author"> Guangzhao Dai, Shuo Wang, Hao Zhao, <em>Bin Zhu</em>, Qianru Sun, and Xiangbo Shu </div> <div class="periodical"> <em>IEEE Transactions on Image Processing</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://ieeexplore.ieee.org/document/11367385" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ICMR Oral</abbr> <figure> <picture> <img src="/assets/img/publication_preview/GaslightingBench.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="GaslightingBench.jpg" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure
3> </div> <div id="zhu2025calling" class="col-sm-8"> <div class="title">Benchmarking Gaslighting Negation Attacks Against Multimodal Large Language Models</div> <div class="author"> <em>Bin Zhu</em>, Yinxuan Gui, Huiyan Qi, Jingjing Chen, Chong-Wah Ngo, and Ee-Peng Lim </div> <div class="periodical"> <em>In ACM International Conference on Multimedia Retrieval</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2501.19017" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://yxg1005.github.io/GaslightingNegationAttacks/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">🏆 MMM Oral</abbr> <figure> <picture> <img src="/assets/img/publication_preview/GaslightingBench-R.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="GaslightingBench-R.png" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="zhu2025reasoning" class="col-sm-8"> <div class="title">Benchmarking Gaslighting Negation Attacks Against Reasoning Models</div> <div class="author"> <em>Bin Zhu</em>, Hailong Yin, Jingjing Chen, and Yu-Gang Jiang </div> <div class="periodical"> <em>In International Conference on Multimedia Modeling (Best Paper Candidate)</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2506.09677" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://binzhubz.github.io/GaslightingBench-R/" class="btn btn-sm z-depth-0" role="button">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">NC</abbr> <figure> <picture> <img src="/assets/img/publication_preview/NC2025.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="NC2025.png" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="LLiM" class="col-sm-8"> <div class="title">LLiM: Large Lithium-ion Battery Model for Secure Shared E-bike Battery in Smart Cities</div> <div class="author"> Donghui Ding, Zhao Li, Linhao Luo, Ming Jin, <em>Bin Zhu</em>, Yichen Zhong, Junhao Hu, Peng Cai, and Huiqi Hu </div> <div class="periodical"> <em>Nature Communications</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://www.nature.com/articles/s41467-025-63678-7" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://github.com/ddhdzt/LLiM" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Code</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ICCV</abbr> <figure> <picture> <img src="/assets/img/publication_preview/ICCV25-Dual-LoRA.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="ICCV25-Dual-LoRA.jpg" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="Dual-LoRA" class="col-sm-8"> <div class="title">From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning</div> <div class="author"> Pengkun Jiao, <em>Bin Zhu</em>, Jingjing Chen, Chong-Wah Ngo, and Yugang Jiang </div> <div class="periodical"> <em>In International Conference on Computer Vision</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2411.12787" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://github.com/pengkun-jiao/Dual-LoRA" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">MM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/MM2025.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="MM2025.png" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="li2024look" class="col-sm-8"> <div class="title">Look before you decide: Prompting active deduction of mllms for assumptive reasoning</div> <div class="author"> Yian Li, Wentao Tian, Yang Jiao, Jingjing Chen, Tianwen Qian, <em>Bin Zhu</em>, Na Zhao, and Yu-Gang Jiang </div> <div class="periodical"> <em>In Proceedings of the 33rd ACM International Conference on Multimedia</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2404.12966" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="#" class="btn btn-sm z-depth-0" role="button">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">CVPR</abbr> <figure> <picture> <img src="/assets/img/publication_preview/HDEPIC.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="HDEPIC.jpg" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="HDEPIC" class="col-sm-8"> <div class="title">HD-EPIC: A highly-detailed egocentric video dataset</div> <div class="author"> Toby Perrett, Ahmad Darkhalil, Saptarshi Sinha, Omar Emara, Sam Pollard, Kranti Kumar Parida, Kaiting Liu, Prajwal Gatti, Siddhant Bansal, Kevin Flanagan, Jacob Chalk, Zhifan Zhu, Rhodri Guerrier, Fahd Abdelazim, <em>Bin Zhu</em>, Davide Moltisanti, Michael Wray, Hazel Doughty, and Dima Damen </div> <div class="periodical"> <em>In Proceedings of the IEEE/CVF Computer Vision and Pattern Recognition Conference</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://openaccess.thecvf.com/content/CVPR2025/papers/Perrett_HD-EPIC_A_Highly-Detailed_Egocentric_Video_Dataset_CVPR_2025_paper.pdf" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://hd-epic.github.io/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">AAAI</abbr> <figure> <picture> <img src="/assets/img/publication_preview/Hand1000.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="Hand1000.jpg" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="zhang2025hand1000" class="col-sm-8"> <div class="title">
3Hand1000: Generating realistic hands from text with only 1,000 images</div> <div class="author"> Haozhuo Zhang, <em>Bin Zhu</em>, Yu Cao, and Yanbin Hao </div> <div class="periodical"> <em>In Proceedings of the AAAI Conference on Artificial Intelligence</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://ojs.aaai.org/index.php/AAAI/article/view/33074" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://github.com/Haozhuo-Zhang/Hand1000" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Code</a> <a href="https://haozhuo-zhang.github.io/Hand1000-project-page/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">AAAI</abbr> <figure> <picture> <img src="/assets/img/publication_preview/HandGrasp.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="HandGrasp.jpg" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="tang2025ragg" class="col-sm-8"> <div class="title">RAGG: Retrieval-Augmented Grasp Generation Model</div> <div class="author"> Zhenhua Tang, <em>Bin Zhu</em>, Yanbin Hao, Chong-Wah Ngo, and Richang Hong </div> <div class="periodical"> <em>In Proceedings of the AAAI Conference on Artificial Intelligence</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://ojs.aaai.org/index.php/AAAI/article/view/32786" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">TMM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/FoodLMM.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="FoodLMM.jpg" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="yin2023foodlmm" class="col-sm-8"> <div class="title">FoodLMM: A versatile food assistant using large multi-modal model</div> <div class="author"> Yuehao Yin, Huiyan Qi, <em>Bin Zhu</em>, Jingjing Chen, Yu-Gang Jiang, and Chong-Wah Ngo </div> <div class="periodical"> <em>IEEE Transactions on Multimedia</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2312.14991" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://github.com/YuehaoYin/FoodLMM" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Code</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ArXiv</abbr> <figure> <picture> <img src="/assets/img/publication_preview/Gaslighting-Attention.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="Gaslighting-Attention.jpg" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="jiao2025don" class="col-sm-8"> <div class="title">Don’t Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs</div> <div class="author"> Pengkun Jiao, <em>Bin Zhu</em>, Jingjing Chen, Chong-Wah Ngo, and Yu-Gang Jiang </div> <div class="periodical"> <em>arXiv preprint arXiv:2504.09456</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2504.09456" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">MM Oral</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_weightprediction.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_weightprediction.jpg" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="gui2024navigating" class="col-sm-8"> <div class="title">Navigating weight prediction with diet diary</div> <div class="author"> Yinxuan Gui, <em>Bin Zhu</em>, Jingjing Chen, Chong Wah Ngo, and Yu-Gang Jiang </div> <div class="periodical"> <em>In Proceedings of the 32nd ACM International Conference on Multimedia</em>, 2024 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2408.05445" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://yxg1005.github.io/weight-prediction/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ECCV</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_DAR.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_DAR.jpg" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="song2024enhancing" class="col-sm-8"> <div class="title">Enhancing recipe retrieval with foundation models: A data augmentation perspective</div> <div class="author"> Fangzhou Song, <em>Bin Zhu</em>, Yanbin Hao, and Shuo Wang </div> <div class="periodical"> <em>In European Conference on Computer Vision</em>, 2024 </div> <div class="periodical"> </div> <div class="links"> <a href="https://www.ecva.net/papers/eccv_2024/papers_ECCV/papers/06751.pdf" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://github.com/Noah888/DAR" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Code</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">MM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_CgT-GAN.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_CgT-GAN.jpg" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="yu2023cgt" class="col-sm-8"> <div class="title">CgT-GAN: clip-guided text GAN for image captioning</div> <div class="author"> Jiarui Yu, Haoran Li, Yanbin Hao, <em>Bin Zhu</em>, Tong Xu, and Xiangnan He </div> <div class="periodical"> <em>In Proceedings of the 31st ACM International Conference on Multimedia</em>, 2023 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2308.12045" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://github.com/Lihr747/CgtGAN" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Code</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">NeurIPS</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_VISOR.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_VISOR.jpg" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="darkhalil2022epic" class="col-sm-8"> <div class="title">Epic-kitchens visor benchmark: Video segmentations and object relations</div> <div class="author"> Ahmad Darkhalil, Dandan Shan, <em>Bin Zhu</em>, Jian Ma, Amlan Kar, Richard Higgins, Sanja Fidler, David Fouhey, and Dima Damen </div> <div class="periodical"> <em>In Advances in Neural Information Processing Systems Track on Datasets and Benchmarks</em>, 2022 </div> <div class="periodical"> </div> <div class="links"> <a href="https://proceedings.neurips.cc/paper_files/paper/2022/file/590a7ebe0da1f262c80d0188f5c4c222-Paper-Datasets_and_Benchmarks.pdf" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://epic-kitchens.github.io/VISOR/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">TIP</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_VIREOFood251.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_VIREOFood251.jpg" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="chen2020study" class="col-sm-8"> <div class="title">A study of multi-task and region-wise deep learning for food ingredient recognition</div> <div class="author"> Jingjing Chen, <em>Bin Zhu</em>, Chong-Wah Ngo, Tat-Seng Chua, and Yu-Gang Jiang </div> <div class="periodical"> <em>IEEE Transactions on Image Processing</em>, 2020 </div> <div class="periodical"> </div> <div class="links"> <a href="https://ink.library.smu.edu.sg/cgi/viewcontent.cgi?article=7304&amp;context=sis_research" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">MM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_crossdomain.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_crossdomain.jpg" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="zhu2020cross" class="col-sm-8"> <div class="title">Cross-domain cross-modal food transfer</div> <div class="author"> <em>Bin Zhu</em>, Chong-Wah Ngo, and Jing-jing Chen </div> <div class="periodical"> <em>In Proceedings of the 28th ACM International Conference on Multimedia</em>, 2020 </div> <div class="periodical"> </div> <div class="links"> <a href="https://ink.library.smu.edu.sg/cgi/viewcontent.cgi?article=7500&amp;context=sis_research" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">CVPR</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_CookGAN.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_CookGAN.jpg" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="zhu2020cookgan" class="col-sm-8"> <div class="title">CookGAN: Causality based text-to-image synthesis</div> <div class="author"> <em>Bin Zhu</em> and Chong-Wah Ngo </div> <div class="periodical"> <em>In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</em>, 2020 </div> <div class="periodical"> </div> <div class="links"> <a href="https://openaccess.thecvf.com/content_CVPR_2020/papers/Zhu_CookGAN_Causality_Based_Text-to-Image_Synthesis_CVPR_2020_paper.pdf" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">CVPR</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_R2GAN.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_R2GAN.jpg" data-zoomable loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="zhu2019r2gan" class="col-sm-8"> <div class="title">R2GAN: Cross-modal recipe retrieval with generative adversarial network</div> <div class="author"> <em>Bin Zhu</em>, Chong-Wah Ngo, Jingjing Chen, and Yanbin Hao </div> <div class="periodical"> <em>In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</em>, 2019 </div> <div class="periodical"> </div> <div class="links"> <a href="https://openaccess.thecvf.com/content_CVPR_2019/papers/Zhu_R2GAN_Cross-Modal_Recipe_Retrieval_With_Generative_Adversarial_Network_CVPR_2019_paper.pdf" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> </ol> </div> </article> </div> </div> <footer class="fixed-bottom" role="contentinfo"> <div class="container mt-0">
3 © Copyright 2026 Bin Zhu. </div> </footer> 
3<script src="https://cdn.jsdelivr.net/npm/[email protected]/dist/jquery.min.js" integrity="sha256-/xUj+3OJU5yExlq6GSYGSHk7tPXikynS7ogEvDej/m4=" crossorigin="anonymous"></script>
3 
3<script src="/assets/js/bootstrap.bundle.min.js"></script>
3 
3<script src="https://cdn.jsdelivr.net/npm/[email protected]/js/mdb.min.js" integrity="sha256-NdbiivsvWt7VYCt6hYNT3h/th9vSTL4EDWeGs5SN3DA=" crossorigin="anonymous"></script>
3 
3<script defer src="https://cdn.jsdelivr.net/npm/[email protected]/dist/masonry.pkgd.min.js" integrity="sha256-Nn1q/fx0H7SNLZMQ5Hw5JLaTRZp0yILA/FRexe19VdI=" crossorigin="anonymous"></script>
3 
3<script defer src="https://cdn.jsdelivr.net/npm/[email protected]/imagesloaded.pkgd.min.js" integrity="sha256-htrLFfZJ6v5udOG+3kNLINIKh2gvoKqwEhHYfTTMICc=" crossorigin="anonymous"></script>
3 
3<script defer src="/assets/js/masonry.js?a0db7e5d5c70cc3252b3138b0c91dcaf" type="text/javascript"></script>
3 
3<script defer src="https://cdn.jsdelivr.net/npm/[email protected]/dist/medium-zoom.min.js" integrity="sha256-ZgMyDAIYDYGxbcpJcfUnYwNevG/xi9OHKaR/8GK+jWc=" crossorigin="anonymous"></script>
3 
3<script defer src="/assets/js/zoom.js?85ddb88934d28b74e78031fd54cf8308"></script>
3 
3<script src="/assets/js/no_defer.js?2781658a0a2b13ed609542042a859126"></script>
3 
3<script defer src="/assets/js/common.js?e0514a05c5c95ac1a93a8dfd5249b92e"></script>
3 
3<script defer src="/assets/js/copy_code.js?c8a01c11a92744d44b093fc3bda915df" type="text/javascript"></script>
3 
3<script defer src="/assets/js/jupyter_new_tab.js?d9f17b6adc2311cbabd747f4538bb15f"></script>
3 
3<script async src="https://d1bxh8uas1mnw7.cloudfront.net/assets/embed.js"></script>
3 
3<script async src="https://badge.dimensions.ai/badge.js"></script>
3 
3<script defer type="text/javascript" id="MathJax-script" src="https://cdn.jsdelivr.net/npm/[email protected]/es5/tex-mml-chtml.js" integrity="sha256-MASABpB4tYktI2Oitl4t+78w/lyA+D7b/s9GEP0JOGI=" crossorigin="anonymous"></script>
3 
3<script src="/assets/js/mathjax-setup.js?a5bb4e6a542c546dd929b24b8b236dfd"></script>
3 
3<script defer src="https://cdnjs.cloudflare.com/polyfill/v3/polyfill.min.js?features=es6" crossorigin="anonymous"></script>
3 
3<script defer src="/assets/js/progress-bar.js?2f30e0e6801ea8f5036fa66e1ab0a71a" type="text/javascript"></script>
3 
3<script src="/assets/js/vanilla-back-to-top.min.js?f40d453793ff4f64e238e420181a1d17"></script>
3 
3<script>
4    addBackToTop();
5  </script>
5 
5<script type="module" src="/assets/js/search/ninja-keys.min.js?a3446f084dcaecc5f75aa1757d087dcf"></script>
5 <ninja-keys hidebreadcrumbs noautoloadmdicons placeholder="Type to start searching"></ninja-keys> 
5<script src="/assets/js/search-setup.js?6c304f7b1992d4b60f7a07956e52f04a"></script>
5 
5<script src="/assets/js/search-data.js"></script>
5 
5<script src="/assets/js/shortcut-key.js?6f508d74becd347268a7f822bca7309d"></script>
5 </body> </html>

Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.