1<!DOCTYPE html> <html lang="en"> <head> <meta http-equiv="Content-Type" content="text/html; charset=UTF-8"> <meta charset="utf-8"> <meta name="viewport" content="width=device-width, initial-scale=1, shrink-to-fit=no"> <meta http-equiv="X-UA-Compatible" content="IE=edge"> <title> publications | Bin Zhu </title> <meta name="author" content="Bin Zhu"> <meta name="description" content=""> <meta name="keywords" content="Multimedia, Multimodal Large Language Model, Egocentric Video Understanding, AI for Healthcare"> <link rel="stylesheet" href="/assets/css/bootstrap.min.css?a4b3f509e79c54a512b890d73235ef04"> <link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/[email protected]/css/mdb.min.css" integrity="sha256-jpjYvU3G3N6nrrBwXJoVEYI/0zw8htfFnhT9ljN3JJw=" crossorigin="anonymous"> <link defer rel="stylesheet" href="/assets/css/academicons.min.css?f0b7046b84e425c55f3463ac249818f5"> <link defer rel="stylesheet" href="/assets/css/scholar-icons.css?62b2ac103a88034e6882a5be5f3e2772"> <link defer rel="stylesheet" type="text/css" href="https://fonts.googleapis.com/css?family=Roboto:300,400,500,700|Roboto+Slab:100,300,400,500,700|Material+Icons&display=swap"> <link defer rel="stylesheet" href="/assets/css/jekyll-pygments-themes-github.css?591dab5a4e56573bf4ef7fd332894c99" media="" id="highlight_theme_light"> <link rel="shortcut icon" href="data:image/svg+xml,<svg%20xmlns=%22http://www.w3.org/2000/svg%22%20viewBox=%220%200%20100%20100%22><text%20y=%22.9em%22%20font-size=%2290%22>%E2%9A%9B%EF%B8%8F</text></svg>"> <link rel="stylesheet" href="/assets/css/main.css?d41d8cd98f00b204e9800998ecf8427e"> <link rel="canonical" href="https://binzhubz.github.io/publications/">
1<script src="/assets/js/theme.js?a81d82887dd692e91686b43de4542f18"></script>
1 <link defer rel="stylesheet" href="/assets/css/jekyll-pygments-themes-native.css?5847e5ed4a4568527aa6cfab446049ca" media="none" id="highlight_theme_dark">
1<script> 2 initTheme(); 3 </script>
3 </head> <body class="fixed-top-nav "> <header> <nav id="navbar" class="navbar navbar-light navbar-expand-sm fixed-top" role="navigation"> <div class="container"> <a class="navbar-brand title font-weight-lighter" href="/"> Bin Zhu </a> <button class="navbar-toggler collapsed ml-auto" type="button" data-toggle="collapse" data-target="#navbarNav" aria-controls="navbarNav" aria-expanded="false" aria-label="Toggle navigation"> <span class="sr-only">Toggle navigation</span> <span class="icon-bar top-bar"></span> <span class="icon-bar middle-bar"></span> <span class="icon-bar bottom-bar"></span> </button> <div class="collapse navbar-collapse text-right" id="navbarNav"> <ul class="navbar-nav ml-auto flex-nowrap"> <li class="nav-item "> <a class="nav-link" href="/">about </a> </li> <li class="nav-item active"> <a class="nav-link" href="/publications/">publications <span class="sr-only">(current)</span> </a> </li> <li class="nav-item "> <a class="nav-link" href="/team/">team </a> </li> <li class="nav-item "> <a class="nav-link" href="/service/">service </a> </li> <li class="nav-item "> <a class="nav-link" href="/grants&talks/">grants&talks </a> </li> <li class="nav-item "> <a class="nav-link" href="/teaching/">teaching </a> </li> <li class="nav-item"> <button id="search-toggle" title="Search" onclick="openSearchModal()"> <span class="nav-link">ctrl k <i class="ti ti-search"></i></span> </button> </li> <li class="toggle-container"> <button id="light-toggle" title="Change theme"> <i class="ti ti-sun-moon" id="light-toggle-system"></i> <i class="ti ti-moon-filled" id="light-toggle-dark"></i> <i class="ti ti-sun-filled" id="light-toggle-light"></i> </button> </li> </ul> </div> </div> </nav> <progress id="progress" value="0"> <div class="progress-container"> <span class="progress-bar"></span> </div> </progress> </header> <div class="container mt-5" role="main"> <div class="post"> <header class="post-header"> <h1 class="post-title">publications</h1> <p class="post-description"></p> </header> <article> <p>Please see <a href="https://scholar.google.com/citations?hl=en&user=xn0ZcJQAAAAJ&view_op=list_works&sortby=pubdate" rel="external nofollow noopener" target="_blank">Google Scholar</a> for more recent works and arXiv papers. </p>
3<script src="/assets/js/bibsearch.js?1bc438ca9037884cc579601c09afd847" type="module"></script>
3 <p><input type="text" id="bibsearch" spellcheck="false" autocomplete="off" class="search bibsearch-form-input" placeholder="Type to filter"></p> <div class="publications"> <h2 class="bibliography">2026</h2> <ol class="bibliography"> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">arXiv</abbr> <figure> <picture> <img src="/assets/img/publication_preview/LIBEROVPro.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="LIBEROVPro.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="li2026LIBERO-VPro" class="col-sm-8"> <div class="title">LIBERO-VPro: Benchmarking Closed-Loop Visual Robustness of Robotic Foundation Models</div> <div class="author"> Huiqiong Li, Zhiting Mei, Anirudha Majumdar, Jingjing Chen, Yu-Gang Jiang, and <em>Bin Zhu</em> </div> <div class="periodical"> <em>arXiv preprint arXiv:2609.24350</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2609.24350" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://huiqiongli.github.io/LIBERO-VPro/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">arXiv</abbr> <figure> <picture> <img src="/assets/img/publication_preview/MoWAM.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="MoWAM.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="wang2026mowam" class="col-sm-8"> <div class="title">MoWAM: Explicit Future Motion Prediction for Efficient World Action Models</div> <div class="author"> Jiayu Wang, <em>Bin Zhu</em>, Yue Yu, and Jingjing Chen </div> <div class="periodical"> <em>arXiv preprint arXiv:2609.20709</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2609.20709" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="#" class="btn btn-sm z-depth-0" role="button">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">EMNLP Findings</abbr> <figure> <picture> <img src="/assets/img/publication_preview/RoboTrust.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="RoboTrust.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="li2026robotrustbench" class="col-sm-8"> <div class="title">RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation</div> <div class="author"> Huiqiong Li, Jiayu Wang, Zhiting Mei, Anirudha Majumdar, Jingjing Chen, and <em>Bin Zhu</em> </div> <div class="periodical"> <em>In Findings of the Conference on Empirical Methods in Natural Language Processing</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2606.01600" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://huiqiongli.github.io/RoboTrustBench/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">EMNLP Main</abbr> <figure> <picture> <img src="/assets/img/publication_preview/CultureVidBench.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="CultureVidBench.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="han2026culturevidbench" class="col-sm-8"> <div class="title">CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation</div> <div class="author"> Xianjing Han, Yuhan Su, Yang Deng, Dong Ma, Wee Peng Tay, and <em>Bin Zhu</em> </div> <div class="periodical"> <em>In Proceedings of the Conference on Empirical Methods in Natural Language Processing</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2608.01942" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://hanxjing.github.io/CultureVidBench/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">MM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/ACMMM2026-gaslighting.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="ACMMM2026-gaslighting.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="gaslightingMM2026" class="col-sm-8"> <div class="title">Semantic-Structural Decoupling: Disentangling Semantic Attention from Structural Bias in the Attention Manifold</div> <div class="author"> Pengkun Jiao, <em>Bin Zhu</em>, Jingjing Chen, and Yu-Gang Jiang </div> <div class="periodical"> <em>In Proceedings of the 34th ACM International Conference on Multimedia</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2607.24017" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://github.com/pengkun-jiao/SPAR" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Code</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">MM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/ACMMM2026-SpatialImaginer.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="ACMMM2026-SpatialImaginer.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="li2026spatialimaginer" class="col-sm-8"> <div class="title">SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning</div> <div class="author"> Yian Li, Yang Jiao, <em>Bin Zhu</em>, Tianwen Qian, Shaoxiang Chen, Jingjing Chen, and Yu-Gang Jiang </div> <div class="periodical"> <em>In Proceedings of the 34th ACM International Conference on Multimedia</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2604.17385" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="#" class="btn btn-sm z-depth-0" role="button">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ECCV</abbr> <figure> <picture> <img src="/assets/img/publication_preview/IGAR.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="IGAR.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="IGAR" class="col-sm-8"> <div class="title">Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration</div> <div class="author"> Ninghao Zhang, <em>Bin Zhu</em>, Shijie Zhou, and Jingjing Chen </div> <div class="periodical"> <em>In European Conference on Computer Vision</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2603.06001" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://ray-nh.github.io/igar/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ECCV</abbr> <figure> <picture> <img src="/assets/img/publication_preview/VICAL.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="VICAL.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="VICAL" class="col-sm-8"> <div class="title">VICAL: Vicinal Consistency Alignment for Long-Tailed Visual Recognition</div> <div class="author"> Jianggang Zhu, Zheng Wang, <em>Bin Zhu</em>, Yi-Ping Phoebe Chen, and Jingjing Chen </div> <div class="periodical"> <em>In European Conference on Computer Vision</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="/assets/pdf/" class="btn btn-sm z-depth-0" role="button">PDF</a> <a href="#" class="btn btn-sm z-depth-0" role="button">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">CVPR Main</abbr> <figure> <picture> <img src="/assets/img/publication_preview/CVPR2026.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="CVPR2026.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="zhouCVPR" class="col-sm-8"> <div class="title">RC-NF: Robot-Conditioned Normalizing Flow for Real-Time Anomaly Detection in Robotic Manipulation</div> <div class="author"> Shijie Zhou, <em>Bin Zhu</em>, Jiarui Yang, Xiangyu Zhao, Jingjing Chen, and Yu-Gang Jiang </div> <div class="periodical"> <em>In Proceedings of the IEEE/CVF Computer Vision and Pattern Recognition Conference</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2603.11106" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://heikaishuizz.github.io/RC-NF/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ACL Main</abbr> <figure> <picture> <img src="/assets/img/publication_preview/OSCBench.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="OSCBench.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="OSCBench" class="col-sm-8"> <div class="title">OSCBench: Benchmarking Object State Change in Text-to-Video Generation</div> <div class="author"> Xianjing Han, <em>Bin Zhu</em>, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, and Jingjing Chen </div> <div class="periodical"> <em>In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2603.11698" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://hanxjing.github.io/OSCBench/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ACL Findings</abbr> <figure> <picture> <img src="/assets/img/publication_preview/VideoLLM-ACL26.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="VideoLLM-ACL26.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="VideoLLM-Gaslighting" class="col-sm-8"> <div class="title">Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models</div> <div class="author"> Ziyao Tang, Pengkun Jiao, <em>Bin Zhu</em>, Huiyan Qi, Jingjing Chen, and Yu-Gang Jiang </div> <div class="periodical"> <em>In Findings of the 64th Annual Meeting of the Association for Computational Linguistics</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2604.17873" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://pengkun-jiao.github.io/GasVideo-1000/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">AAAI Oral</abbr> <figure> <picture> <img src="/assets/img/publication_preview/AAAI2026.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="AAAI2026.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="yang2025actor" class="col-sm-8"> <div class="title">Actor-Critic for Continuous Action Chunks: A Reinforcement Learning Framework for Long-Horizon Robotic Manipulation with Sparse Reward</div> <div class="author"> Jiarui Yang, <em>Bin Zhu</em>, Jingjing Chen, and Yu-Gang Jiang </div> <div class="periodical"> <em>In Proceedings of the AAAI Conference on Artificial Intelligence</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2508.11143" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://github.com/flyfaerss/ac3" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Code</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">TIP</abbr> <figure> <picture> <img src="/assets/img/publication_preview/VLN-TIP2026.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="VLN-TIP2026.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="VLN-TIP2026" class="col-sm-8"> <div class="title">ThinkMatter: Panoramic-Aware Instructional Semantics for Monocular Vision-and-Language Navigation</div> <div class="author"> Guangzhao Dai, Shuo Wang, Hao Zhao, <em>Bin Zhu</em>, Qianru Sun, and Xiangbo Shu </div> <div class="periodical"> <em>IEEE Transactions on Image Processing</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://ieeexplore.ieee.org/document/11367385" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ICMR Oral</abbr> <figure> <picture> <img src="/assets/img/publication_preview/GaslightingBench.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="GaslightingBench.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="zhu2025calling" class="col-sm-8"> <div class="title">Benchmarking Gaslighting Negation Attacks Against Multimodal Large Language Models</div> <div class="author"> <em>Bin Zhu</em>, Yinxuan Gui, Huiyan Qi, Jingjing Chen, Chong-Wah Ngo, and Ee-Peng Lim </div> <div class="periodical"> <em>In ACM International Conference on Multimedia Retrieval</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2501.19017" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://yxg1005.github.io/GaslightingNegationAttacks/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ICMR Oral</abbr> <figure> <picture> <img src="/assets/img/publication_preview/ICMR2026-RODE.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="ICMR2026-RODE.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="RODE" class="col-sm-8"> <div class="title">Rode: Linear rectified mixture of diverse experts for food large multi-modal models</d
3iv> <div class="author"> Pengkun Jiao, Xinlan Wu, <em>Bin Zhu</em>, Jingjing Chen, Chong-Wah Ngo, and Yugang Jiang </div> <div class="periodical"> <em>In ACM International Conference on Multimedia Retrieval</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2407.12730?" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://pengkunjiao.github.io/UniFood-project/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ICMR Oral</abbr> <figure> <picture> <img src="/assets/img/publication_preview/ICMR26-SAM.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="ICMR26-SAM.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="zeng2026sam3" class="col-sm-8"> <div class="title">SAM3-LiteText: An Anatomical Study of the SAM3 Text Encoder for Efficient Vision-Language Segmentation</div> <div class="author"> Chengxi Zeng, Yuxuan Jiang, Ge Gao, Shuai Wang, Duolikun Danier, <em>Bin Zhu</em>, Stevan Rudinac, David Bull, and Fan Zhang </div> <div class="periodical"> <em>In ACM International Conference on Multimedia Retrieval</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2602.12173" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://github.com/SimonZeng7108/efficientsam3/tree/sam3_litetext" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Code</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">TOMM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/TOMM2026-ETTRAG.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="TOMM2026-ETTRAG.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="yin2025efficient" class="col-sm-8"> <div class="title">Efficient Test-Time Retrieval Augmented Generation</div> <div class="author"> Hailong Yin, <em>Bin Zhu</em>, Jingjing Chen, and Chong-Wah Ngo </div> <div class="periodical"> <em>ACM Transactions on Multimedia Computing, Communications and Applications</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2511.01059" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="#" class="btn btn-sm z-depth-0" role="button">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">TOMM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/ModelInversionAttacks.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="ModelInversionAttacks.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="li2024model" class="col-sm-8"> <div class="title">Model inversion attacks through target-specific conditional diffusion models</div> <div class="author"> Ouxiang Li, Yanbin Hao, Zhicai Wang, <em>Bin Zhu</em>, Shuo Wang, Zaixi Zhang, and Fuli Feng </div> <div class="periodical"> <em>ACM Transactions on Multimedia Computing, Communications and Applications</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2407.11424" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="#" class="btn btn-sm z-depth-0" role="button">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">PG</abbr> <figure> <picture> <img src="/assets/img/publication_preview/HeteroDiff-PG2026.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="HeteroDiff-PG2026.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="HeteroDiff-PG2026" class="col-sm-8"> <div class="title">Unveiling Silent Hill with HeteroDiff: Dense Haze Removal via Heterogeneous Diffusion</div> <div class="author"> Xiao Lv, Xiang Tao, Ying Yang, Chuan Ma, Shengfeng He, and <em>Bin Zhu</em> </div> <div class="periodical"> <em>In Pacific Graphics (Journal Track)</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="/assets/pdf/" class="btn btn-sm z-depth-0" role="button">PDF</a> <a href="#" class="btn btn-sm z-depth-0" role="button">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">PG</abbr> <figure> <picture> <img src="/assets/img/publication_preview/OOD-PG2026.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="OOD-PG2026.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="OOD-PG2026" class="col-sm-8"> <div class="title">Hierarchical Visual Feature Fusion for Robust Out-of-Distribution Detection under Semantic and Image-Level Covariate Shifts</div> <div class="author"> Chao Hou, Chuan Ma, Ziyong Du, Xiao Lv, Keke Tang, <em>Bin Zhu</em>, Shengfeng He, and Xiang Tao </div> <div class="periodical"> <em>In Pacific Graphics (Journal Track)</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="/assets/pdf/" class="btn btn-sm z-depth-0" role="button">PDF</a> <a href="#" class="btn btn-sm z-depth-0" role="button">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ICME</abbr> <figure> <picture> <img src="/assets/img/publication_preview/ICME2026.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="ICME2026.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="Multi-Person-Hand" class="col-sm-8"> <div class="title">Region-Aware Optimization for Multi-Person Hand Generation in Text-to-Image Synthesis</div> <div class="author"> Xiaoyu Chen, <em>Bin Zhu</em>, Xue Song, Pengkun Jiao, Yue Yu, and Jingjing Chen </div> <div class="periodical"> <em>In IEEE International Conference on Multimedia and Expo</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="/assets/pdf/" class="btn btn-sm z-depth-0" role="button">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ICASSP</abbr> <figure> <picture> <img src="/assets/img/publication_preview/ICASSP26-SpeechGaslighting.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="ICASSP26-SpeechGaslighting.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="wu2025benchmarking" class="col-sm-8"> <div class="title">Benchmarking Gaslighting Attacks Against Speech Large Language Models</div> <div class="author"> Jinyang Wu, <em>Bin Zhu</em>, Xiandong Zou, Qiquan Zhang, Xu Fang, and Pan Zhou </div> <div class="periodical"> <em>In IEEE International Conference on Acoustics, Speech and Signal Processing</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2509.19858" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://happyjackdreamer.github.io/Speech_LLM_Gaslighting/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ICASSP</abbr> <figure> <picture> <img src="/assets/img/publication_preview/ICASSP26-Hand.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="ICASSP26-Hand.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="ChengHandMotion" class="col-sm-8"> <div class="title">Teacher-Student Diffusion Model for Text-Driven 3D Hand Motion Generation</div> <div class="author"> Ching Lam Cheng, <em>Bin Zhu</em>, and Shengfeng He </div> <div class="periodical"> <em>In IEEE International Conference on Acoustics, Speech and Signal Processing</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2603.24407" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ICASSP</abbr> <figure> <picture> <img src="/assets/img/publication_preview/ICASSP26-Recipe.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="ICASSP26-Recipe.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="guoshan-ICASSP" class="col-sm-8"> <div class="title">Enhancing Action and Ingredient Modeling for Semantically Grounded Recipe Generation</div> <div class="author"> Guoshan Liu, <em>Bin Zhu</em>, Yian Li, Jingjing Chen, Chong-Wah Ngo, and Yu-Gang Jiang </div> <div class="periodical"> <em>In IEEE International Conference on Acoustics, Speech and Signal Processing</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://www.arxiv.org/pdf/2602.15862" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ð MMM Oral</abbr> <figure> <picture> <img src="/assets/img/publication_preview/GaslightingBench-R.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="GaslightingBench-R.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="zhu2025reasoning" class="col-sm-8"> <div class="title">Benchmarking Gaslighting Negation Attacks Against Reasoning Models</div> <div class="author"> <em>Bin Zhu</em>, Hailong Yin, Jingjing Chen, and Yu-Gang Jiang </div> <div class="periodical"> <em>In International Conference on Multimedia Modeling (Best Paper Candidate)</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2506.09677" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://binzhubz.github.io/GaslightingBench-R/" class="btn btn-sm z-depth-0" role="button">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">MMM Oral</abbr> <figure> <picture> <img src="/assets/img/publication_preview/Dual-Lora-MMM2026.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="Dual-Lora-MMM2026.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="wu2025dual" class="col-sm-8"> <div class="title">Dual-LoRA and Quality-Enhanced Pseudo Replay for Multimodal Continual Food Learning</div> <div class="author"> Xinlan Wu, <em>Bin Zhu</em>, Feng Han, Pengkun Jiao, and Jingjing Chen </div> <div class="periodical"> <em>In International Conference on Multimedia Modeling</em>, 2026 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2511.13351" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> </ol> <h2 class="bibliography">2025</h2> <ol class="bibliography"> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">NC</abbr> <figure> <picture> <img src="/assets/img/publication_preview/NC2025.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="NC2025.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="LLiM" class="col-sm-8"> <div class="title">LLiM: Large Lithium-ion Battery Model for Secure Shared E-bike Battery in Smart Cities</div> <div class="author"> Donghui Ding, Zhao Li, Linhao Luo, Ming Jin, <em>Bin Zhu</em>, Yichen Zhong, Junhao Hu, Peng Cai, and Huiqi Hu </div> <div class="periodical"> <em>Nature Communications</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://www.nature.com/articles/s41467-025-63678-7" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://github.com/ddhdzt/LLiM" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Code</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ICCV</abbr> <figure> <picture> <img src="/assets/img/publication_preview/ICCV25-Dual-LoRA.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="ICCV25-Dual-LoRA.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="Dual-LoRA" class="col-sm-8"> <div class="title">From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning</div> <div class="author"> Pengkun Jiao, <em>Bin Zhu</em>, Jingjing Chen, Chong-Wah Ngo, and Yugang Jiang </div> <div class="periodical"> <em>In International Conference on Computer Vision</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2411.12787" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://github.com/pengkun-jiao/Dual-LoRA" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">MM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/MM2025.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="MM2025.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="li2024look" class="col-sm-8"> <div class="title">Look before you decide: Prompting active deduction of mllms for assumptive reasoning</div> <div class="author"> Yian Li, Wentao Tian, Yang Jiao, Jingjing Chen, Tianwen Qian, <em>Bin Zhu</em>, Na Zhao, and Yu-Gang Jiang </div> <div class="periodical"> <em>In Proceedings of the 33rd ACM International Conference on Multimedia</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2404.12966" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="#" class="btn btn-sm z-depth-0" role="button">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">CVPR</abbr> <figure> <picture> <img src="/assets/img/publication_preview/HDEPIC.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="HDEPIC.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="HDEPIC" class="col-sm-8"> <div class="title">HD-EPIC: A highly-detailed egocentric video dataset</div> <div class="author"> Toby Perrett, Ahmad Darkhalil, Saptarshi Sinha, Omar Emara, Sam Pollard, Kranti Kumar Parida, Kaiting Liu, Prajwal Gatti, Siddhant Bansal, Kevin Flanagan, Jacob Chalk, Zhifan Zhu, Rhodri Guerrier, Fahd Abdelazim, <em>Bin Zhu</em>, Davide Moltisanti, Michael Wray, Hazel Doughty, and Dima Damen </div> <div class="periodical"> <em>In Proceedings of the IEEE/CVF Computer Vision and Pattern Recognition Conference</em>, 2025 </
3div> <div class="periodical"> </div> <div class="links"> <a href="https://openaccess.thecvf.com/content/CVPR2025/papers/Perrett_HD-EPIC_A_Highly-Detailed_Egocentric_Video_Dataset_CVPR_2025_paper.pdf" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://hd-epic.github.io/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">AAAI</abbr> <figure> <picture> <img src="/assets/img/publication_preview/Hand1000.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="Hand1000.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="zhang2025hand1000" class="col-sm-8"> <div class="title">Hand1000: Generating realistic hands from text with only 1,000 images</div> <div class="author"> Haozhuo Zhang, <em>Bin Zhu</em>, Yu Cao, and Yanbin Hao </div> <div class="periodical"> <em>In Proceedings of the AAAI Conference on Artificial Intelligence</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://ojs.aaai.org/index.php/AAAI/article/view/33074" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://github.com/Haozhuo-Zhang/Hand1000" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Code</a> <a href="https://haozhuo-zhang.github.io/Hand1000-project-page/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">AAAI</abbr> <figure> <picture> <img src="/assets/img/publication_preview/HandGrasp.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="HandGrasp.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="tang2025ragg" class="col-sm-8"> <div class="title">RAGG: Retrieval-Augmented Grasp Generation Model</div> <div class="author"> Zhenhua Tang, <em>Bin Zhu</em>, Yanbin Hao, Chong-Wah Ngo, and Richang Hong </div> <div class="periodical"> <em>In Proceedings of the AAAI Conference on Artificial Intelligence</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://ojs.aaai.org/index.php/AAAI/article/view/32786" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">TMM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/FoodLMM.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="FoodLMM.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="yin2023foodlmm" class="col-sm-8"> <div class="title">FoodLMM: A versatile food assistant using large multi-modal model</div> <div class="author"> Yuehao Yin, Huiyan Qi, <em>Bin Zhu</em>, Jingjing Chen, Yu-Gang Jiang, and Chong-Wah Ngo </div> <div class="periodical"> <em>IEEE Transactions on Multimedia</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2312.14991" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://github.com/YuehaoYin/FoodLMM" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Code</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ArXiv</abbr> <figure> <picture> <img src="/assets/img/publication_preview/Gaslighting-Attention.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="Gaslighting-Attention.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="jiao2025don" class="col-sm-8"> <div class="title">Donât Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs</div> <div class="author"> Pengkun Jiao, <em>Bin Zhu</em>, Jingjing Chen, Chong-Wah Ngo, and Yu-Gang Jiang </div> <div class="periodical"> <em>arXiv preprint arXiv:2504.09456</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2504.09456" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">WACV</abbr> <figure> <picture> <img src="/assets/img/publication_preview/RAG-WACV25.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="RAG-WACV25.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="liu2025retrieval" class="col-sm-8"> <div class="title">Retrieval augmented recipe generation</div> <div class="author"> Guoshan Liu, Hailong Yin, <em>Bin Zhu</em>, Jingjing Chen, Chong-Wah Ngo, and Yu-Gang Jiang </div> <div class="periodical"> <em>In IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://openaccess.thecvf.com/content/WACV2025/papers/Liu_Retrieval_Augmented_Recipe_Generation_WACV_2025_paper.pdf" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">TOMM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/TOMM25-cookingdiff.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="TOMM25-cookingdiff.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="wang2025cookingdiffusion" class="col-sm-8"> <div class="title">Cookingdiffusion: Cooking procedural image generation with stable diffusion</div> <div class="author"> Yuan Wang, <em>Bin Zhu</em>, Yanbin Hao, Chong-Wah Ngo, Yi Tan, and Xiang Wang </div> <div class="periodical"> <em>ACM Transactions on Multimedia Computing, Communications and Applications</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2501.09042" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="#" class="btn btn-sm z-depth-0" role="button">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ICMR</abbr> <figure> <picture> <img src="/assets/img/publication_preview/FastFood.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="FastFood.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="qi2025advancing" class="col-sm-8"> <div class="title">Advancing Food Nutrition Estimation via Visual-Ingredient Feature Fusion</div> <div class="author"> Huiyan Qi, <em>Bin Zhu</em>, Chong-Wah Ngo, Jingjing Chen, and Ee-Peng Lim </div> <div class="periodical"> <em>In ACM International Conference on Multimedia Retrieval (ICMR)</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2505.08747" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://huiyanqi.github.io/fastfood-nutrition-estimation/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ICME</abbr> <figure> <picture> <img src="/assets/img/publication_preview/ICME25.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="ICME25.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="gui2025efficient" class="col-sm-8"> <div class="title">Efficient Prompt Tuning for Hierarchical Ingredient Recognition</div> <div class="author"> Yinxuan Gui, <em>Bin Zhu</em>, Jingjing Chen, and Chong-Wah Ngo </div> <div class="periodical"> <em>In IEEE International Conference on Multimedia and Expo (ICME)</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2504.10322" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ASSETS</abbr> <figure> <picture> <img src="/assets/img/publication_preview/Assets25.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="Assets25.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="li2025oscar" class="col-sm-8"> <div class="title">
3Exploring Object Status Recognition for Recipe Progress Tracking in Non-Visual Cooking</div> <div class="author"> Franklin Mingzhe Li, Kaitlyn Ng, <em>Bin Zhu</em>, and Patrick Carrington </div> <div class="periodical"> <em>In International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS)</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2507.03330" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">CHI-LBW</abbr> <figure> <picture> <img src="/assets/img/publication_preview/OSCAR-CHI2025.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="OSCAR-CHI2025.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="li2025oscas" class="col-sm-8"> <div class="title">OSCAR: Object Status and Contextual Awareness for Recipes to Support Non-Visual Cooking</div> <div class="author"> Franklin Mingzhe Li, Kaitlyn Ng, <em>Bin Zhu</em>, and Patrick Carrington </div> <div class="periodical"> <em>In Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems</em>, 2025 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2503.05962" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> </ol> <h2 class="bibliography">2024</h2> <ol class="bibliography"> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">MM Oral</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_weightprediction.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_weightprediction.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="gui2024navigating" class="col-sm-8"> <div class="title">Navigating weight prediction with diet diary</div> <div class="author"> Yinxuan Gui, <em>Bin Zhu</em>, Jingjing Chen, Chong Wah Ngo, and Yu-Gang Jiang </div> <div class="periodical"> <em>In Proceedings of the 32nd ACM International Conference on Multimedia</em>, 2024 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2408.05445" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://yxg1005.github.io/weight-prediction/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ECCV</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_DAR.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_DAR.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="song2024enhancing" class="col-sm-8"> <div class="title">Enhancing recipe retrieval with foundation models: A data augmentation perspective</div> <div class="author"> Fangzhou Song, <em>Bin Zhu</em>, Yanbin Hao, and Shuo Wang </div> <div class="periodical"> <em>In European Conference on Computer Vision</em>, 2024 </div> <div class="periodical"> </div> <div class="links"> <a href="https://www.ecva.net/papers/eccv_2024/papers_ECCV/papers/06751.pdf" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://github.com/Noah888/DAR" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Code</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">TMM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_canteen.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_canteen.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="liu2024canteen" class="col-sm-8"> <div class="title">From canteen food to daily meals: Generalizing food recognition to more practical scenarios</div> <div class="author"> Guoshan Liu, Yang Jiao, Jingjing Chen, <em>Bin Zhu</em>, and Yu-Gang Jiang </div> <div class="periodical"> <em>IEEE Transactions on Multimedia</em>, 2024 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2403.07403" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">TMM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/TMM24-hashing.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="TMM24-hashing.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div>
3 <div id="duan2024efficient" class="col-sm-8"> <div class="title">Efficient Unsupervised Video Hashing with Contextual Modeling and Structural Controlling</div> <div class="author"> Jingru Duan, Yanbin Hao, <em>Bin Zhu</em>, Lechao Cheng, Pengyuan Zhou, and Xiang Wang </div> <div class="periodical"> <em>IEEE Transactions on Multimedia</em>, 2024 </div> <div class="periodical"> </div> <div class="links"> <a href="https://ieeexplore.ieee.org/abstract/document/10443557" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">TOMM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_TVP.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_TVP.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="song2024text" class="col-sm-8"> <div class="title">Text-driven video prediction</div> <div class="author"> Xue Song, Jingjing Chen, <em>Bin Zhu</em>, and Yu-Gang Jiang </div> <div class="periodical"> <em>ACM Transactions on Multimedia Computing, Communications and Applications</em>, 2024 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2210.02872" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">TOMM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/TOMM24-caption.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="TOMM24-caption.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="li2024cvlp" class="col-sm-8"> <div class="title">CVLP-NaVD: Contrastive Visual-Language Pre-training Models for Non-annotated Visual Description</div> <div class="author"> Haoran Li, Yanbin Hao, Jiarui Yu, <em>Bin Zhu</em>, Shuo Wang, and Tong Xu </div> <div class="periodical"> <em>ACM Transactions on Multimedia Computing, Communications and Applications</em>, 2024 </div> <div class="periodical"> </div> <div class="links"> <a href="https://dl.acm.org/doi/10.1145/3708348" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">MM Asia</abbr> <figure> <picture> <img src="/assets/img/publication_preview/MMAsia24.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="MMAsia24.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="ma2024active" class="col-sm-8"> <div class="title">Active Object Segmentation: A New Modality for Egocentric Action Recognition</div> <div class="author"> Jian Ma, <em>Bin Zhu</em>, Kun Li, and Dima Damen </div> <div class="periodical"> <em>In Proceedings of the 6th ACM International Conference on Multimedia in Asia</em>, 2024 </div> <div class="periodical"> </div> <div class="links"> <a href="https://dl.acm.org/doi/10.1145/3696409.3700164" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ECCVW</abbr> <figure> <picture> <img src="/assets/img/publication_preview/ECCV24-videoEditing.png" class="preview z-depth-1 rounded" width="100%" height="auto" alt="ECCV24-videoEditing.png" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="zhu2025video" class="col-sm-8"> <div class="title">Video editing for video retrieval</div> <div class="author"> <em>Bin Zhu</em>, Kevin Flanagan, Adriano Fragomeni, Michael Wray, and Dima Damen </div> <div class="periodical"> <em>In European Conference on Computer Vision Workshop</em>, 2024 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/abs/2402.02335" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> </ol> <h2 class="bibliography">2023</h2> <ol class="bibliography"><li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">MM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_CgT-GAN.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_CgT-GAN.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="yu2023cgt" class="col-sm-8"> <div class="title">CgT-GAN: clip-guided text GAN for image captioning</div> <div class="author"> Jiarui Yu, Haoran Li, Yanbin Hao, <em>Bin Zhu</em>, Tong Xu, and Xiangnan He </div> <div class="periodical"> <em>In Proceedings of the 31st ACM International Conference on Multimedia</em>, 2023 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2308.12045" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://github.com/Lihr747/CgtGAN" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Code</a> </div> </div> </div> </li></ol> <h2 class="bibliography">2022</h2> <ol class="bibliography"> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">NeurIPS</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_VISOR.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_VISOR.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="darkhalil2022epic" class="col-sm-8"> <div class="title">Epic-kitchens visor benchmark: Video segmentations and object relations</div> <div class="author"> Ahmad Darkhalil, Dandan Shan, <em>Bin Zhu</em>, Jian Ma, Amlan Kar, Richard Higgins, Sanja Fidler, David Fouhey, and Dima Damen </div> <div class="periodical"> <em>In Advances in Neural Information Processing Systems Track on Datasets and Benchmarks</em>, 2022 </div> <div class="periodical"> </div> <div class="links"> <a href="https://proceedings.neurips.cc/paper_files/paper/2022/file/590a7ebe0da1f262c80d0188f5c4c222-Paper-Datasets_and_Benchmarks.pdf" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://epic-kitchens.github.io/VISOR/" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Website</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">MM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_hashing.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_hashing.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="hao2022unsupervised" class="col-sm-8"> <div class="title">Unsupervised video hashing with multi-granularity contextualization and multi-structure preservation</div> <div class="author"> Yanbin Hao, Jingru Duan, Hao Zhang, <em>Bin Zhu</em>, Pengyuan Zhou, and Xiangnan He </div> <div class="periodical"> <em>In Proceedings of the 30th ACM International Conference on Multimedia</em>, 2022 </div> <div class="periodical"> </div> <div class="links"> <a href="https://ink.library.smu.edu.sg/cgi/viewcontent.cgi?article=10017&context=sis_research" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> <a href="https://github.com/haoyanbin918/MCMSH" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">Code</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">MM Oral</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_DANN.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_DANN.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="yin2022mix" class="col-sm-8"> <div class="title">Mix-dann and dynamic-modal-distillation for video domain adaptation</div> <div class="author"> Yuehao Yin, <em>Bin Zhu</em>, Jingjing Chen, Lechao Cheng, and Yu-Gang Jiang </div> <div class="periodical"> <em>In Proceedings of the 30th ACM International Conference on Multimedia</em>, 2022 </div> <div class="periodical"> </div> <div class="links"> <a href="https://ink.library.smu.edu.sg/cgi/viewcontent.cgi?article=10018&context=sis_research" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">ICMR</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_mixup.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_mixup.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="zhu2022cross" class="col-sm-8"> <div class="title">Cross-lingual adaptation for recipe retrieval with mixup</div> <div class="author"> <em>Bin Zhu</em>, Chong-Wah Ngo, Jingjing Chen, and Wing-Kwong Chan </div> <div class="periodical"> <em>In Proceedings of the 2022 International Conference on Multimedia Retrieval</em>, 2022 </div> <div class="periodical"> </div> <div class="links"> <a href="https://arxiv.org/pdf/2205.03891" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> </ol> <h2 class="bibliography">2021</h2> <ol class="bibliography"> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">TMM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_TMMsurvey.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_TMMsurvey.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="zhu2021learning" class="col-sm-8"> <div class="title">Learning from web recipe-image pairs for food recognition: Problem, baselines and performance</div> <div class="author"> <em>Bin Zhu</em>, Chong-Wah Ngo, and Wing-Kwong Chan </div> <div class="periodical"> <em>IEEE Transactions on Multimedia</em>, 2021 </div> <div class="periodical"> </div> <div class="links"> <a href="https://ink.library.smu.edu.sg/cgi/viewcontent.cgi?article=8249&context=sis_research" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">TIP</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_hyberlink.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_hyberlink.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="hao2021learning" class="col-sm-8"> <div class="title">Learning to match anchor-target video pairs with dual attentional holographic networks</div> <div class="author"> Yanbin Hao, Chong-Wah Ngo, and <em>Bin Zhu</em> </div> <div class="periodical"> <em>IEEE Transactions on Image Processing</em>, 2021 </div> <div class="periodical"> </div> <div class="links"> <a href="https://ieeexplore.ieee.org/abstract/document/9547825" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> </ol> <h2 class="bibliography">2020</h2> <ol class="bibliography"> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">TIP</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_VIREOFood251.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_VIREOFood251.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="chen2020study" class="col-sm-8"> <div class="title">A study of multi-task and region-wise deep learning for food ingredient recognition</div> <div class="author"> Jingjing Chen, <em>Bin Zhu</em>, Chong-Wah Ngo, Tat-Seng Chua, and Yu-Gang Jiang </div> <div class="periodical"> <em>IEEE Transactions on Image Processing</em>, 2020 </div> <div class="periodical"> </div> <div class="links"> <a href="https://ink.library.smu.edu.sg/cgi/viewcontent.cgi?article=7304&context=sis_research" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">MM Grand Challenge</abbr> <figure> <picture> <img src="/assets/img/publication_preview/TSD-TSM.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="TSD-TSM.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="hao2020person" class="col-sm-8"> <div class="title">Person-level action recognition in complex events via tsd-tsm networks</div> <div class="author"> Yanbin Hao, Zi-Niu Liu, Hao Zhang, <em>Bin Zhu</em>, Jingjing Chen, Yu-Gang Jiang, and Chong-Wah Ngo </div> <div class="periodical"> <em>In Proceedings of the 28th ACM International Conference on Multimedia Grand Challenge: Human Centric Analysis</em>, 2020 </div> <div class="periodical"> </div> <div class="links"> <a href="https://ink.library.smu.edu.sg/cgi/viewcontent.cgi?article=7506&context=sis_research" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">MM</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_crossdomain.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_crossdomain.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="zhu2020cross" class="col-sm-8"> <div class="title">Cross-domain cross-modal food transfer</div> <div class="author"> <em>Bin Zhu</em>, Chong-Wah Ngo, and Jing-jing Chen </div> <div class="periodical"> <em>In Proceedings of the 28th ACM International Conference on Multimedia</em>, 2020 </div> <div class="periodical"> </div> <div class="links"> <a href="https://ink.library.smu.edu.sg/cgi/viewcontent.cgi?article=7500&context=sis_research" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> <li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">CVPR</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_CookGAN.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_CookGAN.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="zhu2020cookgan" class="col-sm-8"> <div class="title">CookGAN: Causality based text-to-image synthesis</div> <div class="author"> <em>Bin Zhu</em> and Chong-Wah Ngo </div> <div class="periodical"> <em>In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</em>, 2020 </div> <div class="periodical"> </div> <div class="links"> <a href="https://openaccess.thecvf.com/content_CVPR_2020/papers/Zhu_CookGAN_Causality_Based_Text-to-Image_Synthesis_CVPR_2020_paper.pdf" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li> </ol> <
3h2 class="bibliography">2019</h2> <ol class="bibliography"><li> <div class="row"> <div class="col col-sm-2 abbr"> <abbr class="badge rounded w-100">CVPR</abbr> <figure> <picture> <img src="/assets/img/publication_preview/cover_R2GAN.jpg" class="preview z-depth-1 rounded" width="100%" height="auto" alt="cover_R2GAN.jpg" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"> </picture> </figure> </div> <div id="zhu2019r2gan" class="col-sm-8"> <div class="title">R2GAN: Cross-modal recipe retrieval with generative adversarial network</div> <div class="author"> <em>Bin Zhu</em>, Chong-Wah Ngo, Jingjing Chen, and Yanbin Hao </div> <div class="periodical"> <em>In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</em>, 2019 </div> <div class="periodical"> </div> <div class="links"> <a href="https://openaccess.thecvf.com/content_CVPR_2019/papers/Zhu_R2GAN_Cross-Modal_Recipe_Retrieval_With_Generative_Adversarial_Network_CVPR_2019_paper.pdf" class="btn btn-sm z-depth-0" role="button" rel="external nofollow noopener" target="_blank">PDF</a> </div> </div> </div> </li></ol> </div> </article> </div> </div> <footer class="fixed-bottom" role="contentinfo"> <div class="container mt-0"> © Copyright 2026 Bin Zhu. </div> </footer>
3<script src="https://cdn.jsdelivr.net/npm/[email protected]/dist/jquery.min.js" integrity="sha256-/xUj+3OJU5yExlq6GSYGSHk7tPXikynS7ogEvDej/m4=" crossorigin="anonymous"></script>
3
3<script src="/assets/js/bootstrap.bundle.min.js"></script>
3
3<script src="https://cdn.jsdelivr.net/npm/[email protected]/js/mdb.min.js" integrity="sha256-NdbiivsvWt7VYCt6hYNT3h/th9vSTL4EDWeGs5SN3DA=" crossorigin="anonymous"></script>
3
3<script defer src="https://cdn.jsdelivr.net/npm/[email protected]/dist/masonry.pkgd.min.js" integrity="sha256-Nn1q/fx0H7SNLZMQ5Hw5JLaTRZp0yILA/FRexe19VdI=" crossorigin="anonymous"></script>
3
3<script defer src="https://cdn.jsdelivr.net/npm/[email protected]/imagesloaded.pkgd.min.js" integrity="sha256-htrLFfZJ6v5udOG+3kNLINIKh2gvoKqwEhHYfTTMICc=" crossorigin="anonymous"></script>
3
3<script defer src="/assets/js/masonry.js?a0db7e5d5c70cc3252b3138b0c91dcaf" type="text/javascript"></script>
3
3<script defer src="https://cdn.jsdelivr.net/npm/[email protected]/dist/medium-zoom.min.js" integrity="sha256-ZgMyDAIYDYGxbcpJcfUnYwNevG/xi9OHKaR/8GK+jWc=" crossorigin="anonymous"></script>
3
3<script defer src="/assets/js/zoom.js?85ddb88934d28b74e78031fd54cf8308"></script>
3
3<script src="/assets/js/no_defer.js?2781658a0a2b13ed609542042a859126"></script>
3
3<script defer src="/assets/js/common.js?e0514a05c5c95ac1a93a8dfd5249b92e"></script>
3
3<script defer src="/assets/js/copy_code.js?c8a01c11a92744d44b093fc3bda915df" type="text/javascript"></script>
3
3<script defer src="/assets/js/jupyter_new_tab.js?d9f17b6adc2311cbabd747f4538bb15f"></script>
3
3<script async src="https://d1bxh8uas1mnw7.cloudfront.net/assets/embed.js"></script>
3
3<script async src="https://badge.dimensions.ai/badge.js"></script>
3
3<script defer type="text/javascript" id="MathJax-script" src="https://cdn.jsdelivr.net/npm/[email protected]/es5/tex-mml-chtml.js" integrity="sha256-MASABpB4tYktI2Oitl4t+78w/lyA+D7b/s9GEP0JOGI=" crossorigin="anonymous"></script>
3
3<script src="/assets/js/mathjax-setup.js?a5bb4e6a542c546dd929b24b8b236dfd"></script>
3
3<script defer src="https://cdnjs.cloudflare.com/polyfill/v3/polyfill.min.js?features=es6" crossorigin="anonymous"></script>
3
3<script defer src="/assets/js/progress-bar.js?2f30e0e6801ea8f5036fa66e1ab0a71a" type="text/javascript"></script>
3
3<script src="/assets/js/vanilla-back-to-top.min.js?f40d453793ff4f64e238e420181a1d17"></script>
3
3<script> 4 addBackToTop(); 5 </script>
5
5<script type="module" src="/assets/js/search/ninja-keys.min.js?a3446f084dcaecc5f75aa1757d087dcf"></script>
5 <ninja-keys hidebreadcrumbs noautoloadmdicons placeholder="Type to start searching"></ninja-keys>
5<script src="/assets/js/search-setup.js?6c304f7b1992d4b60f7a07956e52f04a"></script>
5
5<script src="/assets/js/search-data.js"></script>
5
5<script src="/assets/js/shortcut-key.js?6f508d74becd347268a7f822bca7309d"></script>
5 </body> </html>
Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.