1<!DOCTYPE HTML> 2<html lang="en"> 3 4<head> 5 <title>Yuanhao Cai</title> 6 7 <meta content="text/html; charset=utf-8" http-equiv="Content-Type"> 8 9 <meta name="author" content="Yuanhao Cai" /> 10 <meta name="viewport" content="width=device-width, initial-scale=1"> 11 12 <link rel="stylesheet" type="text/css" href="/style.css" /> 13 <link rel="canonical" href="https://caiyuanhao.github.io/"> 14 <link href="https://fonts.googleapis.com/css?family=Lato:400,700,400italic,700italic" rel="stylesheet" type="text/css"> 15 16</head> 17 18 19 20<body> 21 <table style="width:100%;max-width:1000px;border:0px;border-spacing:0px;border-collapse:separate;margin-right:auto;margin-left:auto;"> 22 <tr style="padding:0px"> 23 <td style="padding:0px"> 24 <table style="width:100%;border:0px;border-spacing:0px;border-collapse:separate;margin-right:auto;margin-left:auto;"> 25 <tr style="padding:0px"> 26 <td style="padding:2.5%;width:65%;vertical-align:middle"> 27 <h1> 28 Yuanhao Cai 29 </h1> 30 <p>I am currently a 3rd year PhD student in the department of <a href="https://www.cs.jhu.edu/">Computer Science</a>, <a href="https://www.jhu.edu/">Johns Hopkins University</a>. 31 I am a member of <a href="https://ccvl.jhu.edu/">CCVL</a>, advised by <a href="https://en.wikipedia.org/wiki/Bloomberg_Distinguished_Professorships">Bloomberg Distinguished Professor</a> 32 <a href="https://www.cs.jhu.edu/~ayuille/"> Dr. Alan Yuille</a>. 33 Previously, I received my MSE and BSE degrees from <a href="https://www.tsinghua.edu.cn/en/">Tsinghua University</a> in 2023 and 2020. My master advisor is prof. <a href="https://scholar.google.com/citations?user=eldgnIYAAAAJ&hl=zh-CN">Haoqian Wang</a>. 34 During my study in Tsinghua University, I spent a good time with prof. <a href="https://yulunzhang.com/">Yulun Zhang</a>, prof. <a href="https://en.westlake.edu.cn/faculty/xin-yuan.html">Xin Yuan</a>, prof. <a href="https://www.informatik.uni-wuerzburg.de/computervision/">Radu Timofte</a>, 35 and prof. <a href="https://ee.ethz.ch/the-department/faculty/professors/person-detail.OTAyMzM=.TGlzdC80MTEsMTA1ODA0MjU5.html">Luc Van Gool</a>. I interned in Adobe Research (2024 - 2025) and Meta Superintelligence Labs (2025 - 2026). 36 </p> 37 <p style="text-align:center"> 38 <a href="https://github.com/caiyuanhao1998">GitHub</a> / 39 <a href="https://scholar.google.com/citations?user=3YozQwcAAAAJ&hl=en">Google Scholar</a> / 40 <a href="https://www.linkedin.com/in/yuanhao-cai-b8463b297/">Linkdin</a> / 41 <a href="https://www.zhihu.com/people/cyh-28-29">ZhiHu</a> 42 </p> 43 <p>I am open to any opportunities of collaboration. Please feel free to contact me if you are interested. 44 </p> 45 </td> 46 <td style="padding:2.5%;width:35%;max-width:35%"> 47 <img style="width:100%;max-width:100%" alt="self photo" src="images/avatar_2026_04_01.jpg"/> 48 </td> 49 </tr> 50 </table> 51 52 <table style="width:100%;border:0px;border-spacing:0px;border-collapse:separate;margin-right:auto;margin-left:auto;"> 53 <h2>Honor and Award</h2> 54 <tr style="padding:0px"> 55 <td style="padding:2.5%;vertical-align:left"> 56 <p> 57 <a href="images/top2_scientist.jpg">Stanford/Elsevier Worldâs Top 2% Scientists</a>, 2025 58 </p> 59 <p> 60 <a href="images/AI4CC_best_paper.png">Best Paper Award</a> of <a href="https://ai4cc.net/">AI4CC Workshop</a> at CVPR, 2025 61 </p> 62 <p> 63 Runner-up Award of <a href="https://codalab.lisn.upsaclay.fr/competitions/17640">NTIRE Low Light Enhancement Challenge</a> at CVPR, 2024 64 </p> 65 <p> 66 <a href="images/cvpr24_outstanding_reviewers.jpg">Outstanding Reviewer Award</a> at CVPR, 2024 67 </p> 68 <p> 69 <a href="https://iclr.cc/Conferences/2024/Reviewers#outstanding-reviewers">Outstanding Reviewer Award</a> at ICLR, 2024 70 </p> 71 <p> 72 Excellent Graduate of <a href="https://www.tsinghua.edu.cn/">Tsinghua University</a>, 2023 73 </p> 74 <p> 75 Excellent Master Thesis Award, <a href="https://www.tsinghua.edu.cn/">Tsinghua University</a>, 2023 76 </p> 77 <p> 78 <a href="https://cvpr2023.thecvf.com/Conferences/2023/OutstandingReviewers">Outstanding Reviewer Award</a> at CVPR, 2023 79 </p> 80 <p> 81 <a href="images/Top_Talented.pdf">Top-10 Talented Graduate Student Finalist</a>, <a href="https://www.tsinghua.edu.cn/">Tsinghua University</a>, 2022 82 </p> 83 <p> 84 National Scholarship, <a href="https://www.tsinghua.edu.cn/">Tsinghua University</a>, 2022 85 </p> 86 <p> 87 Winner of <a href="https://codalab.lisn.upsaclay.fr/competitions/721">NTIRE Spectral Reconstruction Challenge</a> at CVPR, 2022 88 </p> 89 <p> 90 National Scholarship, <a href="https://www.tsinghua.edu.cn/">Tsinghua University</a>, 2021 91 </p> 92 <p> 93 Winner of <a href="https://cocodataset.org/workshop/coco-lvis-eccv-2020.html">COCO Keypoint Detection Challenge</a> at ECCV, 2020 94 </p> 95 <p> 96 Winner of <a href="https://cocodataset.org/workshop/coco-mapillary-iccv-2019.html">COCO Keypoint Detection Challenge</a> at ICCV, 2019 97 </p> 98 <p> 99 Best Paper Award of <a href="https://cocodataset.org/workshop/coco-mapillary-iccv-2019.html">Joint COCO and Mapillary Recognition Challenge Workshop</a> at ICCV, 2019 100 </p> 101 <p> 102 2nd Place of AdultSize Drop-In Challenge at <a href="https://2019.robocup.org/">RoboCup World Final</a>, 2019 103 </p> 104 <p> 105 2nd Place of AdultSize Technical Challenge at <a href="https://2019.robocup.org/">RoboCup World Final</a>, 2019 106 </p> 107 <p> 108 3rd Place of AdultSize Soccer Competition at <a href="https://2019.robocup.org/">RoboCup World Final</a>, 2019 109 </p> 110 <p> 111 Learning Progress Scholarship, <a href="https://www.tsinghua.edu.cn/">Tsinghua University</a>, 2019 112 </p> 113 <p> 114 Science and Technology Innovation Excellence Scholarship, <a href="https://www.tsinghua.edu.cn/">Tsinghua University</a>, 2019 115 </p> 116 </td> 117 </tr> 118 </table> 119 120 <table style="width:100%;border:0px;border-spacing:0px;border-collapse:separate;margin-right:auto;margin-left:auto;"> 121 <h2>Publication</h2> ( * = Equal Contribution, â = Corresponding Author) 122 <br> 123 <br> 124 <br> 125 126 127 128 <tr> 129 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 130 <img src="/images/phygdpo.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 131 </td> 132 <td style="padding:2.5%;width:70%;vertical-align:middle"> 133 <h3>PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation</h3> 134 <br> 135 <br> 136 <strong>Yuanhao Cai</strong>, <a href="https://kunpengli1994.github.io/">Kunpeng Li</a>, <a href="https://kmnp.github.io/">Menglin Jia</a>, <a href="https://sites.google.com/view/jialiangwang/home">Jialiang Wang</a>, <a href="https://scholar.google.com/citations?user=wyi0bX0AAAAJ&hl=en">Junzhe Sun</a>, <a href="https://wfchen-umich.github.io/wfchen.github.io/">Weifeng Chen</a>, <a href="https://xujuefei.com/">Felix Juefei Xu</a>, <br>
136 <a href="https://scholar.google.com/citations?user=5aaOtscAAAAJ&hl=en">Chu Wang</a>, <a href="https://www.alithabet.com/">Ali Thabet</a>, <a href="https://sites.google.com/view/xiaoliangdai">Xiaoliang Dai</a>, <a href="https://juxuan27.github.io/">Xuan Ju</a>, <a href="https://www.cs.jhu.edu/~ayuille/">Alan Yuille</a>, <a href="https://sekunde.github.io/">Ji Hou</a> 137 138 <br> 139 <em>European Conference on Computer Vision (ECCV)</em>, 2026 140 <br> 141 142 <a href="https://arxiv.org/abs/2512.24551">paper</a> 143 144 145 / <a href="https://caiyuanhao1998.github.io/project/PhyGDPO">project</a> 146 147 148 149 150 151 152 153 <p></p> 154 <p>A data construction pipeline and a new diret preference optimization framework for physically consistant video generation</p> 155 156 </td> 157 </tr> 158 159 160 161 162 163 <tr> 164 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 165 <img src="/images/omnivcus.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 166 </td> 167 <td style="padding:2.5%;width:70%;vertical-align:middle"> 168 <h3>OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions</h3> 169 <br> 170 <br> 171 <strong>Yuanhao Cai</strong>, <a href="https://sites.google.com/site/hezhangsprinter">He Zhang</a>, <a href="https://xavierchen34.github.io/">Xi Chen</a>, <a href="https://doubiiu.github.io/">Jinbo Xing</a>, <a href="https://yiweihu.netlify.app/">Yiwei Hu</a>, <a href="https://yzhouas.github.io/">Yuqian Zhou</a>, <a href="https://kai-46.github.io/website/">Kai Zhang</a>, <a href="https://zzutk.github.io/">Zhifei Zhang</a>, <a href="https://sites.google.com/view/sooyekim">Soo Ye Kim</a>, <a href="https://stevewongv.github.io/">Tianyu Wang</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="https://english.seiee.sjtu.edu.cn/english/detail/842_802.htm">Xiaokang Yang</a>, <a href="https://sites.google.com/site/zhelin625/">Zhe Lin</a>, <a href="https://www.cs.jhu.edu/~ayuille/">Alan Yuille</a> 172 173 <br> 174 <em>Advances in Neural Information Processing Systems (NeurIPS)</em>, 2025 175 <br> 176 177 <a href="http://arxiv.org/abs/2506.23361">paper</a> 178 179 180 / <a href="https://caiyuanhao1998.github.io/project/OmniVCus">project</a> 181 182 183 184 185 186 187 188 <p></p> 189 <p>A data construction pipeline and a unified diffusion Transformer for subject-driven video customization under different control conditions; 4D generation</p> 190 191 </td> 192 </tr> 193 194 195 196 197 198 <tr> 199 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 200 <img src="/images/editverse.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 201 </td> 202 <td style="padding:2.5%;width:70%;vertical-align:middle"> 203 <h3>EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning</h3> 204 <br> 205 <br> 206 <a href="https://juxuan27.github.io/">Xuan Ju</a>, <a href="https://scholar.google.com/citations?user=yRwZIN8AAAAJ&hl=zh-CN">Tianyu Wang</a>, <a href="https://yzhouas.github.io/">Yuqian Zhou</a>, <a href="https://sites.google.com/site/hezhangsprinter/">He Zhang</a>, <a href="https://qliu24.github.io/">Qing Liu</a>, <a href="https://www.nxzhao.com/">Nanxuan Zhao</a>, <a href="https://zzutk.github.io/">Zhifei Zhang</a>, <a href="https://yijunmaverick.github.io/">Yijun Li</a>, <strong>Yuanhao Cai</strong>, <a href="https://www.shaotengliu.com/">Shaoteng Liu</a>, <a href="https://scholar.google.com/citations?user=UI10l34AAAAJ&hl=en">Daniil Pakhomov</a>, <a href="https://scholar.google.com/citations?user=UI10l34AAAAJ&hl=en">Daniil Pakhomov</a>, <a href="https://sites.google.com/site/zhelin625/">Zhe Lin</a>, <a href="https://sites.google.com/view/sooyekim">Soo Ye Kim</a>, <a href="https://www.cse.cuhk.edu.hk/people/faculty/qiang-xu/">Qiang Xu</a> 207 208 <br> 209 <em>International Conference on Representation Learning (ICLR)</em>, 2025 210 <br> 211 212 <a href="https://arxiv.org/abs/2509.20360">paper</a> 213 214
215 / <a href="http://editverse.s3-website-us-east-1.amazonaws.com/">project</a> 216 217 218 219 220 221 222 223 <p></p> 224 <p>Unifying a diverse range of generation and editing tasks for both images and videos within a single and powerful model</p> 225 226 </td> 227 </tr> 228 229 230 231 232 233 <tr> 234 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 235 <img src="/images/langsplatv2.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 236 </td> 237 <td style="padding:2.5%;width:70%;vertical-align:middle"> 238 <h3>LangSplatV2: High-dimensional 3D Language Gaussian Splatting with 450+ FPS</h3> 239 <br> 240 <br> 241 <a href="https://li-wanhua.github.io/">Wanhua Li*</a>, <a href="https://github.com/ZhaoYujie2002">Yujie Zhao*</a>, <a href="https://minghanqin.github.io/">Minghan Qin*</a>, <a href="https://github.com/jimmyYliu">Yang Liu</a>, <strong>Yuanhao Cai</strong>, <a href="https://people.csail.mit.edu/ganchuang/">Chuang Gan</a>, <a href="https://vcg.seas.harvard.edu/people">Hanspeter Pfister</a> 242 243 <br> 244 <em>Advances in Neural Information Processing Systems (NeurIPS)</em>, 2025 245 <br> 246 247 <a href="https://arxiv.org/abs/2507.07136">paper</a> 248 249 250 / <a href="https://langsplat-v2.github.io/">project</a> 251 252 253 254 255 256 257 258 <p></p> 259 <p>A super fast framework for real-time 3D open-vocabulary querying and high-dimensional feature splatting</p> 260 261 </td> 262 </tr> 263 264 265 266 267 268 <tr> 269 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 270 <img src="/images/care.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 271 </td> 272 <td style="padding:2.5%;width:70%;vertical-align:middle"> 273 <h3>Are Pixel-Wise Metrics Reliable for Sparse-View Computed Tomography Reconstruction?</h3> 274 <br> 275 <br> 276 <a href="https://lin-tianyu.github.io/">Tianyu Lin</a>, Xinran Li, Chuntung Zhuang, <a href="https://scholar.google.com/citations?user=4Q5gs2MAAAAJ&hl=en">Qi Chen</a>, <strong>Yuanhao Cai</strong>, <a href="https://scholar.google.com/citations?user=OvpsAYgAAAAJ&hl=en&oi=ao">Kai Ding</a>, <a href="https://www.cs.jhu.edu/~ayuille/">Alan Yuille</a>, <a href="https://www.zongweiz.com/">Zongwei Zhou</a> 277 278 <br> 279 <em>Advances in Neural Information Processing Systems (NeurIPS)</em>, 2025 280 <br> 281 282 <a href="https://arxiv.org/abs/2506.02093">paper</a> 283 284 285 / <a href="https://github.com/MrGiovanni/CARE">project</a> 286 287 288 289 290 291 292 293 <p></p> 294 <p>New metrics and diffusion based anatomy-aware enhancement for sparse-view CT reconstruction</p> 295 296 </td> 297 </tr> 298 299 300 301 302 303 <tr> 304 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 305 <img src="/images/xlrm.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 306 </td> 307 <td style="padding:2.5%;width:70%;vertical-align:middle"> 308 <h3>X-LRM: X-ray Large Reconstruction Model for Extremely Sparse-View Computed Tomography Recovery in One Second</h3> 309 <br> 310 <br> 311 <a href="https://scholar.google.com/citations?hl=en&user=vl0mzhEAAAAJ">Guofeng Zhang *</a>, <a href="https://ruyi-zha.github.io/">Ruyi Zha *</a>, <a href="https://heye0507.github.io/">Hao He</a>, <a href="https://yixunliang.github.io/">Yixun Liang</a>, <a href="https://www.cs.jhu.edu/~ayuille/">Alan Yuille</a>, <a href="https://users.cecs.anu.edu.au/~hongdong/">Hongdong Li</a>, <strong>Yuanhao Cai</strong> 312 313 <br> 314 <em>International Conference on 3D Vision (3DV)</em>, 2025 315 <br> 316 317 <a href="https://arxiv.org/abs/2503.06382">paper</a> 318 319 320 / <a href="https://github.com/caiyuanhao1998/X-LRM">project</a> 321 322 323 324 325 326 327 328 <p></p> 329 <p>A feedforward method and a large-scale dataset for instant CT reconstruction</p> 330 331 </td> 332 </tr> 333 334 335 336 337 338 <tr> 339 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 340 <img src="/images/diffusiongs.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 341 </td> 342 <td style="padding:2.5%;width:70%;vertical-align:middle"> 343 <h3>Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and Reconstruction</h3> 344 <br> 345 <br> 346 <strong>Yuanhao Cai</strong>, <a href="https://sites.google.com/site/hezhangsprinter">He Zhang</a>, <a href="https://kai-46.github.io/website/">Kai Zhang</a>, <a href="https://yixunliang.github.io/">Yixun Liang</a>, <a href="https://www.mengweiren.com/">Mengwei Ren</a>, <a href="https://luanfujun.com/">Fujun Luan</a>, <a href="https://qliu24.github.io/">Qing Liu</a>, <a href="https://sites.google.com/view/sooyekim">Soo Ye Kim</a>, <a href="https://jimmie33.github.io/">Jianming Zhang</a>, <a href="https://zzutk.github.io/">Zhifei Zhang</a>, <a href="https://yzhouas.github.io/">Yuqian Zhou</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="https://english.seiee.sjtu.edu.cn/english/detail/842_802.htm">Xiaokang Yang</a>, <a href="https://sites.google.com/site/zhelin625/">Zhe Lin</a>, <a href="https://www.cs.jhu.edu/~ayuille/">Alan Yuille</a> 347 348 <br> 349 <em>International Conference on Computer Vision (ICCV)</em>, 2025 350 <br> 351 352 <a href="https://arxiv.org/abs/2411.14384">paper</a> 353 354 355 / <a href="https://caiyuanhao1998.github.io/project/DiffusionGS/">project</a> 356 357 358 359 360 361 362 / <a href="https://x.com/janusch_patas/status/1859867424859856997?ref_src=twsrc%5Egoogle%7Ctwcamp%5Eserp%7Ctwgr%5Etweet">media (MrNeRF)</a> 363 364 365 <p></p> 366 <p>A 3DGS diffusion generates objects and reconstructs scenes from a single view in 6s</p> 367 368 </td> 369 </tr> 370 371 372 373 374 375 <tr> 376 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 377 <img src="/images/x2gs.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 378 </td> 379 <td style="padding:2.5%;width:70%;vertical-align:middle"> 380 <h3>4D Radiative Gaussian Splatting for Continuous-time Tomographic Reconstruction</h3> 381 <br> 382 <br> 383 <a href="https://scholar.google.com/citations?hl=en&user=fCzlLE4AAAAJ">Weihao Yu</a>, <strong>Yuanhao Cai</strong>, <a href="https://ruyi-zha.github.io/">Ruyi Zha</a>, <a href="https://zhiwenfan.github.io/">Zhiwen Fan</a>, <a href="https://chenxinli001.github.io/">Chenxin Li</a>, <a href="https://www.ee.cuhk.edu.hk/~yxyuan/">Yixuan Yuan</a> 384 385 <br> 386 <em>International Conference on Computer Vision (ICCV)</em>, 2025 387 <br> 388 389 <a href="https://arxiv.org/abs/2503.21779">paper</a> 390 391 392 / <a href="https://x2-gaussian.github.io/">project</a> 393 394 395 396 397 398 399 / <a href="https://x.com/janusch_patas/status/1905528815302087135">media (MrNeRF)</a> 400 401 402 <p></p> 403 <p>A 4D Gaussian Splatting method for dynamic CT reconstruction</p> 404 405 </td> 406 </tr> 407 408 409 410 411 412 <tr> 413 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 414 <img src="/images/motionx_plus.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 415 </td> 416 <td style="padding:2.5%;width:70%;vertical-align:middle"> 417 <h3>Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset</h3> 418 <br> 419 <br> 420 Yuhong Zhang, <a href="https://jinglin7.github.io/">Jing Lin</a>, <a href="https://ailingzeng.site/">Ailing Zeng</a>, <a href="https://guanlinwu123.github.io/">Guanlin Wu</a>, <a href="https://shunlinlu.github.io/">Shunlin Lu</a>, Yurong Fu, <strong>Yuanhao Cai</strong>, <a href="http://zhangruimao.site/">Ruimao Zhang</a>, <a href="https://scholar.google.com/citations?user=eldgnIYAAAAJ&hl=zh-CN">Haoqian Wang</a>, <a href="https://www.leizhang.org/">Lei Zhang</a> 421 422 <br> 423 <em>arxiv</em>, 2024 424 <br> 425 426 <a href="https://arxiv.org/abs/2501.05098">paper</a> 427 428 429 / <a href="https://motion-x-dataset.github.io/">project</a> 430 431 432 433 434 435 436 437 <p></p> 438 <p>A Large-Scale Dataset for Multimodal 3D Whole-body Human Motion Generation</p> 439 440 </td> 441 </tr> 442 443 444 445 446 447 <tr> 448 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 449 <img src="/images/videolifter.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 450 </td> 451 <td style="padding:2.5%;width:70%;vertical-align:middle"> 452 <h3>VideoLifter: Lifting Videos to 3D with Fast Hierarchical Stereo Alignment</h3> 453 <br> 454 <br> 455 <a href="https://wenyancong.com/">Wenyan Cong</a>, <a href="https://www.kevin-ai.com/">Kevin Wang</a>, <a href="https://www.cis.upenn.edu/~leijh/">Jiahui Lei</a>, <a href="https://coltonstearns.github.io/">Colton Stearns</a>, <strong>Yuanhao Cai</strong>, <a href="https://wdilin.github.io/">Dilin Wang</a>, <a href="https://www.linkedin.com/in/rakesh-r-3848538/">Rakesh Ranjan</a>, <a href="https://scholar.google.com/citations?user=A-wA73gAAAAJ&hl=en">Matt Feiszli</a>, <a href="https://profiles.stanford.edu/leonidas-guibas">Leonidas Guibas</a>, <a href="https://express.adobe.com/page/CAdrFMJ9QeI2y/">Zhangyang Wang</a>, <a href="https://scholar.google.com/citations?user=d1-hNQoAAAAJ&hl=en">Weiyao Wang</a>, <a href="https://zhiwenfan.github.io/">Zhiwen Fan</a> 456 457 <br> 458 <em>arxiv</em>, 2024 459 <br> 460 461 <a href="https://arxiv.org/abs/2501.01949">paper</a> 462 463 464 / <a href="https://videolifter.github.io/">project</a> 465 466 467 468 469 470 471 472 <p></p> 473 <p>A 3DGS-based reconstruction method for videos without SfM initilization</p> 474 475 </td> 476 </tr> 477 478 479 480 481 482 <tr> 483 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 484 <img src="/images/lucidfusion.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 485 </td> 486 <td style="padding:2.5%;width:70%;vertical-align:middle"> 487 <h3>LucidFusion: Generating 3D Gaussians with Arbitrary Unposed Images</h3> 488 <br> 489 <br> 490 <a href="https://heye0507.github.io/">Hao He</a>, <a href="https://yixunliang.github.io/">Yixun Liang</a>, <a href="https://wileewang.github.io/">Luozhou Wang</a>, <strong>Yuanhao Cai</strong>, <a href="https://scholar.google.com/citations?user=lrgPuBUAAAAJ&hl=en&inst=1381320739207392350/">
490Xinli Xu</a>, Hao-Xiang Guo, Xiang Wen, <a href="https://www.yingcong.me/">Ying-Cong Chen</a> 491 492 <br> 493 <em>arxiv</em>, 2024 494 <br> 495 496 <a href="https://arxiv.org/abs/2410.15636">paper</a> 497 498 499 / <a href="https://heye0507.github.io/LucidFusion_page/">project</a> 500 501 502 503 504 505 506 507 <p></p> 508 <p>A pose-free multi-view 3D reconstruction methods based on Gaussian Splatting</p> 509 510 </td> 511 </tr> 512 513 514 515 516 517 <tr> 518 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 519 <img src="/images/hdr_gs.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 520 </td> 521 <td style="padding:2.5%;width:70%;vertical-align:middle"> 522 <h3>HDR-GS: Efficient High Dynamic Range Novel View Synthesis at 1000x Speed via Gaussian Splatting</h3> 523 <br> 524 <br> 525 <strong>Yuanhao Cai </strong>, <a href="https://scholar.google.com/citations?user=ucb6UssAAAAJ&hl=en">Zihao Xiao</a>, <a href="https://yixunliang.github.io/">Yixun Liang</a>, <a href="https://minghanqin.github.io/">Minghan Qin</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="https://english.seiee.sjtu.edu.cn/english/detail/842_802.htm">Xiaokang Yang</a>, <a href="https://www.cs.jhu.edu/~yyliu/">Yaoyao Liu</a>, <a href="https://www.cs.jhu.edu/~ayuille/">Alan Yuille</a> 526 527 <br> 528 <em>Advances in Neural Information Processing Systems (NeurIPS)</em>, 2024 529 <br> 530 531 <a href="https://arxiv.org/pdf/2405.15125">paper</a> 532 533 534 / <a href="https://github.com/caiyuanhao1998/HDR-GS">project</a> 535 536 537 / <a href="https://www.youtube.com/watch?v=wtU7Kcwe7ck">video</a> 538 539 540 / <a href="https://zhuanlan.zhihu.com/p/10016024329">zhihu</a> 541 542 543 / <a href="https://paperswithcode.com/sota/novel-view-synthesis-on-hdr-gs">leaderboard</a> 544 545 546 / <a href="https://x.com/_akhaliq/status/1794921228462923925?s=46">media (AK)</a> 547 548 549 / <a href="https://x.com/janusch_patas/status/1794932286397489222?s=46">media (MrNeRF)</a> 550 551 552 / <a href="/bibtex/hdr_gs.txt">bibtex</a> 553 554 <p></p> 555 <p>The first 3D Gaussian splatting-based method for high dynamic range imaging</p> 556 557 </td> 558 </tr> 559 560 561 562 563 564 <tr> 565 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 566 <img src="/images/r2gs.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 567 </td> 568 <td style="padding:2.5%;width:70%;vertical-align:middle"> 569 <h3>R2-Gaussian: Rectifying Radiative Gaussian Splatting for Tomographic Reconstruction</h3> 570 <br> 571 <br> 572 <a href="https://ruyi-zha.github.io/">Ruyi Zha</a>, <a href="https://scholar.google.com/citations?user=BD7Hce0AAAAJ&hl=en">Tao Jun Lin</a>, <strong>Yuanhao Caiâ </strong>, <a href="https://sites.google.com/view/yanhaozhang/home">Jiwen Cao</a>, <a href="https://users.cecs.anu.edu.au/~hongdong/">Hongdong Li</a> 573 574 <br> 575 <em>Advances in Neural Information Processing Systems (NeurIPS)</em>, 2024 576 <br> 577 578 <a href="https://arxiv.org/abs/2405.20693v1">paper</a> 579 580 581 / <a href="https://ruyi-zha.github.io/r2_gaussian/r2_gaussian.html">project</a> 582 583 584 585 586 587 588 / <a href="https://x.com/janusch_patas/status/1797493143094571041">media (MrNeRF)</a> 589 590 591 <p></p> 592 <p>The first 3D Gaussian splatting-based method for CT reconstruction</p> 593 594 </td> 595 </tr> 596 597 598 599 600 601 <tr> 602 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 603 <img src="/images/x-gaussian.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 604 </td> 605 <td style="padding:2.5%;width:70%;vertical-align:middle"> 606 <h3>Radiative Gaussian Splatting for Efficient X-ray Novel View Synthesis</h3> 607 <br> 608 <br> 609 <strong>Yuanhao Cai </strong>, <a href="https://yixunliang.github.io/">Yixun Liang</a>, <a href="https://jiahaoplus.github.io/">jiahao Wang</a>, <a href="https://scholar.google.com/citations?hl=en&user=YR7re-cAAAAJ">Angtian Wang</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="https://english.seiee.sjtu.edu.cn/english/detail/842_802.htm">Xiaokang Yang</a>, <a href="https://www.zongweiz.com/">Zongwei Zhou</a>, <a href="https://www.cs.jhu.edu/~ayuille/">Alan Yuille</a> 610 611 <br> 612 <em>European Conference on Computer Vision (ECCV)</em>, 2024 613 <br> 614 615 <a href="https://arxiv.org/pdf/2403.04116.pdf">paper</a> 616 617 618 / <a href="https://github.com/caiyuanhao1998/X-Gaussian">project</a> 619 620 621 / <a href="https://www.youtube.com/watch?v=gDVf_Ngeghg">video</a> 622 623 624 / <a href="https://zhuanlan.zhihu.com/p/717744222">zhihu</a> 625 626 627 628 / <a href="https://x.com/_akhaliq/status/1765929288044290253?s=46">media (AK)</a> 629 630 631 / <a href="https://x.com/janusch_patas/status/1766446189749150126?s=46">media (MrNeRF)</a> 632 633 634 / <a href="/bibtex/x_gaussian.txt">bibtex</a> 635 636 <p></p> 637 <p>The first 3D Gaussian splatting-based method for X-ray 3D reconstruction</p> 638 639 </td> 640 </tr> 641 642 643 644 645 646 <tr> 647 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 648 <img src="/images/acca.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 649 </td> 650 <td style="padding:2.5%;width:70%;vertical-align:middle"> 651 <h3>Unveiling Advanced Frequency Disentanglement Paradigm for Low-Light Image Enhancement</h3> 652 <br> 653 <br> 654 <a href="https://yixunliang.github.io/">Kun Zhou</a>, <a href="https://sds.cuhk.edu.cn/en/node/678">Xinyu Lin</a>, <a href="https://fenglinglwb.github.io/">Wenbo Li</a>, <a href="https://xuxiaogang.com/">Xiaogang Xu</a>, <strong>Yuanhao Cai </strong>, <a href="https://zhonghang-liu.github.io/homepage/">Zhonghang Liu</a>, <a href="https://scholar.google.com/citations?user=z-rqsR4AAAAJ&hl=zh-CN">Xiaoguang Han</a>, <a href="https://sites.google.com/site/jiangbolu/">Jiangbo Lu</a> 655 656 <br> 657 <em>European Conference on Computer Vision (ECCV)</em>, 2024 658 <br> 659 660 <a href="https://arxiv.org/pdf/2403.04116.pdf">paper</a> 661 662 663 / <a href="https://github.com/redrock303/ADF-LLIE">project</a> 664 665 666 667 668 669 670 671 <p></p> 672 <p>A light-weight network for low-light image enhancement</p> 673 674 </td> 675 </tr> 676 677 678 679 680 681 <tr> 682 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 683 <img src="/images/sax-nerf.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 684 </td> 685 <td style="padding:2.5%;width:70%;vertical-align:middle"> 686 <h3>Structure-Aware Sparse-View X-ray 3D Reconstruction</h3> 687 <br> 688 <br> 689 <strong>Yuanhao Cai </strong>, <a href="https://jiahaoplus.github.io/">jiahao Wang</a>, <a href="https://www.zongweiz.com/">Zongwei Zhou</a>, <a href="https://scholar.google.com/citations?hl=en&user=YR7re-cAAAAJ">Angtian Wang</a>, <a href="https://www.cs.jhu.edu/~ayuille/">Alan Yuille</a> 690 691 <br> 692 <em>Conference on Computer Vision and Pattern Recognition (CVPR)</em>, 2024 693 <br> 694 695 <a href="https://arxiv.org/pdf/2311.10959.pdf">paper</a> 696 697 698 / <a href="https://github.com/caiyuanhao1998/SAX-NeRF">project</a> 699 700 701 / <a href="https://www.youtube.com/watch?v=oVVUaBY61eo">video</a> 702 703 704 / <a href="https://zhuanlan.zhihu.com/p/702702109">zhihu</a> 705 706 707 / <a href="https://paperswithcode.com/dataset/x3d">leaderboard</a> 708 709 710 711 712 / <a href="/bibtex/saxnerf.txt">bibtex</a> 713 714 <p></p> 715 <p>A NeRF algorithm capturing structures for large-scale X-ray 3D reconstruction</p> 716 717 </td> 718 </tr> 719 720 721 722 723 724 <tr> 725 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 726 <img src="/images/bisci.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 727 </td> 728 <td style="padding:2.5%;width:70%;vertical-align:middle"> 729 <h3>Binarized Spectral Compressive Imaging</h3> 730 <br> 731 <br> 732 <strong>Yuanhao Cai </strong>, Yuxin Zheng, <a href="https://jinglin7.github.io/">Jing Lin</a>, <a href="https://en.westlake.edu.cn/faculty/xin-yuan.html">Xin Yuan</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a> 733 734 <br> 735 <em>Advances in Neural Information Processing Systems (NeurIPS)</em>, 2023 736 <br> 737 738 <a href="https://arxiv.org/pdf/2305.10299.pdf">paper</a> 739 740 741 / <a href="https://github.com/caiyuanhao1998/BiSCI">project</a> 742 743 744 745 / <a href="https://zhuanlan.zhihu.com/p/668862020">zhihu</a> 746 747 748 749 750 751 / <a href="/bibtex/bisci.txt">bibtex</a> 752 753 <p></p> 754 <p>An Efficient Retinex-based method for Low-light Image Enhancement</p> 755 756 </td> 757 </tr> 758 759 760 761 762 763 <tr> 764 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 765 <img src="/images/motionx.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 766 </td> 767 <td style="padding:2.5%;width:70%;vertical-align:middle"> 768 <h3>
768Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset</h3> 769 <br> 770 <br> 771 <a href="https://jinglin7.github.io/">Jing Lin *</a>, <a href="https://ailingzeng.site/">Ailing Zeng *</a>, <a href="https://shunlinlu.github.io/">Shunling Lu *</a>, <strong>Yuanhao Cai </strong>, <a href="http://www.zhangruimao.site/">Ruimao Zhang</a>, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://www.leizhang.org/">Lei Zhang</a> 772 773 <br> 774 <em>Advances in Neural Information Processing Systems (NeurIPS)</em>, 2023 775 <br> 776 777 <a href="https://openreview.net/attachment?id=WtajAo0JWU&name=pdf">paper</a> 778 779 780 / <a href="https://motion-x-dataset.github.io/">project</a> 781 782 783 784 785 786 787 788 / <a href="/bibtex/motionx.txt">bibtex</a> 789 790 <p></p> 791 <p>A large-scale human motion benchmark with text description</p> 792 793 </td> 794 </tr> 795 796 797 798 799 800 <tr> 801 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 802 <img src="/images/retinexformer.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 803 </td> 804 <td style="padding:2.5%;width:70%;vertical-align:middle"> 805 <h3>Retinexformer: One-stage Retinex-based Transformer for Low-light Image Enhancement</h3> 806 <br> 807 <br> 808 <strong>Yuanhao Cai </strong>, <a href="https://bianhao123.github.io/">Hao Bian</a>, Jing Lin, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="http://people.ee.ethz.ch/~timofter/">Radu Timofte</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a> 809 810 <br> 811 <em>International Conference on Computer Vision (ICCV)</em>, 2023 812 <br> 813 814 <a href="https://arxiv.org/pdf/2303.06705.pdf">paper</a> 815 816 817 / <a href="https://github.com/caiyuanhao1998/Retinexformer">project</a> 818 819 820 821 / <a href="https://zhuanlan.zhihu.com/p/657927878">zhihu</a> 822 823 824 825 826 827 / <a href="/bibtex/retinexformer.txt">bibtex</a> 828 829 <p></p> 830 <p>An Efficient Retinex-based method for Low-light Image Enhancement</p> 831 832 </td> 833 </tr> 834 835 836 837 838 839 <tr> 840 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 841 <img src="/images/DAUHST.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 842 </td> 843 <td style="padding:2.5%;width:70%;vertical-align:middle"> 844 <h3>Degradation-Aware Unfolding Half-Shuffle Transformer for Spectral Compressive Imaging</h3> 845 <br> 846 <br> 847 <strong>Yuanhao Cai *</strong>, Jing Lin *, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://www.bell-labs.com/about/researcher-profiles/xyuan/">Xin Yuan</a>, <a href="https://henghuiding.github.io/">Henghui Ding</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="http://people.ee.ethz.ch/~timofter/">Radu Timofte</a>, <a href="https://ee.ethz.ch/the-department/faculty/professors/person-detail.OTAyMzM=.TGlzdC80MTEsMTA1ODA0MjU5.html">Luc Van Gool</a> 848 849 <br> 850 <em>Advances in Neural Information Processing Systems (NeurIPS)</em>, 2022 851 <br> 852 853 <a href="https://arxiv.org/pdf/2205.10102.pdf">paper</a> 854 855 856 / <a href="https://github.com/caiyuanhao1998/MST">project</a> 857 858 859 860 / <a href="https://zhuanlan.zhihu.com/p/576280023">zhihu</a> 861 862 863 864 865 866 / <a href="/bibtex/dauhst.txt">bibtex</a> 867 868 <p></p> 869 <p>The first Transformer-based deep unfolding method for spectral compressive imaging</p> 870 871 </td> 872 </tr> 873 874 875 876 877 878 <tr> 879 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 880 <img src="/images/CST.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 881 </td> 882 <td style="padding:2.5%;width:70%;vertical-align:middle"> 883 <h3>Coarse-to-Fine Sparse Transformer for Hyperspectral Image Reconstruction</h3> 884 <br> 885 <br> 886 <strong>Yuanhao Cai *</strong>, Jing Lin *, Xiaowan Hu, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://www.bell-labs.com/about/researcher-profiles/xyuan/">Xin Yuan</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="http://people.ee.ethz.ch/~timofter/">Radu Timofte</a>, <a href="https://ee.ethz.ch/the-department/faculty/professors/person-detail.OTAyMzM=.TGlzdC80MTEsMTA1ODA0MjU5.html">Luc Van Gool</a> 887 888 <br> 889 <em>European Conference on Computer Vision (ECCV)</em>, 2022 890 <br> 891 892 <a href="https://arxiv.org/pdf/2203.04845.pdf">paper</a> 893 894 895 / <a href="https://github.com/caiyuanhao1998/MST">project</a> 896 897 898 899 / <a href="https://zhuanlan.zhihu.com/p/544979161">zhihu</a> 900 901 902 903 904 905 / <a href="/bibtex/cst.txt">bibtex</a> 906 907 <p></p> 908 <p>A novel SOTA Transformer-based method for hyperspectral image reconstruction</p> 909 910 </td> 911 </tr> 912 913 914 915 916 917 <tr> 918 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 919 <img src="/images/FGST.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 920 </td> 921 <td style="padding:2.5%;width:70%;vertical-align:middle"> 922 <h3>Flow-Guided Sparse Transformer for Video Deblurring</h3> 923 <br> 924 <br> 925 Jing Lin *, <strong>Yuanhao Cai *</strong>, Xiaowan Hu, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://scholar.google.com.hk/citations?hl=zh-CN&user=JPUwfAMAAAAJ">Youliang Yan</a>, <a href="https://scholar.google.com.hk/citations?hl=zh-CN&user=0ua28KoAAAAJ">Xueyi Zou</a>, <a href="https://henghuiding.github.io/">Henghui Ding</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="http://people.ee.ethz.ch/~timofter/">Radu Timofte</a>, <a href="https://ee.ethz.ch/the-department/faculty/professors/person-detail.OTAyMzM=.TGlzdC80MTEsMTA1ODA0MjU5.html">Luc Van Gool</a> 926 927 <br> 928 <em>International Conference on Machine Learning (ICML)</em>, 2022 929 <br> 930 931 <a href="https://arxiv.org/pdf/2201.01893.pdf">paper</a> 932 933 934 / <a href="https://github.com/linjing7/VR-Baseline">project</a> 935 936 937 938 939 940 941 942 / <a href="/bibtex/fgst.txt">bibtex</a> 943 944 <p></p> 945 <p>The first Transformer-based method for video deblurring</p> 946 947 </td> 948 </tr> 949 950 951 952 953 954 <tr> 955 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 956 <img src="/images/Seq2Seq.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 957 </td> 958 <td style="padding:2.5%;width:70%;vertical-align:middle"> 959 <h3>Unsupervised Flow-Aligned Sequence-to-Sequence Learning for Video Restoration</h3> 960 <br> 961 <br> 962 Jing Lin *, Xiaowan Hu *, <strong>Yuanhao Cai</strong>, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://scholar.google.com.hk/citations?hl=zh-CN&user=JPUwfAMAAAAJ">Youliang Yan</a>, <a href="https://scholar.google.com.hk/citations?hl=zh-CN&user=0ua28KoAAAAJ">Xueyi Zou</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="https://ee.ethz.ch/the-department/faculty/professors/person-detail.OTAyMzM=.TGlzdC80MTEsMTA1ODA0MjU5.html">Luc Van Gool</a> 963 964 <br> 965 <em>International Conference on Machine Learning (ICML)</em>, 2022 966 <br> 967 968 <a href="https://arxiv.org/pdf/2205.10195.pdf">paper</a> 969 970 971 / <a href="https://github.com/linjing7/VR-Baseline">project</a> 972 973 974 975 976 977 978 979 / <a href="/bibtex/seq2seq.txt">bibtex</a> 980 981 <p></p> 982 <p>The first Sequence-to-Sequence model for video restoration</p> 983 984 </td> 985 </tr> 986 987 988 989 990 991 <tr> 992 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 993 <img src="/images/compare_fig.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 994 </td> 995 <td style="padding:2.5%;width:70%;vertical-align:middle"> 996 <h3>MST++: Multi-stage Spectral-wise Transformer for Efficient Spectral Reconstruction</h3> 997 <br> 998 <br> 999 <strong>Yuanhao Cai *</strong>, Jing Lin *, Zudi Lin, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="https://vcg.seas.harvard.edu/people/hanspeter-pfister">Hanspeter Pfister</a>, <a href="http://people.ee.ethz.ch/~timofter/">Radu Timofte</a>, <a href="https://ee.ethz.ch/the-department/faculty/professors/person-detail.OTAyMzM=.TGlzdC80MTEsMTA1ODA0MjU5.html">Luc Van Gool</a> 1000 1001 <br> 1002 <em>Conference on Computer Vision and Pattern Recognition Workshop (CVPRW)</em>, 2022 1003 <br> 1004 1005 <a href="https://arxiv.org/pdf/2204.07908.pdf">paper</a> 1006 1007 1008 / <a href="https://github.com/caiyuanhao1998/MST-plus-plus">project</a> 1009 1010 1011 1012 / <a href="https://zhuanlan.zhihu.com/p/501101943?utm_source=wechat_session&utm_medium=social&utm_oi=980437177842446336&utm_content=group3_article&utm_campaign=shareopn">zhihu</a> 1013 1014 1015 1016 1017 1018 / <a href="/bibtex/mst_pp.txt">bibtex</a> 1019 1020 <p></p> 1021 <p>Winner of NTIRE 2022 Challenge on Spectral Reconstruction from RGB. The first Transformer-based method for spectral reconstruction. A baseline and toolbox.</p> 1022 1023 </td> 1024 </tr> 1025 1026 1027 1028 1029 1030 <tr> 1031 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 1032 <img src="/images/MST.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 1033 </td> 1034 <td style="padding:2.5%;width:70%;vertical-align:middle"> 1035 <h3>Mask-guided Spectral-wise Transformer for Efficient Hyperspectral Image Reconstruction</h3> 1036 <br> 1037 <br> 1038 <strong>Yuanhao Cai *</strong>, Jing Lin *, Xiaowan Hu, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://www.bell-labs.com/about/researcher-profiles/xyuan/">Xin Yuan</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="http://people.ee.ethz.ch/~timofter/">Radu Timofte</a>, <a href="https://ee.ethz.ch/the-department/faculty/professors/person-detail.OTAyMzM=.TGlzdC80MTEsMTA1ODA0MjU5.html">Luc Van Gool</a> 1039 1040 <br> 1041 <em>Conference on Computer Vision and Pattern Recognition (CVPR)</em>, 2022 1042 <br> 1043 1044 <a href="https://arxiv.org/pdf/2111.07910.pdf">paper</a> 1045 1046 1047 / <a href="https://github.com/caiyuanhao1998/MST">project</a> 1048 1049 1050 1051 / <a href="https://zhuanlan.zhihu.com/p/501101943?utm_source=wechat_session&utm_medium=social&utm_oi=980437177842446336&utm_content=group3_article&utm_campaign=shareopn">zhihu</a> 1052 1053 1054 1055 1056 1057 / <a href="/bibtex/mst.txt">bibtex</a> 1058 1059 <p></p> 1060 <p>The first Transformer-based method for hyperspectral image reconstruction</p> 1061 1062 </td> 1063 </tr> 1064 1065 1066 1067 1068 1069 <tr> 1070 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 1071 <img src="/images/HDNet.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 1072 </td> 1073 <td style="padding:2.5%;width:70%;vertical-align:middle"> 1074 <h3>HDNet: High-resolution Dual-domain Learning for Spectral Compressive Imaging</h3> 1075 <br> 1076 <br> 1077 Xiaowan Hu *, <strong>Yuanhao Cai *</strong>, Jing Lin, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://www.bell-labs.com/about/researcher-profiles/xyuan/">Xin Yuan</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="http://people.ee.ethz.ch/~timofter/">Radu Timofte</a>, <a href="https://ee.ethz.ch/the-department/faculty/professors/person-detail.OTAyMzM=.TGlzdC80MTEsMTA1ODA0MjU5.html">Luc Van Gool</a> 1078 1079 <br> 1080 <em>Conference on Computer Vision and Pattern Recognition (CVPR)</em>, 2022 1081 <br> 1082 1083 <a href="https://arxiv.org/pdf/2203.02149.pdf">paper</a> 1084 1085 1086 / <a href="https://github.com/caiyuanhao1998/MST">project</a> 1087 1088 1089 1090 1091 1092 1093 1094 / <a href="/bibtex/hdnet.txt">bibtex</a> 1095 1096 <p></p> 1097 <p>Dual-domain learning for hyperspectral image reconstruction</p> 1098 1099 </td> 1100 </tr> 1101 1102 1103 1104 1105 1106 <tr> 1107 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 1108 <img src="/images/RFormer.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 1109 </td> 1110 <td style="padding:2.5%;width:70%;vertical-align:middle"> 1111 <h3>RFormer: Transformer-based Generative Adversarial Network for Real Fundus Image Restoration on A New Clinical Benchmark</h3> 1112 <br> 1113 <br> 1114 Zhuo Deng *, <strong>Yuanhao Cai *</strong>, Lu Chen, Zheng Gong, Qiqi Bao, Xue Yao, Dong Fang, Shaochong Zhang, <a href="https://sklco.pkusz.edu.cn/info/1030/1046.htm">Lan Ma</a> 1115 1116 <br> 1117 <em>
1117IEEE Journal of Biomedical and Health Informatics (J-BHI)</em>, 2022 1118 <br> 1119 1120 <a href="https://arxiv.org/pdf/2201.00466.pdf">paper</a> 1121 1122 1123 / <a href="https://github.com/dengzhuo-AI/Real-Fundus">project</a> 1124 1125 1126 1127 1128 1129 1130 1131 / <a href="/bibtex/rformer.txt">bibtex</a> 1132 1133 <p></p> 1134 <p>The first clinical benchmark and Transformer-based method for fundus image restoration</p> 1135 1136 </td> 1137 </tr> 1138 1139 1140 1141 1142 1143 <tr> 1144 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 1145 <img src="/images/PNGAN.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 1146 </td> 1147 <td style="padding:2.5%;width:70%;vertical-align:middle"> 1148 <h3>Learning to Generate Realistic Noisy Images via Pixel-level Noise-aware Adversarial Training</h3> 1149 <br> 1150 <br> 1151 <strong>Yuanhao Cai </strong>, Xiaowan Hu, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="https://vcg.seas.harvard.edu/people/hanspeter-pfister">Hanspeter Pfister</a>, <a href="https://donglaiw.github.io/">Donglai Wei</a> 1152 1153 <br> 1154 <em>Advances in Neural Information Processing Systems (NeurIPS)</em>, 2021 1155 <br> 1156 1157 <a href="https://proceedings.neurips.cc/paper/2021/file/1a5b1e4daae265b790965a275b53ae50-Paper.pdf">paper</a> 1158 1159 1160 / <a href="https://github.com/caiyuanhao1998/PNGAN">project</a> 1161 1162 1163 1164 1165 / <a href="https://noise.visinf.tu-darmstadt.de/benchmark/">leaderboard</a> 1166 1167 1168 1169 1170 / <a href="/bibtex/pngan.txt">bibtex</a> 1171 1172 <p></p> 1173 <p>A GAN for real noisy image generation</p> 1174 1175 </td> 1176 </tr> 1177 1178 1179 1180 1181 1182 <tr> 1183 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 1184 <img src="/images/MSFN.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 1185 </td> 1186 <td style="padding:2.5%;width:70%;vertical-align:middle"> 1187 <h3>Multi-Scale Selective Feedback Network with Dual Loss for Real Image Denoising</h3> 1188 <br> 1189 <br> 1190 Xiaowan Hu, <strong>Yuanhao Cai</strong>, Zhihong Liu, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a> 1191 1192 <br> 1193 <em>International Joint Conference on Artificial Intelligence (IJCAI), <strong>Oral</strong></em>, 2021 1194 <br> 1195 1196 <a href="https://www.ijcai.org/proceedings/2021/0101.pdf">paper</a> 1197 1198 1199 / <a href="https://www.ijcai.org/proceedings/2021/101">project</a> 1200 1201 1202 1203 1204 1205 1206 1207 / <a href="/bibtex/msfn.txt">bibtex</a> 1208 1209 <p></p> 1210 <p>A semi-supervised method for real image denoising</p> 1211 1212 </td> 1213 </tr> 1214 1215 1216 1217 1218 1219 <tr> 1220 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 1221 <img src="/images/P3AN.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 1222 </td> 1223 <td style="padding:2.5%;width:70%;vertical-align:middle"> 1224 <h3>Pseudo 3D Auto-Correlation Network for Real Image Denoising</h3> 1225 <br> 1226 <br> 1227 Xiaowan Hu, <a href="https://scholar.google.com/citations?user=k6irHZ0AAAAJ&hl=en">Ruijun Ma</a>, Zhihong Liu, <strong>Yuanhao Cai </strong>, <a href="https://scholar.google.com/citations?user=xkK4mRUAAAAJ&hl=en">Xiaole Zhao</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a> 1228 1229 <br> 1230 <em>Conference on Computer Vision and Pattern Recognition (CVPR)</em>, 2021 1231 <br> 1232 1233 <a href="https://openaccess.thecvf.com/content/CVPR2021/papers/Hu_Pseudo_3D_Auto-Correlation_Network_for_Real_Image_Denoising_CVPR_2021_paper.pdf">paper</a> 1234 1235 1236 / <a href="https://openaccess.thecvf.com/content/CVPR2021/html/Hu_Pseudo_3D_Auto-Correlation_Network_for_Real_Image_Denoising_CVPR_2021_paper.html">project</a> 1237 1238 1239 1240 1241 1242 1243 1244 / <a href="/bibtex/p3an.txt">bibtex</a> 1245 1246 <p></p> 1247 <p>
1247An efficient architecture for image denoising</p> 1248 1249 </td> 1250 </tr> 1251 1252 1253 1254 1255 1256 <tr> 1257 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 1258 <img src="/images/DANet.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 1259 </td> 1260 <td style="padding:2.5%;width:70%;vertical-align:middle"> 1261 <h3>Efficient Human Pose Estimation by Learning Deeply Aggregated Representations</h3> 1262 <br> 1263 <br> 1264 <a href="https://scholar.google.com.hk/citations?user=Sz1yTZsAAAAJ&hl=zh-CN">Zhengxion Luo</a>, <a href="https://scholar.google.com/citations?user=0QBBNGoAAAAJ&hl=zh-CN">Zhicheng Wang</a>, <strong>Yuanhao Cai </strong>, <a href="https://wangguanan.github.io/">Guan'an Wang</a>, <a href="http://www.cbsr.ia.ac.cn/users/liangwang/">Liang Wang</a>, <a href="https://yanrockhuang.github.io/">Yan Huang</a>, <a href="https://scholar.google.com/citations?user=k2ziPUsAAAAJ&hl=zh-CN">ErJin Zhou</a>, <a href="http://www.cbsr.ia.ac.cn/users/tnt/tnt.htm">Tieniu Tan</a>, <a href="http://www.jiansun.org/">Jian Sun</a> 1265 1266 <br> 1267 <em>International Conference on Multimedia and Expo (ICME), <strong>Oral</strong></em>, 2021 1268 <br> 1269 1270 <a href="https://arxiv.org/pdf/2012.07033.pdf">paper</a> 1271 1272 1273 / <a href="https://ieeexplore.ieee.org/abstract/document/9428206">project</a> 1274 1275 1276 1277 1278 1279 1280 1281 / <a href="/bibtex/danet.txt">bibtex</a> 1282 1283 <p></p> 1284 <p>An DenseNet-like backbone for efficient human pose estimation</p> 1285 1286 </td> 1287 </tr> 1288 1289 1290 1291 1292 1293 <tr> 1294 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 1295 <img src="/images/POAN.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 1296 </td> 1297 <td style="padding:2.5%;width:70%;vertical-align:middle"> 1298 <h3>Pyramid Orthogonal Attention Network based on Dual Self-Similarity for Accurate Mr Image Super-Resolution</h3> 1299 <br> 1300 <br> 1301 Xiaowan Hu, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <strong>Yuanhao Cai </strong>, <a href="https://scholar.google.com/citations?user=xkK4mRUAAAAJ&hl=en">Xiaole Zhao</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a> 1302 1303 <br> 1304 <em>International Conference on Multimedia and Expo (ICME)</em>, 2021 1305 <br> 1306 1307 <a href="https://cloud.tsinghua.edu.cn/f/6662fa8ad15a4c92b082/">paper</a> 1308 1309 1310 / <a href="https://ieeexplore.ieee.org/abstract/document/9428112">project</a> 1311 1312 1313 1314 1315 1316 1317 1318 / <a href="/bibtex/poan.txt">bibtex</a> 1319 1320 <p></p> 1321 <p>A novel self-attention mechanism for MR image super-resolution</p> 1322 1323 </td> 1324 </tr> 1325 1326 1327 1328 1329 1330 <tr> 1331 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 1332 <img src="/images/RSN.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 1333 </td> 1334 <td style="padding:2.5%;width:70%;vertical-align:middle"> 1335 <h3>Learning Delicate Local Representations for Multi-Person Pose Estimation</h3> 1336 <br> 1337 <br> 1338 <strong>Yuanhao Cai *</strong>, <a href="https://scholar.google.com/citations?user=0QBBNGoAAAAJ&hl=zh-CN">Zhicheng Wang</a> *, <a href="https://scholar.google.com.hk/citations?user=Sz1yTZsAAAAJ&hl=zh-CN">Zhengxion Luo</a>, Binyi Yin, Ang'ang Du, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://scholar.google.com/citations?user=yuB-cfoAAAAJ&hl=zh-CN">Xiangyu Zhang</a>, <a href="https://scholar.google.com/citations?user=Jv4LCj8AAAAJ&hl=en">Xinyu Zhou</a>, <a href="https://scholar.google.com/citations?user=k2ziPUsAAAAJ&hl=zh-CN">ErJin Zhou</a>, <a href="http://www.jiansun.org/">Jian Sun</a> 1339 1340 <br> 1341 <em>European Conference on Computer Vision (ECCV), <strong>Spotlight</strong></em>, 2020 1342 <br> 1343 1344 <a href="https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123480460.pdf">paper</a> 1345 1346 1347 / <a href="https://github.com/caiyuanhao1998/RSN">
1347project</a> 1348 1349 1350 1351 / <a href="https://zhuanlan.zhihu.com/p/112297707">zhihu</a> 1352 1353 1354 / <a href="https://cocodataset.org/#keypoints-leaderboard">leaderboard</a> 1355 1356 1357 1358 1359 / <a href="/bibtex/rsn.txt">bibtex</a> 1360 1361 <p></p> 1362 <p>A multi-stage structure and an attention mechanism for accurate human pose estimation</p> 1363 1364 </td> 1365 </tr> 1366 1367 1368 1369 1370 1371 <tr> 1372 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 1373 <img src="/images/UDP.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 1374 </td> 1375 <td style="padding:2.5%;width:70%;vertical-align:middle"> 1376 <h3>UDP++</h3> 1377 <br> 1378 <br> 1379 Junjie Huang *, Zengguang Shan *, <strong>Yuanhao Cai *</strong>, Feng Guo, Yun Ye, Xinze Chen, <a href="http://www.zhengzhu.net/">Zheng Zhu</a>, Guan Huang, <a href="http://ivg.au.tsinghua.edu.cn/Jiwen_Lu/">Jiwen Lu</a>, Dalong Du 1380 1381 <br> 1382 <em>European Conference on Computer Vision Workshop (ECCVW), <strong>Oral</strong></em>, 2020 1383 <br> 1384 1385 <a href="https://s3-us-west-1.amazonaws.com/presentations.cocodataset.org/ECCV20/keypoints/UDP.pdf">paper</a> 1386 1387 1388 / <a href="https://github.com/caiyuanhao1998/UDP-plus-plus">project</a> 1389 1390 1391 1392 / <a href="https://zhuanlan.zhihu.com/p/210199401">zhihu</a> 1393 1394 1395 / <a href="https://cocodataset.org/#keypoints-leaderboard">leaderboard</a> 1396 1397 1398 1399 1400 / <a href="/bibtex/udp_pp.txt">bibtex</a> 1401 1402 <p></p> 1403 <p>Winner of COCO Keypoint Detection Challenge, 2020</p> 1404 1405 </td> 1406 </tr> 1407 1408 1409 1410 1411 1412 <tr> 1413 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 1414 <img src="/images/SPIE.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 1415 </td> 1416 <td style="padding:2.5%;width:70%;vertical-align:middle"> 1417 <h3>EG^2N: Enhanced Gradient Guiding Network for Single MR Image Super-Resolution</h3> 1418 <br> 1419 <br> 1420 Xiaowan Hu, <strong>Yuanhao Cai</strong>, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, Yanbin Peng, <a href="https://yulunzhang.com/">Yulun Zhang</a> 1421 1422 <br> 1423 <em>Optoelectronic Imaging and Multimedia Technology VII 11550, 115500I</em>, 2020 1424 <br> 1425 1426 <a href="https://cloud.tsinghua.edu.cn/f/78a711714198479d991f/">paper</a> 1427 1428 1429 / <a href="https://www.spiedigitallibrary.org/conference-proceedings-of-spie/11550/115500I/EG2N--enhanced-gradient-guiding-network-for-single-MR-image/10.1117/12.2575261.short?SSO=1">project</a> 1430 1431 1432 1433 1434 1435 1436 1437 / <a href="/bibtex/eg2n.txt">bibtex</a> 1438 1439 <p></p> 1440 <p>A novel gradient guiding mechanism for MR image super-resolution</p> 1441 1442 </td> 1443 </tr> 1444 1445 1446 1447 1448 1449 <tr> 1450 <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px"> 1451 <img src="/images/Res-Step-Net.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" /> 1452 </td> 1453 <td style="padding:2.5%;width:70%;vertical-align:middle"> 1454 <h3>Res-Steps-Net for Multi-Person Pose Estimation</h3> 1455 <br> 1456 <br> 1457 <strong>Yuanhao Cai *</strong>, <a href="https://scholar.google.com/citations?user=0QBBNGoAAAAJ&hl=zh-CN">Zhicheng Wang</a> *, Binyi Yin, Ruihao Yin, Angang Du, <a href="https://scholar.google.com.hk/citations?user=Sz1yTZsAAAAJ&hl=zh-CN">Zhengxion Luo</a>, <a href="https://www.zemingli.com/">Zeming Li</a>, <a href="https://scholar.google.com/citations?user=Jv4LCj8AAAAJ&hl=en">Xinyu Zhou</a>, <a href="https://www.skicyyu.org/">Gang Yu</a>, <a href="https://scholar.google.com/citations?user=k2ziPUsAAAAJ&hl=zh-CN">ErJin Zhou</a>, <a href="https://scholar.google.com/citations?user=yuB-cfoAAAAJ&hl=zh-CN">Xiangyu Zhang</a>, <a href="https://yichenwei.github.io/">Yichen Wei</a>, <a href="http://www.jiansun.org/">Jian Sun</a> 1458 1459 <br> 1460 <em>International Conference on Computer Vision Workshop (ICCVW), <strong>Best Paper Award</strong></em>, 2019 1461 <br> 1462 1463 <a href="https://cocodataset.org/files/keypoints_2019_reports/Megvii.pdf">paper</a> 1464 1465 1466 / <a href="https://github.com/caiyuanhao1998/RSN">
1466project</a> 1467 1468 1469 1470 / <a href="https://zhuanlan.zhihu.com/p/112297707">zhihu</a> 1471 1472 1473 / <a href="https://cocodataset.org/#keypoints-leaderboard">leaderboard</a> 1474 1475 1476 1477 1478 / <a href="/bibtex/rsn_iccv.txt">bibtex</a> 1479 1480 <p></p> 1481 <p>Winner of COCO Keypoint Detection Challenge, 2019</p> 1482 1483 </td> 1484 </tr> 1485 1486 1487 1488 </table> 1489 1490 <br> 1491 <br> 1492 1493 <table style="width:100%;border:0px;border-spacing:0px;border-collapse:separate;margin-right:auto;margin-left:auto;"> 1494 <h2>Academic Service</h2> 1495 <tr style="padding:0px"> 1496 <td style="padding:2.5%;vertical-align:left"> 1497 <p> 1498 Area Chair: CVPR 2027 1499 </p> 1500 <p> 1501 Conference Reviewer: CVPR, ECCV, ICCV, NeurIPS, ICML, ICLR, AAAI, IJCAI, ACM MM, etc. 1502 </p> 1503 <p> 1504 Journal Reviewer: TPAMI, IJCV, TIP, TNNLS, Pattern Recognition, etc. 1505 </p> 1506 </td> 1507 </tr> 1508 </table> 1509 1510 <br> 1511 </td> 1512 </tr> 1513 </table> 1514</body> 1515 1516</html> 1517
Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.