PageSourceSearch

https://caiyuanhao1998.github.io/

html caiyuanhao1998.github.io collected 2026-10-03 08:44:53 UTC 69,326 bytes, 1,517 lines download raw bytes

1<!DOCTYPE HTML>
2<html lang="en">
3
4<head>
5  <title>Yuanhao Cai</title>
6
7  <meta content="text/html; charset=utf-8" http-equiv="Content-Type">
8
9  <meta name="author" content="Yuanhao Cai" />
10  <meta name="viewport" content="width=device-width, initial-scale=1">
11
12  <link rel="stylesheet" type="text/css" href="/style.css" />
13  <link rel="canonical" href="https://caiyuanhao.github.io/">
14  <link href="https://fonts.googleapis.com/css?family=Lato:400,700,400italic,700italic" rel="stylesheet" type="text/css">
15
16</head>
17
18
19
20<body>
21  <table style="width:100%;max-width:1000px;border:0px;border-spacing:0px;border-collapse:separate;margin-right:auto;margin-left:auto;">
22    <tr style="padding:0px">
23      <td style="padding:0px">
24        <table style="width:100%;border:0px;border-spacing:0px;border-collapse:separate;margin-right:auto;margin-left:auto;">
25          <tr style="padding:0px">
26            <td style="padding:2.5%;width:65%;vertical-align:middle">
27              <h1>
28                Yuanhao Cai
29              </h1>
30              <p>I am currently a 3rd year PhD student in the department of <a href="https://www.cs.jhu.edu/">Computer Science</a>, <a href="https://www.jhu.edu/">Johns Hopkins University</a>.
31                I am a member of <a href="https://ccvl.jhu.edu/">CCVL</a>, advised by <a href="https://en.wikipedia.org/wiki/Bloomberg_Distinguished_Professorships">Bloomberg Distinguished Professor</a>
32                <a href="https://www.cs.jhu.edu/~ayuille/"> Dr. Alan Yuille</a>.
33                Previously, I received my MSE and BSE degrees from <a href="https://www.tsinghua.edu.cn/en/">Tsinghua University</a> in 2023 and 2020. My master advisor is prof. <a href="https://scholar.google.com/citations?user=eldgnIYAAAAJ&hl=zh-CN">Haoqian Wang</a>.
34                During my study in Tsinghua University, I spent a good time with prof. <a href="https://yulunzhang.com/">Yulun Zhang</a>, prof. <a href="https://en.westlake.edu.cn/faculty/xin-yuan.html">Xin Yuan</a>, prof. <a href="https://www.informatik.uni-wuerzburg.de/computervision/">Radu Timofte</a>, 
35                and prof. <a href="https://ee.ethz.ch/the-department/faculty/professors/person-detail.OTAyMzM=.TGlzdC80MTEsMTA1ODA0MjU5.html">Luc Van Gool</a>. I interned in Adobe Research (2024 - 2025) and Meta Superintelligence Labs (2025 - 2026).
36              </p>
37              <p style="text-align:center">
38                <a href="https://github.com/caiyuanhao1998">GitHub</a> &nbsp;/&nbsp;
39                <a href="https://scholar.google.com/citations?user=3YozQwcAAAAJ&hl=en">Google Scholar</a> &nbsp;/&nbsp;
40                <a href="https://www.linkedin.com/in/yuanhao-cai-b8463b297/">Linkdin</a> &nbsp;/&nbsp;
41                <a href="https://www.zhihu.com/people/cyh-28-29">ZhiHu</a> 
42              </p>
43              <p>I am open to any opportunities of collaboration. Please feel free to contact me if you are interested.
44              </p>
45            </td>
46            <td style="padding:2.5%;width:35%;max-width:35%">
47              <img style="width:100%;max-width:100%" alt="self photo" src="images/avatar_2026_04_01.jpg"/>
48            </td>
49          </tr>
50        </table>
51
52        <table style="width:100%;border:0px;border-spacing:0px;border-collapse:separate;margin-right:auto;margin-left:auto;">
53          <h2>Honor and Award</h2>
54          <tr style="padding:0px">
55            <td style="padding:2.5%;vertical-align:left">
56              <p>
57                <a href="images/top2_scientist.jpg">Stanford/Elsevier World’s Top 2% Scientists</a>, 2025
58              </p>
59              <p>
60                <a href="images/AI4CC_best_paper.png">Best Paper Award</a> of <a href="https://ai4cc.net/">AI4CC Workshop</a> at CVPR, 2025
61              </p>
62              <p>
63                Runner-up Award of <a href="https://codalab.lisn.upsaclay.fr/competitions/17640">NTIRE Low Light Enhancement Challenge</a> at CVPR, 2024
64              </p>
65              <p>
66                <a href="images/cvpr24_outstanding_reviewers.jpg">Outstanding Reviewer Award</a> at CVPR, 2024
67              </p>
68              <p>
69                <a href="https://iclr.cc/Conferences/2024/Reviewers#outstanding-reviewers">Outstanding Reviewer Award</a> at ICLR, 2024
70              </p>
71              <p>
72                Excellent Graduate of <a href="https://www.tsinghua.edu.cn/">Tsinghua University</a>, 2023
73              </p>
74              <p>
75                Excellent Master Thesis Award, <a href="https://www.tsinghua.edu.cn/">Tsinghua University</a>, 2023
76              </p>
77              <p>
78                <a href="https://cvpr2023.thecvf.com/Conferences/2023/OutstandingReviewers">Outstanding Reviewer Award</a> at CVPR, 2023
79              </p>
80              <p>
81                <a href="images/Top_Talented.pdf">Top-10 Talented Graduate Student Finalist</a>, <a href="https://www.tsinghua.edu.cn/">Tsinghua University</a>, 2022
82              </p>
83              <p>
84                National Scholarship, <a href="https://www.tsinghua.edu.cn/">Tsinghua University</a>, 2022
85              </p>
86              <p>
87                Winner of <a href="https://codalab.lisn.upsaclay.fr/competitions/721">NTIRE Spectral Reconstruction Challenge</a> at CVPR, 2022
88              </p>
89              <p>
90                National Scholarship, <a href="https://www.tsinghua.edu.cn/">Tsinghua University</a>, 2021
91              </p>
92              <p>
93                Winner of <a href="https://cocodataset.org/workshop/coco-lvis-eccv-2020.html">COCO Keypoint Detection Challenge</a> at ECCV, 2020
94              </p>
95              <p>
96                Winner of <a href="https://cocodataset.org/workshop/coco-mapillary-iccv-2019.html">COCO Keypoint Detection Challenge</a> at ICCV, 2019
97              </p>
98              <p>
99                Best Paper Award of <a href="https://cocodataset.org/workshop/coco-mapillary-iccv-2019.html">Joint COCO and Mapillary Recognition Challenge Workshop</a> at ICCV, 2019
100              </p>
101              <p>
102                2nd Place of AdultSize Drop-In Challenge at <a href="https://2019.robocup.org/">RoboCup World Final</a>, 2019
103              </p>
104              <p>
105                2nd Place of AdultSize Technical Challenge at <a href="https://2019.robocup.org/">RoboCup World Final</a>, 2019
106              </p>
107              <p>
108                3rd Place of AdultSize Soccer Competition at <a href="https://2019.robocup.org/">RoboCup World Final</a>, 2019
109              </p>
110              <p>
111                Learning Progress Scholarship, <a href="https://www.tsinghua.edu.cn/">Tsinghua University</a>, 2019
112              </p>
113              <p>
114                Science and Technology Innovation Excellence Scholarship, <a href="https://www.tsinghua.edu.cn/">Tsinghua University</a>, 2019
115              </p>
116            </td>
117          </tr>
118        </table>
119
120        <table style="width:100%;border:0px;border-spacing:0px;border-collapse:separate;margin-right:auto;margin-left:auto;">
121          <h2>Publication</h2>   ( * = Equal Contribution, † = Corresponding Author)
122          <br>
123          <br>
124          <br>
125          
126          
127          
128          <tr>
129            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
130              <img src="/images/phygdpo.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
131            </td>
132            <td style="padding:2.5%;width:70%;vertical-align:middle">
133              <h3>PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation</h3>
134              <br>
135              <br>
136              <strong>Yuanhao Cai</strong>, <a href="https://kunpengli1994.github.io/">Kunpeng Li</a>, <a href="https://kmnp.github.io/">Menglin Jia</a>, <a href="https://sites.google.com/view/jialiangwang/home">Jialiang Wang</a>, <a href="https://scholar.google.com/citations?user=wyi0bX0AAAAJ&hl=en">Junzhe Sun</a>, <a href="https://wfchen-umich.github.io/wfchen.github.io/">Weifeng Chen</a>, <a href="https://xujuefei.com/">Felix Juefei Xu</a>, <br>
136 <a href="https://scholar.google.com/citations?user=5aaOtscAAAAJ&hl=en">Chu Wang</a>, <a href="https://www.alithabet.com/">Ali Thabet</a>,  <a href="https://sites.google.com/view/xiaoliangdai">Xiaoliang Dai</a>, <a href="https://juxuan27.github.io/">Xuan Ju</a>, <a href="https://www.cs.jhu.edu/~ayuille/">Alan Yuille</a>, <a href="https://sekunde.github.io/">Ji Hou</a>
137
138              <br>
139              <em>European Conference on Computer Vision (ECCV)</em>, 2026
140              <br>
141              
142              <a href="https://arxiv.org/abs/2512.24551">paper</a>
143              
144              
145              / <a href="https://caiyuanhao1998.github.io/project/PhyGDPO">project</a>
146              
147              
148              
149              
150              
151              
152              
153              <p></p>
154              <p>A data construction pipeline and a new diret preference optimization framework for physically consistant video generation</p>
155
156            </td>
157          </tr>
158          
159          
160          
161          
162          
163          <tr>
164            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
165              <img src="/images/omnivcus.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
166            </td>
167            <td style="padding:2.5%;width:70%;vertical-align:middle">
168              <h3>OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions</h3>
169              <br>
170              <br>
171              <strong>Yuanhao Cai</strong>, <a href="https://sites.google.com/site/hezhangsprinter">He Zhang</a>,  <a href="https://xavierchen34.github.io/">Xi Chen</a>, <a href="https://doubiiu.github.io/">Jinbo Xing</a>, <a href="https://yiweihu.netlify.app/">Yiwei Hu</a>, <a href="https://yzhouas.github.io/">Yuqian Zhou</a>, <a href="https://kai-46.github.io/website/">Kai Zhang</a>,  <a href="https://zzutk.github.io/">Zhifei Zhang</a>, <a href="https://sites.google.com/view/sooyekim">Soo Ye Kim</a>, <a href="https://stevewongv.github.io/">Tianyu Wang</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="https://english.seiee.sjtu.edu.cn/english/detail/842_802.htm">Xiaokang Yang</a>, <a href="https://sites.google.com/site/zhelin625/">Zhe Lin</a>,  <a href="https://www.cs.jhu.edu/~ayuille/">Alan Yuille</a>
172
173              <br>
174              <em>Advances in Neural Information Processing Systems (NeurIPS)</em>, 2025
175              <br>
176              
177              <a href="http://arxiv.org/abs/2506.23361">paper</a>
178              
179              
180              / <a href="https://caiyuanhao1998.github.io/project/OmniVCus">project</a>
181              
182              
183              
184              
185              
186              
187              
188              <p></p>
189              <p>A data construction pipeline and a unified diffusion Transformer for subject-driven video customization under different control conditions; 4D generation</p>
190
191            </td>
192          </tr>
193          
194          
195          
196          
197          
198          <tr>
199            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
200              <img src="/images/editverse.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
201            </td>
202            <td style="padding:2.5%;width:70%;vertical-align:middle">
203              <h3>EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning</h3>
204              <br>
205              <br>
206              <a href="https://juxuan27.github.io/">Xuan Ju</a>, <a href="https://scholar.google.com/citations?user=yRwZIN8AAAAJ&hl=zh-CN">Tianyu Wang</a>, <a href="https://yzhouas.github.io/">Yuqian Zhou</a>, <a href="https://sites.google.com/site/hezhangsprinter/">He Zhang</a>, <a href="https://qliu24.github.io/">Qing Liu</a>, <a href="https://www.nxzhao.com/">Nanxuan Zhao</a>, <a href="https://zzutk.github.io/">Zhifei Zhang</a>, <a href="https://yijunmaverick.github.io/">Yijun Li</a>, <strong>Yuanhao Cai</strong>, <a href="https://www.shaotengliu.com/">Shaoteng Liu</a>, <a href="https://scholar.google.com/citations?user=UI10l34AAAAJ&hl=en">Daniil Pakhomov</a>, <a href="https://scholar.google.com/citations?user=UI10l34AAAAJ&hl=en">Daniil Pakhomov</a>, <a href="https://sites.google.com/site/zhelin625/">Zhe Lin</a>, <a href="https://sites.google.com/view/sooyekim">Soo Ye Kim</a>, <a href="https://www.cse.cuhk.edu.hk/people/faculty/qiang-xu/">Qiang Xu</a>
207
208              <br>
209              <em>International Conference on Representation Learning (ICLR)</em>, 2025
210              <br>
211              
212              <a href="https://arxiv.org/abs/2509.20360">paper</a>
213              
214              
215              / <a href="http://editverse.s3-website-us-east-1.amazonaws.com/">project</a>
216              
217              
218              
219              
220              
221              
222              
223              <p></p>
224              <p>Unifying a diverse range of generation and editing tasks for both images and videos within a single and powerful model</p>
225
226            </td>
227          </tr>
228          
229          
230          
231          
232          
233          <tr>
234            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
235              <img src="/images/langsplatv2.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
236            </td>
237            <td style="padding:2.5%;width:70%;vertical-align:middle">
238              <h3>LangSplatV2: High-dimensional 3D Language Gaussian Splatting with 450+ FPS</h3>
239              <br>
240              <br>
241              <a href="https://li-wanhua.github.io/">Wanhua Li*</a>, <a href="https://github.com/ZhaoYujie2002">Yujie Zhao*</a>, <a href="https://minghanqin.github.io/">Minghan Qin*</a>, <a href="https://github.com/jimmyYliu">Yang Liu</a>, <strong>Yuanhao Cai</strong>, <a href="https://people.csail.mit.edu/ganchuang/">Chuang Gan</a>, <a href="https://vcg.seas.harvard.edu/people">Hanspeter Pfister</a>
242
243              <br>
244              <em>Advances in Neural Information Processing Systems (NeurIPS)</em>, 2025
245              <br>
246              
247              <a href="https://arxiv.org/abs/2507.07136">paper</a>
248              
249              
250              / <a href="https://langsplat-v2.github.io/">project</a>
251              
252              
253              
254              
255              
256              
257              
258              <p></p>
259              <p>A super fast framework for real-time 3D open-vocabulary querying and high-dimensional feature splatting</p>
260
261            </td>
262          </tr>
263          
264          
265          
266          
267          
268          <tr>
269            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
270              <img src="/images/care.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
271            </td>
272            <td style="padding:2.5%;width:70%;vertical-align:middle">
273              <h3>Are Pixel-Wise Metrics Reliable for Sparse-View Computed Tomography Reconstruction?</h3>
274              <br>
275              <br>
276              <a href="https://lin-tianyu.github.io/">Tianyu Lin</a>, Xinran Li, Chuntung Zhuang,  <a href="https://scholar.google.com/citations?user=4Q5gs2MAAAAJ&hl=en">Qi Chen</a>, <strong>Yuanhao Cai</strong>, <a href="https://scholar.google.com/citations?user=OvpsAYgAAAAJ&hl=en&oi=ao">Kai Ding</a>, <a href="https://www.cs.jhu.edu/~ayuille/">Alan Yuille</a>, <a href="https://www.zongweiz.com/">Zongwei Zhou</a>
277
278              <br>
279              <em>Advances in Neural Information Processing Systems (NeurIPS)</em>, 2025
280              <br>
281              
282              <a href="https://arxiv.org/abs/2506.02093">paper</a>
283              
284              
285              / <a href="https://github.com/MrGiovanni/CARE">project</a>
286              
287              
288              
289              
290              
291              
292              
293              <p></p>
294              <p>New metrics and diffusion based anatomy-aware enhancement for sparse-view CT reconstruction</p>
295
296            </td>
297          </tr>
298          
299          
300          
301          
302          
303          <tr>
304            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
305              <img src="/images/xlrm.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
306            </td>
307            <td style="padding:2.5%;width:70%;vertical-align:middle">
308              <h3>X-LRM: X-ray Large Reconstruction Model for Extremely Sparse-View Computed Tomography Recovery in One Second</h3>
309              <br>
310              <br>
311              <a href="https://scholar.google.com/citations?hl=en&user=vl0mzhEAAAAJ">Guofeng Zhang *</a>, <a href="https://ruyi-zha.github.io/">Ruyi Zha *</a>, <a href="https://heye0507.github.io/">Hao He</a>, <a href="https://yixunliang.github.io/">Yixun Liang</a>, <a href="https://www.cs.jhu.edu/~ayuille/">Alan Yuille</a>, <a href="https://users.cecs.anu.edu.au/~hongdong/">Hongdong Li</a>, <strong>Yuanhao Cai</strong>
312
313              <br>
314              <em>International Conference on 3D Vision (3DV)</em>, 2025
315              <br>
316              
317              <a href="https://arxiv.org/abs/2503.06382">paper</a>
318              
319              
320              / <a href="https://github.com/caiyuanhao1998/X-LRM">project</a>
321              
322              
323              
324              
325              
326              
327              
328              <p></p>
329              <p>A feedforward method and a large-scale dataset for instant CT reconstruction</p>
330
331            </td>
332          </tr>
333          
334          
335          
336          
337          
338          <tr>
339            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
340              <img src="/images/diffusiongs.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
341            </td>
342            <td style="padding:2.5%;width:70%;vertical-align:middle">
343              <h3>Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and Reconstruction</h3>
344              <br>
345              <br>
346              <strong>Yuanhao Cai</strong>, <a href="https://sites.google.com/site/hezhangsprinter">He Zhang</a>,  <a href="https://kai-46.github.io/website/">Kai Zhang</a>,  <a href="https://yixunliang.github.io/">Yixun Liang</a>,  <a href="https://www.mengweiren.com/">Mengwei Ren</a>,  <a href="https://luanfujun.com/">Fujun Luan</a>,  <a href="https://qliu24.github.io/">Qing Liu</a>,  <a href="https://sites.google.com/view/sooyekim">Soo Ye Kim</a>,  <a href="https://jimmie33.github.io/">Jianming Zhang</a>,  <a href="https://zzutk.github.io/">Zhifei Zhang</a>,  <a href="https://yzhouas.github.io/">Yuqian Zhou</a>,  <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="https://english.seiee.sjtu.edu.cn/english/detail/842_802.htm">Xiaokang Yang</a>, <a href="https://sites.google.com/site/zhelin625/">Zhe Lin</a>,  <a href="https://www.cs.jhu.edu/~ayuille/">Alan Yuille</a>
347
348              <br>
349              <em>International Conference on Computer Vision (ICCV)</em>, 2025
350              <br>
351              
352              <a href="https://arxiv.org/abs/2411.14384">paper</a>
353              
354              
355              / <a href="https://caiyuanhao1998.github.io/project/DiffusionGS/">project</a>
356              
357              
358              
359              
360              
361              
362              / <a href="https://x.com/janusch_patas/status/1859867424859856997?ref_src=twsrc%5Egoogle%7Ctwcamp%5Eserp%7Ctwgr%5Etweet">media (MrNeRF)</a>
363              
364              
365              <p></p>
366              <p>A 3DGS diffusion generates objects and reconstructs scenes from a single view in 6s</p>
367
368            </td>
369          </tr>
370          
371          
372          
373          
374          
375          <tr>
376            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
377              <img src="/images/x2gs.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
378            </td>
379            <td style="padding:2.5%;width:70%;vertical-align:middle">
380              <h3>4D Radiative Gaussian Splatting for Continuous-time Tomographic Reconstruction</h3>
381              <br>
382              <br>
383              <a href="https://scholar.google.com/citations?hl=en&user=fCzlLE4AAAAJ">Weihao Yu</a>, <strong>Yuanhao Cai</strong>, <a href="https://ruyi-zha.github.io/">Ruyi Zha</a>, <a href="https://zhiwenfan.github.io/">Zhiwen Fan</a>,  <a href="https://chenxinli001.github.io/">Chenxin Li</a>,  <a href="https://www.ee.cuhk.edu.hk/~yxyuan/">Yixuan Yuan</a>
384
385              <br>
386              <em>International Conference on Computer Vision (ICCV)</em>, 2025
387              <br>
388              
389              <a href="https://arxiv.org/abs/2503.21779">paper</a>
390              
391              
392              / <a href="https://x2-gaussian.github.io/">project</a>
393              
394              
395              
396              
397              
398              
399              / <a href="https://x.com/janusch_patas/status/1905528815302087135">media (MrNeRF)</a>
400              
401              
402              <p></p>
403              <p>A 4D Gaussian Splatting method for dynamic CT reconstruction</p>
404
405            </td>
406          </tr>
407          
408          
409          
410          
411          
412          <tr>
413            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
414              <img src="/images/motionx_plus.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
415            </td>
416            <td style="padding:2.5%;width:70%;vertical-align:middle">
417              <h3>Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset</h3>
418              <br>
419              <br>
420              Yuhong Zhang, <a href="https://jinglin7.github.io/">Jing Lin</a>, <a href="https://ailingzeng.site/">Ailing Zeng</a>, <a href="https://guanlinwu123.github.io/">Guanlin Wu</a>, <a href="https://shunlinlu.github.io/">Shunlin Lu</a>,  Yurong Fu,  <strong>Yuanhao Cai</strong>, <a href="http://zhangruimao.site/">Ruimao Zhang</a>,  <a href="https://scholar.google.com/citations?user=eldgnIYAAAAJ&hl=zh-CN">Haoqian Wang</a>,  <a href="https://www.leizhang.org/">Lei Zhang</a>
421
422              <br>
423              <em>arxiv</em>, 2024
424              <br>
425              
426              <a href="https://arxiv.org/abs/2501.05098">paper</a>
427              
428              
429              / <a href="https://motion-x-dataset.github.io/">project</a>
430              
431              
432              
433              
434              
435              
436              
437              <p></p>
438              <p>A Large-Scale Dataset for Multimodal 3D Whole-body Human Motion Generation</p>
439
440            </td>
441          </tr>
442          
443          
444          
445          
446          
447          <tr>
448            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
449              <img src="/images/videolifter.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
450            </td>
451            <td style="padding:2.5%;width:70%;vertical-align:middle">
452              <h3>VideoLifter: Lifting Videos to 3D with Fast Hierarchical Stereo Alignment</h3>
453              <br>
454              <br>
455              <a href="https://wenyancong.com/">Wenyan Cong</a>, <a href="https://www.kevin-ai.com/">Kevin Wang</a>, <a href="https://www.cis.upenn.edu/~leijh/">Jiahui Lei</a>, <a href="https://coltonstearns.github.io/">Colton Stearns</a>, <strong>Yuanhao Cai</strong>, <a href="https://wdilin.github.io/">Dilin Wang</a>,  <a href="https://www.linkedin.com/in/rakesh-r-3848538/">Rakesh Ranjan</a>,  <a href="https://scholar.google.com/citations?user=A-wA73gAAAAJ&hl=en">Matt Feiszli</a>,  <a href="https://profiles.stanford.edu/leonidas-guibas">Leonidas Guibas</a>,  <a href="https://express.adobe.com/page/CAdrFMJ9QeI2y/">Zhangyang Wang</a>,  <a href="https://scholar.google.com/citations?user=d1-hNQoAAAAJ&hl=en">Weiyao Wang</a>,  <a href="https://zhiwenfan.github.io/">Zhiwen Fan</a>
456
457              <br>
458              <em>arxiv</em>, 2024
459              <br>
460              
461              <a href="https://arxiv.org/abs/2501.01949">paper</a>
462              
463              
464              / <a href="https://videolifter.github.io/">project</a>
465              
466              
467              
468              
469              
470              
471              
472              <p></p>
473              <p>A 3DGS-based reconstruction method for videos without SfM initilization</p>
474
475            </td>
476          </tr>
477          
478          
479          
480          
481          
482          <tr>
483            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
484              <img src="/images/lucidfusion.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
485            </td>
486            <td style="padding:2.5%;width:70%;vertical-align:middle">
487              <h3>LucidFusion: Generating 3D Gaussians with Arbitrary Unposed Images</h3>
488              <br>
489              <br>
490              <a href="https://heye0507.github.io/">Hao He</a>, <a href="https://yixunliang.github.io/">Yixun Liang</a>, <a href="https://wileewang.github.io/">Luozhou Wang</a>, <strong>Yuanhao Cai</strong>, <a href="https://scholar.google.com/citations?user=lrgPuBUAAAAJ&hl=en&inst=1381320739207392350/">
490Xinli Xu</a>, Hao-Xiang Guo, Xiang Wen, <a href="https://www.yingcong.me/">Ying-Cong Chen</a>
491
492              <br>
493              <em>arxiv</em>, 2024
494              <br>
495              
496              <a href="https://arxiv.org/abs/2410.15636">paper</a>
497              
498              
499              / <a href="https://heye0507.github.io/LucidFusion_page/">project</a>
500              
501              
502              
503              
504              
505              
506              
507              <p></p>
508              <p>A pose-free multi-view 3D reconstruction methods based on Gaussian Splatting</p>
509
510            </td>
511          </tr>
512          
513          
514          
515          
516          
517          <tr>
518            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
519              <img src="/images/hdr_gs.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
520            </td>
521            <td style="padding:2.5%;width:70%;vertical-align:middle">
522              <h3>HDR-GS: Efficient High Dynamic Range Novel View Synthesis at 1000x Speed via Gaussian Splatting</h3>
523              <br>
524              <br>
525              <strong>Yuanhao Cai </strong>, <a href="https://scholar.google.com/citations?user=ucb6UssAAAAJ&hl=en">Zihao Xiao</a>, <a href="https://yixunliang.github.io/">Yixun Liang</a>, <a href="https://minghanqin.github.io/">Minghan Qin</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="https://english.seiee.sjtu.edu.cn/english/detail/842_802.htm">Xiaokang Yang</a>, <a href="https://www.cs.jhu.edu/~yyliu/">Yaoyao Liu</a>,  <a href="https://www.cs.jhu.edu/~ayuille/">Alan Yuille</a>
526
527              <br>
528              <em>Advances in Neural Information Processing Systems (NeurIPS)</em>, 2024
529              <br>
530              
531              <a href="https://arxiv.org/pdf/2405.15125">paper</a>
532              
533              
534              / <a href="https://github.com/caiyuanhao1998/HDR-GS">project</a>
535              
536              
537              / <a href="https://www.youtube.com/watch?v=wtU7Kcwe7ck">video</a> 
538              
539              
540              / <a href="https://zhuanlan.zhihu.com/p/10016024329">zhihu</a>
541              
542              
543              / <a href="https://paperswithcode.com/sota/novel-view-synthesis-on-hdr-gs">leaderboard</a>
544              
545              
546              / <a href="https://x.com/_akhaliq/status/1794921228462923925?s=46">media (AK)</a>
547              
548              
549              / <a href="https://x.com/janusch_patas/status/1794932286397489222?s=46">media (MrNeRF)</a>
550              
551              
552              / <a href="/bibtex/hdr_gs.txt">bibtex</a>
553              
554              <p></p>
555              <p>The first 3D Gaussian splatting-based method for high dynamic range imaging</p>
556
557            </td>
558          </tr>
559          
560          
561          
562          
563          
564          <tr>
565            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
566              <img src="/images/r2gs.gif" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
567            </td>
568            <td style="padding:2.5%;width:70%;vertical-align:middle">
569              <h3>R2-Gaussian: Rectifying Radiative Gaussian Splatting for Tomographic Reconstruction</h3>
570              <br>
571              <br>
572              <a href="https://ruyi-zha.github.io/">Ruyi Zha</a>, <a href="https://scholar.google.com/citations?user=BD7Hce0AAAAJ&hl=en">Tao Jun Lin</a>, <strong>Yuanhao Cai†</strong>, <a href="https://sites.google.com/view/yanhaozhang/home">Jiwen Cao</a>, <a href="https://users.cecs.anu.edu.au/~hongdong/">Hongdong Li</a>
573
574              <br>
575              <em>Advances in Neural Information Processing Systems (NeurIPS)</em>, 2024
576              <br>
577              
578              <a href="https://arxiv.org/abs/2405.20693v1">paper</a>
579              
580              
581              / <a href="https://ruyi-zha.github.io/r2_gaussian/r2_gaussian.html">project</a>
582              
583              
584              
585              
586              
587              
588              / <a href="https://x.com/janusch_patas/status/1797493143094571041">media (MrNeRF)</a>
589              
590              
591              <p></p>
592              <p>The first 3D Gaussian splatting-based method for CT reconstruction</p>
593
594            </td>
595          </tr>
596          
597          
598          
599          
600          
601          <tr>
602            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
603              <img src="/images/x-gaussian.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
604            </td>
605            <td style="padding:2.5%;width:70%;vertical-align:middle">
606              <h3>Radiative Gaussian Splatting for Efficient X-ray Novel View Synthesis</h3>
607              <br>
608              <br>
609              <strong>Yuanhao Cai </strong>, <a href="https://yixunliang.github.io/">Yixun Liang</a>, <a href="https://jiahaoplus.github.io/">jiahao Wang</a>, <a href="https://scholar.google.com/citations?hl=en&user=YR7re-cAAAAJ">Angtian Wang</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="https://english.seiee.sjtu.edu.cn/english/detail/842_802.htm">Xiaokang Yang</a>, <a href="https://www.zongweiz.com/">Zongwei Zhou</a>,  <a href="https://www.cs.jhu.edu/~ayuille/">Alan Yuille</a>
610
611              <br>
612              <em>European Conference on Computer Vision (ECCV)</em>, 2024
613              <br>
614              
615              <a href="https://arxiv.org/pdf/2403.04116.pdf">paper</a>
616              
617              
618              / <a href="https://github.com/caiyuanhao1998/X-Gaussian">project</a>
619              
620              
621              / <a href="https://www.youtube.com/watch?v=gDVf_Ngeghg">video</a> 
622              
623              
624              / <a href="https://zhuanlan.zhihu.com/p/717744222">zhihu</a>
625              
626              
627              
628              / <a href="https://x.com/_akhaliq/status/1765929288044290253?s=46">media (AK)</a>
629              
630              
631              / <a href="https://x.com/janusch_patas/status/1766446189749150126?s=46">media (MrNeRF)</a>
632              
633              
634              / <a href="/bibtex/x_gaussian.txt">bibtex</a>
635              
636              <p></p>
637              <p>The first 3D Gaussian splatting-based method for X-ray 3D reconstruction</p>
638
639            </td>
640          </tr>
641          
642          
643          
644          
645          
646          <tr>
647            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
648              <img src="/images/acca.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
649            </td>
650            <td style="padding:2.5%;width:70%;vertical-align:middle">
651              <h3>Unveiling Advanced Frequency Disentanglement Paradigm for Low-Light Image Enhancement</h3>
652              <br>
653              <br>
654              <a href="https://yixunliang.github.io/">Kun Zhou</a>, <a href="https://sds.cuhk.edu.cn/en/node/678">Xinyu Lin</a>, <a href="https://fenglinglwb.github.io/">Wenbo Li</a>, <a href="https://xuxiaogang.com/">Xiaogang Xu</a>, <strong>Yuanhao Cai </strong>, <a href="https://zhonghang-liu.github.io/homepage/">Zhonghang Liu</a>, <a href="https://scholar.google.com/citations?user=z-rqsR4AAAAJ&hl=zh-CN">Xiaoguang Han</a>, <a href="https://sites.google.com/site/jiangbolu/">Jiangbo Lu</a>
655
656              <br>
657              <em>European Conference on Computer Vision (ECCV)</em>, 2024
658              <br>
659              
660              <a href="https://arxiv.org/pdf/2403.04116.pdf">paper</a>
661              
662              
663              / <a href="https://github.com/redrock303/ADF-LLIE">project</a>
664              
665              
666              
667              
668              
669              
670              
671              <p></p>
672              <p>A light-weight network for low-light image enhancement</p>
673
674            </td>
675          </tr>
676          
677          
678          
679          
680          
681          <tr>
682            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
683              <img src="/images/sax-nerf.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
684            </td>
685            <td style="padding:2.5%;width:70%;vertical-align:middle">
686              <h3>Structure-Aware Sparse-View X-ray 3D Reconstruction</h3>
687              <br>
688              <br>
689              <strong>Yuanhao Cai </strong>, <a href="https://jiahaoplus.github.io/">jiahao Wang</a>, <a href="https://www.zongweiz.com/">Zongwei Zhou</a>, <a href="https://scholar.google.com/citations?hl=en&user=YR7re-cAAAAJ">Angtian Wang</a>, <a href="https://www.cs.jhu.edu/~ayuille/">Alan Yuille</a>
690
691              <br>
692              <em>Conference on Computer Vision and Pattern Recognition (CVPR)</em>, 2024
693              <br>
694              
695              <a href="https://arxiv.org/pdf/2311.10959.pdf">paper</a>
696              
697              
698              / <a href="https://github.com/caiyuanhao1998/SAX-NeRF">project</a>
699              
700              
701              / <a href="https://www.youtube.com/watch?v=oVVUaBY61eo">video</a> 
702              
703              
704              / <a href="https://zhuanlan.zhihu.com/p/702702109">zhihu</a>
705              
706              
707              / <a href="https://paperswithcode.com/dataset/x3d">leaderboard</a>
708              
709              
710              
711              
712              / <a href="/bibtex/saxnerf.txt">bibtex</a>
713              
714              <p></p>
715              <p>A NeRF algorithm capturing structures for large-scale X-ray 3D reconstruction</p>
716
717            </td>
718          </tr>
719          
720          
721          
722          
723          
724          <tr>
725            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
726              <img src="/images/bisci.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
727            </td>
728            <td style="padding:2.5%;width:70%;vertical-align:middle">
729              <h3>Binarized Spectral Compressive Imaging</h3>
730              <br>
731              <br>
732              <strong>Yuanhao Cai </strong>, Yuxin Zheng, <a href="https://jinglin7.github.io/">Jing Lin</a>, <a href="https://en.westlake.edu.cn/faculty/xin-yuan.html">Xin Yuan</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>
733
734              <br>
735              <em>Advances in Neural Information Processing Systems (NeurIPS)</em>, 2023
736              <br>
737              
738              <a href="https://arxiv.org/pdf/2305.10299.pdf">paper</a>
739              
740              
741              / <a href="https://github.com/caiyuanhao1998/BiSCI">project</a>
742              
743              
744              
745              / <a href="https://zhuanlan.zhihu.com/p/668862020">zhihu</a>
746              
747              
748              
749              
750              
751              / <a href="/bibtex/bisci.txt">bibtex</a>
752              
753              <p></p>
754              <p>An Efficient Retinex-based method for Low-light Image Enhancement</p>
755
756            </td>
757          </tr>
758          
759          
760          
761          
762          
763          <tr>
764            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
765              <img src="/images/motionx.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
766            </td>
767            <td style="padding:2.5%;width:70%;vertical-align:middle">
768              <h3>
768Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset</h3>
769              <br>
770              <br>
771              <a href="https://jinglin7.github.io/">Jing Lin *</a>, <a href="https://ailingzeng.site/">Ailing Zeng *</a>, <a href="https://shunlinlu.github.io/">Shunling Lu *</a>, <strong>Yuanhao Cai </strong>, <a href="http://www.zhangruimao.site/">Ruimao Zhang</a>, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://www.leizhang.org/">Lei Zhang</a>
772
773              <br>
774              <em>Advances in Neural Information Processing Systems (NeurIPS)</em>, 2023
775              <br>
776              
777              <a href="https://openreview.net/attachment?id=WtajAo0JWU&name=pdf">paper</a>
778              
779              
780              / <a href="https://motion-x-dataset.github.io/">project</a>
781              
782              
783              
784              
785              
786              
787              
788              / <a href="/bibtex/motionx.txt">bibtex</a>
789              
790              <p></p>
791              <p>A large-scale human motion benchmark with text description</p>
792
793            </td>
794          </tr>
795          
796          
797          
798          
799          
800          <tr>
801            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
802              <img src="/images/retinexformer.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
803            </td>
804            <td style="padding:2.5%;width:70%;vertical-align:middle">
805              <h3>Retinexformer: One-stage Retinex-based Transformer for Low-light Image Enhancement</h3>
806              <br>
807              <br>
808              <strong>Yuanhao Cai </strong>, <a href="https://bianhao123.github.io/">Hao Bian</a>, Jing Lin, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="http://people.ee.ethz.ch/~timofter/">Radu Timofte</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>
809
810              <br>
811              <em>International Conference on Computer Vision (ICCV)</em>, 2023
812              <br>
813              
814              <a href="https://arxiv.org/pdf/2303.06705.pdf">paper</a>
815              
816              
817              / <a href="https://github.com/caiyuanhao1998/Retinexformer">project</a>
818              
819              
820              
821              / <a href="https://zhuanlan.zhihu.com/p/657927878">zhihu</a>
822              
823              
824              
825              
826              
827              / <a href="/bibtex/retinexformer.txt">bibtex</a>
828              
829              <p></p>
830              <p>An Efficient Retinex-based method for Low-light Image Enhancement</p>
831
832            </td>
833          </tr>
834          
835          
836          
837          
838          
839          <tr>
840            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
841              <img src="/images/DAUHST.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
842            </td>
843            <td style="padding:2.5%;width:70%;vertical-align:middle">
844              <h3>Degradation-Aware Unfolding Half-Shuffle Transformer for Spectral Compressive Imaging</h3>
845              <br>
846              <br>
847              <strong>Yuanhao Cai *</strong>, Jing Lin *, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://www.bell-labs.com/about/researcher-profiles/xyuan/">Xin Yuan</a>, <a href="https://henghuiding.github.io/">Henghui Ding</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="http://people.ee.ethz.ch/~timofter/">Radu Timofte</a>, <a href="https://ee.ethz.ch/the-department/faculty/professors/person-detail.OTAyMzM=.TGlzdC80MTEsMTA1ODA0MjU5.html">Luc Van Gool</a>
848
849              <br>
850              <em>Advances in Neural Information Processing Systems (NeurIPS)</em>, 2022
851              <br>
852              
853              <a href="https://arxiv.org/pdf/2205.10102.pdf">paper</a>
854              
855              
856              / <a href="https://github.com/caiyuanhao1998/MST">project</a>
857              
858              
859              
860              / <a href="https://zhuanlan.zhihu.com/p/576280023">zhihu</a>
861              
862              
863              
864              
865              
866              / <a href="/bibtex/dauhst.txt">bibtex</a>
867              
868              <p></p>
869              <p>The first Transformer-based deep unfolding method for spectral compressive imaging</p>
870
871            </td>
872          </tr>
873          
874          
875          
876          
877          
878          <tr>
879            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
880              <img src="/images/CST.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
881            </td>
882            <td style="padding:2.5%;width:70%;vertical-align:middle">
883              <h3>Coarse-to-Fine Sparse Transformer for Hyperspectral Image Reconstruction</h3>
884              <br>
885              <br>
886              <strong>Yuanhao Cai *</strong>, Jing Lin *, Xiaowan Hu, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://www.bell-labs.com/about/researcher-profiles/xyuan/">Xin Yuan</a>,  <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="http://people.ee.ethz.ch/~timofter/">Radu Timofte</a>, <a href="https://ee.ethz.ch/the-department/faculty/professors/person-detail.OTAyMzM=.TGlzdC80MTEsMTA1ODA0MjU5.html">Luc Van Gool</a>
887
888              <br>
889              <em>European Conference on Computer Vision (ECCV)</em>, 2022
890              <br>
891              
892              <a href="https://arxiv.org/pdf/2203.04845.pdf">paper</a>
893              
894              
895              / <a href="https://github.com/caiyuanhao1998/MST">project</a>
896              
897              
898              
899              / <a href="https://zhuanlan.zhihu.com/p/544979161">zhihu</a>
900              
901              
902              
903              
904              
905              / <a href="/bibtex/cst.txt">bibtex</a>
906              
907              <p></p>
908              <p>A novel SOTA Transformer-based method for hyperspectral image reconstruction</p>
909
910            </td>
911          </tr>
912          
913          
914          
915          
916          
917          <tr>
918            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
919              <img src="/images/FGST.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
920            </td>
921            <td style="padding:2.5%;width:70%;vertical-align:middle">
922              <h3>Flow-Guided Sparse Transformer for Video Deblurring</h3>
923              <br>
924              <br>
925              Jing Lin *, <strong>Yuanhao Cai *</strong>, Xiaowan Hu, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://scholar.google.com.hk/citations?hl=zh-CN&user=JPUwfAMAAAAJ">Youliang Yan</a>, <a href="https://scholar.google.com.hk/citations?hl=zh-CN&user=0ua28KoAAAAJ">Xueyi Zou</a>, <a href="https://henghuiding.github.io/">Henghui Ding</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="http://people.ee.ethz.ch/~timofter/">Radu Timofte</a>, <a href="https://ee.ethz.ch/the-department/faculty/professors/person-detail.OTAyMzM=.TGlzdC80MTEsMTA1ODA0MjU5.html">Luc Van Gool</a>
926
927              <br>
928              <em>International Conference on Machine Learning (ICML)</em>, 2022
929              <br>
930              
931              <a href="https://arxiv.org/pdf/2201.01893.pdf">paper</a>
932              
933              
934              / <a href="https://github.com/linjing7/VR-Baseline">project</a>
935              
936              
937              
938              
939              
940              
941              
942              / <a href="/bibtex/fgst.txt">bibtex</a>
943              
944              <p></p>
945              <p>The first Transformer-based method for video deblurring</p>
946
947            </td>
948          </tr>
949          
950          
951          
952          
953          
954          <tr>
955            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
956              <img src="/images/Seq2Seq.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
957            </td>
958            <td style="padding:2.5%;width:70%;vertical-align:middle">
959              <h3>Unsupervised Flow-Aligned Sequence-to-Sequence Learning for Video Restoration</h3>
960              <br>
961              <br>
962              Jing Lin *, Xiaowan Hu *, <strong>Yuanhao Cai</strong>,  <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://scholar.google.com.hk/citations?hl=zh-CN&user=JPUwfAMAAAAJ">Youliang Yan</a>, <a href="https://scholar.google.com.hk/citations?hl=zh-CN&user=0ua28KoAAAAJ">Xueyi Zou</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="https://ee.ethz.ch/the-department/faculty/professors/person-detail.OTAyMzM=.TGlzdC80MTEsMTA1ODA0MjU5.html">Luc Van Gool</a>
963
964              <br>
965              <em>International Conference on Machine Learning (ICML)</em>, 2022
966              <br>
967              
968              <a href="https://arxiv.org/pdf/2205.10195.pdf">paper</a>
969              
970              
971              / <a href="https://github.com/linjing7/VR-Baseline">project</a>
972              
973              
974              
975              
976              
977              
978              
979              / <a href="/bibtex/seq2seq.txt">bibtex</a>
980              
981              <p></p>
982              <p>The first Sequence-to-Sequence model for video restoration</p>
983
984            </td>
985          </tr>
986          
987          
988          
989          
990          
991          <tr>
992            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
993              <img src="/images/compare_fig.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
994            </td>
995            <td style="padding:2.5%;width:70%;vertical-align:middle">
996              <h3>MST++: Multi-stage Spectral-wise Transformer for Efficient Spectral Reconstruction</h3>
997              <br>
998              <br>
999              <strong>Yuanhao Cai *</strong>, Jing Lin *, Zudi Lin, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>,  <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="https://vcg.seas.harvard.edu/people/hanspeter-pfister">Hanspeter Pfister</a>, <a href="http://people.ee.ethz.ch/~timofter/">Radu Timofte</a>, <a href="https://ee.ethz.ch/the-department/faculty/professors/person-detail.OTAyMzM=.TGlzdC80MTEsMTA1ODA0MjU5.html">Luc Van Gool</a>
1000
1001              <br>
1002              <em>Conference on Computer Vision and Pattern Recognition Workshop (CVPRW)</em>, 2022
1003              <br>
1004              
1005              <a href="https://arxiv.org/pdf/2204.07908.pdf">paper</a>
1006              
1007              
1008              / <a href="https://github.com/caiyuanhao1998/MST-plus-plus">project</a>
1009              
1010              
1011              
1012              / <a href="https://zhuanlan.zhihu.com/p/501101943?utm_source=wechat_session&utm_medium=social&utm_oi=980437177842446336&utm_content=group3_article&utm_campaign=shareopn">zhihu</a>
1013              
1014              
1015              
1016              
1017              
1018              / <a href="/bibtex/mst_pp.txt">bibtex</a>
1019              
1020              <p></p>
1021              <p>Winner of NTIRE 2022 Challenge on Spectral Reconstruction from RGB. The first Transformer-based method for spectral reconstruction. A baseline and toolbox.</p>
1022
1023            </td>
1024          </tr>
1025          
1026          
1027          
1028          
1029          
1030          <tr>
1031            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
1032              <img src="/images/MST.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
1033            </td>
1034            <td style="padding:2.5%;width:70%;vertical-align:middle">
1035              <h3>Mask-guided Spectral-wise Transformer for Efficient Hyperspectral Image Reconstruction</h3>
1036              <br>
1037              <br>
1038              <strong>Yuanhao Cai *</strong>, Jing Lin *,  Xiaowan Hu, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://www.bell-labs.com/about/researcher-profiles/xyuan/">Xin Yuan</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="http://people.ee.ethz.ch/~timofter/">Radu Timofte</a>, <a href="https://ee.ethz.ch/the-department/faculty/professors/person-detail.OTAyMzM=.TGlzdC80MTEsMTA1ODA0MjU5.html">Luc Van Gool</a>
1039
1040              <br>
1041              <em>Conference on Computer Vision and Pattern Recognition (CVPR)</em>, 2022
1042              <br>
1043              
1044              <a href="https://arxiv.org/pdf/2111.07910.pdf">paper</a>
1045              
1046              
1047              / <a href="https://github.com/caiyuanhao1998/MST">project</a>
1048              
1049              
1050              
1051              / <a href="https://zhuanlan.zhihu.com/p/501101943?utm_source=wechat_session&utm_medium=social&utm_oi=980437177842446336&utm_content=group3_article&utm_campaign=shareopn">zhihu</a>
1052              
1053              
1054              
1055              
1056              
1057              / <a href="/bibtex/mst.txt">bibtex</a>
1058              
1059              <p></p>
1060              <p>The first Transformer-based method for hyperspectral image reconstruction</p>
1061
1062            </td>
1063          </tr>
1064          
1065          
1066          
1067          
1068          
1069          <tr>
1070            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
1071              <img src="/images/HDNet.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
1072            </td>
1073            <td style="padding:2.5%;width:70%;vertical-align:middle">
1074              <h3>HDNet: High-resolution Dual-domain Learning for Spectral Compressive Imaging</h3>
1075              <br>
1076              <br>
1077              Xiaowan Hu *, <strong>Yuanhao Cai *</strong>, Jing Lin, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://www.bell-labs.com/about/researcher-profiles/xyuan/">Xin Yuan</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="http://people.ee.ethz.ch/~timofter/">Radu Timofte</a>, <a href="https://ee.ethz.ch/the-department/faculty/professors/person-detail.OTAyMzM=.TGlzdC80MTEsMTA1ODA0MjU5.html">Luc Van Gool</a>
1078
1079              <br>
1080              <em>Conference on Computer Vision and Pattern Recognition (CVPR)</em>, 2022
1081              <br>
1082              
1083              <a href="https://arxiv.org/pdf/2203.02149.pdf">paper</a>
1084              
1085              
1086              / <a href="https://github.com/caiyuanhao1998/MST">project</a>
1087              
1088              
1089              
1090              
1091              
1092              
1093              
1094              / <a href="/bibtex/hdnet.txt">bibtex</a>
1095              
1096              <p></p>
1097              <p>Dual-domain learning for hyperspectral image reconstruction</p>
1098
1099            </td>
1100          </tr>
1101          
1102          
1103          
1104          
1105          
1106          <tr>
1107            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
1108              <img src="/images/RFormer.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
1109            </td>
1110            <td style="padding:2.5%;width:70%;vertical-align:middle">
1111              <h3>RFormer: Transformer-based Generative Adversarial Network for Real Fundus Image Restoration on A New Clinical Benchmark</h3>
1112              <br>
1113              <br>
1114              Zhuo Deng *, <strong>Yuanhao Cai *</strong>, Lu Chen, Zheng Gong, Qiqi Bao, Xue Yao, Dong Fang, Shaochong Zhang, <a href="https://sklco.pkusz.edu.cn/info/1030/1046.htm">Lan Ma</a>
1115
1116              <br>
1117              <em>
1117IEEE Journal of Biomedical and Health Informatics (J-BHI)</em>, 2022
1118              <br>
1119              
1120              <a href="https://arxiv.org/pdf/2201.00466.pdf">paper</a>
1121              
1122              
1123              / <a href="https://github.com/dengzhuo-AI/Real-Fundus">project</a>
1124              
1125              
1126              
1127              
1128              
1129              
1130              
1131              / <a href="/bibtex/rformer.txt">bibtex</a>
1132              
1133              <p></p>
1134              <p>The first clinical benchmark and Transformer-based method for fundus image restoration</p>
1135
1136            </td>
1137          </tr>
1138          
1139          
1140          
1141          
1142          
1143          <tr>
1144            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
1145              <img src="/images/PNGAN.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
1146            </td>
1147            <td style="padding:2.5%;width:70%;vertical-align:middle">
1148              <h3>Learning to Generate Realistic Noisy Images via Pixel-level Noise-aware Adversarial Training</h3>
1149              <br>
1150              <br>
1151              <strong>Yuanhao Cai </strong>, Xiaowan Hu, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="https://vcg.seas.harvard.edu/people/hanspeter-pfister">Hanspeter Pfister</a>, <a href="https://donglaiw.github.io/">Donglai Wei</a>
1152
1153              <br>
1154              <em>Advances in Neural Information Processing Systems (NeurIPS)</em>, 2021
1155              <br>
1156              
1157              <a href="https://proceedings.neurips.cc/paper/2021/file/1a5b1e4daae265b790965a275b53ae50-Paper.pdf">paper</a>
1158              
1159              
1160              / <a href="https://github.com/caiyuanhao1998/PNGAN">project</a>
1161              
1162              
1163              
1164              
1165              / <a href="https://noise.visinf.tu-darmstadt.de/benchmark/">leaderboard</a>
1166              
1167              
1168              
1169              
1170              / <a href="/bibtex/pngan.txt">bibtex</a>
1171              
1172              <p></p>
1173              <p>A GAN for real noisy image generation</p>
1174
1175            </td>
1176          </tr>
1177          
1178          
1179          
1180          
1181          
1182          <tr>
1183            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
1184              <img src="/images/MSFN.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
1185            </td>
1186            <td style="padding:2.5%;width:70%;vertical-align:middle">
1187              <h3>Multi-Scale Selective Feedback Network with Dual Loss for Real Image Denoising</h3>
1188              <br>
1189              <br>
1190              Xiaowan Hu, <strong>Yuanhao Cai</strong>, Zhihong Liu, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>
1191
1192              <br>
1193              <em>International Joint Conference on Artificial Intelligence (IJCAI), <strong>Oral</strong></em>, 2021
1194              <br>
1195              
1196              <a href="https://www.ijcai.org/proceedings/2021/0101.pdf">paper</a>
1197              
1198              
1199              / <a href="https://www.ijcai.org/proceedings/2021/101">project</a>
1200              
1201              
1202              
1203              
1204              
1205              
1206              
1207              / <a href="/bibtex/msfn.txt">bibtex</a>
1208              
1209              <p></p>
1210              <p>A semi-supervised method for real image denoising</p>
1211
1212            </td>
1213          </tr>
1214          
1215          
1216          
1217          
1218          
1219          <tr>
1220            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
1221              <img src="/images/P3AN.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
1222            </td>
1223            <td style="padding:2.5%;width:70%;vertical-align:middle">
1224              <h3>Pseudo 3D Auto-Correlation Network for Real Image Denoising</h3>
1225              <br>
1226              <br>
1227              Xiaowan Hu, <a href="https://scholar.google.com/citations?user=k6irHZ0AAAAJ&hl=en">Ruijun Ma</a>, Zhihong Liu, <strong>Yuanhao Cai </strong>, <a href="https://scholar.google.com/citations?user=xkK4mRUAAAAJ&hl=en">Xiaole Zhao</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>
1228
1229              <br>
1230              <em>Conference on Computer Vision and Pattern Recognition (CVPR)</em>, 2021
1231              <br>
1232              
1233              <a href="https://openaccess.thecvf.com/content/CVPR2021/papers/Hu_Pseudo_3D_Auto-Correlation_Network_for_Real_Image_Denoising_CVPR_2021_paper.pdf">paper</a>
1234              
1235              
1236              / <a href="https://openaccess.thecvf.com/content/CVPR2021/html/Hu_Pseudo_3D_Auto-Correlation_Network_for_Real_Image_Denoising_CVPR_2021_paper.html">project</a>
1237              
1238              
1239              
1240              
1241              
1242              
1243              
1244              / <a href="/bibtex/p3an.txt">bibtex</a>
1245              
1246              <p></p>
1247              <p>
1247An efficient architecture for image denoising</p>
1248
1249            </td>
1250          </tr>
1251          
1252          
1253          
1254          
1255          
1256          <tr>
1257            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
1258              <img src="/images/DANet.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
1259            </td>
1260            <td style="padding:2.5%;width:70%;vertical-align:middle">
1261              <h3>Efficient Human Pose Estimation by Learning Deeply Aggregated Representations</h3>
1262              <br>
1263              <br>
1264              <a href="https://scholar.google.com.hk/citations?user=Sz1yTZsAAAAJ&hl=zh-CN">Zhengxion Luo</a>, <a href="https://scholar.google.com/citations?user=0QBBNGoAAAAJ&hl=zh-CN">Zhicheng Wang</a>, <strong>Yuanhao Cai </strong>, <a href="https://wangguanan.github.io/">Guan'an Wang</a>, <a href="http://www.cbsr.ia.ac.cn/users/liangwang/">Liang Wang</a>, <a href="https://yanrockhuang.github.io/">Yan Huang</a>, <a href="https://scholar.google.com/citations?user=k2ziPUsAAAAJ&hl=zh-CN">ErJin Zhou</a>, <a href="http://www.cbsr.ia.ac.cn/users/tnt/tnt.htm">Tieniu Tan</a>, <a href="http://www.jiansun.org/">Jian Sun</a>
1265
1266              <br>
1267              <em>International Conference on Multimedia and Expo (ICME), <strong>Oral</strong></em>, 2021
1268              <br>
1269              
1270              <a href="https://arxiv.org/pdf/2012.07033.pdf">paper</a>
1271              
1272              
1273              / <a href="https://ieeexplore.ieee.org/abstract/document/9428206">project</a>
1274              
1275              
1276              
1277              
1278              
1279              
1280              
1281              / <a href="/bibtex/danet.txt">bibtex</a>
1282              
1283              <p></p>
1284              <p>An DenseNet-like backbone for efficient human pose estimation</p>
1285
1286            </td>
1287          </tr>
1288          
1289          
1290          
1291          
1292          
1293          <tr>
1294            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
1295              <img src="/images/POAN.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
1296            </td>
1297            <td style="padding:2.5%;width:70%;vertical-align:middle">
1298              <h3>Pyramid Orthogonal Attention Network based on Dual Self-Similarity for Accurate Mr Image Super-Resolution</h3>
1299              <br>
1300              <br>
1301              Xiaowan Hu, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <strong>Yuanhao Cai </strong>, <a href="https://scholar.google.com/citations?user=xkK4mRUAAAAJ&hl=en">Xiaole Zhao</a>, <a href="https://yulunzhang.com/">Yulun Zhang</a>
1302
1303              <br>
1304              <em>International Conference on Multimedia and Expo (ICME)</em>, 2021
1305              <br>
1306              
1307              <a href="https://cloud.tsinghua.edu.cn/f/6662fa8ad15a4c92b082/">paper</a>
1308              
1309              
1310              / <a href="https://ieeexplore.ieee.org/abstract/document/9428112">project</a>
1311              
1312              
1313              
1314              
1315              
1316              
1317              
1318              / <a href="/bibtex/poan.txt">bibtex</a>
1319              
1320              <p></p>
1321              <p>A novel self-attention mechanism for MR image super-resolution</p>
1322
1323            </td>
1324          </tr>
1325          
1326          
1327          
1328          
1329          
1330          <tr>
1331            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
1332              <img src="/images/RSN.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
1333            </td>
1334            <td style="padding:2.5%;width:70%;vertical-align:middle">
1335              <h3>Learning Delicate Local Representations for Multi-Person Pose Estimation</h3>
1336              <br>
1337              <br>
1338              <strong>Yuanhao Cai *</strong>, <a href="https://scholar.google.com/citations?user=0QBBNGoAAAAJ&hl=zh-CN">Zhicheng Wang</a> *, <a href="https://scholar.google.com.hk/citations?user=Sz1yTZsAAAAJ&hl=zh-CN">Zhengxion Luo</a>, Binyi Yin, Ang'ang Du, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, <a href="https://scholar.google.com/citations?user=yuB-cfoAAAAJ&hl=zh-CN">Xiangyu Zhang</a>, <a href="https://scholar.google.com/citations?user=Jv4LCj8AAAAJ&hl=en">Xinyu Zhou</a>, <a href="https://scholar.google.com/citations?user=k2ziPUsAAAAJ&hl=zh-CN">ErJin Zhou</a>, <a href="http://www.jiansun.org/">Jian Sun</a>
1339
1340              <br>
1341              <em>European Conference on Computer Vision (ECCV), <strong>Spotlight</strong></em>, 2020
1342              <br>
1343              
1344              <a href="https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123480460.pdf">paper</a>
1345              
1346              
1347              / <a href="https://github.com/caiyuanhao1998/RSN">
1347project</a>
1348              
1349              
1350              
1351              / <a href="https://zhuanlan.zhihu.com/p/112297707">zhihu</a>
1352              
1353              
1354              / <a href="https://cocodataset.org/#keypoints-leaderboard">leaderboard</a>
1355              
1356              
1357              
1358              
1359              / <a href="/bibtex/rsn.txt">bibtex</a>
1360              
1361              <p></p>
1362              <p>A multi-stage structure and an attention mechanism for accurate human pose estimation</p>
1363
1364            </td>
1365          </tr>
1366          
1367          
1368          
1369          
1370          
1371          <tr>
1372            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
1373              <img src="/images/UDP.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
1374            </td>
1375            <td style="padding:2.5%;width:70%;vertical-align:middle">
1376              <h3>UDP++</h3>
1377              <br>
1378              <br>
1379              Junjie Huang *, Zengguang Shan *, <strong>Yuanhao Cai *</strong>, Feng Guo, Yun Ye, Xinze Chen, <a href="http://www.zhengzhu.net/">Zheng Zhu</a>, Guan Huang, <a href="http://ivg.au.tsinghua.edu.cn/Jiwen_Lu/">Jiwen Lu</a>, Dalong Du
1380
1381              <br>
1382              <em>European Conference on Computer Vision Workshop (ECCVW), <strong>Oral</strong></em>, 2020
1383              <br>
1384              
1385              <a href="https://s3-us-west-1.amazonaws.com/presentations.cocodataset.org/ECCV20/keypoints/UDP.pdf">paper</a>
1386              
1387              
1388              / <a href="https://github.com/caiyuanhao1998/UDP-plus-plus">project</a>
1389              
1390              
1391              
1392              / <a href="https://zhuanlan.zhihu.com/p/210199401">zhihu</a>
1393              
1394              
1395              / <a href="https://cocodataset.org/#keypoints-leaderboard">leaderboard</a>
1396              
1397              
1398              
1399              
1400              / <a href="/bibtex/udp_pp.txt">bibtex</a>
1401              
1402              <p></p>
1403              <p>Winner of COCO Keypoint Detection Challenge, 2020</p>
1404
1405            </td>
1406          </tr>
1407          
1408          
1409          
1410          
1411          
1412          <tr>
1413            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
1414              <img src="/images/SPIE.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
1415            </td>
1416            <td style="padding:2.5%;width:70%;vertical-align:middle">
1417              <h3>EG^2N: Enhanced Gradient Guiding Network for Single MR Image Super-Resolution</h3>
1418              <br>
1419              <br>
1420              Xiaowan Hu, <strong>Yuanhao Cai</strong>, <a href="https://www.sigs.tsinghua.edu.cn/whq/">Haoqian Wang</a>, Yanbin Peng, <a href="https://yulunzhang.com/">Yulun Zhang</a>
1421
1422              <br>
1423              <em>Optoelectronic Imaging and Multimedia Technology VII 11550, 115500I</em>, 2020
1424              <br>
1425              
1426              <a href="https://cloud.tsinghua.edu.cn/f/78a711714198479d991f/">paper</a>
1427              
1428              
1429              / <a href="https://www.spiedigitallibrary.org/conference-proceedings-of-spie/11550/115500I/EG2N--enhanced-gradient-guiding-network-for-single-MR-image/10.1117/12.2575261.short?SSO=1">project</a>
1430              
1431              
1432              
1433              
1434              
1435              
1436              
1437              / <a href="/bibtex/eg2n.txt">bibtex</a>
1438              
1439              <p></p>
1440              <p>A novel gradient guiding mechanism for MR image super-resolution</p>
1441
1442            </td>
1443          </tr>
1444          
1445          
1446          
1447          
1448          
1449          <tr>
1450            <td style="padding:0.1%;width:30%;vertical-align:middle;min-width:120px">
1451              <img src="/images/Res-Step-Net.png" alt="project image" style="width:100%; height:auto; object-fit:contain;" />
1452            </td>
1453            <td style="padding:2.5%;width:70%;vertical-align:middle">
1454              <h3>Res-Steps-Net for Multi-Person Pose Estimation</h3>
1455              <br>
1456              <br>
1457              <strong>Yuanhao Cai *</strong>, <a href="https://scholar.google.com/citations?user=0QBBNGoAAAAJ&hl=zh-CN">Zhicheng Wang</a> *, Binyi Yin, Ruihao Yin, Angang Du, <a href="https://scholar.google.com.hk/citations?user=Sz1yTZsAAAAJ&hl=zh-CN">Zhengxion Luo</a>, <a href="https://www.zemingli.com/">Zeming Li</a>, <a href="https://scholar.google.com/citations?user=Jv4LCj8AAAAJ&hl=en">Xinyu Zhou</a>, <a href="https://www.skicyyu.org/">Gang Yu</a>, <a href="https://scholar.google.com/citations?user=k2ziPUsAAAAJ&hl=zh-CN">ErJin Zhou</a>, <a href="https://scholar.google.com/citations?user=yuB-cfoAAAAJ&hl=zh-CN">Xiangyu Zhang</a>, <a href="https://yichenwei.github.io/">Yichen Wei</a>, <a href="http://www.jiansun.org/">Jian Sun</a>
1458
1459              <br>
1460              <em>International Conference on Computer Vision Workshop (ICCVW),  <strong>Best Paper Award</strong></em>, 2019
1461              <br>
1462              
1463              <a href="https://cocodataset.org/files/keypoints_2019_reports/Megvii.pdf">paper</a>
1464              
1465              
1466              / <a href="https://github.com/caiyuanhao1998/RSN">
1466project</a>
1467              
1468              
1469              
1470              / <a href="https://zhuanlan.zhihu.com/p/112297707">zhihu</a>
1471              
1472              
1473              / <a href="https://cocodataset.org/#keypoints-leaderboard">leaderboard</a>
1474              
1475              
1476              
1477              
1478              / <a href="/bibtex/rsn_iccv.txt">bibtex</a>
1479              
1480              <p></p>
1481              <p>Winner of COCO Keypoint Detection Challenge, 2019</p>
1482
1483            </td>
1484          </tr>
1485          
1486          
1487          
1488        </table>
1489        
1490        <br>
1491        <br>
1492
1493         <table style="width:100%;border:0px;border-spacing:0px;border-collapse:separate;margin-right:auto;margin-left:auto;">
1494          <h2>Academic Service</h2>
1495          <tr style="padding:0px">
1496            <td style="padding:2.5%;vertical-align:left">
1497              <p>
1498                Area Chair: CVPR 2027
1499              </p>
1500              <p>
1501                Conference Reviewer: CVPR, ECCV, ICCV, NeurIPS, ICML, ICLR, AAAI, IJCAI, ACM MM, etc.
1502              </p>
1503              <p>
1504                Journal Reviewer: TPAMI, IJCV, TIP, TNNLS, Pattern Recognition, etc.
1505              </p>
1506            </td>
1507          </tr>
1508        </table>
1509        
1510        <br>
1511      </td>
1512    </tr>
1513  </table>
1514</body>
1515
1516</html>
1517

Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.