| May 07, 2026 | π Thesis proposal defense passed! |
| Apr 24, 2026 | π MMComposition has been accepted by TMLR! [Project Page] [Paper] [Code] |
| Feb 20, 2026 | π Video-R4 and Takusen have been accepted to CVPR 2026! |
| Feb 06, 2026 | π Call for papers is open for AISTORY Workshop @CVPR 2026! Yolo is serving as an organizer for the second edition of the workshop. |
| Jan 27, 2026 | π CAT-V received the Best Demo Award Runner-up at the AAAI-26 Demonstration Program! |
| Dec 02, 2025 | π Yolo is in San Diego for NeurIPS 2025. |
| Nov 18, 2025 | π¬ Video-R4 is now on arXiv. Also see the Project Page and GitHub |
| Oct 19, 2025 | πΊ Yolo is attending ICCV 2025 in Honolulu this week. |
| Oct 06, 2025 | π¬ Video-LMM Post-Training, a deep dive into video reasoning with large multimodal models, is now available. [GitHub] |
| Sep 18, 2025 | π MMPerspective has been accepted to NeurIPS 2025 Datasets & Benchmarks. |
| Jul 17, 2025 | π The GenAI for Cel-Animation survey has been accepted to the ICCV 2025 AISTORY Workshop! [Paper] [GitHub] |
| Jun 11, 2025 | π€ Yolo is attending CVPR 2025 in Nashville. |
| May 31, 2025 | π Introducing MMPerspective, a comprehensive benchmark for MLLMs on perspective understanding. |
| May 27, 2025 | π Started as an Applied Scientist Intern at Amazon in Bellevue, WA. |
| May 03, 2025 | π The Vid-LLM survey has been accepted by the IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)! [IEEE Xplore] [GitHub] |
| Apr 09, 2025 | π· Caption Anything in Video (CAT-V) has been released [arXiv] [GitHub] |
| Feb 26, 2025 | π Two papers have been accepted to CVPR 2025, including the VidComposition benchmark! |
| Feb 24, 2025 | π Yolo will attend AAAI 2025 in Philadelphia and present 3 papers |
| Feb 05, 2025 | Upcoming: Applied Scientist Intern at Amazon this summer. |
| Jan 13, 2025 | π¨ Introducing the survey paper on GenAI for Cel-Animation [arXiv] [GitHub] |
| Dec 09, 2024 | π Three papers on Video-LLMs have been accepted to AAAI 2025! |
| Nov 23, 2024 | VidComposition has been released, a benchmark to evaluate MLLMsβ understanding of video compositions. [Project Page] [Paper] [Leaderboard] |
| Oct 13, 2024 | π MMComposition has been publicly released. Read the Paper, check out the latest πLeaderboard, and access the Code to evaluate models. |
| Aug 23, 2024 | Introducing CaRDiff, a framework for video saliency prediction using MLLM CoT reasoning and diffusion model. |
| Aug 05, 2024 | π
First place in the AIM 2024 Challenge on Video Saliency Prediction @ ECCV Workshop! Thanks to Gen Zhan and Li Yang! |
| Jul 23, 2024 | π’ Survey update: "Video Understanding with Large Language Models: A Survey" |
| Jul 15, 2024 | One paper about egocentric video understanding with LLM has been accepted to ACM MM 2024. |
| Jun 18, 2024 | Introducing Differentiated Beam Decoding (DBD), a novel decoding strategy for LVLM hallucination mitigation. |
| May 20, 2024 | π Started an internship at ByteDance in San Jose, CA, mentored by Yiting Liao & Gen Zhan. |
| Apr 18, 2024 | Introducing V2Xum-LLaMA model and Instruct-V2Xum dataset for cross-modal video summarization. |
| Mar 24, 2024 | Released AVicuna, an Audio-Visual LLM empowered by pseudo-untrimmed video annotations for audio-visual event localization. |
| Feb 09, 2024 | Upcoming: Research Intern at ByteDance this summer. |
| Dec 30, 2023 | π₯π₯π₯ Released a survey for Video Understanding with LLMs [arXiv] [GitHub]. |
| Aug 28, 2023 | Officially joined in the Chenliang Xuβs Group at UR CS as a Ph.D. studentπ. |
| Jul 23, 2023 | One paper accepted to International Computer Music Conference (ICMC) 2023. |
| Jun 29, 2023 | Graduated from SUSTech with a bachelorβs degree and the honor of Excellent Graduate for Exceptional Performance. |
| Jun 18, 2023 | The team won first place in the LOVEU (Long-form Video Understanding) Challenge at CVPRβ23 Workshop. |
| May 25, 2023 | Successfully defended the undergraduate thesis Language-Guided Video Cover Generation, which was awarded the Excellent Undergraduate Thesis! |
| May 04, 2023 | The technical report for Caption Anything has been released! |
| Apr 12, 2023 | Caption Anything has been released! Try the demo and star the GitHub repo! |
| Feb 27, 2023 | Upcoming: Ph.D. in Computer Science at the University of Rochester from Fall 2023, working with Prof. Chenliang Xu! |
| Dec 04, 2022 | π Yolo is attending ACCV 2022 in person in Macau. |
| Sep 16, 2022 | One paper about multimodal Ad video editing has been accepted to Asian Conference on Computer Vision (ACCV) 2022. |
| Aug 16, 2022 | Left Tencent and joined SUSTech VIP Lab as an undergraduate research assistant. |
| Sep 24, 2021 | Started a part-time internship at Tencent in Shenzhen, with supervision from Dr. Wenhao Jiang and Qin Lin, alongside university coursework. |