News

May 07, 2026 πŸŽ“ Thesis proposal defense passed!
Apr 24, 2026 πŸŽ‰ MMComposition has been accepted by TMLR! [Project Page] [Paper] [Code]
Feb 20, 2026 πŸŽ‰ Video-R4 and Takusen have been accepted to CVPR 2026!
Feb 06, 2026 πŸš€ Call for papers is open for AISTORY Workshop @CVPR 2026! Yolo is serving as an organizer for the second edition of the workshop.
Jan 27, 2026 πŸ† CAT-V received the Best Demo Award Runner-up at the AAAI-26 Demonstration Program!
Dec 02, 2025 🌊 Yolo is in San Diego for NeurIPS 2025.
Nov 18, 2025 🎬 Video-R4 is now on arXiv. Also see the Project Page and GitHub
Oct 19, 2025 🌺 Yolo is attending ICCV 2025 in Honolulu this week.
Oct 06, 2025 🎬 Video-LMM Post-Training, a deep dive into video reasoning with large multimodal models, is now available. [GitHub]
Sep 18, 2025 πŸŽ‰ MMPerspective has been accepted to NeurIPS 2025 Datasets & Benchmarks.
Jul 17, 2025 πŸŽ‰ The GenAI for Cel-Animation survey has been accepted to the ICCV 2025 AISTORY Workshop! [Paper] [GitHub]
Jun 11, 2025 🀠 Yolo is attending CVPR 2025 in Nashville.
May 31, 2025 πŸ“ Introducing MMPerspective, a comprehensive benchmark for MLLMs on perspective understanding.
May 27, 2025 🌟 Started as an Applied Scientist Intern at Amazon in Bellevue, WA.
May 03, 2025 πŸŽ‰ The Vid-LLM survey has been accepted by the IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)! [IEEE Xplore] [GitHub]
Apr 09, 2025 πŸ“· Caption Anything in Video (CAT-V) has been released [arXiv] [GitHub]
Feb 26, 2025 πŸŽ‰ Two papers have been accepted to CVPR 2025, including the VidComposition benchmark!
Feb 24, 2025 πŸ“ Yolo will attend AAAI 2025 in Philadelphia and present 3 papers
Feb 05, 2025 Upcoming: Applied Scientist Intern at Amazon this summer.
Jan 13, 2025 🎨 Introducing the survey paper on GenAI for Cel-Animation [arXiv] [GitHub]
Dec 09, 2024 πŸŽ‰ Three papers on Video-LLMs have been accepted to AAAI 2025!
Nov 23, 2024 VidComposition has been released, a benchmark to evaluate MLLMs’ understanding of video compositions. [Project Page] [Paper] [Leaderboard]
Oct 13, 2024 πŸš€ MMComposition has been publicly released. Read the Paper, check out the latest πŸ†Leaderboard, and access the Code to evaluate models.
Aug 23, 2024 Introducing CaRDiff, a framework for video saliency prediction using MLLM CoT reasoning and diffusion model.
Aug 05, 2024 πŸ… First place in the AIM 2024 Challenge on Video Saliency Prediction @ ECCV Workshop! Thanks to Gen Zhan and Li Yang!
Jul 23, 2024 πŸ“’ Survey update: "Video Understanding with Large Language Models: A Survey"
Jul 15, 2024 One paper about egocentric video understanding with LLM has been accepted to ACM MM 2024.
Jun 18, 2024 Introducing Differentiated Beam Decoding (DBD), a novel decoding strategy for LVLM hallucination mitigation.
May 20, 2024 🌟 Started an internship at ByteDance in San Jose, CA, mentored by Yiting Liao & Gen Zhan.
Apr 18, 2024 Introducing V2Xum-LLaMA model and Instruct-V2Xum dataset for cross-modal video summarization.
Mar 24, 2024 Released AVicuna, an Audio-Visual LLM empowered by pseudo-untrimmed video annotations for audio-visual event localization.
Feb 09, 2024 Upcoming: Research Intern at ByteDance this summer.
Dec 30, 2023 πŸ”₯πŸ”₯πŸ”₯ Released a survey for Video Understanding with LLMs [arXiv] [GitHub].
Aug 28, 2023 Officially joined in the Chenliang Xu’s Group at UR CS as a Ph.D. studentπŸŽ“.
Jul 23, 2023 One paper accepted to International Computer Music Conference (ICMC) 2023.
Jun 29, 2023 Graduated from SUSTech with a bachelor’s degree and the honor of Excellent Graduate for Exceptional Performance.
Jun 18, 2023 The team won first place in the LOVEU (Long-form Video Understanding) Challenge at CVPR’23 Workshop.
May 25, 2023 Successfully defended the undergraduate thesis Language-Guided Video Cover Generation, which was awarded the Excellent Undergraduate Thesis!
May 04, 2023 The technical report for Caption Anything has been released!
Apr 12, 2023 Caption Anything has been released! Try the demo and star the GitHub repo!
Feb 27, 2023 Upcoming: Ph.D. in Computer Science at the University of Rochester from Fall 2023, working with Prof. Chenliang Xu!
Dec 04, 2022 πŸ“ Yolo is attending ACCV 2022 in person in Macau.
Sep 16, 2022 One paper about multimodal Ad video editing has been accepted to Asian Conference on Computer Vision (ACCV) 2022.
Aug 16, 2022 Left Tencent and joined SUSTech VIP Lab as an undergraduate research assistant.
Sep 24, 2021 Started a part-time internship at Tencent in Shenzhen, with supervision from Dr. Wenhao Jiang and Qin Lin, alongside university coursework.