Omni-modal orchestration
Training-free coordination of modality experts through LLM routing, persistent memory, and full-duplex interaction.
M.S. student at MAC LabXiamen University
Interactive multimodal intelligence
I study how AI systems can listen, see, reason, act, and revise continuously across time.
Omni models · Agentic systems
Long-video reasoning · Efficient inference
“The task we must set for ourselves is not to feel secure, but to be able to tolerate insecurity.”
我们需要为自己设定的任务,不是拥有安全感,而是能够接受不安全感。
Erich Fromm
Current research
I am an M.S. student admitted in Fall 2025 at the MAC Lab, Xiamen University, advised by Prof. Xiawu Zheng.
Training-free coordination of modality experts through LLM routing, persistent memory, and full-duplex interaction.
Policies, memory, tools, and HCI-oriented evaluation for AI systems that must complete grounded tasks.
Event-aware and semantic-boundary-aware frame selection for efficient long-form video understanding.
Speculative decoding, online drafting, and lightweight adaptation for deployable foundation models.
Selected publications
ICML 2026 · First author · CCF-A
A modular orchestration framework with explicit expert routing, cross-modal memory, and full-duplex streaming interaction.
Audio-visual benchmark · First author
A benchmark for who speaks, when to intervene, and how to generate natural interruptions in audio-visual interaction.
CVPR 2026 · CCF-A
Semantic-boundary-aware frame selection for compact, informative, and coherent long-video understanding.
News
Training-Free Multimodal Large Language Model Orchestration was accepted to ICML 2026 as first-author work.
SocialOmni was released on arXiv as a first-author benchmark for audio-visual social interactivity.
WFS-SB was accepted to CVPR 2026 for efficient long-video understanding.
Joined Xiamen University MAC Lab as a master's student in Artificial Intelligence.
Submitted routing-guided expert selection work for mitigating gradient interference in MoE models.