Sixun Dong (董思勋)
Multimodal Learning, VLM, LLM Agent
University of Central Florida, FL, USA
I am a Ph.D. student at the Institute of Artificial Intelligence (IAI), University of Central Florida, advised by Dr. Chen Chen. I build multimodal AI systems that perceive, generate, and act. My research evolved from visual-temporal perception (CVPR Oral'22, CVPR'23) through generative and world modeling (3DV'24, NeurIPS'25) to my current focus on efficient multimodal and long-sequence inference and LLM-based agentic systems (ICLR'26, NeurIPS'26 ×2, ICASSP'26, ECCV'26), exploring system-algorithm co-design for agent harnesses: the inference, memory, and orchestration layer behind long-horizon multimodal agents. I completed my Master's at ShanghaiTech University under Professor Shenghua Gao.
Recent News
Research Roadmap
Selected Publications
NeurIPS'26 Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency
arXiv (Coming Soon) / OpenReview / Code (Coming Soon)
Reframes long-video VLM efficiency as a joint allocation over frame count, per-frame resolution, and front-end decoding latency, not just which tokens to keep. LoHi is a training-free, single-pass framework that pairs a dense low-resolution video stream with a few high-resolution frames (picked from codec I-frames or by a query-aware CLIP DPP). It gains +10.6% over the native-resolution baseline at matched token budget and +5.2% over the strongest prior efficiency methods, while cutting front-end decoding latency by up to 7x on hour-scale videos. Joint work with UCF, Axon & Meta.
NeurIPS'26 MARS: Multi-resolution Adaptive Routing for Sequential Recommendation
arXiv (Coming Soon) / OpenReview / Code (Coming Soon)
Shows that summarization-based recommenders suffer from temporal aliasing: short-lived intent, medium-term interests, and long-term preferences get entangled in one compressed user memory. MARS writes the full history into gated recurrent (linear-attention-style) state tracks with different half-life priors, and a Top-k router lets each cached seed read only the relevant time scales, keeping candidate scoring fixed-size. It beats strong baselines such as VISTA on Amazon, KuaiRand-1K, and KuaiRec, with the largest gains for long-history users. Joint work with Duke & Meta.
A multi-agent 'film crew' that orchestrates off-the-shelf video generators for long-form narrative-to-film generation, coordinated through FilmDSL—a film-oriented domain-specific language that makes cinematic constraints (shots, camera, assets, character personas) explicit. Generation and critic agents plan, synthesize, and repair clips to keep both visual identity and character behavior consistent across the film. Joint work with UMass & MIT.
arXiv'2604 Rethinking Model Efficiency: Multi-Agent Inference with Large Models
Analyzes VLM end-to-end latency and reveals that output token length dominates inference cost. While large models with short outputs outperform small models with long generations, reasoning remains essential for complex tasks. We bridge this by proposing a multi-agent framework where a small model computes the reasoning tokens and transfers them to a large model, achieving large-model accuracy with minimal latency.
ICASSP'26 Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
Proposes an LLM-agent-based post-ASR correction framework for dysarthric speech recognition that focuses on semantic accuracy rather than just minimizing Word Error Rate (WER).
ICLR'26 MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs
Proposes MMTok to efficiently select informative vision tokens via a multimodal coverage maximization strategy, significantly accelerating VLM inference while maintaining model performance.
Introduces an automated data generation framework and evaluation benchmark for complex logical instructions to assess and enhance the multi-step reasoning capabilities of LLMs.
arXiv'2506 Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives
Proposes TimesCLIP, innovatively aligning time series data with visual and textual multi-modal perspectives to effectively enhance both short-term and long-term forecasting performance.
NeurIPS'25 Sculpting Features from Noise: Reward-Guided Hierarchical Diffusion for Task-Optimal Feature Transformation
Proposes a reward-guided hierarchical diffusion model that "sculpts" data representations from noise to generate optimal feature transformations for specific downstream tasks.
Introduces the MLLM-Tool framework and contributes a specialized dataset to empower Multimodal Large Language Models (MLLMs) with the ability to understand, learn, and invoke external tools.
CVPR'23 Weakly Supervised Video Representation Learning with Unaligned Text for Sequential Videos
Proposes a weakly supervised video representation learning approach that utilizes unaligned text for sequential videos, significantly reducing the reliance on fine-grained video-text alignment annotations.
CVPR'22🏆 Oral TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting
Introduces the RepCount dataset and pioneers the use of regression density maps alongside a Transformer-based architecture to encode multi-scale temporal correlations, significantly improving repetitive action counting.
Experience
Research Intern, Tencent Qingyun Program (Top Young Talent Initiative)
Tencent IEG, Motus AI Team
May 2026 - Aug 2026
Responsible for motion understanding, motion generation, and the evaluation harness.
GenAI Research Intern
Zoom Inc., GenAI Research Group
May 2025 - Aug 2025
Worked on VLM and LLM Agent. Published one first-author paper on efficient VLM inference and two collaborative papers on LLM evaluation.
Research Intern
DGene Digital Technology Co., Ltd., Digital Human Algorithm Department
May 2023 - Jan 2024
Led digital human projects: (1) audio-driven talking head video generation, achieving SoTA on commercial and academic benchmarks; (2) 3D human body reconstruction with <7% measurement error; (3) co-speech gesture generation.
Academic Service
Reviewer
Conferences: CVPR (2023–2026), ICCV (2023, 2025), ECCV (2024, 2026), NeurIPS (2025), ICML (2025, 2026),ICLR (2026), ACM MM (2023–2025), ACCV (2024), KDD (2025)
Journals: IEEE Transactions on Multimedia, Neural Networks(Elsevier), ACM Transactions on Knowledge Discovery from Data
Education
Ph.D. in Computer Science
University of Central Florida, USA
M.S. in Computer Science
ShanghaiTech University, China
SVIP-Lab, Advisor: Prof. Shenghua Gao
B.E. in Computer Science (Dual Degree)
Dalian University of Technology, China
B.E. in Process Equipment and Control Engineering
Dalian University of Technology, China