A Lightweight Appearance-Guided Association Module for Underwater Multi-Fish Tracking
Yanan Zhou, Dlxat Turgun, Jianhua Cao, Zhenyu Zhang, L L Qian
5 papers indexed
Yanan Zhou, Dlxat Turgun, Jianhua Cao, Zhenyu Zhang, L L Qian
Ting Huang, Zhenyu Zhang, Wenyuan Huang, Jian Yang, Hao Tang
Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under changing viewpoints. However, existing multimodal large language models (MLLMs) remain largely semantic-…
Andy Dai, Zexue He, Zhenyu Zhang, Alex Pentland, Jiaxin Pei
Current benchmarks for language models primarily evaluate execution on fully specified tasks. However, real user tasks are often ambiguous. Users arrive with incomplete, exploratory, or even inconsistent goals, requiring the assistant to first determine the intended task before c…
Manni Cui, Ziheng Qin, ZiAn Wang, Ruiqi Liu, Dianyuan Zou, Jianglan Wei, et al.
AI-generated videos (AIGVs) typically contain subtle temporal artifacts that arise from inter-frame inconsistencies rather than within individual frames. A detector that captures such artifacts should therefore benefit from video pretrained backbones over image only ones. In prac…
Haijie Yang, Zhenyu Zhang, Yixuan Dong, Jianjun Qian, Jian Yang
Audio-driven talking head synthesis has achieved impressive progress in lip synchronization and visual quality, yet generating expressive emotional avatars with controllable intensity remains challenging, especially under real-time constraints. In this paper, we present GaussianE…