arxivcs.CV2026-07-09
Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?
Masatoshi Tateno, Alexandros Stergiou, Risa Shinoda, Yoichi Sato, Dima Damen
Hand-object interaction (HOI) recognition requires capturing both hand manipulations and object transformations. However, existing video-language models often fall into shortcuts by relying on spurious correlations among hands, objects, or environmental context, rather than reaso…