arxivcs.CV2026-07-17
Multimodal Ambivalence and Hesitancy Recognition via Cross-Attention and Gated Fusion
Oussama Berhili, Yassine Ouzar, Larbi Boubchir
We present a multimodal framework for Ambivalence/Hesitancy (A/H) recognition in video, developed for the ABAW11 challenge at ECCV 2026. The proposed approach fuses textual, acoustic, and visual modalities extracted from the BAH dataset using three pretrained encoders: F2LLM-v2-0…