arxivcs.CV2026-06-29
Reweighting Framewise Attention in Video Transformers for Facial Expression Understanding
Seongro Yoon, Donghyeon Cho, Jinsun Park, François Brémond
Understanding facial expressions in videos requires modeling subtle and localized facial dynamics under unconstrained conditions. Although recent Vision Transformer (ViT)-based video models have shown strong performance through large-scale self-supervised pretraining, their atten…