arxivcs.CV2026-07-12
Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos
Hao Zheng, Jinyi Huang, Tiantian Zheng, Xun Xu, Tuka Alhanai
Assembly action understanding is a key enabler for effective human-robot collaborative assembly, yet it remains challenging due to subtle motions and fine-grained hand-object interactions. We adapt vision-language models (VLMs) to this challenging domain with Compositional Contex…