Vision-language alignment, the stage that bridges pretrained vision encoders and large language models, is widely treated as a form of pretraining requiring full-parameter updates. We challenge this view and investigate what happens when low-rank adaptation is applied to the LLM…
As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their navigation and task completion performance has emerged as a critical research priority. However, existing benchmarks exhibit fundam…
In the Internet of Vehicles (IoV), transmitting high-dimensional multi-modal sensory data to edge servers for time-sensitive tasks faces severe spectrum bottlenecks. To address this, we propose a foundation model-driven over-the-air token fusion (AirTF) framework for task-oriente…