Abstract The rapid growth of multimodal content has introduced new challenges in cybersecurity, particularly in scenarios such as misinformation detection, multimedia forensics, and open-source intelligence. In these settings, verifying the consistency between textual description…
Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) policies unify target identification and trajectory planning, their chain-of-thought…
Benefiting from the powerful priors embedded in large-scale pre-training data and the emerging commonsense reasoning ability, large language models (LLMs) have shown unprecedented generalization capabilities in many research fields. Recently, projecting visual embeddings into the…
Background: Understanding the genetic mechanisms and identifying potential therapeutic targets are essential for clarifying Autism Spectrum Disorder (ASD) etiology and improving treatments. This study aims to bridge the gap between basic transcriptomic discoveries and clinical ap…