Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requirements. Vision-language models (VLMs) offer a natural way to specify these requirements using language, but existing approaches either use a VLM to predict the g…
In this report, we present Pelican-VLA 0.5, a unified VLA model that integrates vision-language understanding, future-frame generation, and action prediction within a single architecture. Pelican-VLA 0.5 achieves attention-level generalization: without object annotations, segment…
Seismic damage assessment of reinforced concrete (RC) structures is a vital issue for post-earthquake evaluation. Conventional onsite inspection depends greatly on subjective judgments and engineering experiences of human inspectors, and the efficiency is limited to large-scale u…