MEMEs are widely used on the internet and often carry strong elements of sarcasm or irony. Understanding their hidden meanings typically requires a joint interpretation of text and vision. Existing methods focus on the dual-stream vision-language model to extract the visual and t…
Multimodal chain-of-thought (CoT) reasoning integrates visual and textual cues through step-by-step inference. In small models with limited token budgets, modality-interaction fusion often suppresses tiny cross-modal differences. In particular, multimodal CoT often struggles when…
Understanding driver emotion and state is critical for the next generation of intelligent in-cabin systems that ensure safety and enhance human-vehicle interaction. However, existing public datasets for in-cabin affective computing are largely limited to visual modalities and rar…
Pretrained vision-language-action (VLA) policies provide strong language-conditioned manipulation knowledge, but they remain largely vision-driven and can struggle once manipulation enters contact states where the scene is occluded, depth is ambiguous, or small force errors push…
Leaf chlorophyll content (LCC) serves as a vital biochemical indicator of photosynthetic activity and nitrogen status, critical for precision agriculture to optimize crop management. While UAV-based hyperspectral sensing offers maize LCC estimation potential, current methods stru…
Tongue diagnosis is a crucial method in traditional Chinese medicine (TCM) for obtaining information about a patient’s health condition. In this study, we propose a tongue image segmentation method based on deep learning and a pixel-level tongue color classification method utiliz…
Accurate and rapid estimation of the crop yield is essential to precision agriculture. Critical to crop improvement, yield is a primary index for selecting excellent genotypes in crop breeding. Recently developed unmanned aerial vehicle (UAV) platforms and advanced algorithms can…