Multimodal chain-of-thought (CoT) reasoning integrates visual and textual cues through step-by-step inference. In small models with limited token budgets, modality-interaction fusion often suppresses tiny cross-modal differences. In particular, multimodal CoT often struggles when…
MEMEs are widely used on the internet and often carry strong elements of sarcasm or irony. Understanding their hidden meanings typically requires a joint interpretation of text and vision. Existing methods focus on the dual-stream vision-language model to extract the visual and t…
Dairy cow management depends on repeated observations of behavior and physical condition to support health, welfare, and operational decisions, but these observations remain labor-intensive. Deep learning (DL)-based computer vision can automate parts of this work, although deploy…
With the increasing demand for reusing paper documents in educational and office settings, accurate segmentation of handwritten and printed text has become a crucial step in document digitization. Although numerous deep learning models have been developed for this task, their hig…
Machine learning (ML) methods have been widely explored for predicting material properties. However, due to the rapid development of ML techniques and the diversity of available models, performance comparisons between traditional and graph-based machine learning models remain lim…
Hexafluoropropylene Oxide Dimer Acid (HFPO-DA or GenX) is a pervasive perfluorinated compound with scant understood toxic effects. Toxicological studies on GenX have been conducted using animal models. To research deeper into the potential toxicity of GenX in humans and animals,…