arxivcs.CLcs.DLcs.HCcs.IR2026-06-30
Building a Multimodal Dataset of Academic Paper for Keyword Extraction
Jingyu Zhang, Xinyi Yan, Yi Xiang, Yingyi Zhang, Chengzhi Zhang
Up to this point, keyword extraction task typically relies solely on textual data. Neglecting visual details and audio features from image and audio modalities leads to deficiencies in information richness and overlooks potential correlations, thereby constraining the model's abi…