arxivcs.CVcs.AI2026-07-23
DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation
Sung-Hoon Yoon, Hoyong Kwon, Changgyoon Oh, Kuk-Jin Yoon
Open-vocabulary semantic segmentation (OVSS) leverages textual semantics to segment objects beyond predefined categories. While the self-supervised model DINOv3 provides strong structured visual representations, its lack of native textual alignment hinders its direct application…