openalexThe international archives of the photogrammetry, remote sensing and spatial information sciences/International archives of the photogrammetry, remote sensing and spatial information sciences2026-07-23Cited by 0
The Emerging Role of Vision-Language Models in the Automation of Railway Asset Management: A Review and Future Perspective
Abstract. The safety, efficiency, and longevity of global railway networks are directly linked to the rigorous inspection and management of their vast inventory of physical assets. Over the past decade, the field has progressed from manual surveys to automated systems leveraging imagery from track-based or aerial platforms. These systems predominantly built on traditional Computer Vision (CV) models have proven effective at detecting a pre-defined set of common assets. However, this progress has exposed a fundamental architectural and operational ceiling: the closed-world assumption. Current models are constrained to a fixed catalogue of classes defined during their training. It makes the model incapable of identifying novel objects or adapting to environmental changes without costly and continuous cycles of data re-annotation, retraining, and redeployment. This review paper argues that Vision-Language Models (VLMs), a paradigm whose rapid maturation is evidenced by recent comprehensive surveys offer a transformative solution. We provide a focused overview of the limitations of current CV systems and map the mechanics of a VLM-powered approach specifically Open-Vocabulary Detection and Reasoning Segmentation directly to the outstanding challenges in rail asset management. Ultimately, the literature suggests that the adoption of VLMs could catalyze a fundamental shift in railway infrastructure management that serves as a key enabler for next-generation Predictive Maintenance and autonomous Digital Twins.
openalexThe international archives of the photogrammetry, remote sensing and spatial information sciences/International archives of the photogrammetry, remote sensing and spatial information sciences2026-07-23
Abstract. Open-vocabulary 3D segmentation offers an attractive alternative to closed-set scene parsing, yet directly transferring 2D vision-language models to outdoor point clouds remains difficult because projection disrupts geometric continuity and sparse sampling weakens mask…
openalexThe international archives of the photogrammetry, remote sensing and spatial information sciences/International archives of the photogrammetry, remote sensing and spatial information sciences2026-07-23
Abstract. We demonstrate an end-to-end pipeline for 3D scene understanding which integrates unsupervised graph-based point cloud segmentation with LLM-enabled spatial reasoning and editing. A point cloud is segmented into a SemanticPatch decomposition (stage 1), labeled using a z…
openalexThe international archives of the photogrammetry, remote sensing and spatial information sciences/International archives of the photogrammetry, remote sensing and spatial information sciences2026-07-23
Abstract. This paper presents an automated framework for generating semantically labelled building point clouds from their corresponding BIM models. The proposed methodology aims to facilitate the creation of training datasets for deep learning–based indoor semantic segmentation.…
openalexThe international archives of the photogrammetry, remote sensing and spatial information sciences/International archives of the photogrammetry, remote sensing and spatial information sciences2026-07-23
Abstract. Vision-based object detection is a key component of autonomous driving perception systems; however, models pretrained on large-scale generic datasets usually struggles when implemented in automotive environments due to domain shift. This research introduces a comprehens…
openalexThe international archives of the photogrammetry, remote sensing and spatial information sciences/International archives of the photogrammetry, remote sensing and spatial information sciences2026-07-23
Abstract. The rapid urbanization and rising traffic volumes strain transportation infrastructure, demanding efficient road design auditing and asset management. Conventional manual surveys are labor-intensive and lack holistic three-dimensional context. This research presents an…
openalexThe international archives of the photogrammetry, remote sensing and spatial information sciences/International archives of the photogrammetry, remote sensing and spatial information sciences2026-07-23
Abstract. The digitisation of cultural heritage objects is an important procedure to conserve, share and analyse artefacts from the past. Nowadays, it is common practice to digitise artefacts using DSLR cameras and Structure from Motion. For most objects, this is a suitable proce…