CORTEXA
← Browse

Heesang Han

1 paper indexed

arxivcs.CVcs.AI2026-07-21

D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models

Heesang Han, A. Lynn Abbott, Abhijit Sarkar

Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs for autonomous driving. However, the main emphasis to date has been for MLLMs using 2D images and videos. In contrast, this paper considers MLLM effectiveness using 3D s…

View free PDFSource page