arxivcs.CV2026-07-02
ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning
Xuanhua He, Jiaxin Xie, Mingzhe Zheng, Qifeng Chen
Monocular video depth estimation requires temporal consistency, geometric accuracy, and generalization across diverse scenarios, yet existing methods struggle to achieve all three simultaneously. Discriminative models excel at per-frame accuracy but suffer from temporal drift due…