Autoregressive video diffusion models have enabled the generation of arbitrarily long videos by removing conditioning on future frames, thus greatly improving computational efficiency. Yet, they suffer from error accumulation over time, as the denoised sequence gradually drifts a…
Image copy detection is commonly addressed using either local descriptors or deep learning models, which can be computationally expensive and rely on high-dimensional features. In contrast, this work explores copy detection using compact perceptual hash representations and learne…