CORTEXA
← Browse
arxivcs.CV2026-06-30

WarpI2I: Image Warping for Image-to-Image Translation

Shen Zheng, Anurag Ghosh, Gaurav Parmar, Srinivasa Narasimhan

Image-to-image (I2I) translation has achieved strong results in tasks like human relighting and driving scene translation using latent diffusion models (LDMs). However, compact LDMs often struggle to preserve fine-grained structures because the encoder compresses high-resolution inputs into a spatially downsampled latent space. To address this issue, we propose a simple saliency-guided warp-unwarp framework that reallocates spatial representation toward salient regions before encoding, enabling better preservation of structural details without increasing latent resolution. The warped image is processed by the original diffusion model and then mapped back via an inverse warp. In addition, we propose a simple and efficient outpainting-based synthetic data generation pipeline to produce high-quality paired data for image relighting. Our method is model-agnostic, requires no architectural modification, and introduces negligible computational overhead. Experiments on human relighting, driving scene relighting, and translation demonstrate improved structural preservation, lighting faithfulness, and image quality, with our framework extending naturally to video via frame-by-frame application with good temporal stability. Project Webpage: https://shenzheng2000.github.io/WarpI2I.github.io

View free PDFSource page

Related papers

arxivcs.CVcs.AI2026-07-09

TMI: Text-to-Image Meets Image-to-Image for Complementary Data Synthesis to Boost Long-Tailed Instance Segmentation

Hyeonseop Song, Seokhun Choi, Hoseok Do

Large-vocabulary instance segmentation is constrained by long-tailed category distributions and fine-grained inter-class ambiguity. While data synthesis offers a promising alternative, current paradigms have complementary limitations: text-to-image (T2I) methods inherit noisy pse…

View free PDFSource page
arxivcs.ROcs.AIcs.CV2026-07-13

Enabling 24-hour Agricultural Robotics: Unsupervised Day-to-Night Cross-Modal Image Translation for Nighttime Visual Navigation

Robel Mamo, Rajitha de Silva, Grzegorz Cielniak, Taeyeong Choi

While visual navigation has been extensively studied in agricultural robotics, most existing systems assume daytime conditions. In fact, deploying autonomous robots at night offers significant advantages, including 24-hour crop and soil monitoring, fruit harvesting, and nocturnal…

View free PDFSource page
arxivcs.CV2026-07-12

Spectral Consistent Flow for One-step 3D Medical Image Translation

Haoqing Li, Jun Shi, Mingchao Li, Zehua Zhu, Qiwei Jia, Jiong Shi, et al.

We present Spectral Consistent Flow (SC-Flow), a 3D medical image translation framework with a single function evaluation (1-NFE) in the latent space. This approach reformulates medical image translation as a stochastic Brownian bridge process that directly constructs a mapping b…

View free PDFSource page
arxivcs.CV2026-06-30

Do Not Break the Vessels: Structure-Preserving Mean Flow for Vascular Image Translation

Changjin Sun, Zhuo Hu, Kaini Wang, Baixuan Wu, Shuo Gao, Runan Zheng, et al.

Reconstructing anatomically faithful vascular structures from clinically accessible imaging modalities is of substantial clinical significance. However, existing cross-modal translation methods mainly emphasize pixel-level fidelity or visual realism and treat structure preservati…

View free PDFSource page
arxivcs.CV2026-07-02

RTE-FM-Dehazer: Radiative Transfer Equation Inspired Flow Matching for Real-World Image Dehazing

Chenfeng Wei, Chun Wang, Boyang Zhao, Si Zuo, Shenhong Wang, Chenguang Yang

Single-image dehazing aims to recover a clear scene from a hazy image and is generally formulated as an image-to-image translation task; however, it faces two limitations. Its performance depends heavily on the haze-formation priors embedded in the model. Prevailing methods adopt…

View free PDFSource page