CORTEXA
← Browse
openalexOpen MINDCited by 0

Combating Data and Target Shifts in Visual Tasks

Zhaorui Tan

The rapid advancement of deep learning models for visual tasks has led to significant progress in many domains. However, a key challenge remains: ensuring that models can generalize effectively to unseen samples or novel classes, especially in real-world scenarios where training data-target pairs are often limited or unavailable, and test data may exhibit shifts in both data and targets. This thesis addresses the problem of shifts across data and targets, proposing novel methods to improve model generalization and transferability in complex settings. The thesis identifies the different shift scenarios for visual tasks and presents their problem forms. The research first relaxes the static target distribution assumption in Multi-Domain Generalization (mDG) tasks, introducing the General Multi-Domain Generalization (GMDG) objective to improve generalization under varying data shifts. Specifically, Extensive experiments validate that GMDG is feasible for classification, segmentation, and regression tasks. Further considering extreme target distribution shifts where even unknown novel classes emerge, the study also explores logical regularization techniques for tasks such as Generalized Category Discovery (GCD), mDG+GCD, and Class Incremental Learning (CIL), proposing a sample-based logical regularization term (L-Reg) to enhance data-, target-, and all-shift generalization. Additionally, a partial logic framework is introduced to address challenges in CIL, enabling models to retain room for unknown classes during training. This is further extended by the Partial-Logic Regularization (PL-Reg), which improves generalization across GCD, mDG+GCD, and transferability for CIL tasks. Furthermore, the thesis proposes a novel Semantic-aware Data Augmentation (SADA) framework for cross-modality generalization in text-to-image synthesis (T2Isyn), text-image retrieval (TIR), and image-text retrieval (ITR) tasks. This framework ensures semantic preservation across text and image modalities, improving both semantic consistency and image quality in cross-modality tasks. The methods proposed in this thesis offer significant advancements in visual generalization and transfer scenarios, providing theoretical insights and practical solutions for handling data shifts, target shifts, and multi-modalities in complex scenarios. Extensive experiments validate the effectiveness of the proposed methods, demonstrating their superior performance across a range of tasks.

Also available via: Open MIND

View free PDFSource page

Related papers

openalexOpen MIND

SOAR: Smooth Online Activation Routing for Stable Neural Learning from Evolving Streams

Sizhen Niu

Online neural learning requires models that update after each incoming example, remain calibrated under distributional change, and avoid brittle gradient transmission. The original version of this work used a small static benchmark, a shallow model, few random seeds, and no signi…

Also available via: Open MIND

View free PDFSource page
openalexOpen MIND

From Sparse to Dense: Label-Efficient Weakly Supervised Segmentation for Images and Videos

J. Wang

Obtaining high-quality annotated data has become a primary bottleneck for training deep learning models, particularly for dense prediction tasks like semantic segmentation and video salient object segmentation. The demand for meticulous, pixel-level labeling makes fully-supervise…

Also available via: Open MIND

View free PDFSource page
openalexOpen MIND

Pilot Assignment and Channel Estimation for User-Centric Cell-Free Massive MIMO Systems

Bowen Zhong

User-centric (UC) cell-free (CF) massive multiple-input multiple-output (MIMO) systems have emerged as a promising solution for the beyond fifth generation (B5G) and the sixth generation (6G) wireless communication systems, providing enhanced coverage, capacity, and user fairness…

Also available via: Open MIND

View free PDFSource page
openalexOpen MIND

Mechs

francis lee

You can treat this as buildable with today’s tech, but you’re in “prototype MBT + experimental railgun + biped robot” cost territory for a single unit. Below is an order‑of‑magnitude cost breakdown for the first full prototype and a manufacturing/integration roadmap assuming your…

Source page