FedRAPT: federated representation-aligned prototypical contrastive learning
Su-Bin Seo, Junsu Kim, Changsun Shin, Hansung Lee, Chun-Bo Sim, Se-Hoon Jung
Federated learning (FL) is a distributed learning paradigm that enables collaborative model training among multiple clients without sharing raw data. FL has recently attracted attention in privacy-sensitive environments, such as mobile sensing and human activity recognition (HAR). However, in real-world environments, data heterogeneity among clients causes non-independent and identically distributed (non-IID) conditions, which induce class-level representation inconsistency and representation drift, thereby degrading the global model performance and training stability. Existing studies have mainly focused on either global alignment or local discrimination, making it difficult to simultaneously reflect class structural consistency and sample discriminability in non-IID environments. To solve this problem, we propose Federated Representation-Aligned Prototypical Contrastive Learning (FedRAPT). FedRAPT introduced an information-noise-contrastive estimation (InfoNCE)-based Cross-client Contrastive Representation Alignment (CCRA) module to jointly optimize class-level alignment and sample-level discrimination using a unified objective function. In addition, an encoder classifier decoupled structure and an exponential moving average (EMA) based prototype update strategy were designed to induce stable representation learning, even in highly heterogeneous data distributions. Experimental results on the WISDM, UCI HAR, and MotionSense datasets showed that the proposed method achieved accuracies of 96.30%, 96.13%, and 94.00% accuracy, respectively, and maintained an accuracy of approximately 94% even in a strong non-IID environment. In addition, improvements in the loss convergence and forgetting rate confirm that the proposed method effectively alleviates performance degradation and training instability caused by data heterogeneity. These results demonstrate that FedRAPT achieves robust and consistent representation learning under non-IID conditions.