CORTEXA
← Browse
arxivcs.DCcs.LGeess.SP2026-07-13

Decentralized Gradient Descent: Bottleneck Regimes and Budget Complexity

Nicolò Michelusi

Decentralized gradient descent (DGD) is widely used for solving distributed optimization problems over networks of agents. While its convergence properties are well understood, less is known about the communication and computation resources required to attain a prescribed accuracy. In this paper, we study DGD from a resource-aware perspective and characterize the communication-computation budget required to attain a target error level. We develop a bottleneck-centric framework in which different factors dominate the optimization dynamics at different error scales. Specifically, we identify operating regimes governed by initialization, objective heterogeneity and network connectivity, gradient noise, and communication noise. To capture these effects, we introduce two fundamental quantities: the gradient-Diversity-to-Network-connectivity Ratio (DNR) and the Gradient-to-Communication-noise Ratio (GCR). We show that these quantities determine the sequence of bottlenecks encountered during optimization and the corresponding budget-optimal operating strategy. Using a multi-stage analysis, we derive optimal stepsize selections and explicit budget-complexity bounds that quantify the budget resources required to attain a prescribed accuracy. The resulting expressions reveal how the overall budget decomposes into contributions associated with successive bottlenecks and provide insight into the fundamental tradeoffs among objective heterogeneity, network connectivity, gradient noise, and communication noise.

View free PDFSource page

Related papers

arxivcs.ITcs.DCcs.LGeess.SP2026-07-14

Mixed-Timescale Differential Coding for Downlink Model Broadcast in Wireless Federated Learning

Chung-Hsuan Hu, Zheng Chen, Erik G. Larsson

In standard federated learning systems, the parameter server broadcasts the global model to the participating devices in every iteration. Motivated by the temporal correlation between consecutive global models, differential coding can be applied to global model dissemination to r…

View free PDFSource page
arxivcs.DCcs.LG2026-07-08

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining

Jieying Wang, Shuyuan Fan, Mingkai Zheng, Zhao Zhang

Gradient communication is a primary scaling bottleneck in large language model (LLM) pretraining. Communicating gradients in low-precision formats, such as FP8 and NVFP4, can significantly reduce the communication volume. Existing methods quantize gradients via linear or nonlinea…

View free PDFSource page
arxivcs.LGcs.AIcs.DCcs.MA2026-06-26

Towards Value-Constrained Credit Assignment in Fully Delegated AI Cooperatives

Young Yoon, Jimin Kim, Soyeon Park

We propose a framework for reward allocation in fully delegated AI cooperatives where humans are represented by agents that contribute data and participate in model updates under heterogeneous value constraints. The key idea is to credit only those updates that remain admissible…

View free PDFSource page
arxivcs.LGcs.DC2026-07-16

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget

Changhai Zhou, Kieran Liu, Yuhua Zhou, Qian Qiao, Jun Gao, Harry Zhang, et al.

A widening gap separates million-token inference from RL post-training, which remains at 256K tokens or below. The gap matters for AI agents, whose observations, tool outputs, documents, and decisions accumulate over long trajectories. Unlike inference, GRPO scores and backpropag…

View free PDFSource page
arxiveess.SPcs.LG2026-07-18

Hierarchical Wireless Foundation Model for Multi-Task Optimization

Yangjing Wang, Ouya Wang, Shenglong Zhou, Geoffrey Ye Li

The increasing complexity of next-generation wireless networks has driven the integration of artificial intelligence (AI) into wireless communications. However, most existing studies focus on developing task-specific deep learning techniques for single scenarios, which limits the…

View free PDFSource page