arxivcs.LGcs.AI2026-06-30
Compressing What Matters: Neuron Importance Meets Data-Aware Low Rank Approximation for Language Model Compression
Athanasios Ntovas, Alexandros Doumanoglou, Petros Drakoulis, Dimitris Zarpalas
To excel at their domain large language models are comprised of billions of parameters. Yet this comes at the cost of huge memory requirements restricting their applicability in resource-constrained environments. To address the problem of neural network (NN) compression Singular…