CORTEXA
← Browse
arxivcs.AR2026-07-27

A Heterogeneous Neural Network Accelerator for End-to-End Multitask RF Signal Recognition

Zhifan Song, Haralampos-G. Stratigopoulos, Hassan Aboushady

This paper presents a heterogeneous neural network accelerator for multi-task RF signal recognition, supporting automatic modulation recognition (AMR), hardware-Trojan covert channel (HT-CC) detection, and GNSS jamming classification. We introduce a compact attention-enhanced convolutional neural network (CNN) combined with LSDec, a learnable streaming decimator that enables adaptive temporal downsampling and flexible input lengths. The hardware architecture integrates a novel dual-pipeline, fused convolution-pooling engine with DMA-based streaming to minimize memory traffic and latency. Co-execution scheduling on the accelerator and SIMD-optimized CPU kernels reduces hardware resource usage while preserving high performance and task-level flexibility. Across three datasets, the proposed system achieves $\geq$ 99% average accuracy above 4 dB Signal-to-Noise Ratios (SNRs) on the RadioML2018 dataset for AMR, 90% on the HT-CC dataset, and 99.5% on the GNSS-Jamming dataset. The accelerator sustains an end-to-end inference latency of 98 $μ$s per frame, demonstrating its effectiveness for low-power, latency-critical multi-task spectrum-intelligence applications on embedded and edge devices.

View free PDFSource page

Related papers

arxivcs.AR2026-07-27

DICE: Detailed Inter-Chiplet End-to-End PHY Modeling for Accurate Chiplet Simulation

Rashid Aligholipour, Stefanos kaxiras, Yuan Yao

Scaling monolithic multicores is increasingly constrained by power/thermal limits, yield, and rising manufacturing and testing costs. Chiplet designs address these challenges by partitioning large dies into smaller parts (typically multiple core-complex dies and an I/O die) linke…

View free PDFSource page
arxivcs.LGcs.ARcs.OS2026-07-16

PolyQ: Codesigning End-to-End Quantization Framework for Scalable Edge CPU LLM Inference

Hyunwoo Oh, Suyeon Jang, Hanning Chen, KyungIn Nam, Sanggeon Yun, Ryozo Masukawa, et al.

CPUs are the most universal target for on-device LLM inference, but existing low-bit quantization methods offer either coarse operating points or fine-grained mixed precision that is difficult to execute efficiently on CPUs. We present PolyQ, a CPU-oriented compiler/quantization…

View free PDFSource page
arxivcs.LGcs.ARmath.NA2026-07-17

CoG-Guided Weight Correction for Fault-Tolerant Deep Neural Networks

Bahram Parchekani, Samira Nazari, Ali Azarpeyvand, Mohammad Hasan Ahmadilivani, Tara Ghasempouri, Jaan Raik

Deep Neural Networks (DNNs) used in safety-critical applications are vulnerable to hardware and memory faults that corrupt network weights and degrade reliability. In this paper, we propose a Center of Gravity (CoG) guided weight correction method that restores faulty weights bas…

View free PDFSource page
arxivcs.ARcs.LG2026-07-09

FPGN: Redefining Ultra-Fast Programmable Gate-based Neural Acceleration with Differentiable LUTs

Jiawei Liang, Haotong Qin, Linfeng Du, Xingyu Liu, Shangkun Li, Hui Yu, et al.

Achieving nanosecond-scale inference latency for deep neural networks (DNNs) has become a primary architectural concern for latency-critical applications. While Field-Programmable Gate Arrays (FPGAs) offer a promising substrate for low-latency inference, conventional FPGA acceler…

View free PDFSource page
arxivcs.ARcs.NE2026-07-14

A 32-channel event-based bio-signal analog front-end with adaptive delta and pulse frequency encoding

Narayanan Shyam, Saptarshi Ghosh, Giacomo Indiveri

Low-power event-based Analog Front-Ends (AFEs) are essential for building efficient, end-to-end neuromorphic signal processing systems. In this paper, we present an event-based AFE Application-Specific Integrated Circuit (ASIC) optimized for biomedical signal acquisition and enco…

View free PDFSource page
arxivcs.ETcs.AIcs.ARcs.DCcs.LG2026-07-06

Optimizing ML Workload Partitioning between CPUs and CIM Accelerators for Heterogeneous Computing

Joel Klein, Rebecca Pelke, Roberto Laudani, Jan Moritz Joseph, Rainer Leupers

Computing-in-Memory (CIM) accelerators execute Matrix-Vector Multiplications (MVMs) in memory, making them a compelling solution for Machine Learning (ML) workloads. However, existing ML workload partitioning approaches for CIM accelerators do not fully account for Resistive Rand…

View free PDFSource page