CORTEXA
← Browse
arxivcs.SDcs.LGeess.AS2026-07-17

Fretiq: Browser-Native Electric Guitar String Classification via Engineered Spectral Features and Held-Out Free-Play Evaluation

Aadi Garg

Identifying which string produces a given pitch in monophonic electric guitar audio is a fundamental classification challenge: a single pitch can often be produced on multiple strings at different fret positions, with timbral differences that prior listening studies confirm are largely imperceptible to untrained humans. Existing approaches using support vector machines and spectral envelope features have achieved F-measures of 0.90 for six-string electric guitar classification, while String-Inverse Frequency features in earlier work achieved F1 scores up to 0.72. We present Fretiq, a preliminary single-instrument, single-player browser-based string classification system built on a 26-dimensional feature representation integrating frequency band energies, spectral statistics, and 13 Mel-Frequency Cepstral Coefficients, achieving 97.1% shuffled frame-level validation accuracy across 322,215 balanced frames. An ablation study identifies MFCCs as the primary accuracy driver (92.2% to 97.1%). We additionally introduce Comparison Training -- a data collection methodology in which adjacent open-string and fifth-fret string pairs are recorded in deliberate alternation -- and evaluate its contribution via confusion matrix analysis. Comparison Training reduces the D3 to A2 frame-level confusion rate by 44% but shows mixed results on other targeted pairs. A held-out free-play evaluation on 103,000 frames yields 87.8% overall accuracy. We describe the feature extraction pipeline in both Python and TypeScript to guarantee training-inference parity, and document two critical implementation failure modes. The system runs entirely in-browser with no hexaphonic pickup, fretboard sensor, camera, or multi-microphone setup required.

View free PDFSource page

Related papers

arxiveess.AScs.LGcs.SD2026-07-03

Deriving Benchmarking Datasets from Long-Form Recordings: Challenges and Opportunities

Kaveri K. Sheth, Lawrence Borst, Tarek Kunze, Marvin Lavechin, Okko Räsänen, Sho Tsuji, et al.

Long-form recordings (LFRs) of child-centered audio are ecologically valid sources for studying early language development, but three problems limit their use. First, LFR corpora are collected across sites with heterogeneous formats and consent structures, making cross-corpus use…

View free PDFSource page
arxiveess.AScs.AIcs.CLcs.LGcs.SD2026-07-10

Phone Segmentation and Recognition through Phonological Activation Mapping

Shikhar Bharadwaj, Kwanghee Choi, Stephen McIntosh, Chin-Jou Li, Eunjung Yeo, Daisuke Saito, et al.

Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve bot…

View free PDFSource page
arxivcs.SDcs.LGeess.ASq-bio.QM2026-07-03

Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types

Paria Vali Zadeh, Sven Tomforde

Reliable analysis of bird vocalisations in passive acoustic monitoring requires models handling multiple, imbalanced annotation targets. We extend BirdCallNet for joint species and call-type classification on the long-tailed WiWa dataset and investigate how task-loss balancing in…

View free PDFSource page
arxiveess.AScs.LGcs.SDeess.SP2026-07-01

CNN Models for Microphone Array Covariance Matrix Upsampling and Acoustic Imaging

Marianthi Adamopoulou, Parthasaarathy Sudarsanam, David Diaz-Guerra, Meng Jiang, Archontis Politis, Seyed Jalaleddin Mousavirad, et al.

Acoustic imaging visualization is a core methodology in acoustics, enabling spatial analysis of sound sources and acoustic scenes. However, limited sensor availability in practical systems motivate approaches that enhance spatial resolution without increasing the hardware complex…

View free PDFSource page
arxiveess.AScs.LGcs.SD2026-07-05

Weakly Guided and Autoregressive Beamformer Parameterization for Generalizable Moving Speaker Extraction in Higher-Order Ambisonics

Jakob Kienegger, Tal Peer, Sina Khanagha, Timo Gerkmann

Linear spatial filters (beamformers) enable robust, generalizable and interpretable speech enhancement with performance guarantees under ideal parameterization. Modern beamformers are often parameterized by deep neural networks, whose performance degrades in dynamic scenarios wit…

View free PDFSource page
arxiveess.AScs.AIcs.CLcs.LGcs.SD2026-07-15

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

Chun-Yi Kuan, Siwon Kim, Byeonggeun Kim, Suyoun Kim, Bo-Ru Lu, Qinming Tang, et al.

Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly emphasize global similarity or perceptual quality, with limit…

View free PDFSource page