CORTEXA
← Browse
arxiveess.ASeess.SP2026-07-04

CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling

Haolong Zheng, Yuanzhuo Hu, Xinyu Liang, Vishal Sunder, Dancheng Liu, Jinjun Xiong, Samuel Thomas, Brian Kingsbury, Zhizheng Wu, Mark A. Hasegawa-Johnson

CHILDES is a large-scale child speech corpus containing long-form recordings of naturalistic child-adult interactions, making it a valuable resource for studying child speech and language development. However, utterance-level timestamps provided in this corpus are often noisy, incomplete, or misaligned with the audio. As a result, utterances cannot always be reliably localized within long recordings, which limits the direct use of these data for training and evaluating speech models. In this work, we propose BEACON (Boundary Estimation via Alignment CONsensus), an ensemble timestamp-curation framework that refines utterance-level timestamps by aggregating knowledge from multiple off-the-shelf ASR models. Specifically, each model's word-level timestamp predictions are first aligned to provided human transcripts, and the final utterance time boundaries are determined by a consensus voting strategy. The framework is corpus-agnostic and applies to any long-form recording paired with a trusted transcript whose timestamps are unreliable or missing, offering a general recipe for timestamp curation. Leveraging this pipeline, we curate and release a 413-hour general-purpose child-speech dataset with corrected utterance-level timestamps, together with a 283-hour quality-controlled subset for ASR training. Fine-tuning on this subset yields up to an average 19.5% relative WER reduction on four out-of-domain child-speech benchmarks.

View free PDFSource page

Related papers

arxivcs.SDeess.ASeess.SP2026-06-27

Underwater Source Detection and Classification for Signal-based Surveillance: Audio Dataset Curation and Cross-Domain Evaluation

Quoc Thinh Vo, David K. Han

Machine learning for underwater acoustics is constrained by the scarcity of publicly available labeled datasets. In contrast to air-acoustic domains, where large benchmarks enable rapid model development, underwater datasets are typically small and limited in acoustic diversity,…

View free PDFSource page
arxivcs.SDcs.AIeess.ASeess.SP2026-07-10

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models

Sang-Hoon Lee, Ha-Yeong Choi

Representation alignment (REPA) has been investigated to accelerate diffusion training, but we observe that regularizing intermediate representations in diffusion Transformers (DiT) may implicitly entangle latents and limit generative capacity. To address this issue, we propose R…

View free PDFSource page
arxiveess.AScs.LGcs.SDeess.SP2026-07-01

CNN Models for Microphone Array Covariance Matrix Upsampling and Acoustic Imaging

Marianthi Adamopoulou, Parthasaarathy Sudarsanam, David Diaz-Guerra, Meng Jiang, Archontis Politis, Seyed Jalaleddin Mousavirad, et al.

Acoustic imaging visualization is a core methodology in acoustics, enabling spatial analysis of sound sources and acoustic scenes. However, limited sensor availability in practical systems motivate approaches that enhance spatial resolution without increasing the hardware complex…

View free PDFSource page
arxiveess.AScs.AIeess.SP2026-07-04

Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings

Héctor Martel, Joe Hennessy-Priest, Taemin Cho

Audio foundation models are widely adopted as general-purpose feature extractors, yet the internal structure of their learned representations remains insufficiently understood. In this work, we analyze CLAP audio embeddings through a probing framework, studying the encoding of th…

View free PDFSource page
arxiveess.ASeess.SP2026-07-04

TRACE-EVC: Text-Guided Relative Affective Control for Zero-Shot Emotional Voice Conversion

Zihan Zhang, Shreeram Suresh Chandra, Zongyang Du, Xiutian Zhao, Aurosweta Mahapatra, Hao Zhang, et al.

Traditional emotional voice conversion (EVC) conditions generation on explicit target emotions like labels or references, defining the target affective state but omitting the direction or nature of the transition. We introduce instruction-guided relative emotional voice conversio…

View free PDFSource page
arxivcs.SDcs.AIcs.LGeess.ASeess.SP2026-06-29

BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations

Ludovic K. Tuncay, Etienne Labbé, Thomas Pellegrini

Self-supervised learning enables audio representations that transfer across domains and tasks. We present BEST-RQ-2, an evolution of BEST-RQ that retains frozen randomprojection-based discrete targets while introducing a two-step contextualize-then-predict pretraining scheme. A V…

View free PDFSource page