Skip to main content
archive
Search Submit Donate Log in
Press Enter to search · Advanced search

Audio and Speech Processing

  • New submissions
  • Cross-lists
  • Replacements

See recent articles

Showing new listings for Thursday, 1 October 2026

Total of 22 entries
Showing up to 2000 entries per page: fewer | more | all

New submissions (showing 11 of 11 entries)

[1] arXiv:2609.38440 [pdf, html, other]
Title: Monotonicity-Guided Semantic Alignment for Zero-shot Multispeaker Image-to-Speech Synthesis
Lijun Wang, Yixian Lu, Shogo Okada
Comments: 5-pages, 1 figures
Subjects: Audio and Speech Processing (eess.AS)

Direct image-to-speech (Img2Sp) poses an alignment challenge in mapping visual content to ordered speech sequences, as images permit multiple spoken descriptions and lack monotonic correspondence with speech sequences. We propose Monotonicity-Guided Semantic Alignment (MGSA), to the best of our knowledge, the first framework for zero-shot multispeaker Img2Sp synthesis. We use semantic speech units to provide shared content targets across speakers with reference speech for speaker conditioning. A query aligner maps semantic memory learned from visual content to speech unit positions via a soft monotonic prior, which yields position-specific conditioning states. A blockwise masked diffusion generator is employed for the speech unit generation conditioning on these states. Experiments on Flickr8k-Audio show competitive captioning performance against single-speaker baselines, while evaluation with LibriTTS-R references supports zero-shot synthesis for unseen speakers. Ablations validate the effectiveness of aligner and block diffusion. Audio samples are available at this https URL.

[2] arXiv:2609.38501 [pdf, html, other]
Title: Voices as Handles: Reasoning about Speaker Identity with Frozen Text LLMs
Runqiu Xu, Zhisheng Zheng, David Harwath
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)

Multi-user voice agents must track who said what across dialogue sessions. Text LLMs are attractive backbones for such agents, but transcripts alone do not expose acoustic speaker identity, leaving the model without a persistent reference for linking information to speakers across sessions. We address this gap by introducing Speaker Handles, soft-token representations that expose acoustic speaker identity to a frozen text LLM for cross-session speaker-dependent reasoning. A three-stage curriculum trains a lightweight projector, with fewer than 0.1% of the backbone's parameters, to map speaker embeddings into these handles. Establishing whether the resulting handles truly support cross-session speaker-dependent reasoning is challenging with existing benchmarks because textual cues can partially reveal fact ownership. We therefore present SpeakerBind, a controlled shared-agent benchmark in which overlapping facts across users require correct cross-session speaker attribution. Speaker Handles achieve 97.40-98.36% accuracy on VoxCeleb1 and 70.40% on SpeakerBind, close to the 71.88% topline. These results show that the proposed Speaker Handles provide an efficient way to integrate acoustic speaker identity into frozen text LLMs for speaker-content reasoning.

[3] arXiv:2609.38658 [pdf, html, other]
Title: Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Comments: Under Review
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)

TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning, such as requiring transcripts of the reference speech during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. Our model replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation. Across two English and two Mandarin datasets, Tacit-TTS achieves competitive zero-shot quality while generating speech over 10x faster than IndexTTS2 for utterances longer than 5 seconds. Its transcript-free conditioning further supports cross-lingual and non-lexical references. We validate this capability using references from eight other languages, infant babble, and synthetic gibberish, where transcript-dependent systems often degrade or fail due to unreliable ASR transcripts.

[4] arXiv:2609.38794 [pdf, other]
Title: A barrier or a booster? Familiarity effects on Mandarin emotion prosody recognition using AI-powered voice cloning
Feng Xu, Gaoyuan Zhang, Shanshan Xue, Yixiang Chen, Hanrui Zhou, Xurong Xie, Hui Chen
Comments: Accepted by Interspeech 2026
Subjects: Audio and Speech Processing (eess.AS); Human-Computer Interaction (cs.HC); Sound (cs.SD)

Emotion prosody perception requires simultaneous processing of acoustic cues and speaker identity. While listeners effortlessly decode natural speech, AI synthetic voices introduce cognitive complexities due to subtle acoustic atypicalities. It remains unclear how these synthetic features interact with a listener's prior social knowledge and memory of a familiar speaker. This study investigated how speech sources (human vs. AI) and speaker familiarity affect emotion recognition accuracy and cognitive load. A within-subject task with Mandarin-speaking adults evaluated behavioral (accuracy, reaction time) and physiological data (heart rate variability). Results showed that human voices yielded significantly higher accuracy and faster processing times than AI voices, while HRV did not significantly differentiate between conditions. These findings show that decoding synthetic speech is gated by top-down social cognition, highlighting limitations in current AI synthesis technologies.

[5] arXiv:2609.38887 [pdf, html, other]
Title: VOSSA: Voiceprint Optimization for Streaming Speech Architectures
Mu-Ruei Tseng, Waris Quamer, Ghady Nasrallah, Ricardo Gutierrez-Osuna
Comments: Published in Proceedings of Interspeech 2026
Journal-ref: Tseng, M.-R., Quamer, W., Nasrallah, G., Gutierrez-Osuna, R. (2026) VOSSA: Voiceprint Optimization for Streaming Speech Architectures
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Machine Learning (cs.LG)

Real-time voice conversion (VC) systems commonly rely on pretrained speaker embeddings from automatic speaker verification (ASV) models. While effective for speaker discrimination, these embeddings are trained to remain stable across phonetic and prosodic variations within-speaker, which may conflict with frame-level acoustic generation in streaming constraints. To address this issue, we propose VOSSA (Voiceprint Optimization for Streaming Speech Architectures), a speaker representation framework that extracts speaker information from intermediate content encoder layers and aggregates using attentive statistics pooling. The embedding is trained jointly with VC objectives, removing the need for a separate speaker encoder. Across six datasets, VOSSA improves F0 dynamics and vowel-discriminative acoustic cues while maintaining comparable NISQA-MOS, WER, and speaker similarity. Perceptual tests further indicate improvements in naturalness, speaker similarity, intelligibility, and vibrancy.

[6] arXiv:2609.39028 [pdf, html, other]
Title: Improving Predicted MOS Scores, Not Perceived Quality: Multi-Predictor Test-Time Optimization of Enhanced Speech
Tsubasa Ochiai, Marc Delcroix, Nahomi Kusunoki, Rintaro Ikeshita, Naohiro Tawara, Naoyuki Kamo, Tetsuji Ogawa, Shoko Araki
Comments: 5 pages, 2 tables
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)

Non-intrusive MOS predictors are widely used instead of subjective listening tests to evaluate and rank speech enhancement (SE) systems. If they accurately reflect perceived quality, raising their scores should lead to higher-quality speech. We present the first comprehensive analysis of test-time optimization for the SE task, which directly modifies the enhanced signal to raise the average of multiple MOS predictor scores. On seven systems from the URGENT 2026 challenge, we find that 1)~all the optimized predicted scores increase while reference-based metrics remain nearly unchanged, 2)~a non-optimized predicted score does not increase, and 3)~a MUSHRA listening test shows no improvement in perceived quality. These findings reveal a risk that such optimization can distort evaluations, e.g., biasing comparisons of SE systems regardless of their perceived quality. We believe these findings can inform future evaluation practices: they suggest that predictors used for optimization should not be used for evaluation, and that challenges should keep the predictors used for ranking undisclosed.

[7] arXiv:2609.39030 [pdf, html, other]
Title: SURE-EVAL: A Systematic and Unified Agentic Framework for Reproducible Evaluation
Jing Peng, Junhao Du, Yixuan Wang, Bowen Wang, Hanqi Li, Chaolei Liu, Weihan Chen, Haohui Xie, Ruichen Sun, Chenghao Wang, Wen Wen, Guanyu Chen, Xiaoyu Gu, Haoyu Li, Yiwei Guo, Bohan Li, Tao Liu, Yucheng Wang, Yu Xi, Yihua Zhou, Qiang Zhou, Feng Lu, Shuai Wang, Kai Yu
Comments: This paper is accecpted to NCMMSC2026 as an oral presentation
Subjects: Audio and Speech Processing (eess.AS)

Audio and speech models are released rapidly, but reported scores often conflate model capability with deployment and evaluation choices. The same checkpoint can produce different predictions under different runtimes, hardware, decoding settings, or fallback policies. Even fixed predictions can receive different scores under different normalization and metric implementations. Existing speech benchmarks standardize selected datasets or scoring procedures, but rarely connect heterogeneous model onboarding, controlled inference, and versioned scoring in one executable workflow. We introduce SURE-EVAL, a Systematic and Unified Agentic framework for Reproducible Evaluation of audio and speech systems. A Tool Agent Workflow converts model releases into isolated, verified callable tools. A Main Agent Workflow commits task-specific inference and scoring protocols, then delegates all score-bearing operations to versioned deterministic programs. Each result retains its runtime, protocol, pipeline nodes, predictions, and audit artifacts. Across 18 public releases covering automatic speech recognition, text-to-speech, voice conversion, speaker diarization, speaker-attributed recognition, and multi-task audio understanding, a Codex-only baseline completes 12 models in one shot, while the same agent with the SURE-EVAL Tool Agent Workflow completes all 18. We also conduct unified evaluations over seven ASR test conditions and two TTS subsets. A protocol analysis of three TTS systems finds absolute differences of 0.02-0.52 points between paper-reported and unified results, with the direction varying by model and language. These results show that reproducible evaluation requires controlling both model execution and output scoring.

[8] arXiv:2609.39216 [pdf, html, other]
Title: PG-SELD: Physics-Guided Sound Event Localization and Detection
Elad Cohen, Elad Dror Cohen, Arnon Netzer, Hai Victor Habi
Subjects: Audio and Speech Processing (eess.AS)

Sound event localization and detection (SELD) aims to jointly recognize sound events and estimate their directions of arrival from multichannel audio. Although recent deep learning approaches have achieved strong performance, their ability to generalize across acoustic environments remains limited, as room reverberation introduces environment-specific characteristics into the learned representations. In this work, we address this challenge by leveraging a physical free-field model as a room-independent reference. Specifically, we propose PG-SELD, a training framework that combines free-field with physics-guided knowledge distillation. Our approach aligns intermediate representations extracted from reverberant signals with those produced by a free-field teacher for matched acoustic scenes. This guidance encourages the model to preserve event- and localization-relevant information while reducing sensitivity to room-specific characteristics. Experimental results on the STARSS23 benchmark show that PG-SELD consistently improves the generalization performance of multiple baseline SELD architectures.

[9] arXiv:2609.39237 [pdf, html, other]
Title: DuSpaR: Dual-State Sparsifying Recurrent Unit with Feedback Modulation for Compute-Efficient Speech Processing
Zixiao Li, Sheng Zhou, Longbiao Cheng, Shih-Chii Liu
Subjects: Audio and Speech Processing (eess.AS)

We introduce the Dual-state Sparsifying Recurrent Unit (DuSpaR) as a computationally efficient building block for speech processing models on resource-constrained edge devices. It employs dual-state recurrence to modulate its input vectors in a stateful feedback loop. Its recurrent cells sparsify the input vector operand involved in matrix-vector multiplication using ReLU activation. By skipping the zero entries dynamically, inference-time savings in multiply-accumulate operations and weight memory fetches can be achieved. We evaluate DuSpaR on three speech tasks: keyword spotting (KWS) on the Google Speech Commands dataset, spoken language understanding (SLU) on the Fluent Speech Commands dataset, and speech enhancement (SE) on the Voice Bank + Demand (VBD) dataset. At similar parameter counts, DuSpaR requires 51.0% and 68.4% less computation than Gated Recurrent Unit (GRU) on KWS and SLU, respectively, while achieving higher accuracy, and 50.1% less computation on SE while maintaining similar quality. At comparable computational cost and across a range of model sizes, DuSpaR also achieves higher KWS/SLU accuracy and better SE quality than other sparsity-aware recurrent models. Ablation studies show that compared with the single-state recurrence baseline, dual-state recurrence reduces the effective compute by factors of 3.1 to 11.3 at similar task performance.

[10] arXiv:2609.39732 [pdf, html, other]
Title: Bin2Ambi: Learning Ambisonic Soundfield Reconstruction from Head-Tracked Binaural Audio
Gavin Milner, Nils Peters
Subjects: Audio and Speech Processing (eess.AS)

User-generated content has become one of the most-consumed content types. However, capturing spatial audio with consumer hardware is still challenging. Given the widespread success of smart earbuds, binaural audio could be a promising option to capture spatial audio on consumer devices, but its inherent signal characteristic limits its usability as a recording format. In this paper, we propose and define a new task, Binaural to Ambisonics conversion (Bin2Ambi). In our proposed system, we exploit simultaneously captured head-tracking data provided from the motion sensors in smart earbuds. We show that this motion data help resolve the inherent directional uncertainty of two-channel binaural audio due to front-back localization ambiguities and lateral errors in the cone of confusion. Our results show that our system learns directional and diffuse-field information and that head-tracking especially reduces extreme localization errors. Objective metrics and a subjective listening test suggest that the converted Ambisonics soundfield achieves an average directional error of up to $11.8^\circ$ and a perceived spatial quality similar to a DirAC ground-truth model. The proposed algorithm can serve as a baseline for future improvements to this novel Bin2Ambi task.

[11] arXiv:2609.39852 [pdf, html, other]
Title: Pitch Smoothing Using Relative Interval Networks
Chin-Yun Yu, Chi-Jen Peng, Li Su, György Fazekas
Comments: Submitted to ICASSP 2027
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)

Pitch tracking systems typically couple a per-frame fundamental frequency ($F_0$) estimator with a temporal smoothing stage to obtain continuous trajectories. Conventional Viterbi smoothers enforce first-order continuity but lack long-term temporal awareness and could lock into octave errors across corrupted frames. We propose Relative Interval Networks (RIN), a trajectory smoothing framework that reconciles per-frame pitch estimates with data-driven multi-hop pitch differences. We extract robust relative pitch intervals across arbitrary frame offsets using Variable-Q Transform cross-correlation. We formulate pitch smoothing as an $L_1$-norm optimization problem and prove its equivalence to a minimum cost circulation problem, solved efficiently via linear programming. Evaluations across speech, singing, and instrumental datasets show that RIN substantially improves weak estimators, matches or outperforms Viterbi decoding at a comparable computational cost, and provides superior robustness under certain acoustic degradation.

Cross submissions (showing 6 of 6 entries)

[12] arXiv:2609.38203 (cross-list from cs.CL) [pdf, html, other]
Title: Automatic estimation of verbal fluency index in people with Motor Neuron Disease using ASR alignment and pause modelling
Bahman Mirheidari, Leslie Ing, Daniel Blackburn, Sharon Abrahams, Christopher McDermott, Heidi Christensen
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)

Monitoring cognitive impairment (CI) in motor neuron disease (MND) is essential for timely treatment and care, yet challenging due to co-occurring speech difficulties. The Edinburgh Cognitive and Behavioural ALS Screen (ECAS) provides a robust metric for CI assessment, with the Verbal Fluency Index (VFI) a central element. Building on recent advances in automated speech analysis, this study proposes a system for estimating VFI. It leverages a unique MND dataset and combines ASR (WhisperX) and VAD (Silero) with refined timestamping to predict the VFI and extract several clinically interpretable measures. Our approach outperformed systems based on traditional acoustic features and self-supervised embeddings, evaluated using multiple regression algorithms. Clinically inspired features consistently outperformed the other sets, with the best models achieving strong results (P-words: R2 0.9, NRMSE 0.05; S-words: R2 0.8, NRMSE 0.08), demonstrating the feasibility of automated VFI estimation.

[13] arXiv:2609.38232 (cross-list from cs.SD) [pdf, html, other]
Title: When Does a Spoken Agent Have Enough Evidence to Act? The PACT-SLM Contract Test
Mengzhe Geng
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)

Streaming spoken agents may take an external action before the available speech supports it, yet final-turn scores do not reveal whether each observed prefix supports that action. We introduce the Partial Speech Action Contract for Turn Taking in Speech Language Models (PACT-SLM), a controlled evaluation that assigns a first valid action time and measures action identity and timing separately. The primary diagnostic contains 80 paired contrast groups from four held-out semantic families and 1,600 prefix predictions across clean and 15 dB noise renderings. After correcting a mismatch between randomized branch codes and semantic labels, a refitted WavLM Base Plus probe reaches 26.03% pooled post-onset semantic-label accuracy (95% group-bootstrap interval: 22.14%-29.68%), exposes an action on 18.99% of pre-onset prefixes, and predicts 5.94% of complete trajectories exactly. It exceeds matched text, scalar-acoustic, and shuffled-representation probes in post-onset label accuracy, but its score is at the 96th percentile of 100 within-prefix label permutations and below the 97.5th-percentile reference (26.73%). Elapsed time is more onset-exact than WavLM Base Plus (36.25% vs. 23.13%) but less accurate about action identity (9.92% vs. 26.03%). These results show that action identity and timing measure distinct aspects of partial-speech decision behavior.

[14] arXiv:2609.38867 (cross-list from cs.AI) [pdf, html, other]
Title: Talk2Agent: Benchmarking Voice Interfaces for Text Agents
Terumi Chiba, Guangzhi Sun, Zheqi Yuan, Chao Zhang
Subjects: Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)

Large language model (LLM) computer-use agents are typically evaluated with clean written instructions, despite speech being an increasingly popular interface for interacting with such systems. Speech input introduces an additional failure point: transcription errors can alter task-critical entities, constraints, or targets before the agent begins reasoning, while conventional ASR metrics do not directly measure whether the information required for successful execution has been preserved. We introduce Talk2Agent, a benchmark for evaluating how effectively voice interfaces convey human-spoken instructions to LLM-based computer-use agents. Talk2Agent builds human-spoken versions of tasks from WildClawBench and OSWorld and evaluates a range of voice interfaces, including dedicated ASR models, audio-capable LLMs, contextual biasing, and LLM-based ontology repair. Because repeatedly executing long-horizon computer-use tasks is costly and stochastic, we further propose an execution-free, task-conditioned evaluation framework that projects the original task grader onto prompt-addressable intentions and measures how much task-relevant information is retained after the voice interface. On WildClawBench, Talk2Agent's execution-free native projection provides a practical, execution-grounded measure of voice-interface quality, correlating with downstream task completion and improving Pearson correlation by 0.246 over WER/CER on 32 hours of real human speech.

[15] arXiv:2609.39032 (cross-list from cs.SD) [pdf, html, other]
Title: How Reliable Are Predicted MOS for Reproducing Human System-Level Preferences in Speech Enhancement?
Nahomi Kusunoki, Tsubasa Ochiai, Naohiro Tawara, Marc Delcroix, Naoyuki Kamo, Tetsuji Ogawa, Shoko Araki
Comments: 5 pages, 2 figures, 2 tables
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)

We investigate whether predicted Mean Opinion Scores (MOS) can reliably support system-level comparisons of speech enhancement (SE) methods by introducing system-level preference accuracy (SPA). Although MOS prediction models are widely used to evaluate SE systems, their performance is typically assessed by correlation with human-rated MOS, which does not guarantee agreement on which system is better. SPA addresses this gap by directly evaluating whether predicted and human-rated MOS yield the same system preferences. Using SPA, we systematically evaluate three settings: single prediction models, ensembling, and domain adaptation. Through experiments, SPA varies substantially across single prediction models, from 9.4% to 76.8%. Even the best model disagrees with human judgments in approximately 23% of system comparisons. Ensembling yields only limited improvement, while domain adaptation tends to substantially improve SPA in the closed condition but brings only modest gains in the more practical open condition, where neither the target systems nor the speakers are known. These results suggest that SPA can reveal errors correlation-based evaluation alone does not expose, and that predicted MOS alone can lead to unreliable conclusions in practical SE system comparison.

[16] arXiv:2609.39453 (cross-list from cs.SD) [pdf, html, other]
Title: From Speech to Editable Concepts: Probing Emotion Recognition with Concept Bottleneck Models
Hezhao Zhang, Thomas Hain
Comments: 5 pages, 2 figures. Submitted to ICASSP 2027
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)

Speech emotion recognition (SER) is the task of assigning emotion labels to utterances. Early systems relied on acoustic features, whereas recent approaches combine multiple modalities, most commonly speech and text. Still, performance remains poor on many datasets. Large language models (LLMs) have therefore attracted interest for SER, as they can process diverse inputs jointly with instructions. However, direct audio input raises questions of explainability. To address similar questions in image classification, concept bottleneck models were introduced. This work adapts concept bottlenecks to SER to examine how individual predictions depend on transcripts, acoustic descriptions and speaker attributes. Experiments test three LLMs on CREMA-D, IEMOCAP and MELD, with concepts extracted by separate tools. On scripted corpora, LLMs are strongly biased towards the transcript in the zero-shot setting, which lowers Macro-F1 from 27.8 to 5.8 on CREMA-D. Fine-tuning removes this bias, and the transcript raises Macro-F1 from 41.8 to 45.1. Removing speech rate changes 48% of Neutral predictions to Disgust on CREMA-D; removing intensity level on MELD changes predictions despite little change in Macro-F1. These findings show that aggregate performance changes alone do not capture the effects of concept removal on individual predictions.

[17] arXiv:2609.40087 (cross-list from cs.SD) [pdf, html, other]
Title: MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo
Comments: Accepted to Interspeech 2026. Project page: this https URL
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)

Flow-matching approaches to voice conversion (VC) have gained attention owing to their high speech quality and strong speaker similarity. Among them, one-step models such as MeanVoiceFlow are particularly attractive because they enable efficient inference; however, their reliance on a computationally intensive content encoder remains a bottleneck. We therefore propose MeanVoiceFlow2, a framework that jointly optimizes a flow-based conversion module and a computationally efficient content encoder. The model is trained through conversion distillation using MeanVoiceFlow and the reconstruction of real data. We further incorporate diffusion-GAN training with sample mixing and teacher-guided conditioning augmentation to enhance realism and disentanglement. Experiments on zero-shot VC showed that MeanVoiceFlow2 achieved higher perceptual quality and approximately $9\times$ faster inference than MeanVoiceFlow while maintaining comparable speaker similarity. Audio samples are available at this https URL.

Replacement submissions (showing 5 of 5 entries)

[18] arXiv:2603.22252 (replaced) [pdf, html, other]
Title: SelfTTS: cross-speaker style transfer through explicit embedding disentanglement and self-refinement using self-augmentation
Lucas H. Ueda, João G. T. Lima, Pedro R. Corrêa, Flávio O. Simões, Mário U. Neto, Paula D. P. Costa
Comments: Submitted to OJSP
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)

This work presents SelfTTS, a text-to-speech (TTS) model designed for cross-speaker style transfer that eliminates the need for external pre-trained speaker or emotion encoders. The architecture achieves emotional expressivity in neutral speakers through an explicit disentanglement strategy utilizing Gradient Reversal Layers (GRL) combined with cosine similarity loss to decouple speaker and emotion information. We introduce Multi Positive Contrastive Learning (MPCL) to induce clustered representations of speaker and emotion embeddings based on their respective labels. Furthermore, SelfTTS employs a self-refinement strategy via Self-Augmentation, exploiting the model's voice conversion capabilities to enhance the naturalness of synthesized speech. Experimental results demonstrate that SelfTTS achieves superior emotional naturalness (eMOS) and robust stability in target timbre and emotion compared to state-of-the-art baselines.

[19] arXiv:2603.23673 (replaced) [pdf, html, other]
Title: Crab: Multi Layer Contrastive Supervision to Improve Speech Emotion Recognition Under Both Acted and Natural Speech Condition
Lucas H. Ueda, João G. T. Lima, Paula D. P. Costa
Comments: Submitted to Speech Communication
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)

Speech Emotion Recognition (SER) in real-world scenarios remains challenging due to severe class imbalance and the prevalence of spontaneous, natural speech. While recent approaches leverage self-supervised learning (SSL) representations and multimodal fusion of speech and text, most existing methods apply supervision only at the final classification layer, limiting the discriminative power of intermediate representations. In this work, we propose Crab (Contrastive Representation and Multimodal Aligned Bottleneck), a bimodal Cross-Modal Transformer architecture that integrates speech representations from WavLM and textual representations from RoBERTa, together with a novel \textit{Multi Layer Contrastive Supervision} (MLCS) strategy. MLCS injects multi-positive contrastive learning signals at multiple layers of the network, encouraging emotionally discriminative representations throughout the model without introducing additional parameters at inference time. To further address data imbalance, we adopt weighted cross-entropy during training. We evaluate the proposed approach on three benchmark datasets covering different degrees of emotional naturalness: IEMOCAP, MELD, and MSP-Podcast 2.0. Experimental results demonstrate that Crab consistently outperforms strong unimodal and multimodal baselines across all datasets, with particularly large gains under naturalistic and highly imbalanced conditions. These findings highlight the effectiveness of \textit{Multi Layer Contrastive Supervision} as a general and robust strategy for SER. Official implementation can be found in this https URL.

[20] arXiv:2609.09940 (replaced) [pdf, html, other]
Title: NVV-Locator: From Transcript Tags to Acoustic Boundaries for Fine-Grained Nonverbal Vocalization Grounding
Yuang Cao, Bingshen Mu, Zhennan Lin, Guojian Li, Haoyue Zhan, Jie Liu, Chuan Xie, Qiang Zhang, Liumeng Xue, Lei Xie
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)

Human speech includes nonverbal vocalizations (NVVs), such as laughter, sighs, breaths, and coughs, which convey affective and interactional information. Existing approaches typically represent NVVs as transcript-level tags, providing limited supervision for their waveform-time boundaries. We present NVV-Locator for fine-grained NVV temporal grounding. We first unify 26 NVV categories across public resources and construct large-scale timestamp-supervised training data through dual-LLM verification, transcript-guided forced alignment, and energy-based boundary refinement. We further introduce NVV-TimeBench, an expert-refined benchmark with 667 utterances and 1,094 events. NVV-Locator uses a non-autoregressive slot-filling architecture to jointly predict lexical timestamps, NVV categories, and event boundaries. On NVV-TimeBench, it achieves 71.0% Micro F1, 70.2% Macro F1, 80.4% Macro mIoU, and 59.6 ms Macro mMAE, outperforming the evaluated large audio model counterparts. Evaluation on an external corpus further demonstrates the cross-corpus generalization of NVV-Locator.

[21] arXiv:2609.26183 (replaced) [pdf, other]
Title: A Hybrid Classical-Learning Framework for Adaptive Decision Directed Speech Enhancement
Ali Rajabi, Xiangwei Zhou
Comments: Novelty is not clear
Subjects: Audio and Speech Processing (eess.AS)

Speech enhancement aims to recover clean speech signals from noisy observations while preserving speech quality and intelligibility. Classical methods such as Spectral Subtraction and Decision-Directed (DD) enhancement remain widely used because of their interpretability and low computational complexity, but they may suffer from musical-noise artifacts or excessive attenuation of weak speech components under low signal-to-noise ratio (SNR) conditions. This paper proposes an Adaptive Beta-Constrained Decision-Directed (ABCDD) speech enhancement framework that extends the conventional DD method through a frame-dependent lower gain bound. The introduced beta parameter controls the tradeoff between noise suppression and speech preservation. To automate parameter selection for large and diverse datasets, a lightweight multilayer perceptron (MLP) model is further developed to predict frame-level beta values directly from noisy-speech features. The proposed framework is evaluated using both a representative speech example and large-scale testing on the VoiceBank-DEMAND dataset. In the representative example, ABCDD outperformed conventional Spectral Subtraction and classical DD across multiple objective metrics, including SNR, Log-Spectral Distance (LSD), Root-Mean-Square Error (RMSE), correlation, and Scale-Invariant Signal-to-Distortion Ratio (SI-SDR). On 100 unseen VoiceBank-DEMAND test files, the proposed MLP-beta ABCDD method improved average scale-aligned SNR from 9.41 dB to 13.82 dB, corresponding to an average gain of 4.41 dB. The results indicate that combining interpretable classical enhancement structure with lightweight machine-learning-based parameter adaptation provides an effective and practical direction for robust speech enhancement.

[22] arXiv:2310.11379 (replaced) [pdf, html, other]
Title: Robust Wake-Up Word Detection by Two-stage Multi-resolution Ensembles
Fernando López, Jordi Luque, Carlos Segura, Pablo Gómez
Comments: Accepted in IberSPEECH 2026
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)

Voice-based interfaces rely on a wake-up word mechanism to initiate communication with devices. However, achieving a robust, energy-efficient, and fast detection remains a challenge. This paper addresses these real production needs by enhancing data with temporal alignments and using detection based on two phases with multi-resolution. It employs two models: a lightweight on-device model for real-time processing of the audio stream and a verification model on the server-side, which is an ensemble of heterogeneous architectures that refine detection. This scheme allows the optimization of two operating points. To protect privacy, audio features are sent to the cloud instead of raw audio. The study investigated different parametric configurations for feature extraction to select one for on-device detection and another for the verification model. Furthermore, thirteen different audio classifiers were compared in terms of performance and inference time. The proposed ensemble outperforms our stronger classifier in every noise condition.

Total of 22 entries
Showing up to 2000 entries per page: fewer | more | all
We gratefully acknowledge support from our major funders, member institutions, , and all contributors.
About · Help · Contact · Subscribe · Copyright · Privacy · Accessibility · Operational Status (opens in new tab)
Major funding support from
Simons Foundation Simons Foundation International Schmidt Sciences