今日论文合集:cs.SD语音17篇,eess.AS音频处理24篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 From Tens of Hours to Tens of Thousands: Scaling Back-Translation for  Speech Recognition
标题: 从数十小时到数万小时:语音识别的回缩翻译
链接:https://arxiv.org/abs/2505.16972
作者: Tianduo Wang,  Lu Xu,  Wei Lu,  Shanbo Cheng 
摘要:自动语音识别(ASR)的最新进展在很大程度上是由大量的语音语料库推动的。然而,在资源有限的情况下将覆盖面扩大到各种语文仍然是一项艰巨的挑战。本文介绍了Speech Back-Translation,这是一个可扩展的管道,通过现成的文本到语音(TTS)模型将大规模文本语料库转换为合成语音来改进多语言ASR模型。我们证明,只有几十个小时的真实转录语音可以有效地训练TTS模型生成合成语音在数百倍的原始音量,同时保持高质量。为了评估合成语音质量,我们开发了一个基于可懂度的评估框架,并为合成数据何时有利于ASR训练建立了明确的阈值。使用Speech Back-Translation,我们以10种语言生成了超过500,000小时的合成语音,并继续预训练Whisper-large-v3,实现了超过30%的平均转录错误减少。这些结果突出了语音回译增强多语言ASR系统的可扩展性和有效性。
摘要:Recent advances in Automatic Speech Recognition (ASR) have been largely fueled by massive speech corpora. However, extending coverage to diverse languages with limited resources remains a formidable challenge. This paper introduces Speech Back-Translation, a scalable pipeline that improves multilingual ASR models by converting large-scale text corpora into synthetic speech via off-the-shelf text-to-speech (TTS) models. We demonstrate that just tens of hours of real transcribed speech can effectively train TTS models to generate synthetic speech at hundreds of times the original volume while maintaining high quality. To evaluate synthetic speech quality, we develop an intelligibility-based assessment framework and establish clear thresholds for when synthetic data benefits ASR training. Using Speech Back-Translation, we generate more than 500,000 hours of synthetic speech in ten languages and continue pre-training Whisper-large-v3, achieving average transcription error reductions of over 30\%. These results highlight the scalability and effectiveness of Speech Back-Translation for enhancing multilingual ASR systems.


【2】 EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion

标题: EZ-VC:轻松Zero-Shot任意语音转换
链接:https://arxiv.org/abs/2505.16691
作者: Advait Joglekar,  Divyanshu Singh,  Rooshil Rohit Bhatia,  S. Umesh 
备注:Submitted to EMNLP 2025, 7 pages, 2 figures, 5 Tables
摘要:近年来,语音转换研究越来越多地集中在提高现有方法的zero-shot能力上。尽管取得了显著的进步,但当前的架构仍然倾向于在跨语言设置的zero-shot中挣扎。他们也经常无法概括出那些看不见的语言和口音的人。在本文中,我们采用了一种简单而有效的方法,结合了离散语音表示的自监督模型与非自回归扩散变压器为基础的条件流匹配语音解码器。我们表明,这种架构使我们能够训练一个纯粹的无文本,自我监督的方式语音转换模型。我们的技术不需要多个编码器来解开语音特征。我们的模型还设法在zero-shot跨语言设置中表现出色,即使是看不见的语言。
摘要:Voice Conversion research in recent times has increasingly focused on improving the zero-shot capabilities of existing methods. Despite remarkable advancements, current architectures still tend to struggle in zero-shot cross-lingual settings. They are also often unable to generalize for speakers of unseen languages and accents. In this paper, we adopt a simple yet effective approach that combines discrete speech representations from self-supervised models with a non-autoregressive Diffusion-Transformer based conditional flow matching speech decoder. We show that this architecture allows us to train a voice-conversion model in a purely textless, self-supervised fashion. Our technique works without requiring multiple encoders to disentangle speech features. Our model also manages to excel in zero-shot cross-lingual settings even for unseen languages.


【3】 X-ARES: A Comprehensive Framework for Assessing Audio Encoder  Performance

标题: X-ARES:一个评估音频编码器性能的综合框架
链接:https://arxiv.org/abs/2505.16369
作者: Junbo Zhang,  Heinrich Dinkel,  Yadong Niu,  Chenyu Liu,  Si Cheng,  Anbei Zhao,  Jian Luan 
备注:Accepted by Interspeech 2025
摘要:我们介绍了X-ARES(eXtensive Audio Representation and Evaluation Suite),这是一种新颖的开源基准测试,旨在系统地评估不同领域的音频编码器性能。通过涵盖语音,环境声音和音乐的任务,X-ARES提供了两种评估音频表示的评估方法:线性微调和非参数化评估。该框架包括22个不同的任务,涵盖了音频处理的基本方面,从语音识别和情感检测到声音事件分类和音乐流派识别。我们对最先进的音频编码器的广泛评估揭示了不同任务和领域的显著性能差异,突出了一般音频表示学习的复杂性。
摘要:We introduces X-ARES (eXtensive Audio Representation and Evaluation Suite), a novel open-source benchmark designed to systematically assess audio encoder performance across diverse domains. By encompassing tasks spanning speech, environmental sounds, and music, X-ARES provides two evaluation approaches for evaluating audio representations: linear fine-tuning and unparameterized evaluation. The framework includes 22 distinct tasks that cover essential aspects of audio processing, from speech recognition and emotion detection to sound event classification and music genre identification. Our extensive evaluation of state-of-the-art audio encoders reveals significant performance variations across different tasks and domains, highlighting the complexity of general audio representation learning.


【4】 Layer-wise Investigation of Large-Scale Self-Supervised Music  Representation Models

标题: 大规模自我监督音乐表现模型的分层研究
链接:https://arxiv.org/abs/2505.16306
作者: Yizhi Zhou,  Haina Zhu,  Hangting Chen 
摘要:最近,基于自监督学习(SSL)的音乐信息检索预训练模型越来越流行,在各种下游任务中取得了成功。然而,对编码信息的具体含义及其适用性的研究有限。探索这些方面可以帮助我们更好地了解它们的能力和局限性,从而在下游任务中更有效地使用。   在这项研究中,我们分析了先进的音乐表示模型MusicFM和新出现的SSL模型MuQ。我们专注于三个主要方面:(i)验证SSL模型在多个下游任务中的优势,(ii)探索不同任务的分层信息的专业化,以及(iii)在选择特定层时比较性能差异。通过这些分析,我们揭示了SSL模型在音乐信息检索中的结构和潜在应用。
摘要:Recently, pre-trained models for music information retrieval based on self-supervised learning (SSL) are becoming popular, showing success in various downstream tasks. However, there is limited research on the specific meanings of the encoded information and their applicability. Exploring these aspects can help us better understand their capabilities and limitations, leading to more effective use in downstream tasks.   In this study, we analyze the advanced music representation model MusicFM and the newly emerged SSL model MuQ. We focus on three main aspects: (i) validating the advantages of SSL models across multiple downstream tasks, (ii) exploring the specialization of layer-wise information for different tasks, and (iii) comparing performance differences when selecting specific layers. Through this analysis, we reveal insights into the structure and potential applications of SSL models in music information retrieval.


【5】 Dialogue in Resonance: An Interactive Music Piece for Piano and  Real-Time Automatic Transcription System

标题: 共鸣中的对话:钢琴和实时自动抄写系统的互动音乐作品
链接:https://arxiv.org/abs/2505.16259
作者: Hayeon Bang,  Taegyun Kwon,  Juhan Nam 
摘要:本文介绍了一个交互式的音乐作品,为人类钢琴家和计算机控制的钢琴,集成了实时自动音乐转录到一个分数驱动的框架。与以前的方法,主要集中在即兴为基础的互动,我们的工作建立了一个平衡的框架,结合组成结构与动态互动。通过实时自动转录作为其核心机制,计算机实时解释和响应人类表演者的输入,创建一个音乐对话,平衡作曲意图与现场互动,同时融入不可预测的元素。本文介绍了从创作到首演的发展过程,包括技术实现、排练过程和演出考虑。
摘要:This paper presents, an interactive music piece for a human pianist and a computer-controlled piano that integrates real-time automatic music transcription into a score-driven framework. Unlike previous approaches that primarily focus on improvisation-based interactions, our work establishes a balanced framework that combines composed structure with dynamic interaction. Through real-time automatic transcription as its core mechanism, the computer interprets and responds to the human performer's input in real time, creating a musical dialogue that balances compositional intent with live interaction while incorporating elements of unpredictability. In this paper, we present the development process from composition to premiere performance, including technical implementation, rehearsal process, and performance considerations.


【6】 AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large  Language Models

标题: AudioTrust:对音频大型语言模型的多方面可信度进行基准测试
链接:https://arxiv.org/abs/2505.16211
作者: Kai Li,  Can Shen,  Yile Liu,  Jirui Han,  Kelong Zheng,  Xuechao Zou,  Zhe Wang,  Xingjian Du,  Shun Zhang,  Hanjun Luo,  Yingbin Jin,  Xinxin Xing,  Ziyang Ma,  Yue Liu,  Xiaojun Jia,  Yifan Zhang,  Junfeng Fang,  Kun Wang,  Yibo Yan,  Haoyang Li,  Yiming Li,  Xiaobin Zhuang,  Yang Liu,  Haibo Hu,  Zhuo Chen,  Zhizheng Wu,  Xiaolin Hu,  Eng-Siong Chng,  XiaoFeng Wang,  Wenyuan Xu,  Wei Dong,  Xinfeng Li 
备注:Technical Report
摘要:音频大语言模型(ALLM)的快速发展和不断扩大的应用要求严格理解其可信度。然而,评估这些模型的系统研究,特别是关于音频模式特有的风险,在很大程度上仍然是未开发的。现有的评价框架主要侧重于文本模态或仅涉及有限的一组安全维度,未能充分考虑音频模态固有的独特特征和应用场景。我们引入AudioTrust-第一个专门为ALLM设计的多方面可信度评估框架和基准。AudioTrust促进了六个关键维度的评估:公平性,幻觉,安全性,隐私性,鲁棒性和身份验证。为了全面评估这些维度,AudioTrust围绕18个不同的实验设置进行了构建。它的核心是一个精心构建的数据集,包含超过4,420个音频/文本样本,来自真实世界的场景(例如,日常对话、紧急呼叫、语音助理交互),专门用于探测ALLM的多方面可信度。在评估方面,该基准精心设计了9个音频特定的评估指标,我们采用了一个大规模的自动化管道,对模型输出进行客观和可扩展的评分。实验结果揭示了当前最先进的开源和闭源ALLM在面对各种高风险音频场景时的可信度边界和局限性,为未来音频模型的安全和可信部署提供了有价值的见解。我们的平台和基准可在https://github.com/JusperLee/AudioTrust上获得。
摘要:The rapid advancement and expanding applications of Audio Large Language Models (ALLMs) demand a rigorous understanding of their trustworthiness. However, systematic research on evaluating these models, particularly concerning risks unique to the audio modality, remains largely unexplored. Existing evaluation frameworks primarily focus on the text modality or address only a restricted set of safety dimensions, failing to adequately account for the unique characteristics and application scenarios inherent to the audio modality. We introduce AudioTrust-the first multifaceted trustworthiness evaluation framework and benchmark specifically designed for ALLMs. AudioTrust facilitates assessments across six key dimensions: fairness, hallucination, safety, privacy, robustness, and authentication. To comprehensively evaluate these dimensions, AudioTrust is structured around 18 distinct experimental setups. Its core is a meticulously constructed dataset of over 4,420 audio/text samples, drawn from real-world scenarios (e.g., daily conversations, emergency calls, voice assistant interactions), specifically designed to probe the multifaceted trustworthiness of ALLMs. For assessment, the benchmark carefully designs 9 audio-specific evaluation metrics, and we employ a large-scale automated pipeline for objective and scalable scoring of model outputs. Experimental results reveal the trustworthiness boundaries and limitations of current state-of-the-art open-source and closed-source ALLMs when confronted with various high-risk audio scenarios, offering valuable insights for the secure and trustworthy deployment of future audio models. Our platform and benchmark are available at https://github.com/JusperLee/AudioTrust.


【7】 Differentiable K-means for Fully-optimized Discrete Token-based ASR

标题: 完全优化的基于离散令牌的ASB的差异K均值
链接:https://arxiv.org/abs/2505.16207
作者: Kentaro Onda,  Yosuke Kashiwagi,  Emiru Tsunoo,  Hayato Futami,  Shinji Watanabe 
备注:Accepted by Interspeech2025
摘要:最近的研究强调了从自监督学习(SSL)模型中获得的离散令牌在各种语音相关任务中的潜力。这些令牌不仅可以作为语言建模中文本的替代品,还可以作为自动语音识别(ASR)等任务的中间表示。然而,离散令牌通常是通过独立于下游任务的SSL特征的k均值聚类获得的,这使得它们对于特定应用而言不是最佳的。本文提出了使用可微k-means,实现标记化和下游任务的联合优化。这种方法能够微调SSL参数和多个SSL层输出的学习权重。以ASR作为下游任务进行实验。由于优化的令牌,ASR准确性成功提高。所获得的令牌也表现出更大的语音信息的纯度,这被发现是有用的,即使在语音再合成。
摘要:Recent studies have highlighted the potential of discrete tokens derived from self-supervised learning (SSL) models for various speech-related tasks. These tokens serve not only as substitutes for text in language modeling but also as intermediate representations for tasks such as automatic speech recognition (ASR). However, discrete tokens are typically obtained via k-means clustering of SSL features independently of downstream tasks, making them suboptimal for specific applications. This paper proposes the use of differentiable k-means, enabling the joint optimization of tokenization and downstream tasks. This approach enables the fine-tuning of the SSL parameters and learning weights for outputs from multiple SSL layers. Experiments were conducted with ASR as a downstream task. ASR accuracy successfully improved owing to the optimized tokens. The acquired tokens also exhibited greater purity of phonetic information, which were found to be useful even in speech resynthesis.


【8】 SpecMaskFoley: Steering Pretrained Spectral Masked Generative  Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet

标题: SpecMaskFoley:引导预训练的频谱屏蔽生成Transformer器通过控制Net实现同步视频到音频合成
链接:https://arxiv.org/abs/2505.16195
作者: Zhi Zhong,  Akira Takahashi,  Shuyang Cui,  Keisuke Toyama,  Shusuke Takahashi,  Yuki Mitsufuji 
备注:4 pages, 2 figures, 2 tables. Demo page: this https URL
摘要:Foley合成旨在合成高质量的音频,其在语义和时间上都与视频帧对齐。鉴于其在创意产业中的广泛应用,该任务在研究界得到了越来越多的关注。为了避免从头开始训练音频生成模型的非平凡任务,将预训练的音频生成模型用于视频同步的Foley合成呈现出一个有吸引力的方向。ControlNet是一种将细粒度控件添加到预训练生成模型的方法,已被应用于Foley合成,但其使用仅限于手工制作的人类可读的时间条件。相比之下,从头开始的模型通过利用使用预训练的视频编码器提取的高维深度特征取得了成功。我们已经观察到基于ControlNet和从头开始的Foley模型之间的性能差距。为了缩小这一差距,我们提出了SpecMaskFoley,这是一种通过ControlNet将预训练的SpecMaskGIT模型转向视频同步Foley合成的方法。为了释放单个ControlNet分支的潜力,我们通过频率感知的时间特征对齐器解决了时间视频特征与预训练SpecMaskGIT的时间-频率性质之间的差异,从而消除了对现有技术中广泛使用的复杂调节机制的需求。对一个常见的福利合成基准的评估表明,SpecMaskFoley甚至可以超越强大的从头开始的基线,大大推进了基于ControlNet的福利合成模型的发展。演示页面:https://zzaudio.github.io/SpecMaskFoley_Demo/
摘要:Foley synthesis aims to synthesize high-quality audio that is both semantically and temporally aligned with video frames. Given its broad application in creative industries, the task has gained increasing attention in the research community. To avoid the non-trivial task of training audio generative models from scratch, adapting pretrained audio generative models for video-synchronized foley synthesis presents an attractive direction. ControlNet, a method for adding fine-grained controls to pretrained generative models, has been applied to foley synthesis, but its use has been limited to handcrafted human-readable temporal conditions. In contrast, from-scratch models achieved success by leveraging high-dimensional deep features extracted using pretrained video encoders. We have observed a performance gap between ControlNet-based and from-scratch foley models. To narrow this gap, we propose SpecMaskFoley, a method that steers the pretrained SpecMaskGIT model toward video-synchronized foley synthesis via ControlNet. To unlock the potential of a single ControlNet branch, we resolve the discrepancy between the temporal video features and the time-frequency nature of the pretrained SpecMaskGIT via a frequency-aware temporal feature aligner, eliminating the need for complicated conditioning mechanisms widely used in prior arts. Evaluations on a common foley synthesis benchmark demonstrate that SpecMaskFoley could even outperform strong from-scratch baselines, substantially advancing the development of ControlNet-based foley synthesis models. Demo page: https://zzaudio.github.io/SpecMaskFoley_Demo/


【9】 Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based  Resynthesis Only with Native Speech Corpora

标题: 仅使用母语语音库通过基于离散令牌的再合成来增强的外国口音模拟
链接:https://arxiv.org/abs/2505.16191
作者: Kentaro Onda,  Keisuke Imoto,  Satoru Fukayama,  Daisuke Saito,  Nobuaki Minematsu 
备注:Accepted by Interspeech2025
摘要:最近,提出了一种仅使用从自监督学习(SSL)模型获得的离散令牌与母语语音数据合成外国口音语音的方法。考虑到口音语音数据的有限性,该方法有望使模拟外国口音变得更加容易。通过将合成的带口音语音作为人类的听力材料或自动语音识别(ASR)的训练数据,两者都将获得对外国口音的更高鲁棒性。然而,以前的方法有一个致命的缺陷,它不能再现持续时间相关的口音。持续性口音通常出现在母语为音节计时或mora-timed节奏的第二语言使用者说重音计时语言时,如英语。在本文中,我们集成了持续时间修改到以前的方法,以更准确地模拟外国口音。实验表明,该方法成功地复制了真实的L2语音中看到的持续性口音。
摘要:Recently, a method for synthesizing foreign-accented speech only with native speech data using discrete tokens obtained from self-supervised learning (SSL) models was proposed. Considering limited availability of accented speech data, this method is expected to make it much easier to simulate foreign accents. By using the synthesized accented speech as listening materials for humans or training data for automatic speech recognition (ASR), both of them will acquire higher robustness against foreign accents. However, the previous method has a fatal flaw that it cannot reproduce duration-related accents. Durational accents are commonly seen when L2 speakers, whose native language has syllable-timed or mora-timed rhythm, speak stress-timed languages, such as English. In this paper, we integrate duration modification to the previous method to simulate foreign accents more accurately. Experiments show that the proposed method successfully replicates durational accents seen in real L2 speech.


【10】 Discrete Tokens Exhibit Interlanguage Speech Intelligibility Benefit: an  Analytical Study Towards Accent-robust ASR Only with Native Speech Data

标题: 离散令牌展现出中介语语音可理解性优势:仅使用母语语音数据实现口音稳健的ASB的分析研究
链接:https://arxiv.org/abs/2505.16182
作者: Kentaro Onda,  Keisuke Imoto,  Satoru Fukayama,  Daisuke Saito,  Nobuaki Minematsu 
备注:Accepted by Interspeech2025
摘要:在这项研究中,我们获得的见解,有助于实现口音鲁棒ASR只使用母语数据。在人类对非母语语音的感知中,观察到被称为“中介语语音可理解性益处”(ISIB)的现象,其中与说话者共享母语的非母语听者甚至比母语听者更好地理解语音。基于从自监督学习(SSL)模型中提取的离散令牌代表人类对语音的感知的想法,我们对基于离散令牌的ASR对非母语语音的鲁棒性进行了分析研究,改变了用于训练令牌化的语言,这被视为ISIB的技术实现。结果表明,ISIB实际上发生在离散的基于令牌的ASR。由于我们的方法只依赖于本地语音数据来模拟人类感知的行为,预计将适用于广泛的口音,语音数据是稀缺的。
摘要:In this study, we gained insight that contributes to achieving accent-robust ASR using only native speech data. In human perception of non-native speech, the phenomenon known as "interlanguage speech intelligibility benefit" (ISIB) is observed, where non-native listeners who share the native language with the speaker understand the speech better compared even to native listeners. Based on the idea that discrete tokens extracted from self-supervised learning (SSL) models represent the human perception of speech, we conducted an analytical study on the robustness of discrete token-based ASR to non-native speech, varying the language used for training the tokenization, which is viewed as a technical implementation of ISIB. The results showed that ISIB actually occurred in the discrete token-based ASR. Since our approach relies only on native speech data to simulate the behavior of human perception, it is expected to be applicable to a wide range of accents for which speech data is scarce.


【11】 Selective Invocation for Multilingual ASR: A Cost-effective Approach  Adapting to Speech Recognition Difficulty

标题: 多语言ASB的选择性调用:一种适应语音识别困难的经济有效方法
链接:https://arxiv.org/abs/2505.16168
作者: Hongfei Xue,  Yufeng Tang,  Jun Zhang,  Xuelong Geng,  Lei Xie 
备注:Accepted by INTERSPEECH 2025
摘要:虽然多语言自动语音识别(ASR)系统已经取得了显著的进步,使单一模型能够处理多种语言,但固有的语言差异和数据不平衡挑战了SOTA在所有语言中的性能。虽然语言识别(LID)模型可以将语音路由到适当的ASR模型,但它们会因调用SOTA商业模型而产生高昂的成本,并且由于错误分类而导致不准确。为了克服这些问题,我们提出SIMA,一个选择性调用多语言ASR,适应输入语音的难度水平。SIMA建立在口语大语言模型(SLLM)上,它评估输入是否足够简单,可以直接转录,或者需要调用SOTA ASR模型。与SLLM相比,我们的方法将字错误率降低了18.7%,与基于LID的方法相比,调用成本降低了一半。三个数据集上的测试表明,SIMA是一个可扩展的,成本效益高的多语言ASR应用程序的解决方案。
摘要:Although multilingual automatic speech recognition (ASR) systems have significantly advanced, enabling a single model to handle multiple languages, inherent linguistic differences and data imbalances challenge SOTA performance across all languages. While language identification (LID) models can route speech to the appropriate ASR model, they incur high costs from invoking SOTA commercial models and suffer from inaccuracies due to misclassification. To overcome these, we propose SIMA, a selective invocation for multilingual ASR that adapts to the difficulty level of the input speech. Built on a spoken large language model (SLLM), SIMA evaluates whether the input is simple enough for direct transcription or requires the invocation of a SOTA ASR model. Our approach reduces word error rates by 18.7% compared to the SLLM and halves invocation costs compared to LID-based methods. Tests on three datasets show that SIMA is a scalable, cost-effective solution for multilingual ASR applications.


【12】 Source Separation by Flow Matching

标题: 通过流量匹配进行源分离
链接:https://arxiv.org/abs/2505.16119
作者: Robin Scheibler,  John R. Hershey,  Arnaud Doucet,  Henry Li 
备注:5 pages, 3 figures, 2 tables
摘要:我们考虑单声道音频源分离的问题,其目标是从它们的混合物中重建$K$个源。我们解决这个不适定的问题与FLOSS(流匹配源分离),基于流匹配的约束生成方法,确保严格的混合一致性。流匹配是一种通用方法,当给定来自定义在同一空间上的两个概率分布的样本时,学习常微分方程以在提供来自另一个分布的样本时输出来自其中一个分布的样本。在我们的上下文中,我们可以访问来自$K$源的联合分布的样本,因此相应的样本来自其混合物的低维分布。为了应用流匹配,我们用人工噪声分量增强这些混合样本,以确保得到的“增强”分布与$K$源分布的维度相匹配。此外,由于源的任何排列产生相同的混合物,我们采用了依赖于合适的定制设计的神经网络架构的流量匹配的等变公式。我们证明了重叠语音分离的方法的性能。
摘要:We consider the problem of single-channel audio source separation with the goal of reconstructing $K$ sources from their mixture. We address this ill-posed problem with FLOSS (FLOw matching for Source Separation), a constrained generation method based on flow matching, ensuring strict mixture consistency. Flow matching is a general methodology that, when given samples from two probability distributions defined on the same space, learns an ordinary differential equation to output a sample from one of the distributions when provided with a sample from the other. In our context, we have access to samples from the joint distribution of $K$ sources and so the corresponding samples from the lower-dimensional distribution of their mixture. To apply flow matching, we augment these mixture samples with artificial noise components to ensure the resulting "augmented" distribution matches the dimensionality of the $K$ source distribution. Additionally, as any permutation of the sources yields the same mixture, we adopt an equivariant formulation of flow matching which relies on a suitable custom-designed neural network architecture. We demonstrate the performance of the method for the separation of overlapping speech.


【13】 A Novel Deep Learning Framework for Efficient Multichannel Acoustic  Feedback Control

标题: 用于高效多通道声反馈控制的新型深度学习框架
链接:https://arxiv.org/abs/2505.15914
作者: Yuan-Kuei Wu,  Juan Azcarreta,  Kashyap Patel,  Buye Xu,  Jung-Suk Lee,  Sanha Lee,  Ashutosh Pandey 
备注:Accepted by Interspeech 2025
摘要:这项研究提出了一个深度学习框架,用于控制音频设备中的多通道声反馈。传统的数字信号处理方法在处理高度相关的噪声(如反馈)时难以收敛。我们引入了一个卷积递归网络,它有效地结合了空间和时间处理,大大提高了语音增强能力,降低了计算需求。我们的方法利用三种训练方法:在环训练,教师强迫,和一个混合策略与多通道维纳滤波器,优化性能在复杂的声学环境。这个可扩展的框架为现实世界的应用提供了一个强大的解决方案,使声反馈控制技术取得了重大进展。
摘要:This study presents a deep-learning framework for controlling multichannel acoustic feedback in audio devices. Traditional digital signal processing methods struggle with convergence when dealing with highly correlated noise such as feedback. We introduce a Convolutional Recurrent Network that efficiently combines spatial and temporal processing, significantly enhancing speech enhancement capabilities with lower computational demands. Our approach utilizes three training methods: In-a-Loop Training, Teacher Forcing, and a Hybrid strategy with a Multichannel Wiener Filter, optimizing performance in complex acoustic environments. This scalable framework offers a robust solution for real-world applications, making significant advances in Acoustic Feedback Control technology.


【14】 Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame  Rate

标题: 解锁时间灵活性:可变帧率的神经语音编解码器
链接:https://arxiv.org/abs/2505.16845
作者: Hanglei Zhang,  Yiwei Guo,  Zhihan Li,  Xiang Hao,  Xie Chen,  Kai Yu 
备注:Accepted to Interspeech 2025
摘要:大多数神经语音编解码器通过帧内机制(例如码本丢弃)以恒定帧速率(CFR)实现比特率调整。然而,语音段固有地具有时变信息密度(例如,无声间隔对有声区域)。该属性使得CFR在比特率和令牌序列长度方面不是最优的,从而阻碍了实时应用的效率。在这项工作中,我们提出了一个时间灵活的编码(TFC)技术,引入可变帧速率(VFR)的神经语音编解码器的第一次。TFC支持无缝可调的平均帧速率,并基于时间熵动态分配帧速率。实验结果表明,采用TFC的编解码器具有较高的灵活性,可实现最佳的重建质量,即使在较低的帧速率下也能保持有竞争力的性能。我们的方法是有希望的集成与其他努力,开发低帧速率的神经语音编解码器,更有效的下游任务。
摘要:Most neural speech codecs achieve bitrate adjustment through intra-frame mechanisms, such as codebook dropout, at a Constant Frame Rate (CFR). However, speech segments inherently have time-varying information density (e.g., silent intervals versus voiced regions). This property makes CFR not optimal in terms of bitrate and token sequence length, hindering efficiency in real-time applications. In this work, we propose a Temporally Flexible Coding (TFC) technique, introducing variable frame rate (VFR) into neural speech codecs for the first time. TFC enables seamlessly tunable average frame rates and dynamically allocates frame rates based on temporal entropy. Experimental results show that a codec with TFC achieves optimal reconstruction quality with high flexibility, and maintains competitive performance even at lower frame rates. Our approach is promising for the integration with other efforts to develop low-frame-rate neural speech codecs for more efficient downstream tasks.


【15】 Attractor-Based Speech Separation of Multiple Utterances by Unknown  Number of Speakers

标题: 基于吸引子的未知说话人多段语音分离
链接:https://arxiv.org/abs/2505.16607
作者: Yuzhu Wang,  Archontis Politis,  Konstantinos Drossos,  Tuomas Virtanen 
备注:5 pages, 4 figures, accepted by Interspeech 2025
摘要:本文研究了单通道语音分离问题,其中说话人的数目是未知的,每个说话人可以说多个话语。我们提出了一个语音分离模型,同时进行分离,动态估计扬声器的数量,并检测单个扬声器活动集成的吸引子模块。所提出的系统优于现有的方法,通过引入一个吸引子为基础的架构,有效地结合了本地和全球多话语的情况下的时间建模。为了在混响和噪声条件下评估该方法,通过将Librispeech语音信号与WHAM!噪声信号结果表明,所提出的系统准确地估计源的数量。该系统有效地检测源活动,并将相应的话语分离成正确的输出,在已知和未知的源计数的情况下。
摘要:This paper addresses the problem of single-channel speech separation, where the number of speakers is unknown, and each speaker may speak multiple utterances. We propose a speech separation model that simultaneously performs separation, dynamically estimates the number of speakers, and detects individual speaker activities by integrating an attractor module. The proposed system outperforms existing methods by introducing an attractor-based architecture that effectively combines local and global temporal modeling for multi-utterance scenarios. To evaluate the method in reverberant and noisy conditions, a multi-speaker multi-utterance dataset was synthesized by combining Librispeech speech signals with WHAM! noise signals. The results demonstrate that the proposed system accurately estimates the number of sources. The system effectively detects source activities and separates the corresponding utterances into correct outputs in both known and unknown source count scenarios.


【16】 UBGAN: Enhancing Coded Speech with Blind and Guided Bandwidth Extension

标题: UBGAN:用盲和引导带宽扩展增强编码语音
链接:https://arxiv.org/abs/2505.16404
作者: Kishan Gupta,  Srikanth Korse,  Andreas Brendel,  Nicola Pia,  Guillaume Fuchs 
摘要:在语音编解码器的实际应用中,诸如无线电连接的质量、限制硬件或所需的用户体验的众多因素使得在可实现的感知质量、产生的比特率和计算复杂度之间进行权衡成为必要。大多数传统和神经语音编解码器对宽带(WB)语音信号进行操作以实现这种折衷。为了进一步提高编码语音的感知质量,传输语音的带宽扩展(BWE)是传统语音编码中有吸引力和流行的技术。相比之下,神经语音编解码器通常是针对特定要求集进行端到端训练的,并且通常不容易适应。特别地,它们通常被训练为以单个固定采样率操作。通过通用带宽扩展生成对抗网络(UBGAN),我们提出了一种模块化和轻量级的基于GAN的解决方案,可以提高各种传统和神经编解码器的操作灵活性。我们的模型在子带域中操作,并将WB信号的带宽从8 kHz扩展到16 kHz,从而产生超宽带(SWB)信号。我们进一步介绍了两个变体,引导UBGAN和盲UBGAN,其中引导版本以非常低的比特率传输量化的学习表示作为边信息,除了编解码器的比特率之外,而盲BWE在没有这样的边信息的情况下操作。我们的主观评估证明了UBGAN应用于WB编解码器的优势,并强调了我们提出的方法在多个编解码器和比特率上的泛化能力。
摘要:In practical application of speech codecs, a multitude of factors such as the quality of the radio connection, limiting hardware or required user experience necessitate trade-offs between achievable perceptual quality, engendered bitrate and computational complexity. Most conventional and neural speech codecs operate on wideband (WB) speech signals to achieve this compromise. To further enhance the perceptual quality of coded speech, bandwidth extension (BWE) of the transmitted speech is an attractive and popular technique in conventional speech coding. In contrast, neural speech codecs are typically trained end-to-end to a specific set of requirements and are often not easily adaptable. In particular, they are typically trained to operate at a single fixed sampling rate. With the Universal Bandwidth Extension Generative Adversarial Network (UBGAN), we propose a modular and lightweight GAN-based solution that increases the operational flexibility of a wide range of conventional and neural codecs. Our model operates in the subband domain and extends the bandwidth of WB signals from 8 kHz to 16 kHz, resulting in super-wideband (SWB) signals. We further introduce two variants, guided-UBGAN and blind-UBGAN, where the guided version transmits quantized learned representation as a side information at a very low bitrate additional to the bitrate of the codec, while blind-BWE operates without such side-information. Our subjective assessments demonstrate the advantage of UBGAN applied to WB codecs and highlight the generalization capacity of our proposed method across multiple codecs and bitrates.


【17】 Towards Holistic Evaluation of Large Audio-Language Models: A  Comprehensive Survey

标题: 大型音频语言模型的整体评估:全面调查
链接:https://arxiv.org/abs/2505.15957
作者: Chih-Kai Yang,  Neo S. Ho,  Hung-yi Lee 
备注:Project Website: this https URL
摘要:随着大型音频语言模型(LALM)的进步,这些模型增强了具有听觉能力的大型语言模型(LLM),预计这些模型将在各种听觉任务中表现出普遍的熟练程度。虽然已经出现了许多基准来评估LALM的性能,但它们仍然是零散的,缺乏结构化的分类。为了弥补这一差距,我们进行了一项全面的调查,并提出了一个系统的分类法LALM评估,分为四个方面的基础上,他们的目标:(1)一般的听觉意识和处理,(2)知识和推理,(3)对话导向的能力,(4)公平,安全和可信度。我们提供了每个类别的详细概述,并强调了该领域的挑战,为有前途的未来方向提供了见解。据我们所知,这是第一次专门针对LALM评估的调查,为社区提供了明确的指导方针。我们将发布调查论文集,并积极维护它,以支持该领域的持续进步。
摘要:With advancements in large audio-language models (LALMs), which enhance large language models (LLMs) with auditory capabilities, these models are expected to demonstrate universal proficiency across various auditory tasks. While numerous benchmarks have emerged to assess LALMs' performance, they remain fragmented and lack a structured taxonomy. To bridge this gap, we conduct a comprehensive survey and propose a systematic taxonomy for LALM evaluations, categorizing them into four dimensions based on their objectives: (1) General Auditory Awareness and Processing, (2) Knowledge and Reasoning, (3) Dialogue-oriented Ability, and (4) Fairness, Safety, and Trustworthiness. We provide detailed overviews within each category and highlight challenges in this field, offering insights into promising future directions. To the best of our knowledge, this is the first survey specifically focused on the evaluations of LALMs, providing clear guidelines for the community. We will release the collection of the surveyed papers and actively maintain it to support ongoing advancements in the field.


eess.AS音频处理

【1】 Active Speech Enhancement: Active Speech Denoising Decliping and  Deveraberation
标题: 主动语音增强:主动语音去噪去倾斜和去失真
链接:https://arxiv.org/abs/2505.16911
作者: Ofir Yaish,  Yehuda Mishaly,  Eliya Nachmani 
摘要:我们介绍了一种新的范例主动声音修改:主动语音增强(ASE)。虽然主动噪声消除(ANC)算法专注于抑制外部干扰,但ASE更进一步,主动塑造语音信号-既衰减不需要的噪声分量,又放大语音相关频率-以提高清晰度和感知质量。为了实现这一点,我们提出了一种新的基于Transformer-Mamba的架构,以及一个特定于任务的损失函数,旨在联合优化干扰抑制和信号富集。我们的方法在多个语音处理任务(包括去噪、去混响和去唇)中优于现有的基线,证明了主动、有针对性的调制在具有挑战性的声学环境中的有效性。
摘要:We introduce a new paradigm for active sound modification: Active Speech Enhancement (ASE). While Active Noise Cancellation (ANC) algorithms focus on suppressing external interference, ASE goes further by actively shaping the speech signal -- both attenuating unwanted noise components and amplifying speech-relevant frequencies -- to improve intelligibility and perceptual quality. To enable this, we propose a novel Transformer-Mamba-based architecture, along with a task-specific loss function designed to jointly optimize interference suppression and signal enrichment. Our method outperforms existing baselines across multiple speech processing tasks -- including denoising, dereverberation, and declipping -- demonstrating the effectiveness of active, targeted modulation in challenging acoustic environments.


【2】 Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame  Rate

标题: 解锁时间灵活性:可变帧率的神经语音编解码器
链接:https://arxiv.org/abs/2505.16845
作者: Hanglei Zhang,  Yiwei Guo,  Zhihan Li,  Xiang Hao,  Xie Chen,  Kai Yu 
备注:Accepted to Interspeech 2025
摘要:大多数神经语音编解码器通过帧内机制(例如码本丢弃)以恒定帧速率(CFR)实现比特率调整。然而,语音段固有地具有时变信息密度(例如,无声间隔对有声区域)。该属性使得CFR在比特率和令牌序列长度方面不是最优的,从而阻碍了实时应用的效率。在这项工作中,我们提出了一个时间灵活的编码(TFC)技术,引入可变帧速率(VFR)的神经语音编解码器的第一次。TFC支持无缝可调的平均帧速率,并基于时间熵动态分配帧速率。实验结果表明,采用TFC的编解码器具有较高的灵活性,可实现最佳的重建质量,即使在较低的帧速率下也能保持有竞争力的性能。我们的方法是有希望的集成与其他努力,开发低帧速率的神经语音编解码器,更有效的下游任务。
摘要:Most neural speech codecs achieve bitrate adjustment through intra-frame mechanisms, such as codebook dropout, at a Constant Frame Rate (CFR). However, speech segments inherently have time-varying information density (e.g., silent intervals versus voiced regions). This property makes CFR not optimal in terms of bitrate and token sequence length, hindering efficiency in real-time applications. In this work, we propose a Temporally Flexible Coding (TFC) technique, introducing variable frame rate (VFR) into neural speech codecs for the first time. TFC enables seamlessly tunable average frame rates and dynamically allocates frame rates based on temporal entropy. Experimental results show that a codec with TFC achieves optimal reconstruction quality with high flexibility, and maintains competitive performance even at lower frame rates. Our approach is promising for the integration with other efforts to develop low-frame-rate neural speech codecs for more efficient downstream tasks.


【3】 SEED: Speaker Embedding Enhancement Diffusion Model

标题: SEED:说话者嵌入增强扩散模型
链接:https://arxiv.org/abs/2505.16798
作者: KiHyun Nam,  Jungwoo Heo,  Jee-weon Jung,  Gangin Park,  Chaeyoung Jung,  Ha-Jin Yu,  Joon Son Chung 
备注:Accepted to Interspeech 2025. The official code can be found at this https URL
摘要:在实际应用中部署说话人识别系统的主要挑战是环境失配引起的性能下降。我们提出了一种基于扩散的方法,从预先训练的说话人识别模型中提取说话人嵌入,并生成细化的嵌入。对于训练,我们的方法通过扩散模型的前向过程分别将高斯噪声添加到从干净和有噪语音中提取的干净和有噪说话人嵌入中,然后在反向过程中将其重建为干净的嵌入。在推理过程中,所有嵌入都通过扩散过程重新生成。我们的方法既不需要说话人标签,也不需要对现有的说话人识别管道进行任何修改。在模拟环境失配场景的评估集上的实验表明,我们的方法可以将识别准确率提高19.6%,同时保留传统场景的性能。我们在这里发布代码https://github.com/kaistmm/seed-pytorch
摘要:A primary challenge when deploying speaker recognition systems in real-world applications is performance degradation caused by environmental mismatch. We propose a diffusion-based method that takes speaker embeddings extracted from a pre-trained speaker recognition model and generates refined embeddings. For training, our approach progressively adds Gaussian noise to both clean and noisy speaker embeddings extracted from clean and noisy speech, respectively, via forward process of a diffusion model, and then reconstructs them to clean embeddings in the reverse process. While inferencing, all embeddings are regenerated via diffusion process. Our method needs neither speaker label nor any modification to the existing speaker recognition pipeline. Experiments on evaluation sets simulating environment mismatch scenarios show that our method can improve recognition accuracy by up to 19.6% over baseline models while retaining performance on conventional scenarios. We publish our code here https://github.com/kaistmm/seed-pytorch


【4】 Adversarial Deep Metric Learning for Cross-Modal Audio-Text Alignment in  Open-Vocabulary Keyword Spotting

标题: 对抗性深度度量学习用于开放词汇关键词发现中的跨模式音频文本对齐
链接:https://arxiv.org/abs/2505.16735
作者: Youngmoon Jung,  Yong-Hyeok Lee,  Myunghun Jung,  Jaeyoung Roh,  Chang Woo Han,  Hoon-Young Cho 
备注:5 pages, 1 figures, Accepted at Interspeech 2025
摘要:对于基于文本注册的开放词汇关键词发现(KWS),声学和文本嵌入通常在音素或话语级别进行比较。为了实现这一点,我们使用深度度量学习(DML)优化声学和文本编码器,从而在共享嵌入空间中直接比较多模态嵌入。然而,音频和文本模态之间的固有异质性提出了一个重大挑战。为了解决这个问题,我们提出了模态对抗学习(MAL),它减少了异构模态表示中的域间隙。具体来说,我们训练一个模态分类器adversarially鼓励两个编码器生成模态不变的嵌入。此外,我们应用DML来实现音频和文本之间的音素级对齐,并在各种DML目标之间进行全面的比较。在华尔街日报(WSJ)和LibriPhrase数据集上的实验证明了该方法的有效性。
摘要:For text enrollment-based open-vocabulary keyword spotting (KWS), acoustic and text embeddings are typically compared at either the phoneme or utterance level. To facilitate this, we optimize acoustic and text encoders using deep metric learning (DML), enabling direct comparison of multi-modal embeddings in a shared embedding space. However, the inherent heterogeneity between audio and text modalities presents a significant challenge. To address this, we propose Modality Adversarial Learning (MAL), which reduces the domain gap in heterogeneous modality representations. Specifically, we train a modality classifier adversarially to encourage both encoders to generate modality-invariant embeddings. Additionally, we apply DML to achieve phoneme-level alignment between audio and text, and conduct comprehensive comparisons across various DML objectives. Experiments on the Wall Street Journal (WSJ) and LibriPhrase datasets demonstrate the effectiveness of the proposed approach.


【5】 Performance of Objective Speech Quality Metrics on Languages Beyond  Validation Data: A Study of Turkish and Korean

标题: 验证数据之外的语言的客观语音质量测试性能:土耳其语和韩语的研究
链接:https://arxiv.org/abs/2505.16616
作者: Javier Perez,  Dimme de Groot,  Jorge Martinez 
摘要:客观的语音质量测量被广泛用于评估视频会议平台和电信系统的性能。它们预测人类评价的语音质量,对于评估系统体验质量至关重要。尽管广泛使用,但质量措施是在有限的一组语言上制定的。这可能是有问题的,因为在看不见的语言上的性能因此得不到保证,甚至无法研究。在这里,我们通过调查土耳其语和韩语的两个客观语音质量指标(PESQ和ViSQOL)的性能来提高对这个问题的认识。使用英语作为基线,我们发现土耳其样本的ViSQOL评分显着更高,并且对于土耳其男性使用者来说,PESQ和ViSQOL之间的相关性最高。这些结果强调了探索跨指标偏差的必要性,并开发了具有多种语言的标记语音质量数据集。
摘要:Objective speech quality measures are widely used to assess the performance of video conferencing platforms and telecommunication systems. They predict human-rated speech quality and are crucial for assessing the systems quality of experience. Despite the widespread use, the quality measures are developed on a limited set of languages. This can be problematic since the performance on unseen languages is consequently not guaranteed or even studied. Here we raise awareness to this issue by investigating the performance of two objective speech quality measures (PESQ and ViSQOL) on Turkish and Korean. Using English as baseline, we show that Turkish samples have significantly higher ViSQOL scores and that for Turkish male speakers the correlation between PESQ and ViSQOL is highest. These results highlight the need to explore biases across metrics and to develop a labeled speech quality dataset with a variety of languages.


【6】 Attractor-Based Speech Separation of Multiple Utterances by Unknown  Number of Speakers

标题: 基于吸引子的未知说话人多段语音分离
链接:https://arxiv.org/abs/2505.16607
作者: Yuzhu Wang,  Archontis Politis,  Konstantinos Drossos,  Tuomas Virtanen 
备注:5 pages, 4 figures, accepted by Interspeech 2025
摘要:本文研究了单通道语音分离问题,其中说话人的数目是未知的,每个说话人可以说多个话语。我们提出了一个语音分离模型,同时进行分离,动态估计扬声器的数量,并检测单个扬声器活动集成的吸引子模块。所提出的系统优于现有的方法,通过引入一个吸引子为基础的架构,有效地结合了本地和全球多话语的情况下的时间建模。为了在混响和噪声条件下评估该方法,通过将Librispeech语音信号与WHAM!噪声信号结果表明,所提出的系统准确地估计源的数量。该系统有效地检测源活动,并将相应的话语分离成正确的输出,在已知和未知的源计数的情况下。
摘要:This paper addresses the problem of single-channel speech separation, where the number of speakers is unknown, and each speaker may speak multiple utterances. We propose a speech separation model that simultaneously performs separation, dynamically estimates the number of speakers, and detects individual speaker activities by integrating an attractor module. The proposed system outperforms existing methods by introducing an attractor-based architecture that effectively combines local and global temporal modeling for multi-utterance scenarios. To evaluate the method in reverberant and noisy conditions, a multi-speaker multi-utterance dataset was synthesized by combining Librispeech speech signals with WHAM! noise signals. The results demonstrate that the proposed system accurately estimates the number of sources. The system effectively detects source activities and separates the corresponding utterances into correct outputs in both known and unknown source count scenarios.


【7】 HPP-Voice: A Large-Scale Evaluation of Speech Embeddings for  Multi-Phenotypic Classification

标题: HPP-Voice:用于多表型分类的言语嵌入的大规模评估
链接:https://arxiv.org/abs/2505.16490
作者: David Krongauz,  Hido Pinto,  Sarah Kohn,  Yanir Marmor,  Eran Segal 
摘要:人类语音包含反映说话者的生理和神经状态的非语言学线索,可能使各种医学表型的非侵入性检测成为可能。我们介绍了人类表型项目语音语料库(HPP-语音):一个包含7,188个录音的数据集,其中讲希伯来语的成年人计数30秒,每个说话者与多达15个潜在的语音相关表型相关,包括呼吸,睡眠,心理健康,代谢,免疫和神经系统疾病。我们对14种现代语音嵌入模型进行了系统的比较,其中来自这些30秒计数任务的现代语音嵌入优于下游健康状况分类的MFCC和人口统计学。我们发现,从说话人识别模型中学习的嵌入可以预测客观测量的中度至重度睡眠呼吸暂停男性的AUC为0.64 $\pm $0.03,而MFCC和人口统计特征分别导致AUC为0.56 $\pm $0.02和0.57 $\pm $0.02。此外,我们的研究结果揭示了不同医学领域模型有效性的性别特异性模式。对于男性,说话者识别和日记模型始终优于呼吸条件的语音基础模型(例如,哮喘:0.61 $\pm $0.03 vs. 0.56 $\pm $0.02)和睡眠相关疾病(失眠:0.65 $\pm $0.04 vs. 0.59 $\pm $0.05)。对于女性,说话人日记模型表现最好的吸烟状态(0.61 $\pm $0.02比0.55 $\pm $0.02),而希伯来语特定的模型表现最好(0.59 $\pm $0.02比0.58 $\pm $0.02)在分类焦虑相比,语音基础模型。我们的研究结果提供了证据,证明一个简单的计数任务可以支持大规模的多表型语音筛查,并突出嵌入家族最适合特定条件,这些见解可以指导未来的声音生物标志物研究和临床部署。
摘要:Human speech contains paralinguistic cues that reflect a speaker's physiological and neurological state, potentially enabling non-invasive detection of various medical phenotypes. We introduce the Human Phenotype Project Voice corpus (HPP-Voice): a dataset of 7,188 recordings in which Hebrew-speaking adults count for 30 seconds, with each speaker linked to up to 15 potentially voice-related phenotypes spanning respiratory, sleep, mental health, metabolic, immune, and neurological conditions. We present a systematic comparison of 14 modern speech embedding models, where modern speech embeddings from these 30-second counting tasks outperform MFCCs and demographics for downstream health condition classifications. We found that embedding learned from a speaker identification model can predict objectively measured moderate to severe sleep apnea in males with an AUC of 0.64 $\pm$ 0.03, while MFCC and demographic features led to AUCs of 0.56 $\pm$ 0.02 and 0.57 $\pm$ 0.02, respectively. Additionally, our results reveal gender-specific patterns in model effectiveness across different medical domains. For males, speaker identification and diarization models consistently outperformed speech foundation models for respiratory conditions (e.g., asthma: 0.61 $\pm$ 0.03 vs. 0.56 $\pm$ 0.02) and sleep-related conditions (insomnia: 0.65 $\pm$ 0.04 vs. 0.59 $\pm$ 0.05). For females, speaker diarization models performed best for smoking status (0.61 $\pm$ 0.02 vs 0.55 $\pm$ 0.02), while Hebrew-specific models performed best (0.59 $\pm$ 0.02 vs. 0.58 $\pm$ 0.02) in classifying anxiety compared to speech foundation models. Our findings provide evidence that a simple counting task can support large-scale, multi-phenotypic voice screening and highlight which embedding families generalize best to specific conditions, insights that can guide future vocal biomarker research and clinical deployment.


【8】 UBGAN: Enhancing Coded Speech with Blind and Guided Bandwidth Extension

标题: UBGAN:用盲和引导带宽扩展增强编码语音
链接:https://arxiv.org/abs/2505.16404
作者: Kishan Gupta,  Srikanth Korse,  Andreas Brendel,  Nicola Pia,  Guillaume Fuchs 
摘要:在语音编解码器的实际应用中,诸如无线电连接的质量、限制硬件或所需的用户体验的众多因素使得在可实现的感知质量、产生的比特率和计算复杂度之间进行权衡成为必要。大多数传统和神经语音编解码器对宽带(WB)语音信号进行操作以实现这种折衷。为了进一步提高编码语音的感知质量,传输语音的带宽扩展(BWE)是传统语音编码中一种有吸引力且流行的技术。相比之下,神经语音编解码器通常是针对特定要求集进行端到端训练的,并且通常不容易适应。特别地,它们通常被训练为以单个固定采样率操作。通过通用带宽扩展生成对抗网络(UBGAN),我们提出了一种模块化和轻量级的基于GAN的解决方案,可以提高各种传统和神经编解码器的操作灵活性。我们的模型在子带域中操作,并将WB信号的带宽从8 kHz扩展到16 kHz,从而产生超宽带(SWB)信号。我们进一步介绍了两个变体,引导UBGAN和盲UBGAN,其中引导版本以非常低的比特率传输量化的学习表示作为边信息,除了编解码器的比特率之外,而盲BWE在没有这样的边信息的情况下操作。我们的主观评估证明了UBGAN应用于WB编解码器的优势,并强调了我们提出的方法在多个编解码器和比特率上的泛化能力。
摘要:In practical application of speech codecs, a multitude of factors such as the quality of the radio connection, limiting hardware or required user experience necessitate trade-offs between achievable perceptual quality, engendered bitrate and computational complexity. Most conventional and neural speech codecs operate on wideband (WB) speech signals to achieve this compromise. To further enhance the perceptual quality of coded speech, bandwidth extension (BWE) of the transmitted speech is an attractive and popular technique in conventional speech coding. In contrast, neural speech codecs are typically trained end-to-end to a specific set of requirements and are often not easily adaptable. In particular, they are typically trained to operate at a single fixed sampling rate. With the Universal Bandwidth Extension Generative Adversarial Network (UBGAN), we propose a modular and lightweight GAN-based solution that increases the operational flexibility of a wide range of conventional and neural codecs. Our model operates in the subband domain and extends the bandwidth of WB signals from 8 kHz to 16 kHz, resulting in super-wideband (SWB) signals. We further introduce two variants, guided-UBGAN and blind-UBGAN, where the guided version transmits quantized learned representation as a side information at a very low bitrate additional to the bitrate of the codec, while blind-BWE operates without such side-information. Our subjective assessments demonstrate the advantage of UBGAN applied to WB codecs and highlight the generalization capacity of our proposed method across multiple codecs and bitrates.


【9】 Multi-Channel Sequence-to-Sequence Neural Diarization: Experimental  Results for The MISP 2025 Challenge

标题: 多通道序列到序列神经日志:MISP 2025挑战的实验结果
链接:https://arxiv.org/abs/2505.16387
作者: Ming Cheng,  Fei Su,  Cancan Li,  Juan Liu,  Ming Li 
备注:Accepted by Interspeech2025
摘要:本文介绍了为多模态基于信息的语音处理(MISP)2025挑战赛开发的扬声器日志化系统。首先,我们利用序列到序列神经日志(S2 SND)框架使用单声道音频生成初始预测。然后,我们扩展了原始的S2 SND框架,创建了一个新的版本,多通道序列到序列神经日记(MC-S2 SND),它使用多通道音频细化了初始结果。最终系统在比赛数据库的评测集上实现了8.09%的日志化错误率(DER),在MISP 2025挑战赛的说话人日志化任务中排名第一。
摘要:This paper describes the speaker diarization system developed for the Multimodal Information-Based Speech Processing (MISP) 2025 Challenge. First, we utilize the Sequence-to-Sequence Neural Diarization (S2SND) framework to generate initial predictions using single-channel audio. Then, we extend the original S2SND framework to create a new version, Multi-Channel Sequence-to-Sequence Neural Diarization (MC-S2SND), which refines the initial results using multi-channel audio. The final system achieves a diarization error rate (DER) of 8.09% on the evaluation set of the competition database, ranking first place in the speaker diarization task of the MISP 2025 Challenge.


【10】 Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency  Transcription and Detection

标题: Dysfluid WFST:Zero-Shot语音不流利转录和检测的框架
链接:https://arxiv.org/abs/2505.16351
作者: Chenxu Guo,  Jiachen Lian,  Xuanru Zhou,  Jinming Zhang,  Shuhe Li,  Zongli Ye,  Hwi Joo Park,  Anaisha Das,  Zoe Ezzes,  Jet Vonk,  Brittany Morin,  Rian Bogley,  Lisa Wauters,  Zachary Miller,  Maria Gorno-Tempini,  Gopala Anumanchipalli 
备注:None
摘要:语音不流利的自动检测有助于语音语言病理学家有效地转录紊乱的语音,增强诊断和治疗计划。传统的方法,往往局限于分类,提供不足的临床洞察力,和文本无关的模型错误分类不流畅,特别是在上下文相关的情况下。这项工作介绍了Dysfluent-WFST,一个zero-shot解码器,同时转录音素和检测不流利。与以前的模型不同,Dysfluent-WFST与WavLM等上游编码器一起运行,不需要额外的训练。它实现了最先进的性能,在语音错误率和不流利检测模拟和真实的语音数据。我们的方法是轻量级的,可解释的,有效的,表明在解码发音行为的明确建模,而不是复杂的架构,是改善不流利处理系统的关键。
摘要:Automatic detection of speech dysfluency aids speech-language pathologists in efficient transcription of disordered speech, enhancing diagnostics and treatment planning. Traditional methods, often limited to classification, provide insufficient clinical insight, and text-independent models misclassify dysfluency, especially in context-dependent cases. This work introduces Dysfluent-WFST, a zero-shot decoder that simultaneously transcribes phonemes and detects dysfluency. Unlike previous models, Dysfluent-WFST operates with upstream encoders like WavLM and requires no additional training. It achieves state-of-the-art performance in both phonetic error rate and dysfluency detection on simulated and real speech data. Our approach is lightweight, interpretable, and effective, demonstrating that explicit modeling of pronunciation behavior in decoding, rather than complex architectures, is key to improving dysfluency processing systems.


【11】 Meta-PerSER: Few-Shot Listener Personalized Speech Emotion Recognition  via Meta-learning

标题: Meta-Perser:Few-Shot预设通过元学习的个性化语音情感识别
链接:https://arxiv.org/abs/2505.16220
作者: Liang-Yeh Shen,  Shi-Xin Fang,  Yi-Cheng Lin,  Huang-Cheng Chou,  Hung-yi Lee 
备注:Accepted by INTERSPEECH 2025. 7 pages, including 2 pages of appendix
摘要:本文介绍了Meta-PerSER,一种新的元学习框架,个性化的语音情感识别(SER),适应每个听众的独特的方式来解释情感。传统的SER系统依赖于聚合的注释,这通常会忽略个别的微妙之处,并导致不一致的预测。相比之下,Meta-PerSER利用模型不可知元学习(MAML)方法,通过组合集元训练,导数退火和每层每步学习率进行增强,仅用几个标记的示例即可实现快速适应。通过整合来自预训练的自监督模型的鲁棒表示,我们的框架首先捕获一般的情感线索,然后根据个人注释风格进行微调。在IEMOCAP语料库上的实验表明,Meta-PerSER在可见和不可见数据场景中的性能明显优于基线方法,突出了其对个性化情感识别的承诺。
摘要:This paper introduces Meta-PerSER, a novel meta-learning framework that personalizes Speech Emotion Recognition (SER) by adapting to each listener's unique way of interpreting emotion. Conventional SER systems rely on aggregated annotations, which often overlook individual subtleties and lead to inconsistent predictions. In contrast, Meta-PerSER leverages a Model-Agnostic Meta-Learning (MAML) approach enhanced with Combined-Set Meta-Training, Derivative Annealing, and per-layer per-step learning rates, enabling rapid adaptation with only a few labeled examples. By integrating robust representations from pre-trained self-supervised models, our framework first captures general emotional cues and then fine-tunes itself to personal annotation styles. Experiments on the IEMOCAP corpus demonstrate that Meta-PerSER significantly outperforms baseline methods in both seen and unseen data scenarios, highlighting its promise for personalized emotion recognition.


【12】 AudioMorphix: Training-free audio editing with diffusion probabilistic  models

标题: AudioMorphix:使用扩散概率模型进行免训练音频编辑
链接:https://arxiv.org/abs/2505.16076
作者: Jinhua Liang,  Yuanzhe Chen,  Yi Yuan,  Dongya Jia,  Xiaobin Zhuang,  Zhuo Chen,  Yuping Wang,  Yuxuan Wang 
摘要:精确编辑声音是音频内容创作中一个至关重要但尚未探索的挑战。虽然现有的作品可以通过文本指令或音频样本对来操纵声音,但它们往往难以精确地修改音频内容,同时保持对原始录音的保真度。在这项工作中,我们介绍了一种新的编辑方法,该方法可以对特定的时频区域进行局部修改,同时通过直接对频谱图进行操作来保持音频的其余部分不变。为了实现这一目标,我们提出了AudioMorphix,一个免训练的音频编辑器,它通过参考另一个录音来操纵声谱图上的目标区域。受变形理论的启发,我们将混音概念化为一个过程,在这个过程中,不同的声音通过变形无缝融合,并可以通过变形分解回单独的分量。我们的AudioMorphix优化了以原始输入和参考音频为条件的噪声潜伏,同时通过一系列能量函数纠正了引导扩散过程。此外,我们增强了自我注意层的缓存机制,以保留原始记录的详细特征。为了推进音频编辑研究,我们设计了一个新的评估基准,其中包括一个具有各种编辑指令的精选数据集。大量的实验表明,AudioMorphix在各种音频编辑任务上都有很好的表现,包括添加、删除、时间平移和拉伸以及音高平移,实现了高保真度和精确度。演示和代码可以在这个URL上找到。
摘要:Editing sound with precision is a crucial yet underexplored challenge in audio content creation. While existing works can manipulate sounds by text instructions or audio exemplar pairs, they often struggled to modify audio content precisely while preserving fidelity to the original recording. In this work, we introduce a novel editing approach that enables localized modifications to specific time-frequency regions while keeping the remaining of the audio intact by operating on spectrograms directly. To achieve this, we propose AudioMorphix, a training-free audio editor that manipulates a target region on the spectrogram by referring to another recording. Inspired by morphing theory, we conceptualize audio mixing as a process where different sounds blend seamlessly through morphing and can be decomposed back into individual components via demorphing. Our AudioMorphix optimizes the noised latent conditioned on raw input and reference audio while rectifying the guided diffusion process through a series of energy functions. Additionally, we enhance self-attention layers with a cache mechanism to preserve detailed characteristics from the original recordings. To advance audio editing research, we devise a new evaluation benchmark, which includes a curated dataset with a variety of editing instructions. Extensive experiments demonstrate that AudioMorphix yields promising performance on various audio editing tasks, including addition, removal, time shifting and stretching, and pitch shifting, achieving high fidelity and precision. Demo and code are available at this url.


【13】 Multimodal Biomarkers for Schizophrenia: Towards Individual Symptom  Severity Estimation

标题: 精神分裂症的多模态生物标志物:个体症状严重程度估计
链接:https://arxiv.org/abs/2505.16044
作者: Gowtham Premananth,  Philip Resnik,  Sonia Bansal,  Deanna L.Kelly,  Carol Espy-Wilson 
备注:Accepted to be presented at Interspeech 2025
摘要:使用深度学习进行精神分裂症评估的研究通常将其视为一项分类任务,以检测疾病的存在或不存在,过度简化了病情并降低了其临床适用性。这种传统的方法忽视了精神分裂症的复杂性,限制了其在医疗保健环境中的实用价值。本研究将重点转移到个人症状的严重程度估计使用多模态的方法,集成语音,视频和文本输入。我们为每种模态和多模态框架开发了单峰模型,以提高准确性和鲁棒性。通过捕获更详细的症状特征,这种方法可以帮助提高诊断精度并支持个性化治疗,为心理健康评估提供可扩展的客观工具。
摘要:Studies on schizophrenia assessments using deep learning typically treat it as a classification task to detect the presence or absence of the disorder, oversimplifying the condition and reducing its clinical applicability. This traditional approach overlooks the complexity of schizophrenia, limiting its practical value in healthcare settings. This study shifts the focus to individual symptom severity estimation using a multimodal approach that integrates speech, video, and text inputs. We develop unimodal models for each modality and a multimodal framework to improve accuracy and robustness. By capturing a more detailed symptom profile, this approach can help in enhancing diagnostic precision and support personalized treatment, offering a scalable and objective tool for mental health assessment.


【14】 Analyzing the Impact of Accent on English Speech: Acoustic and  Articulatory Perspectives

标题: 分析口音对英语演讲的影响:声学和发音的角度
链接:https://arxiv.org/abs/2505.15965
作者: Gowtham Premananth,  Vinith Kugathasan,  Carol Espy-Wilson 
备注:Accepted to be presented at Interspeech 2025
摘要:人工智能驱动的基于语音的应用程序的进步已经改变了从医疗保健到客户服务的各个行业。然而,全球交互中非母语口音语音的日益流行对语音处理系统提出了重大挑战,这些系统通常在以母语语音为主的数据集上进行训练。本研究通过发音和声学分析来研究口音英语语音,发现口音英语语音的协调模式比本族语语音简单,平均音高比本族语语音高。利用本征谱和基于声道变量的协调特征,我们建立了一种有效的方法来量化口音强度,而不依赖于资源密集型的语音翻译。我们的研究结果为研究口音对语音清晰度的影响提供了新的途径,并为开发包容性强、功能强大的语音处理系统提供了见解,这些系统可以适应不同的语言社区。
摘要:Advancements in AI-driven speech-based applications have transformed diverse industries ranging from healthcare to customer service. However, the increasing prevalence of non-native accented speech in global interactions poses significant challenges for speech-processing systems, which are often trained on datasets dominated by native speech. This study investigates accented English speech through articulatory and acoustic analysis, identifying simpler coordination patterns and higher average pitch than native speech. Using eigenspectra and Vocal Tract Variable-based coordination features, we establish an efficient method for quantifying accent strength without relying on resource-intensive phonetic transcriptions. Our findings provide a new avenue for research on the impacts of accents on speech intelligibility and offer insights for developing inclusive, robust speech processing systems that accommodate diverse linguistic communities.


【15】 Towards Holistic Evaluation of Large Audio-Language Models: A  Comprehensive Survey

标题: 大型音频语言模型的整体评估:全面调查
链接:https://arxiv.org/abs/2505.15957
作者: Chih-Kai Yang,  Neo S. Ho,  Hung-yi Lee 
备注:Project Website: this https URL
摘要:随着大型音频语言模型(LALM)的进步,这些模型增强了具有听觉能力的大型语言模型(LLM),预计这些模型将在各种听觉任务中表现出普遍的熟练程度。虽然已经出现了许多基准来评估LALM的性能,但它们仍然是零散的,缺乏结构化的分类。为了弥补这一差距,我们进行了一项全面的调查,并提出了一个系统的分类法LALM评估,分为四个方面的基础上,他们的目标:(1)一般的听觉意识和处理,(2)知识和推理,(3)对话导向的能力,(4)公平,安全和可信度。我们提供了每个类别的详细概述,并强调了该领域的挑战,为有前途的未来方向提供了见解。据我们所知,这是第一次专门针对LALM评估的调查,为社区提供了明确的指导方针。我们将发布调查论文集,并积极维护它,以支持该领域的持续进步。
摘要:With advancements in large audio-language models (LALMs), which enhance large language models (LLMs) with auditory capabilities, these models are expected to demonstrate universal proficiency across various auditory tasks. While numerous benchmarks have emerged to assess LALMs' performance, they remain fragmented and lack a structured taxonomy. To bridge this gap, we conduct a comprehensive survey and propose a systematic taxonomy for LALM evaluations, categorizing them into four dimensions based on their objectives: (1) General Auditory Awareness and Processing, (2) Knowledge and Reasoning, (3) Dialogue-oriented Ability, and (4) Fairness, Safety, and Trustworthiness. We provide detailed overviews within each category and highlight challenges in this field, offering insights into promising future directions. To the best of our knowledge, this is the first survey specifically focused on the evaluations of LALMs, providing clear guidelines for the community. We will release the collection of the surveyed papers and actively maintain it to support ongoing advancements in the field.


【16】 ASVspoof2019 vs. ASVspoof5: Assessment and Comparison

标题: ASVspoof 2019与ASVspoof 5:评估和比较
链接:https://arxiv.org/abs/2505.15911
作者: Avishai Weizman,  Yehuda Ben-Shimol,  Itshak Lapidot 
备注:5 pages, 3 figures. Accepted to Interspeech 2025 Conference
摘要:ASVspoof挑战旨在促进对欺骗语音攻击的理解,并鼓励开发强大的对策系统。这些挑战提供了一个标准化的数据库,用于评估和比较欺骗鲁棒的自动说话人验证解决方案。与ASVspoof2019相比,ASVspoof5挑战引入了数据库条件的转变。虽然ASVspoof2019仅在评估集中的欺骗攻击中存在不匹配条件,但ASVspoof5在真实和欺骗的语音统计中都存在不匹配。本文探讨了这些不匹配的影响,提出了定性和定量比较内和两个数据库之间。我们展示了真实和欺骗性语音的难度增加,并证明在ASVspoof5中,不仅攻击更具挑战性,而且与ASVspoof2019相比,真实语音也转向了欺骗性语音。
摘要:ASVspoof challenges are designed to advance the understanding of spoofing speech attacks and encourage the development of robust countermeasure systems. These challenges provide a standardized database for assessing and comparing spoofing-robust automatic speaker verification solutions. The ASVspoof5 challenge introduces a shift in database conditions compared to ASVspoof2019. While ASVspoof2019 has mismatched conditions only in spoofing attacks in the evaluation set, ASVspoof5 incorporates mismatches in both bona fide and spoofed speech statistics. This paper examines the impact of these mismatches, presenting qualitative and quantitative comparisons within and between the two databases. We show the increased difficulty for genuine and spoofed speech and demonstrate that in ASVspoof5, not only are the attacks more challenging, but the genuine speech also shifts toward spoofed speech compared to ASVspoof2019.


【17】 From Tens of Hours to Tens of Thousands: Scaling Back-Translation for  Speech Recognition

标题: 从数十小时到数万小时:语音识别的回缩翻译
链接:https://arxiv.org/abs/2505.16972
作者: Tianduo Wang,  Lu Xu,  Wei Lu,  Shanbo Cheng 
摘要:自动语音识别(ASR)的最新进展在很大程度上是由大量的语音语料库推动的。然而,在资源有限的情况下将覆盖面扩大到各种语文仍然是一项艰巨的挑战。本文介绍了Speech Back-Translation,这是一个可扩展的管道,通过现成的文本到语音(TTS)模型将大规模文本语料库转换为合成语音来改进多语言ASR模型。我们证明,只有几十个小时的真实转录语音可以有效地训练TTS模型生成合成语音在数百倍的原始音量,同时保持高质量。为了评估合成语音质量,我们开发了一个基于可懂度的评估框架,并为合成数据何时有利于ASR训练建立了明确的阈值。使用Speech Back-Translation,我们以10种语言生成了超过500,000小时的合成语音,并继续预训练Whisper-large-v3,实现了超过30%的平均转录错误减少。这些结果突出了语音回译增强多语言ASR系统的可扩展性和有效性。
摘要:Recent advances in Automatic Speech Recognition (ASR) have been largely fueled by massive speech corpora. However, extending coverage to diverse languages with limited resources remains a formidable challenge. This paper introduces Speech Back-Translation, a scalable pipeline that improves multilingual ASR models by converting large-scale text corpora into synthetic speech via off-the-shelf text-to-speech (TTS) models. We demonstrate that just tens of hours of real transcribed speech can effectively train TTS models to generate synthetic speech at hundreds of times the original volume while maintaining high quality. To evaluate synthetic speech quality, we develop an intelligibility-based assessment framework and establish clear thresholds for when synthetic data benefits ASR training. Using Speech Back-Translation, we generate more than 500,000 hours of synthetic speech in ten languages and continue pre-training Whisper-large-v3, achieving average transcription error reductions of over 30\%. These results highlight the scalability and effectiveness of Speech Back-Translation for enhancing multilingual ASR systems.


【18】 On the reliability of feature attribution methods for speech  classification

标题: 语音分类特征归因方法的可靠性
链接:https://arxiv.org/abs/2505.16406
作者: Gaofei Shen,  Hosein Mohebbi,  Arianna Bisazza,  Afra Alishahi,  Grzegorz Chrupała 
摘要:随着大规模预训练模型能力的发展,理解其输出的决定因素变得更加重要。特征属性旨在揭示输入元素的哪些部分对模型输出的贡献最大。在语音处理中,输入信号的独特性使得特征属性方法的应用具有挑战性。我们研究了输入类型和聚合和扰动时间跨度等因素如何影响标准特征归因方法的可靠性,以及这些因素如何与每个分类任务的特征相互作用。我们发现,标准的方法来功能归属一般是不可靠的,当应用到语音域,与字对齐的扰动方法时,适用于基于字的分类任务的例外。
摘要:As the capabilities of large-scale pre-trained models evolve, understanding the determinants of their outputs becomes more important. Feature attribution aims to reveal which parts of the input elements contribute the most to model outputs. In speech processing, the unique characteristics of the input signal make the application of feature attribution methods challenging. We study how factors such as input type and aggregation and perturbation timespan impact the reliability of standard feature attribution methods, and how these factors interact with characteristics of each classification task. We find that standard approaches to feature attribution are generally unreliable when applied to the speech domain, with the exception of word-aligned perturbation methods when applied to word-based classification tasks.


【19】 Large Language Models based ASR Error Correction for Child Conversations

标题: 基于大型语言模型的儿童对话的ASB错误纠正
链接:https://arxiv.org/abs/2505.16212
作者: Anfeng Xu,  Tiantian Feng,  So Hyun Kim,  Somer Bishop,  Catherine Lord,  Shrikanth Narayanan 
备注:Accepted to Interspeech 2025
摘要:自动语音识别(ASR)最近取得了显著的进展,但准确地转录儿童的语音仍然是一个重大的挑战。大型语言模型(LLM)的最新发展在提高ASR的透明度方面显示出了希望。然而,他们的应用在儿童语音,包括会话场景的探索不足。在这项研究中,我们探讨了使用LLM在纠正会话儿童语音的ASR错误。我们通过实验证明了LLM的承诺和挑战,两个孩子的会话语音数据集与zero-shot和微调ASR输出。我们发现,虽然LLM有助于校正zero-shot ASR输出和微调的基于CTC的ASR输出,但当结合上下文信息或使用微调的自回归ASR(例如,Whisper)输出。
摘要:Automatic Speech Recognition (ASR) has recently shown remarkable progress, but accurately transcribing children's speech remains a significant challenge. Recent developments in Large Language Models (LLMs) have shown promise in improving ASR transcriptions. However, their applications in child speech including conversational scenarios are underexplored. In this study, we explore the use of LLMs in correcting ASR errors for conversational child speech. We demonstrate the promises and challenges of LLMs through experiments on two children's conversational speech datasets with both zero-shot and fine-tuned ASR outputs. We find that while LLMs are helpful in correcting zero-shot ASR outputs and fine-tuned CTC-based ASR outputs, it remains challenging for LLMs to improve ASR performance when incorporating contextual information or when using fine-tuned autoregressive ASR (e.g., Whisper) outputs.


【20】 AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large  Language Models

标题: AudioTrust:对音频大型语言模型的多方面可信度进行基准测试
链接:https://arxiv.org/abs/2505.16211
作者: Kai Li,  Can Shen,  Yile Liu,  Jirui Han,  Kelong Zheng,  Xuechao Zou,  Zhe Wang,  Xingjian Du,  Shun Zhang,  Hanjun Luo,  Yingbin Jin,  Xinxin Xing,  Ziyang Ma,  Yue Liu,  Xiaojun Jia,  Yifan Zhang,  Junfeng Fang,  Kun Wang,  Yibo Yan,  Haoyang Li,  Yiming Li,  Xiaobin Zhuang,  Yang Liu,  Haibo Hu,  Zhuo Chen,  Zhizheng Wu,  Xiaolin Hu,  Eng-Siong Chng,  XiaoFeng Wang,  Wenyuan Xu,  Wei Dong,  Xinfeng Li 
备注:Technical Report
摘要:音频大语言模型(ALLM)的快速发展和不断扩大的应用要求严格理解其可信度。然而,评估这些模型的系统研究,特别是关于音频模式特有的风险,在很大程度上仍然是未开发的。现有的评价框架主要侧重于文本模态或仅涉及有限的一组安全维度,未能充分考虑音频模态固有的独特特征和应用场景。我们引入AudioTrust-第一个专门为ALLM设计的多方面可信度评估框架和基准。AudioTrust促进了六个关键维度的评估:公平性,幻觉,安全性,隐私性,鲁棒性和身份验证。为了全面评估这些维度,AudioTrust围绕18个不同的实验设置进行了构建。它的核心是一个精心构建的数据集,包含超过4,420个音频/文本样本,来自真实世界的场景(例如,日常对话、紧急呼叫、语音助理交互),专门用于探测ALLM的多方面可信度。在评估方面,该基准精心设计了9个音频特定的评估指标,我们采用了一个大规模的自动化管道,对模型输出进行客观和可扩展的评分。实验结果揭示了当前最先进的开源和闭源ALLM在面对各种高风险音频场景时的可信度边界和局限性,为未来音频模型的安全和可信部署提供了有价值的见解。我们的平台和基准可在https://github.com/JusperLee/AudioTrust上获得。
摘要:The rapid advancement and expanding applications of Audio Large Language Models (ALLMs) demand a rigorous understanding of their trustworthiness. However, systematic research on evaluating these models, particularly concerning risks unique to the audio modality, remains largely unexplored. Existing evaluation frameworks primarily focus on the text modality or address only a restricted set of safety dimensions, failing to adequately account for the unique characteristics and application scenarios inherent to the audio modality. We introduce AudioTrust-the first multifaceted trustworthiness evaluation framework and benchmark specifically designed for ALLMs. AudioTrust facilitates assessments across six key dimensions: fairness, hallucination, safety, privacy, robustness, and authentication. To comprehensively evaluate these dimensions, AudioTrust is structured around 18 distinct experimental setups. Its core is a meticulously constructed dataset of over 4,420 audio/text samples, drawn from real-world scenarios (e.g., daily conversations, emergency calls, voice assistant interactions), specifically designed to probe the multifaceted trustworthiness of ALLMs. For assessment, the benchmark carefully designs 9 audio-specific evaluation metrics, and we employ a large-scale automated pipeline for objective and scalable scoring of model outputs. Experimental results reveal the trustworthiness boundaries and limitations of current state-of-the-art open-source and closed-source ALLMs when confronted with various high-risk audio scenarios, offering valuable insights for the secure and trustworthy deployment of future audio models. Our platform and benchmark are available at https://github.com/JusperLee/AudioTrust.


【21】 Differentiable K-means for Fully-optimized Discrete Token-based ASR

标题: 完全优化的基于离散令牌的ASB的差异K均值
链接:https://arxiv.org/abs/2505.16207
作者: Kentaro Onda,  Yosuke Kashiwagi,  Emiru Tsunoo,  Hayato Futami,  Shinji Watanabe 
备注:Accepted by Interspeech2025
摘要:最近的研究强调了从自监督学习(SSL)模型中获得的离散令牌在各种语音相关任务中的潜力。这些令牌不仅可以作为语言建模中文本的替代品,还可以作为自动语音识别(ASR)等任务的中间表示。然而,离散令牌通常是通过独立于下游任务的SSL特征的k均值聚类获得的,这使得它们对于特定应用而言不是最佳的。本文提出了使用可微k-means,实现标记化和下游任务的联合优化。这种方法能够微调SSL参数和多个SSL层输出的学习权重。以ASR作为下游任务进行实验。由于优化的令牌,ASR准确性成功提高。所获得的令牌也表现出更大的语音信息的纯度,这被发现是有用的,即使在语音再合成。
摘要:Recent studies have highlighted the potential of discrete tokens derived from self-supervised learning (SSL) models for various speech-related tasks. These tokens serve not only as substitutes for text in language modeling but also as intermediate representations for tasks such as automatic speech recognition (ASR). However, discrete tokens are typically obtained via k-means clustering of SSL features independently of downstream tasks, making them suboptimal for specific applications. This paper proposes the use of differentiable k-means, enabling the joint optimization of tokenization and downstream tasks. This approach enables the fine-tuning of the SSL parameters and learning weights for outputs from multiple SSL layers. Experiments were conducted with ASR as a downstream task. ASR accuracy successfully improved owing to the optimized tokens. The acquired tokens also exhibited greater purity of phonetic information, which were found to be useful even in speech resynthesis.


【22】 SpecMaskFoley: Steering Pretrained Spectral Masked Generative  Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet

标题: SpecMaskFoley:引导预训练的频谱屏蔽生成Transformer器通过控制Net实现同步视频到音频合成
链接:https://arxiv.org/abs/2505.16195
作者: Zhi Zhong,  Akira Takahashi,  Shuyang Cui,  Keisuke Toyama,  Shusuke Takahashi,  Yuki Mitsufuji 
备注:4 pages, 2 figures, 2 tables. Demo page: this https URL
摘要:Foley合成旨在合成高质量的音频,其在语义和时间上都与视频帧对齐。鉴于其在创意产业中的广泛应用,该任务在研究界得到了越来越多的关注。为了避免从头开始训练音频生成模型的非平凡任务,将预训练的音频生成模型用于视频同步的Foley合成呈现出一个有吸引力的方向。ControlNet是一种将细粒度控件添加到预训练生成模型的方法,已被应用于Foley合成,但其使用仅限于手工制作的人类可读的时间条件。相比之下,从头开始的模型通过利用使用预训练的视频编码器提取的高维深度特征取得了成功。我们已经观察到基于ControlNet和从头开始的Foley模型之间的性能差距。为了缩小这一差距,我们提出了SpecMaskFoley,这是一种通过ControlNet将预训练的SpecMaskGIT模型转向视频同步Foley合成的方法。为了释放单个ControlNet分支的潜力,我们通过频率感知的时间特征对齐器解决了时间视频特征与预训练SpecMaskGIT的时间-频率性质之间的差异,从而消除了对现有技术中广泛使用的复杂调节机制的需求。对一个常见的福利合成基准的评估表明,SpecMaskFoley甚至可以超越强大的从头开始的基线,大大推进了基于ControlNet的福利合成模型的发展。演示页面:https://zzaudio.github.io/SpecMaskFoley_Demo/
摘要:Foley synthesis aims to synthesize high-quality audio that is both semantically and temporally aligned with video frames. Given its broad application in creative industries, the task has gained increasing attention in the research community. To avoid the non-trivial task of training audio generative models from scratch, adapting pretrained audio generative models for video-synchronized foley synthesis presents an attractive direction. ControlNet, a method for adding fine-grained controls to pretrained generative models, has been applied to foley synthesis, but its use has been limited to handcrafted human-readable temporal conditions. In contrast, from-scratch models achieved success by leveraging high-dimensional deep features extracted using pretrained video encoders. We have observed a performance gap between ControlNet-based and from-scratch foley models. To narrow this gap, we propose SpecMaskFoley, a method that steers the pretrained SpecMaskGIT model toward video-synchronized foley synthesis via ControlNet. To unlock the potential of a single ControlNet branch, we resolve the discrepancy between the temporal video features and the time-frequency nature of the pretrained SpecMaskGIT via a frequency-aware temporal feature aligner, eliminating the need for complicated conditioning mechanisms widely used in prior arts. Evaluations on a common foley synthesis benchmark demonstrate that SpecMaskFoley could even outperform strong from-scratch baselines, substantially advancing the development of ControlNet-based foley synthesis models. Demo page: https://zzaudio.github.io/SpecMaskFoley_Demo/


【23】 Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based  Resynthesis Only with Native Speech Corpora

标题: 仅使用母语语音库通过基于离散令牌的再合成来增强的外国口音模拟
链接:https://arxiv.org/abs/2505.16191
作者: Kentaro Onda,  Keisuke Imoto,  Satoru Fukayama,  Daisuke Saito,  Nobuaki Minematsu 
备注:Accepted by Interspeech2025
摘要:最近,提出了一种仅使用从自监督学习(SSL)模型获得的离散令牌与母语语音数据合成外国口音语音的方法。考虑到口音语音数据的有限性,该方法有望使模拟外国口音变得更加容易。通过将合成的带口音语音作为人类的听力材料或自动语音识别(ASR)的训练数据,两者都将获得对外国口音的更高鲁棒性。然而,以前的方法有一个致命的缺陷,它不能再现持续时间相关的口音。持续性口音通常出现在母语为音节计时或mora-timed节奏的第二语言使用者说重音计时语言时,如英语。在本文中,我们集成了持续时间修改到以前的方法,以更准确地模拟外国口音。实验表明,该方法成功地复制了真实的L2语音中看到的持续性口音。
摘要:Recently, a method for synthesizing foreign-accented speech only with native speech data using discrete tokens obtained from self-supervised learning (SSL) models was proposed. Considering limited availability of accented speech data, this method is expected to make it much easier to simulate foreign accents. By using the synthesized accented speech as listening materials for humans or training data for automatic speech recognition (ASR), both of them will acquire higher robustness against foreign accents. However, the previous method has a fatal flaw that it cannot reproduce duration-related accents. Durational accents are commonly seen when L2 speakers, whose native language has syllable-timed or mora-timed rhythm, speak stress-timed languages, such as English. In this paper, we integrate duration modification to the previous method to simulate foreign accents more accurately. Experiments show that the proposed method successfully replicates durational accents seen in real L2 speech.


【24】 Discrete Tokens Exhibit Interlanguage Speech Intelligibility Benefit: an  Analytical Study Towards Accent-robust ASR Only with Native Speech Data

标题: 离散令牌展现出中介语语音可理解性优势:仅使用母语语音数据实现口音稳健的ASB的分析研究
链接:https://arxiv.org/abs/2505.16182
作者: Kentaro Onda,  Keisuke Imoto,  Satoru Fukayama,  Daisuke Saito,  Nobuaki Minematsu 
备注:Accepted by Interspeech2025
摘要:在这项研究中,我们获得的见解,有助于实现口音鲁棒ASR只使用母语数据。在人类对非母语语音的感知中,观察到被称为“中介语语音可理解性益处”(ISIB)的现象,其中与说话者共享母语的非母语听者甚至比母语听者更好地理解语音。基于从自监督学习(SSL)模型中提取的离散令牌代表人类对语音的感知的想法,我们对基于离散令牌的ASR对非母语语音的鲁棒性进行了分析研究,改变了用于训练令牌化的语言,这被视为ISIB的技术实现。结果表明,ISIB实际上发生在离散的基于令牌的ASR。由于我们的方法只依赖于本地语音数据来模拟人类感知的行为,预计将适用于广泛的口音,语音数据是稀缺的。
摘要:In this study, we gained insight that contributes to achieving accent-robust ASR using only native speech data. In human perception of non-native speech, the phenomenon known as "interlanguage speech intelligibility benefit" (ISIB) is observed, where non-native listeners who share the native language with the speaker understand the speech better compared even to native listeners. Based on the idea that discrete tokens extracted from self-supervised learning (SSL) models represent the human perception of speech, we conducted an analytical study on the robustness of discrete token-based ASR to non-native speech, varying the language used for training the tokenization, which is viewed as a technical implementation of ISIB. The results showed that ISIB actually occurred in the discrete token-based ASR. Since our approach relies only on native speech data to simulate the behavior of human perception, it is expected to be applicable to a wide range of accents for which speech data is scarce.


【25】 Selective Invocation for Multilingual ASR: A Cost-effective Approach  Adapting to Speech Recognition Difficulty

标题: 多语言ASB的选择性调用:一种适应语音识别困难的经济有效方法
链接:https://arxiv.org/abs/2505.16168
作者: Hongfei Xue,  Yufeng Tang,  Jun Zhang,  Xuelong Geng,  Lei Xie 
备注:Accepted by INTERSPEECH 2025
摘要:虽然多语言自动语音识别(ASR)系统已经取得了显著的进步,使单一模型能够处理多种语言,但固有的语言差异和数据不平衡挑战了SOTA在所有语言中的性能。虽然语言识别(LID)模型可以将语音路由到适当的ASR模型,但它们会因调用SOTA商业模型而产生高昂的成本,并且由于错误分类而导致不准确。为了克服这些问题,我们提出SIMA,一个选择性调用多语言ASR,适应输入语音的难度水平。SIMA建立在口语大语言模型(SLLM)上,它评估输入是否足够简单,可以直接转录,或者需要调用SOTA ASR模型。与SLLM相比,我们的方法将字错误率降低了18.7%,与基于LID的方法相比,调用成本降低了一半。三个数据集上的测试表明,SIMA是一个可扩展的,成本效益高的多语言ASR应用程序的解决方案。
摘要:Although multilingual automatic speech recognition (ASR) systems have significantly advanced, enabling a single model to handle multiple languages, inherent linguistic differences and data imbalances challenge SOTA performance across all languages. While language identification (LID) models can route speech to the appropriate ASR model, they incur high costs from invoking SOTA commercial models and suffer from inaccuracies due to misclassification. To overcome these, we propose SIMA, a selective invocation for multilingual ASR that adapts to the difficulty level of the input speech. Built on a spoken large language model (SLLM), SIMA evaluates whether the input is simple enough for direct transcription or requires the invocation of a SOTA ASR model. Our approach reduces word error rates by 18.7% compared to the SLLM and halves invocation costs compared to LID-based methods. Tests on three datasets show that SIMA is a scalable, cost-effective solution for multilingual ASR applications.


【26】 Source Separation by Flow Matching

标题: 通过流量匹配进行源分离
链接:https://arxiv.org/abs/2505.16119
作者: Robin Scheibler,  John R. Hershey,  Arnaud Doucet,  Henry Li 
备注:5 pages, 3 figures, 2 tables
摘要:我们考虑单声道音频源分离的问题,其目标是从它们的混合物中重建$K$个源。我们解决这个不适定的问题与FLOSS(流匹配源分离),基于流匹配的约束生成方法,确保严格的混合一致性。流匹配是一种通用方法,当给定来自定义在同一空间上的两个概率分布的样本时,学习常微分方程以在提供来自另一个分布的样本时输出来自其中一个分布的样本。在我们的上下文中,我们可以访问来自$K$源的联合分布的样本,因此相应的样本来自它们的混合物的低维分布。为了应用流匹配,我们用人工噪声分量增强这些混合样本,以确保得到的“增强”分布与$K$源分布的维度相匹配。此外,由于源的任何排列产生相同的混合物,我们采用了依赖于合适的定制设计的神经网络架构的流量匹配的等变公式。我们证明了重叠语音分离的方法的性能。
摘要:We consider the problem of single-channel audio source separation with the goal of reconstructing $K$ sources from their mixture. We address this ill-posed problem with FLOSS (FLOw matching for Source Separation), a constrained generation method based on flow matching, ensuring strict mixture consistency. Flow matching is a general methodology that, when given samples from two probability distributions defined on the same space, learns an ordinary differential equation to output a sample from one of the distributions when provided with a sample from the other. In our context, we have access to samples from the joint distribution of $K$ sources and so the corresponding samples from the lower-dimensional distribution of their mixture. To apply flow matching, we augment these mixture samples with artificial noise components to ensure the resulting "augmented" distribution matches the dimensionality of the $K$ source distribution. Additionally, as any permutation of the sources yields the same mixture, we adopt an equivariant formulation of flow matching which relies on a suitable custom-designed neural network architecture. We demonstrate the performance of the method for the separation of overlapping speech.


【27】 A Novel Deep Learning Framework for Efficient Multichannel Acoustic  Feedback Control

标题: 用于高效多通道声反馈控制的新型深度学习框架
链接:https://arxiv.org/abs/2505.15914
作者: Yuan-Kuei Wu,  Juan Azcarreta,  Kashyap Patel,  Buye Xu,  Jung-Suk Lee,  Sanha Lee,  Ashutosh Pandey 
备注:Accepted by Interspeech 2025
摘要:这项研究提出了一个深度学习框架,用于控制音频设备中的多通道声反馈。传统的数字信号处理方法在处理高度相关的噪声(如反馈)时难以收敛。我们引入了一个卷积递归网络,它有效地结合了空间和时间处理,大大提高了语音增强能力,降低了计算需求。我们的方法利用三种训练方法:在环训练,教师强迫,和一个混合策略与多通道维纳滤波器,优化性能在复杂的声学环境。这个可扩展的框架为现实世界的应用提供了一个强大的解决方案,使声反馈控制技术取得了重大进展。
摘要:This study presents a deep-learning framework for controlling multichannel acoustic feedback in audio devices. Traditional digital signal processing methods struggle with convergence when dealing with highly correlated noise such as feedback. We introduce a Convolutional Recurrent Network that efficiently combines spatial and temporal processing, significantly enhancing speech enhancement capabilities with lower computational demands. Our approach utilizes three training methods: In-a-Loop Training, Teacher Forcing, and a Hybrid strategy with a Multichannel Wiener Filter, optimizing performance in complex acoustic environments. This scalable framework offers a robust solution for real-world applications, making significant advances in Acoustic Feedback Control technology.


机器翻译由腾讯交互翻译提供,仅供参考