今日论文合集:cs.SD语音32篇,eess.AS音频处理38篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 Distilling a speech and music encoder with task arithmetic
标题: 使用任务算法提取语音和音乐编码器
链接:https://arxiv.org/abs/2505.13270
作者: Fabian Ritter-Gutierrez,  Yi-Cheng Lin,  Jui-Chiang Wei,  Jeremy H.M Wong,  Eng Siong Chng,  Nancy F. Chen,  Hung-yi Lee 
备注:Accepted at INTERSPEECH 2025
摘要:尽管语音和音乐的自监督学习(SSL)取得了进展,但现有模型将这些领域分开处理,限制了它们统一音频理解的能力。统一模型对于需要通用表示的应用是理想的,例如音频大语言模型。尽管如此,直接训练语音和音乐的通用模型在计算上是昂贵的。教师合奏的知识蒸馏可能是一个自然的解决方案,但我们认为,分离语音和音乐SSL模型的蒸馏允许更大的灵活性。因此,我们建议学习提取的任务向量,然后对其进行线性插值,以形成统一的语音+音乐模型。该策略通过可调权重实现灵活的域强调,并且训练也更简单。语音和音乐基准的实验表明,我们的方法产生优越的整体性能相比,集成蒸馏。
摘要:Despite the progress in self-supervised learning (SSL) for speech and music, existing models treat these domains separately, limiting their capacity for unified audio understanding. A unified model is desirable for applications that require general representations, e.g. audio large language models. Nonetheless, directly training a general model for speech and music is computationally expensive. Knowledge Distillation of teacher ensembles may be a natural solution, but we posit that decoupling the distillation of the speech and music SSL models allows for more flexibility. Thus, we propose to learn distilled task vectors and then linearly interpolate them to form a unified speech+music model. This strategy enables flexible domain emphasis through adjustable weights and is also simpler to train. Experiments on speech and music benchmarks demonstrate that our method yields superior overall performance compared to ensemble distillation.

【2】 Efficient Speech Language Modeling via Energy Distance in Continuous  Latent Space
标题: 连续潜在空间中通过能量距离进行高效语音语言建模
链接:https://arxiv.org/abs/2505.13181
作者: Zhengrui Ma,  Yang Feng,  Chenze Shao,  Fandong Meng,  Jie Zhou,  Min Zhang 
备注:Demos and code are available at this https URL
摘要:我们介绍SLED,语音语言建模的另一种方法,通过将语音波形编码成连续的潜在表示序列,并使用能量距离目标对其进行自回归建模。能量距离通过对比模拟样本和目标样本提供了分布差距的分析度量,从而实现有效的训练以捕获潜在的连续自回归分布。通过绕过对残差矢量量化的依赖,SLED避免了离散化错误,并消除了对现有语音语言模型中常见的复杂分层结构的需要。它简化了整个建模流程,同时保留了语音信息的丰富性并保持了推理效率。实验结果表明,SLED在zero-shot和流式语音合成中均取得了较好的性能,显示了其在通用语音语言模型中的广泛应用潜力。
摘要:We introduce SLED, an alternative approach to speech language modeling by encoding speech waveforms into sequences of continuous latent representations and modeling them autoregressively using an energy distance objective. The energy distance offers an analytical measure of the distributional gap by contrasting simulated and target samples, enabling efficient training to capture the underlying continuous autoregressive distribution. By bypassing reliance on residual vector quantization, SLED avoids discretization errors and eliminates the need for the complicated hierarchical architectures common in existing speech language models. It simplifies the overall modeling pipeline while preserving the richness of speech information and maintaining inference efficiency. Empirical results demonstrate that SLED achieves strong performance in both zero-shot and streaming speech synthesis, showing its potential for broader applications in general-purpose speech language models.

【3】 Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
标题: 用于时态推理的LALM的基准和置信度评估
链接:https://arxiv.org/abs/2505.13115
作者: Debarpan Bhattacharya,  Apoorva Kulkarni,  Sriram Ganapathy 
备注:Accepted in INTERSPEECH, 2025, Rotterdam, The Netherlands
摘要:基于文本的大型语言模型(LLM)的流行成功简化了多模态社区的注意力,将视觉和音频等其他模态与文本结合起来,以实现类似的多模态功能。在这一探索中,大型音频语言模型(LALM)必须评估推理相关的任务,这是不同于传统的分类或生成任务。为了实现这一目标,我们提出了一种新的数据集,称为时间推理评估音频(TREA)。   我们对开源LALM进行了基准测试,并观察到它们在TREA数据集中的任务上始终落后于人类能力。在评估LALM的同时,我们还提出了一个不确定性度量,该度量计算模型对输入的语义相同扰动的不变性。我们的分析表明,准确性和不确定性指标不一定是相关的,因此,点需要健康的评估LALM高风险的应用程序。
摘要:The popular success of text-based large language models (LLM) has streamlined the attention of the multimodal community to combine other modalities like vision and audio along with text to achieve similar multimodal capabilities. In this quest, large audio language models (LALMs) have to be evaluated on reasoning related tasks which are different from traditional classification or generation tasks. Towards this goal, we propose a novel dataset called temporal reasoning evaluation of audio (TREA).   We benchmark open-source LALMs and observe that they are consistently behind human capabilities on the tasks in the TREA dataset. While evaluating LALMs, we also propose an uncertainty metric, which computes the invariance of the model to semantically identical perturbations of the input. Our analysis shows that the accuracy and uncertainty metrics are not necessarily correlated and thus, points to a need for wholesome evaluation of LALMs for high-stakes applications.

【4】 Time-Frequency-Based Attention Cache Memory Model for Real-Time Speech  Separation
标题: 基于时频的实时语音分离注意力缓存模型
链接:https://arxiv.org/abs/2505.13094
作者: Guo Chen,  Kai Li,  Runxuan Yang,  Xiaolin Hu 
摘要:现有的因果语音分离模型往往表现不佳相比,非因果模型,由于在保留历史信息的困难。为了解决这个问题,我们提出了时间-频率注意力缓存(TFACM)模型,它有效地捕捉时空关系,通过注意力机制和缓存(CM)的历史信息存储。在TFACM中,LSTM层捕获频率相对位置,而因果建模则使用局部和全局表示应用于时间维度。CM模块存储过去的信息,并且因果注意细化(CAR)模块进一步增强基于时间的特征表示以获得更精细的粒度。实验结果表明,TFACM需要与SOTA TF-GridNet-Causal模型相当的性能,具有显着更低的复杂性和更少的可训练参数。有关更多详细信息,请访问项目页面:https://cslikai.cn/TFACM/。
摘要:Existing causal speech separation models often underperform compared to non-causal models due to difficulties in retaining historical information. To address this, we propose the Time-Frequency Attention Cache Memory (TFACM) model, which effectively captures spatio-temporal relationships through an attention mechanism and cache memory (CM) for historical information storage. In TFACM, an LSTM layer captures frequency-relative positions, while causal modeling is applied to the time dimension using local and global representations. The CM module stores past information, and the causal attention refinement (CAR) module further enhances time-based feature representations for finer granularity. Experimental results showed that TFACM achieveed comparable performance to the SOTA TF-GridNet-Causal model, with significantly lower complexity and fewer trainable parameters. For more details, visit the project page: https://cslikai.cn/TFACM/.

【5】 MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and  Voices of Multiple Speakers
标题: MultiActor有声读物:具有多个发言者的面孔和声音的Zero-Shot有声读物生成
链接:https://arxiv.org/abs/2505.13082
作者: Kyeongman Park,  Seongho Joo,  Kyomin Jung 
摘要:我们引入了MultiActor-Audiobook,这是一种用于生成有声读物的零拍摄(zero-shot)方法,可自动生成一致、富有表现力且适合说话者的韵律,包括语调和情感。以前的有声读物系统有几个局限性:它们需要用户手动配置说话者的韵律,与配音演员相比,用单调的音调阅读每个句子,或者依赖昂贵的培训。然而,我们的多演员有声读物通过引入两个新的过程来解决这些问题:(1)MSP(** 多模态扬声器角色生成 **)和(2)LSI(** 基于LLM的脚本指令生成 **)。通过这两个过程,MultiActor-Audiobook可以生成更具情感表达力的有声读物,具有一致的说话者韵律,而无需额外的训练。我们将我们的系统与商业产品进行比较,通过人体和MLLM评估,获得具有竞争力的结果。此外,我们通过消融研究证明了MSP和LSI的有效性。
摘要:We introduce MultiActor-Audiobook, a zero-shot approach for generating audiobooks that automatically produces consistent, expressive, and speaker-appropriate prosody, including intonation and emotion. Previous audiobook systems have several limitations: they require users to manually configure the speaker's prosody, read each sentence with a monotonic tone compared to voice actors, or rely on costly training. However, our MultiActor-Audiobook addresses these issues by introducing two novel processes: (1) MSP (**Multimodal Speaker Persona Generation**) and (2) LSI (**LLM-based Script Instruction Generation**). With these two processes, MultiActor-Audiobook can generate more emotionally expressive audiobooks with a consistent speaker prosody without additional training. We compare our system with commercial products, through human and MLLM evaluations, achieving competitive results. Furthermore, we demonstrate the effectiveness of MSP and LSI through ablation studies.

【6】 Suicide Risk Assessment Using Multimodal Speech Features: A Study on the  SW1 Challenge Dataset
标题: 使用多模式语音特征的自杀风险评估:SW 1挑战数据集的研究
链接:https://arxiv.org/abs/2505.13069
作者: Ambre Marie,  Ilias Maoudj,  Guillaume Dardenne,  Gwenolé Quellec 
备注:Submitted to the SpeechWellness Challenge at Interspeech 2025; 5 pages, 2 figures, 2 tables
摘要:第一届SpeechWellness Challenge传达了对青少年进行基于语言的自杀风险评估的必要性。本研究探讨了一种多模态方法来应对这一挑战,将自动转录与WhisperX、来自中国RoberTa的语言嵌入和来自WavLM的音频嵌入相结合。此外,手工制作的声学功能-包括MFCC,频谱对比度和音高相关的统计-被纳入。我们探索了三种融合策略:早期拼接,特定模态处理和混合正则化加权注意力。结果表明,加权注意力提供了最好的泛化能力,在开发集上达到了69%的准确率,尽管开发集和测试集之间的性能差距突出了泛化的挑战。我们的研究结果,严格绑定到MINI-KID框架,强调细化嵌入表示和融合机制,以提高分类可靠性的重要性。
摘要:The 1st SpeechWellness Challenge conveys the need for speech-based suicide risk assessment in adolescents. This study investigates a multimodal approach for this challenge, integrating automatic transcription with WhisperX, linguistic embeddings from Chinese RoBERTa, and audio embeddings from WavLM. Additionally, handcrafted acoustic features -- including MFCCs, spectral contrast, and pitch-related statistics -- were incorporated. We explored three fusion strategies: early concatenation, modality-specific processing, and weighted attention with mixup regularization. Results show that weighted attention provided the best generalization, achieving 69% accuracy on the development set, though a performance gap between development and test sets highlights generalization challenges. Our findings, strictly tied to the MINI-KID framework, emphasize the importance of refining embedding representations and fusion mechanisms to enhance classification reliability.

【7】 Hearing from Silence: Reasoning Audio Descriptions from Silent Videos  via Vision-Language Model
标题: 从沉默中倾听:通过视觉语言模型从无声视频中推理音频描述
链接:https://arxiv.org/abs/2505.13062
作者: Yong Ren,  Chenxing Li,  Le Xu,  Hao Gu,  Duzhen Zhang,  Yujie Chen,  Manjie Xu,  Ruibo Fu,  Shan Yang,  Dong Yu 
备注:Accepted by Interspeech 2025
摘要:人类可以从无声视频中直观地推断出声音,但多模态大型语言模型是否可以在不访问目标模态的情况下执行模态失配推理仍然相对未被探索。目前的文本辅助视频到音频(VT 2A)的方法在视频Foley任务中表现出色,但在推理过程中难以获得音频描述。我们引入了从无声视频中推理音频描述(SVAD)的任务来解决这一挑战,并研究了视觉语言模型(VLM)在这一任务上的能力。为了进一步增强VLM在SVAD任务中的推理能力,我们构建了一个CoT-AudioCaps数据集,并提出了一种基于权值链的监督微调策略。SVAD和随后的VT 2A任务的实验表明,我们的方法的有效性在两个关键方面:显着提高VLM的模态失配推理SVAD和有效地解决在VT 2A推理过程中获取音频描述的挑战。
摘要:Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remains relatively unexplored. Current text-assisted-video-to-audio (VT2A) methods excel in video foley tasks but struggle to acquire audio descriptions during inference. We introduce the task of Reasoning Audio Descriptions from Silent Videos (SVAD) to address this challenge and investigate vision-language models' (VLMs) capabilities on this task. To further enhance the VLMs' reasoning capacity for the SVAD task, we construct a CoT-AudioCaps dataset and propose a Chain-of-Thought-based supervised fine-tuning strategy. Experiments on SVAD and subsequent VT2A tasks demonstrate our method's effectiveness in two key aspects: significantly improving VLMs' modal-mismatch reasoning for SVAD and effectively addressing the challenge of acquiring audio descriptions during VT2A inference.

【8】 MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio,  Music, and Their Mix
标题: MVAR:语音、音频、音乐及其混合中深度推理的领先基准
链接:https://arxiv.org/abs/2505.13032
作者: Ziyang Ma,  Yinghao Ma,  Yanqiao Zhu,  Chen Yang,  Yi-Wen Chao,  Ruiyang Xu,  Wenxi Chen,  Yuanzhe Chen,  Zhuo Chen,  Jian Cong,  Kai Li,  Keliang Li,  Siyou Li,  Xinfeng Li,  Xiquan Li,  Zheng Lian,  Yuzhe Liang,  Minghao Liu,  Zhikang Niu,  Tianrui Wang,  Yuping Wang,  Yuxuan Wang,  Yihao Wu,  Guanrou Yang,  Jianwei Yu,  Ruibin Yuan,  Zhisheng Zheng,  Ziya Zhou,  Haina Zhu,  Wei Xue,  Emmanouil Benetos,  Kai Yu,  Eng-Siong Chng,  Xie Chen 
备注:Open-source at this https URL
摘要:我们引入了MMAR,这是一个新的基准测试,旨在评估音频语言模型(ALM)在大规模多学科任务中的深度推理能力。MMAR包含1,000个精心策划的音频问答三元组,从真实世界的互联网视频中收集,并通过迭代纠错和质量检查进行优化,以确保高质量。与现有的仅限于声音、音乐或语音的特定领域的基准测试不同,MMAR将它们扩展到广泛的真实世界音频场景,包括声音、音乐和语音的混合模态组合。MMAR中的每个问题都分为四个推理层:信号,感知,语义和文化,每个层中都有额外的子类别,以反映任务的多样性和复杂性。为了进一步促进这一领域的研究,我们用思想链(CoT)理论来注释每个问题,以促进音频推理的未来发展。基准测试中的每个项目都需要进行多步深入推理,而不仅仅是表面的理解。此外,部分问题需要研究生水平的感性和特定领域的知识,提高了基准的难度和深度。我们使用一组广泛的模型来评估MMAR,包括大型音频语言模型(LALM),大型音频推理模型(LARM),Omni语言模型(OLM),大型语言模型(LLM)和大型推理模型(LRM),并带有音频字幕输入。这些模型在MMAR上的表现突出了基准测试的挑战性,我们的分析进一步揭示了当前模型在理解和推理能力方面的关键局限性。我们希望MMAR将成为这一重要但很少探索的领域未来进步的催化剂。
摘要:We introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets, collected from real-world internet videos and refined through iterative error corrections and quality checks to ensure high quality. Unlike existing benchmarks that are limited to specific domains of sound, music, or speech, MMAR extends them to a broad spectrum of real-world audio scenarios, including mixed-modality combinations of sound, music, and speech. Each question in MMAR is hierarchically categorized across four reasoning layers: Signal, Perception, Semantic, and Cultural, with additional sub-categories within each layer to reflect task diversity and complexity. To further foster research in this area, we annotate every question with a Chain-of-Thought (CoT) rationale to promote future advancements in audio reasoning. Each item in the benchmark demands multi-step deep reasoning beyond surface-level understanding. Moreover, a part of the questions requires graduate-level perceptual and domain-specific knowledge, elevating the benchmark's difficulty and depth. We evaluate MMAR using a broad set of models, including Large Audio-Language Models (LALMs), Large Audio Reasoning Models (LARMs), Omni Language Models (OLMs), Large Language Models (LLMs), and Large Reasoning Models (LRMs), with audio caption inputs. The performance of these models on MMAR highlights the benchmark's challenging nature, and our analysis further reveals critical limitations of understanding and reasoning capabilities among current models. We hope MMAR will serve as a catalyst for future advances in this important but little-explored area.

【9】 DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec  for Speech Generation
标题: DualCodec:一种用于语音生成的低帧率、语义增强的神经音频编解码器
链接:https://arxiv.org/abs/2505.13000
作者: Jiaqi Li,  Xiaolong Lin,  Zhekai Li,  Shixi Huang,  Yuancheng Wang,  Chaoren Wang,  Zhenpeng Zhan,  Zhizheng Wu 
备注:Accepted to Interspeech 2025. Github: this https URL
摘要:神经音频编解码器形成基于语言模型(LM)的语音生成的基础构建块。通常,在帧速率和音频质量之间存在折衷。本研究提出一种低帧率、语意增强的编解码器模型。现有方法将语义丰富的自监督(SSL)表示提取到第一层编解码器令牌中。这项工作提出了DualCodec,一个双流编码方法,集成了SSL和波形表示在一个端到端的编解码器框架。在这种设置下,DualCodec增强了第一层编解码器中的语义信息,使编解码器系统能够在低帧速率下运行时保持高音频质量。注意,低帧速率编解码器提高了语音生成的效率。音频编解码器和语音生成任务的实验结果证实了所提出的DualCodec的有效性相比,国家的最先进的编解码器系统,如Mimi编解码器,SpeechTokenizer,DAC,和Encodec。演示和代码可在https://dualcodec.github.io上获得
摘要:Neural audio codecs form the foundational building blocks for language model (LM)-based speech generation. Typically, there is a trade-off between frame rate and audio quality. This study introduces a low-frame-rate, semantically enhanced codec model. Existing approaches distill semantically rich self-supervised (SSL) representations into the first-layer codec tokens. This work proposes DualCodec, a dual-stream encoding approach that integrates SSL and waveform representations within an end-to-end codec framework. In this setting, DualCodec enhances the semantic information in the first-layer codec and enables the codec system to maintain high audio quality while operating at a low frame rate. Note that a low-frame-rate codec improves the efficiency of speech generation. Experimental results on audio codec and speech generation tasks confirm the effectiveness of the proposed DualCodec compared to state-of-the-art codec systems, such as Mimi Codec, SpeechTokenizer, DAC, and Encodec. Demos and codes are available at: https://dualcodec.github.io

【10】 Codec-Based Deepfake Source Tracing via Neural Audio Codec Taxonomy
标题: 通过神经音频编解码分类法基于编解码器的Deepfake源跟踪
链接:https://arxiv.org/abs/2505.12994
作者: Xuanjun Chen,  I-Ming Lin,  Lin Zhang,  Jiawei Du,  Haibin Wu,  Hung-yi Lee,  Jyh-Shing Roger Jang 
备注:Accepted by Interspeech 2025
摘要:基于神经音频编解码器的语音生成(CoSG)模型的最新进展已经产生了非常逼真的音频deepfake。我们将CoSG系统生成的deepfake语音称为基于编解码器的deepfake或CodecFake。尽管现有的CodecFake反欺骗研究主要集中在验证音频样本的真实性上,但几乎没有注意跟踪用于生成这些deepfake的CoSG。在CodecFake生成中,语音到单元编码、离散单元建模和单元到语音解码等过程基本上基于神经音频编解码器。受此启发,我们通过神经音频编解码器分类法引入CodecFake的源跟踪,该分类法剖析神经音频编解码器以跟踪CoSG。我们在CodecFake+数据集上的实验结果为CodecFake源跟踪的可行性提供了有希望的初步证据,同时也强调了需要进一步研究的几个挑战。
摘要:Recent advances in neural audio codec-based speech generation (CoSG) models have produced remarkably realistic audio deepfakes. We refer to deepfake speech generated by CoSG systems as codec-based deepfake, or CodecFake. Although existing anti-spoofing research on CodecFake predominantly focuses on verifying the authenticity of audio samples, almost no attention was given to tracing the CoSG used in generating these deepfakes. In CodecFake generation, processes such as speech-to-unit encoding, discrete unit modeling, and unit-to-speech decoding are fundamentally based on neural audio codecs. Motivated by this, we introduce source tracing for CodecFake via neural audio codec taxonomy, which dissects neural audio codecs to trace CoSG. Our experimental results on the CodecFake+ dataset provide promising initial evidence for the feasibility of CodecFake source tracing while also highlighting several challenges that warrant further investigation.

【11】 Personalized Fine-Tuning with Controllable Synthetic Speech from  LLM-Generated Transcripts for Dysarthric Speech Recognition
标题: 利用LLM生成的脚本中的可控合成语音进行个性化微调,用于发音障碍语音识别
链接:https://arxiv.org/abs/2505.12991
作者: Dominik Wagner,  Ilja Baumann,  Natalie Engert,  Seanie Lee,  Elmar Nöth,  Korbinian Riedhammer,  Tobias Bocklet 
备注:Accepted at Interspeech 2025
摘要:在这项工作中,我们提出了我们提交的语音无障碍项目的挑战构音障碍语音识别。我们将参数有效的微调与潜在的音频表示相结合,以改善编码器-解码器ASR系统。合成训练数据是通过微调Parler-TTS来模拟构音障碍的语音,使用LLM生成的语料库一致的目标成绩单的提示。与非个性化微调相比,使用x向量的个性化始终降低了单词错误率(WER)。AdaLoRA适配器的性能优于完全微调和标准低秩适配,分别实现了约23%和约22%的相对WER降低。进一步的改进(约5%的WER降低)来自于整合基于wav 2 vec 2.0的音频表示。与单独的个性化微调相比,使用合成构音障碍语音的训练产生高达~7%的相对WER改善。
摘要:In this work, we present our submission to the Speech Accessibility Project challenge for dysarthric speech recognition. We integrate parameter-efficient fine-tuning with latent audio representations to improve an encoder-decoder ASR system. Synthetic training data is generated by fine-tuning Parler-TTS to mimic dysarthric speech, using LLM-generated prompts for corpus-consistent target transcripts. Personalization with x-vectors consistently reduces word error rates (WERs) over non-personalized fine-tuning. AdaLoRA adapters outperform full fine-tuning and standard low-rank adaptation, achieving relative WER reductions of ~23% and ~22%, respectively. Further improvements (~5% WER reduction) come from incorporating wav2vec 2.0-based audio representations. Training with synthetic dysarthric speech yields up to ~7% relative WER improvement over personalized fine-tuning alone.

【12】 The Computation of Generalized Embeddings for Underwater Acoustic Target  Recognition using Contrastive Learning
标题: 基于对比学习的水声目标识别中广义嵌入的计算
链接:https://arxiv.org/abs/2505.12904
作者: Hilde I. Hummel,  Arwin Gansekoele,  Sandjai Bhulai,  Rob van der Mei 
摘要:海洋环境中日益严重的声音污染对海洋健康构成越来越大的威胁,因此监测水下噪音至关重要。通过监测这种噪音,可以绘制出造成这种污染的来源。通过被动地听这些声音来执行监测。这会产生大量的数据记录,捕获混合的声源,如船舶活动和海洋哺乳动物的发声。虽然机器学习为自动声音分类提供了一个很有前途的解决方案,但目前最先进的方法实现了监督学习。这需要大量高质量的标记数据,这些数据不是公开的。相比之下,大量低质量的未标记数据是公开的,这为探索无监督学习技术提供了机会。本研究通过实施无监督对比学习方法来探索这种可能性。在这里,基于Conformer的编码器通过所谓的方差-不变性-协方差正则化损失函数对这些较低质量的未标记数据进行优化,并转换为标记数据。通过识别船舶类型和海洋哺乳动物发声的分类任务,我们的方法证明了产生强大的和广义的嵌入。这显示了各种自动水声分析任务的无监督方法的潜力。
摘要:The increasing level of sound pollution in marine environments poses an increased threat to ocean health, making it crucial to monitor underwater noise. By monitoring this noise, the sources responsible for this pollution can be mapped. Monitoring is performed by passively listening to these sounds. This generates a large amount of data records, capturing a mix of sound sources such as ship activities and marine mammal vocalizations. Although machine learning offers a promising solution for automatic sound classification, current state-of-the-art methods implement supervised learning. This requires a large amount of high-quality labeled data that is not publicly available. In contrast, a massive amount of lower-quality unlabeled data is publicly available, offering the opportunity to explore unsupervised learning techniques. This research explores this possibility by implementing an unsupervised Contrastive Learning approach. Here, a Conformer-based encoder is optimized by the so-called Variance-Invariance-Covariance Regularization loss function on these lower-quality unlabeled data and the translation to the labeled data is made. Through classification tasks involving recognizing ship types and marine mammal vocalizations, our method demonstrates to produce robust and generalized embeddings. This shows to potential of unsupervised methods for various automatic underwater acoustic analysis tasks.

【13】 Unified Cross-modal Translation of Score Images, Symbolic Music, and  Performance Audio
标题: 乐谱图像、象征性音乐和表演音频的统一跨模式翻译
链接:https://arxiv.org/abs/2505.12863
作者: Jongmin Jung,  Dongmin Kim,  Sihun Lee,  Seola Cho,  Hyungjoon Soh,  Irmak Bukey,  Chris Donahue,  Dasaem Jeong 
备注:Submitted to IEEE Transactions on Audio, Speech and Language Processing (TASLPRO)
摘要:音乐以各种形式存在,例如乐谱图像、符号乐谱、乐谱和音频。每种模态之间的翻译被确立为音乐信息检索的核心任务,例如自动音乐转录(音频到音频)和光学音乐识别(乐谱图像到符号乐谱)。然而,过去大多数关于多模态翻译的工作都是针对单个翻译任务训练专门的模型。在本文中,我们提出了一个统一的方法,我们训练一个通用的模型上的许多翻译任务同时进行。两个关键因素使这种统一的方法可行:一个新的大规模数据集和每个模态的标记化。首先,我们提出了一个新的数据集,该数据集由从YouTube视频中收集的超过1,300小时的配对音频图像数据组成,这比任何现有的音乐模态翻译数据集都大一个数量级。其次,我们的统一标记化框架将乐谱图像、音频、乐谱和MusicXML离散化为标记序列,使单个编码器-解码器Transformer能够将多个跨模态翻译作为一个连贯的序列到序列任务来处理。实验结果证实,我们的统一多任务模型在几个关键领域比单任务基线有所改进,特别是将光学音乐识别的符号错误率从24.58%降低到最先进的13.67%,而在其他翻译任务中也观察到了类似的实质性改进。值得注意的是,我们的方法实现了第一次成功的分数图像调节音频生成,标志着跨模态音乐生成的重大突破。
摘要:Music exists in various modalities, such as score images, symbolic scores, MIDI, and audio. Translations between each modality are established as core tasks of music information retrieval, such as automatic music transcription (audio-to-MIDI) and optical music recognition (score image to symbolic score). However, most past work on multimodal translation trains specialized models on individual translation tasks. In this paper, we propose a unified approach, where we train a general-purpose model on many translation tasks simultaneously. Two key factors make this unified approach viable: a new large-scale dataset and the tokenization of each modality. Firstly, we propose a new dataset that consists of more than 1,300 hours of paired audio-score image data collected from YouTube videos, which is an order of magnitude larger than any existing music modal translation datasets. Secondly, our unified tokenization framework discretizes score images, audio, MIDI, and MusicXML into a sequence of tokens, enabling a single encoder-decoder Transformer to tackle multiple cross-modal translation as one coherent sequence-to-sequence task. Experimental results confirm that our unified multitask model improves upon single-task baselines in several key areas, notably reducing the symbol error rate for optical music recognition from 24.58% to a state-of-the-art 13.67%, while similarly substantial improvements are observed across the other translation tasks. Notably, our approach achieves the first successful score-image-conditioned audio generation, marking a significant breakthrough in cross-modal music generation.

【14】 OZSpeech: One-step Zero-shot Speech Synthesis with  Learned-Prior-Conditioned Flow Matching
标题: OZSpeech:具有学习先验条件流匹配的一步零激发语音合成
链接:https://arxiv.org/abs/2505.12800
作者: Hieu-Nghia Huynh-Nguyen,  Ngoc Son Nguyen,  Huynh Nguyen Dang,  Thieu Vo,  Truong-Son Hy,  Van Nguyen 
摘要:近年来,在深度学习和神经网络架构的改进的推动下,文本到语音(TTS)系统取得了重大进展。将输出语音视为数据分布,以前的方法通常在流匹配框架内采用传统的语音表示,例如波形或频谱图。然而,这些方法具有局限性,包括忽略各种语音属性,以及由于在训练期间引入的额外约束而导致高计算成本。为了解决这些挑战,我们引入了OZSpeech,这是第一种TTS方法,可以探索最佳的传输条件流匹配,并以一步采样和先验知识为条件,有效地忽略了先前的状态并减少了采样步骤的数量。我们的方法在令牌格式的语音的分解,分解组件上操作,使得每个语音属性的准确建模成为可能,这增强了TTS系统精确克隆提示语音的能力。实验结果表明,我们的方法在内容准确性,自然度,韵律生成和说话人风格保持方面取得了令人满意的性能。音频样本可在我们的演示页面https://ozspeech.github.io/OZSpeech_Web/上获得。
摘要:Text-to-speech (TTS) systems have seen significant advancements in recent years, driven by improvements in deep learning and neural network architectures. Viewing the output speech as a data distribution, previous approaches often employ traditional speech representations, such as waveforms or spectrograms, within the Flow Matching framework. However, these methods have limitations, including overlooking various speech attributes and incurring high computational costs due to additional constraints introduced during training. To address these challenges, we introduce OZSpeech, the first TTS method to explore optimal transport conditional flow matching with one-step sampling and a learned prior as the condition, effectively disregarding preceding states and reducing the number of sampling steps. Our approach operates on disentangled, factorized components of speech in token format, enabling accurate modeling of each speech attribute, which enhances the TTS system's ability to precisely clone the prompt speech. Experimental results show that our method achieves promising performance over existing methods in content accuracy, naturalness, prosody generation, and speaker style preservation. Audio samples are available at our demo page https://ozspeech.github.io/OZSpeech_Web/.

【15】 SounDiT: Geo-Contextual Soundscape-to-Landscape Generation
标题: SoundDiT:地理上下文声景到景观生成
链接:https://arxiv.org/abs/2505.12734
作者: Junbo Wang,  Haofeng Tan,  Bowen Liao,  Albert Jiang,  Teng Fei,  Qixing Huang,  Zhengzhong Tu,  Shan Ye,  Yuhao Kang 
备注:14 pages, 5 figures
摘要:我们提出了一个新的和具有实际意义的问题-地理背景声景景观(GeoS 2L)的生成,其目的是合成地理现实的景观图像从环境音景。现有的音频到图像生成方法通常依赖于通用数据集并忽略地理和环境背景,从而导致与真实世界环境设置不一致的不真实图像。为了解决这个问题,我们引入了一个新的地理上下文计算框架,明确地将地理知识集成到多模态生成建模。我们构建了两个大规模的地理环境多模态数据集,SoundingSVI和SonicUrban,将不同的声音景观与真实世界的景观图像配对。我们提出了SoundDiT,一种新的扩散Transformer(DiT)为基础的模型,采用地理背景的场景条件合成地理上连贯的景观图像。此外,我们提出了一个实际知情的地理环境评估框架,地方相似性分数(PSS),跨元素,场景和人类感知水平,以衡量输入音景和生成的景观图像之间的一致性。大量的实验表明,SoundDiT在视觉保真度和地理环境方面都优于现有的基线。我们的工作不仅为GeoS 2L生成建立了基础基准,还强调了在推进多模态生成模型中融入地理领域知识的重要性,在生成人工智能,地理,城市规划和环境科学的交叉点上开辟了新的方向。
摘要:We present a novel and practically significant problem-Geo-Contextual Soundscape-to-Landscape (GeoS2L) generation-which aims to synthesize geographically realistic landscape images from environmental soundscapes. Prior audio-to-image generation methods typically rely on general-purpose datasets and overlook geographic and environmental contexts, resulting in unrealistic images that are misaligned with real-world environmental settings. To address this limitation, we introduce a novel geo-contextual computational framework that explicitly integrates geographic knowledge into multimodal generative modeling. We construct two large-scale geo-contextual multimodal datasets, SoundingSVI and SonicUrban, pairing diverse soundscapes with real-world landscape images. We propose SounDiT, a novel Diffusion Transformer (DiT)-based model that incorporates geo-contextual scene conditioning to synthesize geographically coherent landscape images. Furthermore, we propose a practically-informed geo-contextual evaluation framework, the Place Similarity Score (PSS), across element-, scene-, and human perception-levels to measure consistency between input soundscapes and generated landscape images. Extensive experiments demonstrate that SounDiT outperforms existing baselines in both visual fidelity and geographic settings. Our work not only establishes foundational benchmarks for GeoS2L generation but also highlights the importance of incorporating geographic domain knowledge in advancing multimodal generative models, opening new directions at the intersection of generative AI, geography, urban planning, and environmental sciences.

【16】 RoVo: Robust Voice Protection Against Unauthorized Speech Synthesis with  Embedding-Level Perturbations
标题: RoVo:针对未经授权的语音合成的强大语音保护,具有嵌入级扰动
链接:https://arxiv.org/abs/2505.12686
作者: Seungmin Kim,  Sohee Park,  Donghyun Kim,  Jisu Lee,  Daeseon Choi 
摘要:随着Deep Voice等基于人工智能的语音合成技术的进步,语音欺骗攻击的风险越来越大,包括未经授权使用他人语音的语音钓鱼和假新闻。现有的直接将对抗性扰动注入音频信号的防御措施效果有限,因为这些扰动可以很容易地通过语音增强方法来中和。为了克服这一限制,我们提出了RoVo(鲁棒语音),这是一种新型的主动防御技术,它将对抗性扰动注入音频信号的高维嵌入向量,将其重建为受保护的语音。这种方法有效地抵御语音合成攻击,也提供了强大的抵抗语音增强模型,这代表了二次攻击的威胁。   在广泛的实验中,RoVo在四种最先进的语音合成模型中,与无保护语音相比,防御成功率(DSR)提高了70%以上。具体来说,RoVo在商业说话者验证API上实现了99.5%的DSR,有效地中和了语音合成攻击。此外,RoVo的扰动即使在强语音增强条件下也保持鲁棒性,优于传统方法。一项用户研究证实,RoVo保留了受保护语音的自然性和可用性,突出了其在复杂和不断变化的威胁场景中的有效性。
摘要:With the advancement of AI-based speech synthesis technologies such as Deep Voice, there is an increasing risk of voice spoofing attacks, including voice phishing and fake news, through unauthorized use of others' voices. Existing defenses that inject adversarial perturbations directly into audio signals have limited effectiveness, as these perturbations can easily be neutralized by speech enhancement methods. To overcome this limitation, we propose RoVo (Robust Voice), a novel proactive defense technique that injects adversarial perturbations into high-dimensional embedding vectors of audio signals, reconstructing them into protected speech. This approach effectively defends against speech synthesis attacks and also provides strong resistance to speech enhancement models, which represent a secondary attack threat.   In extensive experiments, RoVo increased the Defense Success Rate (DSR) by over 70% compared to unprotected speech, across four state-of-the-art speech synthesis models. Specifically, RoVo achieved a DSR of 99.5% on a commercial speaker-verification API, effectively neutralizing speech synthesis attack. Moreover, RoVo's perturbations remained robust even under strong speech enhancement conditions, outperforming traditional methods. A user study confirmed that RoVo preserves both naturalness and usability of protected speech, highlighting its effectiveness in complex and evolving threat scenarios.

【17】 Text2midi-InferAlign: Improving Symbolic Music Generation with  Inference-Time Alignment
标题: 文本2 midi-InferAlign:通过推理时间对齐改进符号音乐生成
链接:https://arxiv.org/abs/2505.12669
作者: Abhinaba Roy,  Geeta Puri,  Dorien Herremans 
备注:7 pages, 1 figure, 5 tables
摘要:我们提出了Text 2 midi-InferAlign,一种新的技术,用于提高在推理时的符号音乐生成。我们的方法在推理过程中利用文本到音频对齐和音乐结构对齐奖励,以鼓励生成的音乐与输入标题一致。具体来说,我们引入了两个目标分数:一个文本音频一致性分数,测量生成的音乐和原始文本标题之间的节奏对齐,和一个谐波一致性分数,惩罚生成的音乐包含不一致的音符的关键。通过在生成过程中优化这些基于字幕的目标,我们的模型产生了与输入字幕更紧密联系的符号音乐,从而提高了生成的作品的整体质量和连贯性。我们的方法可以扩展任何现有的自回归模型,而不需要进一步的训练或微调。我们评估我们的工作上的Text 2 midi-现有的文本到midi生成模型,表现出显着的改善,在客观和主观的评价指标。
摘要:We present Text2midi-InferAlign, a novel technique for improving symbolic music generation at inference time. Our method leverages text-to-audio alignment and music structural alignment rewards during inference to encourage the generated music to be consistent with the input caption. Specifically, we introduce two objectives scores: a text-audio consistency score that measures rhythmic alignment between the generated music and the original text caption, and a harmonic consistency score that penalizes generated music containing notes inconsistent with the key. By optimizing these alignment-based objectives during the generation process, our model produces symbolic music that is more closely tied to the input captions, thereby improving the overall quality and coherence of the generated compositions. Our approach can extend any existing autoregressive model without requiring further training or fine-tuning. We evaluate our work on top of Text2midi - an existing text-to-midi generation model, demonstrating significant improvements in both objective and subjective evaluation metrics.

【18】 Chain-Talker: Chain Understanding and Rendering for Empathetic  Conversational Speech Synthesis
标题: Chain-Talker:同理心对话语音合成的链理解和渲染
链接:https://arxiv.org/abs/2505.12597
作者: Yifan Hu,  Rui Liu,  Yi Ren,  Xiang Yin,  Haizhou Li 
备注:16 pages, 5 figures, 5 tables. Accepted by ACL 2025 (Findings)
摘要:会话语音合成(CSS)旨在将合成的语音与用户-代理交互的情感和风格背景对齐,以实现移情。目前的生成CSS模型面临的可解释性的限制,由于不足的情感感知和冗余的离散语音编码。为了解决上述问题,我们提出了Chain-Talker,一个模仿人类认知的三阶段框架:情感理解从对话历史中获得上下文感知的情感描述符;语义理解通过序列化预测生成紧凑的语义代码;移情渲染通过整合两个组件来合成表达性语音。为了支持情感建模,我们开发了CSS-EmCap,这是一个LLM驱动的自动化管道,用于生成精确的会话语音情感字幕。在三个基准数据集上的实验表明,Chain-Talker比现有方法产生更有表现力和同情心的语音,CSS-EmCap有助于可靠的情感建模。代码和演示可在https://github.com/AI-S2-Lab/Chain-Talker上获得。
摘要:Conversational Speech Synthesis (CSS) aims to align synthesized speech with the emotional and stylistic context of user-agent interactions to achieve empathy. Current generative CSS models face interpretability limitations due to insufficient emotional perception and redundant discrete speech coding. To address the above issues, we present Chain-Talker, a three-stage framework mimicking human cognition: Emotion Understanding derives context-aware emotion descriptors from dialogue history; Semantic Understanding generates compact semantic codes via serialized prediction; and Empathetic Rendering synthesizes expressive speech by integrating both components. To support emotion modeling, we develop CSS-EmCap, an LLM-driven automated pipeline for generating precise conversational speech emotion captions. Experiments on three benchmark datasets demonstrate that Chain-Talker produces more expressive and empathetic speech than existing methods, with CSS-EmCap contributing to reliable emotion modeling. The code and demos are available at: https://github.com/AI-S2-Lab/Chain-Talker.

【19】 VoiceCloak: A Multi-Dimensional Defense Framework against Unauthorized  Diffusion-based Voice Cloning
标题: VoiceCloak:针对未经授权的基于扩散的语音克隆的多维防御框架
链接:https://arxiv.org/abs/2505.12332
作者: Qianyue Hu,  Junyan Wu,  Wei Lu,  Xiangyang Luo 
摘要:扩散模型(DM)在真实语音克隆(VC)方面取得了显着的成功,但也增加了恶意滥用的风险。现有的主动防御设计为传统的VC模型的目的是破坏伪造过程,但他们已被证明与DM由于复杂的生成机制的扩散不兼容。为了弥合这一差距,我们引入了VoiceCloak,这是一个多维的主动防御框架,其目标是在潜在的未经授权的VC中混淆扬声器身份并降低感知质量。为了实现这些目标,我们进行了重点分析,以确定DM中的特定漏洞,允许VoiceCloak通过将对抗性扰动引入参考音频来破坏克隆过程。具体来说,为了混淆说话人身份,VoiceCloak首先通过扭曲表示学习嵌入来最大化身份变化来定位说话人身份,这是由听觉感知原理指导的。此外,VoiceCloak会扰乱关键的条件引导过程,特别是注意力情境,从而阻止对实现令人信服的克隆至关重要的声音特征的对齐。然后,为了解决第二个目标,VoiceCloak引入了分数幅度放大,以主动引导反向轨迹远离高质量语音的生成。噪声引导的语义损坏被进一步用于破坏DM捕获的结构化语音语义,从而降低输出质量。大量的实验突出了VoiceCloak对未经授权的基于扩散的语音克隆的出色防御成功率。VoiceCloak的音频样本可在https://voice-cloak.github.io/VoiceCloak/上获得。
摘要:Diffusion Models (DMs) have achieved remarkable success in realistic voice cloning (VC), while they also increase the risk of malicious misuse. Existing proactive defenses designed for traditional VC models aim to disrupt the forgery process, but they have been proven incompatible with DMs due to the intricate generative mechanisms of diffusion. To bridge this gap, we introduce VoiceCloak, a multi-dimensional proactive defense framework with the goal of obfuscating speaker identity and degrading perceptual quality in potential unauthorized VC. To achieve these goals, we conduct a focused analysis to identify specific vulnerabilities within DMs, allowing VoiceCloak to disrupt the cloning process by introducing adversarial perturbations into the reference audio. Specifically, to obfuscate speaker identity, VoiceCloak first targets speaker identity by distorting representation learning embeddings to maximize identity variation, which is guided by auditory perception principles. Additionally, VoiceCloak disrupts crucial conditional guidance processes, particularly attention context, thereby preventing the alignment of vocal characteristics that are essential for achieving convincing cloning. Then, to address the second objective, VoiceCloak introduces score magnitude amplification to actively steer the reverse trajectory away from the generation of high-quality speech. Noise-guided semantic corruption is further employed to disrupt structural speech semantics captured by DMs, degrading output quality. Extensive experiments highlight VoiceCloak's outstanding defense success rate against unauthorized diffusion-based voice cloning. Audio samples of VoiceCloak are available at https://voice-cloak.github.io/VoiceCloak/.

【20】 BenSParX: A Robust Explainable Machine Learning Framework for  Parkinson's Disease Detection from Bengali Conversational Speech
标题: BenSParX:一个稳健的可解释机器学习框架,用于从孟加拉语对话语音中检测帕金森病
链接:https://arxiv.org/abs/2505.12192
作者: Riad Hossain,  Muhammad Ashad Kabir,  Arat Ibne Golam Mowla,  Animesh Chandra Roy,  Ranjit Kumar Ghosh 
备注:46 pages, 16 figures
摘要:帕金森病(PD)对全球健康构成了日益严重的挑战,孟加拉国与PD相关的死亡率显着上升。在资源有限的环境中,PD的早期检测仍然特别具有挑战性,其中基于语音的分析已成为一种有前途的非侵入性和成本效益的替代方案。然而,现有的研究主要集中在英语或其他主要语言上;值得注意的是,孟加拉语没有PD语音数据集,这对文化包容性和可访问的医疗保健解决方案构成了重大障碍。此外,大多数先前的研究只采用了一组狭窄的声学特征,有限或没有超参数调整和特征选择策略,很少注意模型的可解释性。这限制了一个强大的和可推广的机器学习模型的发展。为了解决这一差距,我们提出了BenSparX,这是第一个用于PD检测的孟加拉语会话语音数据集,以及为早期诊断量身定制的强大且可解释的机器学习框架。所提出的框架结合了不同的声学特征类别,系统的特征选择方法,以及最先进的机器学习算法和广泛的超参数优化。此外,为了增强模型预测的可解释性和可信度,该框架采用SHAP(SHapley Additive exPlanations)分析来量化单个声学特征对PD检测的贡献。我们的框架实现了最先进的性能,准确率为95.77%,F1评分为95.57%,AUC-ROC为0.982。我们通过将框架应用于其他语言的现有PD数据集,进一步从外部验证了我们的方法,在这些语言中,它始终优于最先进的方法。为了促进进一步的研究和再现性,数据集已在https://github.com/Riad071/BenSParX上公开。
摘要:Parkinson's disease (PD) poses a growing global health challenge, with Bangladesh experiencing a notable rise in PD-related mortality. Early detection of PD remains particularly challenging in resource-constrained settings, where voice-based analysis has emerged as a promising non-invasive and cost-effective alternative. However, existing studies predominantly focus on English or other major languages; notably, no voice dataset for PD exists for Bengali - posing a significant barrier to culturally inclusive and accessible healthcare solutions. Moreover, most prior studies employed only a narrow set of acoustic features, with limited or no hyperparameter tuning and feature selection strategies, and little attention to model explainability. This restricts the development of a robust and generalizable machine learning model. To address this gap, we present BenSparX, the first Bengali conversational speech dataset for PD detection, along with a robust and explainable machine learning framework tailored for early diagnosis. The proposed framework incorporates diverse acoustic feature categories, systematic feature selection methods, and state-of-the-art machine learning algorithms with extensive hyperparameter optimization. Furthermore, to enhance interpretability and trust in model predictions, the framework incorporates SHAP (SHapley Additive exPlanations) analysis to quantify the contribution of individual acoustic features toward PD detection. Our framework achieves state-of-the-art performance, yielding an accuracy of 95.77%, F1 score of 95.57%, and AUC-ROC of 0.982. We further externally validated our approach by applying the framework to existing PD datasets in other languages, where it consistently outperforms state-of-the-art approaches. To facilitate further research and reproducibility, the dataset has been made publicly available at https://github.com/Riad071/BenSParX.

【21】 Learning to Highlight Audio by Watching Movies
标题: 学习通过观看电影来突出音频
链接:https://arxiv.org/abs/2505.12154
作者: Chao Huang,  Ruohan Gao,  J. M. F. Tsang,  Jan Kurcius,  Cagdas Bilen,  Chenliang Xu,  Anurag Kumar,  Sanjeel Parekh 
备注:CVPR 2025. Project page: this https URL
摘要:近年来,视频内容的创建和消费显著增加。制作引人入胜的内容需要仔细策划视觉和音频元素。虽然视觉提示策展,通过最佳视点选择或后期编辑等技术,一直是媒体制作的核心,但其自然对应物音频却没有经历过同等的进步。这通常会导致视觉和听觉显着性之间的脱节。为了弥合这一差距,我们引入了一项新的任务:视觉引导的声学高亮,其目的是转换音频,以提供由随附视频引导的适当高亮效果,最终创建更和谐的视听体验。我们提出了一个灵活的,基于transformer的多模态框架来解决这个任务。为了训练我们的模型,我们还引入了一个新的数据集--muddy mix数据集,它利用了电影中细致的音频和视频制作,提供了一种免费的监督形式。我们开发了一个伪数据生成过程来模拟混合不好的音频,通过三步过程-分离,调整和混音来模仿现实世界的场景。我们的方法在定量和主观评估方面始终优于几个基线。我们还系统地研究了不同类型的上下文指导和数据集的难度水平的影响。我们的项目页面在这里:https://wikichao.github.io/VisAH/。
摘要:Recent years have seen a significant increase in video content creation and consumption. Crafting engaging content requires the careful curation of both visual and audio elements. While visual cue curation, through techniques like optimal viewpoint selection or post-editing, has been central to media production, its natural counterpart, audio, has not undergone equivalent advancements. This often results in a disconnect between visual and acoustic saliency. To bridge this gap, we introduce a novel task: visually-guided acoustic highlighting, which aims to transform audio to deliver appropriate highlighting effects guided by the accompanying video, ultimately creating a more harmonious audio-visual experience. We propose a flexible, transformer-based multimodal framework to solve this task. To train our model, we also introduce a new dataset -- the muddy mix dataset, leveraging the meticulous audio and video crafting found in movies, which provides a form of free supervision. We develop a pseudo-data generation process to simulate poorly mixed audio, mimicking real-world scenarios through a three-step process -- separation, adjustment, and remixing. Our approach consistently outperforms several baselines in both quantitative and subjective evaluation. We also systematically study the impact of different types of contextual guidance and difficulty levels of the dataset. Our project page is here: https://wikichao.github.io/VisAH/.

【22】 SepPrune: Structured Pruning for Efficient Deep Speech Separation
标题: SepPrune:用于高效深度语音分离的结构化修剪
链接:https://arxiv.org/abs/2505.12079
作者: Yuqi Li,  Kai Li,  Xin Yin,  Zhifei Yang,  Junhao Dong,  Zeyu Dong,  Chuanguang Yang,  Yingli Tian,  Yao Lu 
摘要:尽管近年来深度学习大大推进了语音分离,但大多数现有研究仍然优先考虑分离质量,而忽略了计算效率,这是实时应用中低延迟语音处理的一个重要因素。在本文中,我们提出了SepPrune,这是第一个专门设计用于压缩深度语音分离模型并降低其计算成本的结构化修剪框架。SepPrune首先分析给定模型的计算结构,以确定计算负担最高的层。然后,它引入了一个微分掩蔽策略,使梯度驱动的通道选择。基于学习的掩码,SepPrune修剪冗余通道并微调剩余参数以恢复性能。大量的实验表明,这种可学习的修剪范式产生实质性的优势,在语音分离模型的通道修剪,优于现有的方法。值得注意的是,使用SepPrune修剪的模型可以恢复85%的预训练模型(经过数百个epoch训练)的性能,只需一个epoch的微调,并且比从头开始训练快36倍。代码可在https://github.com/itsnotacie/SepPrune上获得。
摘要:Although deep learning has substantially advanced speech separation in recent years, most existing studies continue to prioritize separation quality while overlooking computational efficiency, an essential factor for low-latency speech processing in real-time applications. In this paper, we propose SepPrune, the first structured pruning framework specifically designed to compress deep speech separation models and reduce their computational cost. SepPrune begins by analyzing the computational structure of a given model to identify layers with the highest computational burden. It then introduces a differentiable masking strategy to enable gradient-driven channel selection. Based on the learned masks, SepPrune prunes redundant channels and fine-tunes the remaining parameters to recover performance. Extensive experiments demonstrate that this learnable pruning paradigm yields substantial advantages for channel pruning in speech separation models, outperforming existing methods. Notably, a model pruned with SepPrune can recover 85% of the performance of a pre-trained model (trained over hundreds of epochs) with only one epoch of fine-tuning, and achieves convergence 36$\times$ faster than training from scratch. Code is available at https://github.com/itsnotacie/SepPrune.

【23】 Automatic Speech Recognition for African Low-Resource Languages:  Challenges and Future Directions
标题: 非洲低资源语言的自动语音识别:挑战和未来方向
链接:https://arxiv.org/abs/2505.11690
作者: Sukairaj Hafiz Imam,  Babangida Sani,  Dawit Ketema Gete,  Bedru Yimam Ahamed,  Ibrahim Said Ahmad,  Idris Abdulmumin,  Seid Muhie Yimam,  Muhammad Yahuza Bello,  Shamsuddeen Hassan Muhammad 
摘要:自动语音识别(ASR)技术已经改变了人机交互;然而,非洲的低资源语言在研究和实际应用中仍然严重不足。这项研究调查了阻碍这些语言的ASR系统开发的主要挑战,包括数据稀缺,语言复杂性,有限的计算资源,声学可变性以及围绕偏见和隐私的道德问题。主要目标是批判性地分析这些障碍,并确定实用的,包容性的战略,以推进非洲范围内的ASR技术。最近的进展和案例研究强调了有前途的策略,如社区驱动的数据收集,自我监督和多语言学习,轻量级模型架构和优先考虑隐私的技术。来自涉及各种非洲语言的试点项目的证据展示了定制解决方案的可行性和影响,其中包括基于词素的建模和医疗保健和教育等领域的特定领域ASR应用。研究结果强调了跨学科合作和持续投资的重要性,以应对非洲大陆面临的独特语言和基础设施挑战。这项研究为创建道德,高效和包容性的ASR系统提供了一个渐进的路线图,该系统不仅保护语言多样性,还改善了数字可访问性,并促进了非洲语言使用者的社会经济参与。
摘要:Automatic Speech Recognition (ASR) technologies have transformed human-computer interaction; however, low-resource languages in Africa remain significantly underrepresented in both research and practical applications. This study investigates the major challenges hindering the development of ASR systems for these languages, which include data scarcity, linguistic complexity, limited computational resources, acoustic variability, and ethical concerns surrounding bias and privacy. The primary goal is to critically analyze these barriers and identify practical, inclusive strategies to advance ASR technologies within the African context. Recent advances and case studies emphasize promising strategies such as community-driven data collection, self-supervised and multilingual learning, lightweight model architectures, and techniques that prioritize privacy. Evidence from pilot projects involving various African languages showcases the feasibility and impact of customized solutions, which encompass morpheme-based modeling and domain-specific ASR applications in sectors like healthcare and education. The findings highlight the importance of interdisciplinary collaboration and sustained investment to tackle the distinct linguistic and infrastructural challenges faced by the continent. This study offers a progressive roadmap for creating ethical, efficient, and inclusive ASR systems that not only safeguard linguistic diversity but also improve digital accessibility and promote socioeconomic participation for speakers of African languages.

【24】 ASR-FAIRBENCH: Measuring and Benchmarking Equity Across Speech  Recognition Systems
标题: ASR-FAIRBENCH:测量和基准测量语音识别系统的公平性
链接:https://arxiv.org/abs/2505.11572
作者: Anand Rai,  Satyam Rahangdale,  Utkarsh Anand,  Animesh Mukherjee 
备注:Paper accepted at INTERSPEECH 2025
摘要:自动语音识别(ASR)系统在日常应用中已经无处不在,但在不同的人口统计群体之间的性能仍然存在显着差异。在这项工作中,我们介绍了ASR-FAIRBENCH排行榜,该排行榜旨在实时评估ASR模型的准确性和公平性。利用Meta的公平言论数据集,它捕捉了不同的人口统计特征,我们采用了混合效应泊松回归模型来获得整体公平得分。该分数与单词错误率(WER)等传统指标相结合,以计算公平调整的ASR分数(FAAS),提供全面的评估框架。我们的方法揭示了SOTA ASR模型在不同人口群体中的显著性能差异,并提供了一个基准,以推动更具包容性的ASR技术的发展。
摘要:Automatic Speech Recognition (ASR) systems have become ubiquitous in everyday applications, yet significant disparities in performance across diverse demographic groups persist. In this work, we introduce the ASR-FAIRBENCH leaderboard which is designed to assess both the accuracy and equity of ASR models in real-time. Leveraging the Meta's Fair-Speech dataset, which captures diverse demographic characteristics, we employ a mixed-effects Poisson regression model to derive an overall fairness score. This score is integrated with traditional metrics like Word Error Rate (WER) to compute the Fairness Adjusted ASR Score (FAAS), providing a comprehensive evaluation framework. Our approach reveals significant performance disparities in SOTA ASR models across demographic groups and offers a benchmark to drive the development of more inclusive ASR technologies.

【25】 SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based  on Speech and Audio Information
标题: SAKURA:基于语音和音频信息的大型音频语言模型的多跳推理
链接:https://arxiv.org/abs/2505.13237
作者: Chih-Kai Yang,  Neo Ho,  Yen-Ting Piao,  Hung-yi Lee 
备注:Accepted to Interspeech 2025
摘要:大型音频语言模型(LALM)扩展了大型语言模型,在语音,音频等多模态理解,而他们的语音和音频处理任务的性能进行了广泛的研究,他们的推理能力仍然未被探索。特别是,他们的多跳推理,回忆和整合多个事实的能力,缺乏系统的评估。现有的基准集中在一般的语音和音频处理任务,会话能力和公平性,但忽略了这方面。为了弥补这一差距,我们介绍了SAKURA,一个基准评估LALM的多跳推理的语音和音频信息的基础上。结果表明,LALM努力整合语音/音频表示的多跳推理,即使他们正确地提取相关信息,突出了多模态推理的根本挑战。我们的研究结果揭示了LALM的关键局限性,为未来的研究提供了见解和资源。
摘要:Large audio-language models (LALMs) extend the large language models with multimodal understanding in speech, audio, etc. While their performances on speech and audio-processing tasks are extensively studied, their reasoning abilities remain underexplored. Particularly, their multi-hop reasoning, the ability to recall and integrate multiple facts, lacks systematic evaluation. Existing benchmarks focus on general speech and audio-processing tasks, conversational abilities, and fairness but overlook this aspect. To bridge this gap, we introduce SAKURA, a benchmark assessing LALMs' multi-hop reasoning based on speech and audio information. Results show that LALMs struggle to integrate speech/audio representations for multi-hop reasoning, even when they extract the relevant information correctly, highlighting a fundamental challenge in multimodal reasoning. Our findings expose a critical limitation in LALMs, offering insights and resources for future research.

【26】 Optimal Scalogram for Computational Complexity Reduction in Acoustic  Recognition Using Deep Learning
标题: 使用深度学习降低声学识别中计算复杂性的最佳比例图
链接:https://arxiv.org/abs/2505.13017
作者: Dang Thoai Phan,  Tuan Anh Huynh,  Van Tuan Pham,  Cao Minh Tran,  Van Thuan Mai,  Ngoc Quy Tran 
摘要:连续小波变换(CWT)是利用卷积神经网络(CNN)进行声学识别中特征提取的有效工具,特别是当应用于非平稳音频时。然而,它的高计算成本构成了一个重大挑战,往往导致研究人员更喜欢替代方法,如短时傅立叶变换(STFT)。为了解决这个问题,本文提出了一种方法,以减少计算复杂度的连续小波变换的小波核的长度和输出尺度图的跳数大小进行优化。实验结果表明,该方法显着降低了计算成本,同时保持在声学识别任务的训练模型的鲁棒性。
摘要:The Continuous Wavelet Transform (CWT) is an effective tool for feature extraction in acoustic recognition using Convolutional Neural Networks (CNNs), particularly when applied to non-stationary audio. However, its high computational cost poses a significant challenge, often leading researchers to prefer alternative methods such as the Short-Time Fourier Transform (STFT). To address this issue, this paper proposes a method to reduce the computational complexity of CWT by optimizing the length of the wavelet kernel and the hop size of the output scalogram. Experimental results demonstrate that the proposed approach significantly reduces computational cost while maintaining the robust performance of the trained model in acoustic recognition tasks.

【27】 Acoustic Field Reconstruction in Tubes via Physics-Informed Neural  Networks
标题: 通过物理信息神经网络重建管内的声学场
链接:https://arxiv.org/abs/2505.12557
作者: Xinmeng Luan,  Kazuya Yokota,  Gary Scavone 
备注:8 pages, 5 figures, conference
摘要:本研究探讨物理信息神经网络(PINNs)的应用程序在声管分析的逆问题,重点是重建声场从嘈杂和有限的观测数据。具体来说,我们解决的辐射模型是未知的情况下,和压力数据只在管的辐射端。提出了一种PINN框架来重建声场,以及PINN微调方法(PINN-FTM)和用于预测辐射模型系数的传统优化方法(TOM)。结果表明,PINNs可以有效地重建管的声场在噪声条件下,即使与未知的辐射参数。PINN-FTM通过提供平衡和可靠的预测并表现出强大的抗噪能力而优于TOM。
摘要:This study investigates the application of Physics-Informed Neural Networks (PINNs) to inverse problems in acoustic tube analysis, focusing on reconstructing acoustic fields from noisy and limited observation data. Specifically, we address scenarios where the radiation model is unknown, and pressure data is only available at the tube's radiation end. A PINNs framework is proposed to reconstruct the acoustic field, along with the PINN Fine-Tuning Method (PINN-FTM) and a traditional optimization method (TOM) for predicting radiation model coefficients. The results demonstrate that PINNs can effectively reconstruct the tube's acoustic field under noisy conditions, even with unknown radiation parameters. PINN-FTM outperforms TOM by delivering balanced and reliable predictions and exhibiting robust noise-tolerance capabilities.

【28】 Unified Architecture and Unsupervised Speech Disentanglement for Speaker  Embedding-Free Enrollment in Personalized Speech Enhancement
标题: 个性化语音增强中的说话人嵌入免注册的统一架构和无监督语音解纠缠
链接:https://arxiv.org/abs/2505.12288
作者: Ziling Huang,  Haixin Guan,  Yanhua Long 
备注:Submitted to the IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP)
摘要:传统的语音增强(SE)旨在通过抑制噪声而不需要注册语音作为参考来改善语音感知和可懂度,而个性化的SE(PSE)通过使用注册语音提取目标说话人的语音来解决鸡尾酒会问题。虽然这两个任务解决了语音信号处理中不同但互补的挑战,但它们通常共享类似的模型架构,PSE包含一个额外的分支来处理注册语音。这表明开发一个能够有效处理SE和PSE任务的统一模型,从而简化部署,同时保持高性能。然而,PSE性能对注册语音的变化(如情绪语调)敏感,这限制了现实世界应用中的鲁棒性。为了解决这些挑战,我们提出了两个新的模型,USEF-PNet和DSEF-PNet,都扩展了我们以前的SEF-PNet框架。USEF-PNet引入了一个统一的架构来处理注册语音,将SE和PSE集成到一个框架中,以提高性能并简化部署。同时,DSEF-PNet通过将混合语音与两个不同的注册话语配对并在提取的目标语音中执行一致性来结合无监督语音解纠缠方法。该策略有效地将高质量的说话人身份信息从注册语音中分离出来,减少了情感和内容等因素的干扰,从而提高了PSE的鲁棒性。此外,我们探索了一个长短注册配对(LSEP)策略,以检查注册语音持续时间在训练和评估过程中的影响。在Libri 2 Mix和VoiceBank DEMAND上进行的大量实验表明,我们提出的USEF-PNet,DSEF-PNet都实现了实质性的性能改进,随机注册持续时间表现略好。
摘要:Conventional speech enhancement (SE) aims to improve speech perception and intelligibility by suppressing noise without requiring enrollment speech as reference, whereas personalized SE (PSE) addresses the cocktail party problem by extracting a target speaker's speech using enrollment speech. While these two tasks tackle different yet complementary challenges in speech signal processing, they often share similar model architectures, with PSE incorporating an additional branch to process enrollment speech. This suggests developing a unified model capable of efficiently handling both SE and PSE tasks, thereby simplifying deployment while maintaining high performance. However, PSE performance is sensitive to variations in enrollment speech, like emotional tone, which limits robustness in real-world applications. To address these challenges, we propose two novel models, USEF-PNet and DSEF-PNet, both extending our previous SEF-PNet framework. USEF-PNet introduces a unified architecture for processing enrollment speech, integrating SE and PSE into a single framework to enhance performance and streamline deployment. Meanwhile, DSEF-PNet incorporates an unsupervised speech disentanglement approach by pairing a mixture speech with two different enrollment utterances and enforcing consistency in the extracted target speech. This strategy effectively isolates high-quality speaker identity information from enrollment speech, reducing interference from factors such as emotion and content, thereby improving PSE robustness. Additionally, we explore a long-short enrollment pairing (LSEP) strategy to examine the impact of enrollment speech duration during both training and evaluation. Extensive experiments on the Libri2Mix and VoiceBank DEMAND demonstrate that our proposed USEF-PNet, DSEF-PNet all achieve substantial performance improvements, with random enrollment duration performing slightly better.

【29】 Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
标题: 用于粗到细文本到语音合成的浅层流匹配
链接:https://arxiv.org/abs/2505.12226
作者: Dong Yang,  Yiyi Cai,  Yuki Saito,  Lixu Wang,  Hiroshi Saruwatari 
摘要:我们提出了一个浅流匹配(SFM)机制,以增强流匹配(FM)为基础的文本到语音(TTS)模型在一个由粗到细的生成范式。SFM使用粗略的输出表示沿着FM路径构造中间状态。在训练过程中,我们引入了一种正交投影方法来自适应地确定这些状态的时间位置,并应用了基于单段分段流的原则性构造策略。SFM推断从中间状态而不是纯噪声开始,并且将计算集中在FM路径的后级上。我们将SFM集成到多个TTS模型中,并使用轻量级的SFM头。实验表明,SFM一贯提高合成语音的自然度在客观和主观评价,同时显着减少推理时,使用自适应步长ODE求解器。演示和代码可在https://ydqmkkx.github.io/SFMDemo/上获得。
摘要:We propose a shallow flow matching (SFM) mechanism to enhance flow matching (FM)-based text-to-speech (TTS) models within a coarse-to-fine generation paradigm. SFM constructs intermediate states along the FM paths using coarse output representations. During training, we introduce an orthogonal projection method to adaptively determine the temporal position of these states, and apply a principled construction strategy based on a single-segment piecewise flow. The SFM inference starts from the intermediate state rather than pure noise and focuses computation on the latter stages of the FM paths. We integrate SFM into multiple TTS models with a lightweight SFM head. Experiments show that SFM consistently improves the naturalness of synthesized speech in both objective and subjective evaluations, while significantly reducing inference when using adaptive-step ODE solvers. Demo and codes are available at https://ydqmkkx.github.io/SFMDemo/.

【30】 BINAQUAL: A Full-Reference Objective Localization Similarity Metric for  Binaural Audio
标题: SINAQUAL:一种用于双耳音频的全参考客观定位相似性指标
链接:https://arxiv.org/abs/2505.11915
作者: Davoud Shariat Panah,  Dan Barry,  Alessandro Ragano,  Jan Skoglund,  Andrew Hines 
备注:Submitted to the Journal of Audio Engineering Society (JAES)
摘要:空间音频通过创建三维听觉体验来增强虚拟现实、增强现实、游戏和电影等应用的沉浸感。确保双耳音频的空间保真度至关重要,因为压缩、编码或传输等过程可能会改变定位线索。虽然像MUSHRA这样的主观听力测试仍然是评估空间定位质量的黄金标准,但它们成本高昂且耗时。本文介绍了BINAQUAL,一个完整的参考客观指标,旨在评估双耳录音定位相似性。BINAQUAL将AMBIQUAL度量(最初为在立体混响音频格式中进行定位质量评估而开发)调整为双耳域。我们评估BINAQUAL在五个关键的研究问题,检查其灵敏度的声源位置,角度插值,环绕扬声器布局,音频退化,和内容多样性的变化。结果表明,BINAQUAL有效地区分细微的空间变化,并与主观听力测试强烈相关,使其成为双耳定位质量评估的可靠指标。所提出的度量为确保双耳音频处理的空间准确性提供了一个强大的基准,为改善沉浸式音频应用中的客观评估铺平了道路。
摘要:Spatial audio enhances immersion in applications such as virtual reality, augmented reality, gaming, and cinema by creating a three-dimensional auditory experience. Ensuring the spatial fidelity of binaural audio is crucial, given that processes such as compression, encoding, or transmission can alter localization cues. While subjective listening tests like MUSHRA remain the gold standard for evaluating spatial localization quality, they are costly and time-consuming. This paper introduces BINAQUAL, a full-reference objective metric designed to assess localization similarity in binaural audio recordings. BINAQUAL adapts the AMBIQUAL metric, originally developed for localization quality assessment in ambisonics audio format to the binaural domain. We evaluate BINAQUAL across five key research questions, examining its sensitivity to variations in sound source locations, angle interpolations, surround speaker layouts, audio degradations, and content diversity. Results demonstrate that BINAQUAL effectively differentiates between subtle spatial variations and correlates strongly with subjective listening tests, making it a reliable metric for binaural localization quality assessment. The proposed metric provides a robust benchmark for ensuring spatial accuracy in binaural audio processing, paving the way for improved objective evaluations in immersive audio applications.

【31】 Exploring the Potential of SSL Models for Sound Event Detection
标题: 探索SSL模型用于声音事件检测的潜力
链接:https://arxiv.org/abs/2505.11889
作者: Hanfang Cui,  Longfei Song,  Li Li,  Dongxing Xu,  Yanhua Long 
备注:27 pages, 5 figures, submitted to the Journal of King Saud University - Computer and Information Sciences (under review)
摘要:自监督学习(SSL)模型为声音事件检测(SED)提供了强大的代表性,但其协同潜力仍然未得到充分挖掘。本研究系统地评估了最先进的SSL模型,以指导最佳的模型选择和集成SED。我们提出了一个框架,结合异构SSL表示(例如,Beats、HuBERT、WavLM)通过三种融合策略:单独SSL嵌入集成、双模式融合和完全聚合。DCASE 2023任务4挑战赛的实验表明,双模态融合(例如,CRNN+ BEAT +WavLM)实现了互补的性能增益,而单独使用CRNN+ BEAT可以在各个SSL模型中提供最佳结果。我们进一步引入了归一化的声音事件边界框(nSEBB),这是一种自适应后处理方法,可以动态调整事件边界预测,将独立SSL模型的PSDS 1提高了4%。这些发现突出了SSL架构的兼容性和互补性,为特定任务的融合和强大的SED系统设计提供指导。
摘要:Self-supervised learning (SSL) models offer powerful representations for sound event detection (SED), yet their synergistic potential remains underexplored. This study systematically evaluates state-of-the-art SSL models to guide optimal model selection and integration for SED. We propose a framework that combines heterogeneous SSL representations (e.g., BEATs, HuBERT, WavLM) through three fusion strategies: individual SSL embedding integration, dual-modal fusion, and full aggregation. Experiments on the DCASE 2023 Task 4 Challenge reveal that dual-modal fusion (e.g., CRNN+BEATs+WavLM) achieves complementary performance gains, while CRNN+BEATs alone delivers the best results among individual SSL models. We further introduce normalized sound event bounding boxes (nSEBBs), an adaptive post-processing method that dynamically adjusts event boundary predictions, improving PSDS1 by up to 4% for standalone SSL models. These findings highlight the compatibility and complementarity of SSL architectures, providing guidance for task-specific fusion and robust SED system design.

【32】 AnalyticKWS: Towards Exemplar-Free Analytic Class Incremental Learning  for Small-footprint Keyword Spotting
标题: AnalytyKWS:迈向无示例分析类增量学习,以实现小规模关键词发现
链接:https://arxiv.org/abs/2505.11817
作者: Yang Xiao,  Tianyi Peng,  Rohan Kumar Das,  Yuchen Hu,  Huiping Zhuang 
备注:Accepted by ACL 2025
摘要:关键词定位(KWS)提供了一种重要的机制来识别语音系统中的口头命令,在语音系统中,用户的需求经常变化,需要模型随着时间的推移不断学习新的关键词。然而,一个主要的问题是灾难性的遗忘,模型失去了识别早期关键字的能力。虽然一些持续学习方法已经证明了它们在减少遗忘方面的有效性,但大多数现有方法都依赖于存储和重新访问旧数据来对抗灾难性遗忘。虽然有效,但这些方法面临两个实际挑战:1)保留用户数据的隐私风险和2)限制在小型设备上部署的大内存和时间消耗。为了解决这些问题,我们提出了一个无范例的分析持续学习(AnalyticKWS)方法,更新模型参数,而无需重新访问以前的数据。受高效学习原则的启发,AnalyticKWS计算模型更新的封闭形式解析解,并且只需要对传入的关键字进行单次适应。AnalyticKWS通过避免基于梯度的更新而需要更少的计算资源,并且不存储旧数据。通过消除增量学习过程中对反向传播的需求,该模型保持了轻量级和高效。因此,AnalyticKWS满足了前面提到的挑战,并且非常适合资源有限的环境。在各种数据集和设置上的广泛实验表明,AnalyticKWS始终优于现有的持续学习方法。
摘要:Keyword spotting (KWS) offers a vital mechanism to identify spoken commands in voice-enabled systems, where user demands often shift, requiring models to learn new keywords continually over time. However, a major problem is catastrophic forgetting, where models lose their ability to recognize earlier keywords. Although several continual learning methods have proven their usefulness for reducing forgetting, most existing approaches depend on storing and revisiting old data to combat catastrophic forgetting. Though effective, these methods face two practical challenges: 1) privacy risks from keeping user data and 2) large memory and time consumption that limit deployment on small devices. To address these issues, we propose an exemplar-free Analytic Continual Learning (AnalyticKWS) method that updates model parameters without revisiting earlier data. Inspired by efficient learning principles, AnalyticKWS computes a closed-form analytical solution for model updates and requires only a single epoch of adaptation for incoming keywords. AnalyticKWS demands fewer computational resources by avoiding gradient-based updates and does not store old data. By eliminating the need for back-propagation during incremental learning, the model remains lightweight and efficient. As a result, AnalyticKWS meets the challenges mentioned earlier and suits resource-limited settings well. Extensive experiments on various datasets and settings show that AnalyticKWS consistently outperforms existing continual learning methods.


eess.AS音频处理

【1】 SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based  on Speech and Audio Information
标题: SAKURA:基于语音和音频信息的大型音频语言模型的多跳推理
链接:https://arxiv.org/abs/2505.13237
作者: Chih-Kai Yang,  Neo Ho,  Yen-Ting Piao,  Hung-yi Lee 
备注:Accepted to Interspeech 2025
摘要:大型音频语言模型(LALM)扩展了大型语言模型,在语音,音频等多模态理解,而他们的语音和音频处理任务的性能进行了广泛的研究,他们的推理能力仍然未被探索。特别是,他们的多跳推理,回忆和整合多个事实的能力,缺乏系统的评估。现有的基准集中在一般的语音和音频处理任务,会话能力和公平性,但忽略了这方面。为了弥补这一差距,我们介绍了SAKURA,一个基准评估LALM的多跳推理的语音和音频信息的基础上。结果表明,LALM努力整合语音/音频表示的多跳推理,即使他们正确地提取相关信息,突出了多模态推理的根本挑战。我们的研究结果揭示了LALM的关键局限性,为未来的研究提供了见解和资源。
摘要:Large audio-language models (LALMs) extend the large language models with multimodal understanding in speech, audio, etc. While their performances on speech and audio-processing tasks are extensively studied, their reasoning abilities remain underexplored. Particularly, their multi-hop reasoning, the ability to recall and integrate multiple facts, lacks systematic evaluation. Existing benchmarks focus on general speech and audio-processing tasks, conversational abilities, and fairness but overlook this aspect. To bridge this gap, we introduce SAKURA, a benchmark assessing LALMs' multi-hop reasoning based on speech and audio information. Results show that LALMs struggle to integrate speech/audio representations for multi-hop reasoning, even when they extract the relevant information correctly, highlighting a fundamental challenge in multimodal reasoning. Our findings expose a critical limitation in LALMs, offering insights and resources for future research.

【2】 Universal Semantic Disentangled Privacy-preserving Speech Representation  Learning
标题: 通用语义解开隐私保护语音表示学习
链接:https://arxiv.org/abs/2505.13085
作者: Biel Tura Vecino,  Subhadeep Maji,  Aravind Varier,  Antonio Bonafonte,  Ivan Valles,  Michael Owen,  Leif Radel,  Grant Strimmel,  Seyi Feyisetan,  Roberto Barra Chicote,  Ariya Rastrow,  Constantinos Papayiannis,  Volker Leutnant,  Trevor Wood 
备注:Accepted at Interspeech 2025
摘要:使用人类语音的音频记录来训练LLM引起了隐私问题,因为这些模型可能生成与训练数据中的伪像非常相似的输出。在这项研究中,我们提出了一个扬声器的隐私保护表示学习方法,通过通用语音编解码器(USC),一个计算效率高的编码器-解码器模型,将语音分解为:$\texit {(i)}$隐私保护语义丰富的表示,捕获内容和语音非语言学,和$\texit {(ii)}$残留的声学和扬声器表示,使高保真重建。广泛的评估表明,南加州大学的语义表示保留内容,韵律和情感,同时删除潜在的可识别的扬声器属性。结合这两种表示,USC实现了最先进的语音重建。此外,我们引入了一种评估方法来衡量隐私保护属性,与感知测试。我们将USC与文献中的其他编解码器进行了比较,并证明了其在隐私保护表示学习方面的有效性,说明了在学习的语义表示中说话人匿名化,语言保留和内容保留的权衡。音频样本在$\href{https://www.amazon.science/usc-samples}{https://www.amazon.science/usc-samples}$中共享。
摘要:The use of audio recordings of human speech to train LLMs poses privacy concerns due to these models' potential to generate outputs that closely resemble artifacts in the training data. In this study, we propose a speaker privacy-preserving representation learning method through the Universal Speech Codec (USC), a computationally efficient encoder-decoder model that disentangles speech into: $\textit{(i)}$ privacy-preserving semantically rich representations, capturing content and speech paralinguistics, and $\textit{(ii)}$ residual acoustic and speaker representations that enables high-fidelity reconstruction. Extensive evaluations presented show that USC's semantic representation preserves content, prosody, and sentiment, while removing potentially identifiable speaker attributes. Combining both representations, USC achieves state-of-the-art speech reconstruction. Additionally, we introduce an evaluation methodology for measuring privacy-preserving properties, aligning with perceptual tests. We compare USC against other codecs in the literature and demonstrate its effectiveness on privacy-preserving representation learning, illustrating the trade-offs of speaker anonymization, paralinguistics retention and content preservation in the learned semantic representations. Audio samples are shared in $\href{https://www.amazon.science/usc-samples}{https://www.amazon.science/usc-samples}$.

【3】 Cross-modal Knowledge Transfer Learning as Graph Matching Based on  Optimal Transport for ASR
标题: 基于ASB最优传输的跨模式知识转移学习作为图匹配
链接:https://arxiv.org/abs/2505.13079
作者: Xugang Lu,  Peng Shen,  Yu Tsao,  Hisashi Kawai 
备注:To appear in Interspeech 2025
摘要:将语言知识从预训练语言模型(PLM)转移到声学特征学习已被证明在增强端到端自动语音识别(E2 E-ASR)方面是有效的。然而,由于固有的模态差距,语言和声学模态之间的对齐表示仍然是一个挑战。最优传输(OT)已经显示出希望,在减轻这些差距,最大限度地减少语言和声学特征分布之间的Wasserstein距离(WD)。然而,以前基于OT的方法忽略了结构关系,将特征向量视为无序集合。为了解决这个问题,我们提出了图匹配最优传输(GM-OT),它将语言和声学序列建模为结构化图。节点表示特征嵌入,而边缘捕获时间和顺序关系。GM-OT最小化WD(节点之间)和Gromov-Wasserstein距离(GWD)(边缘之间),导致融合Gromov-Wasserstein距离(FGWD)公式。与现有的基于OT的方法相比,这能够实现结构化对齐和更有效的知识转移。理论分析进一步表明,先前的基于OT的语言知识迁移方法可以被视为我们的GM-OT框架中的一个特例。我们使用基于CTC的E2 E-ASR系统评估GM-OT对普通话ASR的影响,并使用PLM进行知识转移。实验结果表明,显着的性能增益超过国家的最先进的模型,验证了我们的方法的有效性。
摘要:Transferring linguistic knowledge from a pretrained language model (PLM) to acoustic feature learning has proven effective in enhancing end-to-end automatic speech recognition (E2E-ASR). However, aligning representations between linguistic and acoustic modalities remains a challenge due to inherent modality gaps. Optimal transport (OT) has shown promise in mitigating these gaps by minimizing the Wasserstein distance (WD) between linguistic and acoustic feature distributions. However, previous OT-based methods overlook structural relationships, treating feature vectors as unordered sets. To address this, we propose Graph Matching Optimal Transport (GM-OT), which models linguistic and acoustic sequences as structured graphs. Nodes represent feature embeddings, while edges capture temporal and sequential relationships. GM-OT minimizes both WD (between nodes) and Gromov-Wasserstein distance (GWD) (between edges), leading to a fused Gromov-Wasserstein distance (FGWD) formulation. This enables structured alignment and more efficient knowledge transfer compared to existing OT-based approaches. Theoretical analysis further shows that prior OT-based methods in linguistic knowledge transfer can be viewed as a special case within our GM-OT framework. We evaluate GM-OT on Mandarin ASR using a CTC-based E2E-ASR system with a PLM for knowledge transfer. Experimental results demonstrate significant performance gains over state-of-the-art models, validating the effectiveness of our approach.

【4】 MDDM: A Multi-view Discriminative Enhanced Diffusion-based Model for  Speech Enhancement
标题: MDDM:一种基于多视点鉴别增强扩散的语音增强模型
链接:https://arxiv.org/abs/2505.13029
作者: Nan Xu,  Zhaolong Huang,  Xiaonan Zhi 
备注:6 pages, 2 figures
摘要:随着深度学习的发展,语音增强在语音质量方面得到了极大的优化。以前的方法通常集中在判别式监督学习或生成式建模,这往往会引入语音失真或高计算成本。在本文中,我们提出了MDDM,一个多视图的歧视性增强扩散为基础的模型。具体来说,我们将三个域(时间,频率和噪声)的特征作为判别预测网络的输入,生成初步的频谱图。然后,通过几个推理采样步骤,可以将有区别的输出转换为干净的语音。由于区分输出和干净目标之间的分布相交,较小的采样步长可以实现与其他基于扩散的方法相比具有竞争力的性能。在公共数据集和真实数据集上进行的实验验证了MDDM的有效性,无论是在主观还是客观度量。
摘要:With the development of deep learning, speech enhancement has been greatly optimized in terms of speech quality. Previous methods typically focus on the discriminative supervised learning or generative modeling, which tends to introduce speech distortions or high computational cost. In this paper, we propose MDDM, a Multi-view Discriminative enhanced Diffusion-based Model. Specifically, we take the features of three domains (time, frequency and noise) as inputs of a discriminative prediction network, generating the preliminary spectrogram. Then, the discriminative output can be converted to clean speech by several inference sampling steps. Due to the intersection of the distributions between discriminative output and clean target, the smaller sampling steps can achieve the competitive performance compared to other diffusion-based methods. Experiments conducted on a public dataset and a realworld dataset validate the effectiveness of MDDM, either on subjective or objective metric.

【5】 Optimal Scalogram for Computational Complexity Reduction in Acoustic  Recognition Using Deep Learning
标题: 使用深度学习降低声学识别中计算复杂性的最佳比例图
链接:https://arxiv.org/abs/2505.13017
作者: Dang Thoai Phan,  Tuan Anh Huynh,  Van Tuan Pham,  Cao Minh Tran,  Van Thuan Mai,  Ngoc Quy Tran 
摘要:连续小波变换(CWT)是利用卷积神经网络(CNN)进行声学识别中特征提取的有效工具,特别是当应用于非平稳音频时。然而,它的高计算成本构成了一个重大挑战,往往导致研究人员更喜欢替代方法,如短时傅立叶变换(STFT)。为了解决这个问题,本文提出了一种方法,以减少计算复杂度的连续小波变换的小波核的长度和输出尺度图的跳数大小进行优化。实验结果表明,该方法显着降低了计算成本,同时保持在声学识别任务的训练模型的鲁棒性。
摘要:The Continuous Wavelet Transform (CWT) is an effective tool for feature extraction in acoustic recognition using Convolutional Neural Networks (CNNs), particularly when applied to non-stationary audio. However, its high computational cost poses a significant challenge, often leading researchers to prefer alternative methods such as the Short-Time Fourier Transform (STFT). To address this issue, this paper proposes a method to reduce the computational complexity of CWT by optimizing the length of the wavelet kernel and the hop size of the output scalogram. Experimental results demonstrate that the proposed approach significantly reduces computational cost while maintaining the robust performance of the trained model in acoustic recognition tasks.

【6】 Acoustic Field Reconstruction in Tubes via Physics-Informed Neural  Networks
标题: 通过物理信息神经网络重建管内的声学场
链接:https://arxiv.org/abs/2505.12557
作者: Xinmeng Luan,  Kazuya Yokota,  Gary Scavone 
备注:8 pages, 5 figures, conference
摘要:本研究探讨物理信息神经网络(PINNs)的应用程序在声管分析的逆问题,重点是重建声场从嘈杂和有限的观测数据。具体来说,我们解决的辐射模型是未知的情况下,和压力数据只在管的辐射端。提出了一种PINN框架来重建声场,以及PINN微调方法(PINN-FTM)和用于预测辐射模型系数的传统优化方法(TOM)。结果表明,PINNs可以有效地重建管的声场在噪声条件下,即使与未知的辐射参数。PINN-FTM通过提供平衡和可靠的预测并表现出强大的抗噪能力而优于TOM。
摘要:This study investigates the application of Physics-Informed Neural Networks (PINNs) to inverse problems in acoustic tube analysis, focusing on reconstructing acoustic fields from noisy and limited observation data. Specifically, we address scenarios where the radiation model is unknown, and pressure data is only available at the tube's radiation end. A PINNs framework is proposed to reconstruct the acoustic field, along with the PINN Fine-Tuning Method (PINN-FTM) and a traditional optimization method (TOM) for predicting radiation model coefficients. The results demonstrate that PINNs can effectively reconstruct the tube's acoustic field under noisy conditions, even with unknown radiation parameters. PINN-FTM outperforms TOM by delivering balanced and reliable predictions and exhibiting robust noise-tolerance capabilities.

【7】 Unified Architecture and Unsupervised Speech Disentanglement for Speaker  Embedding-Free Enrollment in Personalized Speech Enhancement
标题: 个性化语音增强中的说话人嵌入免注册的统一架构和无监督语音解纠缠
链接:https://arxiv.org/abs/2505.12288
作者: Ziling Huang,  Haixin Guan,  Yanhua Long 
备注:Submitted to the IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP)
摘要:传统的语音增强(SE)旨在通过抑制噪声而不需要注册语音作为参考来改善语音感知和可懂度,而个性化的SE(PSE)通过使用注册语音提取目标说话人的语音来解决鸡尾酒会问题。虽然这两个任务解决了语音信号处理中不同但互补的挑战,但它们通常共享类似的模型架构,PSE包含一个额外的分支来处理注册语音。这表明开发一个能够有效处理SE和PSE任务的统一模型,从而简化部署,同时保持高性能。然而,PSE性能对注册语音的变化(如情绪语调)敏感,这限制了现实世界应用中的鲁棒性。为了解决这些挑战,我们提出了两个新的模型,USEF-PNet和DSEF-PNet,都扩展了我们以前的SEF-PNet框架。USEF-PNet引入了一个统一的架构来处理注册语音,将SE和PSE集成到一个框架中,以提高性能并简化部署。同时,DSEF-PNet通过将混合语音与两个不同的注册话语配对并在提取的目标语音中执行一致性来结合无监督语音解纠缠方法。该策略有效地将高质量的说话人身份信息从注册语音中分离出来,减少了情感和内容等因素的干扰,从而提高了PSE的鲁棒性。此外,我们探索了一个长短注册配对(LSEP)策略,以检查注册语音持续时间在训练和评估过程中的影响。在Libri 2 Mix和VoiceBank DEMAND上进行的大量实验表明,我们提出的USEF-PNet,DSEF-PNet都实现了实质性的性能改进,随机注册持续时间表现略好。
摘要:Conventional speech enhancement (SE) aims to improve speech perception and intelligibility by suppressing noise without requiring enrollment speech as reference, whereas personalized SE (PSE) addresses the cocktail party problem by extracting a target speaker's speech using enrollment speech. While these two tasks tackle different yet complementary challenges in speech signal processing, they often share similar model architectures, with PSE incorporating an additional branch to process enrollment speech. This suggests developing a unified model capable of efficiently handling both SE and PSE tasks, thereby simplifying deployment while maintaining high performance. However, PSE performance is sensitive to variations in enrollment speech, like emotional tone, which limits robustness in real-world applications. To address these challenges, we propose two novel models, USEF-PNet and DSEF-PNet, both extending our previous SEF-PNet framework. USEF-PNet introduces a unified architecture for processing enrollment speech, integrating SE and PSE into a single framework to enhance performance and streamline deployment. Meanwhile, DSEF-PNet incorporates an unsupervised speech disentanglement approach by pairing a mixture speech with two different enrollment utterances and enforcing consistency in the extracted target speech. This strategy effectively isolates high-quality speaker identity information from enrollment speech, reducing interference from factors such as emotion and content, thereby improving PSE robustness. Additionally, we explore a long-short enrollment pairing (LSEP) strategy to examine the impact of enrollment speech duration during both training and evaluation. Extensive experiments on the Libri2Mix and VoiceBank DEMAND demonstrate that our proposed USEF-PNet, DSEF-PNet all achieve substantial performance improvements, with random enrollment duration performing slightly better.

【8】 Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
标题: 用于粗到细文本到语音合成的浅层流匹配
链接:https://arxiv.org/abs/2505.12226
作者: Dong Yang,  Yiyi Cai,  Yuki Saito,  Lixu Wang,  Hiroshi Saruwatari 
摘要:我们提出了一个浅流匹配(SFM)机制,以增强流匹配(FM)为基础的文本到语音(TTS)模型在一个由粗到细的生成范式。SFM使用粗略的输出表示沿着FM路径构造中间状态。在训练过程中,我们引入了一种正交投影方法来自适应地确定这些状态的时间位置,并应用了基于单段分段流的原则性构造策略。SFM推断从中间状态而不是纯噪声开始,并且将计算集中在FM路径的后级上。我们将SFM集成到多个TTS模型中,并使用轻量级的SFM头。实验表明,SFM一贯提高合成语音的自然度在客观和主观评价,同时显着减少推理时,使用自适应步长ODE求解器。演示和代码可在https://ydqmkkx.github.io/SFMDemo/上获得。
摘要:We propose a shallow flow matching (SFM) mechanism to enhance flow matching (FM)-based text-to-speech (TTS) models within a coarse-to-fine generation paradigm. SFM constructs intermediate states along the FM paths using coarse output representations. During training, we introduce an orthogonal projection method to adaptively determine the temporal position of these states, and apply a principled construction strategy based on a single-segment piecewise flow. The SFM inference starts from the intermediate state rather than pure noise and focuses computation on the latter stages of the FM paths. We integrate SFM into multiple TTS models with a lightweight SFM head. Experiments show that SFM consistently improves the naturalness of synthesized speech in both objective and subjective evaluations, while significantly reducing inference when using adaptive-step ODE solvers. Demo and codes are available at https://ydqmkkx.github.io/SFMDemo/.

【9】 WaLRUS: Wavelets for Long-range Representation Using SSMs
标题: WaLRUS:使用SSD进行远程表示的波浪
链接:https://arxiv.org/abs/2505.12161
作者: Hossein Babaei,  Mel White,  Sina Alemohammad,  Richard G. Baraniuk 
备注:15 pages, 8 figures. Submitted to Neurips 2025
摘要:状态空间模型(SSM)已被证明是建模序列数据中的长期依赖关系的强大工具。虽然最近被称为HiPPO的方法表现出了强大的性能,并形成了机器学习模型S4和Mamba的基础,但它仍然受到一些特定的、行为良好的基础的封闭形式解决方案的限制。SaFARI框架推广了这种方法,使得能够从任意帧(包括非正交和冗余帧)构建SSM,从而允许SSM家族中可能的“物种”的无限多样性。在本文中,我们介绍了WaLRUS(Wavelength for Long-range Representation Using SSM),一个新的实现SaFARI建立从Daubechies小波。
摘要:State-Space Models (SSMs) have proven to be powerful tools for modeling long-range dependencies in sequential data. While the recent method known as HiPPO has demonstrated strong performance, and formed the basis for machine learning models S4 and Mamba, it remains limited by its reliance on closed-form solutions for a few specific, well-behaved bases. The SaFARi framework generalized this approach, enabling the construction of SSMs from arbitrary frames, including non-orthogonal and redundant ones, thus allowing an infinite diversity of possible "species" within the SSM family. In this paper, we introduce WaLRUS (Wavelets for Long-range Representation Using SSMs), a new implementation of SaFARi built from Daubechies wavelets.

【10】 BINAQUAL: A Full-Reference Objective Localization Similarity Metric for  Binaural Audio
标题: SINAQUAL:一种用于双耳音频的全参考客观定位相似性指标
链接:https://arxiv.org/abs/2505.11915
作者: Davoud Shariat Panah,  Dan Barry,  Alessandro Ragano,  Jan Skoglund,  Andrew Hines 
备注:Submitted to the Journal of Audio Engineering Society (JAES)
摘要:空间音频通过创建三维听觉体验来增强虚拟现实、增强现实、游戏和电影等应用的沉浸感。确保双耳音频的空间保真度至关重要,因为压缩、编码或传输等过程可能会改变定位线索。虽然像MUSHRA这样的主观听力测试仍然是评估空间定位质量的黄金标准,但它们成本高昂且耗时。本文介绍了BINAQUAL,一个完整的参考客观指标,旨在评估双耳录音定位相似性。BINAQUAL将AMBIQUAL度量(最初为在立体混响音频格式中进行定位质量评估而开发)调整为双耳域。我们评估BINAQUAL在五个关键的研究问题,检查其灵敏度的声源位置,角度插值,环绕扬声器布局,音频退化,和内容多样性的变化。结果表明,BINAQUAL有效地区分细微的空间变化,并与主观听力测试强烈相关,使其成为双耳定位质量评估的可靠指标。所提出的度量为确保双耳音频处理的空间准确性提供了一个强大的基准,为改善沉浸式音频应用中的客观评估铺平了道路。
摘要:Spatial audio enhances immersion in applications such as virtual reality, augmented reality, gaming, and cinema by creating a three-dimensional auditory experience. Ensuring the spatial fidelity of binaural audio is crucial, given that processes such as compression, encoding, or transmission can alter localization cues. While subjective listening tests like MUSHRA remain the gold standard for evaluating spatial localization quality, they are costly and time-consuming. This paper introduces BINAQUAL, a full-reference objective metric designed to assess localization similarity in binaural audio recordings. BINAQUAL adapts the AMBIQUAL metric, originally developed for localization quality assessment in ambisonics audio format to the binaural domain. We evaluate BINAQUAL across five key research questions, examining its sensitivity to variations in sound source locations, angle interpolations, surround speaker layouts, audio degradations, and content diversity. Results demonstrate that BINAQUAL effectively differentiates between subtle spatial variations and correlates strongly with subjective listening tests, making it a reliable metric for binaural localization quality assessment. The proposed metric provides a robust benchmark for ensuring spatial accuracy in binaural audio processing, paving the way for improved objective evaluations in immersive audio applications.

【11】 Exploring the Potential of SSL Models for Sound Event Detection
标题: 探索SSL模型用于声音事件检测的潜力
链接:https://arxiv.org/abs/2505.11889
作者: Hanfang Cui,  Longfei Song,  Li Li,  Dongxing Xu,  Yanhua Long 
备注:27 pages, 5 figures, submitted to the Journal of King Saud University - Computer and Information Sciences (under review)
摘要:自监督学习(SSL)模型为声音事件检测(SED)提供了强大的代表性,但其协同潜力仍然未得到充分挖掘。本研究系统地评估了最先进的SSL模型,以指导最佳的模型选择和集成SED。我们提出了一个框架,结合异构SSL表示(例如,Beats、HuBERT、WavLM)通过三种融合策略:单独SSL嵌入集成、双模式融合和完全聚合。DCASE 2023任务4挑战赛的实验表明,双模态融合(例如,CRNN+ BEAT +WavLM)实现了互补的性能提升,而CRNN+ BEAT单独提供了各个SSL模型中的最佳结果。我们进一步引入了归一化的声音事件边界框(nSEBB),这是一种自适应后处理方法,可以动态调整事件边界预测,将独立SSL模型的PSDS 1提高了4%。这些发现突出了SSL架构的兼容性和互补性,为特定任务的融合和强大的SED系统设计提供指导。
摘要:Self-supervised learning (SSL) models offer powerful representations for sound event detection (SED), yet their synergistic potential remains underexplored. This study systematically evaluates state-of-the-art SSL models to guide optimal model selection and integration for SED. We propose a framework that combines heterogeneous SSL representations (e.g., BEATs, HuBERT, WavLM) through three fusion strategies: individual SSL embedding integration, dual-modal fusion, and full aggregation. Experiments on the DCASE 2023 Task 4 Challenge reveal that dual-modal fusion (e.g., CRNN+BEATs+WavLM) achieves complementary performance gains, while CRNN+BEATs alone delivers the best results among individual SSL models. We further introduce normalized sound event bounding boxes (nSEBBs), an adaptive post-processing method that dynamically adjusts event boundary predictions, improving PSDS1 by up to 4% for standalone SSL models. These findings highlight the compatibility and complementarity of SSL architectures, providing guidance for task-specific fusion and robust SED system design.

【12】 AnalyticKWS: Towards Exemplar-Free Analytic Class Incremental Learning  for Small-footprint Keyword Spotting
标题: AnalytyKWS:迈向无示例分析类增量学习,以实现小规模关键词发现
链接:https://arxiv.org/abs/2505.11817
作者: Yang Xiao,  Tianyi Peng,  Rohan Kumar Das,  Yuchen Hu,  Huiping Zhuang 
备注:Accepted by ACL 2025
摘要:关键词定位(KWS)提供了一种重要的机制来识别语音系统中的口头命令,在语音系统中,用户的需求经常变化,需要模型随着时间的推移不断学习新的关键词。然而,一个主要的问题是灾难性的遗忘,模型失去了识别早期关键字的能力。虽然一些持续学习方法已经证明了它们在减少遗忘方面的有效性,但大多数现有方法都依赖于存储和重新访问旧数据来对抗灾难性遗忘。虽然有效,但这些方法面临两个实际挑战:1)保留用户数据的隐私风险和2)限制在小型设备上部署的大内存和时间消耗。为了解决这些问题,我们提出了一个无范例的分析持续学习(AnalyticKWS)方法,更新模型参数,而无需重新访问以前的数据。受高效学习原则的启发,AnalyticKWS计算模型更新的封闭形式解析解,并且只需要对传入的关键字进行单次适应。AnalyticKWS通过避免基于梯度的更新而需要更少的计算资源,并且不存储旧数据。通过消除增量学习过程中对反向传播的需求,该模型保持了轻量级和高效。因此,AnalyticKWS满足了前面提到的挑战,并且非常适合资源有限的环境。在各种数据集和设置上的广泛实验表明,AnalyticKWS始终优于现有的持续学习方法。
摘要:Keyword spotting (KWS) offers a vital mechanism to identify spoken commands in voice-enabled systems, where user demands often shift, requiring models to learn new keywords continually over time. However, a major problem is catastrophic forgetting, where models lose their ability to recognize earlier keywords. Although several continual learning methods have proven their usefulness for reducing forgetting, most existing approaches depend on storing and revisiting old data to combat catastrophic forgetting. Though effective, these methods face two practical challenges: 1) privacy risks from keeping user data and 2) large memory and time consumption that limit deployment on small devices. To address these issues, we propose an exemplar-free Analytic Continual Learning (AnalyticKWS) method that updates model parameters without revisiting earlier data. Inspired by efficient learning principles, AnalyticKWS computes a closed-form analytical solution for model updates and requires only a single epoch of adaptation for incoming keywords. AnalyticKWS demands fewer computational resources by avoiding gradient-based updates and does not store old data. By eliminating the need for back-propagation during incremental learning, the model remains lightweight and efficient. As a result, AnalyticKWS meets the challenges mentioned earlier and suits resource-limited settings well. Extensive experiments on various datasets and settings show that AnalyticKWS consistently outperforms existing continual learning methods.

【13】 Granary: Speech Recognition and Translation Dataset in 25 European  Languages
标题: 粮仓:25种欧洲语言的语音识别和翻译数据集
链接:https://arxiv.org/abs/2505.13404
作者: Nithin Rao Koluguri,  Monica Sekoyan,  George Zelenfroynd,  Sasha Meister,  Shuoyang Ding,  Sofia Kostandian,  He Huang,  Nikolay Karpov,  Jagadeesh Balam,  Vitaly Lavrukhin,  Yifan Peng,  Sara Papi,  Marco Gaido,  Alessio Brutti,  Boris Ginsburg 
备注:Accepted at Interspeech 2025
摘要:多任务和多语言方法使大型模型受益,但由于数据稀缺,低资源语言的语音处理仍然没有得到充分探索。为了解决这个问题,我们提出了Granary,这是一个大规模的语音数据集集合,用于识别和翻译25种欧洲语言。这是第一次在转录和翻译方面进行这种规模的开源努力。我们使用具有分割、两遍推理、幻觉过滤和标点恢复的伪标记管道来提高数据质量。我们进一步使用EuroLLM从伪标记的transmittance生成翻译对,然后是数据过滤管道。我们的管道专为提高效率而设计,可在数小时内处理大量数据。我们通过比较它们在高资源和低资源语言的先前策划的数据集上的性能来评估在处理数据上训练的模型。我们的研究结果表明,这些模型实现了类似的性能使用约。数据减少50%数据集将在https://hf.co/datasets/nvidia/Granary上提供
摘要:Multi-task and multilingual approaches benefit large models, yet speech processing for low-resource languages remains underexplored due to data scarcity. To address this, we present Granary, a large-scale collection of speech datasets for recognition and translation across 25 European languages. This is the first open-source effort at this scale for both transcription and translation. We enhance data quality using a pseudo-labeling pipeline with segmentation, two-pass inference, hallucination filtering, and punctuation restoration. We further generate translation pairs from pseudo-labeled transcriptions using EuroLLM, followed by a data filtration pipeline. Designed for efficiency, our pipeline processes vast amount of data within hours. We assess models trained on processed data by comparing their performance on previously curated datasets for both high- and low-resource languages. Our findings show that these models achieve similar performance using approx. 50% less data. Dataset will be made available at https://hf.co/datasets/nvidia/Granary

【14】 Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data  Condensation and Spoken QA Generation
标题: 用于多模式语音的上下文副语言数据创建-LLM:数据压缩和口语QA生成
链接:https://arxiv.org/abs/2505.13338
作者: Qiongqiong Wang,  Hardik B. Sailor,  Tianchi Liu,  Ai Ti Aw 
备注:Accepted at Interspeech 2025
摘要:目前的语音LLM在上下文推理以及语言理解方面表现出有限的能力,主要是由于缺乏涵盖这两个方面的问答(QA)数据集。我们提出了一个新的框架,数据集生成在野生语音数据,集成了上下文推理与语言信息。它包括一个伪的基于标签的数据压缩的野生语音和基于LLM的上下文副语言QA(CPQA)的生成。Qwen 2-Audio-7 B-Instruct模型在由我们的框架创建的数据集和人工生成的CPQA数据集上的评估具有很强的相关性,从而验证了其有效性。结果还揭示了语音LLM在处理移情推理任务方面的局限性,突出了对此类数据集和更强大模型的需求。所提出的框架是第一个同类的,并有潜力在训练更强大的语音LLM语言推理能力。
摘要:Current speech-LLMs exhibit limited capability in contextual reasoning alongside paralinguistic understanding, primarily due to the lack of Question-Answer (QA) datasets that cover both aspects. We propose a novel framework for dataset generation from in-the-wild speech data, that integrates contextual reasoning with paralinguistic information. It consists of a pseudo paralinguistic label-based data condensation of in-the-wild speech and LLM-based Contextual Paralinguistic QA (CPQA) generation. The effectiveness is validated by a strong correlation in evaluations of the Qwen2-Audio-7B-Instruct model on a dataset created by our framework and human-generated CPQA dataset. The results also reveal the speech-LLM's limitations in handling empathetic reasoning tasks, highlighting the need for such datasets and more robust models. The proposed framework is first of its kind and has potential in training more robust speech-LLMs with paralinguistic reasoning capabilities.

【15】 Distilling a speech and music encoder with task arithmetic
标题: 使用任务算法提取语音和音乐编码器
链接:https://arxiv.org/abs/2505.13270
作者: Fabian Ritter-Gutierrez,  Yi-Cheng Lin,  Jui-Chiang Wei,  Jeremy H.M Wong,  Eng Siong Chng,  Nancy F. Chen,  Hung-yi Lee 
备注:Accepted at INTERSPEECH 2025
摘要:尽管语音和音乐的自监督学习(SSL)取得了进展,但现有模型将这些领域分开处理,限制了它们统一音频理解的能力。统一模型对于需要通用表示的应用是理想的,例如音频大语言模型。尽管如此,直接训练语音和音乐的通用模型在计算上是昂贵的。教师合奏的知识蒸馏可能是一个自然的解决方案,但我们认为,分离语音和音乐SSL模型的蒸馏允许更大的灵活性。因此,我们建议学习提取的任务向量,然后对其进行线性插值,以形成统一的语音+音乐模型。该策略通过可调权重实现灵活的域强调,并且训练也更简单。语音和音乐基准的实验表明,我们的方法产生优越的整体性能相比,集成蒸馏。
摘要:Despite the progress in self-supervised learning (SSL) for speech and music, existing models treat these domains separately, limiting their capacity for unified audio understanding. A unified model is desirable for applications that require general representations, e.g. audio large language models. Nonetheless, directly training a general model for speech and music is computationally expensive. Knowledge Distillation of teacher ensembles may be a natural solution, but we posit that decoupling the distillation of the speech and music SSL models allows for more flexibility. Thus, we propose to learn distilled task vectors and then linearly interpolate them to form a unified speech+music model. This strategy enables flexible domain emphasis through adjustable weights and is also simpler to train. Experiments on speech and music benchmarks demonstrate that our method yields superior overall performance compared to ensemble distillation.

【16】 Efficient Speech Language Modeling via Energy Distance in Continuous  Latent Space
标题: 连续潜在空间中通过能量距离进行高效语音语言建模
链接:https://arxiv.org/abs/2505.13181
作者: Zhengrui Ma,  Yang Feng,  Chenze Shao,  Fandong Meng,  Jie Zhou,  Min Zhang 
备注:Demos and code are available at this https URL
摘要:我们介绍SLED,语音语言建模的另一种方法,通过将语音波形编码成连续的潜在表示序列,并使用能量距离目标对其进行自回归建模。能量距离通过对比模拟样本和目标样本提供了分布差距的分析度量,从而实现有效的训练以捕获潜在的连续自回归分布。通过绕过对残差矢量量化的依赖,SLED避免了离散化错误,并消除了对现有语音语言模型中常见的复杂分层结构的需要。它简化了整个建模流程,同时保留了语音信息的丰富性并保持了推理效率。实验结果表明,SLED在zero-shot和流式语音合成中均取得了较好的性能,显示了其在通用语音语言模型中的广泛应用潜力。
摘要:We introduce SLED, an alternative approach to speech language modeling by encoding speech waveforms into sequences of continuous latent representations and modeling them autoregressively using an energy distance objective. The energy distance offers an analytical measure of the distributional gap by contrasting simulated and target samples, enabling efficient training to capture the underlying continuous autoregressive distribution. By bypassing reliance on residual vector quantization, SLED avoids discretization errors and eliminates the need for the complicated hierarchical architectures common in existing speech language models. It simplifies the overall modeling pipeline while preserving the richness of speech information and maintaining inference efficiency. Empirical results demonstrate that SLED achieves strong performance in both zero-shot and streaming speech synthesis, showing its potential for broader applications in general-purpose speech language models.

【17】 Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
标题: 用于时态推理的LALM的基准和置信度评估
链接:https://arxiv.org/abs/2505.13115
作者: Debarpan Bhattacharya,  Apoorva Kulkarni,  Sriram Ganapathy 
备注:Accepted in INTERSPEECH, 2025, Rotterdam, The Netherlands
摘要:基于文本的大型语言模型(LLM)的流行成功简化了多模态社区的注意力,将视觉和音频等其他模态与文本结合起来,以实现类似的多模态功能。在这一探索中,大型音频语言模型(LALM)必须评估推理相关的任务,这是不同于传统的分类或生成任务。为了实现这一目标,我们提出了一种新的数据集,称为时间推理评估音频(TREA)。   我们对开源LALM进行了基准测试,并观察到它们在TREA数据集中的任务上始终落后于人类能力。在评估LALM的同时,我们还提出了一个不确定性度量,该度量计算模型对输入的语义相同扰动的不变性。我们的分析表明,准确性和不确定性指标不一定是相关的,因此,点需要健康的评估LALM高风险的应用程序。
摘要:The popular success of text-based large language models (LLM) has streamlined the attention of the multimodal community to combine other modalities like vision and audio along with text to achieve similar multimodal capabilities. In this quest, large audio language models (LALMs) have to be evaluated on reasoning related tasks which are different from traditional classification or generation tasks. Towards this goal, we propose a novel dataset called temporal reasoning evaluation of audio (TREA).   We benchmark open-source LALMs and observe that they are consistently behind human capabilities on the tasks in the TREA dataset. While evaluating LALMs, we also propose an uncertainty metric, which computes the invariance of the model to semantically identical perturbations of the input. Our analysis shows that the accuracy and uncertainty metrics are not necessarily correlated and thus, points to a need for wholesome evaluation of LALMs for high-stakes applications.

【18】 Time-Frequency-Based Attention Cache Memory Model for Real-Time Speech  Separation
标题: 基于时频的实时语音分离注意力缓存模型
链接:https://arxiv.org/abs/2505.13094
作者: Guo Chen,  Kai Li,  Runxuan Yang,  Xiaolin Hu 
摘要:现有的因果语音分离模型往往表现不佳相比,非因果模型,由于在保留历史信息的困难。为了解决这个问题,我们提出了时间-频率注意力缓存(TFACM)模型,它有效地捕捉时空关系,通过注意力机制和缓存(CM)的历史信息存储。在TFACM中,LSTM层捕获频率相对位置,而因果建模则使用局部和全局表示应用于时间维度。CM模块存储过去的信息,并且因果注意细化(CAR)模块进一步增强基于时间的特征表示以获得更精细的粒度。实验结果表明,TFACM需要与SOTA TF-GridNet-Causal模型相当的性能,具有显着更低的复杂性和更少的可训练参数。有关更多详细信息,请访问项目页面:https://cslikai.cn/TFACM/。
摘要:Existing causal speech separation models often underperform compared to non-causal models due to difficulties in retaining historical information. To address this, we propose the Time-Frequency Attention Cache Memory (TFACM) model, which effectively captures spatio-temporal relationships through an attention mechanism and cache memory (CM) for historical information storage. In TFACM, an LSTM layer captures frequency-relative positions, while causal modeling is applied to the time dimension using local and global representations. The CM module stores past information, and the causal attention refinement (CAR) module further enhances time-based feature representations for finer granularity. Experimental results showed that TFACM achieveed comparable performance to the SOTA TF-GridNet-Causal model, with significantly lower complexity and fewer trainable parameters. For more details, visit the project page: https://cslikai.cn/TFACM/.

【19】 MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and  Voices of Multiple Speakers
标题: MultiActor有声读物:具有多个发言者的面孔和声音的Zero-Shot有声读物生成
链接:https://arxiv.org/abs/2505.13082
作者: Kyeongman Park,  Seongho Joo,  Kyomin Jung 
摘要:我们介绍了多演员有声读物,一个zero-shot的方法来生成有声读物,自动产生一致的,富有表现力的,和扬声器适当的韵律,包括语调和情感。以前的有声读物系统有几个局限性:它们需要用户手动配置说话者的韵律,与配音演员相比,用单调的音调阅读每个句子,或者依赖昂贵的培训。然而,我们的多演员有声读物通过引入两个新的过程来解决这些问题:(1)MSP(** 多模态扬声器角色生成 **)和(2)LSI(** 基于LLM的脚本指令生成 **)。通过这两个过程,MultiActor-Audiobook可以生成更具情感表达力的有声读物,具有一致的说话者韵律,而无需额外的训练。我们将我们的系统与商业产品进行比较,通过人体和MLLM评估,获得具有竞争力的结果。此外,我们通过消融研究证明了MSP和LSI的有效性。
摘要:We introduce MultiActor-Audiobook, a zero-shot approach for generating audiobooks that automatically produces consistent, expressive, and speaker-appropriate prosody, including intonation and emotion. Previous audiobook systems have several limitations: they require users to manually configure the speaker's prosody, read each sentence with a monotonic tone compared to voice actors, or rely on costly training. However, our MultiActor-Audiobook addresses these issues by introducing two novel processes: (1) MSP (**Multimodal Speaker Persona Generation**) and (2) LSI (**LLM-based Script Instruction Generation**). With these two processes, MultiActor-Audiobook can generate more emotionally expressive audiobooks with a consistent speaker prosody without additional training. We compare our system with commercial products, through human and MLLM evaluations, achieving competitive results. Furthermore, we demonstrate the effectiveness of MSP and LSI through ablation studies.

【20】 Suicide Risk Assessment Using Multimodal Speech Features: A Study on the  SW1 Challenge Dataset
标题: 使用多模式语音特征的自杀风险评估:SW 1挑战数据集的研究
链接:https://arxiv.org/abs/2505.13069
作者: Ambre Marie,  Ilias Maoudj,  Guillaume Dardenne,  Gwenolé Quellec 
备注:Submitted to the SpeechWellness Challenge at Interspeech 2025; 5 pages, 2 figures, 2 tables
摘要:第一届SpeechWellness Challenge传达了对青少年进行基于语言的自杀风险评估的必要性。本研究针对这一挑战研究了一种多模态方法,将自动转录与WhisperX、来自中国RoberTa的语言嵌入和来自WavLM的音频嵌入相结合。此外,手工制作的声学功能-包括MFCC,频谱对比度和音高相关的统计-被纳入。我们探索了三种融合策略:早期拼接,特定模态处理和混合正则化加权注意力。结果表明,加权注意力提供了最好的泛化能力,在开发集上达到了69%的准确率,尽管开发集和测试集之间的性能差距突出了泛化的挑战。我们的研究结果,严格绑定到MINI-KID框架,强调细化嵌入表示和融合机制,以提高分类可靠性的重要性。
摘要:The 1st SpeechWellness Challenge conveys the need for speech-based suicide risk assessment in adolescents. This study investigates a multimodal approach for this challenge, integrating automatic transcription with WhisperX, linguistic embeddings from Chinese RoBERTa, and audio embeddings from WavLM. Additionally, handcrafted acoustic features -- including MFCCs, spectral contrast, and pitch-related statistics -- were incorporated. We explored three fusion strategies: early concatenation, modality-specific processing, and weighted attention with mixup regularization. Results show that weighted attention provided the best generalization, achieving 69% accuracy on the development set, though a performance gap between development and test sets highlights generalization challenges. Our findings, strictly tied to the MINI-KID framework, emphasize the importance of refining embedding representations and fusion mechanisms to enhance classification reliability.
【21】 Hearing from Silence: Reasoning Audio Descriptions from Silent Videos  via Vision-Language Model
标题: 从沉默中倾听:通过视觉语言模型从无声视频中推理音频描述
链接:https://arxiv.org/abs/2505.13062
作者: Yong Ren,  Chenxing Li,  Le Xu,  Hao Gu,  Duzhen Zhang,  Yujie Chen,  Manjie Xu,  Ruibo Fu,  Shan Yang,  Dong Yu 
备注:Accepted by Interspeech 2025
摘要:人类可以从无声视频中直观地推断出声音,但多模态大型语言模型是否可以在不访问目标模态的情况下执行模态失配推理仍然相对未被探索。目前的文本辅助视频到音频(VT 2A)的方法在视频Foley任务中表现出色,但在推理过程中难以获得音频描述。我们引入了从无声视频中推理音频描述(SVAD)的任务来解决这一挑战,并研究了视觉语言模型(VLM)在这一任务上的能力。为了进一步增强VLM在SVAD任务中的推理能力,我们构建了一个CoT-AudioCaps数据集,并提出了一种基于权值链的监督微调策略。SVAD和随后的VT 2A任务的实验表明,我们的方法的有效性在两个关键方面:显着提高VLM的模态失配推理SVAD和有效地解决在VT 2A推理过程中获取音频描述的挑战。
摘要:Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remains relatively unexplored. Current text-assisted-video-to-audio (VT2A) methods excel in video foley tasks but struggle to acquire audio descriptions during inference. We introduce the task of Reasoning Audio Descriptions from Silent Videos (SVAD) to address this challenge and investigate vision-language models' (VLMs) capabilities on this task. To further enhance the VLMs' reasoning capacity for the SVAD task, we construct a CoT-AudioCaps dataset and propose a Chain-of-Thought-based supervised fine-tuning strategy. Experiments on SVAD and subsequent VT2A tasks demonstrate our method's effectiveness in two key aspects: significantly improving VLMs' modal-mismatch reasoning for SVAD and effectively addressing the challenge of acquiring audio descriptions during VT2A inference.

【22】 MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio,  Music, and Their Mix
标题: MVAR:语音、音频、音乐及其混合中深度推理的领先基准
链接:https://arxiv.org/abs/2505.13032
作者: Ziyang Ma,  Yinghao Ma,  Yanqiao Zhu,  Chen Yang,  Yi-Wen Chao,  Ruiyang Xu,  Wenxi Chen,  Yuanzhe Chen,  Zhuo Chen,  Jian Cong,  Kai Li,  Keliang Li,  Siyou Li,  Xinfeng Li,  Xiquan Li,  Zheng Lian,  Yuzhe Liang,  Minghao Liu,  Zhikang Niu,  Tianrui Wang,  Yuping Wang,  Yuxuan Wang,  Yihao Wu,  Guanrou Yang,  Jianwei Yu,  Ruibin Yuan,  Zhisheng Zheng,  Ziya Zhou,  Haina Zhu,  Wei Xue,  Emmanouil Benetos,  Kai Yu,  Eng-Siong Chng,  Xie Chen 
备注:Open-source at this https URL
摘要:我们引入了MMAR,这是一个新的基准测试,旨在评估音频语言模型(ALM)在大规模多学科任务中的深度推理能力。MMAR由1,000个精心策划的音频问答三重奏组成,这些音频问答三重奏从现实世界的互联网视频中收集,并通过迭代错误纠正和质量检查进行改进,以确保高质量。与现有的仅限于声音、音乐或语音的特定领域的基准测试不同,MMAR将它们扩展到广泛的真实世界音频场景,包括声音、音乐和语音的混合模态组合。MMAR中的每个问题都分为四个推理层:信号,感知,语义和文化,每个层中都有额外的子类别,以反映任务的多样性和复杂性。为了进一步促进这一领域的研究,我们用思想链(CoT)理论来注释每个问题,以促进音频推理的未来发展。基准测试中的每个项目都需要进行多步深入推理,而不仅仅是表面的理解。此外,部分问题需要研究生水平的感性和特定领域的知识,提高了基准的难度和深度。我们使用一组广泛的模型来评估MMAR,包括大型音频语言模型(LALM),大型音频推理模型(LARM),Omni语言模型(OLM),大型语言模型(LLM)和大型推理模型(LRM),并带有音频字幕输入。这些模型在MMAR上的表现突出了基准测试的挑战性,我们的分析进一步揭示了当前模型在理解和推理能力方面的关键局限性。我们希望MMAR将成为这一重要但很少探索的领域未来进步的催化剂。
摘要:We introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets, collected from real-world internet videos and refined through iterative error corrections and quality checks to ensure high quality. Unlike existing benchmarks that are limited to specific domains of sound, music, or speech, MMAR extends them to a broad spectrum of real-world audio scenarios, including mixed-modality combinations of sound, music, and speech. Each question in MMAR is hierarchically categorized across four reasoning layers: Signal, Perception, Semantic, and Cultural, with additional sub-categories within each layer to reflect task diversity and complexity. To further foster research in this area, we annotate every question with a Chain-of-Thought (CoT) rationale to promote future advancements in audio reasoning. Each item in the benchmark demands multi-step deep reasoning beyond surface-level understanding. Moreover, a part of the questions requires graduate-level perceptual and domain-specific knowledge, elevating the benchmark's difficulty and depth. We evaluate MMAR using a broad set of models, including Large Audio-Language Models (LALMs), Large Audio Reasoning Models (LARMs), Omni Language Models (OLMs), Large Language Models (LLMs), and Large Reasoning Models (LRMs), with audio caption inputs. The performance of these models on MMAR highlights the benchmark's challenging nature, and our analysis further reveals critical limitations of understanding and reasoning capabilities among current models. We hope MMAR will serve as a catalyst for future advances in this important but little-explored area.

【23】 DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec  for Speech Generation
标题: DualCodec:一种用于语音生成的低帧率、语义增强的神经音频编解码器
链接:https://arxiv.org/abs/2505.13000
作者: Jiaqi Li,  Xiaolong Lin,  Zhekai Li,  Shixi Huang,  Yuancheng Wang,  Chaoren Wang,  Zhenpeng Zhan,  Zhizheng Wu 
备注:Accepted to Interspeech 2025. Github: this https URL
摘要:神经音频编解码器形成基于语言模型(LM)的语音生成的基础构建块。通常,在帧速率和音频质量之间存在折衷。本研究提出一种低帧率、语意增强的编解码器模型。现有方法将语义丰富的自监督(SSL)表示提取到第一层编解码器令牌中。这项工作提出了DualCodec,一个双流编码方法,集成了SSL和波形表示在一个端到端的编解码器框架。在这种设置下,DualCodec增强了第一层编解码器中的语义信息,使编解码器系统能够在低帧速率下运行时保持高音频质量。注意,低帧速率编解码器提高了语音生成的效率。音频编解码器和语音生成任务的实验结果证实了所提出的DualCodec的有效性相比,国家的最先进的编解码器系统,如Mimi编解码器,SpeechTokenizer,DAC,和Encodec。演示和代码可在https://dualcodec.github.io上获得
摘要:Neural audio codecs form the foundational building blocks for language model (LM)-based speech generation. Typically, there is a trade-off between frame rate and audio quality. This study introduces a low-frame-rate, semantically enhanced codec model. Existing approaches distill semantically rich self-supervised (SSL) representations into the first-layer codec tokens. This work proposes DualCodec, a dual-stream encoding approach that integrates SSL and waveform representations within an end-to-end codec framework. In this setting, DualCodec enhances the semantic information in the first-layer codec and enables the codec system to maintain high audio quality while operating at a low frame rate. Note that a low-frame-rate codec improves the efficiency of speech generation. Experimental results on audio codec and speech generation tasks confirm the effectiveness of the proposed DualCodec compared to state-of-the-art codec systems, such as Mimi Codec, SpeechTokenizer, DAC, and Encodec. Demos and codes are available at: https://dualcodec.github.io

【24】 Codec-Based Deepfake Source Tracing via Neural Audio Codec Taxonomy
标题: 通过神经音频编解码分类法基于编解码器的Deepfake源跟踪
链接:https://arxiv.org/abs/2505.12994
作者: Xuanjun Chen,  I-Ming Lin,  Lin Zhang,  Jiawei Du,  Haibin Wu,  Hung-yi Lee,  Jyh-Shing Roger Jang 
备注:Accepted by Interspeech 2025
摘要:基于神经音频编解码器的语音生成(CoSG)模型的最新进展已经产生了非常逼真的音频deepfake。我们将CoSG系统生成的deepfake语音称为基于编解码器的deepfake或CodecFake。尽管现有的CodecFake反欺骗研究主要集中在验证音频样本的真实性上,但几乎没有注意跟踪用于生成这些deepfake的CoSG。在CodecFake生成中,语音到单元编码、离散单元建模和单元到语音解码等过程基本上基于神经音频编解码器。受此启发,我们通过神经音频编解码器分类法引入CodecFake的源跟踪,该分类法剖析神经音频编解码器以跟踪CoSG。我们在CodecFake+数据集上的实验结果为CodecFake源跟踪的可行性提供了有希望的初步证据,同时也强调了需要进一步研究的几个挑战。
摘要:Recent advances in neural audio codec-based speech generation (CoSG) models have produced remarkably realistic audio deepfakes. We refer to deepfake speech generated by CoSG systems as codec-based deepfake, or CodecFake. Although existing anti-spoofing research on CodecFake predominantly focuses on verifying the authenticity of audio samples, almost no attention was given to tracing the CoSG used in generating these deepfakes. In CodecFake generation, processes such as speech-to-unit encoding, discrete unit modeling, and unit-to-speech decoding are fundamentally based on neural audio codecs. Motivated by this, we introduce source tracing for CodecFake via neural audio codec taxonomy, which dissects neural audio codecs to trace CoSG. Our experimental results on the CodecFake+ dataset provide promising initial evidence for the feasibility of CodecFake source tracing while also highlighting several challenges that warrant further investigation.

【25】 Personalized Fine-Tuning with Controllable Synthetic Speech from  LLM-Generated Transcripts for Dysarthric Speech Recognition
标题: 利用LLM生成的脚本中的可控合成语音进行个性化微调,用于发音障碍语音识别
链接:https://arxiv.org/abs/2505.12991
作者: Dominik Wagner,  Ilja Baumann,  Natalie Engert,  Seanie Lee,  Elmar Nöth,  Korbinian Riedhammer,  Tobias Bocklet 
备注:Accepted at Interspeech 2025
摘要:在这项工作中,我们提出了我们提交的语音无障碍项目的挑战构音障碍语音识别。我们将参数有效的微调与潜在的音频表示相结合,以改善编码器-解码器ASR系统。合成训练数据是通过微调Parler-TTS来模拟构音障碍的语音,使用LLM生成的语料库一致的目标成绩单的提示。与非个性化微调相比,使用x向量的个性化始终降低了单词错误率(WER)。AdaLoRA适配器的性能优于完全微调和标准低秩适配,分别实现了约23%和约22%的相对WER降低。进一步的改进(约5%的WER降低)来自于整合基于wav2vec 2.0的音频表示。与单独的个性化微调相比,使用合成构音障碍语音的训练产生高达~7%的相对WER改善。
摘要:In this work, we present our submission to the Speech Accessibility Project challenge for dysarthric speech recognition. We integrate parameter-efficient fine-tuning with latent audio representations to improve an encoder-decoder ASR system. Synthetic training data is generated by fine-tuning Parler-TTS to mimic dysarthric speech, using LLM-generated prompts for corpus-consistent target transcripts. Personalization with x-vectors consistently reduces word error rates (WERs) over non-personalized fine-tuning. AdaLoRA adapters outperform full fine-tuning and standard low-rank adaptation, achieving relative WER reductions of ~23% and ~22%, respectively. Further improvements (~5% WER reduction) come from incorporating wav2vec 2.0-based audio representations. Training with synthetic dysarthric speech yields up to ~7% relative WER improvement over personalized fine-tuning alone.

【26】 The Computation of Generalized Embeddings for Underwater Acoustic Target  Recognition using Contrastive Learning
标题: 基于对比学习的水声目标识别中广义嵌入的计算
链接:https://arxiv.org/abs/2505.12904
作者: Hilde I. Hummel,  Arwin Gansekoele,  Sandjai Bhulai,  Rob van der Mei 
摘要:海洋环境中日益严重的声音污染对海洋健康构成越来越大的威胁,因此监测水下噪音至关重要。通过监测这种噪音,可以绘制出造成这种污染的来源。通过被动地听这些声音来执行监测。这会产生大量的数据记录,捕获混合的声源,如船舶活动和海洋哺乳动物的发声。虽然机器学习为自动声音分类提供了一个很有前途的解决方案,但目前最先进的方法实现了监督学习。这需要大量高质量的标记数据,这些数据不是公开的。相比之下,大量低质量的未标记数据是公开的,这为探索无监督学习技术提供了机会。本研究通过实施无监督对比学习方法来探索这种可能性。在这里,基于Conformer的编码器通过所谓的方差-不变性-协方差正则化损失函数对这些较低质量的未标记数据进行优化,并转换为标记数据。通过识别船舶类型和海洋哺乳动物发声的分类任务,我们的方法证明了产生强大的和广义的嵌入。这显示了各种自动水声分析任务的无监督方法的潜力。
摘要:The increasing level of sound pollution in marine environments poses an increased threat to ocean health, making it crucial to monitor underwater noise. By monitoring this noise, the sources responsible for this pollution can be mapped. Monitoring is performed by passively listening to these sounds. This generates a large amount of data records, capturing a mix of sound sources such as ship activities and marine mammal vocalizations. Although machine learning offers a promising solution for automatic sound classification, current state-of-the-art methods implement supervised learning. This requires a large amount of high-quality labeled data that is not publicly available. In contrast, a massive amount of lower-quality unlabeled data is publicly available, offering the opportunity to explore unsupervised learning techniques. This research explores this possibility by implementing an unsupervised Contrastive Learning approach. Here, a Conformer-based encoder is optimized by the so-called Variance-Invariance-Covariance Regularization loss function on these lower-quality unlabeled data and the translation to the labeled data is made. Through classification tasks involving recognizing ship types and marine mammal vocalizations, our method demonstrates to produce robust and generalized embeddings. This shows to potential of unsupervised methods for various automatic underwater acoustic analysis tasks.

【27】 Unified Cross-modal Translation of Score Images, Symbolic Music, and  Performance Audio
标题: 乐谱图像、象征性音乐和表演音频的统一跨模式翻译
链接:https://arxiv.org/abs/2505.12863
作者: Jongmin Jung,  Dongmin Kim,  Sihun Lee,  Seola Cho,  Hyungjoon Soh,  Irmak Bukey,  Chris Donahue,  Dasaem Jeong 
备注:Submitted to IEEE Transactions on Audio, Speech and Language Processing (TASLPRO)
摘要:音乐以各种形式存在,例如乐谱图像、符号乐谱、乐谱和音频。每种模态之间的翻译被确立为音乐信息检索的核心任务,例如自动音乐转录(音频到音频)和光学音乐识别(乐谱图像到符号乐谱)。然而,过去大多数关于多模态翻译的工作都是针对单个翻译任务训练专门的模型。在本文中,我们提出了一个统一的方法,我们训练一个通用的模型上的许多翻译任务同时进行。两个关键因素使这种统一的方法可行:一个新的大规模数据集和每个模态的标记化。首先,我们提出了一个新的数据集,由从YouTube视频中收集的超过1,300小时的配对音频图像数据组成,这比任何现有的音乐模态翻译数据集都要大一个数量级。其次,我们的统一标记化框架将乐谱图像、音频、乐谱和MusicXML离散化为标记序列,使单个编码器-解码器Transformer能够将多个跨模态翻译作为一个连贯的序列到序列任务来处理。实验结果证实,我们的统一多任务模型在几个关键领域比单任务基线有所改进,特别是将光学音乐识别的符号错误率从24.58%降低到最先进的13.67%,而在其他翻译任务中也观察到了类似的实质性改进。值得注意的是,我们的方法实现了第一次成功的分数图像调节音频生成,标志着跨模态音乐生成的重大突破。
摘要:Music exists in various modalities, such as score images, symbolic scores, MIDI, and audio. Translations between each modality are established as core tasks of music information retrieval, such as automatic music transcription (audio-to-MIDI) and optical music recognition (score image to symbolic score). However, most past work on multimodal translation trains specialized models on individual translation tasks. In this paper, we propose a unified approach, where we train a general-purpose model on many translation tasks simultaneously. Two key factors make this unified approach viable: a new large-scale dataset and the tokenization of each modality. Firstly, we propose a new dataset that consists of more than 1,300 hours of paired audio-score image data collected from YouTube videos, which is an order of magnitude larger than any existing music modal translation datasets. Secondly, our unified tokenization framework discretizes score images, audio, MIDI, and MusicXML into a sequence of tokens, enabling a single encoder-decoder Transformer to tackle multiple cross-modal translation as one coherent sequence-to-sequence task. Experimental results confirm that our unified multitask model improves upon single-task baselines in several key areas, notably reducing the symbol error rate for optical music recognition from 24.58% to a state-of-the-art 13.67%, while similarly substantial improvements are observed across the other translation tasks. Notably, our approach achieves the first successful score-image-conditioned audio generation, marking a significant breakthrough in cross-modal music generation.

【28】 OZSpeech: One-step Zero-shot Speech Synthesis with  Learned-Prior-Conditioned Flow Matching
标题: OZSpeech:具有学习先验条件流匹配的一步零激发语音合成
链接:https://arxiv.org/abs/2505.12800
作者: Hieu-Nghia Huynh-Nguyen,  Ngoc Son Nguyen,  Huynh Nguyen Dang,  Thieu Vo,  Truong-Son Hy,  Van Nguyen 
摘要:近年来,在深度学习和神经网络架构的改进的推动下,文本到语音(TTS)系统取得了重大进展。将输出语音视为数据分布,以前的方法通常在流匹配框架内采用传统的语音表示,例如波形或频谱图。然而,这些方法具有局限性,包括忽略各种语音属性,以及由于在训练期间引入的额外约束而导致高计算成本。为了解决这些挑战,我们引入了OZSpeech,这是第一种TTS方法,可以探索最佳的传输条件流匹配,并以一步采样和先验知识为条件,有效地忽略了先前的状态并减少了采样步骤的数量。我们的方法在令牌格式的语音的分解,分解组件上操作,使得每个语音属性的准确建模成为可能,这增强了TTS系统精确克隆提示语音的能力。实验结果表明,我们的方法在内容准确性,自然度,韵律生成和说话人风格保持方面取得了令人满意的性能。音频样本可在我们的演示页面https://ozspeech.github.io/OZSpeech_Web/上获得。
摘要:Text-to-speech (TTS) systems have seen significant advancements in recent years, driven by improvements in deep learning and neural network architectures. Viewing the output speech as a data distribution, previous approaches often employ traditional speech representations, such as waveforms or spectrograms, within the Flow Matching framework. However, these methods have limitations, including overlooking various speech attributes and incurring high computational costs due to additional constraints introduced during training. To address these challenges, we introduce OZSpeech, the first TTS method to explore optimal transport conditional flow matching with one-step sampling and a learned prior as the condition, effectively disregarding preceding states and reducing the number of sampling steps. Our approach operates on disentangled, factorized components of speech in token format, enabling accurate modeling of each speech attribute, which enhances the TTS system's ability to precisely clone the prompt speech. Experimental results show that our method achieves promising performance over existing methods in content accuracy, naturalness, prosody generation, and speaker style preservation. Audio samples are available at our demo page https://ozspeech.github.io/OZSpeech_Web/.

【29】 SounDiT: Geo-Contextual Soundscape-to-Landscape Generation
标题: SoundDiT:地理上下文声景到景观生成
链接:https://arxiv.org/abs/2505.12734
作者: Junbo Wang,  Haofeng Tan,  Bowen Liao,  Albert Jiang,  Teng Fei,  Qixing Huang,  Zhengzhong Tu,  Shan Ye,  Yuhao Kang 
备注:14 pages, 5 figures
摘要:我们提出了一个新的和具有实际意义的问题-地理背景声景景观(GeoS 2L)的生成,其目的是合成地理现实的景观图像从环境音景。现有的音频到图像生成方法通常依赖于通用数据集并忽略地理和环境背景,从而导致与真实世界环境设置不一致的不真实图像。为了解决这个问题,我们引入了一个新的地理上下文计算框架,明确地将地理知识集成到多模态生成建模。我们构建了两个大规模的地理环境多模态数据集,SoundingSVI和SonicUrban,将不同的声音景观与真实世界的景观图像配对。我们提出了SoundDiT,一种新的扩散Transformer(DiT)为基础的模型,采用地理背景的场景条件合成地理上连贯的景观图像。此外,我们提出了一个实际知情的地理环境评估框架,地方相似性分数(PSS),跨元素,场景和人类感知水平,以衡量输入音景和生成的景观图像之间的一致性。大量的实验表明,SoundDiT在视觉保真度和地理环境方面都优于现有的基线。我们的工作不仅为GeoS 2L生成建立了基础基准,还强调了在推进多模态生成模型中融入地理领域知识的重要性,在生成人工智能,地理,城市规划和环境科学的交叉点上开辟了新的方向。
摘要:We present a novel and practically significant problem-Geo-Contextual Soundscape-to-Landscape (GeoS2L) generation-which aims to synthesize geographically realistic landscape images from environmental soundscapes. Prior audio-to-image generation methods typically rely on general-purpose datasets and overlook geographic and environmental contexts, resulting in unrealistic images that are misaligned with real-world environmental settings. To address this limitation, we introduce a novel geo-contextual computational framework that explicitly integrates geographic knowledge into multimodal generative modeling. We construct two large-scale geo-contextual multimodal datasets, SoundingSVI and SonicUrban, pairing diverse soundscapes with real-world landscape images. We propose SounDiT, a novel Diffusion Transformer (DiT)-based model that incorporates geo-contextual scene conditioning to synthesize geographically coherent landscape images. Furthermore, we propose a practically-informed geo-contextual evaluation framework, the Place Similarity Score (PSS), across element-, scene-, and human perception-levels to measure consistency between input soundscapes and generated landscape images. Extensive experiments demonstrate that SounDiT outperforms existing baselines in both visual fidelity and geographic settings. Our work not only establishes foundational benchmarks for GeoS2L generation but also highlights the importance of incorporating geographic domain knowledge in advancing multimodal generative models, opening new directions at the intersection of generative AI, geography, urban planning, and environmental sciences.

【30】 RoVo: Robust Voice Protection Against Unauthorized Speech Synthesis with  Embedding-Level Perturbations
标题: RoVo:针对未经授权的语音合成的强大语音保护,具有嵌入级扰动
链接:https://arxiv.org/abs/2505.12686
作者: Seungmin Kim,  Sohee Park,  Donghyun Kim,  Jisu Lee,  Daeseon Choi 
摘要:随着Deep Voice等基于人工智能的语音合成技术的进步,语音欺骗攻击的风险越来越大,包括未经授权使用他人语音的语音钓鱼和假新闻。现有的直接将对抗性扰动注入音频信号的防御措施效果有限,因为这些扰动可以很容易地通过语音增强方法来中和。为了克服这一限制,我们提出了RoVo(鲁棒语音),这是一种新型的主动防御技术,它将对抗性扰动注入音频信号的高维嵌入向量,将其重建为受保护的语音。这种方法有效地抵御语音合成攻击,也提供了强大的抵抗语音增强模型,这代表了二次攻击的威胁。   在广泛的实验中,RoVo在四种最先进的语音合成模型中,与无保护语音相比,防御成功率(DSR)提高了70%以上。具体来说,RoVo在商业说话人验证API上实现了99.5%的DSR,有效地中和了语音合成攻击。此外,RoVo的扰动即使在强语音增强条件下也保持鲁棒性,优于传统方法。一项用户研究证实,RoVo保留了受保护语音的自然性和可用性,突出了其在复杂和不断变化的威胁场景中的有效性。
摘要:With the advancement of AI-based speech synthesis technologies such as Deep Voice, there is an increasing risk of voice spoofing attacks, including voice phishing and fake news, through unauthorized use of others' voices. Existing defenses that inject adversarial perturbations directly into audio signals have limited effectiveness, as these perturbations can easily be neutralized by speech enhancement methods. To overcome this limitation, we propose RoVo (Robust Voice), a novel proactive defense technique that injects adversarial perturbations into high-dimensional embedding vectors of audio signals, reconstructing them into protected speech. This approach effectively defends against speech synthesis attacks and also provides strong resistance to speech enhancement models, which represent a secondary attack threat.   In extensive experiments, RoVo increased the Defense Success Rate (DSR) by over 70% compared to unprotected speech, across four state-of-the-art speech synthesis models. Specifically, RoVo achieved a DSR of 99.5% on a commercial speaker-verification API, effectively neutralizing speech synthesis attack. Moreover, RoVo's perturbations remained robust even under strong speech enhancement conditions, outperforming traditional methods. A user study confirmed that RoVo preserves both naturalness and usability of protected speech, highlighting its effectiveness in complex and evolving threat scenarios.

【31】 Text2midi-InferAlign: Improving Symbolic Music Generation with  Inference-Time Alignment
标题: 文本2 midi-InferAlign:通过推理时间对齐改进符号音乐生成
链接:https://arxiv.org/abs/2505.12669
作者: Abhinaba Roy,  Geeta Puri,  Dorien Herremans 
备注:7 pages, 1 figure, 5 tables
摘要:我们提出了Text 2 midi-InferAlign,一种新的技术,用于提高在推理时的符号音乐生成。我们的方法在推理过程中利用文本到音频对齐和音乐结构对齐奖励,以鼓励生成的音乐与输入标题一致。具体来说,我们引入了两个目标分数:一个文本音频一致性分数,测量生成的音乐和原始文本标题之间的节奏对齐,和一个谐波一致性分数,惩罚生成的音乐包含不一致的音符的关键。通过在生成过程中优化这些基于字幕的目标,我们的模型产生了与输入字幕更紧密联系的符号音乐,从而提高了生成的作品的整体质量和连贯性。我们的方法可以扩展任何现有的自回归模型,而不需要进一步的训练或微调。我们评估我们的工作上的Text 2 midi-现有的文本到midi生成模型,表现出显着的改善,在客观和主观的评价指标。
摘要:We present Text2midi-InferAlign, a novel technique for improving symbolic music generation at inference time. Our method leverages text-to-audio alignment and music structural alignment rewards during inference to encourage the generated music to be consistent with the input caption. Specifically, we introduce two objectives scores: a text-audio consistency score that measures rhythmic alignment between the generated music and the original text caption, and a harmonic consistency score that penalizes generated music containing notes inconsistent with the key. By optimizing these alignment-based objectives during the generation process, our model produces symbolic music that is more closely tied to the input captions, thereby improving the overall quality and coherence of the generated compositions. Our approach can extend any existing autoregressive model without requiring further training or fine-tuning. We evaluate our work on top of Text2midi - an existing text-to-midi generation model, demonstrating significant improvements in both objective and subjective evaluation metrics.

【32】 Chain-Talker: Chain Understanding and Rendering for Empathetic  Conversational Speech Synthesis
标题: Chain-Talker:同理心对话语音合成的链理解和渲染
链接:https://arxiv.org/abs/2505.12597
作者: Yifan Hu,  Rui Liu,  Yi Ren,  Xiang Yin,  Haizhou Li 
备注:16 pages, 5 figures, 5 tables. Accepted by ACL 2025 (Findings)
摘要:会话语音合成(CSS)旨在将合成的语音与用户-代理交互的情感和风格背景对齐,以实现移情。目前的生成CSS模型面临的可解释性的限制,由于不足的情感感知和冗余的离散语音编码。为了解决上述问题,我们提出了Chain-Talker,一个模仿人类认知的三阶段框架:情感理解从对话历史中获得上下文感知的情感描述符;语义理解通过序列化预测生成紧凑的语义代码;移情渲染通过整合两个组件来合成表达性语音。为了支持情感建模,我们开发了CSS-EmCap,这是一个LLM驱动的自动化管道,用于生成精确的会话语音情感字幕。在三个基准数据集上的实验表明,Chain-Talker比现有方法产生更有表现力和同情心的语音,CSS-EmCap有助于可靠的情感建模。代码和演示可在https://github.com/AI-S2-Lab/Chain-Talker上获得。
摘要:Conversational Speech Synthesis (CSS) aims to align synthesized speech with the emotional and stylistic context of user-agent interactions to achieve empathy. Current generative CSS models face interpretability limitations due to insufficient emotional perception and redundant discrete speech coding. To address the above issues, we present Chain-Talker, a three-stage framework mimicking human cognition: Emotion Understanding derives context-aware emotion descriptors from dialogue history; Semantic Understanding generates compact semantic codes via serialized prediction; and Empathetic Rendering synthesizes expressive speech by integrating both components. To support emotion modeling, we develop CSS-EmCap, an LLM-driven automated pipeline for generating precise conversational speech emotion captions. Experiments on three benchmark datasets demonstrate that Chain-Talker produces more expressive and empathetic speech than existing methods, with CSS-EmCap contributing to reliable emotion modeling. The code and demos are available at: https://github.com/AI-S2-Lab/Chain-Talker.

【33】 VoiceCloak: A Multi-Dimensional Defense Framework against Unauthorized  Diffusion-based Voice Cloning
标题: VoiceCloak:针对未经授权的基于扩散的语音克隆的多维防御框架
链接:https://arxiv.org/abs/2505.12332
作者: Qianyue Hu,  Junyan Wu,  Wei Lu,  Xiangyang Luo 
摘要:扩散模型(DM)在真实语音克隆(VC)方面取得了显着的成功,但也增加了恶意滥用的风险。现有的主动防御设计为传统的VC模型的目的是破坏伪造过程,但他们已被证明与DM由于复杂的生成机制的扩散不兼容。为了弥合这一差距,我们引入了VoiceCloak,这是一个多维的主动防御框架,其目标是在潜在的未经授权的VC中混淆扬声器身份并降低感知质量。为了实现这些目标,我们进行了重点分析,以确定DM中的特定漏洞,允许VoiceCloak通过将对抗性扰动引入参考音频来破坏克隆过程。具体来说,为了混淆说话人身份,VoiceCloak首先通过扭曲表示学习嵌入来最大化身份变化来定位说话人身份,这是由听觉感知原理指导的。此外,VoiceCloak破坏了关键的条件指导过程,特别是注意力上下文,从而阻止了对实现令人信服的克隆至关重要的声音特征的对齐。然后,为了解决第二个目标,VoiceCloak引入了分数幅度放大,以主动引导反向轨迹远离高质量语音的生成。噪声引导的语义损坏被进一步用于破坏DM捕获的结构化语音语义,从而降低输出质量。大量的实验突出了VoiceCloak对未经授权的基于扩散的语音克隆的出色防御成功率。VoiceCloak的音频样本可在https://voice-cloak.github.io/VoiceCloak/上获得。
摘要:Diffusion Models (DMs) have achieved remarkable success in realistic voice cloning (VC), while they also increase the risk of malicious misuse. Existing proactive defenses designed for traditional VC models aim to disrupt the forgery process, but they have been proven incompatible with DMs due to the intricate generative mechanisms of diffusion. To bridge this gap, we introduce VoiceCloak, a multi-dimensional proactive defense framework with the goal of obfuscating speaker identity and degrading perceptual quality in potential unauthorized VC. To achieve these goals, we conduct a focused analysis to identify specific vulnerabilities within DMs, allowing VoiceCloak to disrupt the cloning process by introducing adversarial perturbations into the reference audio. Specifically, to obfuscate speaker identity, VoiceCloak first targets speaker identity by distorting representation learning embeddings to maximize identity variation, which is guided by auditory perception principles. Additionally, VoiceCloak disrupts crucial conditional guidance processes, particularly attention context, thereby preventing the alignment of vocal characteristics that are essential for achieving convincing cloning. Then, to address the second objective, VoiceCloak introduces score magnitude amplification to actively steer the reverse trajectory away from the generation of high-quality speech. Noise-guided semantic corruption is further employed to disrupt structural speech semantics captured by DMs, degrading output quality. Extensive experiments highlight VoiceCloak's outstanding defense success rate against unauthorized diffusion-based voice cloning. Audio samples of VoiceCloak are available at https://voice-cloak.github.io/VoiceCloak/.

【34】 BenSParX: A Robust Explainable Machine Learning Framework for  Parkinson's Disease Detection from Bengali Conversational Speech
标题: BenSParX:一个稳健的可解释机器学习框架,用于从孟加拉语对话语音中检测帕金森病
链接:https://arxiv.org/abs/2505.12192
作者: Riad Hossain,  Muhammad Ashad Kabir,  Arat Ibne Golam Mowla,  Animesh Chandra Roy,  Ranjit Kumar Ghosh 
备注:46 pages, 16 figures
摘要:帕金森病(PD)对全球健康构成了日益严重的挑战,孟加拉国与PD相关的死亡率显着上升。在资源有限的环境中,PD的早期检测仍然特别具有挑战性,其中基于语音的分析已成为一种有前途的非侵入性和成本效益的替代方案。然而,现有的研究主要集中在英语或其他主要语言上;值得注意的是,孟加拉语没有PD语音数据集,这对文化包容性和可访问的医疗保健解决方案构成了重大障碍。此外,大多数先前的研究只采用了一组狭窄的声学特征,有限或没有超参数调整和特征选择策略,很少注意模型的可解释性。这限制了一个强大的和可推广的机器学习模型的发展。为了解决这一差距,我们提出了BenSparX,这是第一个用于PD检测的孟加拉语会话语音数据集,以及为早期诊断量身定制的强大且可解释的机器学习框架。所提出的框架结合了不同的声学特征类别,系统的特征选择方法,以及最先进的机器学习算法和广泛的超参数优化。此外,为了增强模型预测的可解释性和可信度,该框架采用SHAP(SHapley Additive exPlanations)分析来量化单个声学特征对PD检测的贡献。我们的框架实现了最先进的性能,准确率为95.77%,F1评分为95.57%,AUC-ROC为0.982。我们通过将框架应用于其他语言的现有PD数据集,进一步从外部验证了我们的方法,在这些语言中,它始终优于最先进的方法。为了促进进一步的研究和再现性,数据集已在https://github.com/Riad071/BenSParX上公开。
摘要:Parkinson's disease (PD) poses a growing global health challenge, with Bangladesh experiencing a notable rise in PD-related mortality. Early detection of PD remains particularly challenging in resource-constrained settings, where voice-based analysis has emerged as a promising non-invasive and cost-effective alternative. However, existing studies predominantly focus on English or other major languages; notably, no voice dataset for PD exists for Bengali - posing a significant barrier to culturally inclusive and accessible healthcare solutions. Moreover, most prior studies employed only a narrow set of acoustic features, with limited or no hyperparameter tuning and feature selection strategies, and little attention to model explainability. This restricts the development of a robust and generalizable machine learning model. To address this gap, we present BenSparX, the first Bengali conversational speech dataset for PD detection, along with a robust and explainable machine learning framework tailored for early diagnosis. The proposed framework incorporates diverse acoustic feature categories, systematic feature selection methods, and state-of-the-art machine learning algorithms with extensive hyperparameter optimization. Furthermore, to enhance interpretability and trust in model predictions, the framework incorporates SHAP (SHapley Additive exPlanations) analysis to quantify the contribution of individual acoustic features toward PD detection. Our framework achieves state-of-the-art performance, yielding an accuracy of 95.77%, F1 score of 95.57%, and AUC-ROC of 0.982. We further externally validated our approach by applying the framework to existing PD datasets in other languages, where it consistently outperforms state-of-the-art approaches. To facilitate further research and reproducibility, the dataset has been made publicly available at https://github.com/Riad071/BenSParX.

【35】 Learning to Highlight Audio by Watching Movies
标题: 学习通过观看电影来突出音频
链接:https://arxiv.org/abs/2505.12154
作者: Chao Huang,  Ruohan Gao,  J. M. F. Tsang,  Jan Kurcius,  Cagdas Bilen,  Chenliang Xu,  Anurag Kumar,  Sanjeel Parekh 
备注:CVPR 2025. Project page: this https URL
摘要:近年来,视频内容的创建和消费显著增加。制作引人入胜的内容需要仔细策划视觉和音频元素。虽然视觉提示策展,通过最佳视点选择或后期编辑等技术,一直是媒体制作的核心,但其自然对应物音频却没有经历过同等的进步。这通常会导致视觉和听觉显着性之间的脱节。为了弥合这一差距,我们引入了一项新的任务:视觉引导的声学高亮,其目的是转换音频,以提供由随附视频引导的适当高亮效果,最终创建更和谐的视听体验。我们提出了一个灵活的,基于transformer的多模态框架来解决这个任务。为了训练我们的模型,我们还引入了一个新的数据集--muddy mix数据集,它利用了电影中细致的音频和视频制作,提供了一种免费的监督形式。我们开发了一个伪数据生成过程来模拟混合不好的音频,通过三步过程-分离,调整和混音来模仿现实世界的场景。我们的方法在定量和主观评估方面始终优于几个基线。我们还系统地研究了不同类型的上下文指导和数据集的难度水平的影响。我们的项目页面在这里:https://wikichao.github.io/VisAH/。
摘要:Recent years have seen a significant increase in video content creation and consumption. Crafting engaging content requires the careful curation of both visual and audio elements. While visual cue curation, through techniques like optimal viewpoint selection or post-editing, has been central to media production, its natural counterpart, audio, has not undergone equivalent advancements. This often results in a disconnect between visual and acoustic saliency. To bridge this gap, we introduce a novel task: visually-guided acoustic highlighting, which aims to transform audio to deliver appropriate highlighting effects guided by the accompanying video, ultimately creating a more harmonious audio-visual experience. We propose a flexible, transformer-based multimodal framework to solve this task. To train our model, we also introduce a new dataset -- the muddy mix dataset, leveraging the meticulous audio and video crafting found in movies, which provides a form of free supervision. We develop a pseudo-data generation process to simulate poorly mixed audio, mimicking real-world scenarios through a three-step process -- separation, adjustment, and remixing. Our approach consistently outperforms several baselines in both quantitative and subjective evaluation. We also systematically study the impact of different types of contextual guidance and difficulty levels of the dataset. Our project page is here: https://wikichao.github.io/VisAH/.

【36】 SepPrune: Structured Pruning for Efficient Deep Speech Separation
标题: SepPrune:用于高效深度语音分离的结构化修剪
链接:https://arxiv.org/abs/2505.12079
作者: Yuqi Li,  Kai Li,  Xin Yin,  Zhifei Yang,  Junhao Dong,  Zeyu Dong,  Chuanguang Yang,  Yingli Tian,  Yao Lu 
摘要:尽管近年来深度学习大大推进了语音分离,但大多数现有研究仍然优先考虑分离质量,而忽略了计算效率,这是实时应用中低延迟语音处理的一个重要因素。在本文中,我们提出了SepPrune,这是第一个专门设计用于压缩深度语音分离模型并降低其计算成本的结构化修剪框架。SepPrune首先分析给定模型的计算结构,以确定计算负担最高的层。然后,它引入了一个微分掩蔽策略,使梯度驱动的通道选择。基于学习的掩码,SepPrune修剪冗余通道并微调剩余参数以恢复性能。大量的实验表明,这种可学习的修剪范式产生实质性的优势,在语音分离模型的通道修剪,优于现有的方法。值得注意的是,使用SepPrune修剪的模型可以恢复85%的预训练模型(经过数百个epoch训练)的性能,只需一个epoch的微调,并且比从头开始训练快36倍。代码可在https://github.com/itsnotacie/SepPrune上获得。
摘要:Although deep learning has substantially advanced speech separation in recent years, most existing studies continue to prioritize separation quality while overlooking computational efficiency, an essential factor for low-latency speech processing in real-time applications. In this paper, we propose SepPrune, the first structured pruning framework specifically designed to compress deep speech separation models and reduce their computational cost. SepPrune begins by analyzing the computational structure of a given model to identify layers with the highest computational burden. It then introduces a differentiable masking strategy to enable gradient-driven channel selection. Based on the learned masks, SepPrune prunes redundant channels and fine-tunes the remaining parameters to recover performance. Extensive experiments demonstrate that this learnable pruning paradigm yields substantial advantages for channel pruning in speech separation models, outperforming existing methods. Notably, a model pruned with SepPrune can recover 85% of the performance of a pre-trained model (trained over hundreds of epochs) with only one epoch of fine-tuning, and achieves convergence 36$\times$ faster than training from scratch. Code is available at https://github.com/itsnotacie/SepPrune.

【37】 Automatic Speech Recognition for African Low-Resource Languages:  Challenges and Future Directions
标题: 非洲低资源语言的自动语音识别:挑战和未来方向
链接:https://arxiv.org/abs/2505.11690
作者: Sukairaj Hafiz Imam,  Babangida Sani,  Dawit Ketema Gete,  Bedru Yimam Ahamed,  Ibrahim Said Ahmad,  Idris Abdulmumin,  Seid Muhie Yimam,  Muhammad Yahuza Bello,  Shamsuddeen Hassan Muhammad 
摘要:自动语音识别(ASR)技术已经改变了人机交互;然而,非洲的低资源语言在研究和实际应用中仍然严重不足。这项研究调查了阻碍这些语言的ASR系统开发的主要挑战,包括数据稀缺,语言复杂性,有限的计算资源,声学可变性以及围绕偏见和隐私的道德问题。主要目标是批判性地分析这些障碍,并确定实用的,包容性的战略,以推进非洲范围内的ASR技术。最近的进展和案例研究强调了有前途的策略,如社区驱动的数据收集,自我监督和多语言学习,轻量级模型架构和优先考虑隐私的技术。来自涉及各种非洲语言的试点项目的证据展示了定制解决方案的可行性和影响,其中包括基于词素的建模和医疗保健和教育等领域的特定领域ASR应用。研究结果强调了跨学科合作和持续投资的重要性,以应对非洲大陆面临的独特语言和基础设施挑战。这项研究为创建道德,高效和包容性的ASR系统提供了一个渐进的路线图,该系统不仅保护语言多样性,还改善了数字可访问性,并促进了非洲语言使用者的社会经济参与。
摘要:Automatic Speech Recognition (ASR) technologies have transformed human-computer interaction; however, low-resource languages in Africa remain significantly underrepresented in both research and practical applications. This study investigates the major challenges hindering the development of ASR systems for these languages, which include data scarcity, linguistic complexity, limited computational resources, acoustic variability, and ethical concerns surrounding bias and privacy. The primary goal is to critically analyze these barriers and identify practical, inclusive strategies to advance ASR technologies within the African context. Recent advances and case studies emphasize promising strategies such as community-driven data collection, self-supervised and multilingual learning, lightweight model architectures, and techniques that prioritize privacy. Evidence from pilot projects involving various African languages showcases the feasibility and impact of customized solutions, which encompass morpheme-based modeling and domain-specific ASR applications in sectors like healthcare and education. The findings highlight the importance of interdisciplinary collaboration and sustained investment to tackle the distinct linguistic and infrastructural challenges faced by the continent. This study offers a progressive roadmap for creating ethical, efficient, and inclusive ASR systems that not only safeguard linguistic diversity but also improve digital accessibility and promote socioeconomic participation for speakers of African languages.

【38】 ASR-FAIRBENCH: Measuring and Benchmarking Equity Across Speech  Recognition Systems
标题: ASR-FAIRBENCH:测量和基准测量语音识别系统的公平性
链接:https://arxiv.org/abs/2505.11572
作者: Anand Rai,  Satyam Rahangdale,  Utkarsh Anand,  Animesh Mukherjee 
备注:Paper accepted at INTERSPEECH 2025
摘要:自动语音识别(ASR)系统在日常应用中已经无处不在,但在不同的人口统计群体之间的性能仍然存在显着差异。在这项工作中,我们介绍了ASR-FAIRBENCH排行榜,该排行榜旨在实时评估ASR模型的准确性和公平性。利用Meta的公平言论数据集,它捕捉了不同的人口统计特征,我们采用了混合效应泊松回归模型来获得整体公平得分。该分数与单词错误率(WER)等传统指标相结合,以计算公平调整的ASR分数(FAAS),提供全面的评估框架。我们的方法揭示了SOTA ASR模型在不同人口群体中的显著性能差异,并提供了一个基准,以推动更具包容性的ASR技术的发展。
摘要:Automatic Speech Recognition (ASR) systems have become ubiquitous in everyday applications, yet significant disparities in performance across diverse demographic groups persist. In this work, we introduce the ASR-FAIRBENCH leaderboard which is designed to assess both the accuracy and equity of ASR models in real-time. Leveraging the Meta's Fair-Speech dataset, which captures diverse demographic characteristics, we employ a mixed-effects Poisson regression model to derive an overall fairness score. This score is integrated with traditional metrics like Word Error Rate (WER) to compute the Fairness Adjusted ASR Score (FAAS), providing a comprehensive evaluation framework. Our approach reveals significant performance disparities in SOTA ASR models across demographic groups and offers a benchmark to drive the development of more inclusive ASR technologies.


机器翻译由腾讯交互翻译提供,仅供参考