今日论文合集:cs.SD语音5篇,eess.AS音频处理4篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】UniTok-Audio: A Unified Audio Generation Framework via Generative Modeling on Discrete Codec Tokens
标题:UniTok-音频:通过对离散编解码器令牌进行生成建模的统一音频生成框架
链接:https://arxiv.org/abs/2510.26372

作者:Chengwei Liu, Haoyin Yan, Shaofei Xue, Xiaotao Liang, Yinghao Liu, Zheng Xue, Gang Song, Boyang Zhou
备注:21 pages, 3 figures
摘要:生成建模最近在文本、图像和音频领域取得了显著的成功,展示了统一表示学习的强大功能。然而,音频生成模型在音频质量和跨任务的泛化能力方面仍然面临挑战。这种碎片化会导致冗余的开发工作、不一致的性能和有限的可扩展性。为了解决这些问题,我们提出了\textbf{UniTok-Audio},一个可扩展和可扩展的框架,用于统一的音频生成任务。具体而言,1)UniTok-Audio提取条件的连续特征,以自回归方式生成目标音频的离散令牌; 2)特殊的任务标识符令牌将多个任务的不同学习模式统一在单个框架中; 3)开发了涉及声学和语义分支的双流音频编解码器,用于高保真波形重建。实验结果表明,UniTok-Audio在语音恢复、目标说话人提取、语音分离、语音转换和语言查询音频源分离这五个时间对齐任务上,与最先进的特定任务或多任务系统相比,具有竞争力的性能。为了促进未来的研究,我们将开源我们的代码库。我们工作的演示页面可以在这里找到:https://alibaba.github.io/unified-audio。
摘要:Generative modeling has recently achieved remarkable success across text, image, and audio domains, demonstrating powerful capabilities for unified representation learning. However, audio generation models still face challenges in terms of audio quality and generalization ability across tasks. This fragmentation results in redundant development efforts, inconsistent performance, and limited extensibility. To address these issues, we propose \textbf{UniTok-Audio}, a scalable and extensible framework for unified audio generation tasks. Specifically, 1) UniTok-Audio extracts continuous feature of conditions to generates discrete tokens of target audio in an autoregressive manner; 2) a special task identifier token unifies different learning patterns of multiple tasks in a single framework; 3) a dual-stream audio codec involving acoustic and semantic branch is developed for high-fidelity waveform reconstruction. Experimental results demonstrate that UniTok-Audio achieves competitive performance in comparation with state-of-the-art task-specific or multi-task systems across five time-aligned tasks: speech restoration, target speaker extraction, speech separation, voice conversion, and language-queried audio source separation. To foster future research, we will open-source our codebase. The demo page of our work can be found here: https://alibaba.github.io/unified-audio.


【2】Modeling strategies for speech enhancement in the latent space of a neural audio codec
标题:神经音频编解码器潜在空间中语音增强的建模策略
链接:https://arxiv.org/abs/2510.26299

作者:Sofiene Kammoun, Xavier Alameda-Pineda, Simon Leglaive
摘要:神经音频编解码器(NAC)以连续向量或离散令牌的序列的形式提供紧凑的潜在语音表示。在这项工作中,我们调查如何这两种类型的语音表示比较时,作为训练目标的监督语音增强。我们考虑基于Conformer架构的自回归和非自回归语音增强模型,以及NAC编码器简单地进行语音增强微调的简单基线。我们的实验揭示了三个关键发现:预测连续潜在表示始终优于离散令牌预测;自回归模型实现了更高的质量,但以牺牲可理解性和效率为代价,使非自回归模型在实践中更具吸引力;编码器微调产生了最强的增强指标总体上,尽管以降级的编解码器重建为代价。代码和音频示例可以在线获得。
摘要:Neural audio codecs (NACs) provide compact latent speech representations in the form of sequences of continuous vectors or discrete tokens. In this work, we investigate how these two types of speech representations compare when used as training targets for supervised speech enhancement. We consider both autoregressive and non-autoregressive speech enhancement models based on the Conformer architecture, as well as a simple baseline where the NAC encoder is simply fine-tuned for speech enhancement. Our experiments reveal three key findings: predicting continuous latent representations consistently outperforms discrete token prediction; autoregressive models achieve higher quality but at the expense of intelligibility and efficiency, making non-autoregressive models more attractive in practice; and encoder fine-tuning yields the strongest enhancement metrics overall, though at the cost of degraded codec reconstruction. The code and audio samples are available online.


【3】SP-MCQA: Evaluating Intelligibility of TTS Beyond the Word Level
标题:SP-MCQA:超越文字级别评估TTC的可理解性
链接:https://arxiv.org/abs/2510.26190

作者:Hitomi Jin Ling Tee, Chaoren Wang, Zijie Zhang, Zhizheng Wu
摘要:TTS可懂度的评估已经达到了瓶颈,因为现有的评估严重依赖于逐词的准确性指标,如WER,无法捕捉真实世界语音的复杂性或反映人类的理解需求。为了解决这个问题,我们提出了口语通道多项选择题测试,这是一种新的主观方法,用于评估合成语音中关键信息的准确性,并发布了SP-MCQA-Eval,这是一个用于SP-MCQA评估的8.76小时新闻风格基准数据集。我们的实验表明,低WER并不一定保证高的关键信息的准确性,暴露了传统的指标和实际的可理解性之间的差距。SP-MCQA表明,即使是最先进的(SOTA)模型仍然缺乏强大的文本规范化和语音准确性。这项工作强调了迫切需要高层次的,更逼真的评估标准,现在许多系统已经在WER上表现出色,但可能在现实世界的可理解性方面有所欠缺。
摘要:The evaluation of intelligibility for TTS has reached a bottleneck, as existing assessments heavily rely on word-by-word accuracy metrics such as WER, which fail to capture the complexity of real-world speech or reflect human comprehension needs. To address this, we propose Spoken-Passage Multiple-Choice Question Answering, a novel subjective approach evaluating the accuracy of key information in synthesized speech, and release SP-MCQA-Eval, an 8.76-hour news-style benchmark dataset for SP-MCQA evaluation. Our experiments reveal that low WER does not necessarily guarantee high key-information accuracy, exposing a gap between traditional metrics and practical intelligibility. SP-MCQA shows that even state-of-the-art (SOTA) models still lack robust text normalization and phonetic accuracy. This work underscores the urgent need for high-level, more life-like evaluation criteria now that many systems already excel at WER yet may fall short on real-world intelligibility.


【4】ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models
标题:ALMGGuard:安全捷径以及在哪里找到它们作为音频语言模型的护栏
链接:https://arxiv.org/abs/2510.26096

作者:Weifei Jin, Yuxin Cao, Junjie Su, Minhui Xue, Jie Hao, Ke Xu, Jin Song Dong, Derui Wang
备注:Accepted to NeurIPS 2025
摘要:音频语言模型(ALM)的最新进展显着提高了多模态理解能力。然而,音频模态的引入也带来了新的和独特的脆弱性向量。之前的研究提出了专门针对ALM的越狱攻击,表明直接从传统音频对抗性攻击或基于文本的大型语言模型(LLM)越狱转移的防御对于这些特定于ALM的威胁基本上无效。为了解决这个问题,我们提出了ALMGuard,第一个防御框架量身定制的ALM。基于安全对齐的快捷方式自然存在于ALM中的假设,我们设计了一种方法来识别通用的触发器激活扰动(SAP),该触发器激活安全快捷方式以在推理时保护ALM。为了更好地筛选出有效的触发器,同时保留模型在良性任务上的实用性,我们进一步提出了Mel梯度稀疏掩码(M-GSM),它将扰动限制在对越狱敏感但对语音理解不敏感的Mel频率区间。理论分析和实验结果表明,我们的方法对可见和不可见的攻击的鲁棒性。总的来说,\MethodName在四种模型中将高级特定于ALM的越狱攻击的平均成功率降低到4.6%,同时在良性基准测试中保持相当的实用性,将其确立为最新的技术水平。我们的代码和数据可在https://github.com/WeifeiJin/ALMGuard上获得。
摘要:Recent advances in Audio-Language Models (ALMs) have significantly improved multimodal understanding capabilities. However, the introduction of the audio modality also brings new and unique vulnerability vectors. Previous studies have proposed jailbreak attacks that specifically target ALMs, revealing that defenses directly transferred from traditional audio adversarial attacks or text-based Large Language Model (LLM) jailbreaks are largely ineffective against these ALM-specific threats. To address this issue, we propose ALMGuard, the first defense framework tailored to ALMs. Based on the assumption that safety-aligned shortcuts naturally exist in ALMs, we design a method to identify universal Shortcut Activation Perturbations (SAPs) that serve as triggers that activate the safety shortcuts to safeguard ALMs at inference time. To better sift out effective triggers while preserving the model's utility on benign tasks, we further propose Mel-Gradient Sparse Mask (M-GSM), which restricts perturbations to Mel-frequency bins that are sensitive to jailbreaks but insensitive to speech understanding. Both theoretical analyses and empirical results demonstrate the robustness of our method against both seen and unseen attacks. Overall, \MethodName reduces the average success rate of advanced ALM-specific jailbreak attacks to 4.6% across four models, while maintaining comparable utility on benign benchmarks, establishing it as the new state of the art. Our code and data are available at https://github.com/WeifeiJin/ALMGuard.


【5】Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders
标题:通过稀疏自动编码器学习音频潜在空间中的可解释特征
链接:https://arxiv.org/abs/2510.23802

作者:Nathan Paek, Yongyi Zang, Qihui Yang, Randal Leistikow
备注:Accepted to NeurIPS 2025 Mechanistic Interpretability Workshop
摘要:虽然稀疏自动编码器(SAE)成功地从语言模型中提取了可解释的特征,但将其应用于音频生成面临着独特的挑战:音频的密集性质需要压缩,从而模糊了语义含义,并且自动特征表征仍然有限。我们提出了一个框架来解释音频生成模型映射其潜在的表示人类可解释的声学概念。我们训练音频自动编码器潜伏期的SAE,然后学习从SAE特征到离散声学属性(音高,振幅和音色)的线性映射。这使得人工智能音乐生成过程的可控操作和分析成为可能,揭示了合成过程中声学特性的出现。我们验证了我们的方法上连续(DiffRhythm-VAE)和离散(EnCodec,WavTokenizer)音频潜在空间,并分析DiffRhythm,一个国家的最先进的文本到音乐模型,演示如何音高,音色和响度在整个生成过程中演变。虽然我们的工作只是在音频模态上完成,但我们的框架可以扩展到视觉潜在空间生成模型的可解释性分析。
摘要:While sparse autoencoders (SAEs) successfully extract interpretable features from language models, applying them to audio generation faces unique challenges: audio's dense nature requires compression that obscures semantic meaning, and automatic feature characterization remains limited. We propose a framework for interpreting audio generative models by mapping their latent representations to human-interpretable acoustic concepts. We train SAEs on audio autoencoder latents, then learn linear mappings from SAE features to discretized acoustic properties (pitch, amplitude, and timbre). This enables both controllable manipulation and analysis of the AI music generation process, revealing how acoustic properties emerge during synthesis. We validate our approach on continuous (DiffRhythm-VAE) and discrete (EnCodec, WavTokenizer) audio latent spaces, and analyze DiffRhythm, a state-of-the-art text-to-music model, to demonstrate how pitch, timbre, and loudness evolve throughout generation. While our work is only done on audio modality, our framework can be extended to interpretable analysis of visual latent space generation models.


eess.AS音频处理


【1】SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations
标题:SPSYS:用于学习语音和音频表示的统一SSL框架
链接:https://arxiv.org/abs/2510.25955

作者:Xiaoyu Yang, Yifan Yang, Zengrui Jin, Ziyun Cui, Wen Wu, Baoxiang Li, Chao Zhang, Phil Woodland
摘要:自监督学习(SSL)擅长学习声学信号的通用表示,但流行的方法仍然是特定于领域的,专为语音或通用音频量身定制,阻碍了在这两个领域具有综合能力的统一表示模型的开发。为了解决这个问题,我们提出了SPECTRA(语音和音频表示),第一个SSL框架,成功地从语音和音频数据的混合物中学习统一的语音和音频表示。SPARTERS提出了一个统一的预训练目标,基于语音和一般音频的细粒度离散令牌的掩码预测。这些令牌来自连续语音和音频表示使用多码本矢量量化(MVQ)方法,保留丰富的声学细节建模语音和复杂的音频事件必不可少的。SPRINT用于预训练单域和统一的语音和音频SSL模型。我们的语音域模型在SUPERB基准测试(SSL模型的语音处理基准测试)上建立了一个新的最先进的水平,在具有相同预训练语料库和类似模型大小的15个任务中,有12个匹配或超过了极具竞争力的WavLM Large。至关重要的是,我们的统一模型可以学习互补功能,并在SUPERB和HEAR两个主要基准测试中展示全面的功能,以评估音频表示。通过进一步扩大模型大小和预训练数据,我们提出了一个具有600 M参数的统一模型,在这两个领域都表现出色,使其成为听觉理解中最强大和最通用的开源SSL模型之一。推理代码和预训练模型将公开提供。
摘要:Self-Supervised Learning (SSL) excels at learning generic representations of acoustic signals, yet prevailing methods remain domain-specific, tailored to either speech or general audio, hindering the development of a unified representation model with a comprehensive capability over both domains. To address this, we present SPEAR (SPEech and Audio Representations), the first SSL framework to successfully learn unified speech and audio representations from a mixture of speech and audio data. SPEAR proposes a unified pre-training objective based on masked prediction of fine-grained discrete tokens for both speech and general audio. These tokens are derived from continuous speech and audio representations using a Multi-codebook Vector Quantisation (MVQ) method, retaining rich acoustic detail essential for modelling both speech and complex audio events. SPEAR is applied to pre-train both single-domain and unified speech-and-audio SSL models. Our speech-domain model establishes a new state-of-the-art on the SUPERB benchmark, a speech processing benchmark for SSL models, matching or surpassing the highly competitive WavLM Large on 12 out of 15 tasks with the same pre-training corpora and a similar model size. Crucially, our unified model learns complementary features and demonstrates comprehensive capabilities across two major benchmarks, SUPERB and HEAR, for evaluating audio representations. By further scaling up the model size and pre-training data, we present a unified model with 600M parameters that excels in both domains, establishing it as one of the most powerful and versatile open-source SSL models for auditory understanding. The inference code and pre-trained models will be made publicly available.


【2】Modeling strategies for speech enhancement in the latent space of a neural audio codec
标题:神经音频编解码器潜在空间中语音增强的建模策略
链接:https://arxiv.org/abs/2510.26299

作者:Sofiene Kammoun, Xavier Alameda-Pineda, Simon Leglaive
摘要:神经音频编解码器(NAC)以连续向量或离散令牌的序列的形式提供紧凑的潜在语音表示。在这项工作中,我们调查如何这两种类型的语音表示比较时,作为训练目标的监督语音增强。我们考虑基于Conformer架构的自回归和非自回归语音增强模型,以及NAC编码器简单地进行语音增强微调的简单基线。我们的实验揭示了三个关键发现:预测连续潜在表示始终优于离散令牌预测;自回归模型实现了更高的质量,但以牺牲可理解性和效率为代价,使非自回归模型在实践中更具吸引力;编码器微调产生了最强的增强指标,尽管以降级的编解码器重建为代价。代码和音频示例可以在线获得。
摘要:Neural audio codecs (NACs) provide compact latent speech representations in the form of sequences of continuous vectors or discrete tokens. In this work, we investigate how these two types of speech representations compare when used as training targets for supervised speech enhancement. We consider both autoregressive and non-autoregressive speech enhancement models based on the Conformer architecture, as well as a simple baseline where the NAC encoder is simply fine-tuned for speech enhancement. Our experiments reveal three key findings: predicting continuous latent representations consistently outperforms discrete token prediction; autoregressive models achieve higher quality but at the expense of intelligibility and efficiency, making non-autoregressive models more attractive in practice; and encoder fine-tuning yields the strongest enhancement metrics overall, though at the cost of degraded codec reconstruction. The code and audio samples are available online.


【3】SP-MCQA: Evaluating Intelligibility of TTS Beyond the Word Level
标题:SP-MCQA:超越文字级别评估TTC的可理解性
链接:https://arxiv.org/abs/2510.26190

作者:Hitomi Jin Ling Tee, Chaoren Wang, Zijie Zhang, Zhizheng Wu
摘要:TTS可懂度的评估已经达到了瓶颈,因为现有的评估严重依赖于逐词的准确性指标,如WER,无法捕捉真实世界语音的复杂性或反映人类的理解需求。为了解决这个问题,我们提出了口语通道多项选择题测试,这是一种新的主观方法,用于评估合成语音中关键信息的准确性,并发布了SP-MCQA-Eval,这是一个用于SP-MCQA评估的8.76小时新闻风格基准数据集。我们的实验表明,低WER并不一定保证高的关键信息的准确性,暴露了传统的指标和实际的可理解性之间的差距。SP-MCQA表明,即使是最先进的(SOTA)模型仍然缺乏强大的文本规范化和语音准确性。这项工作强调了迫切需要高层次的,更逼真的评估标准,现在许多系统已经在WER上表现出色,但可能在现实世界的可理解性方面有所欠缺。
摘要:The evaluation of intelligibility for TTS has reached a bottleneck, as existing assessments heavily rely on word-by-word accuracy metrics such as WER, which fail to capture the complexity of real-world speech or reflect human comprehension needs. To address this, we propose Spoken-Passage Multiple-Choice Question Answering, a novel subjective approach evaluating the accuracy of key information in synthesized speech, and release SP-MCQA-Eval, an 8.76-hour news-style benchmark dataset for SP-MCQA evaluation. Our experiments reveal that low WER does not necessarily guarantee high key-information accuracy, exposing a gap between traditional metrics and practical intelligibility. SP-MCQA shows that even state-of-the-art (SOTA) models still lack robust text normalization and phonetic accuracy. This work underscores the urgent need for high-level, more life-like evaluation criteria now that many systems already excel at WER yet may fall short on real-world intelligibility.


【4】POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
标题:POWSM:一个语音开放耳语风格语音基础模型
链接:https://arxiv.org/abs/2510.24992

作者:Chin-Jou Li, Kalvin Chang, Shikhar Bharadwaj, Eunjung Yeo, Kwanghee Choi, Jian Zhu, David Mortensen, Shinji Watanabe
备注:14 pages, under review
摘要:口语处理的最新进展已经导致语音任务的实质性进展,例如自动语音识别(ASR)、音素识别(PR)、字素到音素转换(G2P)和音素到字素转换(P2G)。尽管它们在概念上相似,但这些任务在很大程度上是孤立地研究的,每个任务都依赖于特定于任务的架构和数据集。在本文中,我们介绍了POWSM(语音开放耳语式语音模型),第一个统一的框架,能够共同执行多个电话相关的任务。POWSM支持音频、文本(字素)和音素之间的无缝转换,为通用和低资源语音处理开辟了新的可能性。我们的模型优于或匹配类似规模的专用PR模型(Wav2Vec2Phoneme和ZIPA),同时共同支持G2P,P2G和ASR。我们发布的训练数据、代码和模型旨在促进开放科学。
摘要:Recent advances in spoken language processing have led to substantial progress in phonetic tasks such as automatic speech recognition (ASR), phone recognition (PR), grapheme-to-phoneme conversion (G2P), and phoneme-to-grapheme conversion (P2G). Despite their conceptual similarity, these tasks have largely been studied in isolation, each relying on task-specific architectures and datasets. In this paper, we introduce POWSM (Phonetic Open Whisper-style Speech Model), the first unified framework capable of jointly performing multiple phone-related tasks. POWSM enables seamless conversion between audio, text (graphemes), and phones, opening up new possibilities for universal and low-resource speech processing. Our model outperforms or matches specialized PR models of similar size (Wav2Vec2Phoneme and ZIPA) while jointly supporting G2P, P2G, and ASR. Our training data, code and models are released to foster open science.


机器翻译由腾讯交互翻译提供,仅供参考