今日论文合集:cs.SD语音9篇,eess.AS音频处理6篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】A Comparative Analysis on ASR System Combination for Attention, CTC, Factored Hybrid, and Transducer Models
标题:注意力、CSC、因子混合和传感器模型的ASB系统组合比较分析
链接:https://arxiv.org/abs/2508.09880

作者: Noureldin Bayoumi, Robin Schmitt, Tina Raissi, Albert Zeyer, Ralf Schlüter, Hermann Ney
备注:Accepted for presentation at IEEE Speech Communication; 16th ITG Conference
摘要:语音识别(ASR)系统的组合方法涵盖了结构化的语音级或基于单词的合并技术,以及在波束搜索期间组合模型分数。在这项工作中,我们比较流行的ASR架构的模型组合。我们的方法利用不同模型的互补优势,探索搜索空间的不同部分。我们对两个候选模型的联合假设列表进行重新评分。然后,我们通过这些序列水平分数的对数线性组合来确定最佳假设。虽然第一遍识别期间的模型组合可能会提高性能,但由于解码方法不同,它会引入可变性,使直接比较更具挑战性。我们的两次通过方法确保了本研究中所有系统组合结果的一致性比较。我们评估模型对候选人不同的架构和标签拓扑结构和单位。实验结果提供了Librispeech 960h任务。
摘要:Combination approaches for speech recognition (ASR) systems cover structured sentence-level or word-based merging techniques as well as combination of model scores during beam search. In this work, we compare model combination across popular ASR architectures. Our method leverages the complementary strengths of different models in exploring diverse portions of the search space. We rescore a joint hypothesis list of two model candidates. We then identify the best hypothesis through log-linear combination of these sequence-level scores. While model combination during first-pass recognition may yield improved performance, it introduces variability due to differing decoding methods, making direct comparison more challenging. Our two-pass method ensures consistent comparisons across all system combination results presented in this study. We evaluate model pair candidates with varying architectures and label topologies and units. Experimental results are provided for the Librispeech 960h task.


【2】Analysis of Domain Shift across ASR Architectures via TTS-Enabled Separation of Target Domain and Acoustic Conditions
标题:通过TTS实现的目标域和声学条件分离分析ASR架构中的域偏移
链接:https://arxiv.org/abs/2508.09868

作者:Tina Raissi, Nick Rossenbach, Ralf Schlüter
备注:Accepted for presentation at IEEE ASRU 2025
摘要:我们分析了自动语音识别(ASR)的建模选择域不匹配,比较经典的模块化和新的序列到序列(seq 2seq)架构。在不同的ASR架构中,我们研究了一系列建模选择,包括标签单元、上下文长度和拓扑结构。为了从声学变化中分离语言域效应,我们使用在LibriSpeech上训练的文本到语音系统合成目标域音频。我们将目标域n-gram和神经语言模型用于域适应,而无需重新训练声学模型。据我们所知,这是第一次在域转移下跨最先进架构对优化的ASR系统进行受控比较,从而深入了解其泛化情况。结果表明,在域移位下,影响性能的不是解码器架构的选择或经典模块化和新颖seq 2seq模型之间的区别,而是特定的建模选择。
摘要:We analyze automatic speech recognition (ASR) modeling choices under domain mismatch, comparing classic modular and novel sequence-to-sequence (seq2seq) architectures. Across the different ASR architectures, we examine a spectrum of modeling choices, including label units, context length, and topology. To isolate language domain effects from acoustic variation, we synthesize target domain audio using a text-to-speech system trained on LibriSpeech. We incorporate target domain n-gram and neural language models for domain adaptation without retraining the acoustic model. To our knowledge, this is the first controlled comparison of optimized ASR systems across state-of-the-art architectures under domain shift, offering insights into their generalization. The results show that, under domain shift, rather than the decoder architecture choice or the distinction between classic modular and novel seq2seq models, it is specific modeling choices that influence performance.


【3】BeatFM: Improving Beat Tracking with Pre-trained Music Foundation Model
标题:BeatFM:利用预先训练的音乐基金会模型改进节拍跟踪
链接:https://arxiv.org/abs/2508.09790

作者:Ganghui Ru, Jieying Wang, Jiahao Zhao, Yulun Wu, Yi Yu, Nannan Jiang, Wei Wang, Wei Li
备注:This paper has been accepted by ICME2025
摘要:节拍跟踪是音乐信息检索中的一个研究热点。然而,当前的节拍跟踪方法面临的挑战,由于标记数据的稀缺,这限制了他们的能力,以概括不同的音乐风格和准确地捕捉复杂的节奏结构。为了克服这些挑战,我们提出了一种新的节拍跟踪范例BeatFM,它引入了一个预先训练的音乐基础模型,并利用其丰富的语义知识来提高节拍跟踪性能。在不同音乐数据集上进行预训练,使音乐基础模型对音乐有着强大的理解,从而有效地应对这些挑战。为了进一步适应节拍跟踪,我们设计了一个即插即用的多维语义聚合模块,它由三个并行的子模块组成,每个子模块分别专注于时间,频率和信道域的语义聚合。大量的实验表明,我们的方法在多个基准数据集上的节拍和下拍跟踪中达到了最先进的性能。
摘要:Beat tracking is a widely researched topic in music information retrieval. However, current beat tracking methods face challenges due to the scarcity of labeled data, which limits their ability to generalize across diverse musical styles and accurately capture complex rhythmic structures. To overcome these challenges, we propose a novel beat tracking paradigm BeatFM, which introduces a pre-trained music foundation model and leverages its rich semantic knowledge to improve beat tracking performance. Pre-training on diverse music datasets endows music foundation models with a robust understanding of music, thereby effectively addressing these challenges. To further adapt it for beat tracking, we design a plug-and-play multi-dimensional semantic aggregation module, which is composed of three parallel sub-modules, each focusing on semantic aggregation in the temporal, frequency, and channel domains, respectively. Extensive experiments demonstrate that our method achieves state-of-the-art performance in beat and downbeat tracking across multiple benchmark datasets.


【4】HingeNet: A Harmonic-Aware Fine-Tuning Approach for Beat Tracking
标题:HingeNet:一种用于节拍跟踪的和声感知微调方法
链接:https://arxiv.org/abs/2508.09788

作者:Ganghui Ru, Jieying Wang, Jiahao Zhao, Yulun Wu, Yi Yu, Nannan Jiang, Wei Wang, Wei Li
备注:This paper has been accepted by ICME2025
摘要:微调预训练的基础模型在音乐信息检索方面取得了重大进展。然而,将这些模型应用于节拍跟踪任务仍然是未知的,因为有限的注释数据使得传统的微调方法无效。为了应对这一挑战,我们提出了HingeNet,一种专门为节拍跟踪任务设计的新颖且通用的参数有效的微调方法。HingeNet是一个轻量级和可分离的网络,视觉上类似于铰链,旨在通过使用其中间特征表示作为输入与预训练的基础模型紧密对接。这种独特的架构赋予了HingeNet广泛的通用性,使其能够与各种预先训练的基础模型进行有效集成。此外,考虑到谐波在节拍跟踪中的重要性,我们在微调过程中引入谐波感知机制,以更好地捕获和强调音乐信号中的谐波结构。在基准数据集上的实验表明,HingeNet在节拍和下拍跟踪方面达到了最先进的性能
摘要:Fine-tuning pre-trained foundation models has made significant progress in music information retrieval. However, applying these models to beat tracking tasks remains unexplored as the limited annotated data renders conventional fine-tuning methods ineffective. To address this challenge, we propose HingeNet, a novel and general parameter-efficient fine-tuning method specifically designed for beat tracking tasks. HingeNet is a lightweight and separable network, visually resembling a hinge, designed to tightly interface with pre-trained foundation models by using their intermediate feature representations as input. This unique architecture grants HingeNet broad generalizability, enabling effective integration with various pre-trained foundation models. Furthermore, considering the significance of harmonics in beat tracking, we introduce harmonic-aware mechanism during the fine-tuning process to better capture and emphasize the harmonic structures in musical signals. Experiments on benchmark datasets demonstrate that HingeNet achieves state-of-the-art performance in beat and downbeat tracking


【5】MetaGuardian: Enhancing Voice Assistant Security through Advanced Acoustic Metamaterials
标题:MetaGuardian:通过先进的声学超材料增强语音助手的安全性
链接:https://arxiv.org/abs/2508.09728

作者:Zhiyuan Ning, Zheng Wang, Zhanyong Tang
摘要:我们提出MetaGuardian,一个基于声学超材料的语音助手(VA)保护系统。MetaGuardian可以直接集成到各种智能设备的外壳中,有效地防御听不见的,敌对的和激光攻击,而无需依赖额外的软件支持或改变底层硬件,确保可用性。为了实现这一目标,MetaGuardian利用超材料单元之间的互阻抗效应,将信号滤波范围扩展到16-40 kHz,以有效阻止宽带听不见的攻击。此外,它还采用了精心设计的螺旋空间结构,可以精确干扰敌方攻击,同时确保VA的正常工作。此外,MetaGuardian提供了通用的结构设计,使其能够灵活地适应各种智能设备,在便携性和保护有效性之间取得平衡。在受控的评估环境中,MetaGuardian对各种攻击类型(包括对抗性攻击、听不见的攻击和激光攻击)实现了高防御成功率。
摘要:We present MetaGuardian, a voice assistant (VA) protection system based on acoustic metamaterials. MetaGuardian can be directly integrated into the enclosures of various smart devices, effectively defending against inaudible, adversarial and laser attacks without relying on additional software support or altering the underlying hardware, ensuring usability. To achieve this, MetaGuardian leverages the mutual impedance effects between metamaterial units to extend the signal filtering range to 16-40 kHz to effectively block wide-band inaudible attacks. Additionally, it adopts a carefully designed coiled space structure to precisely interfere with adversarial attacks while ensuring the normal functioning of VAs. Furthermore, MetaGuardian offers a universal structural design, allowing itself to be flexibly adapted to various smart devices, striking a balance between portability and protection effectiveness. In controled evaluation environments, MetaGuardian achieves a high defense success rate against various attack types, including adversarial, inaudible and laser attacks.


【6】OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
标题:OSUM-EChat:通过理解驱动的口语对话增强端到端的移情口语聊天机器人
链接:https://arxiv.org/abs/2508.09600

作者:Xuelong Geng, Qijie Shao, Hongfei Xue, Shuiyuan Wang, Hanke Xie, Zhao Guo, Yi Zhao, Guojian Li, Wenjie Tian, Chengyou Wang, Zhixian Zhao, Kangxiang Xia, Ziyu Zhang, Zhennan Lin, Tianlun Zuo, Mingchen Shao, Yuang Cao, Guobin Ma, Longhao Li, Yuhang Dai, Dehui Gao, Dake Guo, Lei Xie
摘要:同理心对于在口语对话系统中实现自然交互至关重要,允许机器识别并适当地响应年龄,性别和情绪等非语言线索。端到端语音语言模型的最新进展,统一了语音理解和生成,提供了有前途的解决方案。然而,一些挑战仍然存在,包括过度依赖大规模对话数据集,对传达同理心至关重要的非语言线索提取不足,以及缺乏同理心特定数据集和评估框架。为了解决这些问题,我们引入OSUM-EChat,这是一个开源的端到端口语对话系统,旨在增强移情互动,特别是在资源有限的环境中。OSUM-EChat引入了两个关键创新:(1)三阶段理解驱动的口语对话训练策略,将大型语音理解模型的功能扩展到口语对话任务,以及(2)语言-非语言双重思维机制,通过思维链将非语言理解与对话生成相结合,使系统能够产生更多的移情反应。这种方法减少了对大规模对话数据集的依赖,同时保持了高质量的移情互动。此外,我们还介绍了EChat-200 K数据集,这是一个丰富的移情语音对话语料库,以及ECat-eval基准,这是一个用于评估对话系统移情能力的综合框架。实验结果表明,OSUM-EChat优于端到端的口语对话模型的移情反应,验证了其有效性。
摘要:Empathy is crucial in enabling natural interactions within spoken dialogue systems, allowing machines to recognize and respond appropriately to paralinguistic cues such as age, gender, and emotion. Recent advancements in end-to-end speech language models, which unify speech understanding and generation, provide promising solutions. However, several challenges persist, including an over-reliance on large-scale dialogue datasets, insufficient extraction of paralinguistic cues vital for conveying empathy, and the lack of empathy-specific datasets and evaluation frameworks. To address these issues, we introduce OSUM-EChat, an open-source, end-to-end spoken dialogue system designed to enhance empathetic interactions, particularly in resource-limited settings. OSUM-EChat introduces two key innovations: (1) a three-stage understanding-driven spoken dialogue training strategy that extends the capabilities of a large speech understanding model to spoken dialogue tasks, and (2) a linguistic-paralinguistic dual thinking mechanism that integrates paralinguistic understanding through a chain of thought with dialogue generation, enabling the system to produce more empathetic responses. This approach reduces reliance on large-scale dialogue datasets while maintaining high-quality empathetic interactions. Additionally, we introduce the EChat-200K dataset, a rich corpus of empathetic speech-to-speech dialogues, and the EChat-eval benchmark, a comprehensive framework for evaluating the empathetic capabilities of dialogue systems. Experimental results demonstrate that OSUM-EChat outperforms end-to-end spoken dialogue models regarding empathetic responsiveness, validating its effectiveness.


【7】Leveraging Zipformer Model for Effective Language Identification in Code-Switched Child-Directed Speech
标题:利用Zipformer模型在代码切换儿童引导语音中进行有效语言识别
链接:https://arxiv.org/abs/2508.09430

作者:Lavanya Shankar, Leibny Paola Garcia Perera
摘要:在儿童导向的场景中,语码转换和语言识别提出了重大挑战,特别是在双语环境中。本文通过使用Zipformer来处理语音的细微差别来解决这一挑战,其中包含两种不平衡的语言,普通话和英语,在一个话语中。这项工作表明,Zipformer的内部层有效地编码的语言特征,可以利用在语言识别。我们提出了内层的选择方法来提取嵌入,并与不同的后端进行比较。我们的分析表明,Zipformer在这些后端上都是健壮的。我们的方法有效地处理了不平衡的数据,实现了81.89%的平衡准确率(BAC),比语言识别基线提高了15.47%。这些发现突出了Transformer编码器架构模型在实际场景中的潜力。
摘要:Code-switching and language identification in child-directed scenarios present significant challenges, particularly in bilingual environments. This paper addresses this challenge by using Zipformer to handle the nuances of speech, which contains two imbalanced languages, Mandarin and English, in an utterance. This work demonstrates that the internal layers of the Zipformer effectively encode the language characteristics, which can be leveraged in language identification. We present the selection methodology of the inner layers to extract the embeddings and make a comparison with different back-ends. Our analysis shows that Zipformer is robust across these backends. Our approach effectively handles imbalanced data, achieving a Balanced Accuracy (BAC) of 81.89%, a 15.47% improvement over the language identification baseline. These findings highlight the potential of the transformer encoder architecture model in real scenarios.


【8】$\text{M}^3\text{PDB}$: A Multimodal, Multi-Label, Multilingual Prompt Database for Speech Generation
标题:$ ext{M}#3 ext{DBC}$:用于语音生成的多模式、多标签、多语言提示数据库
链接:https://arxiv.org/abs/2508.09702

作者: Boyu Zhu, Cheng Gong, Muyang Wu, Ruihao Jing, Fan Liu, Xiaolei Zhang, Chi Zhang, Xuelong Li
摘要:zero-shot语音生成的最新进展使得模型能够从语音提示合成模仿说话者身份和说话风格的语音。然而,这些模型的有效性在高质量的语音提示不存在、不完整或不在域的真实世界场景中受到显著限制。这个问题主要是由于用于模型训练的语音数据和推理期间的输入提示语音之间的显著质量不匹配引起的。为了解决这个问题,我们引入了$\text{M}^3\text{PDB}$,这是第一个大规模、多模态、多标签和多语言的提示数据库,专为语音生成中的鲁棒提示选择而设计。我们的数据集构建利用了一种新颖的多模态、多主体注释框架,能够在不同的模态中进行精确和分层的标记。此外,我们提出了一个轻量级的,但有效的提示选择策略量身定制的实时,资源受限的推理设置。实验结果表明,我们提出的数据库和选择策略有效地支持各种具有挑战性的语音生成场景。我们希望我们的工作可以激励社区将重点从提高标准基准测试的性能转移到解决语音生成中更现实和多样化的应用场景。代码和数据集可在https://github.com/hizening/M3PDB上获得。
摘要:Recent advancements in zero-shot speech generation have enabled models to synthesize speech that mimics speaker identity and speaking style from speech prompts. However, these models' effectiveness is significantly limited in real-world scenarios where high-quality speech prompts are absent, incomplete, or out of domain. This issue arises primarily from a significant quality mismatch between the speech data utilized for model training and the input prompt speech during inference. To address this, we introduce $\text{M}^3\text{PDB}$, the first large-scale, multi-modal, multi-label, and multilingual prompt database designed for robust prompt selection in speech generation. Our dataset construction leverages a novel multi-modal, multi-agent annotation framework, enabling precise and hierarchical labeling across diverse modalities. Furthermore, we propose a lightweight yet effective prompt selection strategy tailored for real-time, resource-constrained inference settings. Experimental results demonstrate that our proposed database and selection strategy effectively support various challenging speech generation scenarios. We hope our work can inspire the community to shift focus from improving performance on standard benchmarks to addressing more realistic and diverse application scenarios in speech generation. Code and dataset are available at: https://github.com/hizening/M3PDB.


【9】ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
标题:ProMode:一个以声学和文本输入为条件的语音韵律模型
链接:https://arxiv.org/abs/2508.09389

作者:Eray Eren, Qingju Liu, Hyeongwoo Kim, Pablo Garrido, Abeer Alwan
备注:Interspeech 2025; demo page at this https URL
摘要:韵律不仅传达了语音信号丰富的情感和语义信息,而且也传达了个体的个性特征。我们提出了一个独立的模型,将文本映射到韵律特征,如F0和能量,并可用于下游任务,如TTS。ProMode编码器将声学特征和时间对齐的文本内容作为输入,两者都被部分掩蔽,并获得固定长度的潜在韵律嵌入。解码器使用编码的韵律输入和未掩蔽的文本内容两者来预测掩蔽区域中的声学。在GigaSpeech数据集上训练,我们将我们的方法与最先进的风格编码器进行了比较。对于F0和能量预测,我们在不同的粒度级别上显示了我们的模型的一致改进。我们还将这些预测的韵律特征集成到TTS系统中,并进行感知测试,与基线相比,这些测试显示出更高的韵律偏好,表明该模型在韵律建模很重要的任务中具有潜力。
摘要:Prosody conveys rich emotional and semantic information of the speech signal as well as individual idiosyncrasies. We propose a stand-alone model that maps text-to-prosodic features such as F0 and energy and can be used in downstream tasks such as TTS. The ProMode encoder takes as input acoustic features and time-aligned textual content, both are partially masked, and obtains a fixed-length latent prosodic embedding. The decoder predicts acoustics in the masked region using both the encoded prosody input and unmasked textual content. Trained on the GigaSpeech dataset, we compare our method with state-of-the-art style encoders. For F0 and energy predictions, we show consistent improvements for our model at different levels of granularity. We also integrate these predicted prosodic features into a TTS system and conduct perceptual tests, which show higher prosody preference compared to the baselines, demonstrating the model's potential in tasks where prosody modeling is important.


eess.AS音频处理


【1】Improving the Speaker Anonymization Evaluation's Robustness to Target Speakers with Adversarial Learning
标题:利用对抗学习提高说话人识别评价对目标说话人的鲁棒性
链接:https://arxiv.org/abs/2508.09803

作者:Carlos Franzreb, Arnab Das, Tim Polzehl, Sebastian Möller
摘要:目前的隐私评估扬声器匿名往往高估隐私时,使用同性目标选择算法(TSA),虽然TSA泄漏扬声器的性别,因此应该更容易受到攻击。我们假设发生这种情况是因为评估没有考虑到匿名语音包含来自源和目标说话者的信息这一事实。为了解决这个问题,我们建议添加一个目标分类器,它可以测量目标说话者信息在评估中的影响,这也可以通过对抗学习来消除。实验表明,这种方法对于多个匿名者是有效的,特别是当使用同性TSA时,可以进行更可靠的评估。
摘要:The current privacy evaluation for speaker anonymization often overestimates privacy when a same-gender target selection algorithm (TSA) is used, although this TSA leaks the speaker's gender and should hence be more vulnerable. We hypothesize that this occurs because the evaluation does not account for the fact that anonymized speech contains information from both the source and target speakers. To address this, we propose to add a target classifier that measures the influence of target speaker information in the evaluation, which can also be removed with adversarial learning. Experiments demonstrate that this approach is effective for multiple anonymizers, particularly when using a same-gender TSA, leading to a more reliable assessment.


【2】$\text{M}^3\text{PDB}$: A Multimodal, Multi-Label, Multilingual Prompt Database for Speech Generation
标题:$ ext{M}#3 ext{DBC}$:用于语音生成的多模式、多标签、多语言提示数据库
链接:https://arxiv.org/abs/2508.09702

作者:Boyu Zhu, Cheng Gong, Muyang Wu, Ruihao Jing, Fan Liu, Xiaolei Zhang, Chi Zhang, Xuelong Li
摘要:zero-shot语音生成的最新进展使得模型能够从语音提示合成模仿说话者身份和说话风格的语音。然而,这些模型的有效性在高质量的语音提示不存在、不完整或不在域的真实世界场景中受到显著限制。这个问题主要是由于用于模型训练的语音数据和推理期间的输入提示语音之间的显著质量不匹配引起的。为了解决这个问题,我们引入了$\text{M}^3\text{PDB}$,这是第一个大规模、多模态、多标签和多语言的提示数据库,专为语音生成中的鲁棒提示选择而设计。我们的数据集构建利用了一种新颖的多模态、多主体注释框架,能够在不同的模态中进行精确和分层的标记。此外,我们提出了一个轻量级的,但有效的提示选择策略量身定制的实时,资源受限的推理设置。实验结果表明,我们提出的数据库和选择策略有效地支持各种具有挑战性的语音生成场景。我们希望我们的工作可以激励社区将重点从提高标准基准测试的性能转移到解决语音生成中更现实和多样化的应用场景。代码和数据集可在https://github.com/hizening/M3PDB上获得。
摘要:Recent advancements in zero-shot speech generation have enabled models to synthesize speech that mimics speaker identity and speaking style from speech prompts. However, these models' effectiveness is significantly limited in real-world scenarios where high-quality speech prompts are absent, incomplete, or out of domain. This issue arises primarily from a significant quality mismatch between the speech data utilized for model training and the input prompt speech during inference. To address this, we introduce $\text{M}^3\text{PDB}$, the first large-scale, multi-modal, multi-label, and multilingual prompt database designed for robust prompt selection in speech generation. Our dataset construction leverages a novel multi-modal, multi-agent annotation framework, enabling precise and hierarchical labeling across diverse modalities. Furthermore, we propose a lightweight yet effective prompt selection strategy tailored for real-time, resource-constrained inference settings. Experimental results demonstrate that our proposed database and selection strategy effectively support various challenging speech generation scenarios. We hope our work can inspire the community to shift focus from improving performance on standard benchmarks to addressing more realistic and diverse application scenarios in speech generation. Code and dataset are available at: https://github.com/hizening/M3PDB.


【3】ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
标题:ProMode:一个以声学和文本输入为条件的语音韵律模型
链接:https://arxiv.org/abs/2508.09389

作者:Eray Eren, Qingju Liu, Hyeongwoo Kim, Pablo Garrido, Abeer Alwan
备注:Interspeech 2025; demo page at this https URL
摘要:韵律不仅传达了语音信号丰富的情感和语义信息,而且也传达了个体的个性特征。我们提出了一个独立的模型,将文本映射到韵律特征,如F0和能量,并可用于下游任务,如TTS。ProMode编码器将声学特征和时间对齐的文本内容作为输入,两者都被部分掩蔽,并获得固定长度的潜在韵律嵌入。解码器使用编码的韵律输入和未掩蔽的文本内容两者来预测掩蔽区域中的声学。在GigaSpeech数据集上训练,我们将我们的方法与最先进的风格编码器进行了比较。对于F0和能量预测,我们在不同的粒度级别上显示了我们的模型的一致改进。我们还将这些预测的韵律特征集成到TTS系统中,并进行感知测试,与基线相比,这些测试显示出更高的韵律偏好,表明该模型在韵律建模很重要的任务中具有潜力。
摘要:Prosody conveys rich emotional and semantic information of the speech signal as well as individual idiosyncrasies. We propose a stand-alone model that maps text-to-prosodic features such as F0 and energy and can be used in downstream tasks such as TTS. The ProMode encoder takes as input acoustic features and time-aligned textual content, both are partially masked, and obtains a fixed-length latent prosodic embedding. The decoder predicts acoustics in the masked region using both the encoded prosody input and unmasked textual content. Trained on the GigaSpeech dataset, we compare our method with state-of-the-art style encoders. For F0 and energy predictions, we show consistent improvements for our model at different levels of granularity. We also integrate these predicted prosodic features into a TTS system and conduct perceptual tests, which show higher prosody preference compared to the baselines, demonstrating the model's potential in tasks where prosody modeling is important.


【4】Fake-Mamba: Real-Time Speech Deepfake Detection Using Bidirectional Mamba as Self-Attention's Alternative
标题:Fake-Mamba:使用双向Mamba作为自我注意替代方案的实时语音深度伪造检测
链接:https://arxiv.org/abs/2508.09294

作者:Xi Xuan, Zimo Zhu, Wenxin Zhang, Yi-Cheng Lin, Tomi Kinnunen
备注:Accepted at IEEE ASRU 2025
摘要:语音合成的进步加剧了安全威胁,激发了实时深度造假检测研究。我们调查是否双向曼巴可以作为一个有竞争力的替代自我注意力在检测合成语音。我们的解决方案Fake-Mamba将XLSR前端与双向Mamba集成,以捕获本地和全局工件。我们的核心创新引入了三种高效编码器:TransBiMamba、ConBiMamba和PN-BiMamba。利用XLSR丰富的语言表示,PN-BiMamba可以有效地捕捉合成语音的微妙线索。在ASVspoof 21 LA、21 DF和In-The-Wild基准测试中,Fake-Mamba分别达到0.97%、1.74%和5.85%的EER,这代表了相对于SOTA模型XLSR-Conformer和XLSR-Mamba的实质性相对增益。该框架保持跨话语长度的实时推理,表现出很强的泛化能力和实际可行性。该代码可在https://github.com/xuanxixi/Fake-Mamba上获得。
摘要:Advances in speech synthesis intensify security threats, motivating real-time deepfake detection research. We investigate whether bidirectional Mamba can serve as a competitive alternative to Self-Attention in detecting synthetic speech. Our solution, Fake-Mamba, integrates an XLSR front-end with bidirectional Mamba to capture both local and global artifacts. Our core innovation introduces three efficient encoders: TransBiMamba, ConBiMamba, and PN-BiMamba. Leveraging XLSR's rich linguistic representations, PN-BiMamba can effectively capture the subtle cues of synthetic speech. Evaluated on ASVspoof 21 LA, 21 DF, and In-The-Wild benchmarks, Fake-Mamba achieves 0.97%, 1.74%, and 5.85% EER, respectively, representing substantial relative gains over SOTA models XLSR-Conformer and XLSR-Mamba. The framework maintains real-time inference across utterance lengths, demonstrating strong generalization and practical viability. The code is available at https://github.com/xuanxixi/Fake-Mamba.


【5】Objective Soups: Multilingual Multi-Task Modeling for Speech Processing
标题:目标汤:语音处理的多语言多任务建模
链接:https://arxiv.org/abs/2508.09228

作者:A F M Saiff, Lisha Chen, Xiaodong Cui, Songtao Lu, Brian Kingsbury, Tianyi Chen
摘要:训练多语言、多任务语音处理(MSP)的单一模型受到语音识别和翻译等任务之间目标冲突的严重阻碍。虽然多目标优化(MOO)旨在对齐梯度更新,但其有效性随着任务数量的增加而降低,因此很难找到共同的下降方向。这就提出了一个根本性的问题:高度冲突的目标应该联合优化还是分成层次结构优化?为了解决这个问题,本文研究了三个多目标MSP配方,我们称之为\textbf{目标汤配方}。这些公式在不同的优化级别应用多目标优化,以减轻所有目标之间的潜在冲突。为了确保效率,我们引入了一个轻量级的层选择机制,只使用最有问题的层来计算避免冲突的梯度,从而最大限度地减少计算和内存开销。在CoVoST v2、LibriSpeech和AISHELL-1上进行的大量实验表明,将识别和翻译任务分开的双层配方始终优于标准的平面优化。我们的工作表明,分层MOO是一个更有效的和可扩展的方法,为建设国家的最先进的MSP模型。我们的代码已在https://github.com/afmsaif/Objective_Soups上发布。
摘要:Training a single model for multilingual, multi-task speech processing (MSP) is severely hampered by conflicting objectives between tasks like speech recognition and translation. While multi-objective optimization (MOO) aims to align gradient updates, its effectiveness diminishes as the number of tasks grows, making it difficult to find a common descent direction. This raises a fundamental question: should highly conflicting objectives be optimized jointly or separated into a hierarchical structure? To address this question, this paper investigates three multi-objective MSP formulations, which we refer to as \textbf{objective soup recipes}. These formulations apply multi-objective optimization at different optimization levels to mitigate potential conflicts among all objectives. To ensure efficiency, we introduce a lightweight layer-selection mechanism that computes the conflict-avoiding gradient using only the most problematic layers, minimizing computational and memory overhead. Extensive experiments on CoVoST v2, LibriSpeech, and AISHELL-1 reveal that a bi-level recipe separating recognition and translation tasks consistently outperforms standard flat optimization. Our work demonstrates that hierarchical MOO is a more effective and scalable approach for building state-of-the-art MSP models. Our code has been released at https://github.com/afmsaif/Objective_Soups.


【6】UtterTune: LoRA-Based Target-Language Pronunciation Edit and Control in Multilingual Text-to-Speech
标题:UtterButton:多语言文本到语音中基于LoRA的目标语言发音编辑和控制
链接:https://arxiv.org/abs/2508.09767

作者:Shuhei Kato
摘要:我们提出了UtterTune,一个轻量级的适应方法,微调多语言的文本到语音(TTS)系统的基础上,一个大的语言模型(LLM)架构,旨在提高目标语言的发音的可控性,同时保留在其他性能。虽然LLM架构已经使TTS模型能够实现显著的自然度,但是准确地建模字素到音素(G2P)映射和韵律仍然具有挑战性,特别是当模型省略显式G2P模块并且直接处理最小编码文本(例如,字节对编码)。UtterTune利用低秩自适应,使控制的分段发音和音高口音在音素水平的日语语音,本文中的目标语言,同时保持自然度和扬声器相似性在zero-shot设置。客观和主观评价证实了其有效性。
摘要:We propose UtterTune, a lightweight adaptation method that fine-tunes a multilingual text-to-speech (TTS) system based on a large language model (LLM) architecture, designed to enhance the controllability of pronunciation in a target language while preserving performance in others. While LLM architectures have enabled TTS models to achieve remarkable naturalness, accurately modeling grapheme-to-phoneme (G2P) mapping and prosody remains challenging, especially when the model omits an explicit G2P module and directly processes minimally encoded text (e.g., byte-pair encoding). UtterTune leverages low-rank adaptation to enable the control of segmental pronunciation and pitch accent at the phoneme level for Japanese speech, the target language in this paper, while maintaining naturalness and speaker similarity in a zero-shot setting. Objective and subjective evaluations confirm its effectiveness.


机器翻译由腾讯交互翻译提供,仅供参考