今日论文合集:cs.SD语音7篇,eess.AS音频处理5篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pretraining
标题:Fineguardly:驯服异类监督以实现细粒度语音预训练
链接:https://arxiv.org/abs/2604.01155

作者:Xiquan Li,Xuenan Xu,Ziyang Ma,Wenxi Chen,Haolin He,Qiuqiang Kong,Xie Chen
摘要:对比预训练的音频语言模型(例如,CLAP)擅长于剪辑级的理解,但在框架级的任务上却很吃力。现有的扩展未能利用真实世界的音频文本数据,其中大量的剪辑级文本描述与有限的帧级注释共存的粒度变化。本文提出了一种新的训练范式,即细粒度的音频预训练(Fine-grained audio Pretraining,Fine-grained audio pretraining),它可以在CLAP中利用异构数据实现剪辑级和帧级对齐。FineEncoder引入了一种双流S形损失和基于聚类的采样策略,以共同从剪辑级和帧级监督中学习。为了捕获全局语义和局部细节,FineTime在自监督编码器之上使用解耦的音频投影器。为了缓解时间注释数据的稀缺性,我们提出了FineWorks-100 k,这是一个通过可扩展的策展管道构建的大规模合成SED数据集。大量的实验表明,FineReader在多个音频理解任务中实现了SOTA性能,包括检索,分类,声音事件检测和文本到音频接地。消融研究进一步表明,粗粒度和细粒度对齐是互利的,为构建更好的音频语言模型(ALM)提供了见解。
摘要:Contrastively pretrained audio-language models (e.g., CLAP) excel at clip-level understanding but struggle with frame-level tasks. Existing extensions fail to exploit the varying granularity of real-world audio-text data, where massive clip-level textual descriptions coexist with limited frame-level annotations. This paper proposes Fine-grained Language-Audio Pretraining (FineLAP), a novel training paradigm that advances both clip- and frame-level alignment in CLAP with heterogeneous data. FineLAP introduces a dual-stream sigmoid loss with a cluster-based sampling strategy to jointly learn from clip- and frame-level supervision. To capture both global semantics and local details, FineLAP uses a decoupled audio projector on top of a self-supervised encoder. To alleviate the scarcity of temporally annotated data, we present FineLAP-100k, a large-scale synthetic SED dataset constructed through a scalable curation pipeline. Extensive experiments demonstrate that FineLAP achieves SOTA performance across multiple audio understanding tasks, including retrieval, classification, sound event detection, and text-to-audio grounding. Ablation studies further show that coarse- and fine-grained alignment are mutually beneficial, providing insights for building better audio-language models (ALMs).


【2】TRACE: Training-Free Partial Audio Deepfake Detection via Embedding Trajectory Analysis of Speech Foundation Models
标题:TRACE:通过嵌入语音基础模型轨迹分析进行免训练部分音频深度伪造检测
链接:https://arxiv.org/abs/2604.01083

作者:Awais Khan,Muhammad Umar Farooq,Kutub Uddin,Khalid Malik
摘要:部分音频deepfakes,其中合成片段拼接到真正的录音中,特别具有欺骗性,因为大多数音频仍然是真实的。现有的探测器受到监督:它们需要帧级注释、对特定合成流水线的过拟合,并且必须随着新的生成模型的出现而重新训练。我们认为这种监督是不必要的。我们假设语音基础模型隐含地编码了一个法医信号:真正的语音形式平滑,缓慢变化的嵌入轨迹,而拼接边界在帧级转换中引入突然中断。在此基础上,我们提出了TRACE(Training-free Representation-based Audio Countermeasure via Embedding dynamics),这是一个无需训练的框架,通过分析冻结语音基础模型表示的一阶动态来检测部分音频deepfake,而无需任何训练,标记数据或架构修改。我们评估跟踪四个基准,跨越两种语言使用六个语音基础模型。在PartialSpoof中,TRACE实现了8.08%的EER,与微调的监督基线相比具有竞争力。在LlamaPartialSpoof中,最具挑战性的基准是LLM驱动的商业合成,TRACE完全超过了监督基线(24.12% vs. 24.49%EER),没有任何目标域数据。这些结果表明,语音基础模型中的时间动力学为免训练音频取证提供了有效的,概括的信号。
摘要:Partial audio deepfakes, where synthesized segments are spliced into genuine recordings, are particularly deceptive because most of the audio remains authentic. Existing detectors are supervised: they require frame-level annotations, overfit to specific synthesis pipelines, and must be retrained as new generative models emerge. We argue that this supervision is unnecessary. We hypothesize that speech foundation models implicitly encode a forensic signal: genuine speech forms smooth, slowly varying embedding trajectories, while splice boundaries introduce abrupt disruptions in frame-level transitions. Building on this, we propose TRACE (Training-free Representation-based Audio Countermeasure via Embedding dynamics), a training-free framework that detects partial audio deepfakes by analyzing the first-order dynamics of frozen speech foundation model representations without any training, labeled data, or architectural modification. We evaluate TRACE on four benchmarks that span two languages using six speech foundation models. In PartialSpoof, TRACE achieves 8.08% EER, competitive with fine-tuned supervised baselines. In LlamaPartialSpoof, the most challenging benchmark featuring LLM-driven commercial synthesis, TRACE surpasses a supervised baseline outright (24.12% vs. 24.49% EER) without any target-domain data. These results show that temporal dynamics in speech foundation models provide an effective, generalize signal for training-free audio forensics.


【3】Sona: Real-Time Multi-Target Sound Attenuation for Noise Sensitivity
标题:Sona:实时多目标声音衰减,提高噪音敏感度
链接:https://arxiv.org/abs/2604.00447

作者:Jeremy Zhengqi Huang,Emani Hicks,Sidharth,Gillian R. Hayes,Dhruv Jain
备注:12 pages, 6 figures
摘要:对于对噪音敏感的人来说,日常的音景可能会让人不知所措。现有的工具,如主动噪声消除,通过抑制整个声学环境来减少不适,通常是以周围的人和事件的意识为代价的。我们提出了索纳,一个互动的移动系统,实时音景调解,选择性地衰减烦人的声音,同时保留所需的音频。Sona建立在目标调节神经管道上,支持多个重叠声源的同时衰减,克服了现有系统的单目标限制。它在设备上实时运行,并通过现场音频示例支持用户可扩展的声音类,无需重新训练。Sona通过对68名噪音敏感者的形成性研究获得了信息。通过技术基准测试和10名参与者的现场研究,我们表明Sona实现了适合现场聆听的低延迟,多目标衰减,并在保持对周围环境的感知的同时,有效减少了令人烦恼的声音。这些结果指向了一类新的个人AI系统,通过调解现实世界的声学环境来支持舒适和社会参与。
摘要:For people with noise sensitivity, everyday soundscapes can be overwhelming. Existing tools such as active noise cancellation reduce discomfort by suppressing the entire acoustic environment, often at the cost of awareness of surrounding people and events. We present Sona, an interactive mobile system for real-time soundscape mediation that selectively attenuates bothersome sounds while preserving desired audio. Sona is built on a target-conditioned neural pipeline that supports simultaneous attenuation of multiple overlapping sound sources, overcoming the single-target limitation of prior systems. It runs in real time on-device and supports user-extensible sound classes through in-situ audio examples, without retraining. Sona is informed by a formative study with 68 noise-sensitive individuals. Through technical benchmarking and an in-situ study with 10 participants, we show that Sona achieves low-latency, multi-target attenuation suitable for live listening, and enables meaningful reductions in bothersome sounds while maintaining awareness of surroundings. These results point toward a new class of personal AI systems that support comfort and social participation by mediating real-world acoustic environments.


【4】Vocal Prognostic Digital Biomarkers in Monitoring Chronic Heart Failure: A Longitudinal Observational Study
标题:监测慢性心力衰竭的声音预后数字生物标志物:一项纵向观察研究
链接:https://arxiv.org/abs/2604.00308

作者:Fan Wu,Matthias P. Nägele,Daryush D. Mehta,Elgar Fleisch,Frank Ruschitzka,Andreas J. Flammer,Filipe Barata
摘要:目的:本研究旨在评估哪些声音特征可以预测慢性心力衰竭患者的健康恶化。   背景资料:心力衰竭(HF)是一种慢性疾病,具有进行性恶化和急性失代偿,通常需要住院治疗,并造成大量的医疗保健和经济负担。当前的标准护理(SoC)家庭监测(例如体重跟踪)缺乏预测准确性,并且需要患者高度参与。声音是一种有前途的非侵入性生物标志物,尽管先前的研究主要集中在急性HF阶段。   研究方法:在一项为期2个月的纵向研究中,32名HF患者在家中收集了每日语音记录和体重和血压的SoC测量值,并每两周进行一次健康状况问卷调查。声学分析生成详细的元音和语音特征。时间序列特征从聚合的回顾窗口(例如,7天),以预测第二天的健康状况。可解释的机器学习与嵌套交叉验证确定了顶级的声音生物标志物,案例研究说明了模型应用。   结果:共分析了21,863份记录。声学元音特征与健康状况有很强的相关性。回顾窗口内的时间序列语音特征优于相应的标准护理措施,SoC指标的峰值灵敏度和特异性分别为0.826和0.782,而0.783和0.567。识别恶化的关键预后语音特征包括延迟能量转移、低能量变异性和元音的较高微光变异性,以及说话和清晰度降低、发声比降低、语音质量降低和语音共振峰变异性增加。   结论:基于语音的监测提供了一种非侵入性方法来检测慢性HF的早期健康变化,支持主动和个性化的护理。
摘要:Objective: This study aimed to evaluate which voice features can predict health deterioration in patients with chronic HF.   Background: Heart failure (HF) is a chronic condition with progressive deterioration and acute decompensations, often requiring hospitalization and imposing substantial healthcare and economic burdens. Current standard-of-care (SoC) home monitoring, such as weight tracking, lacks predictive accuracy and requires high patient engagement. Voice is a promising non-invasive biomarker, though prior studies have mainly focused on acute HF stages.   Methods: In a 2-month longitudinal study, 32 patients with HF collected daily voice recordings and SoC measures of weight and blood pressure at home, with biweekly questionnaires for health status. Acoustic analysis generated detailed vowel and speech features. Time-series features were extracted from aggregated lookback windows (e.g., 7 days) to predict next-day health status. Explainable machine learning with nested cross-validation identified top vocal biomarkers, and a case study illustrated model application.   Results: A total of 21,863 recordings were analyzed. Acoustic vowel features showed strong correlations with health status. Time-series voice features within the lookback window outperformed corresponding standard care measures, achieving peak sensitivity and specificity of 0.826 and 0.782 versus 0.783 and 0.567 for SoC metrics. Key prognostic voice features identifying deterioration included delayed energy shift, low energy variability, and higher shimmer variability in vowels, along with reduced speaking and articulation rate, lower phonation ratio, decreased voice quality, and increased formant variability in speech.   Conclusion: Voice-based monitoring offers a non-invasive approach to detect early health changes in chronic HF, supporting proactive and personalized care.


【5】MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control
标题:MambaVoiceCloning:通过状态空间建模和扩散控制实现高效且富有表达力的文本到语音
链接:https://arxiv.org/abs/2604.00292

作者:Sahil Kumar,Namrataben Patel,Honggang Wang,Youshan Zhang
备注:Accepted at ICLR 2026
摘要:MambaVoiceCloning(MVC)询问基于扩散的TTS的调节路径是否可以在推理时完全只进行SSM,消除所有注意力和跨文本,节奏和韵律的显式RNN风格的递归层,同时在受控条件下保持或提高质量。MVC结合了门控双向Mamba文本编码器,由训练后丢弃的轻量级对齐教师监督的Temporal Bi-Mamba,以及具有AdaLN调制的Expressive Mamba,产生具有有界激活记忆和实际有限前瞻流的线性时间O(T)条件。与在推理时保持混合的现有Mamba-TTS系统不同,MVC在固定的StyleTTS 2梅尔扩散声码器骨干下去除了基于注意力的持续时间和风格模块。在LJSpeech/LibriTTS上训练并在VCTK,CSS 10(ES/DE/FR)和长格式Gutenberg通道上进行评估,MVC在MOS/CMOS,F0 RMSE,MCD和WER中实现了适度但统计上可靠的增益,同时将编码器参数减少到21 M并将吞吐量提高了1.6倍。扩散仍然是主要的延迟来源,但仅SSM调节可以改善内存占用、稳定性和可部署性。
摘要:MambaVoiceCloning (MVC) asks whether the conditioning path of diffusion-based TTS can be made fully SSM-only at inference, removing all attention and explicit RNN-style recurrence layers across text, rhythm, and prosody, while preserving or improving quality under controlled conditions. MVC combines a gated bidirectional Mamba text encoder, a Temporal Bi-Mamba supervised by a lightweight alignment teacher discarded after training, and an Expressive Mamba with AdaLN modulation, yielding linear-time O(T) conditioning with bounded activation memory and practical finite look-ahead streaming. Unlike prior Mamba-TTS systems that remain hybrid at inference, MVC removes attention-based duration and style modules under a fixed StyleTTS2 mel-diffusion-vocoder backbone. Trained on LJSpeech/LibriTTS and evaluated on VCTK, CSS10 (ES/DE/FR), and long-form Gutenberg passages, MVC achieves modest but statistically reliable gains over StyleTTS2, VITS, and Mamba-attention hybrids in MOS/CMOS, F0 RMSE, MCD, and WER, while reducing encoder parameters to 21M and improving throughput by 1.6x. Diffusion remains the dominant latency source, but SSM-only conditioning improves memory footprint, stability, and deployability.


【6】An Empirical Recipe for Universal Phone Recognition
标题:通用手机识别的经验秘诀
链接:https://arxiv.org/abs/2603.29042

作者:Shikhar Bharadwaj,Chin-Jou Li,Kwanghee Choi,Eunjung Yeo,William Chen,Shinji Watanabe,David R. Mortensen
备注:Submitted to Interspeech 2026. Code: https://github.com/changelinglab/PhoneticXeus
摘要:电话识别(PR)是多语言和低资源语音处理任务的关键推动因素,但强大的性能仍然难以实现。高性能的以英语为中心的模型不会跨语言泛化,而多语言模型则没有充分利用预训练的表示。目前还不清楚数据规模,架构和训练目标如何有助于多语言PR。我们提出PhoneticXEUS -在大规模多语言数据上训练,并在多语言(17.7%PFER)和带口音的英语语音(10.6%PFER)上实现最先进的性能。通过在统一方案下对100多种语言进行评估的受控消融,我们根据经验建立了我们的训练配方,并量化了SSL表示、数据规模和损失目标的影响。此外,我们还分析了跨语系、口音和发音特征的错误模式。所有数据和代码都是公开发布的。
摘要:Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive. Highly performant English-focused models do not generalize across languages, while multilingual models underutilize pretrained representations. It also remains unclear how data scale, architecture, and training objective contribute to multilingual PR. We present PhoneticXEUS -- trained on large-scale multilingual data and achieving state-of-the-art performance on both multilingual (17.7% PFER) and accented English speech (10.6% PFER). Through controlled ablations with evaluations across 100+ languages under a unified scheme, we empirically establish our training recipe and quantify the impact of SSL representations, data scale, and loss objectives. In addition, we analyze error patterns across language families, accented speech, and articulatory features. All data and code are released openly.


【7】Semantic Audio-Visual Navigation in Continuous Environments
标题:连续环境中的语义视听导航
链接:https://arxiv.org/abs/2603.19660

作者:Yichen Zeng,Hebaixu Wang,Meng Liu,Yu Zhou,Chen Gao,Kehan Chen,Gongping Huang
备注:This paper has been accepted to CVPR 2026
摘要:视听导航使体现代理导航到声音发射目标,利用听觉和视觉线索。然而,大多数现有的方法依赖于预先计算的房间脉冲响应(RIR)的双耳音频渲染,限制代理离散的网格位置,并导致空间上不连续的观察。为了建立一个更现实的设置,我们引入了连续环境中的语义视听导航(SAVN-CE),代理可以在3D空间中自由移动,并感知时间和空间相干的视听流。在这种情况下,目标可能会间歇性地变得沉默或完全停止发出声音,导致智能体丢失目标信息。为了应对这一挑战,我们提出了MAGNet,一个多模态的基于变换的模型,它联合编码空间和语义目标表示,并将历史背景与自我运动线索相结合,以实现记忆增强的目标推理。综合实验表明,MAGNet显着优于国家的最先进的方法,实现了高达12.1%的绝对提高成功率。这些结果还突出了它对短时间声音和长距离导航场景的鲁棒性。该代码可在https://github.com/yichenzeng24/SAVN-CE上获得。
摘要:Audio-visual navigation enables embodied agents to navigate toward sound-emitting targets by leveraging both auditory and visual cues. However, most existing approaches rely on precomputed room impulse responses (RIRs) for binaural audio rendering, restricting agents to discrete grid positions and leading to spatially discontinuous observations. To establish a more realistic setting, we introduce Semantic Audio-Visual Navigation in Continuous Environments (SAVN-CE), where agents can move freely in 3D spaces and perceive temporally and spatially coherent audio-visual streams. In this setting, targets may intermittently become silent or stop emitting sound entirely, causing agents to lose goal information. To tackle this challenge, we propose MAGNet, a multimodal transformer-based model that jointly encodes spatial and semantic goal representations and integrates historical context with self-motion cues to enable memory-augmented goal reasoning. Comprehensive experiments demonstrate that MAGNet significantly outperforms state-of-the-art methods, achieving up to a 12.1\% absolute improvement in success rate. These results also highlight its robustness to short-duration sounds and long-distance navigation scenarios. The code is available at https://github.com/yichenzeng24/SAVN-CE.


eess.AS音频处理


【1】Diff-VS: Efficient Audio-Aware Diffusion U-Net for Vocals Separation
标题:迪夫-VS:用于人声分离的高效音频感知扩散U-Net
链接:https://arxiv.org/abs/2604.01120

作者:Yun-Ning,Hung,Richard Vogl,Filip Korzeniowski,Igor Pereira
备注:Accepted at ICASSP 2026
摘要:虽然扩散模型以其在生成任务中的性能而闻名,但它们也已成功应用于许多其他任务,包括音频源分离。然而,目前的音乐源分离的生成方法往往在标准的客观指标上表现不佳。在本文中,我们解决这个问题,通过引入一种新的生成声乐分离模型的基础上阐明扩散模型(EDM)的框架。我们的模型处理复杂的短时傅立叶变换频谱图,并采用基于音乐知情设计选择的改进的U-Net架构。我们的方法在客观指标上匹配有区别的基线,并实现了与最先进的系统相当的感知质量,如代理主观指标所评估的那样。我们希望这些结果鼓励更广泛的探索音乐源分离的生成方法
摘要:While diffusion models are best known for their performance in generative tasks, they have also been successfully applied to many other tasks, including audio source separation. However, current generative approaches to music source separation often underperform on standard objective metrics. In this paper, we address this issue by introducing a novel generative vocal separation model based on the Elucidated Diffusion Model (EDM) framework. Our model processes complex short-time Fourier transform spectrograms and employs an improved U-Net architecture based on music-informed design choices. Our approach matches discriminative baselines on objective metrics and achieves perceptual quality comparable to state-of-the-art systems, as assessed by proxy subjective metrics. We hope these results encourage broader exploration of generative methods for music source separation


【2】VisG AV-HuBERT: Viseme-Guided AV-HuBERT
标题:VisG AV-HuBERT:Viseme-Guided AV-HuBERT
链接:https://arxiv.org/abs/2604.00982

作者:Aristeidis Papadopoulos,Rishabh Jain,Naomi Harte
备注:Includes Supplementary Material. Accepted for Publication at International Conference on Pattern Recognition 2026 - ICPR 2026. Code is available at https://github.com/aristosp/visg_avhubert
摘要:视听语音识别(AVSR)系统目前集成了大语言模型(LLM)解码器和基于变换器的编码器,实现了最先进的结果。然而,改进的语言建模与增强的视听编码的相对贡献仍然不清楚。我们提出了Viseme-Guided AV-HuBERT(VisG AV-HuBERT),这是一个多任务微调框架,它结合了辅助视位分类来加强模型对视觉发音特征的依赖。该方法通过对AV-HuBERT算法进行扩展,引入一个轻量级的视位预测子网络,显式地引导编码器保留语音的视觉信息。在LRS 3上进行评估,VisG AV-HuBERT实现了与基线AV-HuBERT相当或更高的性能,在重噪声条件下具有显着的增益。对于语音噪声,在-10 dB信噪比(SNR)下,WER从13.59%降低到6.60%(相对改善51.4%)。更深入的分析表明,在噪声类型的替代错误大幅减少,表现出改善语音单元的歧视。对LRS 2的评估证实了泛化能力。我们的研究结果表明,显式视位建模增强了编码器表示,并提供了一个基础,通过编码器级的改进,提高噪声鲁棒AVSR。
摘要:Audio-Visual Speech Recognition (AVSR) systems nowadays integrate Large Language Model (LLM) decoders with transformer-based encoders, achieving state-of-the-art results. However, the relative contributions of improved language modelling versus enhanced audiovisual encoding remain unclear. We propose Viseme-Guided AV-HuBERT (VisG AV-HuBERT), a multi-task fine-tuning framework that incorporates auxiliary viseme classification to strengthen the model's reliance on visual articulatory features. By extending AV-HuBERT with a lightweight viseme prediction sub-network, this method explicitly guides the encoder to preserve visual speech information. Evaluated on LRS3, VisG AV-HuBERT achieves comparable or improved performance over the baseline AV-HuBERT, with notable gains under heavy noise conditions. WER reduces from 13.59% to 6.60% (51.4% relative improvement) at -10 dB Signal-to-Noise Ratio (SNR) for Speech noise. Deeper analysis reveals substantial reductions in substitution errors across noise types, demonstrating improved speech unit discrimination. Evaluation on LRS2 confirms generalization capability. Our results demonstrate that explicit viseme modelling enhances encoder representations, and provides a foundation for enhancing noise-robust AVSR through encoder-level improvements.


【3】Description and Discussion on DCASE 2026 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes
标题:DUSE 2026挑战任务4:声音场景的空间语义分割描述与讨论
链接:https://arxiv.org/abs/2604.00776

作者:Masahiro Yasuda,Binh Thien Nguyen,Noboru Harada,Romain Serizel,Mayank Mishra,Marc Delcroix,Carlos Hernandez-Olivan,Shoko Araki,Daiki Takeuchi,Tomohiro Nakatani,Nobutaka Ono
摘要:本文概述了声学场景和事件的检测和分类(DCASE)2026年挑战任务4,声音场景的空间语义分割(S5)。S5任务专注于复杂空间音频混合中声音事件的联合检测和分离,为沉浸式通信奠定基础。S5任务在DCASE 2025中首次引入,在DCASE 2026任务4中继续进行关键更改,以更好地反映现实世界的条件,包括允许混合物包含同一类别的多个源,并且不包含目标源。在本文中,我们描述了任务设置,以及对评估指标和数据集的相应更新。文中还对所提交系统的实验结果进行了报告和分析。数据和代码的官方访问点是https://github.com/nttcslab/dcase2026_task4_baseline。
摘要:This paper presents an overview of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2026 Challenge Task 4, Spatial Semantic Segmentation of Sound Scenes (S5). The S5 task focuses on the joint detection and separation of sound events in complex spatial audio mixtures, contributing to the foundation of immersive communication. First introduced in DCASE 2025, the S5 task continues in DCASE 2026 Task 4 with key changes to better reflect real-world conditions, including allowing mixtures to contain multiple sources of the same class and to contain no target sources. In this paper, we describe task setting, along with the corresponding updates to the evaluation metrics and dataset. The experimental results of the submitted systems are also reported and analyzed. The official access point for data and code is https://github.com/nttcslab/dcase2026_task4_baseline.


【4】OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models
标题:OmniVoice:利用扩散语言模型实现全语言Zero-Shot文本到语音
链接:https://arxiv.org/abs/2604.00688

作者:Han Zhu,Lingxuan Ye,Wei Kang,Zengwei Yao,Liyong Guo,Fangjun Kuang,Zhifeng Han,Weiji Zhuang,Long Lin,Daniel Povey
摘要:我们提出了OmniVoice,一个大规模的多语言zero-shot文本到语音(TTS)模型,可扩展到600多种语言。其核心是一种新型的扩散语言模型风格的离散非自回归(NAR)架构。与传统的离散NAR模型在复杂的两阶段(文本到语义到声学)管道中遇到性能瓶颈不同,OmniVoice直接将文本映射到多码本声学令牌。这种简化的方法通过两个关键的技术创新来促进:(1)用于有效训练的全码本随机掩蔽策略,以及(2)从预训练的LLM初始化以确保卓越的可懂度。通过利用完全由开源数据管理的581k小时多语言数据集,OmniVoice实现了迄今为止最广泛的语言覆盖,并在中文,英文和多种多语言基准中提供了最先进的性能。我们的代码和预训练模型可在https://github.com/k2-fsa/OmniVoice上公开获取。
摘要:We present OmniVoice, a massive multilingual zero-shot text-to-speech (TTS) model that scales to over 600 languages. At its core is a novel diffusion language model-style discrete non-autoregressive (NAR) architecture. Unlike conventional discrete NAR models that suffer from performance bottlenecks in complex two-stage (text-to-semantic-to-acoustic) pipelines, OmniVoice directly maps text to multi-codebook acoustic tokens. This simplified approach is facilitated by two key technical innovations: (1) a full-codebook random masking strategy for efficient training, and (2) initialization from a pre-trained LLM to ensure superior intelligibility. By leveraging a 581k-hour multilingual dataset curated entirely from open-source data, OmniVoice achieves the broadest language coverage to date and delivers state-of-the-art performance across Chinese, English, and diverse multilingual benchmarks. Our code and pre-trained models are publicly available at https://github.com/k2-fsa/OmniVoice.


【5】An Empirical Recipe for Universal Phone Recognition
标题:通用手机识别的经验秘诀
链接:https://arxiv.org/abs/2603.29042

作者:Shikhar Bharadwaj,Chin-Jou Li,Kwanghee Choi,Eunjung Yeo,William Chen,Shinji Watanabe,David R. Mortensen
备注:Submitted to Interspeech 2026. Code: https://github.com/changelinglab/PhoneticXeus
摘要:电话识别(PR)是多语言和低资源语音处理任务的关键推动因素,但强大的性能仍然难以实现。高性能的以英语为中心的模型不会跨语言泛化,而多语言模型则没有充分利用预训练的表示。目前还不清楚数据规模,架构和训练目标如何有助于多语言PR。我们提出PhoneticXEUS -在大规模多语言数据上训练,并在多语言(17.7%PFER)和带口音的英语语音(10.6%PFER)上实现最先进的性能。通过在统一方案下对100多种语言进行评估的受控消融,我们根据经验建立了我们的训练配方,并量化了SSL表示、数据规模和损失目标的影响。此外,我们还分析了跨语系、口音和发音特征的错误模式。所有数据和代码都是公开发布的。
摘要:Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive. Highly performant English-focused models do not generalize across languages, while multilingual models underutilize pretrained representations. It also remains unclear how data scale, architecture, and training objective contribute to multilingual PR. We present PhoneticXEUS -- trained on large-scale multilingual data and achieving state-of-the-art performance on both multilingual (17.7% PFER) and accented English speech (10.6% PFER). Through controlled ablations with evaluations across 100+ languages under a unified scheme, we empirically establish our training recipe and quantify the impact of SSL representations, data scale, and loss objectives. In addition, we analyze error patterns across language families, accented speech, and articulatory features. All data and code are released openly.


机器翻译由腾讯交互翻译提供,仅供参考