微信公众号:arXiv_Daily
cs.SD语音
【1】VibeVoice Technical Report
标题:VibeVoice技术报告
链接:https://arxiv.org/abs/2508.19205
摘要:本报告介绍了VibeVoice,一种新的模型,旨在通过采用下一个令牌扩散来合成多个扬声器的长格式语音,这是一种通过扩散自回归生成潜在向量来建模连续数据的统一方法。为了实现这一点,我们引入了一种新的连续语音标记器,与流行的Encodec模型相比,它将数据压缩提高了80倍,同时保持了相当的性能。标记器有效地保持了音频保真度,同时显著提高了处理长序列的计算效率。因此,VibeVoice可以合成长达90分钟的长格式语音(在64K上下文窗口长度中),最多4个扬声器,捕捉真实的对话“vibe”,并超越开源和专有对话模型。
摘要:This report presents VibeVoice, a novel model designed to synthesize long-form speech with multiple speakers by employing next-token diffusion, which is a unified method for modeling continuous data by autoregressively generating latent vectors via diffusion. To enable this, we introduce a novel continuous speech tokenizer that, when compared to the popular Encodec model, improves data compression by 80 times while maintaining comparable performance. The tokenizer effectively preserves audio fidelity while significantly boosting computational efficiency for processing long sequences. Thus, VibeVoice can synthesize long-form speech for up to 90 minutes (in a 64K context window length) with a maximum of 4 speakers, capturing the authentic conversational ``vibe'' and surpassing open-source and proprietary dialogue models.
【2】DESAMO: A Device for Elder-Friendly Smart Homes Powered by Embedded LLM with Audio Modality
标题:DESAMO:一款适合老年人的智能家居设备,由具有音频模式的嵌入式LLM提供支持
链接:https://arxiv.org/abs/2508.18918
备注:2 pages, 2 figures. Accepted for presentation as a UIST 2025 Poster
摘要:我们介绍DESAMO,这是一个由Audio LLM提供支持的老年人友好使用的设备上智能家居系统,支持自然和私人互动。传统的语音助手依赖于基于ASR的管道或ASR-LLM级联,通常难以处理老年用户中常见的不清晰语音,无法处理非语音音频,而DESAMO利用Audio LLM直接处理原始音频输入,从而能够强大地理解用户意图和关键事件,例如跌倒或呼救。
摘要:We present DESAMO, an on-device smart home system for elder-friendly use powered by Audio LLM, that supports natural and private interactions. While conventional voice assistants rely on ASR-based pipelines or ASR-LLM cascades, often struggling with the unclear speech common among elderly users and unable to handle non-speech audio, DESAMO leverages an Audio LLM to process raw audio input directly, enabling a robust understanding of user intent and critical events, such as falls or calls for help.
【3】SegReConcat: A Data Augmentation Method for Voice Anonymization Attack
标题:SegReConcat:一种用于语音匿名攻击的数据增强方法
链接:https://arxiv.org/abs/2508.18907
备注:The Paper has been accepted by APCIPA ASC 2025
摘要:语音的匿名化试图隐藏说话者的身份,同时保持语音数据的实用性。然而,剩余的说话人线索往往会持续存在,这会带来隐私风险。我们提出了SegReConcat,攻击方增强自动说话人确认系统的数据增强方法。SegReConcat在单词级别对匿名语音进行分段,使用随机或基于相似性的策略重新排列片段以破坏长期上下文线索,并将它们与原始话语连接起来,使攻击者能够从多个角度学习源说话者的特征。所提出的方法已在七个匿名系统的VoicePrivacy Attacker Challenge 2024框架中进行了评估,SegReConcat改进了七个系统中五个系统的去匿名化。
摘要:Anonymization of voice seeks to conceal the identity of the speaker while maintaining the utility of speech data. However, residual speaker cues often persist, which pose privacy risks. We propose SegReConcat, a data augmentation method for attacker-side enhancement of automatic speaker verification systems. SegReConcat segments anonymized speech at the word level, rearranges segments using random or similarity-based strategies to disrupt long-term contextual cues, and concatenates them with the original utterance, allowing an attacker to learn source speaker traits from multiple perspectives. The proposed method has been evaluated in the VoicePrivacy Attacker Challenge 2024 framework across seven anonymization systems, SegReConcat improves de-anonymization on five out of seven systems.
【4】Cross-Learning Fine-Tuning Strategy for Dysarthric Speech Recognition Via CDSD database
标题:通过CDSD数据库进行结构障碍语音识别的交叉学习微调策略
链接:https://arxiv.org/abs/2508.18732
摘要:构音障碍语音识别面临着相对于正常语音的严重程度变化和差异的挑战。传统的方法单独地微调ASR模型,该模型是在每个患者的正常语音上预先训练的,以防止特征冲突。与直觉相反,实验表明,多说话者微调(同时对多个构音障碍的扬声器)提高了识别个人的语音模式。该策略通过更广泛的病理特征学习增强了泛化能力,减轻了特定于说话者的过度拟合,降低了对每个患者数据的依赖性,并提高了目标说话者的准确性-与单说话者微调相比,WER降低了13.15%。
摘要:Dysarthric speech recognition faces challenges from severity variations and disparities relative to normal speech. Conventional approaches individually fine-tune ASR models pre-trained on normal speech per patient to prevent feature conflicts. Counter-intuitively, experiments reveal that multi-speaker fine-tuning (simultaneously on multiple dysarthric speakers) improves recognition of individual speech patterns. This strategy enhances generalization via broader pathological feature learning, mitigates speaker-specific overfitting, reduces per-patient data dependence, and improves target-speaker accuracy - achieving up to 13.15% lower WER versus single-speaker fine-tuning.
【5】Emotion Omni: Enabling Empathetic Speech Response Generation through Large Language Models
标题:情感Omni:通过大型语言模型实现同理心语音响应生成
链接:https://arxiv.org/abs/2508.18655
备注:5 pages, 1 figure, submitted to ICASSP 2026
摘要:随着语音大语言模型(speech LLM)的发展,用户现在可以通过语音直接与助手交互。然而,大多数现有的模型只是简单地将响应内容转换为语音,而没有充分理解用户查询中嵌入的丰富的情感和语言学线索。在许多情况下,同一个句子可以有不同的含义取决于情感表达。此外,情感理解对于改善人机交互中的用户体验至关重要。目前,大多数具有移情能力的语音LLM都是在大量数据集上训练的。这种方法需要大量的数据和大量的计算资源。因此,一个关键的挑战在于如何开发一个语音LLM能够产生移情反应有限的数据,而不需要大规模的培训。为了解决这一挑战,我们提出了情感Omni,一种新的模型架构,旨在了解用户语音输入的情感内容,并生成移情语音响应。此外,我们开发了一个基于开源TTS框架的数据生成管道,以构建一个20万的情感对话数据集,该数据集支持构建一个移情语音助手。这些演示可在https://w311411.github.io/omni_demo/上获得
摘要:With the development of speech large language models (speech LLMs), users can now interact directly with assistants via speech. However, most existing models simply convert the response content into speech without fully understanding the rich emotional and paralinguistic cues embedded in the user's query. In many cases, the same sentence can have different meanings depending on the emotional expression. Furthermore, emotional understanding is essential for improving user experience in human-machine interaction. Currently, most speech LLMs with empathetic capabilities are trained on massive datasets. This approach requires vast amounts of data and significant computational resources. Therefore, a key challenge lies in how to develop a speech LLM capable of generating empathetic responses with limited data and without the need for large-scale training. To address this challenge, we propose Emotion Omni, a novel model architecture designed to understand the emotional content of user speech input and generate empathetic speech responses. Additionally, we developed a data generation pipeline based on an open-source TTS framework to construct a 200k emotional dialogue dataset, which supports the construction of an empathetic speech assistant. The demos are available at https://w311411.github.io/omni_demo/
【6】The Sound of Risk: A Multimodal Physics-Informed Acoustic Model for Forecasting Market Volatility and Enhancing Market Interpretability
标题:风险之声:用于预测市场波动性和增强市场解释性的多模式物理信息声学模型
链接:https://arxiv.org/abs/2508.18653
备注:9 pages, 6 figures
摘要:金融市场中的信息不对称,往往被精心策划的企业叙事放大,破坏了传统文本分析的有效性。我们提出了一个新的多模态框架,金融风险评估,整合文本情感与来自执行声道动态的盈利电话的非语言线索。该框架的核心是物理信息声学模型(PIAM),它应用非线性声学从受到信号削波等失真的原始电话会议声音中鲁棒地提取情感签名。听觉和文本情感状态都被投射到一个可解释的三维情感状态标签(ASL)空间-紧张,稳定和唤醒。使用1,795个财报电话(约1,800小时)的数据集,我们构建了捕捉脚本演示和自发问答交流之间的执行影响动态变化的功能。我们的主要发现揭示了预测能力的明显差异:虽然多模态特征不能预测方向性股票收益,但它们解释了30天已实现波动率中高达43.8%的样本外方差。重要的是,波动性预测强烈驱动的情绪动态执行过渡期间,从脚本到自发的讲话,特别是减少文本的稳定性和提高声学不稳定性,从首席执行官,和显着的唤醒变化。一项消融研究证实,我们的多模式方法大大优于仅财务基线,强调了声学和文本模式的互补贡献。通过从可验证的生物特征信号中解码潜在的不确定性标记,我们的方法为投资者和监管机构提供了一个强大的工具,用于增强市场的可解释性和识别隐藏的公司不确定性。
摘要:Information asymmetry in financial markets, often amplified by strategically crafted corporate narratives, undermines the effectiveness of conventional textual analysis. We propose a novel multimodal framework for financial risk assessment that integrates textual sentiment with paralinguistic cues derived from executive vocal tract dynamics in earnings calls. Central to this framework is the Physics-Informed Acoustic Model (PIAM), which applies nonlinear acoustics to robustly extract emotional signatures from raw teleconference sound subject to distortions such as signal clipping. Both acoustic and textual emotional states are projected onto an interpretable three-dimensional Affective State Label (ASL) space-Tension, Stability, and Arousal. Using a dataset of 1,795 earnings calls (approximately 1,800 hours), we construct features capturing dynamic shifts in executive affect between scripted presentation and spontaneous Q&A exchanges. Our key finding reveals a pronounced divergence in predictive capacity: while multimodal features do not forecast directional stock returns, they explain up to 43.8% of the out-of-sample variance in 30-day realized volatility. Importantly, volatility predictions are strongly driven by emotional dynamics during executive transitions from scripted to spontaneous speech, particularly reduced textual stability and heightened acoustic instability from CFOs, and significant arousal variability from CEOs. An ablation study confirms that our multimodal approach substantially outperforms a financials-only baseline, underscoring the complementary contributions of acoustic and textual modalities. By decoding latent markers of uncertainty from verifiable biometric signals, our methodology provides investors and regulators a powerful tool for enhancing market interpretability and identifying hidden corporate uncertainty.
【7】SwiftF0: Fast and Accurate Monophonic Pitch Detection
标题:SwiftF0:快速准确的单音音调检测
链接:https://arxiv.org/abs/2508.18440
摘要:在噪声条件下,特别是在资源受限的设备上,准确和实时的单声道音高估计仍然是音频处理中的一个公开挑战。我们提出了一种新的轻量级神经模型SwiftF 0,它为单声道音高估计提供了一种新的最先进的方法。通过对不同的语音、音乐和合成数据集进行训练,并进行广泛的数据增强,SwiftF 0在保持计算效率的同时,实现了跨声学域的鲁棒泛化。SwiftF 0在10 dB SNR时达到91.80\%的谐波平均值(HM),比CREPE等基线性能高出12个百分点以上,仅比干净音频降低2.3个百分点。SwiftF 0仅需要95,842个参数,在CPU上的运行速度比CREPE快约42倍,使其成为高效实时部署的理想选择。为了解决语音语料库(通常依赖于算法估计器或喉镜信号)中缺乏完全准确的地面真实音调的问题,我们引入了\{SpeechSynth}。这个由音素级TTS模型生成的合成语音数据集提供了精确的、按需的真实音调曲线,从而实现了更强大的模型训练和评估。此外,我们提出了一个统一的指标,结合六个互补的性能指标,全面和可靠的音高评估,并发布了一个开源的音高基准套件。SwiftF 0的实时演示可在https://swift-f0.github.io/获得,源代码可在https://github.com/lars76/swift-f0获得,基准框架可在https://github.com/lars76/pitch-benchmark获得。
摘要:Accurate and real-time monophonic pitch estimation in noisy conditions, particularly on resource-constrained devices, remains an open challenge in audio processing. We present \emph{SwiftF0}, a novel, lightweight neural model that sets a new state-of-the-art for monophonic pitch estimation. Through training on diverse speech, music, and synthetic datasets with extensive data augmentation, SwiftF0 achieves robust generalization across acoustic domains while maintaining computational efficiency. SwiftF0 achieves a 91.80\% harmonic mean (HM) at 10 dB SNR, outperforming baselines like CREPE by over 12 percentage points and degrading by only 2.3 points from clean audio. SwiftF0 requires only 95,842 parameters and runs approximately 42x faster than CREPE on CPU, making it ideal for efficient, real-time deployment. To address the critical lack of perfectly accurate ground truth pitch in speech corpora (which typically rely on algorithmic estimators or laryngograph signals), we introduce \emph{SpeechSynth}. This synthetic speech dataset, generated by a phoneme-level TTS model, provides exact, on-demand ground-truth pitch curves, enabling more robust model training and evaluation. Furthermore, we propose a unified metric, combining six complementary performance measures for comprehensive and reliable pitch evaluation, and release an open-source pitch benchmark suite. A live demo of SwiftF0 is available at https://swift-f0.github.io/, the source code at https://github.com/lars76/swift-f0, and the benchmark framework at https://github.com/lars76/pitch-benchmark.
【8】H-PRM: A Pluggable Hotword Pre-Retrieval Module for Various Speech Recognition Systems
标题:H-PRM:用于各种语音识别系统的可插入热词预检索模块
链接:https://arxiv.org/abs/2508.18295
摘要:热词定制在ASR中至关重要,以提高特定领域术语的准确性。它主要是由传统模型和音频大语言模型(LLM)的进步推动的。然而,现有的模型往往与大规模的热词斗争,因为识别率随着热词数量的增加而急剧下降。在本文中,我们介绍了一种新的热词定制系统,利用一个热词预检索模块(H-PRM),以确定最相关的热词候选人通过测量之间的声学相似性的热词和语音段。这种即插即用的解决方案可以很容易地集成到传统的模型中,如SeACo-Paraformer,显著提高热词后召回率(PRR)。此外,我们还通过一种基于H-PRM的方法将H-PRM整合到Audio LLM中,从而实现热词的无缝定制。大量的测试验证了H-PRM可以优于现有的方法,为ASR中的热词定制提供了新的方向。
摘要:Hotword customization is crucial in ASR to enhance the accuracy of domain-specific terms. It has been primarily driven by the advancements in traditional models and Audio large language models (LLMs). However, existing models often struggle with large-scale hotwords, as the recognition rate drops dramatically with the number of hotwords increasing. In this paper, we introduce a novel hotword customization system that utilizes a hotword pre-retrieval module (H-PRM) to identify the most relevant hotword candidate by measuring the acoustic similarity between the hotwords and the speech segment. This plug-and-play solution can be easily integrated into traditional models such as SeACo-Paraformer, significantly enhancing hotwords post-recall rate (PRR). Additionally, we incorporate H-PRM into Audio LLMs through a prompt-based approach, enabling seamless customization of hotwords. Extensive testing validates that H-PRM can outperform existing methods, showing a new direction for hotword customization in ASR.
【9】MDD: a Mask Diffusion Detector to Protect Speaker Verification Systems from Adversarial Perturbations
标题:DDD:一种保护说话人验证系统免受对抗性扰动的屏蔽扩散检测器
链接:https://arxiv.org/abs/2508.19180
备注:Accepted by APSIPA ASC 2025
摘要:说话人确认系统越来越多地部署在安全敏感的应用中,但仍然非常容易受到对抗性扰动的影响。在这项工作中,我们提出了掩模扩散检测器(MDD),一种新的对抗性检测和净化框架的基础上\textit{文本条件掩模扩散模型}。在训练过程中,MDD对Mel频谱图应用部分掩蔽,并通过前向扩散过程逐步添加噪声,模拟干净语音特征的退化。然后,逆过程以输入转录为条件重建干净的表示。与以前的方法不同,MDD不需要对抗性示例或大规模预训练。实验结果表明,MDD实现了强大的对抗性检测性能,并优于现有的最先进的方法,包括基于扩散和基于神经编解码器的方法。此外,MDD有效地净化了不利的操纵语音,恢复说话人验证性能接近干净的条件下观察到的水平。这些研究结果表明,基于扩散的掩蔽策略的安全和可靠的说话人确认系统的潜力。
摘要:Speaker verification systems are increasingly deployed in security-sensitive applications but remain highly vulnerable to adversarial perturbations. In this work, we propose the Mask Diffusion Detector (MDD), a novel adversarial detection and purification framework based on a \textit{text-conditioned masked diffusion model}. During training, MDD applies partial masking to Mel-spectrograms and progressively adds noise through a forward diffusion process, simulating the degradation of clean speech features. A reverse process then reconstructs the clean representation conditioned on the input transcription. Unlike prior approaches, MDD does not require adversarial examples or large-scale pretraining. Experimental results show that MDD achieves strong adversarial detection performance and outperforms prior state-of-the-art methods, including both diffusion-based and neural codec-based approaches. Furthermore, MDD effectively purifies adversarially-manipulated speech, restoring speaker verification performance to levels close to those observed under clean conditions. These findings demonstrate the potential of diffusion-based masking strategies for secure and reliable speaker verification systems.
【1】Interpolating Speaker Identities in Embedding Space for Data Expansion
标题:在嵌入空间中插入说话者身份以进行数据扩展
链接:https://arxiv.org/abs/2508.19210
备注:accepted by APSIPA ASC 2025
摘要:基于深度学习的说话人验证系统的成功在很大程度上归功于对大规模和多样化说话人身份数据的访问。然而,从更多的身份收集数据是昂贵的,具有挑战性的,并且往往受到隐私问题的限制。为了解决这一限制,我们提出了内插扬声器身份嵌入空间(INSIDE),一种新的数据扩展方法,通过在现有的扬声器嵌入之间进行内插来合成新的扬声器身份。具体来说,我们从预训练的说话者嵌入空间中选择附近的说话者嵌入对,并使用球面线性插值计算中间嵌入。这些内插嵌入然后被馈送到文本到语音系统以生成相应的语音波形。将得到的数据与原始数据集组合以训练下游模型。实验结果表明,使用INSIDE扩展数据训练的模型性能优于仅使用真实数据训练的模型,相对性能提高了3.06 ~ 5.24%.虽然INSIDE主要是为说话人确认而设计的,但我们也验证了其在性别分类方面的有效性,相对提高了13.44%。此外,INSIDE与其他增强技术兼容,可以作为现有培训管道的灵活,可扩展的补充。
摘要:The success of deep learning-based speaker verification systems is largely attributed to access to large-scale and diverse speaker identity data. However, collecting data from more identities is expensive, challenging, and often limited by privacy concerns. To address this limitation, we propose INSIDE (Interpolating Speaker Identities in Embedding Space), a novel data expansion method that synthesizes new speaker identities by interpolating between existing speaker embeddings. Specifically, we select pairs of nearby speaker embeddings from a pretrained speaker embedding space and compute intermediate embeddings using spherical linear interpolation. These interpolated embeddings are then fed to a text-to-speech system to generate corresponding speech waveforms. The resulting data is combined with the original dataset to train downstream models. Experiments show that models trained with INSIDE-expanded data outperform those trained only on real data, achieving 3.06\% to 5.24\% relative improvements. While INSIDE is primarily designed for speaker verification, we also validate its effectiveness on gender classification, where it yields a 13.44\% relative improvement. Moreover, INSIDE is compatible with other augmentation techniques and can serve as a flexible, scalable addition to existing training pipelines.
【2】MDD: a Mask Diffusion Detector to Protect Speaker Verification Systems from Adversarial Perturbations
标题:DDD:一种保护说话人验证系统免受对抗性扰动的屏蔽扩散检测器
链接:https://arxiv.org/abs/2508.19180
备注:Accepted by APSIPA ASC 2025
摘要:说话人确认系统越来越多地部署在安全敏感的应用中,但仍然非常容易受到对抗性扰动的影响。在这项工作中,我们提出了掩模扩散检测器(MDD),一种新的对抗性检测和净化框架的基础上\textit{文本条件掩模扩散模型}。在训练过程中,MDD对Mel频谱图应用部分掩蔽,并通过前向扩散过程逐步添加噪声,模拟干净语音特征的退化。然后,逆过程以输入转录为条件重建干净的表示。与以前的方法不同,MDD不需要对抗性示例或大规模预训练。实验结果表明,MDD实现了强大的对抗性检测性能,并优于现有的最先进的方法,包括基于扩散和基于神经编解码器的方法。此外,MDD有效地净化了不利的操纵语音,恢复说话人验证性能接近干净的条件下观察到的水平。这些研究结果表明,基于扩散的掩蔽策略的安全和可靠的说话人确认系统的潜力。
摘要:Speaker verification systems are increasingly deployed in security-sensitive applications but remain highly vulnerable to adversarial perturbations. In this work, we propose the Mask Diffusion Detector (MDD), a novel adversarial detection and purification framework based on a \textit{text-conditioned masked diffusion model}. During training, MDD applies partial masking to Mel-spectrograms and progressively adds noise through a forward diffusion process, simulating the degradation of clean speech features. A reverse process then reconstructs the clean representation conditioned on the input transcription. Unlike prior approaches, MDD does not require adversarial examples or large-scale pretraining. Experimental results show that MDD achieves strong adversarial detection performance and outperforms prior state-of-the-art methods, including both diffusion-based and neural codec-based approaches. Furthermore, MDD effectively purifies adversarially-manipulated speech, restoring speaker verification performance to levels close to those observed under clean conditions. These findings demonstrate the potential of diffusion-based masking strategies for secure and reliable speaker verification systems.
【3】CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
标题:CLEAR:用于高质量和低延迟语音合成的连续潜在自回归模型
链接:https://arxiv.org/abs/2508.19098
备注:Preprint
摘要:自回归(AR)语言模型已经成为zero-shot文本到语音(TTS)合成的强大解决方案,能够从几秒钟的音频提示生成自然语音。然而,传统的基于AR的TTS系统依赖于离散的音频令牌面临着令牌化过程中有损压缩的挑战,需要更长的离散令牌序列来捕获与连续令牌序列相同的信息,这增加了推理延迟并使AR建模复杂化。为了解决这个问题,本文提出了连续潜在自回归模型(CLEAR),一个统一的zero-shot TTS框架,直接模拟连续的音频表示。更具体地说,CLEAR引入了一种具有快捷连接的增强型变分自动编码器,该编码器实现了高压缩比,可将波形映射为紧凑的连续潜伏期。一个轻量级的MLP为基础的整流头,独立运作的每个隐藏状态的连续潜在概率分布建模,并在一个单阶段的框架内与AR模型共同训练。实验结果表明,提出的zero-shot CLEAR文语转换系统能够以较低的时延合成出高质量的语音。与最先进的(SOTA)TTS模型相比,CLEAR在鲁棒性,说话人相似性和自然度方面具有竞争力的性能,同时提供较低的实时因素(RTF)。特别是,CLEAR在LibriSpeech测试干净数据集上实现了SOTA结果,单词错误率为1.88%,RTF为0.29。此外,CLEAR有助于以96 ms的第一帧延迟进行流式语音合成,同时保持高质量的语音合成。
摘要:Autoregressive (AR) language models have emerged as powerful solutions for zero-shot text-to-speech (TTS) synthesis, capable of generating natural speech from a few seconds of audio prompts. However, conventional AR-based TTS systems relying on discrete audio tokens face the challenge of lossy compression during tokenization, requiring longer discrete token sequences to capture the same information as continuous ones, which adds inference latency and complicates AR modeling. To address this challenge, this paper proposes the Continuous Latent Autoregressive model (CLEAR), a unified zero-shot TTS framework that directly models continuous audio representations. More specifically, CLEAR introduces an enhanced variational autoencoder with shortcut connections, which achieves a high compression ratio to map waveforms into compact continuous latents. A lightweight MLP-based rectified flow head that operates independently for each hidden state is presented to model the continuous latent probability distribution, and trained jointly with the AR model within a single-stage framework. Experiments show that the proposed zero-shot CLEAR TTS can synthesize high-quality speech with low latency. Compared to state-of-the-art (SOTA) TTS models, CLEAR delivers competitive performance in robustness, speaker similarity and naturalness, while offering a lower real-time factor (RTF). In particular, CLEAR achieves SOTA results on the LibriSpeech test-clean dataset, with a word error rate of 1.88\% and an RTF of 0.29. Moreover, CLEAR facilitates streaming speech synthesis with a first-frame delay of 96ms, while maintaining high-quality speech synthesis.
【4】MOSA: Mixtures of Simple Adapters Outperform Monolithic Approaches in LLM-based Multilingual ASR
标题:MOSA:在基于LLM的多语言ASB中,简单适配器的混合优于整体方法
链接:https://arxiv.org/abs/2508.18998
摘要:端到端的多语言ASR旨在将不同语言的语音转录为相应的文本,但往往受到稀缺的多语言数据的限制。基于LLM的ASR通过投影仪将语音编码器输出与LLM输入空间对齐,并取得了显着的成功。然而,之前的工作主要通过增加数据来提高性能,很少关注跨语言知识共享。此外,单个复杂的投影仪难以有效地捕获共享和特定语言的功能。在这项工作中,我们提出了MOSA(简单适配器混合),利用专家混合机制来组合学习共享和特定语言知识的轻量级适配器。这可以更好地利用高资源语言数据来支持低资源语言,缓解数据稀缺问题。实验结果表明,MOSA-Base实现了15.4%的平均WER相对减少相比,基线基础,并始终优于它在所有语言。值得注意的是,MOSA-Base甚至在仅使用其参数的60%进行训练时也优于Baseline-Base。同样,MOSA-Large在平均WER方面优于Baseline-Large,并对数据不平衡表现出更强的鲁棒性。消融研究进一步表明,MOSA在处理个人语言和学习语言特定和共享的语言知识方面更有效。这些发现支持,在基于LLM的ASR中,简单适配器的混合比单一复杂的适配器设计更有效。
摘要:End-to-end multilingual ASR aims to transcribe speech from different languages into corresponding text, but is often limited by scarce multilingual data. LLM-based ASR aligns speech encoder outputs with LLM input space via a projector and has achieved notable success. However, prior work mainly improves performance by increasing data, with little focus on cross-lingual knowledge sharing. Moreover, a single complex projector struggles to capture both shared and language-specific features effectively. In this work, we propose MOSA (Mixture of Simple Adapters), leveraging a Mixture-of-Experts mechanism to combine lightweight adapters that learn shared and language-specific knowledge. This enables better utilization of high-resource language data to support low-resource languages, mitigating data scarcity issues. Experimental results show that MOSA-Base achieves a 15.4\% relative reduction in average WER compared to the Baseline-Base and consistently outperforms it across all languages. Remarkably, MOSA-Base surpasses the Baseline-Base even when trained with only 60\% of its parameters. Similarly, MOSA-Large outperforms the Baseline-Large in average WER and demonstrates greater robustness to data imbalance. Ablation studies further indicate that MOSA is more effective at handling individual languages and learning both language-specific and shared linguistic knowledge. These findings support that, in LLM-based ASR, a mixture of simple adapters is more effective than a single, complex adapter design.
【5】A Framework for Robust Speaker Verification in Highly Noisy Environments Leveraging Both Noisy and Enhanced Audio
标题:利用噪音和增强音频的高噪音环境中稳健的说话者验证框架
链接:https://arxiv.org/abs/2508.18913
备注:5 pages, 2 figures, 1 table. Submitted to EUSIPCO 2025. Keywords: speaker verification, speaker recognition, speaker embedding, speech enhancement, ECAPA-TDNN, SpeakerNet, x-vectors, noisy speech, robust embeddings
摘要:说话人确认技术的最新进展显示出了希望,但它们的性能往往在具有挑战性的声学环境中显着恶化。虽然语音增强方法可以提高感知的音频质量,但它们可能会无意中扭曲说话者特定的信息,这可能会影响验证的准确性。随着越来越多地使用生成深度神经网络(DNN)进行语音增强,这个问题变得更加明显。虽然这些网络即使在非常低的信噪比(SNR)的条件下也可以产生可理解的语音,但它们也可能严重改变独特的说话者特征。为了解决这个问题,我们提出了一种新的神经网络框架,有效地结合了扬声器嵌入提取的噪声和增强语音使用暹罗架构。这种架构使我们能够利用来自两个来源的互补信息,增强了在严重噪声条件下说话人验证的鲁棒性。我们的框架是轻量级的,对特定的说话人验证和语音增强技术是不可知的,使使用范围广泛的最先进的解决方案,而无需修改。实验结果表明,我们提出的框架的优越性能。
摘要:Recent advancements in speaker verification techniques show promise, but their performance often deteriorates significantly in challenging acoustic environments. Although speech enhancement methods can improve perceived audio quality, they may unintentionally distort speaker-specific information, which can affect verification accuracy. This problem has become more noticeable with the increasing use of generative deep neural networks (DNNs) for speech enhancement. While these networks can produce intelligible speech even in conditions of very low signal-to-noise ratio (SNR), they may also severely alter distinctive speaker characteristics. To tackle this issue, we propose a novel neural network framework that effectively combines speaker embeddings extracted from both noisy and enhanced speech using a Siamese architecture. This architecture allows us to leverage complementary information from both sources, enhancing the robustness of speaker verification under severe noise conditions. Our framework is lightweight and agnostic to specific speaker verification and speech enhancement techniques, enabling the use of a wide range of state-of-the-art solutions without modification. Experimental results demonstrate the superior performance of our proposed framework.
【6】On the Application of Diffusion Models for Simultaneous Denoising and Dereverberation
标题:扩散模型在同时去噪和去回响中的应用
链接:https://arxiv.org/abs/2508.18833
备注:Accepted at 16th ITG Conference on Speech Communication 2025
摘要:扩散模型已被证明可以实现自然的声音增强的语音退化的噪声或混响。然而,到目前为止,它们的同时去噪和去混响能力还没有得到太多的研究,尽管这可以说是实际应用中最常见的情况。在这项工作中,我们研究不同的方法来增强嘈杂和/或混响的语音。我们研究模型的级联应用,每个模型只在一个失真上训练,并将其与单个模型进行比较,该模型仅在噪声和混响的数据上训练,或者在包括纯噪声,纯混响和噪声混响语音的子集的数据上训练。测试是在人工生成的和真实的噪音和/或混响数据记录上进行的。结果表明,当使用级联模型时,只有按主导畸变的顺序应用才能获得满意的结果。如果只需要一个可以在所有失真场景下运行的模型,那么最好的折衷方案似乎是在上述三个降级语音数据子集上训练的模型。
摘要:Diffusion models have been shown to achieve natural-sounding enhancement of speech degraded by noise or reverberation. However, their simultaneous denoising and dereverberation capability has so far not been studied much, although this is arguably the most common scenario in a practical application. In this work, we investigate different approaches to enhance noisy and/or reverberant speech. We examine the cascaded application of models, each trained on only one of the distortions, and compare it with a single model, trained either solely on data that is both noisy and reverberated, or trained on data comprising subsets of purely noisy, of purely reverberated, and of noisy reverberant speech. Tests are performed both on artificially generated and real recordings of noisy and/or reverberant data. The results show that, when using the cascade of models, satisfactory results are only achieved if they are applied in the order of the dominating distortion. If only a single model is desired that can operate on all distortion scenarios, the best compromise appears to be a model trained on the aforementioned three subsets of degraded speech data.
【7】EAI-Avatar: Emotion-Aware Interactive Talking Head Generation
标题:EAI-Avatar:具有感知力的互动会说话的头一代
链接:https://arxiv.org/abs/2508.18337
摘要:生成模型发展迅速,能够生成令人印象深刻的会说话的头部,将AI带入生活。然而,大多数现有的方法只关注单向肖像动画。即使是少数支持双向会话交互的设备也缺乏精确的情感自适应能力,这大大限制了它们的实用性。在本文中,我们提出了EAI-Avatar,一种新的情感感知的二元交互的说话头生成框架。利用大型语言模型(LLM,例如,GPT-4),我们的方法产生时间上一致的虚拟化身,具有丰富的情感变化,在说话和倾听状态之间无缝过渡。具体来说,我们设计了一个基于变换器的头部掩模生成器,学习时间上一致的运动特征在一个潜在的掩模空间,能够生成任意长度,时间上一致的掩模序列来约束头部运动。此外,我们引入了一个交互式的谈话树结构来表示对话状态转换,其中每个树节点包含的信息,如孩子/父母/兄弟姐妹节点和当前字符的情绪状态。通过执行反向遍历,我们从当前节点提取丰富的历史情感线索,以指导表情合成。大量的实验证明了我们的方法的优越性能和有效性。
摘要:Generative models have advanced rapidly, enabling impressive talking head generation that brings AI to life. However, most existing methods focus solely on one-way portrait animation. Even the few that support bidirectional conversational interactions lack precise emotion-adaptive capabilities, significantly limiting their practical applicability. In this paper, we propose EAI-Avatar, a novel emotion-aware talking head generation framework for dyadic interactions. Leveraging the dialogue generation capability of large language models (LLMs, e.g., GPT-4), our method produces temporally consistent virtual avatars with rich emotional variations that seamlessly transition between speaking and listening states. Specifically, we design a Transformer-based head mask generator that learns temporally consistent motion features in a latent mask space, capable of generating arbitrary-length, temporally consistent mask sequences to constrain head motions. Furthermore, we introduce an interactive talking tree structure to represent dialogue state transitions, where each tree node contains information such as child/parent/sibling nodes and the current character's emotional state. By performing reverse-level traversal, we extract rich historical emotional cues from the current node to guide expression synthesis. Extensive experiments demonstrate the superior performance and effectiveness of our method.
【8】Toward Responsible ASR for African American English Speakers: A Scoping Review of Bias and Equity in Speech Technology
标题:迈向非裔美国英语使用者负责任的ASB:言语技术偏见和公平的范围审查
链接:https://arxiv.org/abs/2508.18288
备注:10 pages, 9 Pages (References and Appendices). The archival version has been accepted to AAAI (AIES 2025) without the extended Appendices. This extended version includes Appendices
摘要:这一范围文献综述探讨了公平,偏见和公平是如何在自动语音识别(ASR)和邻近的语音和语言技术(ESTA)的非裔美国人英语(AAE)的扬声器和其他语言多样化的社区概念化和操作。从人机交互(HCI),机器学习/自然语言处理(ML/NLP)和社会语言学的44篇同行评议的出版物中,我们确定了四个主要的研究领域:(1)研究人员如何理解与ASR相关的危害;(2)涵盖收集,策展,注释和模型训练的包容性数据实践;(3)语言包容的方法和理论方法;(4)研究人员如何理解与ASR相关的危害。以及(4)新兴实践和设计更公平系统的建议。虽然技术公平的干预措施越来越多,我们的审查突出了一个关键的差距,以治理为中心的方法,前景社区机构,语言正义和参与式问责制。我们提出了一个以治理为中心的ASR生命周期,作为负责任的ASR开发的新兴跨学科框架,并为寻求解决语音AI系统中语言边缘化问题的研究人员、从业人员和政策制定者提供了启示。
摘要:This scoping literature review examines how fairness, bias, and equity are conceptualized and operationalized in Automatic Speech Recognition (ASR) and adjacent speech and language technologies (SLT) for African American English (AAE) speakers and other linguistically diverse communities. Drawing from 44 peer-reviewed publications across Human-Computer Interaction (HCI), Machine Learning/Natural Language Processing (ML/NLP), and Sociolinguistics, we identify four major areas of inquiry: (1) how researchers understand ASR-related harms; (2) inclusive data practices spanning collection, curation, annotation, and model training; (3) methodological and theoretical approaches to linguistic inclusion; and (4) emerging practices and design recommendations for more equitable systems. While technical fairness interventions are growing, our review highlights a critical gap in governance-centered approaches that foreground community agency, linguistic justice, and participatory accountability. We propose a governance-centered ASR lifecycle as an emergent interdisciplinary framework for responsible ASR development and offer implications for researchers, practitioners, and policymakers seeking to address language marginalization in speech AI systems.
【9】Improving Noise Robust Audio-Visual Speech Recognition via Router-Gated Cross-Modal Feature Fusion
标题:通过路由器选通跨模式特征融合提高噪音鲁棒的视听语音识别
链接:https://arxiv.org/abs/2508.18734
备注:Accepted to IEEE ASRU 2025
摘要:噪声环境中的鲁棒视听语音识别(AVSR)仍然具有挑战性,因为现有系统难以估计音频可靠性并动态调整模态依赖性。我们提出了路由器门控跨模态特征融合,一种新的AVSR框架,自适应地重新加权音频和视觉功能的基础上令牌级的声学腐败分数。使用基于视听特征融合的路由器,我们的方法降低了不可靠的音频令牌的权重,并通过每个解码器层中的门控交叉注意来加强视觉线索。这使得模型能够在音频质量恶化时转向视觉模态。在LRS 3上的实验表明,与AV-HuBERT相比,我们的方法实现了16.51-42.67%的相对减少字错误率。消融研究证实,路由器和门控机制都有助于提高真实世界噪声下的鲁棒性。
摘要:Robust audio-visual speech recognition (AVSR) in noisy environments remains challenging, as existing systems struggle to estimate audio reliability and dynamically adjust modality reliance. We propose router-gated cross-modal feature fusion, a novel AVSR framework that adaptively reweights audio and visual features based on token-level acoustic corruption scores. Using an audio-visual feature fusion-based router, our method down-weights unreliable audio tokens and reinforces visual cues through gated cross-attention in each decoder layer. This enables the model to pivot toward the visual modality when audio quality deteriorates. Experiments on LRS3 demonstrate that our approach achieves an 16.51-42.67% relative reduction in word error rate compared to AV-HuBERT. Ablation studies confirm that both the router and gating mechanism contribute to improved robustness under real-world acoustic noise.
机器翻译由腾讯交互翻译提供,仅供参考
