微信公众号:arXiv_Daily
cs.SD语音
【1】Advances in Speech Separation: Techniques, Challenges, and Future Trends
标题:语音分离的进展:技术、挑战和未来趋势
链接:https://arxiv.org/abs/2508.10830
备注:34 pages, 10 figures
摘要:语音分离领域,解决“鸡尾酒会问题”,已经看到了DNN的革命性进展。语音分离是语音识别和说话人识别的关键预处理技术,它能提高复杂声学环境下的语音清晰度。然而,目前的文献狭隘地关注于特定的架构或孤立的方法,造成了零散的理解。本调查通过对基于DNN的语音分离技术进行系统检查来解决这一差距。我们的工作通过以下方面区分开来:(I)全面的视角:我们系统地研究了学习范式,已知/未知扬声器的分离场景,监督/自监督/无监督框架的比较分析,以及从编码器到估计策略的架构组件。(II)及时性:对前沿发展的报道确保获得当前的创新和基准。(III)独特的见解:除了总结,我们评估技术轨迹,识别新兴模式,并强调有前途的方向,包括领域强大的框架,高效的架构,多模式集成和新颖的自我监督范式。(IV)公平评估:我们对标准数据集进行定量评估,揭示不同方法的真实能力和局限性。这项全面的调查为经验丰富的研究人员和新来者导航语音分离的复杂景观提供了一个可访问的参考。
摘要:The field of speech separation, addressing the "cocktail party problem", has seen revolutionary advances with DNNs. Speech separation enhances clarity in complex acoustic environments and serves as crucial pre-processing for speech recognition and speaker recognition. However, current literature focuses narrowly on specific architectures or isolated approaches, creating fragmented understanding. This survey addresses this gap by providing systematic examination of DNN-based speech separation techniques. Our work differentiates itself through: (I) Comprehensive perspective: We systematically investigate learning paradigms, separation scenarios with known/unknown speakers, comparative analysis of supervised/self-supervised/unsupervised frameworks, and architectural components from encoders to estimation strategies. (II) Timeliness: Coverage of cutting-edge developments ensures access to current innovations and benchmarks. (III) Unique insights: Beyond summarization, we evaluate technological trajectories, identify emerging patterns, and highlight promising directions including domain-robust frameworks, efficient architectures, multimodal integration, and novel self-supervised paradigms. (IV) Fair evaluation: We provide quantitative evaluations on standard datasets, revealing true capabilities and limitations of different methods. This comprehensive survey serves as an accessible reference for experienced researchers and newcomers navigating speech separation's complex landscape.
【2】Ensembling Synchronisation-based and Face-Voice Association Paradigms for Robust Active Speaker Detection in Egocentric Recordings
标题:集成基于同步和面部语音关联的范式,用于自我中心录音中的鲁棒性活动说话人检测
链接:https://arxiv.org/abs/2508.10580
备注:Accepted to SPECOM 2025, 13 pages, 4 figures. To appear in the Proceedings of the 27th International Conference on Speech and Computer (SPECOM) 2025, October 13-14, 2025, Szeged, Hungary
摘要:在以自我为中心的录音中,视听主动说话人检测(ASD)受到频繁的遮挡、运动模糊和音频干扰的挑战,这些干扰破坏了嘴唇运动和语音之间的时间同步的可辨别性。传统的基于同步的系统在干净的条件下表现良好,但在第一人称录音中急剧下降。相反,基于人脸-语音关联(FVA)的方法放弃同步建模,支持跨模态生物特征匹配,表现出对瞬时视觉损坏的鲁棒性,但在出现重叠语音或前端分割错误时会受到影响。在本文中,提出了一种简单而有效的集成方法,通过加权平均融合同步相关和同步不可知的模型输出,从而利用互补的线索,而不引入复杂的融合架构。一个完善的预处理流水线的FVA为基础的组件也被引入到优化集成。在Ego 4D-AVD验证集上的实验表明,使用TalkNet和Light-ASD主干的集成分别达到70.2%和66.7%的平均精度(mAP)。一个定性分析分层的人脸图像质量和话语掩蔽患病率进一步证实了每个组件的互补优势。
摘要:Audiovisual active speaker detection (ASD) in egocentric recordings is challenged by frequent occlusions, motion blur, and audio interference, which undermine the discernability of temporal synchrony between lip movement and speech. Traditional synchronisation-based systems perform well under clean conditions but degrade sharply in first-person recordings. Conversely, face-voice association (FVA)-based methods forgo synchronisation modelling in favour of cross-modal biometric matching, exhibiting robustness to transient visual corruption but suffering when overlapping speech or front-end segmentation errors occur. In this paper, a simple yet effective ensemble approach is proposed to fuse synchronisation-dependent and synchronisation-agnostic model outputs via weighted averaging, thereby harnessing complementary cues without introducing complex fusion architectures. A refined preprocessing pipeline for the FVA-based component is also introduced to optimise ensemble integration. Experiments on the Ego4D-AVD validation set demonstrate that the ensemble attains 70.2% and 66.7% mean Average Precision (mAP) with TalkNet and Light-ASD backbones, respectively. A qualitative analysis stratified by face image quality and utterance masking prevalence further substantiates the complementary strengths of each component.
【3】Fake Speech Wild: Detecting Deepfake Speech on Social Media Platform
标题:Fake Speech Wild:在社交媒体平台上检测Deepfake Speech
链接:https://arxiv.org/abs/2508.10559
摘要:语音生成技术的快速发展导致deepfake语音在社交媒体平台上广泛传播。虽然Deepfake音频对策(CM)在公共数据集上取得了令人鼓舞的结果,但它们的性能在跨域场景中会显着下降。为了推进用于真实世界deepfake检测的CM,我们首先提出了Fake Speech Wild(FSW)数据集,其中包括来自四个不同媒体平台的254小时真实和deepfake音频,重点关注社交媒体。作为CM,我们使用公共数据集和基于高级自监督学习(SSL)的CM建立了一个基准,以评估现实场景中的当前CM。我们还评估了数据增强策略在增强CM鲁棒性以检测社交媒体上的深度虚假语音方面的有效性。最后,通过扩展公共数据集并结合FSW训练集,我们显著提高了真实世界的deepfake音频检测性能,在所有评估集上实现了3.54%的平均等错误率(EER)。
摘要:The rapid advancement of speech generation technology has led to the widespread proliferation of deepfake speech across social media platforms. While deepfake audio countermeasures (CMs) achieve promising results on public datasets, their performance degrades significantly in cross-domain scenarios. To advance CMs for real-world deepfake detection, we first propose the Fake Speech Wild (FSW) dataset, which includes 254 hours of real and deepfake audio from four different media platforms, focusing on social media. As CMs, we establish a benchmark using public datasets and advanced selfsupervised learning (SSL)-based CMs to evaluate current CMs in real-world scenarios. We also assess the effectiveness of data augmentation strategies in enhancing CM robustness for detecting deepfake speech on social media. Finally, by augmenting public datasets and incorporating the FSW training set, we significantly advanced real-world deepfake audio detection performance, achieving an average equal error rate (EER) of 3.54% across all evaluation sets.
【4】Motive-level Analysis of Form-functions Association in Korean Folk song
标题:韩国民歌形式功能联想的动机层面分析
链接:https://arxiv.org/abs/2508.10472
备注:None
摘要:民歌音频的计算分析是具有挑战性的,由于结构的不规则性和手工注释的需要。我们提出了一种方法,自动动机分割在韩国民歌微调语音转录模型的音频歌词与模体边界注释。应用这856首歌曲,我们提取主题计数和持续时间熵的结构特征。统计分析表明,这些特征根据歌曲的社会功能而系统地变化。例如,与集体劳动有关的歌曲表现出与娱乐或个人背景不同的结构模式。这项工作提供了一个可扩展的方法,口头音乐传统的定量结构分析。
摘要:Computational analysis of folk song audio is challenging due to structural irregularities and the need for manual annotation. We propose a method for automatic motive segmentation in Korean folk songs by fine-tuning a speech transcription model on audio lyric with motif boundary annotation. Applying this to 856 songs, we extracted motif count and duration entropy as structural features. Statistical analysis revealed that these features vary systematically according to the social function of the songs. Songs associated with collective labor, for instance, showed different structural patterns from those for entertainment or personal settings. This work offers a scalable approach for quantitative structural analysis of oral music traditions.
【5】Alternating Approach-Putt Models for Multi-Stage Speech Enhancement
标题:交替逼近-Putt模型的多级语音增强
链接:https://arxiv.org/abs/2508.10436
备注:This work has been submitted to the IEEE for possible publication
摘要:使用人工神经网络的语音增强旨在从含噪语音信号中去除噪声,同时保留语音内容。然而,语音增强网络经常向语音信号引入失真,称为伪像,这会降低音频质量。在这项工作中,我们提出了一个后处理神经网络,旨在减轻语音增强模型引入的文物。灵感来自于在高尔夫中的“接近”之后进行“推杆”的类比,我们将我们的模型命名为PuttNet。我们证明,交替之间的语音增强模型和Putt模型,导致改善语音质量,感知质量分数(PESQ),客观清晰度(STOI),和背景噪声侵入(CBAK)分数。此外,我们用图形分析说明了为什么这种交替的方法优于单独使用任何一种模型的重复应用。
摘要:Speech enhancement using artificial neural networks aims to remove noise from noisy speech signals while preserving the speech content. However, speech enhancement networks often introduce distortions to the speech signal, referred to as artifacts, which can degrade audio quality. In this work, we propose a post-processing neural network designed to mitigate artifacts introduced by speech enhancement models. Inspired by the analogy of making a `Putt' after an `Approach' in golf, we name our model PuttNet. We demonstrate that alternating between a speech enhancement model and the proposed Putt model leads to improved speech quality, as measured by perceptual quality scores (PESQ), objective intelligibility (STOI), and background noise intrusiveness (CBAK) scores. Furthermore, we illustrate with graphical analysis why this alternating Approach outperforms repeated application of either model alone.
【6】MCP2OSC: Parametric Control by Natural Language
标题:MPP 2OSC:自然语言参数控制
链接:https://arxiv.org/abs/2508.10414
摘要:文本提示可以实现直观的内容创建,但可能无法实现复杂任务的高精度;旋钮或滑块控件提供精确的调整,但代价是增加了复杂性。为了解决旋钮和提示之间的差距,一个新的MCP(模型上下文协议)服务器和一组独特的提示设计标准,使探索参数OSC(OpenSoundControl)控制的自然语言提示。通过14个具有最佳实践和通用提示模板的实际QA示例,本研究发现Claude与MCP 2 OSC服务器集成,有效地通过自然语言生成OSC消息,解释,搜索和可视化OSC消息,验证和调试OSC消息,以及管理OSC地址模式。MCP 2 OSC通过利用LLM(大型语言模型)来处理复杂的OSC开发任务,并通过具有灵活精度控制的直观语言界面来增强人类的创造力,从而增强了人机协作:一个基于JavaScript的OSC工具。这项研究提供了一个新的视角,创造性的MCP应用程序在网络协议层面上,利用LLM的实力,直接处理和生成人类可读的OSC消息。结果表明,它的潜在的基于LLM的多媒体设备的通用控制机制。
摘要:Text prompts enable intuitive content creation but may fall short in achieving high precision for intricate tasks; knob or slider controls offer precise adjustments at the cost of increased complexity. To address the gap between knobs and prompts, a new MCP (Model Context Protocol) server and a unique set of prompt design criteria are presented to enable exploring parametric OSC (OpenSoundControl) control by natural language prompts. Demonstrated by 14 practical QA examples with best practices and the generalized prompt templates, this study finds Claude integrated with the MCP2OSC server effective in generating OSC messages by natural language, interpreting, searching, and visualizing OSC messages, validating and debugging OSC messages, and managing OSC address patterns. MCP2OSC enhances human-machine collaboration by leveraging LLM (Large Language Model) to handle intricate OSC development tasks, and by empowering human creativity with an intuitive language interface featuring flexible precision controls: a prompt-based OSC tool. This study provides a novel perspective on the creative MCP application at the network protocol level by utilizing LLM's strength in directly processing and generating human-readable OSC messages. The results suggest its potential for a LLM-based universal control mechanism for multimedia devices.
【7】Facilitating Personalized TTS for Dysarthric Speakers Using Knowledge Anchoring and Curriculum Learning
标题:利用知识托管和课程学习促进发音障碍者的个性化TTC
链接:https://arxiv.org/abs/2508.10412
备注:Interspeech 2025
摘要:由于言语器官的运动控制受损,导致言语清晰度降低,因此关节功能障碍的说话者经历了实质性的沟通挑战。这在数据集管理中造成了重大障碍,因为为了训练个性化TTS模型而实际记录长的、清晰的句子变得不可行。因此,除了音频内存在的发音错误之外,音频数据的有限可用性使用于目标构音障碍说话者适应的个性化语音合成复杂化。为了解决这个问题,我们框架的问题作为一个域转移任务,并引入了一个知识锚定框架,利用师生模型,通过音频增强课程学习增强。实验结果表明,所提出的zero-shot多说话人TTS模型在保持韵律自然度的同时,能够有效地生成清晰度误差明显减小、说话人保真度较高的合成语音。
摘要:Dysarthric speakers experience substantial communication challenges due to impaired motor control of the speech apparatus, which leads to reduced speech intelligibility. This creates significant obstacles in dataset curation since actual recording of long, articulate sentences for the objective of training personalized TTS models becomes infeasible. Thus, the limited availability of audio data, in addition to the articulation errors that are present within the audio, complicates personalized speech synthesis for target dysarthric speaker adaptation. To address this, we frame the issue as a domain transfer task and introduce a knowledge anchoring framework that leverages a teacher-student model, enhanced by curriculum learning through audio augmentation. Experimental results show that the proposed zero-shot multi-speaker TTS model effectively generates synthetic speech with markedly reduced articulation errors and high speaker fidelity, while maintaining prosodic naturalness.
【8】A dataset and model for recognition of audiologically relevant environments for hearing aids: AHEAD-DS and YAMNet+
标题:识别助听器听力相关环境的数据集和模型:AHEAD-DS和YAMNet+
链接:https://arxiv.org/abs/2508.10360
摘要:听觉相关环境的场景识别对于助听器很重要;然而,它具有挑战性,部分原因是现有数据集的局限性。数据集通常缺乏公共可访问性,完整性或听觉相关标签,阻碍了机器学习模型的系统比较。在资源受限的边缘设备上部署这些模型是另一个挑战。我们的解决方案是双重的:我们利用几个开源数据集创建AHEAD-DS,一个为听觉相关环境的场景识别设计的数据集,并引入YAMNet+,一个声音识别模型。AHEAD-DS旨在提供一个标准化的、公开可用的数据集,具有与助听器相关的一致标签,便于模型比较。YAMNet+旨在部署在边缘设备上,如连接到听力设备的智能手机,如助听器和具有助听器功能的无线耳机;作为基于声音的场景识别的基线模型。YAMNet+在14类听觉相关环境中的AHEAD-DS测试集上实现了平均0.83的平均精度和0.93的准确度。我们发现,从预训练的YAMNet模型中应用迁移学习是必不可少的。我们通过将YAMNet+部署到Android智能手机上,在边缘设备上展示了基于声音的实时场景识别功能。即使使用Google Pixel 3(2018年发布的一款规格适中的手机),该模型处理音频的延迟也约为50 ms以加载模型,并且每1秒音频近似线性增加30 ms。我们的网站和代码www.example.com。
摘要:Scene recognition of audiologically relevant environments is important for hearing aids; however, it is challenging, in part because of the limitations of existing datasets. Datasets often lack public accessibility, completeness, or audiologically relevant labels, hindering systematic comparison of machine learning models. Deploying these models on resource-constrained edge devices presents another challenge. Our solution is two-fold: we leverage several open source datasets to create AHEAD-DS, a dataset designed for scene recognition of audiologically relevant environments, and introduce YAMNet+, a sound recognition model. AHEAD-DS aims to provide a standardised, publicly available dataset with consistent labels relevant to hearing aids, facilitating model comparison. YAMNet+ is designed for deployment on edge devices like smartphones connected to hearing devices, such as hearing aids and wireless earphones with hearing aid functionality; serving as a baseline model for sound-based scene recognition. YAMNet+ achieved a mean average precision of 0.83 and accuracy of 0.93 on the testing set of AHEAD-DS across fourteen categories of audiologically relevant environments. We found that applying transfer learning from the pretrained YAMNet model was essential. We demonstrated real-time sound-based scene recognition capabilities on edge devices by deploying YAMNet+ to an Android smartphone. Even with a Google Pixel 3 (a phone with modest specifications, released in 2018), the model processes audio with approximately 50ms of latency to load the model, and an approximate linear increase of 30ms per 1 second of audio. Our website and code https://github.com/Australian-Future-Hearing-Initiative .
【9】No Free Lunch from Audio Pretraining in Bioacoustics: A Benchmark Study of Embeddings
标题:生物声学音频预训练没有免费午餐:嵌入式基准研究
链接:https://arxiv.org/abs/2508.10230
摘要:生物声学是对动物声音的研究,为监测生态系统提供了一种非侵入性的方法。从音频预训练的深度学习(DL)模型中提取嵌入而无需微调已经成为获取任务生物声学特征的流行方法。然而,最近的一项基准研究表明,虽然微调的音频预训练VGG和Transformer模型在某些任务中实现了最先进的性能,但在其他任务中却失败了。本研究通过降低学习嵌入的维度并通过聚类对其进行评估,对11个DL模型进行了相同任务的基准测试。我们发现,音频预训练的DL模型1)没有微调甚至不如微调的AlexNet,2)有和没有微调都无法将背景从标记的声音中分离出来,但ResNet做到了,3)当微调过程中包含较少的背景声音时,表现优于其他模型。这项研究强调了微调音频预训练模型并在微调后检查嵌入的必要性。我们的代码可在以下网址获得:https://github.com/NeuroscienceAI/Audio\_Embeddings
摘要:Bioacoustics, the study of animal sounds, offers a non-invasive method to monitor ecosystems. Extracting embeddings from audio-pretrained deep learning (DL) models without fine-tuning has become popular for obtaining bioacoustic features for tasks. However, a recent benchmark study reveals that while fine-tuned audio-pretrained VGG and transformer models achieve state-of-the-art performance in some tasks, they fail in others. This study benchmarks 11 DL models on the same tasks by reducing their learned embeddings' dimensionality and evaluating them through clustering. We found that audio-pretrained DL models 1) without fine-tuning even underperform fine-tuned AlexNet, 2) both with and without fine-tuning fail to separate the background from labeled sounds, but ResNet does, and 3) outperform other models when fewer background sounds are included during fine-tuning. This study underscores the necessity of fine-tuning audio-pretrained models and checking the embeddings after fine-tuning. Our codes are available: https://github.com/NeuroscienceAI/Audio\_Embeddings
【10】Dynamic Synchronization and Resonance as a Universal Origin of 1/f Fluctuations -- Amplitude Modulation Across Music and Nature
标题:动态同步和共振是1/f波动的普遍起源--跨越音乐和自然的幅度调制
链接:https://arxiv.org/abs/2508.10049
备注:14 pages, 10 figures
摘要:我们提出了一个普遍的物理机制出现的1/f波动,在广泛的系统中观察到。特别是,我们验证这声学的情况下。该机制基于幅度调制(AM)和解调(DM),其中1/f频谱定律不是在原始波形中而是在其解调的幅度包络中出现。两个不同但互补的过程产生所需的AM:(i)振荡器之间的随机同步,通过扩展的仓本框架,捕获永久同步-去同步周期建模,以及(ii)频率选择性共振,通过声学或结构环境中本征模式的频谱积累建模。数值模拟表明,这两种机制,单独或组合作用,稳健地产生1/f谱在几十年内,当DM应用,经典的仓本临界点是不必要的,他们的出现。我们通过对音乐表演、地震记录和天体物理时间序列的分析,展示了AM/DM框架的跨域相关性,揭示了一个共同的基本结构。这项工作建立了解调作为1/f波动的一般途径,为它在自然和工程系统中的普遍存在提供了一个简单和可扩展的解释。 关键词:1/f涨落,调幅,同步,共振,仓本模型,音乐,自然噪声,解调
摘要:We propose a universal physical mechanism for the emergence of 1/f fluctuations, observed across a wide range of systems. In particular, we verify this on acoustic cases. The mechanism is based on amplitude modulation (AM) and demodulation (DM), where the 1/f spectral law arises not in the raw waveform but in its demodulated amplitude envelope. Two distinct yet complementary processes generate the required AM: (i) stochastic synchronization among oscillators, modeled via an extended Kuramoto framework that captures perpetual synchronization-desynchronization cycles, and (ii) frequency-selective resonance, modeled by spectral accumulation of eigenmodes in acoustic or structural environments. Numerical simulations demonstrate that both mechanisms, acting separately or in combination, robustly produce 1/f spectra over several decades when DM is applied, and that the classical Kuramoto critical point is not necessary for their emergence. We demonstrate the cross-domain relevance of this AM/DM framework through analyses of musical performances, seismic records, and astrophysical time series, revealing a common underlying structure. This work establishes demodulation as a general route to 1/f fluctuations, providing a simple and scalable explanation for its ubiquity in both natural and engineered systems. Keywords: 1/f fluctuation, amplitude modulation, synchronization, resonance, Kuramoto model, music, natural noise, demodulation
【11】Beyond Hard Sharing: Efficient Multi-Task Speech-to-Text Modeling with Supervised Mixture of Experts
标题:超越硬共享:具有受监督混合专家的高效多任务语音到文本建模
链接:https://arxiv.org/abs/2508.10009
备注:Accepted to Interspeech 2025
摘要:硬参数共享是跨不同任务联合训练单个模型的常见策略。然而,这通常会导致任务干扰,阻碍整体模型性能。为了解决这个问题,我们提出了一个简单而有效的监督混合专家(S-MoE)。与传统的专家混合模型不同,S-MoE通过利用特殊的引导令牌将每个任务路由到其指定的专家,消除了对训练门控函数的需要。通过将每个任务分配给一个单独的前馈网络,S-MoE克服了硬参数共享的限制。我们进一步将S-MoE应用于语音到文本模型,使该模型能够处理混合带宽输入,同时联合执行自动语音识别(ASR)和语音翻译(ST)。实验结果表明,所提出的S-MoE的有效性,实现了6.35%的字错误率(WER)的相对改善时,适用于编码器和解码器。
摘要:Hard-parameter sharing is a common strategy to train a single model jointly across diverse tasks. However, this often leads to task interference, impeding overall model performance. To address the issue, we propose a simple yet effective Supervised Mixture of Experts (S-MoE). Unlike traditional Mixture of Experts models, S-MoE eliminates the need for training gating functions by utilizing special guiding tokens to route each task to its designated expert. By assigning each task to a separate feedforward network, S-MoE overcomes the limitations of hard-parameter sharing. We further apply S-MoE to a speech-to-text model, enabling the model to process mixed-bandwidth input while jointly performing automatic speech recognition (ASR) and speech translation (ST). Experimental results demonstrate the effectiveness of the proposed S-MoE, achieving a 6.35% relative improvement in Word Error Rate (WER) when applied to both the encoder and decoder.
【12】Whisper Smarter, not Harder: Adversarial Attack on Partial Suppression
标题:低语更聪明,而不是更难:对部分抑制的对抗攻击
链接:https://arxiv.org/abs/2508.09994
备注:13 pages, 7 figures
摘要:目前,自动语音识别(ASR)模型被部署在广泛的应用中。然而,最近的研究已经证明了对这些模型进行对抗性攻击的可能性,这可能会抑制或破坏模型输出。我们调查和验证这些攻击的鲁棒性,并探讨是否有可能提高其不可感知性。我们还发现,通过放松的优化目标从完全抑制部分抑制,我们可以进一步降低攻击的不可感知性。我们还探讨了针对这些攻击的可能防御措施,并表明低通滤波器防御可能是一种有效的防御措施。
摘要:Currently, Automatic Speech Recognition (ASR) models are deployed in an extensive range of applications. However, recent studies have demonstrated the possibility of adversarial attack on these models which could potentially suppress or disrupt model output. We investigate and verify the robustness of these attacks and explore if it is possible to increase their imperceptibility. We additionally find that by relaxing the optimisation objective from complete suppression to partial suppression, we can further decrease the imperceptibility of the attack. We also explore possible defences against these attacks and show a low-pass filter defence could potentially serve as an effective defence.
【13】Layer-Wise Analysis of Self-Supervised Representations for Age and Gender Classification in Children's Speech
标题:儿童言语中年龄和性别分类的自我监督表征的分层分析
链接:https://arxiv.org/abs/2508.10332
备注:Accepted at Workshop on Child Computer Interaction (WOCCI 2025)
摘要:儿童的语音提出了挑战,年龄和性别分类,由于高度的变化,在音高,清晰度和发展特点。虽然自我监督学习(SSL)模型在成人语音任务中表现良好,但它们对儿童说话者特征进行编码的能力仍然未得到充分研究。本文使用PFSTAR和CMU Kids数据集对四种Wav 2 Vec 2变体进行了详细的逐层分析。结果表明,早期层(1-7)比更深的层更有效地捕捉特定于说话者的线索,而更深的层越来越关注语言信息。应用PCA进一步改进了分类,减少了冗余并突出了信息量最大的组件。Wav 2 Vec 2-large-lv 60模型在CMU Kids上达到97.14%(年龄)和98.20%(性别); base-100 h和large-lv 60模型在PFSTAR上达到86.05%和95.00%。这些结果揭示了说话人特征是如何在SSL模型深度上进行结构化的,并支持针对儿童感知语音界面的更有针对性的自适应策略。
摘要:Children's speech presents challenges for age and gender classification due to high variability in pitch, articulation, and developmental traits. While self-supervised learning (SSL) models perform well on adult speech tasks, their ability to encode speaker traits in children remains underexplored. This paper presents a detailed layer-wise analysis of four Wav2Vec2 variants using the PFSTAR and CMU Kids datasets. Results show that early layers (1-7) capture speaker-specific cues more effectively than deeper layers, which increasingly focus on linguistic information. Applying PCA further improves classification, reducing redundancy and highlighting the most informative components. The Wav2Vec2-large-lv60 model achieves 97.14% (age) and 98.20% (gender) on CMU Kids; base-100h and large-lv60 models reach 86.05% and 95.00% on PFSTAR. These results reveal how speaker traits are structured across SSL model depth and support more targeted, adaptive strategies for child-aware speech interfaces.
【1】Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
标题:探索适形器-传感器语音识别系统的跨言语语音上下文
链接:https://arxiv.org/abs/2508.10456
摘要:本文研究了流和非流Conformer-Transformer(C-T)ASR系统中四种类型的跨话语语音上下文建模方法:i)输入音频特征级联; ii)跨话语编码器嵌入级联; iii)跨话语编码器嵌入池化投影; iv)首次应用于C-T模型的新的基于块的方法。提出了一种有效的批量训练方案,上下文C-T,使用拼接的语音话语内的每个minibatch,以尽量减少同步开销,同时保持跨话语语音上下文的顺序。实验在三种语言的四个基准语音数据集上进行:用于上下文C-T模型预训练的英语GigaSpeech和普通话Wenetspeech语料库;以及用于领域微调的英语DementiaBank Pitt和粤语JCCOCC MoCA老年语音数据集。在预训练和微调阶段,最好的上下文C-T系统在没有交叉话语语音上下文的情况下始终优于其各自的基线,统计上显着的平均单词错误率(WER)或字符错误率(CER)降低了0.9%,1.1%,0.51%和0.98%的绝对值(6.0%、5.4%、2.0%和3.4%)。它们与Wav2vec2.0-Conformer、XLSR-128和Whisper模型的性能竞争力突出了将跨话语语音上下文纳入当前语音基础模型的潜在好处。
摘要:This paper investigates four types of cross-utterance speech contexts modeling approaches for streaming and non-streaming Conformer-Transformer (C-T) ASR systems: i) input audio feature concatenation; ii) cross-utterance Encoder embedding concatenation; iii) cross-utterance Encoder embedding pooling projection; or iv) a novel chunk-based approach applied to C-T models for the first time. An efficient batch-training scheme is proposed for contextual C-Ts that uses spliced speech utterances within each minibatch to minimize the synchronization overhead while preserving the sequential order of cross-utterance speech contexts. Experiments are conducted on four benchmark speech datasets across three languages: the English GigaSpeech and Mandarin Wenetspeech corpora used in contextual C-T models pre-training; and the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech datasets used in domain fine-tuning. The best performing contextual C-T systems consistently outperform their respective baselines using no cross-utterance speech contexts in pre-training and fine-tuning stages with statistically significant average word error rate (WER) or character error rate (CER) reductions up to 0.9%, 1.1%, 0.51%, and 0.98% absolute (6.0%, 5.4%, 2.0%, and 3.4% relative) on the four tasks respectively. Their performance competitiveness against Wav2vec2.0-Conformer, XLSR-128, and Whisper models highlights the potential benefit of incorporating cross-utterance speech contexts into current speech foundation models.
【2】Towards Frame-level Quality Predictions of Synthetic Speech
标题:合成语音的帧级质量预测
链接:https://arxiv.org/abs/2508.10374
备注:Accepted at Interspeech 2025
摘要:虽然自动主观语音质量评估已经取得了很大的进展,一个悬而未决的问题是,是否有可能在帧分辨率的自动质量评估。这将是非常可取的,因为它增加了可解释性的语音合成系统的评估。在这里,我们通过识别现有质量预测器的问题来实现这一目标的第一步,这些问题会阻止合理的帧级预测。此外,我们定义了帧级预测器应该满足的标准。我们还提出了一个基于块的处理,避免了局部失真的影响,对相邻帧的分数。最后,我们在实验中与本地化的人工失真的一组帧级质量预测的本地化性能进行测量,并表明它们可以优于从众包感知实验中获得的人类注释的检测性能。
摘要:While automatic subjective speech quality assessment has witnessed much progress, an open question is whether an automatic quality assessment at frame resolution is possible. This would be highly desirable, as it adds explainability to the assessment of speech synthesis systems. Here, we take first steps towards this goal by identifying issues of existing quality predictors that prevent sensible frame-level prediction. Further, we define criteria that a frame-level predictor should fulfill. We also suggest a chunk-based processing that avoids the impact of a localized distortion on the score of neighboring frames. Finally, we measure in experiments with localized artificial distortions the localization performance of a set of frame-level quality predictors and show that they can outperform detection performance of human annotations obtained from a crowd-sourced perception experiment.
【3】Layer-Wise Analysis of Self-Supervised Representations for Age and Gender Classification in Children's Speech
标题:儿童言语中年龄和性别分类的自我监督表征的分层分析
链接:https://arxiv.org/abs/2508.10332
备注:Accepted at Workshop on Child Computer Interaction (WOCCI 2025)
摘要:儿童的语音提出了挑战,年龄和性别分类,由于高度的变化,在音高,清晰度和发展特点。虽然自我监督学习(SSL)模型在成人语音任务中表现良好,但它们对儿童说话者特征进行编码的能力仍然未得到充分研究。本文使用PFSTAR和CMU Kids数据集对四种Wav 2 Vec 2变体进行了详细的逐层分析。结果表明,早期的层(1-7)捕捉特定于说话人的线索比更深的层,越来越多地集中在语言信息更有效。应用PCA进一步改进了分类,减少了冗余并突出了信息量最大的组件。Wav 2 Vec 2-large-lv 60模型在CMU Kids上的成绩为97.14%(年龄)和98.20%(性别); base-100 h和large-lv 60模型在PFSTAR上的成绩为86.05%和95.00%。这些结果揭示了说话人特征是如何在SSL模型深度上进行结构化的,并支持针对儿童感知语音界面的更有针对性的自适应策略。
摘要:Children's speech presents challenges for age and gender classification due to high variability in pitch, articulation, and developmental traits. While self-supervised learning (SSL) models perform well on adult speech tasks, their ability to encode speaker traits in children remains underexplored. This paper presents a detailed layer-wise analysis of four Wav2Vec2 variants using the PFSTAR and CMU Kids datasets. Results show that early layers (1-7) capture speaker-specific cues more effectively than deeper layers, which increasingly focus on linguistic information. Applying PCA further improves classification, reducing redundancy and highlighting the most informative components. The Wav2Vec2-large-lv60 model achieves 97.14% (age) and 98.20% (gender) on CMU Kids; base-100h and large-lv60 models reach 86.05% and 95.00% on PFSTAR. These results reveal how speaker traits are structured across SSL model depth and support more targeted, adaptive strategies for child-aware speech interfaces.
【4】Advances in Speech Separation: Techniques, Challenges, and Future Trends
标题:语音分离的进展:技术、挑战和未来趋势
链接:https://arxiv.org/abs/2508.10830
备注:34 pages, 10 figures
摘要:语音分离领域,解决“鸡尾酒会问题”,已经看到了DNN的革命性进展。语音分离是语音识别和说话人识别的关键预处理技术,它能提高复杂声学环境下的语音清晰度。然而,目前的文献狭隘地关注于特定的架构或孤立的方法,造成了零散的理解。本调查通过对基于DNN的语音分离技术进行系统检查来解决这一差距。我们的工作通过以下方面区分开来:(I)全面的视角:我们系统地研究了学习范式,已知/未知扬声器的分离场景,监督/自监督/无监督框架的比较分析,以及从编码器到估计策略的架构组件。(II)及时性:对前沿发展的报道确保获得当前的创新和基准。(III)独特的见解:除了总结,我们评估技术轨迹,识别新兴模式,并强调有前途的方向,包括领域强大的框架,高效的架构,多模式集成和新颖的自我监督范式。(IV)公平评估:我们对标准数据集进行定量评估,揭示不同方法的真实能力和局限性。这项全面的调查为经验丰富的研究人员和新来者导航语音分离的复杂景观提供了一个可访问的参考。
摘要:The field of speech separation, addressing the "cocktail party problem", has seen revolutionary advances with DNNs. Speech separation enhances clarity in complex acoustic environments and serves as crucial pre-processing for speech recognition and speaker recognition. However, current literature focuses narrowly on specific architectures or isolated approaches, creating fragmented understanding. This survey addresses this gap by providing systematic examination of DNN-based speech separation techniques. Our work differentiates itself through: (I) Comprehensive perspective: We systematically investigate learning paradigms, separation scenarios with known/unknown speakers, comparative analysis of supervised/self-supervised/unsupervised frameworks, and architectural components from encoders to estimation strategies. (II) Timeliness: Coverage of cutting-edge developments ensures access to current innovations and benchmarks. (III) Unique insights: Beyond summarization, we evaluate technological trajectories, identify emerging patterns, and highlight promising directions including domain-robust frameworks, efficient architectures, multimodal integration, and novel self-supervised paradigms. (IV) Fair evaluation: We provide quantitative evaluations on standard datasets, revealing true capabilities and limitations of different methods. This comprehensive survey serves as an accessible reference for experienced researchers and newcomers navigating speech separation's complex landscape.
【5】Alternating Approach-Putt Models for Multi-Stage Speech Enhancement
标题:交替逼近-Putt模型的多级语音增强
链接:https://arxiv.org/abs/2508.10436
备注:This work has been submitted to the IEEE for possible publication
摘要:使用人工神经网络的语音增强旨在从含噪语音信号中去除噪声,同时保留语音内容。然而,语音增强网络经常向语音信号引入失真,称为伪像,这会降低音频质量。在这项工作中,我们提出了一个后处理神经网络,旨在减轻语音增强模型引入的文物。灵感来自于在高尔夫中的“接近”之后进行“推杆”的类比,我们将我们的模型命名为PuttNet。我们证明,交替之间的语音增强模型和Putt模型,导致改善语音质量,感知质量分数(PESQ),客观清晰度(STOI),和背景噪声侵入(CBAK)分数。此外,我们用图形分析说明了为什么这种交替的方法优于单独使用任何一种模型的重复应用。
摘要:Speech enhancement using artificial neural networks aims to remove noise from noisy speech signals while preserving the speech content. However, speech enhancement networks often introduce distortions to the speech signal, referred to as artifacts, which can degrade audio quality. In this work, we propose a post-processing neural network designed to mitigate artifacts introduced by speech enhancement models. Inspired by the analogy of making a `Putt' after an `Approach' in golf, we name our model PuttNet. We demonstrate that alternating between a speech enhancement model and the proposed Putt model leads to improved speech quality, as measured by perceptual quality scores (PESQ), objective intelligibility (STOI), and background noise intrusiveness (CBAK) scores. Furthermore, we illustrate with graphical analysis why this alternating Approach outperforms repeated application of either model alone.
【6】MCP2OSC: Parametric Control by Natural Language
标题:MPP 2OSC:自然语言参数控制
链接:https://arxiv.org/abs/2508.10414
摘要:文本提示可以实现直观的内容创建,但可能无法实现复杂任务的高精度;旋钮或滑块控件提供精确的调整,但代价是增加了复杂性。为了解决旋钮和提示之间的差距,一个新的MCP(模型上下文协议)服务器和一组独特的提示设计标准,使探索参数OSC(OpenSoundControl)控制的自然语言提示。通过14个具有最佳实践和通用提示模板的实际QA示例,本研究发现Claude与MCP 2 OSC服务器集成,有效地通过自然语言生成OSC消息,解释,搜索和可视化OSC消息,验证和调试OSC消息,以及管理OSC地址模式。MCP 2 OSC通过利用LLM(大型语言模型)来处理复杂的OSC开发任务,并通过具有灵活精度控制的直观语言界面来增强人类的创造力,从而增强了人机协作:一个基于JavaScript的OSC工具。这项研究提供了一个新的视角,创造性的MCP应用程序在网络协议层面上,利用LLM的实力,直接处理和生成人类可读的OSC消息。结果表明,它的潜在的基于LLM的多媒体设备的通用控制机制。
摘要:Text prompts enable intuitive content creation but may fall short in achieving high precision for intricate tasks; knob or slider controls offer precise adjustments at the cost of increased complexity. To address the gap between knobs and prompts, a new MCP (Model Context Protocol) server and a unique set of prompt design criteria are presented to enable exploring parametric OSC (OpenSoundControl) control by natural language prompts. Demonstrated by 14 practical QA examples with best practices and the generalized prompt templates, this study finds Claude integrated with the MCP2OSC server effective in generating OSC messages by natural language, interpreting, searching, and visualizing OSC messages, validating and debugging OSC messages, and managing OSC address patterns. MCP2OSC enhances human-machine collaboration by leveraging LLM (Large Language Model) to handle intricate OSC development tasks, and by empowering human creativity with an intuitive language interface featuring flexible precision controls: a prompt-based OSC tool. This study provides a novel perspective on the creative MCP application at the network protocol level by utilizing LLM's strength in directly processing and generating human-readable OSC messages. The results suggest its potential for a LLM-based universal control mechanism for multimedia devices.
【7】A dataset and model for recognition of audiologically relevant environments for hearing aids: AHEAD-DS and YAMNet+
标题:识别助听器听力相关环境的数据集和模型:AHEAD-DS和YAMNet+
链接:https://arxiv.org/abs/2508.10360
摘要:听觉相关环境的场景识别对于助听器很重要;然而,它具有挑战性,部分原因是现有数据集的局限性。数据集通常缺乏公共可访问性,完整性或听觉相关标签,阻碍了机器学习模型的系统比较。在资源受限的边缘设备上部署这些模型是另一个挑战。我们的解决方案是双重的:我们利用几个开源数据集创建AHEAD-DS,一个为听觉相关环境的场景识别设计的数据集,并引入YAMNet+,一个声音识别模型。AHEAD-DS旨在提供一个标准化的、公开可用的数据集,具有与助听器相关的一致标签,便于模型比较。YAMNet+旨在部署在边缘设备上,如连接到听力设备的智能手机,如助听器和具有助听器功能的无线耳机;作为基于声音的场景识别的基线模型。YAMNet+在14类听觉相关环境中的AHEAD-DS测试集上实现了平均0.83的平均精度和0.93的准确度。我们发现,从预训练的YAMNet模型中应用迁移学习是必不可少的。我们通过将YAMNet+部署到Android智能手机上,在边缘设备上展示了基于声音的实时场景识别功能。即使使用Google Pixel 3(2018年发布的一款规格适中的手机),该模型处理音频的延迟也约为50 ms以加载模型,并且每1秒音频近似线性增加30 ms。我们的网站和代码https://github.com/Australian-Future-Hearing-Initiative。
摘要:Scene recognition of audiologically relevant environments is important for hearing aids; however, it is challenging, in part because of the limitations of existing datasets. Datasets often lack public accessibility, completeness, or audiologically relevant labels, hindering systematic comparison of machine learning models. Deploying these models on resource-constrained edge devices presents another challenge. Our solution is two-fold: we leverage several open source datasets to create AHEAD-DS, a dataset designed for scene recognition of audiologically relevant environments, and introduce YAMNet+, a sound recognition model. AHEAD-DS aims to provide a standardised, publicly available dataset with consistent labels relevant to hearing aids, facilitating model comparison. YAMNet+ is designed for deployment on edge devices like smartphones connected to hearing devices, such as hearing aids and wireless earphones with hearing aid functionality; serving as a baseline model for sound-based scene recognition. YAMNet+ achieved a mean average precision of 0.83 and accuracy of 0.93 on the testing set of AHEAD-DS across fourteen categories of audiologically relevant environments. We found that applying transfer learning from the pretrained YAMNet model was essential. We demonstrated real-time sound-based scene recognition capabilities on edge devices by deploying YAMNet+ to an Android smartphone. Even with a Google Pixel 3 (a phone with modest specifications, released in 2018), the model processes audio with approximately 50ms of latency to load the model, and an approximate linear increase of 30ms per 1 second of audio. Our website and code https://github.com/Australian-Future-Hearing-Initiative .
【8】Beyond Hard Sharing: Efficient Multi-Task Speech-to-Text Modeling with Supervised Mixture of Experts
标题:超越硬共享:具有受监督混合专家的高效多任务语音到文本建模
链接:https://arxiv.org/abs/2508.10009
备注:Accepted to Interspeech 2025
摘要:硬参数共享是跨不同任务联合训练单个模型的常见策略。然而,这通常会导致任务干扰,从而阻碍整体模型性能。为了解决这个问题,我们提出了一个简单而有效的监督混合专家(S-MoE)。与传统的专家混合模型不同,S-MoE通过利用特殊的引导令牌将每个任务路由到其指定的专家,消除了对训练门控函数的需要。通过将每个任务分配给一个单独的前馈网络,S-MoE克服了硬参数共享的限制。我们进一步将S-MoE应用于语音到文本模型,使该模型能够处理混合带宽输入,同时联合执行自动语音识别(ASR)和语音翻译(ST)。实验结果表明,所提出的S-MoE的有效性,实现了6.35%的字错误率(WER)的相对改善时,适用于编码器和解码器。
摘要:Hard-parameter sharing is a common strategy to train a single model jointly across diverse tasks. However, this often leads to task interference, impeding overall model performance. To address the issue, we propose a simple yet effective Supervised Mixture of Experts (S-MoE). Unlike traditional Mixture of Experts models, S-MoE eliminates the need for training gating functions by utilizing special guiding tokens to route each task to its designated expert. By assigning each task to a separate feedforward network, S-MoE overcomes the limitations of hard-parameter sharing. We further apply S-MoE to a speech-to-text model, enabling the model to process mixed-bandwidth input while jointly performing automatic speech recognition (ASR) and speech translation (ST). Experimental results demonstrate the effectiveness of the proposed S-MoE, achieving a 6.35% relative improvement in Word Error Rate (WER) when applied to both the encoder and decoder.
【9】Whisper Smarter, not Harder: Adversarial Attack on Partial Suppression
标题:低语更聪明,而不是更难:对部分抑制的对抗攻击
链接:https://arxiv.org/abs/2508.09994
备注:13 pages, 7 figures
摘要:目前,自动语音识别(ASR)模型被部署在广泛的应用中。然而,最近的研究已经证明了对这些模型进行对抗性攻击的可能性,这可能会抑制或破坏模型输出。我们调查和验证这些攻击的鲁棒性,并探讨是否有可能提高其不可感知性。我们还发现,通过将优化目标从完全抑制放宽到部分抑制,我们可以进一步降低攻击的不可感知性。我们还探讨了针对这些攻击的可能防御措施,并表明低通滤波器防御可能是一种有效的防御措施。
摘要:Currently, Automatic Speech Recognition (ASR) models are deployed in an extensive range of applications. However, recent studies have demonstrated the possibility of adversarial attack on these models which could potentially suppress or disrupt model output. We investigate and verify the robustness of these attacks and explore if it is possible to increase their imperceptibility. We additionally find that by relaxing the optimisation objective from complete suppression to partial suppression, we can further decrease the imperceptibility of the attack. We also explore possible defences against these attacks and show a low-pass filter defence could potentially serve as an effective defence.
机器翻译由腾讯交互翻译提供,仅供参考
