今日论文合集:cs.SD语音6篇,eess.AS音频处理7篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction
标题:BRASP:用于自动MOS预测的双分辨率专注统计池框架
链接:https://arxiv.org/abs/2508.21407

作者: Cheng-Yeh Yang, Kuan-Tang Huang, Chien-Chun Wang, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen
备注:Accepted to APSIPA ASC 2025
摘要:池机制是必不可少的平均意见得分(MOS)预测,促进可变长度的音频特征转换成一个简洁的固定大小的表示,有效地编码语音质量。现有的池化方法通常以单一粒度操作,集中于全面的全局视角或详细的帧级分析,这可能会忽略互补的感知见解。为了解决这个问题,我们引入了双分辨率注意统计池(Dual-Resolution Attentive Statistics Pooling,DRASP)框架。DRASP集成了粗粒度的全球统计摘要和细粒度的对感知重要部分的仔细分析。这种双视图架构使我们的模型能够制定一个更全面和强大的表示,同时捕获总体结构背景和突出的局部细节。大量的实验验证了该框架的有效性和较强的泛化能力。它在不同的数据集(MusicEval和AES-Natural),MOS预测骨干(包括基于CLAP的模型和AudioBox美学)和不同的音频生成系统中始终优于各种基线方法,与广泛使用的平均池化方法相比,系统级Spearman秩相关系数(SRCC)相对提高了10.39%。
摘要:A pooling mechanism is essential for mean opinion score (MOS) prediction, facilitating the transformation of variable-length audio features into a concise fixed-size representation that effectively encodes speech quality. Existing pooling methods typically operate at a singular granularity, concentrating either on a comprehensive global perspective or a detailed frame-level analysis, which may overlook complementary perceptual insights. To address this limitation, we introduce the Dual-Resolution Attentive Statistics Pooling (DRASP) framework. DRASP integrates both coarse-grained, global statistical summaries and fine-grained, attentive analyses of perceptually significant segments. This dual-view architecture empowers our model to formulate a more thorough and robust representation, capturing both the overarching structural context and salient local details concurrently. Extensive experiments validate the effectiveness and strong generalization ability of the proposed framework. It consistently outperforms various baseline methods across diverse datasets (MusicEval and AES-Natural), MOS prediction backbones (including a CLAP-based model and AudioBox-Aesthetics), and different audio generation systems, achieving a relative improvement of 10.39% in system-level Spearman's rank correlation coefficient (SRCC) over the widely-used average pooling approach.


【2】Full-Frequency Temporal Patching and Structured Masking for Enhanced Audio Classification
标题:用于增强音频分类的全频率时间修补和结构化掩蔽
链接:https://arxiv.org/abs/2508.21243

作者:Aditya Makineni, Baocheng Geng, Qing Tian
摘要:Transformers和状态空间模型(SSM)通过将频谱图建模为补丁序列来改进音频分类。然而,现有的模型,如音频频谱图Transformer(AST)和音频曼巴(AuM)采用来自计算机视觉的方形修补,这会破坏连续的频率模式并产生过多的补丁,减慢训练速度并增加计算量。我们提出了全频时间修补(FFTP),修补策略,更好地匹配的时间-频率不对称的频谱图跨越全频带与本地化的时间背景,保留谐波结构,并显着减少补丁计数和计算。我们还介绍了SpecMask,一种补丁对齐的频谱图增强,它在固定的掩蔽预算下结合了全频和局部时频掩模,在保持频谱连续性的同时增强了时间鲁棒性。当应用于AST和AuM时,我们使用SpecMask的修补方法在AudioSet-18 k上将mAP提高了+6.76,在SpeechCommandsV 2上将准确度提高了+8.46,同时将计算量减少了83.26%,证明了性能和效率的提高。
摘要:Transformers and State-Space Models (SSMs) have advanced audio classification by modeling spectrograms as sequences of patches. However, existing models such as the Audio Spectrogram Transformer (AST) and Audio Mamba (AuM) adopt square patching from computer vision, which disrupts continuous frequency patterns and produces an excessive number of patches, slowing training, and increasing computation. We propose Full-Frequency Temporal Patching (FFTP), a patching strategy that better matches the time-frequency asymmetry of spectrograms by spanning full frequency bands with localized temporal context, preserving harmonic structure, and significantly reducing patch count and computation. We also introduce SpecMask, a patch-aligned spectrogram augmentation that combines full-frequency and localized time-frequency masks under a fixed masking budget, enhancing temporal robustness while preserving spectral continuity. When applied on both AST and AuM, our patching method with SpecMask improves mAP by up to +6.76 on AudioSet-18k and accuracy by up to +8.46 on SpeechCommandsV2, while reducing computation by up to 83.26%, demonstrating both performance and efficiency gains.


【3】RARR : Robust Real-World Activity Recognition with Vibration by Scavenging Near-Surface Audio Online
标题:RARR:通过在线清理近表面音频来实现具有振动的鲁棒的现实世界活动识别
链接:https://arxiv.org/abs/2508.21167

作者: Dong Yoon Lee, Alyssa Weakley, Hui Wei, Blake Brown, Keyana Carrion, Shijia Pan
摘要:四分之一的痴呆症患者独自生活,导致家庭成员从远处承担起照顾的角色。许多研究人员已经开发了远程监控解决方案来减少监控需求;然而,仍然存在局限性,包括隐私保护解决方案,活动识别以及对新用户和环境的模型推广。结构振动传感器系统是一种不显眼的解决方案,已被证明可以通过感测活动产生的表面振动,在受控环境中准确监测人类信息,例如身份识别和活动识别。然而,当部署在最终用户的家中时,当前的解决方案需要大量的标记数据来进行准确的活动识别。我们的可扩展解决方案采用来自近表面声学音频的合成数据来预训练模型,并允许使用非常有限的数据进行微调,以便为日常跟踪创建一个强大的框架。
摘要:One in four people dementia live alone, leading family members to take on caregiving roles from a distance. Many researchers have developed remote monitoring solutions to lessen caregiving needs; however, limitations remain including privacy preserving solutions, activity recognition, and model generalizability to new users and environments. Structural vibration sensor systems are unobtrusive solutions that have been proven to accurately monitor human information, such as identification and activity recognition, in controlled settings by sensing surface vibrations generated by activities. However, when deploying in an end user's home, current solutions require a substantial amount of labeled data for accurate activity recognition. Our scalable solution adapts synthesized data from near-surface acoustic audio to pretrain a model and allows fine tuning with very limited data in order to create a robust framework for daily routine tracking.


【4】WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration
标题:WaveLLDM:用于语音增强和恢复的轻量级潜在扩散模型的设计和开发
链接:https://arxiv.org/abs/2508.21153

作者:Kevin Putra Santoso, Rizka Wakhidatus Sholikah, Raden Venantius Hari Ginardi
摘要:高质量的音频在广泛的应用中至关重要,包括在线通信、虚拟助手和多媒体行业。然而,由噪声、压缩和传输伪影引起的降级仍然是一个主要挑战。虽然扩散模型已被证明对音频恢复是有效的,但它们通常需要大量的计算资源,并且难以处理较长的缺失片段。本研究介绍了WaveLLDM(Wave Lightweight Latent Diffusion Model),这是一种集成了高效神经音频编解码器和潜在扩散的架构,用于音频恢复和去噪。与在时域或频谱域中操作的传统方法不同,WaveLLDM在压缩的潜在空间中处理音频,从而在保持重建质量的同时降低计算复杂度。在Voicebank+DEMAND测试集上的经验评估表明,WaveLLDM实现了准确的频谱重建,具有低对数频谱距离(LSD)分数(0.48至0.60)和对未知数据的良好适应性。然而,与最先进的方法相比,它在感知质量和语音清晰度方面仍然表现不佳,WB-PESQ得分范围为1.62至1.71,STOI得分在0.76至0.78之间。这些限制归因于次优的架构调优、缺乏微调以及训练持续时间不足。尽管如此,结合神经音频编解码器和潜在扩散模型的灵活架构为未来的开发提供了坚实的基础。
摘要:High-quality audio is essential in a wide range of applications, including online communication, virtual assistants, and the multimedia industry. However, degradation caused by noise, compression, and transmission artifacts remains a major challenge. While diffusion models have proven effective for audio restoration, they typically require significant computational resources and struggle to handle longer missing segments. This study introduces WaveLLDM (Wave Lightweight Latent Diffusion Model), an architecture that integrates an efficient neural audio codec with latent diffusion for audio restoration and denoising. Unlike conventional approaches that operate in the time or spectral domain, WaveLLDM processes audio in a compressed latent space, reducing computational complexity while preserving reconstruction quality. Empirical evaluations on the Voicebank+DEMAND test set demonstrate that WaveLLDM achieves accurate spectral reconstruction with low Log-Spectral Distance (LSD) scores (0.48 to 0.60) and good adaptability to unseen data. However, it still underperforms compared to state-of-the-art methods in terms of perceptual quality and speech clarity, with WB-PESQ scores ranging from 1.62 to 1.71 and STOI scores between 0.76 and 0.78. These limitations are attributed to suboptimal architectural tuning, the absence of fine-tuning, and insufficient training duration. Nevertheless, the flexible architecture that combines a neural audio codec and latent diffusion model provides a strong foundation for future development.


【5】Zero-Shot KWS for Children's Speech using Layer-Wise Features from SSL Models
标题:使用SSL模型的分层功能针对儿童语音的Zero-ShotKWS
链接:https://arxiv.org/abs/2508.21248

作者:Subham Kutum, Abhijit Sinha, Hemant Kumar Kathania, Sudarsana Reddy Kadiri, Mahesh Chandra Govil
备注:Accepted
摘要:已经提出了许多方法来增强成人语音中的关键词识别(KWS),但儿童语音由于其独特的声学和语言特征,对KWS系统提出了独特的挑战。本文介绍了一种零触发KWS方法(zero-shot KWS),该方法利用了最先进的自监督学习(SSL)模型,包括Wave 2 Vec 2、HuBERT和Data 2 Vec。从这些SSL模型中逐层提取特征,并用于训练基于Kaldi的DNN KWS系统。WSJCAM 0成人语音数据集用于训练,而PFSTAR儿童语音数据集用于测试,证明了我们的方法的zero-shot能力。我们的方法在儿童语音的所有关键词集上都取得了最先进的结果。值得注意的是,Wav 2 Vec 2模型,特别是第22层,表现最好,对于一组30个关键字,ATWV得分为0.691,MTWV得分为0.7003,虚警概率和未命中概率分别为0.0164和0.0547。此外,针对不同年龄段儿童的绩效评估证实了该系统的有效性。为了评估系统对噪声的鲁棒性,使用性能最好的Wav 2 Vec 2模型的性能最好的层进行了额外的实验。结果表明,与传统的基于MFCC的基线相比,有了显着的改进,强调了SSL嵌入的潜力,即使在嘈杂的条件下。为了进一步推广KWS框架,针对额外的CMU数据集重复实验。总体而言,结果突出了SSL功能在增强儿童语音的Zero-Shot KWS性能方面的重大贡献,有效地解决了与儿童说话者的独特特征相关的挑战。
摘要:Numerous methods have been proposed to enhance Keyword Spotting (KWS) in adult speech, but children's speech presents unique challenges for KWS systems due to its distinct acoustic and linguistic characteristics. This paper introduces a zero-shot KWS approach that leverages state-of-the-art self-supervised learning (SSL) models, including Wav2Vec2, HuBERT and Data2Vec. Features are extracted layer-wise from these SSL models and used to train a Kaldi-based DNN KWS system. The WSJCAM0 adult speech dataset was used for training, while the PFSTAR children's speech dataset was used for testing, demonstrating the zero-shot capability of our method. Our approach achieved state-of-the-art results across all keyword sets for children's speech. Notably, the Wav2Vec2 model, particularly layer 22, performed the best, delivering an ATWV score of 0.691, a MTWV score of 0.7003 and probability of false alarm and probability of miss of 0.0164 and 0.0547 respectively, for a set of 30 keywords. Furthermore, age-specific performance evaluation confirmed the system's effectiveness across different age groups of children. To assess the system's robustness against noise, additional experiments were conducted using the best-performing layer of the best-performing Wav2Vec2 model. The results demonstrated a significant improvement over traditional MFCC-based baseline, emphasizing the potential of SSL embeddings even in noisy conditions. To further generalize the KWS framework, the experiments were repeated for an additional CMU dataset. Overall the results highlight the significant contribution of SSL features in enhancing Zero-Shot KWS performance for children's speech, effectively addressing the challenges associated with the distinct characteristics of child speakers.


【6】Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
标题:分层SSL功能能否提高儿童语音的Zero-ShotASB性能?
链接:https://arxiv.org/abs/2508.21225

作者:Abhijit Sinha, Hemant Kumar Kathania, Sudarsana Reddy Kadiri, Shrikanth Narayanan
备注:Accepted
摘要:自动语音识别(ASR)系统往往很难准确地处理儿童的语音,由于其独特的和高度可变的声学和语言特征。虽然自我监督学习(SSL)模型的最新进展极大地增强了成人语音的转录,但准确地转录儿童语音仍然是一个重大挑战。这项研究调查了从最先进的SSL预训练模型中提取的逐层特征的有效性-特别是Wav 2 Vec 2,HuBERT,Data 2 Vec和WavLM,以提高zero-shot场景中儿童语音的ASR性能。对从这些模型中提取的特征进行了详细的分析,并使用Kaldi工具包将其集成到简化的基于DNN的ASR系统中。该分析确定了在zero-shot场景中增强儿童语音ASR性能的最有效层,其中WSJCAM 0成人语音用于训练,PFSTAR儿童语音用于测试。实验结果表明,Wav 2 Vec 2模型的第22层实现了5.15%的最低字错误率(WER),表示比使用Wav 2 Vec 2的直接zero-shot解码(WER为10.65%)相对改善了51.64%。此外,年龄组分析表明,随着年龄的增长,性能得到了一致的改善,即使在使用SSL功能的年轻年龄组中也观察到了显着的收益。对CMU Kids数据集的进一步实验证实了类似的趋势,突出了所提出方法的普遍性。
摘要:Automatic Speech Recognition (ASR) systems often struggle to accurately process children's speech due to its distinct and highly variable acoustic and linguistic characteristics. While recent advancements in self-supervised learning (SSL) models have greatly enhanced the transcription of adult speech, accurately transcribing children's speech remains a significant challenge. This study investigates the effectiveness of layer-wise features extracted from state-of-the-art SSL pre-trained models - specifically, Wav2Vec2, HuBERT, Data2Vec, and WavLM in improving the performance of ASR for children's speech in zero-shot scenarios. A detailed analysis of features extracted from these models was conducted, integrating them into a simplified DNN-based ASR system using the Kaldi toolkit. The analysis identified the most effective layers for enhancing ASR performance on children's speech in a zero-shot scenario, where WSJCAM0 adult speech was used for training and PFSTAR children speech for testing. Experimental results indicated that Layer 22 of the Wav2Vec2 model achieved the lowest Word Error Rate (WER) of 5.15%, representing a 51.64% relative improvement over the direct zero-shot decoding using Wav2Vec2 (WER of 10.65%). Additionally, age group-wise analysis demonstrated consistent performance improvements with increasing age, along with significant gains observed even in younger age groups using the SSL features. Further experiments on the CMU Kids dataset confirmed similar trends, highlighting the generalizability of the proposed approach.


eess.AS音频处理


【1】Towards Improved Speech Recognition through Optimized Synthetic Data Generation
标题:通过优化合成数据生成改进语音识别
链接:https://arxiv.org/abs/2508.21631

作者:Yanis Perrin, Gilles Boulianne
备注:12 pages, 3 figures
摘要:语音识别模型的监督训练需要访问转录的音频数据,由于保密性问题,这通常是不可能的。我们解决这个问题的方法是使用具有语音克隆功能的最先进的文本到语音模型从纯文本语料库生成合成音频。我们的目标是实现与在真实数据上训练的模型相当的自动语音识别(ASR)性能。我们探索了通过微调、过滤和评估优化合成数据生成的方法,以及它用于训练端到端编码器-解码器ASR模型的方法。实验使用两个数据集的自发,会话语音在魁北克法语。我们表明,改进数据生成导致在合成数据上训练的最终ASR系统的大幅改进。
摘要:Supervised training of speech recognition models requires access to transcribed audio data, which often is not possible due to confidentiality issues. Our approach to this problem is to generate synthetic audio from a text-only corpus using a state-of-the-art text-to-speech model with voice cloning capabilities. Our goal is to achieve automatic speech recognition (ASR) performance comparable to models trained on real data. We explore ways to optimize synthetic data generation through finetuning, filtering and evaluation, and its use for training an end-to-end encoder-decoder ASR model. Experiments were conducted using two datasets of spontaneous, conversational speech in Qu\'ebec French. We show that improving data generation leads to large improvements in the final ASR system trained on synthetic data.


【2】Fundamentals of Data-Driven Approaches to Acoustic Signal Detection, Filtering, and Transformation
标题:声信号检测、过滤和转换的数据驱动方法的基础
链接:https://arxiv.org/abs/2508.21470

作者:Chao Pan
摘要:近几十年来,由于各种应用需求,信号处理领域迅速发展,产生了丰富的科学问题和研究领域。信号的形式、形成机制和信息提取方法因应用而异,导致信号处理技术多种多样。常见的技术可以分为三种类型:转换、检测和过滤。信号变换将信号从其原始域转换到更合适的目标域进行分析;信号检测旨在识别信号中相关信息的存在及其特定时间和位置;信号滤波侧重于从观测信号中提取或分离感兴趣的源信号。在声学信号处理中,技术包括声源定位,声音事件检测,声纹提取和识别,降噪和源分离,应用于语音通信,语音交互,智能医疗保健和工业诊断。最近,深度学习技术的进步已经将声学信号处理的方法从知识驱动转变为数据驱动的方法,从而产生了重大的研究成果。本文旨在系统总结数据驱动的声信号处理原理和方法,为学术探索和实际应用提供一个全面的理解框架。
摘要:In recent decades, the field of signal processing has rapidly evolved due to diverse application demands, leading to a rich array of scientific questions and research areas. The forms of signals, their formation mechanisms, and the information extraction methods vary by application, resulting in diverse signal processing techniques. Common techniques can be categorized into three types: transformation, detection, and filtering. Signal transformation converts signals from their original domain to a more suitable target domain for analysis; signal detection aims to identify the existence of relevant information within a signal and its specific time and location; and signal filtering focuses on extracting or separating source signals of interest from observed signals. In acoustic signal processing, techniques include sound source localization, sound event detection, voiceprint extraction and recognition, noise reduction, and source separation, with applications in speech communication, voice interaction, smart healthcare, and industrial diagnostics. Recently, the advancement of deep learning technologies has shifted methodologies in acoustic signal processing from knowledge-driven to data-driven approaches, leading to significant research outcomes. This paper aims to systematically summarize the principles and methods of data-driven acoustic signal processing, providing a comprehensive understanding framework for academic exploration and practical applications.


【3】Cochleagram-based Noise Adapted Speaker Identification System for Distorted Speech
标题:失真语音的基于高斯图的噪音自适应说话人识别系统
链接:https://arxiv.org/abs/2508.21347

作者:Sabbir Ahmed, Nursadul Mamun, Md Azad Hossain
备注:10 pages, 10 figures, 4 tables
摘要:说话人识别是指使用一个人的声音从已知说话人的集合中识别一个人的过程。环境噪声、混响和失真使得说话人自动识别的任务变得具有挑战性,因为提取的特征变得退化,从而影响说话人识别(SID)系统的性能。本文提出了一种在噪声、失配、混响和失真环境下的鲁棒噪声自适应SID系统。该方法利用一种称为耳蜗图的听觉特征来提取说话人特征,从而识别说话人。一个128美元的通道伽玛滤波器组的频率范围从50美元到8000美元赫兹被用来产生2-D耳蜗图。宽带以及窄带噪声与干净的语音一起使用,以获得在不同水平的信噪比(SNR)的噪声耳蜗图。然后将信噪比仅为$-5$ dB的干净和有噪声的耳蜗图馈送到卷积神经网络(CNN)中以构建说话人模型,以执行SID,这被称为噪声自适应说话人模型(NASM)。NASM使用一定的噪声进行训练,然后使用干净的和各种类型的噪声进行评估。此外,所提出的系统的鲁棒性进行了测试,使用混响以及失真的测试数据。所提出的系统的性能表现出了可衡量的准确性改善现有的神经图为基础的SID系统。
摘要:Speaker Identification refers to the process of identifying a person using one's voice from a collection of known speakers. Environmental noise, reverberation and distortion make the task of automatic speaker identification challenging as extracted features get degraded thus affecting the performance of the speaker identification (SID) system. This paper proposes a robust noise adapted SID system under noisy, mismatched, reverberated and distorted environments. This method utilizes an auditory features called cochleagram to extract speaker characteristics and thus identify the speaker. A $128$ channel gammatone filterbank with a frequency range from $50$ to $8000$ Hz was used to generate 2-D cochleagrams. Wideband as well as narrowband noises were used along with clean speech to obtain noisy cochleagrams at various levels of signal to noise ratio (SNR). Both clean and noisy cochleagrams of only $-5$ dB SNR were then fed into a convolutional neural network (CNN) to build a speaker model in order to perform SID which is referred as noise adapted speaker model (NASM). The NASM was trained using a certain noise and then was evaluated using clean and various types of noises. Moreover, the robustness of the proposed system was tested using reverberated as well as distorted test data. Performance of the proposed system showed a measurable accuracy improvement over existing neurogram based SID system.


【4】Zero-Shot KWS for Children's Speech using Layer-Wise Features from SSL Models
标题:使用SSL模型的分层功能针对儿童语音的Zero-ShotKWS
链接:https://arxiv.org/abs/2508.21248

作者:Subham Kutum, Abhijit Sinha, Hemant Kumar Kathania, Sudarsana Reddy Kadiri, Mahesh Chandra Govil
备注:Accepted
摘要:已经提出了许多方法来增强成人语音中的关键词识别(KWS),但儿童语音由于其独特的声学和语言特征,对KWS系统提出了独特的挑战。本文介绍了一种零触发KWS方法(zero-shot KWS),该方法利用了最先进的自监督学习(SSL)模型,包括Wave 2 Vec 2、HuBERT和Data 2 Vec。从这些SSL模型中逐层提取特征,并用于训练基于Kaldi的DNN KWS系统。WSJCAM 0成人语音数据集用于训练,而PFSTAR儿童语音数据集用于测试,证明了我们的方法的zero-shot能力。我们的方法在儿童语音的所有关键词集上都取得了最先进的结果。值得注意的是,Wav 2 Vec 2模型,特别是第22层,表现最好,对于一组30个关键字,ATWV得分为0.691,MTWV得分为0.7003,虚警概率和未命中概率分别为0.0164和0.0547。此外,针对不同年龄段儿童的绩效评估证实了该系统的有效性。为了评估系统对噪声的鲁棒性,使用性能最好的Wav 2 Vec 2模型的性能最好的层进行了额外的实验。结果表明,与传统的基于MFCC的基线相比,有了显着的改进,强调了SSL嵌入的潜力,即使在嘈杂的条件下。为了进一步推广KWS框架,针对额外的CMU数据集重复实验。总体而言,结果突出了SSL功能在增强儿童语音的Zero-Shot KWS性能方面的重大贡献,有效地解决了与儿童说话者的独特特征相关的挑战。
摘要:Numerous methods have been proposed to enhance Keyword Spotting (KWS) in adult speech, but children's speech presents unique challenges for KWS systems due to its distinct acoustic and linguistic characteristics. This paper introduces a zero-shot KWS approach that leverages state-of-the-art self-supervised learning (SSL) models, including Wav2Vec2, HuBERT and Data2Vec. Features are extracted layer-wise from these SSL models and used to train a Kaldi-based DNN KWS system. The WSJCAM0 adult speech dataset was used for training, while the PFSTAR children's speech dataset was used for testing, demonstrating the zero-shot capability of our method. Our approach achieved state-of-the-art results across all keyword sets for children's speech. Notably, the Wav2Vec2 model, particularly layer 22, performed the best, delivering an ATWV score of 0.691, a MTWV score of 0.7003 and probability of false alarm and probability of miss of 0.0164 and 0.0547 respectively, for a set of 30 keywords. Furthermore, age-specific performance evaluation confirmed the system's effectiveness across different age groups of children. To assess the system's robustness against noise, additional experiments were conducted using the best-performing layer of the best-performing Wav2Vec2 model. The results demonstrated a significant improvement over traditional MFCC-based baseline, emphasizing the potential of SSL embeddings even in noisy conditions. To further generalize the KWS framework, the experiments were repeated for an additional CMU dataset. Overall the results highlight the significant contribution of SSL features in enhancing Zero-Shot KWS performance for children's speech, effectively addressing the challenges associated with the distinct characteristics of child speakers.


【5】Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
标题:分层SSL功能能否提高儿童语音的Zero-ShotASB性能?
链接:https://arxiv.org/abs/2508.21225

作者:Abhijit Sinha, Hemant Kumar Kathania, Sudarsana Reddy Kadiri, Shrikanth Narayanan
备注:Accepted
摘要:自动语音识别(ASR)系统往往很难准确地处理儿童的语音,由于其独特的和高度可变的声学和语言特征。虽然自我监督学习(SSL)模型的最新进展极大地增强了成人语音的转录,但准确转录儿童语音仍然是一个重大挑战。这项研究调查了从最先进的SSL预训练模型中提取的逐层特征的有效性-特别是Wav 2 Vec 2,HuBERT,Data 2 Vec和WavLM,以提高zero-shot场景中儿童语音的ASR性能。对从这些模型中提取的特征进行了详细的分析,并使用Kaldi工具包将其集成到简化的基于DNN的ASR系统中。该分析确定了在zero-shot场景中增强儿童语音ASR性能的最有效层,其中WSJCAM 0成人语音用于训练,PFSTAR儿童语音用于测试。实验结果表明,Wav 2 Vec 2模型的第22层实现了5.15%的最低字错误率(WER),表示比使用Wav 2 Vec 2的直接zero-shot解码(WER为10.65%)相对改善了51.64%。此外,年龄组分析表明,随着年龄的增长,性能得到了一致的改善,即使在使用SSL功能的年轻年龄组中也观察到了显着的收益。对CMU Kids数据集的进一步实验证实了类似的趋势,突出了所提出方法的普遍性。
摘要:Automatic Speech Recognition (ASR) systems often struggle to accurately process children's speech due to its distinct and highly variable acoustic and linguistic characteristics. While recent advancements in self-supervised learning (SSL) models have greatly enhanced the transcription of adult speech, accurately transcribing children's speech remains a significant challenge. This study investigates the effectiveness of layer-wise features extracted from state-of-the-art SSL pre-trained models - specifically, Wav2Vec2, HuBERT, Data2Vec, and WavLM in improving the performance of ASR for children's speech in zero-shot scenarios. A detailed analysis of features extracted from these models was conducted, integrating them into a simplified DNN-based ASR system using the Kaldi toolkit. The analysis identified the most effective layers for enhancing ASR performance on children's speech in a zero-shot scenario, where WSJCAM0 adult speech was used for training and PFSTAR children speech for testing. Experimental results indicated that Layer 22 of the Wav2Vec2 model achieved the lowest Word Error Rate (WER) of 5.15%, representing a 51.64% relative improvement over the direct zero-shot decoding using Wav2Vec2 (WER of 10.65%). Additionally, age group-wise analysis demonstrated consistent performance improvements with increasing age, along with significant gains observed even in younger age groups using the SSL features. Further experiments on the CMU Kids dataset confirmed similar trends, highlighting the generalizability of the proposed approach.


【6】Benchmarking Large Pretrained Multilingual Models on Québec French Speech Recognition
标题:魁北克法语语音识别上的大型预训练多语言模型基准
链接:https://arxiv.org/abs/2508.21193

作者:Coralie Serrand, Gilles Boulianne, Amira Morsli
备注:11 pages, 3 figures
摘要:我们评估了大型预训练的多语言语音识别模型在加拿大魁北克地区法语口语中的速度,单词错误率和语义准确性方面的性能。为此,我们建立了一个基准和评估管道的基础上CommissionsQc数据集,在魁北克最近举行的公开调查期间记录的自发对话语料库。这些模型在FLEURS或CommonVoice等知名基准测试上发布的结果并不能很好地预测我们在CommissionsQC上观察到的性能。我们的研究结果应该感兴趣的从业者有兴趣在现实条件下或区域语言品种的语音应用程序。
摘要:We evaluate the performance of large pretrained multilingual speech recognition models on a regional variety of French spoken in Qu\'ebec, Canada, in terms of speed, word error rate and semantic accuracy. To this end we build a benchmark and evaluation pipeline based on the CommissionsQc datasets, a corpus of spontaneous conversations recorded during public inquiries recently held in Qu\'ebec. Published results for these models on well-known benchmarks such as FLEURS or CommonVoice are not good predictors of the performance we observe on CommissionsQC. Our results should be of interest for practitioners interested in building speech applications for realistic conditions or regional language varieties.


【7】WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration
标题:WaveLLDM:用于语音增强和恢复的轻量级潜在扩散模型的设计和开发
链接:https://arxiv.org/abs/2508.21153

作者:Kevin Putra Santoso, Rizka Wakhidatus Sholikah, Raden Venantius Hari Ginardi
摘要:高质量的音频在广泛的应用中至关重要,包括在线通信、虚拟助手和多媒体行业。然而,由噪声、压缩和传输伪影引起的降级仍然是一个主要挑战。虽然扩散模型已被证明对音频恢复是有效的,但它们通常需要大量的计算资源,并且难以处理较长的缺失片段。本研究介绍了WaveLLDM(Wave Lightweight Latent Diffusion Model),这是一种集成了高效神经音频编解码器和潜在扩散的架构,用于音频恢复和去噪。与在时域或频谱域中操作的传统方法不同,WaveLLDM在压缩的潜在空间中处理音频,从而在保持重建质量的同时降低计算复杂度。在Voicebank+DEMAND测试集上的经验评估表明,WaveLLDM实现了准确的频谱重建,具有低对数频谱距离(LSD)分数(0.48至0.60)和对未知数据的良好适应性。然而,与最先进的方法相比,它在感知质量和语音清晰度方面仍然表现不佳,WB-PESQ得分范围为1.62至1.71,STOI得分在0.76至0.78之间。这些限制归因于次优的架构调优、缺乏微调以及训练持续时间不足。尽管如此,结合神经音频编解码器和潜在扩散模型的灵活架构为未来的开发提供了坚实的基础。
摘要:High-quality audio is essential in a wide range of applications, including online communication, virtual assistants, and the multimedia industry. However, degradation caused by noise, compression, and transmission artifacts remains a major challenge. While diffusion models have proven effective for audio restoration, they typically require significant computational resources and struggle to handle longer missing segments. This study introduces WaveLLDM (Wave Lightweight Latent Diffusion Model), an architecture that integrates an efficient neural audio codec with latent diffusion for audio restoration and denoising. Unlike conventional approaches that operate in the time or spectral domain, WaveLLDM processes audio in a compressed latent space, reducing computational complexity while preserving reconstruction quality. Empirical evaluations on the Voicebank+DEMAND test set demonstrate that WaveLLDM achieves accurate spectral reconstruction with low Log-Spectral Distance (LSD) scores (0.48 to 0.60) and good adaptability to unseen data. However, it still underperforms compared to state-of-the-art methods in terms of perceptual quality and speech clarity, with WB-PESQ scores ranging from 1.62 to 1.71 and STOI scores between 0.76 and 0.78. These limitations are attributed to suboptimal architectural tuning, the absence of fine-tuning, and insufficient training duration. Nevertheless, the flexible architecture that combines a neural audio codec and latent diffusion model provides a strong foundation for future development.


机器翻译由腾讯交互翻译提供,仅供参考