今日论文合集:cs.SD语音10篇,eess.AS音频处理4篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】When Pamplona sounds different: the soundscape transformation of San Fermin through intelligent acoustic sensors and a sound repository
标题:当潘普洛纳听起来不同时:通过智能声学传感器和声音库对圣佛明的声景进行改造
链接:https://arxiv.org/abs/2512.17740

作者:Amaia Sagasti,Frederic Font
备注:46 pages, 27 figures
摘要:这项研究提出了一个用例的低成本声学智能传感器网络部署在潘普洛纳市分析圣佛明节期间城市声景的变化。这些传感器安装在城市的不同区域,在事件发生之前,期间和之后,捕获连续的声学数据。我们的分析揭示了节日期间城市的声音环境发生了重大变化:整体声压级显着增加,声景模式发生变化,声学景观成为与人类活动相关的声音主导。这些发现突出了分布式智能声学监测系统在表征城市声景的时间动态方面的潜力,并强调了圣佛明大规模事件如何彻底重塑潘普洛纳市的整体声学动态。此外,为了补充客观测量,我们还精心制作了一系列真实的圣佛明录音并向公众开放,以保护音乐节独特的声音遗产。
摘要:This study presents a use-case of a network of low-cost acoustic smart sensors deployed in the city of Pamplona to analyse changes in the urban soundscape during the San Fermin Festival. The sensors were installed in different areas of the city before, during, and after the event, capturing continuous acoustic data. Our analysis reveals a significant transformation in the city's sonic environment during the festive period: overall sound pressure levels increase significantly, soundscape patterns change, and the acoustic landscape becomes dominated by sounds associated with human activity. These findings highlight the potential of distributed smart acoustic monitoring systems to characterize the temporal dynamics of urban soundscapes and underscore how the large-scale event of San Fermin drastically reshapes the overall acoustic dynamics of the city of Pamplona. Additionally, to complement the objective measurements, a curated collection of real San Fermin sound recordings has been created and made publicly available, preserving the festival's unique sonic heritage.


【2】When De-noising Hurts: A Systematic Study of Speech Enhancement Effects on Modern Medical ASR Systems
标题:去噪何时会造成伤害:现代医疗ASB系统语音增强效果的系统研究
链接:https://arxiv.org/abs/2512.17562

作者:Sujal Chondhekar,Vasanth Murukuri,Rushabh Vasani,Sanika Goyal,Rajshree Badami,Anushree Rana,Sanjana SN,Karthik Pandia,Sulabh Katiyar,Neha Jagadeesh,Sankalp Gulati
备注:Technical Report
摘要:语音增强方法被普遍认为可以提高噪声环境中自动语音识别(ASR)的性能。然而,对于在多样化、有噪数据上训练的现代大规模ASR模型来说,这些技术的有效性不能被视为理所当然。我们在四个最先进的ASR系统上对MetricGAN加语音库去噪进行了系统评估:OpenAI Whisper,NVIDIA Parakeet,Google Gemini Flash 2.0,Parrotlet-a,在九种噪声条件下使用500个医疗语音记录。ASR性能是使用语义WER(semWER)来衡量的,这是一个归一化的单词错误率(WER)指标,用于解释特定于域的标准化。我们的研究结果揭示了一个违反直觉的发现:语音增强预处理降低了所有噪声条件和模型的ASR性能。在所有40种测试配置(4种型号x10种条件)中,原始嘈杂音频的semWER低于增强音频,衰减范围为1.1%至46.6%的绝对semWER增加。这些研究结果表明,现代ASR模型具有足够的内部噪声鲁棒性,传统的语音增强可能会删除声学特征的关键ASR。对于在嘈杂的临床环境中部署医疗抄写系统的从业者来说,我们的研究结果表明,使用降噪技术对音频进行预处理不仅可能在计算上浪费,而且可能对转录准确性有害。
摘要:Speech enhancement methods are commonly believed to improve the performance of automatic speech recognition (ASR) in noisy environments. However, the effectiveness of these techniques cannot be taken for granted in the case of modern large-scale ASR models trained on diverse, noisy data. We present a systematic evaluation of MetricGAN-plus-voicebank denoising on four state-of-the-art ASR systems: OpenAI Whisper, NVIDIA Parakeet, Google Gemini Flash 2.0, Parrotlet-a using 500 medical speech recordings under nine noise conditions. ASR performance is measured using semantic WER (semWER), a normalized word error rate (WER) metric accounting for domain-specific normalizations. Our results reveal a counterintuitive finding: speech enhancement preprocessing degrades ASR performance across all noise conditions and models. Original noisy audio achieves lower semWER than enhanced audio in all 40 tested configurations (4 models x 10 conditions), with degradations ranging from 1.1% to 46.6% absolute semWER increase. These findings suggest that modern ASR models possess sufficient internal noise robustness and that traditional speech enhancement may remove acoustic features critical for ASR. For practitioners deploying medical scribe systems in noisy clinical environments, our results indicate that preprocessing audio with noise reduction techniques might not just be computationally wasteful but also be potentially harmful to the transcription accuracy.


【3】Training Text-to-Speech Model with Purely Synthetic Data: Feasibility, Sensitivity, and Generalization Capability
标题:用活泼的合成数据训练文本到语音模型:可行性、敏感性和概括能力
链接:https://arxiv.org/abs/2512.17356

作者:Tingxiao Zhou,Leying Zhang,Zhengyang Chen,Yanmin Qian
备注:14 pages, 5 figures, received by National Conference on Man-Machine Speech Communication (NCMMSC2025)
摘要:合成数据在文语转换(TTS)模型训练中的潜力越来越受到关注,但其合理性和有效性需要系统验证。在这项研究中,我们系统地研究了使用纯合成数据进行TTS训练的可行性,并探讨了各种因素-包括文本丰富性,说话者多样性,噪声水平和说话风格-如何影响模型性能。我们的实验表明,增加扬声器和文本的多样性显着提高合成质量和鲁棒性。具有最小噪声的更干净的训练数据进一步提高了性能。此外,我们发现,标准的说话风格有利于更有效的模型学习。我们的实验表明,由于没有真实世界的缺陷和噪声,在类似条件下,在合成数据上训练的模型有很大的潜力超过在真实数据上训练的模型。
摘要:The potential of synthetic data in text-to-speech (TTS) model training has gained increasing attention, yet its rationality and effectiveness require systematic validation. In this study, we systematically investigate the feasibility of using purely synthetic data for TTS training and explore how various factors--including text richness, speaker diversity, noise levels, and speaking styles--affect model performance. Our experiments reveal that increasing speaker and text diversity significantly enhances synthesis quality and robustness. Cleaner training data with minimal noise further improves performance. Moreover, we find that standard speaking styles facilitate more effective model learning. Our experiments indicate that models trained on synthetic data have great potential to outperform those trained on real data under similar conditions, due to the absence of real-world imperfections and noise.


【4】Robust TTS Training via Self-Purifying Flow Matching for the WildSpoof 2026 TTS Track
标题:通过WildSpoof 2026 TTC赛道的自我净化流匹配进行稳健的TTC训练
链接:https://arxiv.org/abs/2512.17293

作者:June Young Yi,Hyeongju Kim,Juheon Lee
备注:2 pages, preprint, This work has been submitted to the IEEE for possible publication. Submitted to ICASSP 2026 SPGC (WildSpoof Challenge, TTS track)
摘要:本文提出了一个轻量级的文本到语音(TTS)系统开发的WildSpoof挑战TTS轨道。我们的方法微调了最近发布的开放权重TTS模型,\textit{Supertonic}\footnote{\url{https://github.com/supertone-inc/supertonic}},具有自净化流匹配(SPFM),以实现对野生语音的鲁棒适应。SPFM通过比较每个样本上的有条件和无条件流匹配损失来减轻标签噪声,将可疑的文本-语音对路由到无条件训练,同时仍然利用它们的声学信息。由此产生的模型在所有参与团队中实现了最低的单词错误率(WER),同时在UTMOS和DNSMOS等感知指标中排名第二。这些发现表明,像Supertonic这样的高效开放权重架构,在与SPFM等显式噪声处理机制相结合时,可以有效地适应各种真实世界的语音条件。
摘要:This paper presents a lightweight text-to-speech (TTS) system developed for the WildSpoof Challenge TTS Track. Our approach fine-tunes the recently released open-weight TTS model, \textit{Supertonic}\footnote{\url{https://github.com/supertone-inc/supertonic}}, with Self-Purifying Flow Matching (SPFM) to enable robust adaptation to in-the-wild speech. SPFM mitigates label noise by comparing conditional and unconditional flow matching losses on each sample, routing suspicious text--speech pairs to unconditional training while still leveraging their acoustic information. The resulting model achieves the lowest Word Error Rate (WER) among all participating teams, while ranking second in perceptual metrics such as UTMOS and DNSMOS. These findings demonstrate that efficient, open-weight architectures like Supertonic can be effectively adapted to diverse real-world speech conditions when combined with explicit noise-handling mechanisms such as SPFM.


【5】LibriVAD: A Scalable Open Dataset with Deep Learning Benchmarks for Voice Activity Detection
标题:LibriVAR:具有深度学习基准的可扩展开放数据集,用于语音活动检测
链接:https://arxiv.org/abs/2512.17281

作者:Ioannis Stylianou,Achintya kr. Sarkar,Nauman Dawalatabad,James Glass,Zheng-Hua Tan
摘要:鲁棒的语音活动检测(VAD)仍然是一项具有挑战性的任务,特别是在嘈杂,多样化和不可见的声学条件下。除了算法开发之外,推进VAD研究的一个关键限制是缺乏大规模,系统控制和公开可用的数据集。为了解决这个问题,我们引入了LibriVAD -一个可扩展的开源数据集,来自LibriSpeech,并增加了各种真实世界和合成噪声源。LibriVAD能够系统地控制语音噪声比、静音语音比(SSR)和噪声多样性,并以三种大小(15 GB、150 GB和1.5 TB)发布,其中两种变体(LibriVAD-NonConcat和LibriVAD-Concat)支持不同的实验设置。我们对多个特征模型组合进行基准测试,包括波形、Mel频率倒谱系数(MFCC)和Gammatone滤波器组倒谱系数,并介绍了用于VAD的Vision Transformer(ViT)架构。我们的实验表明,具有MFCC特征的ViT在可见、不可见和分布外(OOD)条件下的性能始终优于已建立的VAD模型,如增强型深度神经网络和卷积长短期记忆深度神经网络,包括对真实世界VOiCES数据集的评估。我们进一步分析了数据集大小和SSR对模型泛化的影响,实验表明,在OOD条件下,扩大数据集大小和平衡SSR显著且一致地增强了VAD性能。所有数据集、训练模型和代码都公开发布,以促进可重复性并加速VAD研究的进展。
摘要:Robust Voice Activity Detection (VAD) remains a challenging task, especially under noisy, diverse, and unseen acoustic conditions. Beyond algorithmic development, a key limitation in advancing VAD research is the lack of large-scale, systematically controlled, and publicly available datasets. To address this, we introduce LibriVAD - a scalable open-source dataset derived from LibriSpeech and augmented with diverse real-world and synthetic noise sources. LibriVAD enables systematic control over speech-to-noise ratio, silence-to-speech ratio (SSR), and noise diversity, and is released in three sizes (15 GB, 150 GB, and 1.5 TB) with two variants (LibriVAD-NonConcat and LibriVAD-Concat) to support different experimental setups. We benchmark multiple feature-model combinations, including waveform, Mel-Frequency Cepstral Coefficients (MFCC), and Gammatone filter bank cepstral coefficients, and introduce the Vision Transformer (ViT) architecture for VAD. Our experiments show that ViT with MFCC features consistently outperforms established VAD models such as boosted deep neural network and convolutional long short-term memory deep neural network across seen, unseen, and out-of-distribution (OOD) conditions, including evaluation on the real-world VOiCES dataset. We further analyze the impact of dataset size and SSR on model generalization, experimentally showing that scaling up dataset size and balancing SSR noticeably and consistently enhance VAD performance under OOD conditions. All datasets, trained models, and code are publicly released to foster reproducibility and accelerate progress in VAD research.


【6】Do Foundational Audio Encoders Understand Music Structure?
标题:基础音频编码器了解音乐结构吗?
链接:https://arxiv.org/abs/2512.17209

作者:Keisuke Toyama,Zhi Zhong,Akira Takahashi,Shusuke Takahashi,Yuki Mitsufuji
摘要:在音乐信息检索(MIR)研究中,使用预训练的基础音频编码器(FAE)最近已成为一种趋势。在大量音乐和音频数据上预训练的FAE已被证明可以提高MIR任务的性能,例如音乐标记和自动音乐转录。然而,他们的使用音乐结构分析(MSA)仍然是探索不足。虽然有许多开源FAE模型可用,但只有一小部分已被用于MSA,并且学习方法,训练数据和模型上下文长度等因素对MSA性能的影响仍然不清楚。在这项研究中,我们对11种FAE进行了综合实验,以研究这些因素如何影响MSA性能。我们的研究结果表明,使用自监督学习的FAE与音乐数据上的掩蔽语言建模对MSA特别有效。这些发现为MSA的未来研究铺平了道路。
摘要:In music information retrieval (MIR) research, the use of pretrained foundational audio encoders (FAEs) has recently become a trend. FAEs pretrained on large amounts of music and audio data have been shown to improve performance on MIR tasks such as music tagging and automatic music transcription. However, their use for music structure analysis (MSA) remains underexplored. Although many open-source FAE models are available, only a small subset has been examined for MSA, and the impact of factors such as learning methods, training data, and model context length on MSA performance remains unclear. In this study, we conduct comprehensive experiments on 11 types of FAEs to investigate how these factors affect MSA performance. Our results demonstrate that FAEs using selfsupervised learning with masked language modeling on music data are particularly effective for MSA. These findings pave the way for future research in MSA.


【7】InstructDubber: Instruction-based Alignment for Zero-shot Movie Dubbing
标题:DirectionDubber:基于指令的Zero-Shot电影配音对齐
链接:https://arxiv.org/abs/2512.17154

作者:Zhedong Zhang,Liang Li,Gaoxiang Cong,Chunshan Liu,Yuhan Gao,Xiaowan Wang,Tao Gu,Yuankai Qi
备注:Accepted by AAAI2026
摘要:电影配音试图使用特定的声音从给定的脚本中合成语音,同时确保准确的嘴唇同步和情感韵律与角色的视觉表现一致。然而,现有的基于视觉特征的对齐方法面临两个关键限制:(1)它们依赖于复杂的手工视觉预处理管道,包括面部标志检测和特征提取;(2)它们对看不见的视觉域的泛化能力很差,通常导致对齐和配音质量下降。为了解决这些问题,我们提出了InstructDubber,这是一种新型的基于指令的对齐配音方法,用于鲁棒的域内和零镜头电影配音(zero-shot movie dubbing)。具体来说,我们首先将视频,脚本和相应的提示输入到多模态大型语言模型中,以生成关于视频中描述的语速和情感状态的自然语言配音指令,这对视觉域变化具有鲁棒性。其次,我们设计了一个指令时长提取模块,从语速指令中挖掘区分时长线索,预测唇对齐音素级发音时长。第三,对于情感韵律对齐,我们设计了一个指令情感校准模块,它微调基于LLM的指令分析器使用地面实况配音情感作为监督和预测韵律的基础上校准的情感分析。最后,将预测的持续时间和韵律与脚本一起馈送到音频解码器中以生成视频对齐的配音。在三个主要基准测试上进行的大量实验表明,在域内和zero-shot场景中,InstructDubber的性能优于最先进的方法。
摘要:Movie dubbing seeks to synthesize speech from a given script using a specific voice, while ensuring accurate lip synchronization and emotion-prosody alignment with the character's visual performance. However, existing alignment approaches based on visual features face two key limitations: (1)they rely on complex, handcrafted visual preprocessing pipelines, including facial landmark detection and feature extraction; and (2) they generalize poorly to unseen visual domains, often resulting in degraded alignment and dubbing quality. To address these issues, we propose InstructDubber, a novel instruction-based alignment dubbing method for both robust in-domain and zero-shot movie dubbing. Specifically, we first feed the video, script, and corresponding prompts into a multimodal large language model to generate natural language dubbing instructions regarding the speaking rate and emotion state depicted in the video, which is robust to visual domain variations. Second, we design an instructed duration distilling module to mine discriminative duration cues from speaking rate instructions to predict lip-aligned phoneme-level pronunciation duration. Third, for emotion-prosody alignment, we devise an instructed emotion calibrating module, which finetunes an LLM-based instruction analyzer using ground truth dubbing emotion as supervision and predicts prosody based on the calibrated emotion analysis. Finally, the predicted duration and prosody, together with the script, are fed into the audio decoder to generate video-aligned dubbing. Extensive experiments on three major benchmarks demonstrate that InstructDubber outperforms state-of-the-art approaches across both in-domain and zero-shot scenarios.


【8】Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
标题:Speech-FT:合并预训练和微调的语音表示模型以实现跨任务概括
链接:https://arxiv.org/abs/2502.12672

作者:Tzu-Quan Lin,Wei-Ping Huang,Hao Tang,Hung-yi Lee
备注:Published in IEEE Transactions on Audio, Speech, and Language Processing (TASLP). Model and code available at: https://github.com/nervjack2/Speech-FT
摘要:微调语音表示模型可以提高特定任务的性能,但往往会损害其跨任务的泛化能力。这种退化通常是由表示的过度变化引起的,这使得很难保留在预训练期间学到的信息。现有的方法,例如在微调期间正则化权重变化,可能无法与预训练模型保持足够高的特征相似性,因此可能会失去跨任务泛化。为了解决这个问题,我们提出了Speech-FT,这是一种新的两阶段微调框架,旨在保持跨任务的泛化,同时受益于微调。Speech-FT首先应用专门设计的微调,以减少代表性漂移,然后使用预训练模型进行权重空间插值,以恢复跨任务泛化。在HuBERT、wav 2 vec 2.0、DeCoAR 2.0和WavLM Base+上进行的大量实验表明,Speech-FT在各种监督、无监督和多任务微调场景下都能持续提高性能。此外,与明确约束权重变化的微调基线(如权重空间正则化和LoRA微调)相比,Speech-FT实现了卓越的跨任务泛化。我们的分析表明,与其他策略相比,Speech-FT与预训练模型保持更高的特征相似性,尽管允许更大的权重空间更新。值得注意的是,Speech-FT在SUPERB基准测试中取得了显著的改进。例如,在自动语音识别上微调HuBERT时,Speech-FT能够将电话错误率从5.17%降至3.94%,将单词错误率从6.38%降至5.75%,并将说话人识别准确率从81.86%提高到84.11%。Speech-FT提供了一个简单而强大的解决方案,用于在预训练后进一步细化语音表示模型。
摘要:Fine-tuning speech representation models can enhance performance on specific tasks but often compromises their cross-task generalization ability. This degradation is often caused by excessive changes in the representations, making it difficult to retain information learned during pre-training. Existing approaches, such as regularizing weight changes during fine-tuning, may fail to maintain sufficiently high feature similarity with the pre-trained model, and thus could possibly lose cross-task generalization. To address this issue, we propose Speech-FT, a novel two-stage fine-tuning framework designed to maintain cross-task generalization while benefiting from fine-tuning. Speech-FT first applies fine-tuning specifically designed to reduce representational drift, followed by weight-space interpolation with the pre-trained model to restore cross-task generalization. Extensive experiments on HuBERT, wav2vec 2.0, DeCoAR 2.0, and WavLM Base+ demonstrate that Speech-FT consistently improves performance across a wide range of supervised, unsupervised, and multitask fine-tuning scenarios. Moreover, Speech-FT achieves superior cross-task generalization compared to fine-tuning baselines that explicitly constrain weight changes, such as weight-space regularization and LoRA fine-tuning. Our analysis reveals that Speech-FT maintains higher feature similarity to the pre-trained model compared to alternative strategies, despite allowing larger weight-space updates. Notably, Speech-FT achieves significant improvements on the SUPERB benchmark. For example, when fine-tuning HuBERT on automatic speech recognition, Speech-FT is able to reduce phone error rate from 5.17% to 3.94%, lower word error rate from 6.38% to 5.75%, and increase speaker identification accuracy from 81.86% to 84.11%. Speech-FT provides a simple yet powerful solution for further refining speech representation models after pre-training.


【9】Review of MEMS Speakers for Audio Applications
标题:MEMS音频扬声器综述
链接:https://arxiv.org/abs/2512.17708

作者:Nils Wittek,Anton Melnikov,Bert Kaiser,André Zimmermann
备注:37 pages, 6 figures
摘要:微机电系统(MEMS)扬声器是传统音圈扬声器的紧凑、可扩展的替代品,通过精确的半导体制造有望改善音质。本文综述了MEMS扬声器的研究现状,包括基于超声脉冲和热声的声音产生,并按致动原理对MEMS扬声器进行了分类:电动,压电和静电。对1990-2025年性能指标的比较分析突出了具有直接空气位移的压电MEMS的主导地位,重点是小型化和效率。该评论概述了即将到来的研究挑战,并确定了实现全频谱音频性能的潜在候选人。对创新方法的关注可能会导致MEMS扬声器的宽带采用。
摘要:Microelectromechanical systems (MEMS) speakers are compact, scalable alternatives to traditional voice coil speakers, promising improved sound quality through precise semiconductor manufacturing. This review provides an overview of the research landscape, including ultrasound pulse-based and thermoacoustic sound generation, classifying MEMS speakers by actuation principle: electrodynamic, piezoelectric, and electrostatic. A comparative analysis of performance indicators from 1990-2025 highlights the dominance of piezoelectric MEMS with direct air displacement, focusing on miniaturization and efficiency. The review outlines upcoming research challenges and identifies potential candidates for achieving full-spectrum audio performance. A focus on innovative approaches could lead to wideband adoption of MEMS-only speakers.


【10】Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
标题:使用商业自动语音识别和多模式大语言模型进行发音障碍语音的Zero-Shot识别
链接:https://arxiv.org/abs/2512.17474

作者:Ali Alsayegh,Tariq Masood
摘要:基于语音的人机交互是访问智能系统的主要方式,但构音障碍患者由于识别性能差距而面临系统性排斥。虽然自动语音识别(ASR)在典型语音上实现了低于5%的单词错误率(WER),但对于构音障碍的说话者,性能会显着下降。多模态大型语言模型(MLLM)提供了利用上下文推理来补偿声学退化的潜力,但其zero-shot能力仍然没有被描述。这项研究评估了TORGO构音障碍语音语料库上的八种商业语音转文本服务:四种传统的ASR系统(AssemblyAI,Whisper large-v3,Deepgram Nova-3,Nova-3 Medical)和四种基于MLLM的系统(GPT-4 o,GPT-4 o Mini,Gemini 2.5 Pro,Gemini 2.5 Flash)。评估包括词汇准确性,语义保留和成本延迟权衡。结果表明,严重程度相关的退化:轻度构音障碍达到3-5%的WER接近典型的语音基准,而严重构音障碍超过49%的WER在所有系统。逐字转录提示产生特定于架构的效果:GPT-4 o实现了7.36个百分点的WER降低,所有测试扬声器都有一致的改善,而Gemini变体表现出退化。语义指标表明,尽管词汇错误率升高,但交际意图仍然部分可恢复。这些发现建立了经验基线,使基于证据的技术选择辅助语音接口部署。
摘要:Voice-based human-machine interaction is a primary modality for accessing intelligent systems, yet individuals with dysarthria face systematic exclusion due to recognition performance gaps. Whilst automatic speech recognition (ASR) achieves word error rates (WER) below 5% on typical speech, performance degrades dramatically for dysarthric speakers. Multimodal large language models (MLLMs) offer potential for leveraging contextual reasoning to compensate for acoustic degradation, yet their zero-shot capabilities remain uncharacterised. This study evaluates eight commercial speech-to-text services on the TORGO dysarthric speech corpus: four conventional ASR systems (AssemblyAI, Whisper large-v3, Deepgram Nova-3, Nova-3 Medical) and four MLLM-based systems (GPT-4o, GPT-4o Mini, Gemini 2.5 Pro, Gemini 2.5 Flash). Evaluation encompasses lexical accuracy, semantic preservation, and cost-latency trade-offs. Results demonstrate severity-dependent degradation: mild dysarthria achieves 3-5% WER approaching typical-speech benchmarks, whilst severe dysarthria exceeds 49% WER across all systems. A verbatim-transcription prompt yields architecture-specific effects: GPT-4o achieves 7.36 percentage point WER reduction with consistent improvement across all tested speakers, whilst Gemini variants exhibit degradation. Semantic metrics indicate that communicative intent remains partially recoverable despite elevated lexical error rates. These findings establish empirical baselines enabling evidence-based technology selection for assistive voice interface deployment.


eess.AS音频处理


【1】Review of MEMS Speakers for Audio Applications
标题:MEMS音频扬声器综述
链接:https://arxiv.org/abs/2512.17708

作者:Nils Wittek,Anton Melnikov,Bert Kaiser,André Zimmermann
备注:37 pages, 6 figures
摘要:微机电系统(MEMS)扬声器是传统音圈扬声器的紧凑、可扩展的替代品,通过精确的半导体制造有望改善音质。本文综述了MEMS扬声器的研究现状,包括基于超声脉冲和热声的声音产生,并按致动原理对MEMS扬声器进行了分类:电动,压电和静电。对1990-2025年性能指标的比较分析突出了具有直接空气位移的压电MEMS的主导地位,重点是小型化和效率。该评论概述了即将到来的研究挑战,并确定了实现全频谱音频性能的潜在候选人。对创新方法的关注可能会导致MEMS扬声器的宽带采用。
摘要:Microelectromechanical systems (MEMS) speakers are compact, scalable alternatives to traditional voice coil speakers, promising improved sound quality through precise semiconductor manufacturing. This review provides an overview of the research landscape, including ultrasound pulse-based and thermoacoustic sound generation, classifying MEMS speakers by actuation principle: electrodynamic, piezoelectric, and electrostatic. A comparative analysis of performance indicators from 1990-2025 highlights the dominance of piezoelectric MEMS with direct air displacement, focusing on miniaturization and efficiency. The review outlines upcoming research challenges and identifies potential candidates for achieving full-spectrum audio performance. A focus on innovative approaches could lead to wideband adoption of MEMS-only speakers.


【2】Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
标题:使用商业自动语音识别和多模式大语言模型进行发音障碍语音的Zero-Shot识别
链接:https://arxiv.org/abs/2512.17474

作者:Ali Alsayegh,Tariq Masood
摘要:基于语音的人机交互是访问智能系统的主要方式,但构音障碍患者由于识别性能差距而面临系统性排斥。虽然自动语音识别(ASR)在典型语音上实现了低于5%的单词错误率(WER),但对于构音障碍的说话者,性能会显着下降。多模态大型语言模型(MLLM)提供了利用上下文推理来补偿声学退化的潜力,但其zero-shot能力仍然没有被描述。这项研究评估了TORGO构音障碍语音语料库上的八种商业语音转文本服务:四种传统的ASR系统(AssemblyAI,Whisper large-v3,Deepgram Nova-3,Nova-3 Medical)和四种基于MLLM的系统(GPT-4 o,GPT-4 o Mini,Gemini 2.5 Pro,Gemini 2.5 Flash)。评估包括词汇准确性,语义保留和成本延迟权衡。结果表明,严重程度相关的退化:轻度构音障碍达到3-5%的WER接近典型的语音基准,而严重构音障碍超过49%的WER在所有系统。逐字转录提示产生特定于架构的效果:GPT-4 o实现了7.36个百分点的WER降低,所有测试扬声器都有一致的改善,而Gemini变体表现出退化。语义指标表明,尽管词汇错误率升高,但交际意图仍然部分可恢复。这些发现建立了经验基线,使基于证据的技术选择辅助语音接口部署。
摘要:Voice-based human-machine interaction is a primary modality for accessing intelligent systems, yet individuals with dysarthria face systematic exclusion due to recognition performance gaps. Whilst automatic speech recognition (ASR) achieves word error rates (WER) below 5% on typical speech, performance degrades dramatically for dysarthric speakers. Multimodal large language models (MLLMs) offer potential for leveraging contextual reasoning to compensate for acoustic degradation, yet their zero-shot capabilities remain uncharacterised. This study evaluates eight commercial speech-to-text services on the TORGO dysarthric speech corpus: four conventional ASR systems (AssemblyAI, Whisper large-v3, Deepgram Nova-3, Nova-3 Medical) and four MLLM-based systems (GPT-4o, GPT-4o Mini, Gemini 2.5 Pro, Gemini 2.5 Flash). Evaluation encompasses lexical accuracy, semantic preservation, and cost-latency trade-offs. Results demonstrate severity-dependent degradation: mild dysarthria achieves 3-5% WER approaching typical-speech benchmarks, whilst severe dysarthria exceeds 49% WER across all systems. A verbatim-transcription prompt yields architecture-specific effects: GPT-4o achieves 7.36 percentage point WER reduction with consistent improvement across all tested speakers, whilst Gemini variants exhibit degradation. Semantic metrics indicate that communicative intent remains partially recoverable despite elevated lexical error rates. These findings establish empirical baselines enabling evidence-based technology selection for assistive voice interface deployment.


【3】When De-noising Hurts: A Systematic Study of Speech Enhancement Effects on Modern Medical ASR Systems
标题:去噪何时会造成伤害:现代医疗ASB系统语音增强效果的系统研究
链接:https://arxiv.org/abs/2512.17562

作者:Sujal Chondhekar,Vasanth Murukuri,Rushabh Vasani,Sanika Goyal,Rajshree Badami,Anushree Rana,Sanjana SN,Karthik Pandia,Sulabh Katiyar,Neha Jagadeesh,Sankalp Gulati
备注:Technical Report
摘要:语音增强方法被普遍认为可以提高噪声环境中自动语音识别(ASR)的性能。然而,这些技术的有效性不能被认为是理所当然的情况下,现代大规模的ASR模型训练的多样化,噪声数据。我们在四个最先进的ASR系统上对MetricGAN加语音库去噪进行了系统评估:OpenAI Whisper,NVIDIA Parakeet,Google Gemini Flash 2.0,Parrotlet-a,在九种噪声条件下使用500个医疗语音记录。ASR性能是使用语义WER(semWER)来衡量的,这是一个归一化的单词错误率(WER)指标,用于解释特定于域的标准化。我们的研究结果揭示了一个违反直觉的发现:语音增强预处理降低了所有噪声条件和模型的ASR性能。在所有40种测试配置(4种型号x10种条件)中,原始嘈杂音频的semWER低于增强音频,衰减范围为1.1%至46.6%的绝对semWER增加。这些研究结果表明,现代ASR模型具有足够的内部噪声鲁棒性,传统的语音增强可能会删除声学特征的关键ASR。对于在嘈杂的临床环境中部署医疗抄写系统的从业者来说,我们的研究结果表明,使用降噪技术对音频进行预处理不仅可能在计算上浪费,而且可能对转录准确性有害。
摘要:Speech enhancement methods are commonly believed to improve the performance of automatic speech recognition (ASR) in noisy environments. However, the effectiveness of these techniques cannot be taken for granted in the case of modern large-scale ASR models trained on diverse, noisy data. We present a systematic evaluation of MetricGAN-plus-voicebank denoising on four state-of-the-art ASR systems: OpenAI Whisper, NVIDIA Parakeet, Google Gemini Flash 2.0, Parrotlet-a using 500 medical speech recordings under nine noise conditions. ASR performance is measured using semantic WER (semWER), a normalized word error rate (WER) metric accounting for domain-specific normalizations. Our results reveal a counterintuitive finding: speech enhancement preprocessing degrades ASR performance across all noise conditions and models. Original noisy audio achieves lower semWER than enhanced audio in all 40 tested configurations (4 models x 10 conditions), with degradations ranging from 1.1% to 46.6% absolute semWER increase. These findings suggest that modern ASR models possess sufficient internal noise robustness and that traditional speech enhancement may remove acoustic features critical for ASR. For practitioners deploying medical scribe systems in noisy clinical environments, our results indicate that preprocessing audio with noise reduction techniques might not just be computationally wasteful but also be potentially harmful to the transcription accuracy.


【4】Do Foundational Audio Encoders Understand Music Structure?
标题:基础音频编码器了解音乐结构吗?
链接:https://arxiv.org/abs/2512.17209

作者:Keisuke Toyama,Zhi Zhong,Akira Takahashi,Shusuke Takahashi,Yuki Mitsufuji
摘要:在音乐信息检索(MIR)研究中,使用预训练的基础音频编码器(FAE)最近已成为一种趋势。在大量音乐和音频数据上预训练的FAE已被证明可以提高MIR任务的性能,例如音乐标记和自动音乐转录。然而,他们的使用音乐结构分析(MSA)仍然是探索不足。虽然有许多开源FAE模型可用,但只有一小部分已被用于MSA,并且学习方法,训练数据和模型上下文长度等因素对MSA性能的影响仍然不清楚。在这项研究中,我们对11种FAE进行了综合实验,以研究这些因素如何影响MSA性能。我们的研究结果表明,使用自监督学习的FAE与音乐数据上的掩蔽语言建模对MSA特别有效。这些发现为MSA的未来研究铺平了道路。
摘要:In music information retrieval (MIR) research, the use of pretrained foundational audio encoders (FAEs) has recently become a trend. FAEs pretrained on large amounts of music and audio data have been shown to improve performance on MIR tasks such as music tagging and automatic music transcription. However, their use for music structure analysis (MSA) remains underexplored. Although many open-source FAE models are available, only a small subset has been examined for MSA, and the impact of factors such as learning methods, training data, and model context length on MSA performance remains unclear. In this study, we conduct comprehensive experiments on 11 types of FAEs to investigate how these factors affect MSA performance. Our results demonstrate that FAEs using selfsupervised learning with masked language modeling on music data are particularly effective for MSA. These findings pave the way for future research in MSA.


机器翻译由腾讯交互翻译提供,仅供参考