微信公众号:arXiv_Daily
cs.SD语音
链接:https://arxiv.org/abs/2506.16969
备注:paper is in 4+1 pages
摘要:耳语语音识别对传统的自动语音识别系统提出了重大挑战,特别是当与方言变化相结合时。然而,利用一种有效的方法来解决这个问题,使用低范围的数据集和处理负载是有益的。本文提出了一种解决方案,使用基于Mamba的状态空间模型和四个微调的自监督模型,包括Wav2Vec2,WavLM,HuBERT和Whisper,以解决耳语音和方言多样性的双重挑战。根据我们的知识,这代表了在wTIMIT和CHAINS数据集上报告的用于低声语音识别的最佳性能。我们使用新加坡,美国和爱尔兰方言的耳语和正常语音数据训练模型。研究结果表明,利用提出的基于Mamba的模型可以作为一种高效的模型,用少量的耳语数据训练,同时进行耳语和正常语音识别。这项工作的代码是免费提供的。
摘要:Whispered speech recognition presents significant challenges for conventional automatic speech recognition systems, particularly when combined with dialect variation. However, utilizing an efficient method to solve this problem using a low-range dataset and processing load is beneficial. This paper proposes a solution using a Mamba-based state-space model and four fine-tuned self-supervised models consisting of Wav2Vec2, WavLM, HuBERT, and Whisper to address the dual challenges of whispered speech and dialect diversity. Based on our knowledge, this represents the best performance reported on the wTIMIT and CHAINS datasets for whispered speech recognition. We trained the models using whispered and normal speech data across Singaporean, US, and Irish dialects. The findings demonstrated that utilizing the proposed Mamba-based model could work as a highly efficient model trained with low amounts of whispered data to simultaneously work on whispered and normal speech recognition. The code for this work is freely available.
【2】Implementing Keyword Spotting on the MCUX947 Microcontroller with Integrated NPU
链接:https://arxiv.org/abs/2506.08911
备注: 4 pages
摘要:本文介绍了一种在恩智浦MCXN947微控制器上实现的关键词识别(KWS)系统,该系统集成了神经处理单元(NPU),可在资源受限的设备上实现实时语音交互。该系统将MFCC特征提取与CNN分类器相结合,使用量化感知训练进行优化,以最小的精度下降来减小模型大小。实验结果表明,与仅CPU执行相比,利用NPU的推理时间加快了59倍,在模型大小为30.58 KB的情况下实现了97.06%的准确率,证明了在嵌入式平台上实现高效、低功耗语音接口的可行性。摘要:This paper presents a keyword spotting (KWS) system implemented on the NXP MCXN947 microcontroller with an integrated Neural Processing Unit (NPU), enabling real-time voice interaction on resource-constrained devices. The system combines MFCC feature extraction with a CNN classifier, optimized using Quantization Aware Training to reduce model size with minimal accuracy drop. Experimental results demonstrate a 59x speedup in inference time when leveraging the NPU compared to CPU-only execution, achieving 97.06% accuracy with a model size of 30.58 KB, demonstrating the feasibility of efficient, low-power voice interfaces on embedded platforms.【3】 EDNet: A Distortion-Agnostic Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training
摘要:真实世界环境中的语音信号经常受到各种失真的影响,例如加性噪声、混响和带宽限制,这些失真可能单独出现或组合出现。传统的语音增强方法通常依赖于掩蔽或映射,掩蔽侧重于抑制非语音分量,同时保留可观察的结构,映射试图通过输入的直接变换来恢复干净的语音。每种方法在特定情况下都有优势,但在目标条件之外可能效果不佳。我们提出了一种与失真无关的语音增强框架--EDNet,该框架旨在处理广泛的失真类型,而无需事先假设任务或输入特性。EDNet由两个主要部分组成:(1)门控曼巴(GM)模块,其通过可学习的门控机制自适应地组合掩蔽和映射,所述门控机制基于局部信号特征在抑制(DRAW)和重建(Draw)之间进行选择,以及(2)相移不变训练(PSIT),一种移位容忍监督策略,其通过在训练期间启用动态对准来改进相位估计,同时保持与标准损失函数兼容。去噪、去混响、带宽扩展和多失真增强任务的实验结果表明,EDNet在各种条件下始终实现强劲的性能,展示了其架构灵活性和对不同任务设置的适应性。
摘要:Speech signals in real-world environments are frequently affected by various distortions such as additive noise, reverberation, and bandwidth limitation, which may appear individually or in combination. Traditional speech enhancement methods typically rely on either masking, which focuses on suppressing non-speech components while preserving observable structure, or mapping, which seeks to recover clean speech through direct transformation of the input. Each approach offers strengths in specific scenarios but may be less effective outside its target conditions. We propose the Erase and Draw Network (EDNet), a distortion-agnostic speech enhancement framework designed to handle a broad range of distortion types without prior assumptions about task or input characteristics. EDNet consists of two main components: (1) the Gating Mamba (GM) module, which adaptively combines masking and mapping through a learnable gating mechanism that selects between suppression (Erase) and reconstruction (Draw) based on local signal features, and (2) Phase Shift-Invariant Training (PSIT), a shift tolerant supervision strategy that improves phase estimation by enabling dynamic alignment during training while remaining compatible with standard loss functions. Experimental results on denoising, dereverberation, bandwidth extension, and multi distortion enhancement tasks show that EDNet consistently achieves strong performance across conditions, demonstrating its architectural flexibility and adaptability to diverse task settings.
【4】Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering
备注:Accepted at Interspeech 2025
摘要:我们提出了一个空间光谱,结合基于模型和数据驱动的日志化管道组成的TDOA为基础的分割,然后嵌入为基础的聚类。所提出的系统既不需要访问多通道训练数据,也不需要关于麦克风的数量或位置的先验知识。它适用于紧凑型麦克风阵列和分布式麦克风,只需进行微小的调整。由于其在分割过程中对重叠语音的出色处理,所提出的管道在具有紧凑麦克风阵列的场景和具有分布式麦克风的设置中均显着优于单通道pyannote方法。此外,我们表明,与完全空间的日记管道,所提出的系统可以正确地跟踪扬声器时,他们改变位置。
摘要:We propose a spatio-spectral, combined model-based and data-driven diarization pipeline consisting of TDOA-based segmentation followed by embedding-based clustering. The proposed system requires neither access to multi-channel training data nor prior knowledge about the number or placement of microphones. It works for both a compact microphone array and distributed microphones, with minor adjustments. Due to its superior handling of overlapping speech during segmentation, the proposed pipeline significantly outperforms the single-channel pyannote approach, both in a scenario with a compact microphone array and in a setup with distributed microphones. Additionally, we show that, unlike fully spatial diarization pipelines, the proposed system can correctly track speakers when they change positions.
【5】Universal Music Representations? Evaluating Foundation Models on World Music Corpora
备注: Accepted at ISMIR 2025
摘要:基金会模型已经彻底改变了音乐信息检索,但问题仍然是他们的能力,在不同的音乐传统的概括。本文提出了一个全面的评估五个国家的最先进的音频基础模型在六个音乐语料库跨越西方流行,希腊,土耳其和印度的古典传统。我们采用三种互补的方法来研究这些模型的跨文化能力:探索评估固有的表示,有针对性的监督微调1-2层,和多标签Few-Shot学习低资源的情况。我们的分析显示了不同的跨文化概括,较大的模型通常在非西方音乐上表现出色,但对于文化上遥远的传统,结果会下降。值得注意的是,我们的方法在六个评估数据集中的五个数据集上实现了最先进的性能,证明了世界音乐理解基础模型的有效性。我们还发现,我们有针对性的微调方法并不总是优于所有设置的探测,这表明基础模型已经编码了大量的音乐知识。我们的评估框架和基准测试结果有助于了解当前模型距离实现通用音乐表示有多远,同时为未来的进展建立指标。
摘要:Foundation models have revolutionized music information retrieval, but questions remain about their ability to generalize across diverse musical traditions. This paper presents a comprehensive evaluation of five state-of-the-art audio foundation models across six musical corpora spanning Western popular, Greek, Turkish, and Indian classical traditions. We employ three complementary methodologies to investigate these models' cross-cultural capabilities: probing to assess inherent representations, targeted supervised fine-tuning of 1-2 layers, and multi-label few-shot learning for low-resource scenarios. Our analysis shows varying cross-cultural generalization, with larger models typically outperforming on non-Western music, though results decline for culturally distant traditions. Notably, our approaches achieve state-of-the-art performance on five out of six evaluated datasets, demonstrating the effectiveness of foundation models for world music understanding. We also find that our targeted fine-tuning approach does not consistently outperform probing across all settings, suggesting foundation models already encode substantial musical knowledge. Our evaluation framework and benchmarking results contribute to understanding how far current models are from achieving universal music representations while establishing metrics for future progress.
【6】ITO-Master: Inference-Time Optimization for Audio Effects Modeling of Music Mastering Processors
备注:ISMIR 2025
摘要:音乐掌握风格迁移的目的是将参考曲目的掌握特征建模并应用于目标曲目,模拟专业掌握过程。然而,现有方法基于参考轨道应用固定处理,限制了用户微调结果以匹配其艺术意图的能力。在本文中,我们介绍了ITO-Master框架,一个基于参考的母版制作风格转换系统,集成了推理时间优化(ITO),使用户能够更好地控制母版制作过程。通过在推理过程中优化参考嵌入,我们的方法允许用户动态地优化输出,进行微观调整以实现更精确的母版制作结果。我们探索了用于对母版处理器建模的黑盒和白盒方法,并证明ITO可以提高不同风格的母版性能。通过客观评价,主观听力测试,定性分析使用基于文本的条件反射与CLAP嵌入,我们验证,ITO提高掌握风格的相似性,同时提供更高的适应性。我们的框架提供了一个有效的和用户可控的解决方案,掌握风格转移,允许用户改进他们的结果超出了最初的风格转移。
摘要:Music mastering style transfer aims to model and apply the mastering characteristics of a reference track to a target track, simulating the professional mastering process. However, existing methods apply fixed processing based on a reference track, limiting users' ability to fine-tune the results to match their artistic intent. In this paper, we introduce the ITO-Master framework, a reference-based mastering style transfer system that integrates Inference-Time Optimization (ITO) to enable finer user control over the mastering process. By optimizing the reference embedding during inference, our approach allows users to refine the output dynamically, making micro-level adjustments to achieve more precise mastering results. We explore both black-box and white-box methods for modeling mastering processors and demonstrate that ITO improves mastering performance across different styles. Through objective evaluation, subjective listening tests, and qualitative analysis using text-based conditioning with CLAP embeddings, we validate that ITO enhances mastering style similarity while offering increased adaptability. Our framework provides an effective and user-controllable solution for mastering style transfer, allowing users to refine their results beyond the initial style transfer.
【7】Hybrid-Sep: Language-queried audio source separation via pre-trained Model Fusion and Adversarial Diffusion Training
备注:Submitted to WASAA 2025
摘要:语音查询音频分离(LASS)采用基于语义描述的语言查询来分离目标声音。然而,现有的方法在保持分离精度的同时,将复杂的听觉特征与语言背景对齐面临挑战。目前的研究工作主要集中在文本描述增强和架构创新上,但集成预训练的自监督学习(SSL)音频模型和对比存储音频预训练(CLAP)框架的潜力,能够提取跨模态音频-文本关系,仍然没有得到充分的探索。为了解决这个问题,我们提出了HybridSep,一个两阶段的LASS框架,协同基于SSL的声学表示与CLAP派生的语义嵌入。我们的框架引入了对抗一致性训练(ACT),这是一种新型的优化策略,将扩散视为辅助正则化损失,同时集成对抗性训练以增强分离保真度。实验表明,HybridSep在最先进的基线上实现了显着的性能改进(例如,AudioSep、FlowSep),为LASS任务建立新的基准。
摘要:Language-queried Audio Separation (LASS) employs linguistic queries to isolate target sounds based on semantic descriptions. However, existing methods face challenges in aligning complex auditory features with linguistic context while preserving separation precision. Current research efforts focus primarily on text description augmentation and architectural innovations, yet the potential of integrating pre-trained self-supervised learning (SSL) audio models and Contrastive Language-Audio Pretraining (CLAP) frameworks, capable of extracting cross-modal audio-text relationships, remains underexplored. To address this, we present HybridSep, a two-stage LASS framework that synergizes SSL-based acoustic representations with CLAP-derived semantic embeddings. Our framework introduces Adversarial Consistent Training (ACT), a novel optimization strategy that treats diffusion as an auxiliary regularization loss while integrating adversarial training to enhance separation fidelity. Experiments demonstrate that HybridSep achieves significant performance improvements over state-of-the-art baselines (e.g., AudioSep, FlowSep) across multiple metrics, establishing new benchmarks for LASS tasks.
【8】LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
摘要:随着语音语言模型(SLMs)的快速发展,离散的语音令牌已经成为语音和文本之间的核心接口,使统一建模跨模态。最近的语音标记化方法旨在将语义信息从低级别声学中隔离出来,以更好地与语言模型保持一致。特别是,以前的方法使用SSL教师(如HuBERT)来提取语义表示,然后将其提取到语义量化器中以抑制声学冗余并捕获与内容相关的潜在结构。然而,它们产生的语音标记序列仍然比文本标记序列长得多,这给高效的语音语言建模带来了挑战。降低帧速率是一种自然的解决方案,但标准技术,如跨帧的刚性平均池化,可能会扭曲或稀释有效LM对齐所需的语义结构。为了解决这个问题,我们提出了LM-SPT,一种语音标记化方法,引入了一种新的语义蒸馏。而不是直接匹配教师和学生的功能,通过池,我们重建语音仅从语义令牌和最小化的原始和重建的波形,从冻结的自动语音识别(ASR)编码器获得的编码表示之间的差异。这种间接但数据驱动的监督使分词器能够学习与语言模型语义更一致的离散单元。LM-SPT进一步对语音标记化的编码器和解码器进行了架构改进,并支持多种帧速率,包括25 Hz、12.5Hz和6.25Hz。实验结果表明,与基线相比,LM-SPT实现了更好的重建保真度,并且使用LM-SPT令牌训练的SLM在语音到文本的性能上具有竞争力,并且在文本到语音的任务上始终优于基线。
摘要:With the rapid progress of speech language models (SLMs), discrete speech tokens have emerged as a core interface between speech and text, enabling unified modeling across modalities. Recent speech tokenization approaches aim to isolate semantic information from low-level acoustics to better align with language models. In particular, previous methods use SSL teachers such as HuBERT to extract semantic representations, which are then distilled into a semantic quantizer to suppress acoustic redundancy as well as capture content-related latent structures. However, they still produce speech token sequences significantly longer than their textual counterparts, creating challenges for efficient speech-language modeling. Reducing the frame rate is a natural solution, but standard techniques, such as rigid average pooling across frames, can distort or dilute the semantic structure required for effective LM alignment. To address this, we propose LM-SPT, a speech tokenization method that introduces a novel semantic distillation. Instead of directly matching teacher and student features via pooling, we reconstruct speech solely from semantic tokens and minimize the discrepancy between the encoded representations of the original and reconstructed waveforms, obtained from a frozen automatic speech recognition (ASR) encoder. This indirect yet data-driven supervision enables the tokenizer to learn discrete units that are more semantically aligned with language models. LM-SPT further incorporates architectural improvements to the encoder and decoder for speech tokenization, and supports multiple frame rates, including 25Hz, 12.5Hz, and 6.25Hz. Experimental results show that LM-SPT achieves superior reconstruction fidelity compared to baselines, and that SLMs trained with LM-SPT tokens achieve competitive performances on speech-to-text and consistently outperform baselines on text-to-speech tasks.
【9】Learning Magnitude Distribution of Sound Fields via Conditioned Autoencoder
备注:To appear in Forum Acusticum 2025
摘要:提出了一种基于学习的空间稀疏测量声场幅度分布估计方法。当相位测量不可靠或不可访问时,估计声学传递函数(ATF)的幅度分布是有用的,并且具有与空间音频相关的广泛应用。我们提出了一种基于神经网络的ATF震级估计方法。我们的网络架构的关键特征是输入和输出层的源和接收器的位置和频率和潜在变量的聚合模块,这可以被解释为一个基于自动编码器的扩展声场的基础扩展。数值模拟结果表明,我们提出的方法是准确的估计ATF的震级与少量的接收机。
摘要:A learning-based method for estimating the magnitude distribution of sound fields from spatially sparse measurements is proposed. Estimating the magnitude distribution of acoustic transfer function (ATF) is useful when phase measurements are unreliable or inaccessible and has a wide range of applications related to spatial audio. We propose a neural-network-based method for the ATF magnitude estimation. The key feature of our network architecture is the input and output layers conditioned on source and receiver positions and frequency and the aggregation module of latent variables, which can be interpreted as an autoencoder-based extension of the basis expansion of the sound field. Numerical simulation results indicated that the ATF magnitude is accurately estimated with a small number of receivers by our proposed method.
【10】Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
备注:Accepted to INTERSPEECH 2025
摘要:我们提出了第一个流口音转换(AC)模型,将非母语语音转换为类似母语的口音,同时保留扬声器的身份,韵律和改善发音。我们的方法使流处理通过修改以前的AC架构与Emformer编码器和优化的推理机制。此外,我们集成了一个原生的文本到语音(TTS)模型,以生成理想的地面实况数据,用于有效的训练。我们的流式AC模型实现了与顶级AC模型相当的性能,同时保持稳定的延迟,使其成为第一个能够流式传输的AC系统。
摘要:We propose a first streaming accent conversion (AC) model that transforms non-native speech into a native-like accent while preserving speaker identity, prosody and improving pronunciation. Our approach enables stream processing by modifying a previous AC architecture with an Emformer encoder and an optimized inference mechanism. Additionally, we integrate a native text-to-speech (TTS) model to generate ideal ground-truth data for efficient training. Our streaming AC model achieves comparable performance to the top AC models while maintaining stable latency, making it the first AC system capable of streaming.
【11】Weight Factorization and Centralization for Continual Learning in Speech Recognition
备注:Accepted to INTERSPEECH 2025
摘要:基于现代神经网络的语音识别模型需要不断吸收新数据,而无需重新训练整个系统,特别是在使用基础模型的下游应用中,无法访问原始训练数据。在无排练、多语言和语言不可知的条件下持续训练模型,可能会导致灾难性的遗忘,因为对权重的看似微不足道的破坏可能会破坏模型的质量。受人类大脑通过觉醒-睡眠周期学习和巩固知识的能力的启发,我们提出了一种持续学习方法,该方法具有两个不同的阶段:分解和集中,相应地学习和合并知识。我们在一系列不同的语码转换数据集上的实验表明,集中化阶段可以有效地防止灾难性遗忘,通过在多个分散的低秩适配器中积累知识。
摘要:Modern neural network based speech recognition models are required to continually absorb new data without re-training the whole system, especially in downstream applications using foundation models, having no access to the original training data. Continually training the models in a rehearsal-free, multilingual, and language agnostic condition, likely leads to catastrophic forgetting, when a seemingly insignificant disruption to the weights can destructively harm the quality of the models. Inspired by the ability of human brains to learn and consolidate knowledge through the waking-sleeping cycle, we propose a continual learning approach with two distinct phases: factorization and centralization, learning and merging knowledge accordingly. Our experiments on a sequence of varied code-switching datasets showed that the centralization stage can effectively prevent catastrophic forgetting by accumulating the knowledge in multiple scattering low-rank adapters.
【12】Automatic Speech Recognition Biases in Newcastle English: an Error Analysis
备注:Submitted to Interspeech 2025
摘要:自动语音识别(ASR)系统由于偏向于主流变体的有偏见的训练而与区域方言作斗争。虽然以前的研究已经确定了ASR中的种族,年龄和性别偏见,但区域偏见仍然没有得到充分的研究。本研究调查ASR性能纽卡斯尔英语,一个有据可查的区域方言已知的是具有挑战性的ASR。两个阶段的分析进行:第一,人工错误分析的子样本确定关键的语音,词汇,和形态句法错误背后的ASR mismismismisconnitations;第二,案例研究的重点是系统分析的ASR识别的区域代词“yous”和“wors”。结果表明,ASR错误与方言特征直接相关,而社会因素在ASR错配中的作用较小。我们提倡在ASR训练数据中增加方言多样性,并强调社会语言学分析在诊断和解决区域偏见方面的价值。
摘要:Automatic Speech Recognition (ASR) systems struggle with regional dialects due to biased training which favours mainstream varieties. While previous research has identified racial, age, and gender biases in ASR, regional bias remains underexamined. This study investigates ASR performance on Newcastle English, a well-documented regional dialect known to be challenging for ASR. A two-stage analysis was conducted: first, a manual error analysis on a subsample identified key phonological, lexical, and morphosyntactic errors behind ASR misrecognitions; second, a case study focused on the systematic analysis of ASR recognition of the regional pronouns ``yous'' and ``wor''. Results show that ASR errors directly correlate with regional dialectal features, while social factors play a lesser role in ASR mismatches. We advocate for greater dialectal diversity in ASR training data and highlight the value of sociolinguistic analysis in diagnosing and addressing regional biases.
【13】Towards Bitrate-Efficient and Noise-Robust Speech Coding with Variable Bitrate RVQ
备注:Accepted to Interspeech 2025
摘要:残差矢量量化(RVQ)已经成为神经语音和音频编码中的主导方法,提供高保真压缩。然而,语音编码由于真实世界的噪声而带来额外的挑战,这降低了压缩效率。标准编解码器均匀地分配比特,将比特率浪费在对可懂度没有贡献的噪声分量上。本文介绍了一种可变比特率RVQ(VRVQ)框架的噪声鲁棒性的语音编码,动态调整每帧的比特率,以优化率失真的权衡。与恒定比特率(CBR)RVQ不同,我们的方法优先考虑关键的语音分量,同时抑制残留噪声。此外,我们集成了一个特征去噪器,以进一步提高噪声鲁棒性。实验结果表明,与传统方法相比,VRVQ改进了率失真折衷,在噪声条件下获得了更好的压缩效率和感知质量。样品可在我们的项目页面:www.example.com。
摘要:Residual Vector Quantization (RVQ) has become a dominant approach in neural speech and audio coding, providing high-fidelity compression. However, speech coding presents additional challenges due to real-world noise, which degrades compression efficiency. Standard codecs allocate bits uniformly, wasting bitrate on noise components that do not contribute to intelligibility. This paper introduces a Variable Bitrate RVQ (VRVQ) framework for noise-robust speech coding, dynamically adjusting bitrate per frame to optimize rate-distortion trade-offs. Unlike constant bitrate (CBR) RVQ, our method prioritizes critical speech components while suppressing residual noise. Additionally, we integrate a feature denoiser to further improve noise robustness. Experimental results show that VRVQ improves rate-distortion trade-offs over conventional methods, achieving better compression efficiency and perceptual quality in noisy conditions. Samples are available at our project page: https://yoongi43.github.io/noise_robust_vrvq/.
【14】InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
备注:19 pages, 9 figures
摘要:在现代语音合成中,非语言信息--如说话者的音色、情绪状态和动态韵律--在传达语义之外的细微差别方面起着关键作用。传统的文本到语音(TTS)系统依赖于固定的风格标签或插入语音提示来控制这些提示,这严重限制了灵活性。最近的尝试试图采用自然语言的指令,以调节语言学特征,大大提高了概括的解释驱动的TTS模型。虽然许多TTS系统现在都支持通过文本描述进行定制合成,但它们解释和执行复杂指令的实际能力在很大程度上仍未得到开发。此外,还缺乏专门为基于语音识别的TTS设计的高质量基准和自动化评估指标,这阻碍了对这些模型的准确评估和迭代优化。为了解决这些局限性,我们引入了InstructTTSEval,这是一个衡量复杂自然语言风格控制能力的基准。我们介绍了三个任务,即声学参数规范,描述性风格指令和角色扮演,包括英语和中文子集,每个测试用例1 k(共6 k)与参考音频配对。我们利用双子座作为一个自动判断,以评估他们的预防以下能力。我们的评估无障碍预防以下TTS系统突出了进一步改进的空间。我们预计,InstructTTSEval将推动发展更强大,灵活,准确的解释以下TTS。
摘要:In modern speech synthesis, paralinguistic information--such as a speaker's vocal timbre, emotional state, and dynamic prosody--plays a critical role in conveying nuance beyond mere semantics. Traditional Text-to-Speech (TTS) systems rely on fixed style labels or inserting a speech prompt to control these cues, which severely limits flexibility. Recent attempts seek to employ natural-language instructions to modulate paralinguistic features, substantially improving the generalization of instruction-driven TTS models. Although many TTS systems now support customized synthesis via textual description, their actual ability to interpret and execute complex instructions remains largely unexplored. In addition, there is still a shortage of high-quality benchmarks and automated evaluation metrics specifically designed for instruction-based TTS, which hinders accurate assessment and iterative optimization of these models. To address these limitations, we introduce InstructTTSEval, a benchmark for measuring the capability of complex natural-language style control. We introduce three tasks, namely Acoustic-Parameter Specification, Descriptive-Style Directive, and Role-Play, including English and Chinese subsets, each with 1k test cases (6k in total) paired with reference audio. We leverage Gemini as an automatic judge to assess their instruction-following abilities. Our evaluation of accessible instruction-following TTS systems highlights substantial room for further improvement. We anticipate that InstructTTSEval will drive progress toward more powerful, flexible, and accurate instruction-following TTS.
【15】Optimizing Multilingual Text-To-Speech with Accents & Emotions
备注:12 pages, 8 figures
摘要:最先进的文本到语音(TTS)系统在单语环境中实现了高度的自然性,由于当前框架中的文化差异,合成具有正确多语言口音(特别是印度语)和上下文相关情感的语音仍然存在困难。本文介绍了一种新的TTS体系结构,它将口音与保留音译和多尺度情感建模相结合,特别适合印地语和印度英语口音。我们的方法扩展了Parler-TTS模型,通过集成特定语言的音素对齐混合编码器-解码器架构,以及在母语语料库上训练的文化敏感的情感嵌入层,以及将动态口音代码切换与残余矢量量化相结合。定量测试表明,口音准确率提高了23.7%(单词错误率从15.4%降低到11.8%),本地听众的情感识别准确率提高了85.3%,超过了METTS和VECL-TTS基线。该系统的新颖之处在于,它可以在实时生成语句(如“Namaste,让我们谈谈”)中混合代码
摘要:State-of-the-art text-to-speech (TTS) systems realize high naturalness in monolingual environments, synthesizing speech with correct multilingual accents (especially for Indic languages) and context-relevant emotions still poses difficulty owing to cultural nuance discrepancies in current frameworks. This paper introduces a new TTS architecture integrating accent along with preserving transliteration with multi-scale emotion modelling, in particularly tuned for Hindi and Indian English accent. Our approach extends the Parler-TTS model by integrating A language-specific phoneme alignment hybrid encoder-decoder architecture, and culture-sensitive emotion embedding layers trained on native speaker corpora, as well as incorporating a dynamic accent code switching with residual vector quantization. Quantitative tests demonstrate 23.7% improvement in accent accuracy (Word Error Rate reduction from 15.4% to 11.8%) and 85.3% emotion recognition accuracy from native listeners, surpassing METTS and VECL-TTS baselines. The novelty of the system is that it can mix code in real time - generating statements such as "Namaste, let's talk about
【16】 Advancing Automated Speaking Assessment Leveraging Multifaceted Relevance and Grammar Information
备注: submitted to the ISCA SLaTE-2025 Workshop
摘要:当前用于多方面评估的自动口语评估(ASA)系统通常未能充分利用内容相关性,忽视图像或范例线索,并且采用缺乏详细错误类型的肤浅语法分析。本文通过引入两个新的增强来改进这些不足,以构建一个混合评分模型。首先,一个多方面的相关性模块集成的问题和相关的图像内容,范例,和口语的L2扬声器的内容相关性的全面评估的反应。其次,细粒度的语法错误的功能,使用高级语法错误校正(GEC)和详细的注释,以确定特定的错误类别。实验和消融研究表明,这些组件显着提高内容相关性,语言使用和整体ASA性能的评估,突出使用更丰富,更细致入微的功能集的整体口语评估的好处。
摘要:Current automated speaking assessment (ASA) systems for use in multi-aspect evaluations often fail to make full use of content relevance, overlooking image or exemplar cues, and employ superficial grammar analysis that lacks detailed error types. This paper ameliorates these deficiencies by introducing two novel enhancements to construct a hybrid scoring model. First, a multifaceted relevance module integrates question and the associated image content, exemplar, and spoken response of an L2 speaker for a comprehensive assessment of content relevance. Second, fine-grained grammar error features are derived using advanced grammar error correction (GEC) and detailed annotation to identify specific error categories. Experiments and ablation studies demonstrate that these components significantly improve the evaluation of content relevance, language use, and overall ASA performance, highlighting the benefits of using richer, more nuanced feature sets for holistic speaking assessment.
【17】AeroGPT: Leveraging Large-Scale Audio Model for Aero-Engine Bearing Fault Diagnosis
摘要:航空发动机作为航空航天工业的关键部件,需要连续、准确的故障诊断,以保证运行安全,防止灾难性故障的发生。虽然深度学习技术在这种情况下得到了广泛的研究,但它们输出logits或置信度分数,需要进行后处理以获得可操作的见解。此外,大规模音频模型在这一领域的潜力在很大程度上仍未开发。为了解决这些局限性,本文提出了AeroGPT,一种新的框架,从一般的音频域知识转移到航空发动机轴承故障诊断。AeroGPT是一个基于大规模音频模型的框架,该框架结合了振动信号对齐(VSA)以使通用音频知识适应特定领域的振动模式,并结合生成故障分类(GFC)以直接输出可解释的故障标签。这种方法消除了对故障标签的后处理的需要,支持交互式的、可解释的和可操作的故障诊断,从而大大提高了工业适用性。通过对两个航空发动机轴承数据集的全面实验验证,AeroGPT在DIRG数据集上实现了98.94%的准确率,在HIT轴承数据集上实现了完美的100%分类,超越了传统的深度学习方法。附加的定性分析验证了我们的方法的有效性,并强调了大规模模型的潜力,彻底改变故障诊断。
摘要:Aerospace engines, as critical components in aviation and aerospace industries, require continuous and accurate fault diagnosis to ensure operational safety and prevent catastrophic failures. While deep learning techniques have been extensively studied in this context, they output logits or confidence scores, necessitating post-processing to derive actionable insights. Furthermore, the potential of large-scale audio models in this domain remains largely untapped. To address these limitations, this paper proposes AeroGPT, a novel framework that transfers knowledge from general audio domain to aero-engine bearing fault diagnosis. AeroGPT is a framework based on large-scale audio model that incorporates Vibration Signal Alignment (VSA) to adapt general audio knowledge to domain-specific vibration patterns, and combines Generative Fault Classification (GFC) to directly output interpretable fault labels. This approach eliminates the need for post-processing of fault labels, supports interactive, interpretable, and actionable fault diagnosis, thereby greatly enhancing industrial applicability. Through comprehensive experimental validation on two aero-engine bearing datasets, AeroGPT achieved exceptional performance with 98.94% accuracy on the DIRG dataset and perfect 100% classification on the HIT bearing dataset, surpassing traditional deep learning approaches. Additional Qualitative analysis validates the effectiveness of our approach and highlights the potential of large-scale models to revolutionize fault diagnosis.
【18】ingle-Microphone-Based Sound Source Localization for Mobile Robots in Reverberant Environments
备注:This paper was accepted and going to appear in the 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
摘要:准确估计声源位置对于机器人听力至关重要。然而,现有的声源定位方法通常依赖于具有至少两个空间预配置的麦克风的麦克风阵列。这一要求阻碍了基于麦克风的机器人试听系统和技术的适用性。为了缓解这些挑战,我们提出了一种在线声源定位方法,使用一个单一的麦克风安装在移动机器人在混响环境中。具体来说,我们开发了一个轻量级的神经网络模型,只有43k个参数,通过从混响信号中提取时间信息来执行实时距离估计。估计的距离,然后使用扩展卡尔曼滤波器进行处理,以实现在线声源定位。据我们所知,这是第一个使用移动机器人上的单个麦克风实现在线声源定位的工作,我们的目标是填补这项工作的空白。大量的实验证明了我们的方法的有效性和优点。为了使更广泛的研究社区受益,我们在https://github.com/JiangWAV/single-mic-SSL上开源了我们的代码。
摘要:Accurately estimating sound source positions is crucial for robot audition. However, existing sound source localization methods typically rely on a microphone array with at least two spatially preconfigured microphones. This requirement hinders the applicability of microphone-based robot audition systems and technologies. To alleviate these challenges, we propose an online sound source localization method that uses a single microphone mounted on a mobile robot in reverberant environments. Specifically, we develop a lightweight neural network model with only 43k parameters to perform real-time distance estimation by extracting temporal information from reverberant signals. The estimated distances are then processed using an extended Kalman filter to achieve online sound source localization. To the best of our knowledge, this is the first work to achieve online sound source localization using a single microphone on a moving robot, a gap that we aim to fill in this work. Extensive experiments demonstrate the effectiveness and merits of our approach. To benefit the broader research community, we have open-sourced our code at https://github.com/JiangWAV/single-mic-SSL.
【19】Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching
备注:Accepted at Interspeech 2025
摘要:None
摘要:Dysarthria is a neurological disorder that significantly impairs speech intelligibility, often rendering affected individuals unable to communicate effectively. This necessitates the development of robust dysarthric-to-regular speech conversion techniques. In this work, we investigate the utility and limitations of self-supervised learning (SSL) features and their quantized representations as an alternative to mel-spectrograms for speech generation. Additionally, we explore methods to mitigate speaker variability by generating clean speech in a single-speaker voice using features extracted from WavLM. To this end, we propose a fully non-autoregressive approach that leverages Conditional Flow Matching (CFM) with Diffusion Transformers to learn a direct mapping from dysarthric to clean speech. Our findings highlight the effectiveness of discrete acoustic units in improving intelligibility while achieving faster convergence compared to traditional mel-spectrogram-based approaches.
【20】VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schrödinger Bridge
备注:Accepted by Interspeech 2025
摘要:为了探索利用空间线索从图像中产生立体声歌唱声音与房间混响的潜在优势,我们介绍了VS-Singer,视觉引导模型,旨在从场景图像产生立体声歌唱声音与房间混响。VS-Singer包括三个模块:首先,模态交互网络将空间特征整合到文本编码中,以创建富含空间信息的语言表示。其次,解码器采用一致性薛定谔桥,以促进一步样本生成。此外,我们利用SFE模块,以提高视听匹配的一致性。据我们所知,这项研究是第一个结合立体声歌唱声音合成与视觉声学匹配在一个统一的框架。实验结果表明,VS-Singer可以有效地生成立体声歌唱声音,在一个单一的步骤,符合场景的角度。
摘要:To explore the potential advantages of utilizing spatial cues from images for generating stereo singing voices with room reverberation, we introduce VS-Singer, a vision-guided model designed to produce stereo singing voices with room reverberation from scene images. VS-Singer comprises three modules: firstly, a modal interaction network integrates spatial features into text encoding to create a linguistic representation enriched with spatial information. Secondly, the decoder employs a consistency Schr\"odinger bridge to facilitate one-step sample generation. Moreover, we utilize the SFE module to improve the consistency of audio-visual matching. To our knowledge, this study is the first to combine stereo singing voice synthesis with visual acoustic matching within a unified framework. Experimental results demonstrate that VS-Singer can effectively generate stereo singing voices that align with the scene perspective in a single step.
【21】Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion
备注:Accepted to ACL 2025 Findings
摘要:基于人工智能的音乐生成工具的快速发展正在彻底改变音乐行业,但也给艺术家、版权所有者和提供商带来了挑战。这需要可靠的方法来检测这种AI生成的内容。然而,现有的检测器,依赖于音频或歌词,面临着关键的实际限制:基于音频的检测器无法推广到新的或看不见的发电机,容易受到音频扰动;基于歌词的方法需要干净的格式和准确的歌词,在实践中不可用。为了克服这些限制,我们提出了一种新颖的,实际接地的方法:多模式,模块化后期融合管道,结合自动转录的歌词和语音功能捕捉歌词相关的信息在音频。通过直接依赖于来自音频的抒情方面,我们的方法增强了鲁棒性,减轻了对低级伪影的敏感性,并实现了实用性。实验表明,我们的方法,DE检测,优于现有的基于歌词的检测器,同时也更强大的音频扰动。因此,它为在现实世界中检测人工智能生成的音乐提供了一种有效、强大的解决方案。我们的代码可在https://github.com/deezer/robust-AI-lyrics-detection上获得。
摘要:The rapid advancement of AI-based music generation tools is revolutionizing the music industry but also posing challenges to artists, copyright holders, and providers alike. This necessitates reliable methods for detecting such AI-generated content. However, existing detectors, relying on either audio or lyrics, face key practical limitations: audio-based detectors fail to generalize to new or unseen generators and are vulnerable to audio perturbations; lyrics-based methods require cleanly formatted and accurate lyrics, unavailable in practice. To overcome these limitations, we propose a novel, practically grounded approach: a multimodal, modular late-fusion pipeline that combines automatically transcribed sung lyrics and speech features capturing lyrics-related information within the audio. By relying on lyrical aspects directly from audio, our method enhances robustness, mitigates susceptibility to low-level artifacts, and enables practical applicability. Experiments show that our method, DE-detect, outperforms existing lyrics-based detectors while also being more robust to audio perturbations. Thus, it offers an effective, robust solution for detecting AI-generated music in real-world scenarios. Our code is available at https://github.com/deezer/robust-AI-lyrics-detection.
【22】Early Attentive Sparsification Accelerates Neural Speech Transcription
摘要:基于变换器的神经语音处理已经达到了最先进的性能。由于语音音频信号是高度可压缩的,在这里,我们寻求加速神经语音转录的时域信号稀疏化早期的神经编码阶段,利用自注意机制的可解释性的Transformer音频编码器。使用Whisper系列模型,我们在稀疏化阶段(某个编码器层)和压缩比(稀疏度)的联合空间上执行系统的架构搜索。我们发现,在1%的准确度下降下,最佳结果解决方案选择在早期编码阶段将隐藏状态稀疏化到40-60%的稀疏度,从而在Nvidia GPU上实现高达1.6倍的英语语音转录任务运行时加速,而无需任何微调。
摘要:Transformer-based neural speech processing has achieved state-of-the-art performance. Since speech audio signals are known to be highly compressible, here we seek to accelerate neural speech transcription by time-domain signal sparsification early in the neural encoding stage, taking advantage of the interpretability of the self-attention mechanism in transformer audio encoders. With the Whisper family of models, we perform a systematic architecture search over the joint space of sparsification stage (a certain encoder layer) and compression ratio (sparsity). We found that the best resulting solutions under 1% accuracy degradation choose to sparsify the hidden state to 40-60% sparsity at an early encoding stage, and thereby achieve up to 1.6x runtime acceleration in English speech transcription tasks on Nvidia GPUs without any fine-tuning.
【23】Sonic4D: Spatial Audio Generation for Immersive 4D Scene Exploration
备注: 17 pages, 7 figures. Project page: this https URL
摘要:4D生成的最新进展已经证明了其在合成动态3D场景的真实感渲染方面的显着能力。然而,尽管实现了令人印象深刻的视觉性能,但几乎所有现有的方法都忽略了与对应的4D场景对齐的空间音频的生成,这对真正沉浸式视听体验造成了重大限制。为了缓解这个问题,我们提出了Sonic4D,一个新的框架,使空间音频生成沉浸式探索的4D场景。具体来说,我们的方法由三个阶段组成:1)为了从单目视频中捕获动态视觉内容和原始听觉信息,我们首先采用预先训练的专家模型来生成4D场景及其相应的单声道音频。2)随后,为了将单声道音频转换为空间音频,我们在4D场景中定位和跟踪声源,其中通过像素级视觉接地策略估计它们在不同时间戳的3D空间坐标。3)基于估计的声源位置,我们进一步合成合理的空间音频,不同的观点和时间戳使用基于物理的模拟。大量的实验表明,我们提出的方法生成逼真的空间音频与合成的4D场景在一个训练自由的方式一致,显着提高用户的沉浸式体验。生成的音频和视频示例可在https://x-drunker.github.io/Sonic4D-project-page上获得。
摘要:Recent advancements in 4D generation have demonstrated its remarkable capability in synthesizing photorealistic renderings of dynamic 3D scenes. However, despite achieving impressive visual performance, almost all existing methods overlook the generation of spatial audio aligned with the corresponding 4D scenes, posing a significant limitation to truly immersive audiovisual experiences. To mitigate this issue, we propose Sonic4D, a novel framework that enables spatial audio generation for immersive exploration of 4D scenes. Specifically, our method is composed of three stages: 1) To capture both the dynamic visual content and raw auditory information from a monocular video, we first employ pre-trained expert models to generate the 4D scene and its corresponding monaural audio. 2) Subsequently, to transform the monaural audio into spatial audio, we localize and track the sound sources within the 4D scene, where their 3D spatial coordinates at different timestamps are estimated via a pixel-level visual grounding strategy. 3) Based on the estimated sound source locations, we further synthesize plausible spatial audio that varies across different viewpoints and timestamps using physics-based simulation. Extensive experiments have demonstrated that our proposed method generates realistic spatial audio consistent with the synthesized 4D scene in a training-free manner, significantly enhancing the immersive experience for users. Generated audio and video examples are available at https://x-drunker.github.io/Sonic4D-project-page.
【24】 Explainable speech emotion recognition through attentive pooling: insights from attention-based temporal localization
备注:None
摘要:用于语音情感识别(SER)的最先进的Transformer模型依赖于时间特征聚合,然而高级池化方法仍然未被探索。我们系统地对池化策略进行了基准测试,包括多查询多头注意统计池化,它比平均池化实现了3.5个百分点的宏观F1增益。注意力分析显示,15%的帧捕捉了80%的情感线索,揭示了情感信息的局部模式。高关注帧的分析表明,非语言发声和hyperarticulated音素不成比例地优先在池,反映人类的感知策略。我们的研究结果的位置注意池作为一个性能SER机制和生物学上合理的工具,可解释的情绪定位。在自然条件下的Interspeech 2025语音情感识别挑战赛中,我们的方法获得了0.3649的宏观F1分数。
摘要:State-of-the-art transformer models for Speech Emotion Recognition (SER) rely on temporal feature aggregation, yet advanced pooling methods remain underexplored. We systematically benchmark pooling strategies, including Multi-Query Multi-Head Attentive Statistics Pooling, which achieves a 3.5 percentage point macro F1 gain over average pooling. Attention analysis shows 15 percent of frames capture 80 percent of emotion cues, revealing a localized pattern of emotional information. Analysis of high-attention frames reveals that non-linguistic vocalizations and hyperarticulated phonemes are disproportionately prioritized during pooling, mirroring human perceptual strategies. Our findings position attentive pooling as both a performant SER mechanism and a biologically plausible tool for explainable emotion localization. On Interspeech 2025 Speech Emotion Recognition in Naturalistic Conditions Challenge, our approach obtained a macro F1 score of 0.3649.
eess.AS音频处理
链接:https://arxiv.org/abs/2506.16969
备注:paper is in 4+1 pages
摘要:耳语语音识别对传统的自动语音识别系统提出了重大挑战,特别是当与方言变化相结合时。然而,利用一种有效的方法来解决这个问题,使用低范围的数据集和处理负载是有益的。本文提出了一种解决方案,使用基于Mamba的状态空间模型和四个微调的自监督模型,包括Wav2Vec2,WavLM,HuBERT和Whisper,以解决耳语音和方言多样性的双重挑战。根据我们的知识,这代表了在wTIMIT和CHAINS数据集上报告的用于低声语音识别的最佳性能。我们使用新加坡,美国和爱尔兰方言的耳语和正常语音数据训练模型。研究结果表明,利用提出的基于Mamba的模型可以作为一种高效的模型,用少量的耳语数据训练,同时进行耳语和正常语音识别。这项工作的代码是免费提供的。
摘要:Whispered speech recognition presents significant challenges for conventional automatic speech recognition systems, particularly when combined with dialect variation. However, utilizing an efficient method to solve this problem using a low-range dataset and processing load is beneficial. This paper proposes a solution using a Mamba-based state-space model and four fine-tuned self-supervised models consisting of Wav2Vec2, WavLM, HuBERT, and Whisper to address the dual challenges of whispered speech and dialect diversity. Based on our knowledge, this represents the best performance reported on the wTIMIT and CHAINS datasets for whispered speech recognition. We trained the models using whispered and normal speech data across Singaporean, US, and Irish dialects. The findings demonstrated that utilizing the proposed Mamba-based model could work as a highly efficient model trained with low amounts of whispered data to simultaneously work on whispered and normal speech recognition. The code for this work is freely available.
【2】H-QuEST: Accelerating Query-by-Example Spoken Term Detection with Hierarchical Indexing
备注:None
摘要:逐例查询口语术语检测(QbE-STD)使用样本口语查询在音频数据集中搜索匹配的单词或短语。当注释数据有限或不可用时,QbE-STD通常使用模板匹配方法(如动态时间规整(DTW))来完成,这些方法计算成本高且无法很好地扩展。为了解决这个问题,我们提出了H-QuEST(分层查询的例子口语术语检测),一种新的框架,加速口语术语检索,利用词频和逆文档频率(TF-IDF)的稀疏表示,通过先进的音频表示学习技术和分层可导航小世界(HNSW)索引与进一步完善。实验结果表明,H-QuEST提供了实质性的改进,在检索速度,而不牺牲准确性相比,现有的方法。
摘要:Query-by-example spoken term detection (QbE-STD) searches for matching words or phrases in an audio dataset using a sample spoken query. When annotated data is limited or unavailable, QbE-STD is often done using template matching methods like dynamic time warping (DTW), which are computationally expensive and do not scale well. To address this, we propose H-QuEST (Hierarchical Query-by-Example Spoken Term Detection), a novel framework that accelerates spoken term retrieval by utilizing Term Frequency and Inverse Document Frequency (TF-IDF)-based sparse representations obtained through advanced audio representation learning techniques and Hierarchical Navigable Small World (HNSW) indexing with further refinement. Experimental results show that H-QuEST delivers substantial improvements in retrieval speed without sacrificing accuracy compared to existing methods.
【3】RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching
备注: Accepted on Interspeech 2025
摘要:我们介绍RapFlow-TTS,一个快速和高保真的TTS声学模型,利用流速一致性约束的流匹配(FM)训练。虽然基于常微分方程(ODE)的TTS生成实现了自然质量的语音,但它通常需要大量的生成步骤,从而导致质量和推理速度之间的权衡。为了应对这一挑战,RapFlow-TTS在FM拉直的ODE轨迹上实现了速度场的一致性,从而以更少的生成步骤实现了一致的合成质量。此外,我们还引入了时间间隔调度和对抗学习等技术,以进一步提高几步合成的质量。实验结果表明,RapFlow-TTS实现了高保真语音合成,合成步骤比传统的FM和分数为基础的方法,分别减少了5倍和10倍。
摘要:We introduce RapFlow-TTS, a rapid and high-fidelity TTS acoustic model that leverages velocity consistency constraints in flow matching (FM) training. Although ordinary differential equation (ODE)-based TTS generation achieves natural-quality speech, it typically requires a large number of generation steps, resulting in a trade-off between quality and inference speed. To address this challenge, RapFlow-TTS enforces consistency in the velocity field along the FM-straightened ODE trajectory, enabling consistent synthetic quality with fewer generation steps. Additionally, we introduce techniques such as time interval scheduling and adversarial learning to further enhance the quality of the few-step synthesis. Experimental results show that RapFlow-TTS achieves high-fidelity speech synthesis with a 5- and 10-fold reduction in synthesis steps than the conventional FM- and score-based approaches, respectively.
【4】 EDNet: A Distortion-Agnostic Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training
摘要:真实世界环境中的语音信号经常受到各种失真的影响,例如加性噪声、混响和带宽限制,这些失真可能单独出现或组合出现。传统的语音增强方法通常依赖于掩蔽或映射,掩蔽侧重于抑制非语音分量,同时保留可观察的结构,映射试图通过输入的直接变换来恢复干净的语音。每种方法在特定情况下都有优势,但在目标条件之外可能效果不佳。我们提出了一种与失真无关的语音增强框架--EDNet,该框架旨在处理广泛的失真类型,而无需事先假设任务或输入特性。EDNet由两个主要部分组成:(1)门控曼巴(GM)模块,其通过可学习的门控机制自适应地组合掩蔽和映射,所述门控机制基于局部信号特征在抑制(DRAW)和重建(Draw)之间进行选择,以及(2)相移不变训练(PSIT),一种移位容忍监督策略,其通过在训练期间启用动态对准来改进相位估计,同时保持与标准损失函数兼容。去噪、去混响、带宽扩展和多失真增强任务的实验结果表明,EDNet在各种条件下始终实现强劲的性能,展示了其架构灵活性和对不同任务设置的适应性。
摘要:Speech signals in real-world environments are frequently affected by various distortions such as additive noise, reverberation, and bandwidth limitation, which may appear individually or in combination. Traditional speech enhancement methods typically rely on either masking, which focuses on suppressing non-speech components while preserving observable structure, or mapping, which seeks to recover clean speech through direct transformation of the input. Each approach offers strengths in specific scenarios but may be less effective outside its target conditions. We propose the Erase and Draw Network (EDNet), a distortion-agnostic speech enhancement framework designed to handle a broad range of distortion types without prior assumptions about task or input characteristics. EDNet consists of two main components: (1) the Gating Mamba (GM) module, which adaptively combines masking and mapping through a learnable gating mechanism that selects between suppression (Erase) and reconstruction (Draw) based on local signal features, and (2) Phase Shift-Invariant Training (PSIT), a shift tolerant supervision strategy that improves phase estimation by enabling dynamic alignment during training while remaining compatible with standard loss functions. Experimental results on denoising, dereverberation, bandwidth extension, and multi distortion enhancement tasks show that EDNet consistently achieves strong performance across conditions, demonstrating its architectural flexibility and adaptability to diverse task settings.
【5】Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering
备注:Accepted at Interspeech 2025
摘要:我们提出了一个空间光谱,结合基于模型和数据驱动的日志化管道组成的TDOA为基础的分割,然后嵌入为基础的聚类。所提出的系统既不需要访问多通道训练数据,也不需要关于麦克风的数量或位置的先验知识。它适用于紧凑型麦克风阵列和分布式麦克风,只需进行微小的调整。由于其在分割过程中对重叠语音的出色处理,所提出的管道在具有紧凑麦克风阵列的场景和具有分布式麦克风的设置中均显着优于单通道pyannote方法。此外,我们表明,与完全空间的日记管道,所提出的系统可以正确地跟踪扬声器时,他们改变位置。
摘要:We propose a spatio-spectral, combined model-based and data-driven diarization pipeline consisting of TDOA-based segmentation followed by embedding-based clustering. The proposed system requires neither access to multi-channel training data nor prior knowledge about the number or placement of microphones. It works for both a compact microphone array and distributed microphones, with minor adjustments. Due to its superior handling of overlapping speech during segmentation, the proposed pipeline significantly outperforms the single-channel pyannote approach, both in a scenario with a compact microphone array and in a setup with distributed microphones. Additionally, we show that, unlike fully spatial diarization pipelines, the proposed system can correctly track speakers when they change positions.
【6】Universal Music Representations? Evaluating Foundation Models on World Music Corpora
备注: Accepted at ISMIR 2025
摘要:基金会模型已经彻底改变了音乐信息检索,但问题仍然是他们的能力,在不同的音乐传统的概括。本文提出了一个全面的评估五个国家的最先进的音频基础模型在六个音乐语料库跨越西方流行,希腊,土耳其和印度的古典传统。我们采用三种互补的方法来研究这些模型的跨文化能力:探索评估固有的表示,有针对性的监督微调1-2层,和多标签Few-Shot学习低资源的情况。我们的分析显示了不同的跨文化概括,较大的模型通常在非西方音乐上表现出色,但对于文化上遥远的传统,结果会下降。值得注意的是,我们的方法在六个评估数据集中的五个数据集上实现了最先进的性能,证明了世界音乐理解基础模型的有效性。我们还发现,我们有针对性的微调方法并不总是优于所有设置的探测,这表明基础模型已经编码了大量的音乐知识。我们的评估框架和基准测试结果有助于了解当前模型距离实现通用音乐表示有多远,同时为未来的进展建立指标。
摘要:Foundation models have revolutionized music information retrieval, but questions remain about their ability to generalize across diverse musical traditions. This paper presents a comprehensive evaluation of five state-of-the-art audio foundation models across six musical corpora spanning Western popular, Greek, Turkish, and Indian classical traditions. We employ three complementary methodologies to investigate these models' cross-cultural capabilities: probing to assess inherent representations, targeted supervised fine-tuning of 1-2 layers, and multi-label few-shot learning for low-resource scenarios. Our analysis shows varying cross-cultural generalization, with larger models typically outperforming on non-Western music, though results decline for culturally distant traditions. Notably, our approaches achieve state-of-the-art performance on five out of six evaluated datasets, demonstrating the effectiveness of foundation models for world music understanding. We also find that our targeted fine-tuning approach does not consistently outperform probing across all settings, suggesting foundation models already encode substantial musical knowledge. Our evaluation framework and benchmarking results contribute to understanding how far current models are from achieving universal music representations while establishing metrics for future progress.
【7】ITO-Master: Inference-Time Optimization for Audio Effects Modeling of Music Mastering Processors
备注:ISMIR 2025
摘要:音乐掌握风格迁移的目的是将参考曲目的掌握特征建模并应用于目标曲目,模拟专业掌握过程。然而,现有方法基于参考轨道应用固定处理,限制了用户微调结果以匹配其艺术意图的能力。在本文中,我们介绍了ITO-Master框架,一个基于参考的母版制作风格转换系统,集成了推理时间优化(ITO),使用户能够更好地控制母版制作过程。通过在推理过程中优化参考嵌入,我们的方法允许用户动态地优化输出,进行微观调整以实现更精确的母版制作结果。我们探索了用于对母版处理器建模的黑盒和白盒方法,并证明ITO可以提高不同风格的母版性能。通过客观评价,主观听力测试,定性分析使用基于文本的条件反射与CLAP嵌入,我们验证,ITO提高掌握风格的相似性,同时提供更高的适应性。我们的框架提供了一个有效的和用户可控的解决方案,掌握风格转移,允许用户改进他们的结果超出了最初的风格转移。
摘要:Music mastering style transfer aims to model and apply the mastering characteristics of a reference track to a target track, simulating the professional mastering process. However, existing methods apply fixed processing based on a reference track, limiting users' ability to fine-tune the results to match their artistic intent. In this paper, we introduce the ITO-Master framework, a reference-based mastering style transfer system that integrates Inference-Time Optimization (ITO) to enable finer user control over the mastering process. By optimizing the reference embedding during inference, our approach allows users to refine the output dynamically, making micro-level adjustments to achieve more precise mastering results. We explore both black-box and white-box methods for modeling mastering processors and demonstrate that ITO improves mastering performance across different styles. Through objective evaluation, subjective listening tests, and qualitative analysis using text-based conditioning with CLAP embeddings, we validate that ITO enhances mastering style similarity while offering increased adaptability. Our framework provides an effective and user-controllable solution for mastering style transfer, allowing users to refine their results beyond the initial style transfer.
【8】Hybrid-Sep: Language-queried audio source separation via pre-trained Model Fusion and Adversarial Diffusion Training
备注:Submitted to WASAA 2025
摘要:语音查询音频分离(LASS)采用基于语义描述的语言查询来分离目标声音。然而,现有的方法在保持分离精度的同时,将复杂的听觉特征与语言背景对齐面临挑战。目前的研究工作主要集中在文本描述增强和架构创新上,但集成预训练的自监督学习(SSL)音频模型和对比存储音频预训练(CLAP)框架的潜力,能够提取跨模态音频-文本关系,仍然没有得到充分的探索。为了解决这个问题,我们提出了HybridSep,一个两阶段的LASS框架,协同基于SSL的声学表示与CLAP派生的语义嵌入。我们的框架引入了对抗一致性训练(ACT),这是一种新型的优化策略,将扩散视为辅助正则化损失,同时集成对抗性训练以增强分离保真度。实验表明,HybridSep在最先进的基线上实现了显着的性能改进(例如,AudioSep、FlowSep),为LASS任务建立新的基准。
摘要:Language-queried Audio Separation (LASS) employs linguistic queries to isolate target sounds based on semantic descriptions. However, existing methods face challenges in aligning complex auditory features with linguistic context while preserving separation precision. Current research efforts focus primarily on text description augmentation and architectural innovations, yet the potential of integrating pre-trained self-supervised learning (SSL) audio models and Contrastive Language-Audio Pretraining (CLAP) frameworks, capable of extracting cross-modal audio-text relationships, remains underexplored. To address this, we present HybridSep, a two-stage LASS framework that synergizes SSL-based acoustic representations with CLAP-derived semantic embeddings. Our framework introduces Adversarial Consistent Training (ACT), a novel optimization strategy that treats diffusion as an auxiliary regularization loss while integrating adversarial training to enhance separation fidelity. Experiments demonstrate that HybridSep achieves significant performance improvements over state-of-the-art baselines (e.g., AudioSep, FlowSep) across multiple metrics, establishing new benchmarks for LASS tasks.
【9】LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
摘要:随着语音语言模型(SLMs)的快速发展,离散的语音令牌已经成为语音和文本之间的核心接口,使统一建模跨模态。最近的语音标记化方法旨在将语义信息从低级别声学中隔离出来,以更好地与语言模型保持一致。特别是,以前的方法使用SSL教师(如HuBERT)来提取语义表示,然后将其提取到语义量化器中以抑制声学冗余并捕获与内容相关的潜在结构。然而,它们产生的语音标记序列仍然比文本标记序列长得多,这给高效的语音语言建模带来了挑战。降低帧速率是一种自然的解决方案,但标准技术,如跨帧的刚性平均池化,可能会扭曲或稀释有效LM对齐所需的语义结构。为了解决这个问题,我们提出了LM-SPT,一种语音标记化方法,引入了一种新的语义蒸馏。而不是直接匹配教师和学生的功能,通过池,我们重建语音仅从语义令牌和最小化的原始和重建的波形,从冻结的自动语音识别(ASR)编码器获得的编码表示之间的差异。这种间接但数据驱动的监督使分词器能够学习与语言模型语义更一致的离散单元。LM-SPT进一步对语音标记化的编码器和解码器进行了架构改进,并支持多种帧速率,包括25 Hz、12.5Hz和6.25Hz。实验结果表明,与基线相比,LM-SPT实现了更好的重建保真度,并且使用LM-SPT令牌训练的SLM在语音到文本的性能上具有竞争力,并且在文本到语音的任务上始终优于基线。
摘要:With the rapid progress of speech language models (SLMs), discrete speech tokens have emerged as a core interface between speech and text, enabling unified modeling across modalities. Recent speech tokenization approaches aim to isolate semantic information from low-level acoustics to better align with language models. In particular, previous methods use SSL teachers such as HuBERT to extract semantic representations, which are then distilled into a semantic quantizer to suppress acoustic redundancy as well as capture content-related latent structures. However, they still produce speech token sequences significantly longer than their textual counterparts, creating challenges for efficient speech-language modeling. Reducing the frame rate is a natural solution, but standard techniques, such as rigid average pooling across frames, can distort or dilute the semantic structure required for effective LM alignment. To address this, we propose LM-SPT, a speech tokenization method that introduces a novel semantic distillation. Instead of directly matching teacher and student features via pooling, we reconstruct speech solely from semantic tokens and minimize the discrepancy between the encoded representations of the original and reconstructed waveforms, obtained from a frozen automatic speech recognition (ASR) encoder. This indirect yet data-driven supervision enables the tokenizer to learn discrete units that are more semantically aligned with language models. LM-SPT further incorporates architectural improvements to the encoder and decoder for speech tokenization, and supports multiple frame rates, including 25Hz, 12.5Hz, and 6.25Hz. Experimental results show that LM-SPT achieves superior reconstruction fidelity compared to baselines, and that SLMs trained with LM-SPT tokens achieve competitive performances on speech-to-text and consistently outperform baselines on text-to-speech tasks.
【10】Learning Magnitude Distribution of Sound Fields via Conditioned Autoencoder
备注:To appear in Forum Acusticum 2025
摘要:提出了一种基于学习的空间稀疏测量声场幅度分布估计方法。当相位测量不可靠或不可访问时,估计声学传递函数(ATF)的幅度分布是有用的,并且具有与空间音频相关的广泛应用。我们提出了一种基于神经网络的ATF震级估计方法。我们的网络架构的关键特征是输入和输出层的源和接收器的位置和频率和潜在变量的聚合模块,这可以被解释为一个基于自动编码器的扩展声场的基础扩展。数值模拟结果表明,我们提出的方法是准确的估计ATF的震级与少量的接收机。
摘要:A learning-based method for estimating the magnitude distribution of sound fields from spatially sparse measurements is proposed. Estimating the magnitude distribution of acoustic transfer function (ATF) is useful when phase measurements are unreliable or inaccessible and has a wide range of applications related to spatial audio. We propose a neural-network-based method for the ATF magnitude estimation. The key feature of our network architecture is the input and output layers conditioned on source and receiver positions and frequency and the aggregation module of latent variables, which can be interpreted as an autoencoder-based extension of the basis expansion of the sound field. Numerical simulation results indicated that the ATF magnitude is accurately estimated with a small number of receivers by our proposed method.
【11】Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
备注:Accepted to INTERSPEECH 2025
摘要:我们提出了第一个流口音转换(AC)模型,将非母语语音转换为类似母语的口音,同时保留扬声器的身份,韵律和改善发音。我们的方法使流处理通过修改以前的AC架构与Emformer编码器和优化的推理机制。此外,我们集成了一个原生的文本到语音(TTS)模型,以生成理想的地面实况数据,用于有效的训练。我们的流式AC模型实现了与顶级AC模型相当的性能,同时保持稳定的延迟,使其成为第一个能够流式传输的AC系统。
摘要:We propose a first streaming accent conversion (AC) model that transforms non-native speech into a native-like accent while preserving speaker identity, prosody and improving pronunciation. Our approach enables stream processing by modifying a previous AC architecture with an Emformer encoder and an optimized inference mechanism. Additionally, we integrate a native text-to-speech (TTS) model to generate ideal ground-truth data for efficient training. Our streaming AC model achieves comparable performance to the top AC models while maintaining stable latency, making it the first AC system capable of streaming.
【12】Weight Factorization and Centralization for Continual Learning in Speech Recognition
备注:Accepted to INTERSPEECH 2025
摘要:基于现代神经网络的语音识别模型需要不断吸收新数据,而无需重新训练整个系统,特别是在使用基础模型的下游应用中,无法访问原始训练数据。在无排练、多语言和语言不可知的条件下持续训练模型,可能会导致灾难性的遗忘,因为对权重的看似微不足道的破坏可能会破坏模型的质量。受人类大脑通过觉醒-睡眠周期学习和巩固知识的能力的启发,我们提出了一种持续学习方法,该方法具有两个不同的阶段:分解和集中,相应地学习和合并知识。我们在一系列不同的语码转换数据集上的实验表明,集中化阶段可以有效地防止灾难性遗忘,通过在多个分散的低秩适配器中积累知识。
摘要:Modern neural network based speech recognition models are required to continually absorb new data without re-training the whole system, especially in downstream applications using foundation models, having no access to the original training data. Continually training the models in a rehearsal-free, multilingual, and language agnostic condition, likely leads to catastrophic forgetting, when a seemingly insignificant disruption to the weights can destructively harm the quality of the models. Inspired by the ability of human brains to learn and consolidate knowledge through the waking-sleeping cycle, we propose a continual learning approach with two distinct phases: factorization and centralization, learning and merging knowledge accordingly. Our experiments on a sequence of varied code-switching datasets showed that the centralization stage can effectively prevent catastrophic forgetting by accumulating the knowledge in multiple scattering low-rank adapters.
【13】Automatic Speech Recognition Biases in Newcastle English: an Error Analysis
备注:Submitted to Interspeech 2025
摘要:自动语音识别(ASR)系统由于偏向于主流变体的有偏见的训练而与区域方言作斗争。虽然以前的研究已经确定了ASR中的种族,年龄和性别偏见,但区域偏见仍然没有得到充分的研究。本研究调查ASR性能纽卡斯尔英语,一个有据可查的区域方言已知的是具有挑战性的ASR。两个阶段的分析进行:第一,人工错误分析的子样本确定关键的语音,词汇,和形态句法错误背后的ASR mismismismisconnitations;第二,案例研究的重点是系统分析的ASR识别的区域代词“yous”和“wors”。结果表明,ASR错误与方言特征直接相关,而社会因素在ASR错配中的作用较小。我们提倡在ASR训练数据中增加方言多样性,并强调社会语言学分析在诊断和解决区域偏见方面的价值。
摘要:Automatic Speech Recognition (ASR) systems struggle with regional dialects due to biased training which favours mainstream varieties. While previous research has identified racial, age, and gender biases in ASR, regional bias remains underexamined. This study investigates ASR performance on Newcastle English, a well-documented regional dialect known to be challenging for ASR. A two-stage analysis was conducted: first, a manual error analysis on a subsample identified key phonological, lexical, and morphosyntactic errors behind ASR misrecognitions; second, a case study focused on the systematic analysis of ASR recognition of the regional pronouns ``yous'' and ``wor''. Results show that ASR errors directly correlate with regional dialectal features, while social factors play a lesser role in ASR mismatches. We advocate for greater dialectal diversity in ASR training data and highlight the value of sociolinguistic analysis in diagnosing and addressing regional biases.
【14】Towards Bitrate-Efficient and Noise-Robust Speech Coding with Variable Bitrate RVQ
备注:Accepted to Interspeech 2025
摘要:残差矢量量化(RVQ)已经成为神经语音和音频编码中的主导方法,提供高保真压缩。然而,语音编码由于真实世界的噪声而带来额外的挑战,这降低了压缩效率。标准编解码器均匀地分配比特,将比特率浪费在对可懂度没有贡献的噪声分量上。本文介绍了一种可变比特率RVQ(VRVQ)框架的噪声鲁棒性的语音编码,动态调整每帧的比特率,以优化率失真的权衡。与恒定比特率(CBR)RVQ不同,我们的方法优先考虑关键的语音分量,同时抑制残留噪声。此外,我们集成了一个特征去噪器,以进一步提高噪声鲁棒性。实验结果表明,与传统方法相比,VRVQ改进了率失真折衷,在噪声条件下获得了更好的压缩效率和感知质量。样品可在我们的项目页面:www.example.com。
摘要:Residual Vector Quantization (RVQ) has become a dominant approach in neural speech and audio coding, providing high-fidelity compression. However, speech coding presents additional challenges due to real-world noise, which degrades compression efficiency. Standard codecs allocate bits uniformly, wasting bitrate on noise components that do not contribute to intelligibility. This paper introduces a Variable Bitrate RVQ (VRVQ) framework for noise-robust speech coding, dynamically adjusting bitrate per frame to optimize rate-distortion trade-offs. Unlike constant bitrate (CBR) RVQ, our method prioritizes critical speech components while suppressing residual noise. Additionally, we integrate a feature denoiser to further improve noise robustness. Experimental results show that VRVQ improves rate-distortion trade-offs over conventional methods, achieving better compression efficiency and perceptual quality in noisy conditions. Samples are available at our project page: https://yoongi43.github.io/noise_robust_vrvq/.
【15】InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
备注:19 pages, 9 figures
摘要:在现代语音合成中,非语言信息--如说话者的音色、情绪状态和动态韵律--在传达语义之外的细微差别方面起着关键作用。传统的文本到语音(TTS)系统依赖于固定的风格标签或插入语音提示来控制这些提示,这严重限制了灵活性。最近的尝试试图采用自然语言的指令,以调节语言学特征,大大提高了概括的解释驱动的TTS模型。虽然许多TTS系统现在都支持通过文本描述进行定制合成,但它们解释和执行复杂指令的实际能力在很大程度上仍未得到开发。此外,还缺乏专门为基于语音识别的TTS设计的高质量基准和自动化评估指标,这阻碍了对这些模型的准确评估和迭代优化。为了解决这些局限性,我们引入了InstructTTSEval,这是一个衡量复杂自然语言风格控制能力的基准。我们介绍了三个任务,即声学参数规范,描述性风格指令和角色扮演,包括英语和中文子集,每个测试用例1 k(共6 k)与参考音频配对。我们利用双子座作为一个自动判断,以评估他们的预防以下能力。我们的评估无障碍预防以下TTS系统突出了进一步改进的空间。我们预计,InstructTTSEval将推动发展更强大,灵活,准确的解释以下TTS。
摘要:In modern speech synthesis, paralinguistic information--such as a speaker's vocal timbre, emotional state, and dynamic prosody--plays a critical role in conveying nuance beyond mere semantics. Traditional Text-to-Speech (TTS) systems rely on fixed style labels or inserting a speech prompt to control these cues, which severely limits flexibility. Recent attempts seek to employ natural-language instructions to modulate paralinguistic features, substantially improving the generalization of instruction-driven TTS models. Although many TTS systems now support customized synthesis via textual description, their actual ability to interpret and execute complex instructions remains largely unexplored. In addition, there is still a shortage of high-quality benchmarks and automated evaluation metrics specifically designed for instruction-based TTS, which hinders accurate assessment and iterative optimization of these models. To address these limitations, we introduce InstructTTSEval, a benchmark for measuring the capability of complex natural-language style control. We introduce three tasks, namely Acoustic-Parameter Specification, Descriptive-Style Directive, and Role-Play, including English and Chinese subsets, each with 1k test cases (6k in total) paired with reference audio. We leverage Gemini as an automatic judge to assess their instruction-following abilities. Our evaluation of accessible instruction-following TTS systems highlights substantial room for further improvement. We anticipate that InstructTTSEval will drive progress toward more powerful, flexible, and accurate instruction-following TTS.
【16】Optimizing Multilingual Text-To-Speech with Accents & Emotions
备注:12 pages, 8 figures
摘要:最先进的文本到语音(TTS)系统在单语环境中实现了高度的自然性,由于当前框架中的文化差异,合成具有正确多语言口音(特别是印度语)和上下文相关情感的语音仍然存在困难。本文介绍了一种新的TTS体系结构,它将口音与保留音译和多尺度情感建模相结合,特别适合印地语和印度英语口音。我们的方法扩展了Parler-TTS模型,通过集成特定语言的音素对齐混合编码器-解码器架构,以及在母语语料库上训练的文化敏感的情感嵌入层,以及将动态口音代码切换与残余矢量量化相结合。定量测试表明,口音准确率提高了23.7%(单词错误率从15.4%降低到11.8%),本地听众的情感识别准确率提高了85.3%,超过了METTS和VECL-TTS基线。该系统的新颖之处在于,它可以在实时生成语句(如“Namaste,让我们谈谈”)中混合代码
摘要:State-of-the-art text-to-speech (TTS) systems realize high naturalness in monolingual environments, synthesizing speech with correct multilingual accents (especially for Indic languages) and context-relevant emotions still poses difficulty owing to cultural nuance discrepancies in current frameworks. This paper introduces a new TTS architecture integrating accent along with preserving transliteration with multi-scale emotion modelling, in particularly tuned for Hindi and Indian English accent. Our approach extends the Parler-TTS model by integrating A language-specific phoneme alignment hybrid encoder-decoder architecture, and culture-sensitive emotion embedding layers trained on native speaker corpora, as well as incorporating a dynamic accent code switching with residual vector quantization. Quantitative tests demonstrate 23.7% improvement in accent accuracy (Word Error Rate reduction from 15.4% to 11.8%) and 85.3% emotion recognition accuracy from native listeners, surpassing METTS and VECL-TTS baselines. The novelty of the system is that it can mix code in real time - generating statements such as "Namaste, let's talk about
【17】 Advancing Automated Speaking Assessment Leveraging Multifaceted Relevance and Grammar Information
备注: submitted to the ISCA SLaTE-2025 Workshop
摘要:当前用于多方面评估的自动口语评估(ASA)系统通常未能充分利用内容相关性,忽视图像或范例线索,并且采用缺乏详细错误类型的肤浅语法分析。本文通过引入两个新的增强来改进这些不足,以构建一个混合评分模型。首先,一个多方面的相关性模块集成的问题和相关的图像内容,范例,和口语的L2扬声器的内容相关性的全面评估的反应。其次,细粒度的语法错误的功能,使用高级语法错误校正(GEC)和详细的注释,以确定特定的错误类别。实验和消融研究表明,这些组件显着提高内容相关性,语言使用和整体ASA性能的评估,突出使用更丰富,更细致入微的功能集的整体口语评估的好处。
摘要:Current automated speaking assessment (ASA) systems for use in multi-aspect evaluations often fail to make full use of content relevance, overlooking image or exemplar cues, and employ superficial grammar analysis that lacks detailed error types. This paper ameliorates these deficiencies by introducing two novel enhancements to construct a hybrid scoring model. First, a multifaceted relevance module integrates question and the associated image content, exemplar, and spoken response of an L2 speaker for a comprehensive assessment of content relevance. Second, fine-grained grammar error features are derived using advanced grammar error correction (GEC) and detailed annotation to identify specific error categories. Experiments and ablation studies demonstrate that these components significantly improve the evaluation of content relevance, language use, and overall ASA performance, highlighting the benefits of using richer, more nuanced feature sets for holistic speaking assessment.
【18】End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data
摘要:缺乏高质量的注释数据是开发有效的端到端语音到文本翻译(ST)系统的一个重大挑战,特别是对于低资源语言。本文探讨了弱标记数据可以用来为低资源语言对建立ST模型的假设。我们使用最先进的句子编码器,在双文本挖掘的帮助下构建了语音到文本的翻译数据集。我们挖掘了多语言Shrutilipi语料库,以构建Shrutilipi-anuvaad,这是一个包含语言对孟加拉语-印地语,马来语-印地语,Odia-印地语和泰卢固语-印地语的ST数据的数据集。我们创建了具有不同质量和数量程度的多个版本的训练数据,以研究弱标记数据的质量与数量对ST模型性能的影响。结果表明,ST系统可以使用弱标记数据构建,其性能可与SONAR和M4 T等大规模多模态多语言基线相媲美。
摘要:The scarcity of high-quality annotated data presents a significant challenge in developing effective end-to-end speech-to-text translation (ST) systems, particularly for low-resource languages. This paper explores the hypothesis that weakly labeled data can be used to build ST models for low-resource language pairs. We constructed speech-to-text translation datasets with the help of bitext mining using state-of-the-art sentence encoders. We mined the multilingual Shrutilipi corpus to build Shrutilipi-anuvaad, a dataset comprising ST data for language pairs Bengali-Hindi, Malayalam-Hindi, Odia-Hindi, and Telugu-Hindi. We created multiple versions of training data with varying degrees of quality and quantity to investigate the effect of quality versus quantity of weakly labeled data on ST model performance. Results demonstrate that ST systems can be built using weakly labeled data, with performance comparable to massive multi-modal multilingual baselines such as SONAR and SeamlessM4T.
【19】AeroGPT: Leveraging Large-Scale Audio Model for Aero-Engine Bearing Fault Diagnosis
摘要:航空发动机作为航空航天工业的关键部件,需要连续、准确的故障诊断,以保证运行安全,防止灾难性故障的发生。虽然深度学习技术在这种情况下得到了广泛的研究,但它们输出logits或置信度分数,需要进行后处理以获得可操作的见解。此外,大规模音频模型在这一领域的潜力在很大程度上仍未开发。为了解决这些局限性,本文提出了AeroGPT,一种新的框架,从一般的音频域知识转移到航空发动机轴承故障诊断。AeroGPT是一个基于大规模音频模型的框架,该框架结合了振动信号对齐(VSA)以使通用音频知识适应特定领域的振动模式,并结合生成故障分类(GFC)以直接输出可解释的故障标签。这种方法消除了对故障标签的后处理的需要,支持交互式的、可解释的和可操作的故障诊断,从而大大提高了工业适用性。通过对两个航空发动机轴承数据集的全面实验验证,AeroGPT在DIRG数据集上实现了98.94%的准确率,在HIT轴承数据集上实现了完美的100%分类,超越了传统的深度学习方法。附加的定性分析验证了我们的方法的有效性,并强调了大规模模型的潜力,彻底改变故障诊断。
摘要:Aerospace engines, as critical components in aviation and aerospace industries, require continuous and accurate fault diagnosis to ensure operational safety and prevent catastrophic failures. While deep learning techniques have been extensively studied in this context, they output logits or confidence scores, necessitating post-processing to derive actionable insights. Furthermore, the potential of large-scale audio models in this domain remains largely untapped. To address these limitations, this paper proposes AeroGPT, a novel framework that transfers knowledge from general audio domain to aero-engine bearing fault diagnosis. AeroGPT is a framework based on large-scale audio model that incorporates Vibration Signal Alignment (VSA) to adapt general audio knowledge to domain-specific vibration patterns, and combines Generative Fault Classification (GFC) to directly output interpretable fault labels. This approach eliminates the need for post-processing of fault labels, supports interactive, interpretable, and actionable fault diagnosis, thereby greatly enhancing industrial applicability. Through comprehensive experimental validation on two aero-engine bearing datasets, AeroGPT achieved exceptional performance with 98.94% accuracy on the DIRG dataset and perfect 100% classification on the HIT bearing dataset, surpassing traditional deep learning approaches. Additional Qualitative analysis validates the effectiveness of our approach and highlights the potential of large-scale models to revolutionize fault diagnosis.
【20】ingle-Microphone-Based Sound Source Localization for Mobile Robots in Reverberant Environments
备注:This paper was accepted and going to appear in the 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
摘要:准确估计声源位置对于机器人听力至关重要。然而,现有的声源定位方法通常依赖于具有至少两个空间预配置的麦克风的麦克风阵列。这一要求阻碍了基于麦克风的机器人试听系统和技术的适用性。为了缓解这些挑战,我们提出了一种在线声源定位方法,使用一个单一的麦克风安装在移动机器人在混响环境中。具体来说,我们开发了一个轻量级的神经网络模型,只有43k个参数,通过从混响信号中提取时间信息来执行实时距离估计。估计的距离,然后使用扩展卡尔曼滤波器进行处理,以实现在线声源定位。据我们所知,这是第一个使用移动机器人上的单个麦克风实现在线声源定位的工作,我们的目标是填补这项工作的空白。大量的实验证明了我们的方法的有效性和优点。为了使更广泛的研究社区受益,我们在https://github.com/JiangWAV/single-mic-SSL上开源了我们的代码。
摘要:Accurately estimating sound source positions is crucial for robot audition. However, existing sound source localization methods typically rely on a microphone array with at least two spatially preconfigured microphones. This requirement hinders the applicability of microphone-based robot audition systems and technologies. To alleviate these challenges, we propose an online sound source localization method that uses a single microphone mounted on a mobile robot in reverberant environments. Specifically, we develop a lightweight neural network model with only 43k parameters to perform real-time distance estimation by extracting temporal information from reverberant signals. The estimated distances are then processed using an extended Kalman filter to achieve online sound source localization. To the best of our knowledge, this is the first work to achieve online sound source localization using a single microphone on a moving robot, a gap that we aim to fill in this work. Extensive experiments demonstrate the effectiveness and merits of our approach. To benefit the broader research community, we have open-sourced our code at https://github.com/JiangWAV/single-mic-SSL.
【21】Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching
备注:Accepted at Interspeech 2025
摘要:None
摘要:Dysarthria is a neurological disorder that significantly impairs speech intelligibility, often rendering affected individuals unable to communicate effectively. This necessitates the development of robust dysarthric-to-regular speech conversion techniques. In this work, we investigate the utility and limitations of self-supervised learning (SSL) features and their quantized representations as an alternative to mel-spectrograms for speech generation. Additionally, we explore methods to mitigate speaker variability by generating clean speech in a single-speaker voice using features extracted from WavLM. To this end, we propose a fully non-autoregressive approach that leverages Conditional Flow Matching (CFM) with Diffusion Transformers to learn a direct mapping from dysarthric to clean speech. Our findings highlight the effectiveness of discrete acoustic units in improving intelligibility while achieving faster convergence compared to traditional mel-spectrogram-based approaches.
【22】VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schrödinger Bridge
备注:Accepted by Interspeech 2025
摘要:为了探索利用空间线索从图像中产生立体声歌唱声音与房间混响的潜在优势,我们介绍了VS-Singer,视觉引导模型,旨在从场景图像产生立体声歌唱声音与房间混响。VS-Singer包括三个模块:首先,模态交互网络将空间特征整合到文本编码中,以创建富含空间信息的语言表示。其次,解码器采用一致性薛定谔桥,以促进一步样本生成。此外,我们利用SFE模块,以提高视听匹配的一致性。据我们所知,这项研究是第一个结合立体声歌唱声音合成与视觉声学匹配在一个统一的框架。实验结果表明,VS-Singer可以有效地生成立体声歌唱声音,在一个单一的步骤,符合场景的角度。
摘要:To explore the potential advantages of utilizing spatial cues from images for generating stereo singing voices with room reverberation, we introduce VS-Singer, a vision-guided model designed to produce stereo singing voices with room reverberation from scene images. VS-Singer comprises three modules: firstly, a modal interaction network integrates spatial features into text encoding to create a linguistic representation enriched with spatial information. Secondly, the decoder employs a consistency Schr\"odinger bridge to facilitate one-step sample generation. Moreover, we utilize the SFE module to improve the consistency of audio-visual matching. To our knowledge, this study is the first to combine stereo singing voice synthesis with visual acoustic matching within a unified framework. Experimental results demonstrate that VS-Singer can effectively generate stereo singing voices that align with the scene perspective in a single step.
【23】Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion
备注:Accepted to ACL 2025 Findings
摘要:基于人工智能的音乐生成工具的快速发展正在彻底改变音乐行业,但也给艺术家、版权所有者和提供商带来了挑战。这需要可靠的方法来检测这种AI生成的内容。然而,现有的检测器,依赖于音频或歌词,面临着关键的实际限制:基于音频的检测器无法推广到新的或看不见的发电机,容易受到音频扰动;基于歌词的方法需要干净的格式和准确的歌词,在实践中不可用。为了克服这些限制,我们提出了一种新颖的,实际接地的方法:多模式,模块化后期融合管道,结合自动转录的歌词和语音功能捕捉歌词相关的信息在音频。通过直接依赖于来自音频的抒情方面,我们的方法增强了鲁棒性,减轻了对低级伪影的敏感性,并实现了实用性。实验表明,我们的方法,DE检测,优于现有的基于歌词的检测器,同时也更强大的音频扰动。因此,它为在现实世界中检测人工智能生成的音乐提供了一种有效、强大的解决方案。我们的代码可在https://github.com/deezer/robust-AI-lyrics-detection上获得。
摘要:The rapid advancement of AI-based music generation tools is revolutionizing the music industry but also posing challenges to artists, copyright holders, and providers alike. This necessitates reliable methods for detecting such AI-generated content. However, existing detectors, relying on either audio or lyrics, face key practical limitations: audio-based detectors fail to generalize to new or unseen generators and are vulnerable to audio perturbations; lyrics-based methods require cleanly formatted and accurate lyrics, unavailable in practice. To overcome these limitations, we propose a novel, practically grounded approach: a multimodal, modular late-fusion pipeline that combines automatically transcribed sung lyrics and speech features capturing lyrics-related information within the audio. By relying on lyrical aspects directly from audio, our method enhances robustness, mitigates susceptibility to low-level artifacts, and enables practical applicability. Experiments show that our method, DE-detect, outperforms existing lyrics-based detectors while also being more robust to audio perturbations. Thus, it offers an effective, robust solution for detecting AI-generated music in real-world scenarios. Our code is available at https://github.com/deezer/robust-AI-lyrics-detection.
【24】Early Attentive Sparsification Accelerates Neural Speech Transcription
摘要:基于变换器的神经语音处理已经达到了最先进的性能。由于语音音频信号是高度可压缩的,在这里,我们寻求加速神经语音转录的时域信号稀疏化早期的神经编码阶段,利用自注意机制的可解释性的Transformer音频编码器。使用Whisper系列模型,我们在稀疏化阶段(某个编码器层)和压缩比(稀疏度)的联合空间上执行系统的架构搜索。我们发现,在1%的准确度下降下,最佳结果解决方案选择在早期编码阶段将隐藏状态稀疏化到40-60%的稀疏度,从而在Nvidia GPU上实现高达1.6倍的英语语音转录任务运行时加速,而无需任何微调。
摘要:Transformer-based neural speech processing has achieved state-of-the-art performance. Since speech audio signals are known to be highly compressible, here we seek to accelerate neural speech transcription by time-domain signal sparsification early in the neural encoding stage, taking advantage of the interpretability of the self-attention mechanism in transformer audio encoders. With the Whisper family of models, we perform a systematic architecture search over the joint space of sparsification stage (a certain encoder layer) and compression ratio (sparsity). We found that the best resulting solutions under 1% accuracy degradation choose to sparsify the hidden state to 40-60% sparsity at an early encoding stage, and thereby achieve up to 1.6x runtime acceleration in English speech transcription tasks on Nvidia GPUs without any fine-tuning.
【25】Sonic4D: Spatial Audio Generation for Immersive 4D Scene Exploration
备注: 17 pages, 7 figures. Project page: this https URL
摘要:4D生成的最新进展已经证明了其在合成动态3D场景的真实感渲染方面的显着能力。然而,尽管实现了令人印象深刻的视觉性能,但几乎所有现有的方法都忽略了与对应的4D场景对齐的空间音频的生成,这对真正沉浸式视听体验造成了重大限制。为了缓解这个问题,我们提出了Sonic4D,一个新的框架,使空间音频生成沉浸式探索的4D场景。具体来说,我们的方法由三个阶段组成:1)为了从单目视频中捕获动态视觉内容和原始听觉信息,我们首先采用预先训练的专家模型来生成4D场景及其相应的单声道音频。2)随后,为了将单声道音频转换为空间音频,我们在4D场景中定位和跟踪声源,其中通过像素级视觉接地策略估计它们在不同时间戳的3D空间坐标。3)基于估计的声源位置,我们进一步合成合理的空间音频,不同的观点和时间戳使用基于物理的模拟。大量的实验表明,我们提出的方法生成逼真的空间音频与合成的4D场景在一个训练自由的方式一致,显着提高用户的沉浸式体验。生成的音频和视频示例可在https://x-drunker.github.io/Sonic4D-project-page上获得。
摘要:Recent advancements in 4D generation have demonstrated its remarkable capability in synthesizing photorealistic renderings of dynamic 3D scenes. However, despite achieving impressive visual performance, almost all existing methods overlook the generation of spatial audio aligned with the corresponding 4D scenes, posing a significant limitation to truly immersive audiovisual experiences. To mitigate this issue, we propose Sonic4D, a novel framework that enables spatial audio generation for immersive exploration of 4D scenes. Specifically, our method is composed of three stages: 1) To capture both the dynamic visual content and raw auditory information from a monocular video, we first employ pre-trained expert models to generate the 4D scene and its corresponding monaural audio. 2) Subsequently, to transform the monaural audio into spatial audio, we localize and track the sound sources within the 4D scene, where their 3D spatial coordinates at different timestamps are estimated via a pixel-level visual grounding strategy. 3) Based on the estimated sound source locations, we further synthesize plausible spatial audio that varies across different viewpoints and timestamps using physics-based simulation. Extensive experiments have demonstrated that our proposed method generates realistic spatial audio consistent with the synthesized 4D scene in a training-free manner, significantly enhancing the immersive experience for users. Generated audio and video examples are available at https://x-drunker.github.io/Sonic4D-project-page.
【26】 Explainable speech emotion recognition through attentive pooling: insights from attention-based temporal localization
备注:None
摘要:用于语音情感识别(SER)的最先进的Transformer模型依赖于时间特征聚合,然而高级池化方法仍然未被探索。我们系统地对池化策略进行了基准测试,包括多查询多头注意统计池化,它比平均池化实现了3.5个百分点的宏观F1增益。注意力分析显示,15%的帧捕捉了80%的情感线索,揭示了情感信息的局部模式。高关注帧的分析表明,非语言发声和hyperarticulated音素不成比例地优先在池,反映人类的感知策略。我们的研究结果的位置注意池作为一个性能SER机制和生物学上合理的工具,可解释的情绪定位。在自然条件下的Interspeech 2025语音情感识别挑战赛中,我们的方法获得了0.3649的宏观F1分数。
摘要:State-of-the-art transformer models for Speech Emotion Recognition (SER) rely on temporal feature aggregation, yet advanced pooling methods remain underexplored. We systematically benchmark pooling strategies, including Multi-Query Multi-Head Attentive Statistics Pooling, which achieves a 3.5 percentage point macro F1 gain over average pooling. Attention analysis shows 15 percent of frames capture 80 percent of emotion cues, revealing a localized pattern of emotional information. Analysis of high-attention frames reveals that non-linguistic vocalizations and hyperarticulated phonemes are disproportionately prioritized during pooling, mirroring human perceptual strategies. Our findings position attentive pooling as both a performant SER mechanism and a biologically plausible tool for explainable emotion localization. On Interspeech 2025 Speech Emotion Recognition in Naturalistic Conditions Challenge, our approach obtained a macro F1 score of 0.3649.
机器翻译由腾讯交互翻译提供,仅供参考
