今日论文合集:cs.SD语音22篇,eess.AS音频处理24篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Accent-Invariant Automatic Speech Recognition via Saliency-Driven Spectrogram Masking
标题:通过显著性驱动频谱图掩蔽的口音不变自动语音识别
链接:https://arxiv.org/abs/2510.09528

作者:Mohammad Hossein Sameti, Sepehr Harfi Moridani, Ali Zarean, Hossein Sameti
备注:Submitted to ICASSP 2026
摘要:预先训练的基于transformer的模型显著提高了自动语音识别(ASR),但它们对口音和方言变化仍然敏感,导致英语和波斯语等语言多样性语言的单词错误率(WER)升高。为了解决这一挑战,我们提出了一个口音不变的ASR框架,将口音和方言分类集成到识别管道中。我们的方法包括训练一个基于谱图的分类器来捕获特定口音的线索,掩盖对其预测最有影响的区域,并使用掩蔽的谱图进行数据增强。这增强了ASR模型对口音变化的鲁棒性。我们使用英语和波斯语的讲话的方法进行评估。对于波斯语,我们引入了一个新收集的跨越多个区域口音的数据集,建立了波斯语ASR中口音变化的第一个系统基准,填补了多语言语音研究的关键空白,并为未来低资源,语言多样性语言的研究提供了基础。Whisper模型的实验结果表明,我们的掩蔽和增强策略在英语和波斯语设置中产生了大量的WER减少,证实了该方法的有效性。这项研究推进了能够适应口音和方言多样性的多语言ASR系统的开发。代码和数据集可在https://github.com/MH-Sameti/Accent_invariant_ASR上公开获取
摘要:Pre-trained transformer-based models have significantly advanced automatic speech recognition (ASR), yet they remain sensitive to accent and dialectal variations, resulting in elevated word error rates (WER) in linguistically diverse languages such as English and Persian. To address this challenge, we propose an accent-invariant ASR framework that integrates accent and dialect classification into the recognition pipeline. Our approach involves training a spectrogram-based classifier to capture accent-specific cues, masking the regions most influential to its predictions, and using the masked spectrograms for data augmentation. This enhances the robustness of ASR models against accent variability. We evaluate the method using both English and Persian speech. For Persian, we introduce a newly collected dataset spanning multiple regional accents, establishing the first systematic benchmark for accent variation in Persian ASR that fills a critical gap in multilingual speech research and provides a foundation for future studies on low-resource, linguistically diverse languages. Experimental results with the Whisper model demonstrate that our masking and augmentation strategy yields substantial WER reductions in both English and Persian settings, confirming the effectiveness of the approach. This research advances the development of multilingual ASR systems that are resilient to accent and dialect diversity. Code and dataset are publicly available at: https://github.com/MH-Sameti/Accent_invariant_ASR


【2】WildElder: A Chinese Elderly Speech Dataset from the Wild with Fine-Grained Manual Annotations
标题:WildElder:来自野外的中国老年人语音数据集,带有细粒度手动注释
链接:https://arxiv.org/abs/2510.09344

作者:Hui Wang, Jiaming Zhou, Jiabei He, Haoqin Sun, Yong Qin
摘要:由于与年龄相关的变化,如发音较慢和声音震颤,老年人的语音对自动处理提出了独特的挑战。现有的中国数据集大多在受控环境中记录,限制了其多样性和现实世界的适用性。为了解决这一差距,我们提出了WildElder,一个从在线视频中收集的普通话老年人语音语料库,并使用细粒度的手动注释进行了丰富,包括转录,说话人年龄,性别和口音强度。WildElder将野外数据的真实性与专家策展相结合,实现了对自动语音识别和说话人分析的强大研究。实验结果揭示了老年人语音识别的困难和WildElder作为一个具有挑战性的新基准的潜力。数据集和代码可在https://github.com/NKU-HLT/WildElder上获得。
摘要:Elderly speech poses unique challenges for automatic processing due to age-related changes such as slower articulation and vocal tremors. Existing Chinese datasets are mostly recorded in controlled environments, limiting their diversity and real-world applicability. To address this gap, we present WildElder, a Mandarin elderly speech corpus collected from online videos and enriched with fine-grained manual annotations, including transcription, speaker age, gender, and accent strength. Combining the realism of in-the-wild data with expert curation, WildElder enables robust research on automatic speech recognition and speaker profiling. Experimental results reveal both the difficulties of elderly speech recognition and the potential of WildElder as a challenging new benchmark. The dataset and code are available at https://github.com/NKU-HLT/WildElder.


【3】SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
标题:SynthVC:利用合成数据进行端到端低延迟流媒体语音转换
链接:https://arxiv.org/abs/2510.09245

作者:Zhao Guo, Ziqian Ning, Guobin Ma, Lei Xie
备注:Accepted by NCMMSC2025
摘要:语音转换(VC)旨在修改说话者的音色,同时保留语言内容。虽然最近的VC模型实现了强大的性能,但由于高延迟,对ASR模块的依赖或复杂的扬声器解缠,大多数在实时流媒体场景中挣扎,这通常会导致音色泄漏或自然度下降。我们提出了SynthVC,一个流媒体端到端VC框架,直接学习扬声器音色转换从合成的并行数据生成的预训练的zero-shot VC模型。这种设计消除了对显式内容-说话者分离或识别模块的需要。基于神经音频编解码器架构,SynthVC支持低延迟流推理,具有高输出保真度。实验结果表明,SynthVC在自然度和说话人相似度方面优于基准流VC系统,实现了仅77.1 ms的端到端延迟。
摘要:Voice Conversion (VC) aims to modify a speaker's timbre while preserving linguistic content. While recent VC models achieve strong performance, most struggle in real-time streaming scenarios due to high latency, dependence on ASR modules, or complex speaker disentanglement, which often results in timbre leakage or degraded naturalness. We present SynthVC, a streaming end-to-end VC framework that directly learns speaker timbre transformation from synthetic parallel data generated by a pre-trained zero-shot VC model. This design eliminates the need for explicit content-speaker separation or recognition modules. Built upon a neural audio codec architecture, SynthVC supports low-latency streaming inference with high output fidelity. Experimental results show that SynthVC outperforms baseline streaming VC systems in both naturalness and speaker similarity, achieving an end-to-end latency of just 77.1 ms.


【4】FLToP CTC: Frame-Level Token Pruning via Relative Threshold for Efficient and Memory-Saving Decoding on Diverse Platforms
标题:FLToP CTC:通过相对阈值进行帧级令牌修剪,以在不同平台上进行高效且节省内存的解码
链接:https://arxiv.org/abs/2510.09085

作者:Atul Shree, Harshith Jupuru
备注:5 pages, 5 figures
摘要:基于CTC的ASR系统在资源有限的环境中面临计算和内存瓶颈。传统的CTC解码器,在系统中需要高达90%的处理时间(例如,wav 2 vec 2-large在L4 GPU上),由于详尽的令牌级操作而面临效率低下的问题。本文介绍了一种新的解码算法,采用帧级令牌修剪的相对阈值概率的指导下,连接时间分类(FLToP CTC)的帧级令牌修剪。通过动态消除每帧的低概率令牌,FLToP CTC减少了计算和内存需求,同时保持可忽略的WER降级。在LibriSpeech上,FLToP CTC与标准CTC解码器相比,实现了10.5倍的运行时加速和2.78倍的内存减少。它的简单性使其能够跨平台(CPU,GPU等)无缝集成到CTC解码器中。FLToP CTC解决了CTC瓶颈,为资源有限的环境和实时应用提供了可扩展性,增强了语音识别的可访问性和效率。
摘要:CTC-based ASR systems face computational and memory bottlenecks in resource-limited environments. Traditional CTC decoders, requiring up to 90% of processing time in systems (e.g., wav2vec2-large on L4 GPUs), face inefficiencies due to exhaustive token-level operations. This paper introduces Frame Level Token Pruning for Connectionist Temporal Classification (FLToP CTC), a novel decoding algorithm that employs frame-level token pruning guided by a relative threshold probability. By dynamically eliminating low-probability tokens per frame, FLToP CTC reduces compute and memory demands while maintaining negligible WER degradation. On LibriSpeech, FLToP CTC achieves a 10.5x runtime speedup and 2.78x memory reduction versus standard CTC decoders. Its simplicity enables seamless integration into CTC decoders across platforms (CPUs, GPUs, etc.). FLToP CTC addresses CTC bottlenecks, offering scalability for resource-limited environments and realtime applications, enhancing speech recognition accessibility and efficiency.


【5】Emotion-Disentangled Embedding Alignment for Noise-Robust and Cross-Corpus Speech Emotion Recognition
标题:用于噪音稳健和跨数据库语音情感识别的描述去纠缠嵌入对齐
链接:https://arxiv.org/abs/2510.09072

作者:Upasana Tiwari, Rupayan Chakraborty, Sunil Kumar Kopparapu
备注:13 pages, 1 figure
摘要:语音情感识别在现实世界中的有效性往往受到噪声环境和数据集差异的阻碍。本文介绍了一个两步的方法,以提高语音情感识别模型的鲁棒性和泛化能力,通过改进的表示学习。首先,我们的模型采用EDRL(推理-分解表示学习)来提取特定于类的判别特征,同时保留跨情感类别的共享相似性。接下来,MEA(多块嵌入对齐)通过将这些表示投影到联合判别潜在子空间中来细化这些表示,该子空间最大化与原始语音输入的协方差。学习的EDRL-MEA嵌入随后用于使用来自公开可用数据集的干净样本来训练情感分类器,并且在看不见的噪声和跨语料库语音样本上进行评估。在这些具有挑战性的条件下,性能的改善证明了所提出的方法的有效性。
摘要:Effectiveness of speech emotion recognition in real-world scenarios is often hindered by noisy environments and variability across datasets. This paper introduces a two-step approach to enhance the robustness and generalization of speech emotion recognition models through improved representation learning. First, our model employs EDRL (Emotion-Disentangled Representation Learning) to extract class-specific discriminative features while preserving shared similarities across emotion categories. Next, MEA (Multiblock Embedding Alignment) refines these representations by projecting them into a joint discriminative latent subspace that maximizes covariance with the original speech input. The learned EDRL-MEA embeddings are subsequently used to train an emotion classifier using clean samples from publicly available datasets, and are evaluated on unseen noisy and cross-corpus speech samples. Improved performance under these challenging conditions demonstrates the effectiveness of the proposed method.


【6】MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation
标题:MMAudioSep:驯服视频到音频生成模型以实现视频/文本查询声音分离
链接:https://arxiv.org/abs/2510.09065

作者:Akira Takahashi, Shusuke Takahashi, Yuki Mitsufuji
备注:4 pages, 4 figures, 2 tables
摘要:我们介绍了MMAudioSep,一个生成模型的视频/文本查询的声音分离,是建立在一个预训练的视频到音频模型。通过利用通过预训练的音频生成模型学习的关于视频/文本和音频之间的关系的知识,我们可以更有效地训练模型,即,该模型不需要从头开始训练。我们通过将MMAudioSep与现有的分离模型(包括基于确定性和生成方法的模型)进行比较来评估MMAudioSep的性能,并发现它优于基线模型。此外,我们证明,即使在通过微调获得声音分离功能后,该模型仍保留了原始视频到音频生成的能力。这突出了基础声音生成模型的潜力,以通过与声音相关的下游任务。我们的代码可在www.example.com上获得。
摘要:We introduce MMAudioSep, a generative model for video/text-queried sound separation that is founded on a pretrained video-to-audio model. By leveraging knowledge about the relationship between video/text and audio learned through a pretrained audio generative model, we can train the model more efficiently, i.e., the model does not need to be trained from scratch. We evaluate the performance of MMAudioSep by comparing it to existing separation models, including models based on both deterministic and generative approaches, and find it is superior to the baseline models. Furthermore, we demonstrate that even after acquiring functionality for sound separation via fine-tuning, the model retains the ability for original video-to-audio generation. This highlights the potential of foundational sound generation models to be adopted for sound-related downstream tasks. Our code is available at https://github.com/sony/mmaudiosep.


【7】O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion
标题:O_O-VC:合成数据驱动的一对一对齐,实现任意语音转换
链接:https://arxiv.org/abs/2510.09061

作者:Huu Tuong Tu, Huan Vu, cuong tien nguyen, Dien Hy Ngo, Nguyen Thi Thu Trang
备注:EMNLP 2025
摘要:传统的语音转换(VC)方法通常试图将说话人身份和语言信息分离成不同的表示,然后将其组合以重建音频。然而,有效地解开这些因素仍然具有挑战性,往往导致培训过程中的信息丢失。在本文中,我们提出了一种新的方法,利用合成语音数据生成的高质量,预训练的多扬声器文本到语音(TTS)模型。具体地,共享相同语言内容但说话者身份不同的合成数据对被用作输入-输出对以训练语音转换模型。这使模型能够学习源和目标语音之间的直接映射,有效地捕捉特定于说话者的特征,同时保留语言内容。此外,我们还引入了一种灵活的任意语音转换训练策略,可以很好地推广到看不见的说话者和新语言,增强了zero-shot场景中的适应性和性能。实验结果表明,该方法在词错误率上相对降低了16.35%,在说话人余弦相似度上相对提高了5.91%,优于几种最先进的方法。语音转换示例可访问:https://oovc-emnlp-2025.github.io/
摘要:Traditional voice conversion (VC) methods typically attempt to separate speaker identity and linguistic information into distinct representations, which are then combined to reconstruct the audio. However, effectively disentangling these factors remains challenging, often leading to information loss during training. In this paper, we propose a new approach that leverages synthetic speech data generated by a high-quality, pretrained multispeaker text-to-speech (TTS) model. Specifically, synthetic data pairs that share the same linguistic content but differ in speaker identity are used as input-output pairs to train the voice conversion model. This enables the model to learn a direct mapping between source and target voices, effectively capturing speaker-specific characteristics while preserving linguistic content. Additionally, we introduce a flexible training strategy for any-to-any voice conversion that generalizes well to unseen speakers and new languages, enhancing adaptability and performance in zero-shot scenarios. Our experiments show that our proposed method achieves a 16.35% relative reduction in word error rate and a 5.91% improvement in speaker cosine similarity, outperforming several state-of-the-art methods. Voice conversion samples can be accessed at: https://oovc-emnlp-2025.github.io/


【8】Déréverbération non-supervisée de la parole par modèle hybride
标题:修改非监督假释模式混合
链接:https://arxiv.org/abs/2510.09025

作者:Louis Bahrman (IDS, S2A), Mathieu Fontaine (IDS, S2A), Gaël Richard (IDS, S2A)
备注:in French language
摘要:本文介绍了一种新的训练策略,以改善语音去混响系统在无监督的方式只使用混响语音。大多数现有的算法依赖于成对的干/混响数据,这是很难获得的。我们的方法使用有限的声学信息,如混响时间(RT 60),训练去混响系统。实验结果表明,我们的方法实现了更一致的性能在各种客观指标比国家的最先进的。
摘要:This paper introduces a new training strategy to improve speech dereverberation systems in an unsupervised manner using only reverberant speech. Most existing algorithms rely on paired dry/reverberant data, which is difficult to obtain. Our approach uses limited acoustic information, like the reverberation time (RT60), to train a dereverberation system. Experimental results demonstrate that our method achieves more consistent performance across various objective metrics than the state-of-the-art.


【9】DiTSinger: Scaling Singing Voice Synthesis with Diffusion Transformer and Implicit Alignment
标题:DiTSinger:使用扩散Transformer和隐式对齐缩放歌唱声音合成
链接:https://arxiv.org/abs/2510.09016

作者:Zongcai Du, Guilin Deng, Xiaofeng Guo, Xin Gao, Linke Li, Kaichang Cheng, Fubo Han, Siyu Yang, Peng Liu, Pan Zhong, Qiang Fu
备注:under review
摘要:基于扩散的歌唱声音合成(SVS)的最新进展表现出很强的表现力,但仍然受到数据稀缺和模型可扩展性的限制。我们引入了一个两阶段的管道:通过将固定的旋律与不同的LLM生成的歌词配对来构建人类演唱录音的紧凑种子集,并训练特定于旋律的模型来合成超过500小时的高质量中文演唱数据。在此语料库的基础上,我们提出了DiTSinger,一个扩散Transformer与ROPE和qk规范,系统地缩放的深度,宽度和分辨率,以提高保真度。此外,我们设计了一个隐式对齐机制,避免音素级的持续时间标签,通过约束音素到声学注意字符级跨度内,从而提高噪声或不确定的对齐下的鲁棒性。大量的实验验证了我们的方法可以实现可扩展的,无干扰的,高保真的SVS。
摘要:Recent progress in diffusion-based Singing Voice Synthesis (SVS) demonstrates strong expressiveness but remains limited by data scarcity and model scalability. We introduce a two-stage pipeline: a compact seed set of human-sung recordings is constructed by pairing fixed melodies with diverse LLM-generated lyrics, and melody-specific models are trained to synthesize over 500 hours of high-quality Chinese singing data. Building on this corpus, we propose DiTSinger, a Diffusion Transformer with RoPE and qk-norm, systematically scaled in depth, width, and resolution for enhanced fidelity. Furthermore, we design an implicit alignment mechanism that obviates phoneme-level duration labels by constraining phoneme-to-acoustic attention within character-level spans, thereby improving robustness under noisy or uncertain alignments. Extensive experiments validate that our approach enables scalable, alignment-free, and high-fidelity SVS.


【10】VM-UNSSOR: Unsupervised Neural Speech Separation Enhanced by Higher-SNR Virtual Microphone Arrays
标题:VM-UNSsor:通过更高的SNR虚拟麦克风阵列增强无监督神经语音分离
链接:https://arxiv.org/abs/2510.08914

作者:Shulin He, Zhong-Qiu Wang
摘要:盲语音分离(BSS)的目的是在未知阵列几何和房间冲激响应的情况下,从多通道、多说话人混合信号中恢复出多个语音源。在无监督设置中,干净的目标语音无法用于模型训练,UNSSOR提出了一种混合一致性(MC)损失,用于在超定训练混合物上训练深度神经网络(DNN),以实现无监督语音分离。然而,当训练混合信号的麦克风数量减少时,MC约束减弱,分离性能急剧下降。为了解决这个问题,我们提出了VM-UNSSOR,增加了观察到的训练混合信号记录的麦克风与几个更高的SNR虚拟麦克风(VM)信号,这是通过应用线性空间解混器(如IVA和空间聚类)到观察到的训练混合信号。作为所观察到的混合物的线性投影,虚拟麦克风信号通常可以增加每个源的SNR,并且可以被利用来计算额外的MC损耗,以改善UNSSOR并解决UNSSOR中的频率排列问题。在SMS-WSJ数据集上,在超定六麦克风、两扬声器分离设置中,VM-UNSSOR达到17.1 dB SI-SDR,而UNSSOR仅获得14.7 dB;在确定的两麦克风、两扬声器情况下,UNSSOR塌陷至-2.7 dB SI-SDR,而VM-UNSSOR达到10.7 dB。
摘要:Blind speech separation (BSS) aims to recover multiple speech sources from multi-channel, multi-speaker mixtures under unknown array geometry and room impulse responses. In unsupervised setup where clean target speech is not available for model training, UNSSOR proposes a mixture consistency (MC) loss for training deep neural networks (DNN) on over-determined training mixtures to realize unsupervised speech separation. However, when the number of microphones of the training mixtures decreases, the MC constraint weakens and the separation performance falls dramatically. To address this, we propose VM-UNSSOR, augmenting the observed training mixture signals recorded by a limited number of microphones with several higher-SNR virtual-microphone (VM) signals, which are obtained by applying linear spatial demixers (such as IVA and spatial clustering) to the observed training mixtures. As linear projections of the observed mixtures, the virtual-microphone signals can typically increase the SNR of each source and can be leveraged to compute extra MC losses to improve UNSSOR and address the frequency permutation problem in UNSSOR. On the SMS-WSJ dataset, in the over-determined six-microphone, two-speaker separation setup, VM-UNSSOR reaches 17.1 dB SI-SDR, while UNSSOR only obtains 14.7 dB; and in the determined two-microphone, two-speaker case, UNSSOR collapses to -2.7 dB SI-SDR, while VM-UNSSOR achieves 10.7 dB.


【11】ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling
标题:ControlAudio:通过渐进扩散模型处理文本引导、定时指示和可理解的音频生成
链接:https://arxiv.org/abs/2510.08878

作者:Yuxuan Jiang, Zehua Chen, Zeqian Ju, Yusheng Dai, Weibei Dou, Jun Zhu
备注:18 pages, 8 tables, 5 figures
摘要:使用细粒度控制信号生成文本到音频(TTA),例如,精确的定时控制或可理解的语音内容在最近的工作中已经被探索。然而,受数据稀缺性的限制,它们的大规模发电性能仍然受到影响。在这项研究中,我们将可控TTA生成重新定义为多任务学习问题,并引入了一种渐进扩散建模方法ControlAudio。我们的方法巧妙地适合分布条件下更细粒度的信息,包括文本,时间和音素特征,通过一步一步的策略。首先,我们提出了一个数据构建方法跨越注释和模拟,增加条件信息的文本,时间和音素的顺序。其次,在模型训练阶段,我们在大规模文本-音频对上预训练扩散Transformer(DiT),实现可扩展的TTA生成,然后增量地将时序和音素特征与统一的语义表示相结合,扩展可控性。最后,在推理阶段,我们提出了逐步引导的生成,它依次强调更细粒度的信息,内在地与DiT的粗到细采样性质相一致。大量的实验表明,ControlAudio在时间准确性和语音清晰度方面达到了最先进的性能,在客观和主观评价方面都明显优于现有方法。演示示例可在以下网址获得:https://control-audio.github.io/Control-Audio。
摘要:Text-to-audio (TTA) generation with fine-grained control signals, e.g., precise timing control or intelligible speech content, has been explored in recent works. However, constrained by data scarcity, their generation performance at scale is still compromised. In this study, we recast controllable TTA generation as a multi-task learning problem and introduce a progressive diffusion modeling approach, ControlAudio. Our method adeptly fits distributions conditioned on more fine-grained information, including text, timing, and phoneme features, through a step-by-step strategy. First, we propose a data construction method spanning both annotation and simulation, augmenting condition information in the sequence of text, timing, and phoneme. Second, at the model training stage, we pretrain a diffusion transformer (DiT) on large-scale text-audio pairs, achieving scalable TTA generation, and then incrementally integrate the timing and phoneme features with unified semantic representations, expanding controllability. Finally, at the inference stage, we propose progressively guided generation, which sequentially emphasizes more fine-grained information, aligning inherently with the coarse-to-fine sampling nature of DiT. Extensive experiments show that ControlAudio achieves state-of-the-art performance in terms of temporal accuracy and speech clarity, significantly outperforming existing methods on both objective and subjective evaluations. Demo samples are available at: https://control-audio.github.io/Control-Audio.


【12】Audible Networks: Deconstructing and Manipulating Sounds with Deep Non-Negative Autoencoders
标题:可听网络:使用深度非负自动编码器解构和操纵声音
链接:https://arxiv.org/abs/2510.08816

作者:Juan José Burred, Carmine-Emanuele Cella
摘要:我们建议使用非负自动编码器(NAE)进行声音解构和用户引导的声音操作,以达到创造性的目的。NAE为传统的基于非负矩阵分解(NMF)的方法提供了一种通用且可扩展的扩展,用于可解释的音频分解。通过执行非负约束,通过投影梯度下降,我们得到的分解,内部权重和激活可以直接解释为频谱形状和时间包络,其中组件本身可以作为单独的声音事件听。特别是,多层Deep NAE架构能够实现具有可调粒度级别的分层表示,允许在多个抽象级别解构声音:从高级音符包络到细粒度频谱细节。该框架实现了广泛的新的表达,可控和随机的声音转换。我们介绍了新的操作,包括跨组件和跨层合成,分层解构,和几个随机化的策略,控制音色和事件密度。通过可视化和重新合成的实际例子,我们展示了如何NAES可以作为灵活和可解释的工具,基于对象的声音编辑。
摘要:We propose the use of Non-Negative Autoencoders (NAEs) for sound deconstruction and user-guided manipulation of sounds for creative purposes. NAEs offer a versatile and scalable extension of traditional Non-Negative Matrix Factorization (NMF)-based approaches for interpretable audio decomposition. By enforcing non-negativity constraints through projected gradient descent, we obtain decompositions where internal weights and activations can be directly interpreted as spectral shapes and temporal envelopes, and where components can themselves be listened to as individual sound events. In particular, multi-layer Deep NAE architectures enable hierarchical representations with an adjustable level of granularity, allowing sounds to be deconstructed at multiple levels of abstraction: from high-level note envelopes down to fine-grained spectral details. This framework enables a wide new range of expressive, controllable, and randomized sound transformations. We introduce novel manipulation operations including cross-component and cross-layer synthesis, hierarchical deconstructions, and several randomization strategies that control timbre and event density. Through visualizations and resynthesis of practical examples, we demonstrate how NAEs can serve as flexible and interpretable tools for object-based sound editing.


【13】Hierarchical Self-Supervised Representation Learning for Depression Detection from Speech
标题:分层自监督表示学习用于语音抑郁检测
链接:https://arxiv.org/abs/2510.08593

作者:Yuxin Li, Eng Siong Chng, Cuntai Guan
摘要:基于语音的抑郁检测(SDD)是传统临床评估的一种有前途的非侵入性替代方法。然而,随着时间的推移,它仍然受到提取有意义的特征和捕获稀疏,异构抑郁线索的困难的限制。预训练的自监督学习(SSL)模型(如WavLM)提供了丰富的多层语音表示,但大多数现有的SDD方法仅依赖于最后一层或搜索单个最佳性能。这些方法通常过度拟合特定的数据集,并且无法利用检测微妙和持续抑郁信号所需的完整层次结构。   为了应对这一挑战,我们提出了HAREN-CTC,一种新的架构,它集成了多层SSL功能,使用多任务学习框架内的交叉注意,结合连接主义时间分类损失来处理稀疏的时间监督。HAREN-CTC包括两个关键模块:一个分层自适应聚类模块,将SSL特征重组为互补的嵌入,以及一个跨模态融合模块,通过交叉注意力对层间依赖关系进行建模。CTC目标支持对齐感知训练,允许模型跟踪抑郁言语线索的不规则时间模式。   我们评估HAREN-CTC下的上限设置与标准的数据分割和泛化设置使用五重交叉验证。该模型在DAIC-WOZ上实现了最先进的宏F1分数0.81,在MODMA上实现了0.82,在两种评估场景中均优于先前的方法。
摘要:Speech-based depression detection (SDD) is a promising, non-invasive alternative to traditional clinical assessments. However, it remains limited by the difficulty of extracting meaningful features and capturing sparse, heterogeneous depressive cues over time. Pretrained self-supervised learning (SSL) models such as WavLM provide rich, multi-layer speech representations, yet most existing SDD methods rely only on the final layer or search for a single best-performing one. These approaches often overfit to specific datasets and fail to leverage the full hierarchical structure needed to detect subtle and persistent depression signals.   To address this challenge, we propose HAREN-CTC, a novel architecture that integrates multi-layer SSL features using cross-attention within a multitask learning framework, combined with Connectionist Temporal Classification loss to handle sparse temporal supervision. HAREN-CTC comprises two key modules: a Hierarchical Adaptive Clustering module that reorganizes SSL features into complementary embeddings, and a Cross-Modal Fusion module that models inter-layer dependencies through cross-attention. The CTC objective enables alignment-aware training, allowing the model to track irregular temporal patterns of depressive speech cues.   We evaluate HAREN-CTC under both an upper-bound setting with standard data splits and a generalization setting using five-fold cross-validation. The model achieves state-of-the-art macro F1-scores of 0.81 on DAIC-WOZ and 0.82 on MODMA, outperforming prior methods across both evaluation scenarios.


【14】EGSTalker: Real-Time Audio-Driven Talking Head Generation with Efficient Gaussian Deformation
标题:EGSTalker:具有高效高斯变形的实时音频驱动说话头生成
链接:https://arxiv.org/abs/2510.08587

作者:Tianheng Zhu, Yinfeng Yu, Liejun Wang, Fuchun Sun, Wendong Zheng
备注:Main paper (6 pages). Accepted for publication by IEEE International Conference on Systems, Man, and Cybernetics 2025
摘要:本文介绍了EGSTalker,一个实时音频驱动的说话人头生成框架的基础上三维高斯飞溅(3DGS)。EGSTalker旨在提高速度和视觉保真度,仅需3-5分钟的训练视频即可合成高质量的面部动画。该框架包括两个关键阶段:静态高斯初始化和音频驱动变形。在第一阶段,多分辨率散列三平面和Kolmogorov-Arnold网络(KAN)被用来提取空间特征,并构造一个紧凑的三维高斯表示。在第二阶段,我们提出了一个有效的空间音频注意力(ESAA)模块来融合音频和空间线索,而KAN预测相应的高斯变形。大量的实验表明,EGSTalker实现了渲染质量和唇同步精度相媲美的国家的最先进的方法,同时显着优于他们的推理速度。这些结果突出了EGSTalker的实时多媒体应用的潜力。
摘要:This paper presents EGSTalker, a real-time audio-driven talking head generation framework based on 3D Gaussian Splatting (3DGS). Designed to enhance both speed and visual fidelity, EGSTalker requires only 3-5 minutes of training video to synthesize high-quality facial animations. The framework comprises two key stages: static Gaussian initialization and audio-driven deformation. In the first stage, a multi-resolution hash triplane and a Kolmogorov-Arnold Network (KAN) are used to extract spatial features and construct a compact 3D Gaussian representation. In the second stage, we propose an Efficient Spatial-Audio Attention (ESAA) module to fuse audio and spatial cues, while KAN predicts the corresponding Gaussian deformations. Extensive experiments demonstrate that EGSTalker achieves rendering quality and lip-sync accuracy comparable to state-of-the-art methods, while significantly outperforming them in inference speed. These results highlight EGSTalker's potential for real-time multimedia applications.


【15】Evaluating Hallucinations in Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions
标题:在不同声学条件下评估带有口语按钮的多模式LLM中的幻觉
链接:https://arxiv.org/abs/2510.08581

作者:Hansol Park, Hoseong Ahn, Junwon Moon, Yejin Lee, Kyuhong Shim
摘要:视觉语言模型中的幻觉已经被广泛研究,使用的基准探测图像-文本设置中的可靠性。相比之下,语音查询对多模态幻觉的影响在很大程度上仍未被探索,尽管语音驱动界面的作用越来越大。在这项工作中,我们研究了口语输入如何影响多模态大型语言模型中的幻觉。我们提出了RePOPE-Spk,音频增强扩展的RePOPE基准,查询提供不同的声学条件下的语音。使用RePOPE-Spk,我们系统地评估专有和开源模型。实验结果表明,当查询是口语而不是书面时,幻觉会升级:在干净的语音下错误率增加3%,在环境噪音下增加20%。输入顺序和查询长度进一步影响鲁棒性,而多镜头提示和思维链推理等策略提供了部分但不充分的缓解。这些发现突出了一个关键的和未充分探索的挑战,为构建可靠的语音接口系统开辟了新的方向。
摘要:Hallucinations in vision-language models have been extensively studied using benchmarks that probe reliability in image-text settings. In contrast, the effect of spoken queries on multimodal hallucinations remains largely unexplored, despite the growing role of voice-driven interfaces. In this work, we investigate how spoken input influences hallucinations in multimodal large language models. We present RePOPE-Spk, an audio-augmented extension of the RePOPE benchmark, where queries are provided as speech under diverse acoustic conditions. Using RePOPE-Spk, we systematically evaluate both proprietary and open-source models. Experimental results show that hallucinations escalate when queries are spoken rather than written: error rates increase by 3% under clean speech and by up to 20% with environmental noise. Input order and query length further affect robustness, while strategies such as many-shot prompting and chain-of-thought reasoning offer partial but insufficient mitigation. These findings highlight a critical and underexplored challenge, opening new directions for building reliable voice interface systems.


【16】LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection
标题:LadderSym:用于音乐练习错误检测的多模式交织Transformer
链接:https://arxiv.org/abs/2510.08580

作者:Benjamin Shiue-Hal Chou, Purvish Jajal, Nick John Eliopoulos, James C. Davis, George K. Thiruvathukal, Kristen Yeon-Ji Yun, Yung-Hsiang Lu
备注:Under Submission
摘要:音乐学习者可以从准确检测练习中错误的工具中受益匪浅。现有的方法通常比较音频记录的乐谱,使用的算法或可学习的模型。本文介绍了\textit{LadderSym},一种新的基于transformer的音乐错误检测方法。\textit{LadderSym}由关于最先进方法的两个关键观察结果指导:(1)后期融合限制了流间对齐和跨模态比较能力;以及(2)对乐谱音频的依赖在频谱中引入了模糊性,降低了具有并发音符的音乐的性能。为了解决这些限制,\textit{LadderSym}引入了(1)一个具有流间对齐模块的双流编码器,以提高音频比较能力和错误检测F1分数,以及(2)一种多模式策略,通过将符号表示作为解码器提示,减少歧义并提高F1分数,从而利用音频和符号分数。我们通过测量每个音符类别的F1得分,在\textit{MAESTRO-E}和\textit{CocoChorales-E}数据集上评估我们的方法。与之前的技术水平相比,\textit{LadderSym}在\textit{MAESTRO-E}上对未命中音符的F1检测能力提高了一倍多(26.8\% $\rightarrow $56.3\%),并将额外音符检测能力提高了14.4点(72.0\% $\rightarrow $86.4\%)。在\textit{CocoChorales-E}上观察到类似的增益。这项工作介绍了有关比较模型的一般见解,这些模型可以为强化学习,人类技能评估和模型评估的序列评估任务提供信息。
摘要:Music learners can greatly benefit from tools that accurately detect errors in their practice. Existing approaches typically compare audio recordings to music scores using heuristics or learnable models. This paper introduces \textit{LadderSym}, a novel Transformer-based method for music error detection. \textit{LadderSym} is guided by two key observations about the state-of-the-art approaches: (1) late fusion limits inter-stream alignment and cross-modality comparison capability; and (2) reliance on score audio introduces ambiguity in the frequency spectrum, degrading performance in music with concurrent notes. To address these limitations, \textit{LadderSym} introduces (1) a two-stream encoder with inter-stream alignment modules to improve audio comparison capabilities and error detection F1 scores, and (2) a multimodal strategy that leverages both audio and symbolic scores by incorporating symbolic representations as decoder prompts, reducing ambiguity and improving F1 scores. We evaluate our method on the \textit{MAESTRO-E} and \textit{CocoChorales-E} datasets by measuring the F1 score for each note category. Compared to the previous state of the art, \textit{LadderSym} more than doubles F1 for missed notes on \textit{MAESTRO-E} (26.8\% $\rightarrow$ 56.3\%) and improves extra note detection by 14.4 points (72.0\% $\rightarrow$ 86.4\%). Similar gains are observed on \textit{CocoChorales-E}. This work introduces general insights about comparison models that could inform sequence evaluation tasks for reinforcement Learning, human skill assessment, and model evaluation.


【17】Effects of automotive microphone frequency response characteristics and noise conditions on speech and ASR quality -- an experimental evaluation
标题:汽车麦克风频率响应特性和噪音条件对语音和ASB质量的影响--实验评估
链接:https://arxiv.org/abs/2510.09236

作者:Michele Buccoli, Yu Du, Jacob Soendergaard, Simone Shawn Cazzaniga
摘要:在选择用于汽车免提通信或自动语音识别(ASR)应用的麦克风时,OEM通常根据已建立的标准建议(例如,ITU-P.1110、ITU-P.1120)。在实践中,当考虑到车厢内麦克风放置的限制和约束以及汽车级环境鲁棒性要求时,实现汽车麦克风的优选带宽通常具有挑战性。另一方面,关于每个麦克风特性对实际性能的影响,似乎没有共识或足够的数据。为了回答这个问题,我们使用在真实车辆和各种驾驶条件下记录的噪声信号来实验研究麦克风特性与语音通信的最终音频质量和ASR引擎性能之间的关系。我们专注于麦克风带宽和幅度频率响应形状的变化如何影响感知语音质量。通过使用ETSI TS 103 281度量(S-MOS、N-MOS、G-MOS)和诸如SNR的辅助度量来比较语音质量结果。ASR结果使用标准指标进行评估,例如单词错误率(WER)。这项研究的结果提供了了解哪些麦克风频率响应特性与音频质量和选择适当的麦克风规格更相关的知识,特别是对于汽车应用。
摘要:Upon choosing microphones for automotive hands-free communication or Automatic Speech Recognition (ASR) applications, OEMs typically specify wideband, super wideband or even fullband requirements following established standard recommendations (e.g., ITU-P.1110, ITU-P.1120). In practice, it is often challenging to achieve the preferred bandwidth for an automotive microphone when considering limitations and constraints on microphone placement inside the cabin, and the automotive grade environmental robustness requirements. On the other hand, there seems to be no consensus or sufficient data on the effect of each microphone characteristic on the actual performance. As an attempt to answer this question, we used noise signals recorded in real vehicles and under various driving conditions to experimentally study the relationship between the microphones' characteristics and the final audio quality of speech communication and performance of ASR engines. We focus on how variations in microphone bandwidth and amplitude frequency response shapes affect the perceptual speech quality. The speech quality results are compared by using ETSI TS 103 281 metrics (S-MOS, N-MOS, G-MOS) and ancillary metrics such as SNR. The ASR results are evaluated with standard metrics such as Word Error Rate (WER). Findings from this study provide knowledge in the understanding of what microphone frequency response characteristics are more relevant for audio quality and choice of proper microphone specifications, particularly for automotive applications.


【18】Unsupervised lexicon learning from speech is limited by representations rather than clustering
标题:从语音中进行的无监督词典学习受到表示而不是集群的限制
链接:https://arxiv.org/abs/2510.09225

作者:Danel Adendorff, Simon Malan, Herman Kamper
备注:Submitted to ICASSP 2026
摘要:零资源分词和聚类系统旨在将语音标记成类似单词的单元,而无需访问文本标签。尽管取得了进步,但诱导词汇仍然远远不够完美。在具有黄金单词边界的理想设置中,我们询问性能是否受到单词段表示的限制,或者受到将它们分组为类似单词类型的聚类方法的限制。我们将一系列自监督语音特征(连续/离散、帧/词级)与英语和普通话数据的不同聚类方法(K均值、分层、基于图)相结合。最好的系统使用图聚类,并对连续特征进行动态时间扭曲。更快的替代方案使用图聚类与余弦距离的平均连续功能或编辑距离的离散单元序列。通过控制实验,隔离的表示或聚类方法,我们证明了相同的词类型,而不是集群的段的表示变异性是限制性能的主要因素。
摘要:Zero-resource word segmentation and clustering systems aim to tokenise speech into word-like units without access to text labels. Despite progress, the induced lexicons are still far from perfect. In an idealised setting with gold word boundaries, we ask whether performance is limited by the representation of word segments, or by the clustering methods that group them into word-like types. We combine a range of self-supervised speech features (continuous/discrete, frame/word-level) with different clustering methods (K-means, hierarchical, graph-based) on English and Mandarin data. The best system uses graph clustering with dynamic time warping on continuous features. Faster alternatives use graph clustering with cosine distance on averaged continuous features or edit distance on discrete unit sequences. Through controlled experiments that isolate either the representations or the clustering method, we demonstrate that representation variability across segments of the same word type -- rather than clustering -- is the primary factor limiting performance.


【19】Look before Transcription: End-to-End SlideASR with Visually-Anchored Policy Optimization
标题:转录前先看:具有视觉锚定策略优化的端到端SlideASB
链接:https://arxiv.org/abs/2510.08618

作者:Rui Hu, Delai Qiu, Yining Wang, Shengping Liu, Jitao Sang
摘要:自动语音识别(ASR)系统经常与特定领域的术语斗争,特别是在学术讲座等专业环境中。为了解决这个问题,我们定义了SlideASR任务,它利用演示幻灯片中丰富的视觉信息来提高转录的准确性。用于此任务的现有流水线方法往往是复杂的并且表现不佳。虽然全模态大型语言模型(OLLM)提供了一个有前途的端到端框架,但它们在实践中经常会退化为简单的光学字符识别(OCR)系统。为了克服这一点,我们提出了视觉锚定策略优化(VAPO),这是一种新的后训练方法,旨在控制模型的推理过程。借鉴思维链推理范式,VAPO使用格式强制执行结构化的“转录前查看”程序。具体而言,模型首先在思考步骤中对幻灯片内容执行OCR,然后在回答步骤中通过引用该识别的视觉信息来生成转录。这个推理过程通过强化学习进行了优化,有四个不同的奖励,分别针对格式合规性、OCR准确性、ASR质量和视觉锚定一致性。为了支持进一步的研究,我们构建了SlideASR-Bench,这是一个新的实体丰富的基准,由用于训练和测试的合成数据集和具有挑战性的真实世界评估集组成。大量的实验表明,VAPO显著提高了特定领域术语的识别,为SlideASR建立了一个有效的端到端范例。
摘要:Automatic speech recognition (ASR) systems often struggle with domain-specific terminology, especially in specialized settings such as academic lectures. To address this, we define the SlideASR task, which leverages the rich visual information from presentation slides to improve transcription accuracy. Existing pipeline methods for this task tend to be complex and underperform. Although omni-modal large language models (OLLMs) provide a promising end-to-end framework, they frequently fail in practice by degenerating into simple optical character recognition (OCR) systems. To overcome this, we propose Visually-Anchored Policy Optimization (VAPO), a novel post-training method designed to control the model's reasoning process. Drawing on the Chain-of-Thought reasoning paradigm, VAPO enforces a structured "Look before Transcription" procedure using a  format. Specifically, the model first performs OCR on the slide content within the think step, then generates the transcription by referencing this recognized visual information in the answer step. This reasoning process is optimized via reinforcement learning with four distinct rewards targeting format compliance, OCR accuracy, ASR quality, and visual anchoring consistency. To support further research, we construct SlideASR-Bench, a new entity-rich benchmark consisting of a synthetic dataset for training and testing, and a challenging real-world set for evaluation. Extensive experiments demonstrate that VAPO significantly improves recognition of domain-specific terms, establishing an effective end-to-end paradigm for SlideASR.


【20】BaldWhisper: Faster Whisper with Head Shearing and Layer Merging
标题:BaldWhisper:具有头部剪切和分层合并的更快Whisper
链接:https://arxiv.org/abs/2510.08599

作者:Yaya Sy, Christophe Cerisara, Irina Illina
摘要:为低资源语言修剪大型预训练的Transformers是一项挑战,因为它通常需要大量的重新训练数据来恢复性能。例如,Distill-Whisper对Whisper进行了40%的删减,并对21,000小时的语音进行了重新训练,远远超过了大多数语言的可用时间。Whisper可以在数据稀缺的环境中为边缘设备提供更轻、更快的速度吗?针对只有32 h语音到文本数据的Bambara,我们提出了一种新的剪枝方法。而不是词汇修剪,这是不合适的,由于频繁的代码切换的班巴拉扬声器,我们压缩嵌入低秩分解和特征蒸馏。而不是删除层,我们合并它们以限制性能损失。最终的模型保留了90%的原始性能,同时在MacBook Air M1上缩小了48%,速度提高了2.15倍。
摘要:Pruning large pre-trained transformers for low-resource languages is challenging, as it often requires massive retraining data to recover performance. For instance, Distill-Whisper prunes Whisper by 40% and retrains on 21,000 hours of speech, far beyond what is available for most languages. Can Whisper be made lighter and faster for edge devices in data-scarce settings? Focusing on Bambara with only 32h of speech-to-text data, we propose a new pruning recipe. Instead of vocabulary pruning, which is unsuitable due to frequent code-switching by Bambara speakers, we compress the embeddings with low-rank decomposition and feature distillation. Rather than removing layers, we merge them to limit performance loss. The final model preserves 90% of the original performance while being 48% smaller and 2.15x faster on a MacBook Air M1.


【21】Dynamic Stress Detection: A Study of Temporal Progression Modelling of Stress in Speech
标题:动态压力检测:言语压力的时间进程模型研究
链接:https://arxiv.org/abs/2510.08586

作者:Vishakha Lall, Yisi Liu
备注:Accepted at IEEE CogMI 2025
摘要:在高压环境下,从言语中检测心理压力至关重要。虽然先前的工作已经利用声学特征进行压力检测,但大多数将压力视为静态标签。在这项工作中,我们的模型压力作为一个时间上不断变化的现象,历史情绪状态的影响。我们提出了一种动态标签策略,从情感标签中获得细粒度的压力注释,并引入基于交叉注意的顺序模型,单向LSTM和Transformer Encoder,以捕获时间压力进展。我们的方法在MuSE(+5%)和StressID(+18%)上比现有基线实现了显着的准确性增益,并很好地推广到自定义的真实世界数据集。这些结果突出强调了建模压力作为一个动态结构在讲话中的价值。
摘要:Detecting psychological stress from speech is critical in high-pressure settings. While prior work has leveraged acoustic features for stress detection, most treat stress as a static label. In this work, we model stress as a temporally evolving phenomenon influenced by historical emotional state. We propose a dynamic labelling strategy that derives fine-grained stress annotations from emotional labels and introduce cross-attention-based sequential models, a Unidirectional LSTM and a Transformer Encoder, to capture temporal stress progression. Our approach achieves notable accuracy gains on MuSE (+5%) and StressID (+18%) over existing baselines, and generalises well to a custom real-world dataset. These results highlight the value of modelling stress as a dynamic construct in speech.


【22】Articulation-Informed ASR: Integrating Articulatory Features into ASR via Auxiliary Speech Inversion and Cross-Attention Fusion
标题:关节语知情的ASB:通过辅助言语倒置和交叉注意融合将关节语特征集成到ASB中
链接:https://arxiv.org/abs/2510.08585

作者:Ahmed Adel Attia, Jing Liu, Carol Espy Wilson
摘要:先前的工作已经研究了使用发音特征作为自动语音识别(ASR)的补充表示,但它们的使用主要限于浅层声学模型。在这项工作中,我们重新审视了深度学习时代的发音信息,并提出了一个框架,该框架利用发音表示作为辅助任务和识别模型的伪输入。具体来说,我们采用语音反转作为辅助预测任务,预测的发音特征作为查询流注入到模型中的交叉注意模块中,声学嵌入作为键和值。LibriSpeech上的实验表明,我们的方法在基于transformer的强基线上得到了一致的改进,特别是在低资源条件下。这些发现表明,发音功能,一旦靠边站的ASR研究,可以提供有意义的好处时,重新引入现代建筑。
摘要:Prior works have investigated the use of articulatory features as complementary representations for automatic speech recognition (ASR), but their use was largely confined to shallow acoustic models. In this work, we revisit articulatory information in the era of deep learning and propose a framework that leverages articulatory representations both as an auxiliary task and as a pseudo-input to the recognition model. Specifically, we employ speech inversion as an auxiliary prediction task, and the predicted articulatory features are injected into the model as a query stream in a cross-attention module with acoustic embeddings as keys and values. Experiments on LibriSpeech demonstrate that our approach yields consistent improvements over strong transformer-based baselines, particularly under low-resource conditions. These findings suggest that articulatory features, once sidelined in ASR research, can provide meaningful benefits when reintroduced with modern architectures.


eess.AS音频处理


【1】Spatially-Augmented Sequence-to-Sequence Neural Diarization for Meetings
标题:会议的空间增强序列到序列神经扩张
链接:https://arxiv.org/abs/2510.09505

作者:Li Li, Ming Cheng, Hongyu Zhang, Juan Liu, Ming Li
备注:This paper has submitted to ICASSP 2026
摘要:本文提出了一种空间增强的序列到序列神经日志(SA-S2 SND)框架,该框架将SRP-DNN估计的到达方向(DOA)线索集成到S2 SND骨干中。该方法采用两阶段训练策略:首先用单通道音频和DOA特征对模型进行训练,然后在DOA引导下用多通道输入对模型进行优化。此外,模拟DOA产生方案,以减轻依赖于匹配的多通道语料库。在AliMeeting数据集上,SA-S2 SND的表现始终优于S2 SND基线,在离线模式下实现了7.4%的相对DER降低,在与渠道注意力相结合时提高了19%以上。这些结果表明,空间线索是高度互补的跨通道建模,产生良好的性能在在线和离线设置。
摘要:This paper proposes a Spatially-Augmented Sequence-to-Sequence Neural Diarization (SA-S2SND) framework, which integrates direction-of-arrival (DOA) cues estimated by SRP-DNN into the S2SND backbone. A two-stage training strategy is adopted: the model is first trained with single-channel audio and DOA features, and then further optimized with multi-channel inputs under DOA guidance. In addition, a simulated DOA generation scheme is introduced to alleviate dependence on matched multi-channel corpora. On the AliMeeting dataset, SA-S2SND consistently outperform the S2SND baseline, achieving a 7.4% relative DER reduction in the offline mode and over 19% improvement when combined with channel attention. These results demonstrate that spatial cues are highly complementary to cross-channel modeling, yielding good performance in both online and offline settings.


【2】A Study of the Removability of Speaker-Adversarial Perturbations
标题:说话者对抗性扰动的可去除性研究
链接:https://arxiv.org/abs/2510.09504

作者:Liping Chen, Chenyang Guo, Kong Aik Lee, Zhen-Hua Ling, Wu Guo
摘要:对抗性攻击的最新进展已经证明了它们在误导说话者识别模型,对说话者身份做出错误预测方面的有效性。另一方面,针对说话人对抗攻击的防御技术侧重于减少说话人对抗扰动对说话人属性提取的影响。这些技术并不寻求完全去除扰动并恢复原始语音。为此,本文研究了说话人对抗扰动的可移除性。具体而言,调查进行假设不同程度的认识扰动发生器在三种情况下:无知,半知情,和知情。此外,我们考虑了基于优化和前馈扰动生成方法。在LibriSpeech数据集上进行的实验表明:1)在无知场景中,说话者-对抗者扰动不能被消除,尽管它们对说话者属性提取的影响被减小,2)在半知情场景中,说话者-对抗者扰动不能被完全去除,而由前馈模型生成的那些扰动可以被显著减小,以及3)在知情场景中,几乎消除了说话者对抗干扰,允许恢复原始语音。音频样本可以在https://voiceprivacy.github.io/Perturbation-Generation-Removal/上找到。
摘要:Recent advancements in adversarial attacks have demonstrated their effectiveness in misleading speaker recognition models, making wrong predictions about speaker identities. On the other hand, defense techniques against speaker-adversarial attacks focus on reducing the effects of speaker-adversarial perturbations on speaker attribute extraction. These techniques do not seek to fully remove the perturbations and restore the original speech. To this end, this paper studies the removability of speaker-adversarial perturbations. Specifically, the investigation is conducted assuming various degrees of awareness of the perturbation generator across three scenarios: ignorant, semi-informed, and well-informed. Besides, we consider both the optimization-based and feedforward perturbation generation methods. Experiments conducted on the LibriSpeech dataset demonstrated that: 1) in the ignorant scenario, speaker-adversarial perturbations cannot be eliminated, although their impact on speaker attribute extraction is reduced, 2) in the semi-informed scenario, the speaker-adversarial perturbations cannot be fully removed, while those generated by the feedforward model can be considerably reduced, and 3) in the well-informed scenario, speaker-adversarial perturbations are nearly eliminated, allowing for the restoration of the original speech. Audio samples can be found in https://voiceprivacy.github.io/Perturbation-Generation-Removal/.


【3】Target speaker anonymization in multi-speaker recordings
标题:多说话人录音中的目标说话人匿名化
链接:https://arxiv.org/abs/2510.09307

作者:Natalia Tomashenko, Junichi Yamagishi, Xin Wang, Yun Liu, Emmanuel Vincent
备注:Submitted to ICASSP 2026
摘要:现有的大多数说话人匿名化研究都集中在单说话人音频上,导致了针对这种情况优化的技术和评估指标的发展。本研究解决了多说话人会话音频中说话人匿名化的重大挑战,特别是当只有一个目标说话人需要匿名化时。这种场景在呼叫中心等环境中高度相关,在这些环境中,客户隐私只需要在与运营商的交互中匿名客户的声音。传统的匿名化方法通常不适合这项任务。此外,目前的评估方法不允许我们准确地评估在这个复杂的多扬声器场景中的隐私保护和实用性。这项工作旨在弥合这些差距,探讨有效的策略,有针对性的说话人匿名在会话音频,突出其发展中的潜在问题,并提出相应的改进评估方法。
摘要:Most of the existing speaker anonymization research has focused on single-speaker audio, leading to the development of techniques and evaluation metrics optimized for such condition. This study addresses the significant challenge of speaker anonymization within multi-speaker conversational audio, specifically when only a single target speaker needs to be anonymized. This scenario is highly relevant in contexts like call centers, where customer privacy necessitates anonymizing only the customer's voice in interactions with operators. Conventional anonymization methods are often not suitable for this task. Moreover, current evaluation methodology does not allow us to accurately assess privacy protection and utility in this complex multi-speaker scenario. This work aims to bridge these gaps by exploring effective strategies for targeted speaker anonymization in conversational audio, highlighting potential problems in their development and proposing corresponding improved evaluation methodologies.


【4】Effects of automotive microphone frequency response characteristics and noise conditions on speech and ASR quality -- an experimental evaluation
标题:汽车麦克风频率响应特性和噪音条件对语音和ASB质量的影响--实验评估
链接:https://arxiv.org/abs/2510.09236

作者:Michele Buccoli, Yu Du, Jacob Soendergaard, Simone Shawn Cazzaniga
摘要:在选择用于汽车免提通信或自动语音识别(ASR)应用的麦克风时,OEM通常根据已建立的标准建议(例如,ITU-P.1110、ITU-P.1120)。在实践中,当考虑到车厢内麦克风放置的限制和约束以及汽车级环境鲁棒性要求时,实现汽车麦克风的优选带宽通常具有挑战性。另一方面,关于每个麦克风特性对实际性能的影响,似乎没有共识或足够的数据。为了回答这个问题,我们使用在真实车辆和各种驾驶条件下记录的噪声信号来实验研究麦克风特性与语音通信的最终音频质量和ASR引擎性能之间的关系。我们专注于麦克风带宽和幅度频率响应形状的变化如何影响感知语音质量。通过使用ETSI TS 103 281度量(S-MOS、N-MOS、G-MOS)和诸如SNR的辅助度量来比较语音质量结果。ASR结果使用标准指标进行评估,例如单词错误率(WER)。这项研究的结果提供了了解哪些麦克风频率响应特性与音频质量和选择适当的麦克风规格更相关的知识,特别是对于汽车应用。
摘要:Upon choosing microphones for automotive hands-free communication or Automatic Speech Recognition (ASR) applications, OEMs typically specify wideband, super wideband or even fullband requirements following established standard recommendations (e.g., ITU-P.1110, ITU-P.1120). In practice, it is often challenging to achieve the preferred bandwidth for an automotive microphone when considering limitations and constraints on microphone placement inside the cabin, and the automotive grade environmental robustness requirements. On the other hand, there seems to be no consensus or sufficient data on the effect of each microphone characteristic on the actual performance. As an attempt to answer this question, we used noise signals recorded in real vehicles and under various driving conditions to experimentally study the relationship between the microphones' characteristics and the final audio quality of speech communication and performance of ASR engines. We focus on how variations in microphone bandwidth and amplitude frequency response shapes affect the perceptual speech quality. The speech quality results are compared by using ETSI TS 103 281 metrics (S-MOS, N-MOS, G-MOS) and ancillary metrics such as SNR. The ASR results are evaluated with standard metrics such as Word Error Rate (WER). Findings from this study provide knowledge in the understanding of what microphone frequency response characteristics are more relevant for audio quality and choice of proper microphone specifications, particularly for automotive applications.


【5】Unsupervised lexicon learning from speech is limited by representations rather than clustering
标题:从语音中进行的无监督词典学习受到表示而不是集群的限制
链接:https://arxiv.org/abs/2510.09225

作者:Danel Adendorff, Simon Malan, Herman Kamper
备注:Submitted to ICASSP 2026
摘要:零资源分词和聚类系统旨在将语音标记成类似单词的单元,而无需访问文本标签。尽管取得了进步,但诱导词汇仍然远远不够完美。在一个理想化的设置与黄金字的边界,我们问性能是否是有限的词段的表示,或聚类方法,将它们分组到词样类型。我们将一系列自监督语音特征(连续/离散、帧/词级)与英语和普通话数据的不同聚类方法(K均值、分层、基于图)相结合。最好的系统使用图聚类,并对连续特征进行动态时间扭曲。更快的替代方案使用图聚类与余弦距离的平均连续功能或编辑距离的离散单元序列。通过控制实验,隔离的表示或聚类方法,我们证明了相同的词类型,而不是集群的段的表示变异性是限制性能的主要因素。
摘要:Zero-resource word segmentation and clustering systems aim to tokenise speech into word-like units without access to text labels. Despite progress, the induced lexicons are still far from perfect. In an idealised setting with gold word boundaries, we ask whether performance is limited by the representation of word segments, or by the clustering methods that group them into word-like types. We combine a range of self-supervised speech features (continuous/discrete, frame/word-level) with different clustering methods (K-means, hierarchical, graph-based) on English and Mandarin data. The best system uses graph clustering with dynamic time warping on continuous features. Faster alternatives use graph clustering with cosine distance on averaged continuous features or edit distance on discrete unit sequences. Through controlled experiments that isolate either the representations or the clustering method, we demonstrate that representation variability across segments of the same word type -- rather than clustering -- is the primary factor limiting performance.


【6】Impact of HRTF individualisation and head movements in a real/virtual localisation task
标题:真实/虚拟本地化任务中HRTF个性化和头部移动的影响
链接:https://arxiv.org/abs/2510.09161

作者:Vincent Martin, Lorenzo Picinali
摘要:音频增强现实(AAR)应用的目标是将虚拟声源无缝集成到真实环境中。对于这些应用来说,虚拟源精确定位在预期位置以及声学环境精确匹配至关重要。   一种用于在耳机上空间化声音的有效方法是通过头部相关传递函数(HRTF)。这些解释了在声波到达耳膜之前,听者的身体特征是如何改变声波的。本研究探讨使用个性化的HRTF的本地化和感知现实主义的虚拟声源与一个真正的视觉对象的影响。   参与者的任务是分别通过耳机和球形扬声器阵列定位虚拟和真实的语音源。评估的重点是感知的现实性和来源的位置。所有的来源都与30个真正的视觉源(扬声器)安排在半消声室之一。   各种声源渲染进行了比较,包括单扬声器渲染和双耳渲染与个性化或非个性化的HRTF。此外,研究人员还探讨了头部运动的影响:10名参与者在头部运动和不运动的情况下完成了相同的任务。   结果表明,使用个人HRTF提高了感知的现实主义,但在静态场景中的本地化性能。令人惊讶的是,当头部运动是可能的和鼓励时,观察到相反的情况。
摘要:The objective of Audio Augmented Reality (AAR) applications are to seamlessly integrate virtual sound sources within a real environment. It is critical for these applications that virtual sources are localised precisely at the intended position, and that the acoustic environments are accurately matched.   One effective method for spatialising sound on headphones is through Head-Related Transfer Functions (HRTFs). These characterise how the physical features of a listener modify sound waves before they reach the eardrum. This study examines the influence of using individualised HRTFs on the localisation and the perceived realism of virtual sound sources associated with a real visual object.   Participants were tasked with localising virtual and real speech sources presented via headphones and through a spherical loudspeaker array, respectively. The assessment focussed on perceived realism and sources location. All sources were associated with one of thirty real visual sources (loudspeakers) arranged in a semi-anechoic room.   Various sound source renderings were compared, including single loudspeaker rendering and binaural rendering with individualised or non-individualised HRTFs. Additionally, the impact of head movements was explored: ten participants completed the same task with and without the possibility to move their head.   The results showed that using individual HRTFs improved perceived realism but not localisation performance in the static scenario. Surprisingly, the opposite was observed when head movements were possible and encouraged.


【7】Look before Transcription: End-to-End SlideASR with Visually-Anchored Policy Optimization
标题:转录前先看:具有视觉锚定策略优化的端到端SlideASB
链接:https://arxiv.org/abs/2510.08618

作者:Rui Hu, Delai Qiu, Yining Wang, Shengping Liu, Jitao Sang
摘要:自动语音识别(ASR)系统经常与特定领域的术语斗争,特别是在学术讲座等专业环境中。为了解决这个问题,我们定义了SlideASR任务,它利用演示幻灯片中丰富的视觉信息来提高转录的准确性。用于此任务的现有流水线方法往往是复杂的并且表现不佳。虽然全模态大型语言模型(OLLM)提供了一个有前途的端到端框架,但它们在实践中经常会退化为简单的光学字符识别(OCR)系统。为了克服这一点,我们提出了视觉锚定策略优化(VAPO),这是一种新的后训练方法,旨在控制模型的推理过程。借鉴思维链推理范式,VAPO使用格式强制执行结构化的“转录前查看”程序。具体而言,模型首先在思考步骤中对幻灯片内容执行OCR,然后在回答步骤中通过引用该识别的视觉信息来生成转录。这个推理过程通过强化学习进行了优化,有四个不同的奖励,分别针对格式合规性、OCR准确性、ASR质量和视觉锚定一致性。为了支持进一步的研究,我们构建了SlideASR-Bench,这是一个新的实体丰富的基准,由用于训练和测试的合成数据集和具有挑战性的真实世界评估集组成。大量的实验表明,VAPO显著提高了特定领域术语的识别,为SlideASR建立了一个有效的端到端范例。
摘要:Automatic speech recognition (ASR) systems often struggle with domain-specific terminology, especially in specialized settings such as academic lectures. To address this, we define the SlideASR task, which leverages the rich visual information from presentation slides to improve transcription accuracy. Existing pipeline methods for this task tend to be complex and underperform. Although omni-modal large language models (OLLMs) provide a promising end-to-end framework, they frequently fail in practice by degenerating into simple optical character recognition (OCR) systems. To overcome this, we propose Visually-Anchored Policy Optimization (VAPO), a novel post-training method designed to control the model's reasoning process. Drawing on the Chain-of-Thought reasoning paradigm, VAPO enforces a structured "Look before Transcription" procedure using a  format. Specifically, the model first performs OCR on the slide content within the think step, then generates the transcription by referencing this recognized visual information in the answer step. This reasoning process is optimized via reinforcement learning with four distinct rewards targeting format compliance, OCR accuracy, ASR quality, and visual anchoring consistency. To support further research, we construct SlideASR-Bench, a new entity-rich benchmark consisting of a synthetic dataset for training and testing, and a challenging real-world set for evaluation. Extensive experiments demonstrate that VAPO significantly improves recognition of domain-specific terms, establishing an effective end-to-end paradigm for SlideASR.


【8】BaldWhisper: Faster Whisper with Head Shearing and Layer Merging
标题:BaldWhisper:具有头部剪切和分层合并的更快Whisper
链接:https://arxiv.org/abs/2510.08599

作者:Yaya Sy, Christophe Cerisara, Irina Illina
摘要:为低资源语言修剪大型预训练的Transformers是一项挑战,因为它通常需要大量的重新训练数据来恢复性能。例如,Distill-Whisper对Whisper进行了40%的删减,并对21,000小时的语音进行了重新训练,远远超过了大多数语言的可用时间。Whisper可以在数据稀缺的环境中为边缘设备提供更轻、更快的速度吗?针对只有32 h语音到文本数据的Bambara,我们提出了一种新的剪枝方法。我们没有进行词汇修剪(由于班巴拉语使用者频繁的代码切换,这是不合适的),而是通过低秩分解和特征蒸馏来压缩嵌入。而不是删除层,我们合并它们以限制性能损失。最终的模型保留了90%的原始性能,同时在MacBook Air M1上缩小了48%,速度提高了2.15倍。
摘要:Pruning large pre-trained transformers for low-resource languages is challenging, as it often requires massive retraining data to recover performance. For instance, Distill-Whisper prunes Whisper by 40% and retrains on 21,000 hours of speech, far beyond what is available for most languages. Can Whisper be made lighter and faster for edge devices in data-scarce settings? Focusing on Bambara with only 32h of speech-to-text data, we propose a new pruning recipe. Instead of vocabulary pruning, which is unsuitable due to frequent code-switching by Bambara speakers, we compress the embeddings with low-rank decomposition and feature distillation. Rather than removing layers, we merge them to limit performance loss. The final model preserves 90% of the original performance while being 48% smaller and 2.15x faster on a MacBook Air M1.


【9】Dynamic Stress Detection: A Study of Temporal Progression Modelling of Stress in Speech
标题:动态压力检测:言语压力的时间进程模型研究
链接:https://arxiv.org/abs/2510.08586

作者:Vishakha Lall, Yisi Liu
备注:Accepted at IEEE CogMI 2025
摘要:在高压环境下,从言语中检测心理压力至关重要。虽然先前的工作已经利用声学特征进行压力检测,但大多数将压力视为静态标签。在这项工作中,我们的模型压力作为一个时间上不断变化的现象,历史情绪状态的影响。我们提出了一种动态标签策略,从情感标签中获得细粒度的压力注释,并引入基于交叉注意的顺序模型,单向LSTM和Transformer Encoder,以捕获时间压力进展。我们的方法在MuSE(+5%)和StressID(+18%)上比现有基线实现了显着的准确性增益,并很好地推广到自定义的真实世界数据集。这些结果突出强调了建模压力作为一个动态结构在讲话中的价值。
摘要:Detecting psychological stress from speech is critical in high-pressure settings. While prior work has leveraged acoustic features for stress detection, most treat stress as a static label. In this work, we model stress as a temporally evolving phenomenon influenced by historical emotional state. We propose a dynamic labelling strategy that derives fine-grained stress annotations from emotional labels and introduce cross-attention-based sequential models, a Unidirectional LSTM and a Transformer Encoder, to capture temporal stress progression. Our approach achieves notable accuracy gains on MuSE (+5%) and StressID (+18%) over existing baselines, and generalises well to a custom real-world dataset. These results highlight the value of modelling stress as a dynamic construct in speech.


【10】Articulation-Informed ASR: Integrating Articulatory Features into ASR via Auxiliary Speech Inversion and Cross-Attention Fusion
标题:关节语知情的ASB:通过辅助言语倒置和交叉注意融合将关节语特征集成到ASB中
链接:https://arxiv.org/abs/2510.08585

作者:Ahmed Adel Attia, Jing Liu, Carol Espy Wilson
摘要:先前的工作已经研究了使用发音特征作为自动语音识别(ASR)的补充表示,但它们的使用主要限于浅层声学模型。在这项工作中,我们重新审视了深度学习时代的发音信息,并提出了一个框架,该框架利用发音表示作为辅助任务和识别模型的伪输入。具体来说,我们采用语音反转作为辅助预测任务,预测的发音特征作为查询流注入到模型中的交叉注意模块中,声学嵌入作为键和值。LibriSpeech上的实验表明,我们的方法在基于transformer的强基线上得到了一致的改进,特别是在低资源条件下。这些发现表明,发音功能,一旦靠边站的ASR研究,可以提供有意义的好处时,重新引入现代建筑。
摘要:Prior works have investigated the use of articulatory features as complementary representations for automatic speech recognition (ASR), but their use was largely confined to shallow acoustic models. In this work, we revisit articulatory information in the era of deep learning and propose a framework that leverages articulatory representations both as an auxiliary task and as a pseudo-input to the recognition model. Specifically, we employ speech inversion as an auxiliary prediction task, and the predicted articulatory features are injected into the model as a query stream in a cross-attention module with acoustic embeddings as keys and values. Experiments on LibriSpeech demonstrate that our approach yields consistent improvements over strong transformer-based baselines, particularly under low-resource conditions. These findings suggest that articulatory features, once sidelined in ASR research, can provide meaningful benefits when reintroduced with modern architectures.


【11】Accent-Invariant Automatic Speech Recognition via Saliency-Driven Spectrogram Masking
标题:通过显著性驱动频谱图掩蔽的口音不变自动语音识别
链接:https://arxiv.org/abs/2510.09528

作者:Mohammad Hossein Sameti, Sepehr Harfi Moridani, Ali Zarean, Hossein Sameti
备注:Submitted to ICASSP 2026
摘要:预先训练的基于transformer的模型显著提高了自动语音识别(ASR),但它们对口音和方言变化仍然敏感,导致英语和波斯语等语言多样性语言的单词错误率(WER)升高。为了解决这一挑战,我们提出了一个口音不变的ASR框架,将口音和方言分类集成到识别管道中。我们的方法包括训练一个基于谱图的分类器来捕获特定口音的线索,掩盖对其预测最有影响的区域,并使用掩蔽的谱图进行数据增强。这增强了ASR模型对口音变化的鲁棒性。我们使用英语和波斯语的讲话的方法进行评估。对于波斯语,我们引入了一个新收集的跨越多个区域口音的数据集,建立了波斯语ASR中口音变化的第一个系统基准,填补了多语言语音研究的关键空白,并为未来低资源,语言多样性语言的研究提供了基础。Whisper模型的实验结果表明,我们的掩蔽和增强策略在英语和波斯语设置中产生了大量的WER减少,证实了该方法的有效性。这项研究促进了多语言ASR系统的发展,这些系统对口音和方言的多样性具有弹性。代码和数据集可在https://github.com/MH-Sameti/Accent_invariant_ASR上公开获取
摘要:Pre-trained transformer-based models have significantly advanced automatic speech recognition (ASR), yet they remain sensitive to accent and dialectal variations, resulting in elevated word error rates (WER) in linguistically diverse languages such as English and Persian. To address this challenge, we propose an accent-invariant ASR framework that integrates accent and dialect classification into the recognition pipeline. Our approach involves training a spectrogram-based classifier to capture accent-specific cues, masking the regions most influential to its predictions, and using the masked spectrograms for data augmentation. This enhances the robustness of ASR models against accent variability. We evaluate the method using both English and Persian speech. For Persian, we introduce a newly collected dataset spanning multiple regional accents, establishing the first systematic benchmark for accent variation in Persian ASR that fills a critical gap in multilingual speech research and provides a foundation for future studies on low-resource, linguistically diverse languages. Experimental results with the Whisper model demonstrate that our masking and augmentation strategy yields substantial WER reductions in both English and Persian settings, confirming the effectiveness of the approach. This research advances the development of multilingual ASR systems that are resilient to accent and dialect diversity. Code and dataset are publicly available at: https://github.com/MH-Sameti/Accent_invariant_ASR


【12】The Speech-LLM Takes It All: A Truly Fully End-to-End Spoken Dialogue State Tracking Approach
标题:演讲LLM囊括一切:真正完全端到端的口语对话状态跟踪方法
链接:https://arxiv.org/abs/2510.09424

作者:Nizar El Ghazal, Antoine Caubrière, Valentin Vielzeuf
摘要:本文提出了一种比较研究的上下文管理策略,端到端的口语对话状态跟踪使用语音LLM。我们系统地评估了传统的多模态上下文(结合文本历史和口语当前的转折),完整的口语历史,压缩口语历史的方法。我们在SpokenWOZ语料库上的实验表明,提供完整的口语对话作为输入,在类似大小的模型中产生最高的性能,显着超过先前的方法。此外,我们表明,注意力池为基础的压缩的口语历史提供了一个很强的权衡,保持竞争力的准确性,减少上下文大小。详细的分析证实,改进源于更有效的上下文利用。
摘要:This paper presents a comparative study of context management strategies for end-to-end Spoken Dialog State Tracking using Speech-LLMs. We systematically evaluate traditional multimodal context (combining text history and spoken current turn), full spoken history, and compressed spoken history approaches. Our experiments on the SpokenWOZ corpus demonstrate that providing the full spoken conversation as input yields the highest performance among models of similar size, significantly surpassing prior methods. Furthermore, we show that attention-pooling-based compression of the spoken history offers a strong trade-off, maintaining competitive accuracy with reduced context size. Detailed analysis confirms that improvements stem from more effective context utilization.


【13】FLToP CTC: Frame-Level Token Pruning via Relative Threshold for Efficient and Memory-Saving Decoding on Diverse Platforms
标题:FLToP CTC:通过相对阈值进行帧级令牌修剪,以在不同平台上进行高效且节省内存的解码
链接:https://arxiv.org/abs/2510.09085

作者:Atul Shree, Harshith Jupuru
备注:5 pages, 5 figures
摘要:基于CTC的ASR系统在资源有限的环境中面临计算和内存瓶颈。传统的CTC解码器,在系统中需要高达90%的处理时间(例如,wav 2 vec 2-large在L4 GPU上),由于详尽的令牌级操作而面临效率低下的问题。本文介绍了一种新的解码算法,采用帧级令牌修剪的相对阈值概率的指导下,连接时间分类(FLToP CTC)的帧级令牌修剪。通过动态消除每帧的低概率令牌,FLToP CTC减少了计算和内存需求,同时保持可忽略的WER降级。在LibriSpeech上,FLToP CTC与标准CTC解码器相比,实现了10.5倍的运行时加速和2.78倍的内存减少。它的简单性使其能够跨平台(CPU,GPU等)无缝集成到CTC解码器中。FLToP CTC解决了CTC瓶颈,为资源有限的环境和实时应用提供了可扩展性,增强了语音识别的可访问性和效率。
摘要:CTC-based ASR systems face computational and memory bottlenecks in resource-limited environments. Traditional CTC decoders, requiring up to 90% of processing time in systems (e.g., wav2vec2-large on L4 GPUs), face inefficiencies due to exhaustive token-level operations. This paper introduces Frame Level Token Pruning for Connectionist Temporal Classification (FLToP CTC), a novel decoding algorithm that employs frame-level token pruning guided by a relative threshold probability. By dynamically eliminating low-probability tokens per frame, FLToP CTC reduces compute and memory demands while maintaining negligible WER degradation. On LibriSpeech, FLToP CTC achieves a 10.5x runtime speedup and 2.78x memory reduction versus standard CTC decoders. Its simplicity enables seamless integration into CTC decoders across platforms (CPUs, GPUs, etc.). FLToP CTC addresses CTC bottlenecks, offering scalability for resource-limited environments and realtime applications, enhancing speech recognition accessibility and efficiency.


【14】MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation
标题:MMAudioSep:驯服视频到音频生成模型以实现视频/文本查询声音分离
链接:https://arxiv.org/abs/2510.09065

作者:Akira Takahashi, Shusuke Takahashi, Yuki Mitsufuji
备注:4 pages, 4 figures, 2 tables
摘要:我们介绍了MMAudioSep,一个生成模型的视频/文本查询的声音分离,是建立在一个预训练的视频到音频模型。通过利用通过预训练的音频生成模型学习的关于视频/文本和音频之间的关系的知识,我们可以更有效地训练模型,即,该模型不需要从头开始训练。我们通过将MMAudioSep与现有的分离模型(包括基于确定性和生成方法的模型)进行比较来评估MMAudioSep的性能,并发现它优于基线模型。此外,我们证明,即使在通过微调获得声音分离功能后,该模型仍保留了原始视频到音频生成的能力。这突出了基础声音生成模型的潜力,以通过与声音相关的下游任务。我们的代码可在https://github.com/sony/mmaudiosep上获得。
摘要:We introduce MMAudioSep, a generative model for video/text-queried sound separation that is founded on a pretrained video-to-audio model. By leveraging knowledge about the relationship between video/text and audio learned through a pretrained audio generative model, we can train the model more efficiently, i.e., the model does not need to be trained from scratch. We evaluate the performance of MMAudioSep by comparing it to existing separation models, including models based on both deterministic and generative approaches, and find it is superior to the baseline models. Furthermore, we demonstrate that even after acquiring functionality for sound separation via fine-tuning, the model retains the ability for original video-to-audio generation. This highlights the potential of foundational sound generation models to be adopted for sound-related downstream tasks. Our code is available at https://github.com/sony/mmaudiosep.


【15】O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion
标题:O_O-VC:合成数据驱动的一对一对齐,实现任意语音转换
链接:https://arxiv.org/abs/2510.09061

作者:Huu Tuong Tu, Huan Vu, cuong tien nguyen, Dien Hy Ngo, Nguyen Thi Thu Trang
备注:EMNLP 2025
摘要:传统的语音转换(VC)方法通常试图将说话人身份和语言信息分离成不同的表示,然后将其组合以重建音频。然而,有效地解开这些因素仍然具有挑战性,往往导致培训过程中的信息丢失。在本文中,我们提出了一种新的方法,利用合成语音数据生成的高质量,预训练的多扬声器文本到语音(TTS)模型。具体地,共享相同语言内容但说话者身份不同的合成数据对被用作输入-输出对以训练语音转换模型。这使模型能够学习源和目标语音之间的直接映射,有效地捕捉特定于说话者的特征,同时保留语言内容。此外,我们还引入了一种灵活的任意语音转换训练策略,可以很好地推广到看不见的说话者和新语言,增强了zero-shot场景中的适应性和性能。实验结果表明,该方法在词错误率上相对降低了16.35%,在说话人余弦相似度上相对提高了5.91%,优于几种最先进的方法。语音转换示例可访问:https://oovc-emnlp-2025.github.io/
摘要:Traditional voice conversion (VC) methods typically attempt to separate speaker identity and linguistic information into distinct representations, which are then combined to reconstruct the audio. However, effectively disentangling these factors remains challenging, often leading to information loss during training. In this paper, we propose a new approach that leverages synthetic speech data generated by a high-quality, pretrained multispeaker text-to-speech (TTS) model. Specifically, synthetic data pairs that share the same linguistic content but differ in speaker identity are used as input-output pairs to train the voice conversion model. This enables the model to learn a direct mapping between source and target voices, effectively capturing speaker-specific characteristics while preserving linguistic content. Additionally, we introduce a flexible training strategy for any-to-any voice conversion that generalizes well to unseen speakers and new languages, enhancing adaptability and performance in zero-shot scenarios. Our experiments show that our proposed method achieves a 16.35% relative reduction in word error rate and a 5.91% improvement in speaker cosine similarity, outperforming several state-of-the-art methods. Voice conversion samples can be accessed at: https://oovc-emnlp-2025.github.io/


【16】Déréverbération non-supervisée de la parole par modèle hybride
标题:修改非监督假释模式混合
链接:https://arxiv.org/abs/2510.09025

作者:Louis Bahrman (IDS, S2A), Mathieu Fontaine (IDS, S2A), Gaël Richard (IDS, S2A)
备注:in French language
摘要:本文介绍了一种新的训练策略,以改善语音去混响系统在无监督的方式只使用混响语音。大多数现有的算法依赖于成对的干/混响数据,这是很难获得的。我们的方法使用有限的声学信息,如混响时间(RT 60),训练去混响系统。实验结果表明,我们的方法实现了更一致的性能在各种客观指标比国家的最先进的。
摘要:This paper introduces a new training strategy to improve speech dereverberation systems in an unsupervised manner using only reverberant speech. Most existing algorithms rely on paired dry/reverberant data, which is difficult to obtain. Our approach uses limited acoustic information, like the reverberation time (RT60), to train a dereverberation system. Experimental results demonstrate that our method achieves more consistent performance across various objective metrics than the state-of-the-art.


【17】DiTSinger: Scaling Singing Voice Synthesis with Diffusion Transformer and Implicit Alignment
标题:DiTSinger:使用扩散Transformer和隐式对齐缩放歌唱声音合成
链接:https://arxiv.org/abs/2510.09016

作者:Zongcai Du, Guilin Deng, Xiaofeng Guo, Xin Gao, Linke Li, Kaichang Cheng, Fubo Han, Siyu Yang, Peng Liu, Pan Zhong, Qiang Fu
备注:under review
摘要:基于扩散的歌唱声音合成(SVS)的最新进展表现出很强的表现力,但仍然受到数据稀缺和模型可扩展性的限制。我们引入了一个两阶段的管道:通过将固定的旋律与不同的LLM生成的歌词配对来构建人类演唱录音的紧凑种子集,并训练特定于旋律的模型来合成超过500小时的高质量中文演唱数据。在此语料库的基础上,我们提出了DiTSinger,一个扩散Transformer与ROPE和qk规范,系统地缩放的深度,宽度和分辨率,以提高保真度。此外,我们设计了一个隐式对齐机制,避免音素级的持续时间标签,通过约束音素到声学注意字符级跨度内,从而提高噪声或不确定的对齐下的鲁棒性。大量的实验验证了我们的方法可以实现可扩展的,无干扰的,高保真的SVS。
摘要:Recent progress in diffusion-based Singing Voice Synthesis (SVS) demonstrates strong expressiveness but remains limited by data scarcity and model scalability. We introduce a two-stage pipeline: a compact seed set of human-sung recordings is constructed by pairing fixed melodies with diverse LLM-generated lyrics, and melody-specific models are trained to synthesize over 500 hours of high-quality Chinese singing data. Building on this corpus, we propose DiTSinger, a Diffusion Transformer with RoPE and qk-norm, systematically scaled in depth, width, and resolution for enhanced fidelity. Furthermore, we design an implicit alignment mechanism that obviates phoneme-level duration labels by constraining phoneme-to-acoustic attention within character-level spans, thereby improving robustness under noisy or uncertain alignments. Extensive experiments validate that our approach enables scalable, alignment-free, and high-fidelity SVS.


【18】VM-UNSSOR: Unsupervised Neural Speech Separation Enhanced by Higher-SNR Virtual Microphone Arrays
标题:VM-UNSsor:通过更高的SNR虚拟麦克风阵列增强无监督神经语音分离
链接:https://arxiv.org/abs/2510.08914

作者:Shulin He, Zhong-Qiu Wang
摘要:盲语音分离(BSS)的目的是在未知阵列几何和房间冲激响应的情况下,从多通道、多说话人混合信号中恢复出多个语音源。在无监督设置中,干净的目标语音无法用于模型训练,UNSSOR提出了一种混合一致性(MC)损失,用于在超定训练混合物上训练深度神经网络(DNN),以实现无监督语音分离。然而,当训练混合信号的麦克风数量减少时,MC约束减弱,分离性能急剧下降。为了解决这个问题,我们提出了VM-UNSSOR,增加了观察到的训练混合信号记录的麦克风与几个更高的SNR虚拟麦克风(VM)信号,这是通过应用线性空间解混器(如IVA和空间聚类)到观察到的训练混合信号。作为所观察到的混合物的线性投影,虚拟麦克风信号通常可以增加每个源的SNR,并且可以被利用来计算额外的MC损耗,以改善UNSSOR并解决UNSSOR中的频率排列问题。在SMS-WSJ数据集上,在超定六麦克风、两扬声器分离设置中,VM-UNSSOR达到17.1 dB SI-SDR,而UNSSOR仅获得14.7 dB;在确定的两麦克风、两扬声器情况下,UNSSOR塌陷至-2.7 dB SI-SDR,而VM-UNSSOR达到10.7 dB。
摘要:Blind speech separation (BSS) aims to recover multiple speech sources from multi-channel, multi-speaker mixtures under unknown array geometry and room impulse responses. In unsupervised setup where clean target speech is not available for model training, UNSSOR proposes a mixture consistency (MC) loss for training deep neural networks (DNN) on over-determined training mixtures to realize unsupervised speech separation. However, when the number of microphones of the training mixtures decreases, the MC constraint weakens and the separation performance falls dramatically. To address this, we propose VM-UNSSOR, augmenting the observed training mixture signals recorded by a limited number of microphones with several higher-SNR virtual-microphone (VM) signals, which are obtained by applying linear spatial demixers (such as IVA and spatial clustering) to the observed training mixtures. As linear projections of the observed mixtures, the virtual-microphone signals can typically increase the SNR of each source and can be leveraged to compute extra MC losses to improve UNSSOR and address the frequency permutation problem in UNSSOR. On the SMS-WSJ dataset, in the over-determined six-microphone, two-speaker separation setup, VM-UNSSOR reaches 17.1 dB SI-SDR, while UNSSOR only obtains 14.7 dB; and in the determined two-microphone, two-speaker case, UNSSOR collapses to -2.7 dB SI-SDR, while VM-UNSSOR achieves 10.7 dB.


【19】ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling
标题:ControlAudio:通过渐进扩散模型处理文本引导、定时指示和可理解的音频生成
链接:https://arxiv.org/abs/2510.08878

作者:Yuxuan Jiang, Zehua Chen, Zeqian Ju, Yusheng Dai, Weibei Dou, Jun Zhu
备注:18 pages, 8 tables, 5 figures
摘要:使用细粒度控制信号生成文本到音频(TTA),例如,精确的定时控制或可理解的语音内容在最近的工作中已经被探索。然而,受数据稀缺性的限制,它们的大规模发电性能仍然受到影响。在这项研究中,我们将可控TTA生成重新定义为多任务学习问题,并引入了一种渐进扩散建模方法ControlAudio。我们的方法巧妙地适合分布条件下更细粒度的信息,包括文本,时间和音素特征,通过一步一步的策略。首先,我们提出了一个数据构建方法跨越注释和模拟,增加条件信息的文本,时间和音素的顺序。其次,在模型训练阶段,我们在大规模文本-音频对上预训练扩散Transformer(DiT),实现可扩展的TTA生成,然后增量地将时序和音素特征与统一的语义表示相结合,扩展可控性。最后,在推理阶段,我们提出了逐步引导的生成,它依次强调更细粒度的信息,内在地与DiT的粗到细采样性质相一致。大量的实验表明,ControlAudio在时间准确性和语音清晰度方面达到了最先进的性能,在客观和主观评价方面都明显优于现有方法。演示示例可在以下网址获得:https://control-audio.github.io/Control-Audio。
摘要:Text-to-audio (TTA) generation with fine-grained control signals, e.g., precise timing control or intelligible speech content, has been explored in recent works. However, constrained by data scarcity, their generation performance at scale is still compromised. In this study, we recast controllable TTA generation as a multi-task learning problem and introduce a progressive diffusion modeling approach, ControlAudio. Our method adeptly fits distributions conditioned on more fine-grained information, including text, timing, and phoneme features, through a step-by-step strategy. First, we propose a data construction method spanning both annotation and simulation, augmenting condition information in the sequence of text, timing, and phoneme. Second, at the model training stage, we pretrain a diffusion transformer (DiT) on large-scale text-audio pairs, achieving scalable TTA generation, and then incrementally integrate the timing and phoneme features with unified semantic representations, expanding controllability. Finally, at the inference stage, we propose progressively guided generation, which sequentially emphasizes more fine-grained information, aligning inherently with the coarse-to-fine sampling nature of DiT. Extensive experiments show that ControlAudio achieves state-of-the-art performance in terms of temporal accuracy and speech clarity, significantly outperforming existing methods on both objective and subjective evaluations. Demo samples are available at: https://control-audio.github.io/Control-Audio.


【20】Audible Networks: Deconstructing and Manipulating Sounds with Deep Non-Negative Autoencoders
标题:可听网络:使用深度非负自动编码器解构和操纵声音
链接:https://arxiv.org/abs/2510.08816

作者:Juan José Burred, Carmine-Emanuele Cella
摘要:我们建议使用非负自动编码器(NAE)进行声音解构和用户引导的声音操作,以达到创造性的目的。NAE为传统的基于非负矩阵分解(NMF)的方法提供了一种通用且可扩展的扩展,用于可解释的音频分解。通过执行非负约束,通过投影梯度下降,我们得到的分解,内部权重和激活可以直接解释为频谱形状和时间包络,其中组件本身可以作为单独的声音事件听。特别是,多层Deep NAE架构能够实现具有可调粒度级别的分层表示,允许在多个抽象级别解构声音:从高级音符包络到细粒度频谱细节。该框架实现了广泛的新的表达,可控和随机的声音转换。我们介绍了新的操作,包括跨组件和跨层合成,分层解构,和几个随机化的策略,控制音色和事件密度。通过可视化和重新合成的实际例子,我们展示了如何NAES可以作为灵活和可解释的工具,基于对象的声音编辑。
摘要:We propose the use of Non-Negative Autoencoders (NAEs) for sound deconstruction and user-guided manipulation of sounds for creative purposes. NAEs offer a versatile and scalable extension of traditional Non-Negative Matrix Factorization (NMF)-based approaches for interpretable audio decomposition. By enforcing non-negativity constraints through projected gradient descent, we obtain decompositions where internal weights and activations can be directly interpreted as spectral shapes and temporal envelopes, and where components can themselves be listened to as individual sound events. In particular, multi-layer Deep NAE architectures enable hierarchical representations with an adjustable level of granularity, allowing sounds to be deconstructed at multiple levels of abstraction: from high-level note envelopes down to fine-grained spectral details. This framework enables a wide new range of expressive, controllable, and randomized sound transformations. We introduce novel manipulation operations including cross-component and cross-layer synthesis, hierarchical deconstructions, and several randomization strategies that control timbre and event density. Through visualizations and resynthesis of practical examples, we demonstrate how NAEs can serve as flexible and interpretable tools for object-based sound editing.


【21】Hierarchical Self-Supervised Representation Learning for Depression Detection from Speech
标题:分层自监督表示学习用于语音抑郁检测
链接:https://arxiv.org/abs/2510.08593

作者:Yuxin Li, Eng Siong Chng, Cuntai Guan
摘要:基于语音的抑郁检测(SDD)是传统临床评估的一种有前途的非侵入性替代方法。然而,随着时间的推移,它仍然受到提取有意义的特征和捕获稀疏,异构抑郁线索的困难的限制。预训练的自监督学习(SSL)模型(如WavLM)提供了丰富的多层语音表示,但大多数现有的SDD方法仅依赖于最后一层或搜索单个最佳性能。这些方法通常过度拟合特定的数据集,并且无法利用检测微妙和持续抑郁信号所需的完整层次结构。   为了应对这一挑战,我们提出了HAREN-CTC,一种新的架构,它集成了多层SSL功能,使用多任务学习框架内的交叉注意,结合连接主义时间分类损失来处理稀疏的时间监督。HAREN-CTC包括两个关键模块:一个分层自适应聚类模块,将SSL特征重组为互补的嵌入,以及一个跨模态融合模块,通过交叉注意力对层间依赖关系进行建模。CTC目标使抑郁感知训练成为可能,允许模型跟踪抑郁言语线索的不规则时间模式。   我们评估HAREN-CTC下的上限设置与标准的数据分割和泛化设置使用五重交叉验证。该模型在DAIC-WOZ上实现了最先进的宏F1分数0.81,在MODMA上实现了0.82,在两种评估场景中均优于先前的方法。
摘要:Speech-based depression detection (SDD) is a promising, non-invasive alternative to traditional clinical assessments. However, it remains limited by the difficulty of extracting meaningful features and capturing sparse, heterogeneous depressive cues over time. Pretrained self-supervised learning (SSL) models such as WavLM provide rich, multi-layer speech representations, yet most existing SDD methods rely only on the final layer or search for a single best-performing one. These approaches often overfit to specific datasets and fail to leverage the full hierarchical structure needed to detect subtle and persistent depression signals.   To address this challenge, we propose HAREN-CTC, a novel architecture that integrates multi-layer SSL features using cross-attention within a multitask learning framework, combined with Connectionist Temporal Classification loss to handle sparse temporal supervision. HAREN-CTC comprises two key modules: a Hierarchical Adaptive Clustering module that reorganizes SSL features into complementary embeddings, and a Cross-Modal Fusion module that models inter-layer dependencies through cross-attention. The CTC objective enables alignment-aware training, allowing the model to track irregular temporal patterns of depressive speech cues.   We evaluate HAREN-CTC under both an upper-bound setting with standard data splits and a generalization setting using five-fold cross-validation. The model achieves state-of-the-art macro F1-scores of 0.81 on DAIC-WOZ and 0.82 on MODMA, outperforming prior methods across both evaluation scenarios.


【22】EGSTalker: Real-Time Audio-Driven Talking Head Generation with Efficient Gaussian Deformation
标题:EGSTalker:具有高效高斯变形的实时音频驱动说话头生成
链接:https://arxiv.org/abs/2510.08587

作者:Tianheng Zhu, Yinfeng Yu, Liejun Wang, Fuchun Sun, Wendong Zheng
备注:Main paper (6 pages). Accepted for publication by IEEE International Conference on Systems, Man, and Cybernetics 2025
摘要:本文介绍了EGSTalker,一个实时音频驱动的说话人头生成框架的基础上三维高斯飞溅(3DGS)。EGSTalker旨在提高速度和视觉保真度,仅需3-5分钟的训练视频即可合成高质量的面部动画。该框架包括两个关键阶段:静态高斯初始化和音频驱动变形。在第一阶段,多分辨率散列三平面和Kolmogorov-Arnold网络(KAN)被用来提取空间特征,并构造一个紧凑的三维高斯表示。在第二阶段,我们提出了一个有效的空间音频注意力(ESAA)模块来融合音频和空间线索,而KAN预测相应的高斯变形。大量的实验表明,EGSTalker实现了渲染质量和唇同步精度相媲美的国家的最先进的方法,同时显着优于他们的推理速度。这些结果突出了EGSTalker的实时多媒体应用的潜力。
摘要:This paper presents EGSTalker, a real-time audio-driven talking head generation framework based on 3D Gaussian Splatting (3DGS). Designed to enhance both speed and visual fidelity, EGSTalker requires only 3-5 minutes of training video to synthesize high-quality facial animations. The framework comprises two key stages: static Gaussian initialization and audio-driven deformation. In the first stage, a multi-resolution hash triplane and a Kolmogorov-Arnold Network (KAN) are used to extract spatial features and construct a compact 3D Gaussian representation. In the second stage, we propose an Efficient Spatial-Audio Attention (ESAA) module to fuse audio and spatial cues, while KAN predicts the corresponding Gaussian deformations. Extensive experiments demonstrate that EGSTalker achieves rendering quality and lip-sync accuracy comparable to state-of-the-art methods, while significantly outperforming them in inference speed. These results highlight EGSTalker's potential for real-time multimedia applications.


【23】Evaluating Hallucinations in Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions
标题:在不同声学条件下评估带有口语按钮的多模式LLM中的幻觉
链接:https://arxiv.org/abs/2510.08581

作者:Hansol Park, Hoseong Ahn, Junwon Moon, Yejin Lee, Kyuhong Shim
摘要:视觉语言模型中的幻觉已经被广泛研究,使用的基准探测图像-文本设置中的可靠性。相比之下,语音查询对多模态幻觉的影响在很大程度上仍未被探索,尽管语音驱动界面的作用越来越大。在这项工作中,我们研究了口语输入如何影响多模态大型语言模型中的幻觉。我们提出了RePOPE-Spk,音频增强扩展的RePOPE基准,查询提供不同的声学条件下的语音。使用RePOPE-Spk,我们系统地评估专有和开源模型。实验结果表明,当查询是口语而不是书面时,幻觉会升级:在干净的语音下错误率增加3%,在环境噪音下增加20%。输入顺序和查询长度进一步影响鲁棒性,而多镜头提示和思维链推理等策略提供了部分但不充分的缓解。这些发现突出了一个关键的和未充分探索的挑战,为构建可靠的语音接口系统开辟了新的方向。
摘要:Hallucinations in vision-language models have been extensively studied using benchmarks that probe reliability in image-text settings. In contrast, the effect of spoken queries on multimodal hallucinations remains largely unexplored, despite the growing role of voice-driven interfaces. In this work, we investigate how spoken input influences hallucinations in multimodal large language models. We present RePOPE-Spk, an audio-augmented extension of the RePOPE benchmark, where queries are provided as speech under diverse acoustic conditions. Using RePOPE-Spk, we systematically evaluate both proprietary and open-source models. Experimental results show that hallucinations escalate when queries are spoken rather than written: error rates increase by 3% under clean speech and by up to 20% with environmental noise. Input order and query length further affect robustness, while strategies such as many-shot prompting and chain-of-thought reasoning offer partial but insufficient mitigation. These findings highlight a critical and underexplored challenge, opening new directions for building reliable voice interface systems.


【24】LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection
标题:LadderSym:用于音乐练习错误检测的多模式交织Transformer
链接:https://arxiv.org/abs/2510.08580

作者:Benjamin Shiue-Hal Chou, Purvish Jajal, Nick John Eliopoulos, James C. Davis, George K. Thiruvathukal, Kristen Yeon-Ji Yun, Yung-Hsiang Lu
备注:Under Submission
摘要:音乐学习者可以从准确检测练习中错误的工具中受益匪浅。现有的方法通常比较音频记录的乐谱,使用的算法或可学习的模型。本文介绍了\textit {LadderSym},一种新的基于transformer的音乐错误检测方法。\textit {LadderSym}由关于最先进方法的两个关键观察结果指导:(1)后期融合限制了流间对齐和跨模态比较能力;以及(2)对乐谱音频的依赖在频谱中引入了模糊性,降低了具有并发音符的音乐的性能。为了解决这些限制,\textit {LadderSym}引入了(1)一个具有流间对齐模块的双流编码器,以提高音频比较能力和错误检测F1分数,以及(2)一种多模式策略,通过将符号表示作为解码器提示,减少歧义并提高F1分数,从而利用音频和符号分数。我们通过测量每个音符类别的F1得分,在\textit {MAESTRO-E}和\textit {CocoChorales-E}数据集上评估我们的方法。与之前的技术水平相比,\textit {LadderSym}在\textit {MAESTRO-E}上对未命中音符的F1检测能力提高了一倍多(26.8\ %$\rightarrow $56.3\ %),并将额外音符检测能力提高了14.4点(72.0\ %$\rightarrow $86.4\ %)。在\textit {CocoChorales-E}上观察到类似的增益。这项工作介绍了有关比较模型的一般见解,这些模型可以为强化学习,人类技能评估和模型评估的序列评估任务提供信息。
摘要:Music learners can greatly benefit from tools that accurately detect errors in their practice. Existing approaches typically compare audio recordings to music scores using heuristics or learnable models. This paper introduces \textit{LadderSym}, a novel Transformer-based method for music error detection. \textit{LadderSym} is guided by two key observations about the state-of-the-art approaches: (1) late fusion limits inter-stream alignment and cross-modality comparison capability; and (2) reliance on score audio introduces ambiguity in the frequency spectrum, degrading performance in music with concurrent notes. To address these limitations, \textit{LadderSym} introduces (1) a two-stream encoder with inter-stream alignment modules to improve audio comparison capabilities and error detection F1 scores, and (2) a multimodal strategy that leverages both audio and symbolic scores by incorporating symbolic representations as decoder prompts, reducing ambiguity and improving F1 scores. We evaluate our method on the \textit{MAESTRO-E} and \textit{CocoChorales-E} datasets by measuring the F1 score for each note category. Compared to the previous state of the art, \textit{LadderSym} more than doubles F1 for missed notes on \textit{MAESTRO-E} (26.8\% $\rightarrow$ 56.3\%) and improves extra note detection by 14.4 points (72.0\% $\rightarrow$ 86.4\%). Similar gains are observed on \textit{CocoChorales-E}. This work introduces general insights about comparison models that could inform sequence evaluation tasks for reinforcement Learning, human skill assessment, and model evaluation.


机器翻译由腾讯交互翻译提供,仅供参考