今日论文合集:cs.SD语音27篇,eess.AS音频处理34篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】SloPalSpeech: A 2,8000-Hour Slovak Speech Corpus from Parliamentary Data
标题:SloPalSpeech:来自议会数据的2,8000小时斯洛伐克演讲数据库
链接:https://arxiv.org/abs/2509.19270

作者:Erik Božík, Marek Šuppa
摘要:像斯洛伐克语这样的低资源语言的自动语音识别(ASR)受到训练数据稀缺的阻碍。为了解决这个问题,我们引入了SlopalSpeech,这是一个新的大规模斯洛伐克ASR数据集,包含来自议会程序的2,806小时的演讲。我们开发了一个强大的处理管道,将长格式的录音对齐并分割成适合模型训练的清晰的30秒音频-转录对。我们使用此数据集来微调几个OpenAI Whisper模型(小型,中型,大型v3和大型v3-turbo),在标准斯洛伐克基准(如Common Voice和FLEURS)上实现显着的字错误率(WER)降低。例如,微调的Whisper-small模型的WER下降了70%,接近更大的Whisper-large-v3模型的基线性能。为了促进未来在低资源语音识别方面的研究,我们公开发布了完整的SlopalSpeech数据集,完全分段的成绩单(6000万字)以及我们所有的微调模型。
摘要:Automatic Speech Recognition (ASR) for low-resource languages like Slovak is hindered by the scarcity of training data. To address this, we introduce SloPalSpeech, a new, large-scale Slovak ASR dataset containing 2,806 hours of speech from parliamentary proceedings. We developed a robust processing pipeline to align and segment long-form recordings into clean, 30-second audio-transcript pairs suitable for model training. We use this dataset to fine-tune several OpenAI Whisper models (small, medium, large-v3, and large-v3-turbo), achieving significant Word Error Rate (WER) reductions on standard Slovak benchmarks like Common Voice and FLEURS. For instance, the fine-tuned Whisper-small model's WER dropped by up to 70\%, approaching the baseline performance of the much larger Whisper-large-v3 model. To foster future research in low-resource speech recognition, we publicly release the complete SloPalSpeech dataset, the fully segmented transcripts (60 million words), and all our fine-tuned models.


【2】Finding My Voice: Generative Reconstruction of Disordered Speech for Automated Clinical Evaluation
标题:寻找我的声音:言语障碍的生成重建,用于自动临床评估
链接:https://arxiv.org/abs/2509.19231

作者:Karen Rosero, Eunjung Yeo, David R. Mortensen, Cortney Van't Slot, Rami R. Hallac, Carlos Busso
摘要:我们提出了ChiReSSD,语音重建框架,保留儿童扬声器的身份,同时抑制发音错误。与之前训练健康成人语音的方法不同,ChiReSSD适应患有语音障碍(SSD)的儿童的声音,特别强调音高和韵律。我们评估我们的方法在STAR数据集上,并报告在词汇准确性和说话人身份保护方面的实质性改进。此外,我们自动预测原始和重建对中的语音内容,其中校正辅音的比例与临床语音评估指标正确辅音的百分比(PCC)相当。我们的实验显示自动和人类专家注释之间的Pearson相关性为0.63,突出了减少手动转录负担的潜力。此外,TORGO数据集上的实验表明,有效的泛化重建成人构音障碍的语音。我们的研究结果表明,解开,基于风格的TTS重建可以提供不同的临床人群的身份保留的讲话。
摘要:We present ChiReSSD, a speech reconstruction framework that preserves children speaker's identity while suppressing mispronunciations. Unlike prior approaches trained on healthy adult speech, ChiReSSD adapts to the voices of children with speech sound disorders (SSD), with particular emphasis on pitch and prosody. We evaluate our method on the STAR dataset and report substantial improvements in lexical accuracy and speaker identity preservation. Furthermore, we automatically predict the phonetic content in the original and reconstructed pairs, where the proportion of corrected consonants is comparable to the percentage of correct consonants (PCC), a clinical speech assessment metric. Our experiments show Pearson correlation of 0.63 between automatic and human expert annotations, highlighting the potential to reduce the manual transcription burden. In addition, experiments on the TORGO dataset demonstrate effective generalization for reconstructing adult dysarthric speech. Our results indicate that disentangled, style-based TTS reconstruction can provide identity-preserving speech across diverse clinical populations.


【3】Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models
标题:更加关注音频:缓解大型音频语言模型中跨模式注意力的不平衡
链接:https://arxiv.org/abs/2509.18816

作者:Junyu Wang, Ziyang Ma, Zhengding Luo, Tianrui Wang, Meng Ge, Xiaobao Wang, Longbiao Wang
备注:Submitted to ICASSP 2026
摘要:大型音频语言模型(LALM)经常遭受音频-文本注意力不平衡,优先考虑文本而不是声学信息,特别是在Transformer架构的多模态融合层中。这种偏见阻碍了他们充分利用声学线索的能力,导致音频推理任务的表现不佳。为了缓解这一问题,我们提出了一种新的无训练方法\textbf{MATA},该方法动态地推动LALM在自注意力机制中支付\textbf{M}或\textbf {A}注意力\textbf{T}或\textbf{A}音频令牌。具体而言,MATA干预后原始注意力评分,仅针对中间层中的最后一个令牌,而不引入额外的参数或计算开销。在MMAU和MMAR基准测试上的实验证实了MATA的有效性,并具有一致的性能增益。值得注意的是,在MMAR上,MATA使开源模型首次超越了专有的Gemini 2.0 Flash。我们的工作提供了一个有效的解决方案,以减轻注意偏差,并打开了一个新的研究方向,提高音频处理能力的多模态模型。
摘要:Large Audio-Language Models (LALMs) often suffer from audio-textual attention imbalance, prioritizing text over acoustic information, particularly in the multi-modal fusion layers of the Transformer architecture. This bias hinders their ability to fully utilize acoustic cues, causing suboptimal performance on audio reasoning tasks. To mitigate this, we propose \textbf{MATA}, a novel training-free method that dynamically pushes LALMs to pay \textbf{M}ore \textbf{A}ttention \textbf{T}o \textbf{A}udio tokens within the self-attention mechanism. Specifically, MATA intervenes post raw attention scoring, targeting only the last token in intermediate layers without introducing additional parameters or computational overhead. Experiments on the MMAU and MMAR benchmarks confirm MATA's effectiveness, with consistent performance gains. Notably, on MMAR, MATA enables an open-source model to surpass the proprietary Gemini 2.0 Flash for the first time. Our work provides an efficient solution to mitigate attention bias and opens a new research direction for enhancing the audio-processing capabilities of multi-modal models.


【4】MECap-R1: Emotion-aware Policy with Reinforcement Learning for Multimodal Emotion Captioning
标题:MECap-R1:多模式情感字幕的强化学习感知策略
链接:https://arxiv.org/abs/2509.18729

作者:Haoqin Sun, Chenyang Lyu, Xiangyu Kong, Shiwan Zhao, Jiaming Zhou, Hui Wang, Aobo Kong, Jinghua Zhao, Longyue Wang, Weihua Luo, Kaifu Zhang, Yong Qin
摘要:语音情感字幕(SEC)已经成为一个值得注意的研究方向。人类语音中情感内容的固有复杂性使得传统的离散分类方法难以提供足够的表示。因此,利用自然语言来描述语音情感提供了一种更有效地捕捉和表达情感的新途径。在本文中,我们提出了MECap-R1,一个开拓性的情感感知政策与强化学习的多模态情感字幕。该框架采用基于情感感知的组相对策略优化(GRE-GRPO),精确地捕捉了字幕的情感和语义特征,从而解决了刚性规则在处理字幕的动态性和灵活性方面的不足。实验结果表明,MECap-R1在生成情感描述方面表现良好,在准确性和多样性方面都有很大的提高。
摘要:Speech Emotion Captioning (SEC) has emerged as a notable research direction. The inherent complexity of emotional content in human speech makes it challenging for traditional discrete classification methods to provide an adequate representation. Consequently, utilizing natural language to describe speech emotions presents a novel avenue for more effectively capturing and expressing affect. In this paper, we propose MECap-R1, a pioneering emotion-aware policy with reinforcement learning for multimodal emotion captioning. By employing Group Relative Policy Optimization with emotion-aware reward (Emo-GRPO), the framework precisely captures the emotion and semantic features, thereby addressing the shortcomings of rigid rules in handling the dynamic and flexible nature of captions. Experimental results on the EmotionTalk dataset demonstrate that MECap-R1 performs well in generating emotion descriptions and achieves substantial gains in both accuracy and diversity.


【5】LOTUSDIS: A Thai far-field meeting corpus for robust conversational ASR
标题:LOTUSDIS:用于强大对话的泰国远场会议文集
链接:https://arxiv.org/abs/2509.18722

作者:Pattara Tipaksorn, Sumonmas Thatphithakkul, Vataya Chunwijitra, Kwanchiva Thangthai
摘要:我们提出LOTUSDIS,一个公开的泰国会议语料库,旨在推进远场会话ASR。该数据集包括114个小时的自发的,无脚本的对话,收集在15-20分钟的会议与三个参与者,其中重叠的语音是频繁和自然的。语音由九个独立的单通道设备同时记录,跨越六种麦克风类型,距离从0.12米到10米,保留混响,噪音和设备着色的真实效果,而不依赖于麦克风阵列。我们提供标准的训练、开发、测试分割并发布可复制的基线系统。我们在zero-shot和微调条件下对几种Whisper变体进行了基准测试。现成的模型随着距离的增加表现出强烈的退化,证实了预训练数据和泰国远场语音之间的不匹配。LOTUSDIS上的微调显著提高了鲁棒性:Thai Whisper基线将总体WER从64.3降至38.3,远场WER从81.6降至49.5,最远麦克风的增益特别大。这些结果强调了距离多样性训练数据对于强大ASR的重要性。该语料库在CC-BY-SA 4.0下可用。我们还发布培训和评估脚本作为基线系统,以促进该领域的可重复研究。
摘要:We present LOTUSDIS, a publicly available Thai meeting corpus designed to advance far-field conversational ASR. The dataset comprises 114 hours of spontaneous, unscripted dialogue collected in 15-20 minute sessions with three participants, where overlapping speech is frequent and natural. Speech was recorded simultaneously by nine independent single-channel devices spanning six microphone types at distances from 0.12 m to 10 m, preserving the authentic effects of reverberation, noise, and device coloration without relying on microphone arrays. We provide standard train, dev, test splits and release a reproducible baseline system. We benchmarked several Whisper variants under zero-shot and fine-tuned conditions. Off-the-shelf models showed strong degradation with distance, confirming a mismatch between pre-training data and Thai far-field speech. Fine-tuning on LOTUSDIS dramatically improved robustness: a Thai Whisper baseline reduced overall WER from 64.3 to 38.3 and far-field WER from 81.6 to 49.5, with especially large gains on the most distant microphones. These results underscore the importance of distance-diverse training data for robust ASR. The corpus is available under CC-BY-SA 4.0. We also release training and evaluation scripts as a baseline system to promote reproducible research in this field.


【6】Enhancing Automatic Chord Recognition through LLM Chain-of-Thought Reasoning
标题:通过LLM思想链推理增强自动和弦识别
链接:https://arxiv.org/abs/2509.18700

作者:Chih-Cheng Chang, Bo-Yu Chen, Lu-Rong Chen, Li Su
摘要:音乐信息检索(MIR)包括用于分析和理解音乐内容的广泛计算技术,最近的深度学习进步推动了实质性的改进。基于这些进展,本文探讨了大型语言模型(LLM)如何作为一个集成的桥梁,连接和集成来自多个MIR工具的信息,重点是提高自动和弦识别性能。我们提出了一种新的方法,将基于文本的LLM定位为智能协调器,该协调器处理和集成来自各种最先进的MIR工具的输出,包括音乐源分离,关键检测,和弦识别和节拍跟踪。我们的方法将音频衍生的音乐信息转换为文本表示,使LLM能够专门针对和弦识别任务进行推理和校正。我们设计了一个5阶段的思想链框架,允许GPT-4 o系统地分析,比较和完善和弦识别结果,通过利用音乐理论知识来整合不同MIR组件的信息。对三个数据集的实验评估表明,多个评估指标的一致改进,MIREX指标的总体准确率提高了1-2.77%。我们的研究结果表明,LLM可以有效地作为MIR管道中的综合桥梁,为音乐信息检索任务中的多工具协调开辟了新的方向。
摘要:Music Information Retrieval (MIR) encompasses a broad range of computational techniques for analyzing and understanding musical content, with recent deep learning advances driving substantial improvements. Building upon these advances, this paper explores how large language models (LLMs) can serve as an integrative bridge to connect and integrate information from multiple MIR tools, with a focus on enhancing automatic chord recognition performance. We present a novel approach that positions text-based LLMs as intelligent coordinators that process and integrate outputs from diverse state-of-the-art MIR tools-including music source separation, key detection, chord recognition, and beat tracking. Our method converts audio-derived musical information into textual representations, enabling LLMs to perform reasoning and correction specifically for chord recognition tasks. We design a 5-stage chain-of-thought framework that allows GPT-4o to systematically analyze, compare, and refine chord recognition results by leveraging music-theoretical knowledge to integrate information across different MIR components. Experimental evaluation on three datasets demonstrates consistent improvements across multiple evaluation metrics, with overall accuracy gains of 1-2.77% on the MIREX metric. Our findings demonstrate that LLMs can effectively function as integrative bridges in MIR pipelines, opening new directions for multi-tool coordination in music information retrieval tasks.


【7】An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
标题:从掩蔽频谱图进行自我监督音频表示学习的神经架构概述
链接:https://arxiv.org/abs/2509.18691

作者:Sarthak Yadav, Sergios Theodoridis, Zheng-Hua Tan
摘要:近年来,自监督学习在没有标记数据的情况下训练深度神经表示方面引起了极大的兴趣。一种这样的自监督学习方法是掩蔽频谱图建模,其中目标是通过预测输入音频频谱图的移除或隐藏部分来学习语义丰富的上下文表示。以Transformer神经架构为核心,掩蔽声谱图建模已成为学习通用音频表示的主要方法,也称为音频基础模型。同时,解决Transformer架构的问题,特别是底层的缩放点积注意力操作,其与输入序列长度成二次缩放,导致了对递归序列建模方法的新兴趣。其中,选择性结构化状态空间模型(如Mamba)和扩展的长短期记忆(xLSTM)是两种最有前途的方法,已被广泛采用。虽然关于这两个专题的工作不断增加,但目前缺乏对这两个专题的交叉点的适当概述。在本文中,我们对上述研究领域进行了全面的概述,涵盖了掩蔽谱图建模和前面提到的神经序列建模架构Mamba和xLSTM。此外,我们比较了Transformers,Mamba和xLSTM基于掩蔽的声谱图模型在一个统一的,可重复的框架上的十个不同的下游音频分类任务,这将有助于感兴趣的读者作出明智的决定,对邻近应用程序的评估方法的适用性。
摘要:In recent years, self-supervised learning has amassed significant interest for training deep neural representations without labeled data. One such self-supervised learning approach is masked spectrogram modeling, where the objective is to learn semantically rich contextual representations by predicting removed or hidden portions of the input audio spectrogram. With the Transformer neural architecture at its core, masked spectrogram modeling has emerged as the prominent approach for learning general purpose audio representations, a.k.a. audio foundation models. Meanwhile, addressing the issues of the Transformer architecture, in particular the underlying Scaled Dot-product Attention operation, which scales quadratically with input sequence length, has led to renewed interest in recurrent sequence modeling approaches. Among them, Selective structured state space models (such as Mamba) and extended Long Short-Term Memory (xLSTM) are the two most promising approaches which have experienced widespread adoption. While the body of work on these two topics continues to grow, there is currently a lack of an adequate overview encompassing the intersection of these topics. In this paper, we present a comprehensive overview of the aforementioned research domains, covering masked spectrogram modeling and the previously mentioned neural sequence modeling architectures, Mamba and xLSTM. Further, we compare Transformers, Mamba and xLSTM based masked spectrogram models in a unified, reproducible framework on ten diverse downstream audio classification tasks, which will help interested readers to make informed decisions regarding suitability of the evaluated approaches to adjacent applications.


【8】Scalable Evaluation for Audio Identification via Synthetic Latent Fingerprint Generation
标题:基于合成潜在指纹生成的可扩展音频识别评估
链接:https://arxiv.org/abs/2509.18620

作者:Aditya Bhattacharjee,, Marco Pasini, Emmanouil Benetos
备注:Under review for International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Barcelona, 2026
摘要:由于缺乏大型公共音乐数据库,现实规模的音频指纹评估受到限制。我们提出了一种无音频的方法,合成的潜在指纹近似真实指纹的分布。我们的方法在预训练的神经音频指纹系统提取的嵌入上训练整流流模型。使用我们的系统生成的合成指纹作为现实的干扰,并在不需要额外的音频的情况下,使检索性能的模拟在大规模。我们通过比较真实数据的分布来评估合成指纹的保真度。我们进一步基准的检索性能在多个国家的最先进的音频指纹识别框架,通过增加真实的参考数据库与合成干扰,并显示与合成干扰获得的缩放趋势密切跟踪与真实干扰。最后,我们规模的合成干扰数据库模型检索性能非常大的数据库,提供了一个实用的度量系统的可扩展性,不依赖于访问音频语料库。
摘要:The evaluation of audio fingerprinting at a realistic scale is limited by the scarcity of large public music databases. We present an audio-free approach that synthesises latent fingerprints which approximate the distribution of real fingerprints. Our method trains a Rectified Flow model on embeddings extracted by pre-trained neural audio fingerprinting systems. The synthetic fingerprints generated using our system act as realistic distractors and enable the simulation of retrieval performance at a large scale without requiring additional audio. We assess the fidelity of synthetic fingerprints by comparing the distributions to real data. We further benchmark the retrieval performances across multiple state-of-the-art audio fingerprinting frameworks by augmenting real reference databases with synthetic distractors, and show that the scaling trends obtained with synthetic distractors closely track those obtained with real distractors. Finally, we scale the synthetic distractor database to model retrieval performance for very large databases, providing a practical metric of system scalability that does not depend on access to audio corpora.


【9】Explore the Reinforcement Learning for the LLM based ASR and TTS system
标题:探索基于LLM的SVR和TTC系统的强化学习
链接:https://arxiv.org/abs/2509.18569

作者: Changfeng Gao, Yabin Li, Keyu An, Zhifu Gao, Zhihao Du, Han Zhao, Xiangang Li
摘要:近年来,大语言模型(LLM)在自动语音识别(ASR)和文本到语音(TTS)系统中发挥了重要作用。虽然强化学习(RL)在基于文本的任务中显着提高了LLM的性能,但由于训练基于音频的模型的复杂性,其在ASR和TTS中的应用仍然没有得到充分的探索。在这项研究中,我们提出了一个轻量级的RL框架,专为基于音频的LLM,可以处理音频输入和生成音频输出。基于这个框架,我们评估了强化学习在ASR和TTS任务上的有效性。对于ASR任务,我们在组相对策略优化(GRPO)框架内实验了不同的基于规则的奖励函数,并研究了RL数据构建的影响。对于TTS任务,我们将GRPO与微分奖励优化(DiffRO)进行比较,并进一步将这两种方法结合起来以提高性能。我们的实验表明,RL可以显着提高ASR和TTS系统的性能,即使在有限的训练数据和少量的优化步骤。
摘要:In recent years, large language models (LLMs) have played an important role in automatic speech recognition (ASR) and text-to-speech (TTS) systems. While reinforcement learning (RL) has significantly enhanced LLM performance in text-based tasks, its application to ASR and TTS remains underexplored due to the complexity of training audio-based models. In this study, we propose a lightweight RL framework tailored for audio-based LLMs that can process audio inputs and generate audio outputs. Based on this framework, we evaluate the effectiveness of reinforcement learning on both ASR and TTS tasks. For the ASR task, we experiment with different rule-based reward functions within the Group Relative Policy Optimization (GRPO) framework and investigate the impact of RL data construction. For the TTS task, we compare GRPO with Differentiable Reward Optimization (DiffRO) and further combine the two approaches to achieve improved performance. Our experiments demonstrate that RL can significantly enhance the performance of both ASR and TTS systems, even with limited training data and a small number of optimization steps.


【10】Scattering Transformer: A Training-Free Transformer Architecture for Heart Murmur Detection
标题:Scattering Transformer:一种用于心脏杂音检测的免训练Transformer架构
链接:https://arxiv.org/abs/2509.18424

作者:Rami Zewail
摘要:为了解决对熟练的临床医生在心音解释方面的需求,最近关于自动心脏听诊的研究工作已经探索了深度学习方法。这些方法中的大多数都是基于监督学习的,在训练数据有限的情况下总是面临挑战。最近,人们对预训练的自监督音频基础模型用于生物医学最终任务的潜力越来越感兴趣。尽管表现出有希望的结果,这些基础模型通常是计算密集型的。在自动心脏听诊的背景下,本研究通过引入散射Transformer(一种用于心脏杂音检测的新型免训练Transformer架构)来探索这些通用音频基础模型的轻量级替代方案。所提出的方法利用标准的小波散射网络,通过在一个类似于变换的架构中引入上下文依赖关系,而无需任何反向传播。我们在公共CirCor DigiScope数据集上评估了我们的方法,直接将其与领先的通用基础模型进行比较。散射Transformer实现了0.786的加权准确率(WAR)和0.697的未加权平均召回率(UAR),证明了与当代最先进方法相比具有高度竞争力的性能。这项研究建立了散射Transformer作为一个可行的和有前途的替代资源有限的设置。
摘要:In an attempt to address the need for skilled clinicians in heart sound interpretation, recent research efforts on automating cardiac auscultation have explored deep learning approaches. The majority of these approaches have been based on supervised learning that is always challenged in occasions where training data is limited. More recently, there has been a growing interest in potentials of pre-trained self-supervised audio foundation models for biomedical end tasks. Despite exhibiting promising results, these foundational models are typically computationally intensive. Within the context of automatic cardiac auscultation, this study explores a lightweight alternative to these general-purpose audio foundation models by introducing the Scattering Transformer, a novel, training-free transformer architecture for heart murmur detection. The proposed method leverages standard wavelet scattering networks by introducing contextual dependencies in a transformer-like architecture without any backpropagation. We evaluate our approach on the public CirCor DigiScope dataset, directly comparing it against leading general-purpose foundational models. The Scattering Transformer achieves a Weighted Accuracy(WAR) of 0.786 and an Unweighted Average Recall(UAR) of 0.697, demonstrating performance highly competitive with contemporary state of the art methods. This study establishes the Scattering Transformer as a viable and promising alternative in resource-constrained setups.


【11】Identifying birdsong syllables without labelled data
标题:在没有标记数据的情况下识别鸟鸣音节
链接:https://arxiv.org/abs/2509.18412

作者:Melisande Teng, Julien Boussard, David Rolnick, Hugo Larochelle
摘要:识别鸟鸣中的音节序列是应对各种挑战的关键,包括鸟类个体识别和更好地理解动物交流和感觉运动学习。最近,机器学习方法在减轻专家手工标记长音频记录的需求方面表现出了巨大的潜力。然而,它们通常仍然依赖于标记数据的可用性来进行模型训练,限制了对少数物种和数据集的适用性。在这项工作中,我们建立了第一个完全无监督的算法分解成音节序列的鸟鸣录音。我们首先检测音节事件,然后将它们聚类以提取模板-音节表示-然后执行匹配追踪以将记录分解为音节序列。我们评估我们的自动注释对人类标签的数据集孟加拉雀的歌曲,发现我们的无监督的方法实现了高性能。我们还表明,我们的方法可以区分单个鸟类在一个物种内,通过其独特的声音签名,孟加拉雀和另一个物种,大山雀。
摘要:Identifying sequences of syllables within birdsongs is key to tackling a wide array of challenges, including bird individual identification and better understanding of animal communication and sensory-motor learning. Recently, machine learning approaches have demonstrated great potential to alleviate the need for experts to label long audio recordings by hand. However, they still typically rely on the availability of labelled data for model training, restricting applicability to a few species and datasets. In this work, we build the first fully unsupervised algorithm to decompose birdsong recordings into sequences of syllables. We first detect syllable events, then cluster them to extract templates --syllable representations-- before performing matching pursuit to decompose the recording as a sequence of syllables. We evaluate our automatic annotations against human labels on a dataset of Bengalese finch songs and find that our unsupervised method achieves high performance. We also demonstrate that our approach can distinguish individual birds within a species through their unique vocal signatures, for both Bengalese finches and another species, the great tit.


【12】A Dimensional Approach to Canine Bark Analysis for Assistance Dog Seizure Signaling
标题:辅助犬癫痫信号的犬吠分析的维度方法
链接:https://arxiv.org/abs/2509.18375

作者:Hailin Song, Shelley Brady, Tomás Ward, Alan F. Smeaton
摘要:犬发声的标准分类对于辅助犬来说是非常有限的,因为样本数据在犬只之间是稀疏和可变的,并且在道德上限制了对所有树皮类型的捕获。我们重新定义这个问题作为一个连续的回归任务在一个二维的唤醒价空间。我们的方法的核心是一个调整后的暹罗网络,它不是基于二进制相似性训练的,而是基于输入样本对之间的顺序和数字距离训练的。经过公共数据集的训练,与回归基线相比,我们的模型在具有挑战性的效价维度上将周转百分比降低了高达50%。在真实世界数据集上的定性验证证实了学习空间在语义上是有意义的,建立了在严重数据限制下分析犬吠的概念验证。
摘要:Standard classification of canine vocalisations is severely limited for assistance dogs, where sample data is sparse and variable across dogs and where capture of the full range of bark types is ethically constrained. We reframe this problem as a continuous regression task within a two-dimensional arousal-valence space. Central to our approach is an adjusted Siamese Network trained not on binary similarity, but on the ordinal and numeric distance between input sample pairs. Trained on a public dataset, our model reduces Turn-around Percentage by up to 50% on the challenging valence dimension compared to a regression baseline. Qualitative validation on a real-world dataset confirms the learned space is semantically meaningful, establishing a proof-of-concept for analysing canine barking under severe data limitations.


【13】StereoFoley: Object-Aware Stereo Audio Generation from Video
标题:StereoFoley:从视频中生成对象感知立体声音频
链接:https://arxiv.org/abs/2509.18272

作者:Tornike Karchkhadze, Kuan-Lin Chen, Mojtaba (Moji)Heydari, Robert Henzel, Alessandro Toso, Mehrez Souden, Joshua Atkins
摘要:我们提出了StereoFoley,一个视频到音频生成框架,产生语义对齐,时间同步,空间准确的立体声在48 kHz。虽然最近的生成视频到音频模型实现了强大的语义和时间保真度,但它们在很大程度上仍然限于单声道或无法提供对象感知的立体声成像,受到缺乏专业混合的空间准确的视频到音频数据集的限制。首先,我们开发并训练一个从视频生成立体声音频的基础模型,在语义准确性和同步方面都达到了最先进的水平。接下来,为了克服数据集的限制,我们引入了一个合成数据生成管道,该管道将视频分析、对象跟踪和音频合成与动态平移和基于距离的响度控制相结合,从而实现空间精确的对象感知声音。最后,我们在这个合成数据集上微调基础模型,产生清晰的对象-音频对应关系。由于没有既定的指标存在,我们引入立体声对象感知措施,并通过人类听力研究验证它,表现出很强的相关性与感知。这项工作为立体声对象感知视频到音频生成建立了第一个端到端框架,解决了一个关键的差距,并在该领域建立了一个新的基准。
摘要:We present StereoFoley, a video-to-audio generation framework that produces semantically aligned, temporally synchronized, and spatially accurate stereo sound at 48 kHz. While recent generative video-to-audio models achieve strong semantic and temporal fidelity, they largely remain limited to mono or fail to deliver object-aware stereo imaging, constrained by the lack of professionally mixed, spatially accurate video-to-audio datasets. First, we develop and train a base model that generates stereo audio from video, achieving state-of-the-art in both semantic accuracy and synchronization. Next, to overcome dataset limitations, we introduce a synthetic data generation pipeline that combines video analysis, object tracking, and audio synthesis with dynamic panning and distance-based loudness controls, enabling spatially accurate object-aware sound. Finally, we fine-tune the base model on this synthetic dataset, yielding clear object-audio correspondence. Since no established metrics exist, we introduce stereo object-awareness measures and validate it through a human listening study, showing strong correlation with perception. This work establishes the first end-to-end framework for stereo object-aware video-to-audio generation, addressing a critical gap and setting a new benchmark in the field.


【14】MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech
标题:MNV-17:用于语音非言语发声识别的高质量表演性普通话数据集
链接:https://arxiv.org/abs/2509.18196

作者:Jialong Mai, Jinxin Ji, Xiaofen Xing, Chen Yang, Weidong Chen, Jingyuan Xing, Xiangmin Xu
备注:Submitted to ICASSP 2026
摘要:主流的自动语音识别(ASR)系统擅长转录词汇内容,但在很大程度上无法识别嵌入在语音中的非言语发声(NV),如叹息,笑声和咳嗽。这种能力对于全面理解人类交流非常重要,因为NVs传达了关键的情感和意图线索。NV-aware ASR的进展一直受到缺乏高质量,注释良好的数据集的阻碍。为了解决这一差距,我们引入了MNV-17,一个7.55小时的表演性普通话语音数据集。与大多数依赖于基于模型的检测的现有语料库不同,MNV-17的表演性质确保了高保真度,清晰的NV实例。据我们所知,MNV-17提供了最广泛的非语言发声类别,包括17个不同的和平衡的常见NV类。我们在四种主流ASR架构上对MNV-17进行了基准测试,评估了它们在语义转录和NV分类方面的联合性能。数据集和预训练模型检查点将公开提供,以促进未来表达性ASR的研究。
摘要:Mainstream Automatic Speech Recognition (ASR) systems excel at transcribing lexical content, but largely fail to recognize nonverbal vocalizations (NVs) embedded in speech, such as sighs, laughs, and coughs. This capability is important for a comprehensive understanding of human communication, as NVs convey crucial emotional and intentional cues. Progress in NV-aware ASR has been hindered by the lack of high-quality, well-annotated datasets. To address this gap, we introduce MNV-17, a 7.55-hour performative Mandarin speech dataset. Unlike most existing corpora that rely on model-based detection, MNV-17's performative nature ensures high-fidelity, clearly articulated NV instances. To the best of our knowledge, MNV-17 provides the most extensive set of nonverbal vocalization categories, comprising 17 distinct and well-balanced classes of common NVs. We benchmarked MNV-17 on four mainstream ASR architectures, evaluating their joint performance on semantic transcription and NV classification. The dataset and the pretrained model checkpoints will be made publicly available to facilitate future research in expressive ASR.


【15】XMUspeech Systems for the ASVspoof 5 Challenge
标题:ASVspoof 5挑战赛的XMUspeech系统
链接:https://arxiv.org/abs/2509.18102

作者:Wangjie Li1, Xingjia Xie, Yishuang Li, Wenhao Guan, Kaidi Wang, Pengyu Ren, Lin Li, Qingyang Hong
摘要:在本文中,我们将我们提交的XMU语音系统提交给ASVspoof 5挑战赛的语音deepfake检测轨道。与之前的挑战相比,ASVspoof 5数据库中的音频持续时间显着增加。我们观察到,仅仅调整输入音频长度就可以大大提高系统性能。为了在多个级别上捕获伪影,我们探索了AASIST、HM-Conformer、Hubert和Wav 2 vec 2在各种输入特征和损失函数下的性能。具体来说,为了获得伪影相关信息,我们在包含欺骗话语的数据集上训练了自监督模型作为特征提取器。我们采用了自适应多尺度特征融合(AMFF)方法,将多个Transformer层的特征与手工特征相结合,以增强检测能力。此外,我们对单类损失函数进行了广泛的实验,并提供了优化的配置,以更好地配合反欺骗任务。我们的融合系统在闭合条件下的minDCF为0.4783,EER为20.45%,在开放条件下的minDCF为0.2245,EER为9.36%。
摘要:In this paper, we present our submitted XMUspeech systems to the speech deepfake detection track of the ASVspoof 5 Challenge. Compared to previous challenges, the audio duration in ASVspoof 5 database has significantly increased. And we observed that merely adjusting the input audio length can substantially improve system performance. To capture artifacts at multiple levels, we explored the performance of AASIST, HM-Conformer, Hubert, and Wav2vec2 with various input features and loss functions. Specifically, in order to obtain artifact-related information, we trained self-supervised models on the dataset containing spoofing utterances as the feature extractors. And we applied an adaptive multi-scale feature fusion (AMFF) method to integrate features from multiple Transformer layers with the hand-crafted feature to enhance the detection capability. In addition, we conducted extensive experiments on one-class loss functions and provided optimized configurations to better align with the anti-spoofing task. Our fusion system achieved a minDCF of 0.4783 and an EER of 20.45% in the closed condition, and a minDCF of 0.2245 and an EER of 9.36% in the open condition.


【16】Audio-Based Pedestrian Detection in the Presence of Vehicular Noise
标题:存在车辆噪音时基于音频的行人检测
链接:https://arxiv.org/abs/2509.19295

作者:Yonghyun Kim, Chaeyeon Han, Akash Sarode, Noah Posner, Subhrajit Guhathakurta, Alexander Lerch
备注:Accepted to the 10th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE), 2025
摘要:基于音频的行人检测是一项具有挑战性的任务,迄今为止,仅在噪声受限的环境中进行了探索。我们提出了一个新的数据集,结果,并详细分析了最先进的基于音频的行人检测中存在的车辆噪声。在我们的研究中,我们进行了三项分析:(i)噪声和噪声限制环境之间的跨数据集评估,(ii)评估噪声数据对模型性能的影响,突出声学背景的影响,以及(iii)评估模型对域外声音的预测鲁棒性。新数据集是一个全面的1321小时路边数据集。它结合了丰富的交通音景。每个记录包括与帧级行人注释和1fps视频缩略图同步的16kHz音频。
摘要:Audio-based pedestrian detection is a challenging task and has, thus far, only been explored in noise-limited environments. We present a new dataset, results, and a detailed analysis of the state-of-the-art in audio-based pedestrian detection in the presence of vehicular noise. In our study, we conduct three analyses: (i) cross-dataset evaluation between noisy and noise-limited environments, (ii) an assessment of the impact of noisy data on model performance, highlighting the influence of acoustic context, and (iii) an evaluation of the model's predictive robustness on out-of-domain sounds. The new dataset is a comprehensive 1321-hour roadside dataset. It incorporates traffic-rich soundscapes. Each recording includes 16kHz audio synchronized with frame-level pedestrian annotations and 1fps video thumbnails.


【17】Improving Test-Time Performance of RVQ-based Neural Codecs
标题:提高基于RVQ的神经编解码器的测试时间性能
链接:https://arxiv.org/abs/2509.19186

作者:Hyeongju Kim, Junhyeok Lee, Jacob Morton, Juheon Lee, Jinhyeok Yang
备注:5 pages, preprint
摘要:残差矢量量化(RVQ)技术在神经音频编解码器的最新进展中起着核心作用。这些模型有效地合成高保真音频从有限数量的代码,由于量化级别之间的层次结构。在本文中,我们提出了一种编码算法,以进一步提高合成质量的RVQ为基础的神经编解码器在测试时。首先,我们指出的次优性质的量化矢量产生的传统方法。我们证明,量化误差可以通过选择一组不同的代码来减轻。随后,我们提出了我们的编码算法,旨在确定一组离散代码,实现较低的量化误差。然后,我们将所提出的方法应用于预先训练的模型,并使用不同的指标评估其有效性。实验结果表明,该方法不仅减少了量化误差,而且提高了合成质量。
摘要:The residual vector quantization (RVQ) technique plays a central role in recent advances in neural audio codecs. These models effectively synthesize high-fidelity audio from a limited number of codes due to the hierarchical structure among quantization levels. In this paper, we propose an encoding algorithm to further enhance the synthesis quality of RVQ-based neural codecs at test-time. Firstly, we point out the suboptimal nature of quantized vectors generated by conventional methods. We demonstrate that quantization error can be mitigated by selecting a different set of codes. Subsequently, we present our encoding algorithm, designed to identify a set of discrete codes that achieve a lower quantization error. We then apply the proposed method to pre-trained models and evaluate its efficacy using diverse metrics. Our experimental findings validate that our method not only reduces quantization errors, but also improves synthesis quality.


【18】Training Flow Matching Models with Reliable Labels via Self-Purification
标题:通过自我净化训练具有可靠标签的流匹配模型
链接:https://arxiv.org/abs/2509.19091

作者:Hyeongju Kim, Yechan Yu, June Young Yi, Juheon Lee
备注:5 pages, 3 figures, preprint
摘要:训练数据集本质上是不完美的,通常包含由于人为注释错误、标记模型的限制和其他噪声源而导致的错误标记样本。这种标签污染会显著降低训练模型的性能。在这项工作中,我们介绍了自净化流匹配(SPFM),过滤不可靠的数据流匹配框架内的原则性方法。SPFM在训练过程中使用模型本身识别可疑数据,绕过了对预训练模型或其他模块的需求。我们的实验表明,使用SPFM训练的模型生成的样本准确地遵守指定的条件,即使在嘈杂的标签上训练。此外,我们验证了SPFM在TITW数据集上的鲁棒性,该数据集由野外语音数据组成,实现了超越现有基线的性能。
摘要:Training datasets are inherently imperfect, often containing mislabeled samples due to human annotation errors, limitations of tagging models, and other sources of noise. Such label contamination can significantly degrade the performance of a trained model. In this work, we introduce Self-Purifying Flow Matching (SPFM), a principled approach to filtering unreliable data within the flow-matching framework. SPFM identifies suspicious data using the model itself during the training process, bypassing the need for pretrained models or additional modules. Our experiments demonstrate that models trained with SPFM generate samples that accurately adhere to the specified conditioning, even when trained on noisy labels. Furthermore, we validate the robustness of SPFM on the TITW dataset, which consists of in-the-wild speech data, achieving performance that surpasses existing baselines.


【19】HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS
标题:HD-PPT:基于指令的TTC的内容和预算偏好令牌的分层解码
链接:https://arxiv.org/abs/2509.19001

作者:Sihang Nie, Xiaofen Xing, Jingyuan Xing, Baiji Liu, Xiangmin Xu
备注:5 pages, 2 figures, submitted to ICASSP2026
摘要:基于大语言模型(LLM)的文语转换(TTS)模型已经达到了很高的自然度。然而,TTS推理的精度控制仍然是一个挑战。虽然基于解释的文本到语音转换(Instruct-TTS)模型被提出,但是由于单级文本指令和多级语音令牌之间的模态间隙,这些模型仍然缺乏细粒度的控制。为了解决这一限制,我们提出了HD-PPT,一个框架,将语音合成转换成一个结构化的,层次化的任务。为了实现细粒度控制,我们引入了一种新的语音编解码器,从复杂的语音令牌中提取不同的语音偏好和内容偏好令牌,由自动语音识别(ASR)和跨语言音频文本预训练(CLAP)目标监督。为了弥合这些令牌的模态差距,我们提出了一种分层解码策略,其中LLM以结构化的顺序生成令牌:首先是语义,然后是细粒度风格,最后是完整的声学表示。大量的实验表明,这种分层模式显着提高了指令的遵守,并实现了最先进的自然,验证了我们的方法,精确和可控的语音合成。音频样本可在https://xxh333.github.io/上获得。
摘要:Large Language Model (LLM)-based Text-to-Speech (TTS) models have already reached a high degree of naturalness. However, the precision control of TTS inference is still challenging. Although instruction-based Text-to-Speech (Instruct-TTS) models are proposed, these models still lack fine-grained control due to the modality gap between single-level text instructions and multilevel speech tokens. To address this limitation, we propose HD-PPT, a framework that transforms speech synthesis into a structured, hierarchical task. To enable fine-grained control, we introduce a novel speech codec to extract distinct prompt-preference and content-preference tokens from the complex speech tokens, supervised by automatic speech recognition (ASR) and cross-lingual audio-text pre-training (CLAP) objectives. To bridge the modality gap of these tokens, we propose a hierarchical decoding strategy, where the LLM generates tokens in a structured order: first semantic, then fine-grained style, and finally complete acoustic representation. Extensive experiments demonstrate that this hierarchical paradigm significantly improves instruction adherence and achieves state-of-the-art naturalness, validating our approach for precise and controllable speech synthesis. Audio samples are available at https://xxh333.github.io/.


【20】FlexSED: Towards Open-Vocabulary Sound Event Detection
标题:FlexMED:迈向开放词汇声音事件检测
链接:https://arxiv.org/abs/2509.18606

作者:Jiarui Hai, Helin Wang, Weizhe Guo, Mounya Elhilali
摘要:尽管最近在能够处理数百个声音类别的大规模声音事件检测(SED)系统中取得了进展,但现有的多类别分类框架仍然从根本上受到限制。它们不能处理自由文本声音查询,而自由文本声音查询可以实现更灵活和用户友好的交互,并且它们缺乏zero-shot功能,并且提供较差的Few-Shot适应性。虽然基于文本查询的分离方法已被探索,他们主要集中在源分离和不适合的SED任务,需要精确的时间定位和高效的检测,在大型和多样化的声音词汇。在本文中,我们提出了FlexSED,一个开放词汇的声音事件检测系统。FlexSED建立在预训练的音频SSL模型和CLAP文本编码器的基础上,引入了编码器-解码器组合和自适应融合策略,以实现从预训练权重进行有效的连续训练。为了确保强大的监督,它还采用了大型语言模型(LLM)来帮助在训练过程中选择事件查询,解决与缺失标签相关的挑战。因此,与AudioSet-Strong上的普通SED模型相比,FlexSED实现了卓越的性能,同时展示了强大的zero-shot和Few-Shot功能。我们发布代码和预训练模型,以支持未来基于FlexSED的研究和应用。
摘要:Despite recent progress in large-scale sound event detection (SED) systems capable of handling hundreds of sound classes, existing multi-class classification frameworks remain fundamentally limited. They cannot process free-text sound queries, which enable more flexible and user-friendly interaction, and they lack zero-shot capabilities and offer poor few-shot adaptability. Although text-query-based separation methods have been explored, they primarily focus on source separation and are ill-suited for SED tasks that require precise temporal localization and efficient detection across large and diverse sound vocabularies. In this paper, we propose FlexSED, an open-vocabulary sound event detection system. FlexSED builds on a pretrained audio SSL model and the CLAP text encoder, introducing an encoder-decoder composition and an adaptive fusion strategy to enable effective continuous training from pretrained weights. To ensure robust supervision, it also employs large language models (LLMs) to assist in event query selection during training, addressing challenges related to missing labels. As a result, FlexSED achieves superior performance compared to vanilla SED models on AudioSet-Strong, while demonstrating strong zero-shot and few-shot capabilities. We release the code and pretrained models to support future research and applications based on FlexSED.


【21】SynSonic: Augmenting Sound Event Detection through Text-to-Audio Diffusion ControlNet and Effective Sample Filtering
标题:SynSonic:通过文本到音频扩散控制网和有效的样本过滤增强声音事件检测
链接:https://arxiv.org/abs/2509.18603

作者:Jiarui Hai, Mounya Elhilali
摘要:由于时间标记数据的稀缺性,数据合成和增强是声音事件检测(SED)的关键。虽然SpecAugment和Mix-up等增强方法可以提高模型性能,但它们仍然受到现有样本多样性的限制。最近的生成模型提供了新的机会,但它们的直接应用到SED是具有挑战性的,由于缺乏精确的时间注释和通过不可靠的过滤引入噪声的风险。为了解决这些挑战,并使生成为基础的增强SED,我们提出了SynSonic,为这项任务量身定制的数据增强方法。SynSonic利用由能量包络ControlNet引导的文本到音频扩散模型来生成时间相干的声音事件。具有双分类器的联合评分过滤策略确保了样本质量,并且我们探索了将其实际集成到训练管道中。实验结果表明,SynSonic提高了复调声音检测分数(PSDS 1和PSDS 2),增强了时间定位和声音类别区分。
摘要:Data synthesis and augmentation are essential for Sound Event Detection (SED) due to the scarcity of temporally labeled data. While augmentation methods like SpecAugment and Mix-up can enhance model performance, they remain constrained by the diversity of existing samples. Recent generative models offer new opportunities, yet their direct application to SED is challenging due to the lack of precise temporal annotations and the risk of introducing noise through unreliable filtering. To address these challenges and enable generative-based augmentation for SED, we propose SynSonic, a data augmentation method tailored for this task. SynSonic leverages text-to-audio diffusion models guided by an energy-envelope ControlNet to generate temporally coherent sound events. A joint score filtering strategy with dual classifiers ensures sample quality, and we explore its practical integration into training pipelines. Experimental results show that SynSonic improves Polyphonic Sound Detection Scores (PSDS1 and PSDS2), enhancing both temporal localization and sound class discrimination.


【22】Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation
标题:教授音频模型推理:源蒸馏和分层蒸馏的统一框架
链接:https://arxiv.org/abs/2509.18579

作者:Runyan Yang, Yuke Si, Yingying Gao, Junlan Feng, Chao Deng, Shilei Zhang
备注:5 pages; submitted to ICASSP 2026
摘要:虽然大型音频语言模型擅长ASR和情感识别等任务,但由于音频和文本之间的模态差距以及缺乏结构化的中间监督,它们仍然难以进行复杂的推理。为了解决这个问题,我们提出了一个统一的知识蒸馏框架,将推理能力从高容量的文本教师模型转移到学生音频模型,同时保留其声学能力。我们的方法引入了两个关键维度:源级蒸馏,它利用文本和声学教师提供互补的特定模态监督;分层蒸馏,它将教师信号与适当的学生层对齐,以提高传输效率。这种双维度策略能够对蒸馏过程进行细粒度控制,有效地弥合了符号推理和语音表示之间的差距。实验结果表明,音频推理性能的显着改善,证明了我们的框架作为音频建模的推理转移解决方案的有效性。
摘要:While large audio language models excel at tasks like ASR and emotion recognition, they still struggle with complex reasoning due to the modality gap between audio and text as well as the lack of structured intermediate supervision. To address this, we propose a unified knowledge distillation framework to transfer reasoning capabilities from a high-capacity textual teacher model to a student audio models while preserving its acoustic competence. Our method introduces two key dimensions: source-wise distillation, which leverages both textual and acoustic teachers to provide complementary modality-specific supervision; and layer-wise distillation, which aligns teacher signals with appropriate student layers to improve transfer efficiency. This dual-dimensional strategy enables fine-grained control over the distillation process, effectively bridging the gap between symbolic reasoning and speech representations. Experimental results show significant improvements in audio reasoning performance, demonstrating the effectiveness of our framework as a reasoning transfer solution for audio modeling.


【23】HarmoniFuse: A Component-Selective and Prompt-Adaptive Framework for Multi-Task Speech Language Modeling
标题:Harmonious:一个用于多任务语音语言建模的对象选择和对象自适应框架
链接:https://arxiv.org/abs/2509.18570

作者:Yuke Si, Runyan Yang, Yingying Gao, Junlan Feng, Chao Deng, Shilei Zhang
备注:5 pages; submitted to ICASSP 2026
摘要:大型语言模型的最新进展促进了能够在共享架构内支持多个语音任务的统一语音语言模型(SLM)的开发。然而,诸如自动语音识别(ASR)和语音情感识别(SER)等任务依赖于不同类型的信息:ASR主要依赖于语言内容,而SER需要整合语言和非语言线索。现有的多任务SLM通常采用朴素的参数共享或基于约束的条件,而没有明确地建模每个任务所需的信息组成的差异。这种设计有任务干扰和性能下降的风险,特别是在有限的数据条件下。为了解决这些局限性,我们提出了Harmoniplane,一个组件选择性和自适应框架的多任务语音语言建模。Harmonietry旨在通过选择和融合语音表示的任务相关组件来协调异构任务需求。具体来说,它集成了一个门控语音编码器,以提取特定于任务的声学特征和一个自适应动态融合模块聚合Transformer层的任务特性的基础上。此外,批量交错训练策略可以利用单独的ASR和SER数据集,而无需联合注释。实验结果表明,谐波改善ASR和SER的性能,提供了一个可扩展的和强大的解决方案,多任务语音理解现实的数据约束下。
摘要:Recent advances in large language models have facilitated the development of unified speech language models (SLMs) capable of supporting multiple speech tasks within a shared architecture. However, tasks such as automatic speech recognition (ASR) and speech emotion recognition (SER) rely on distinct types of information: ASR primarily depends on linguistic content, whereas SER requires the integration of both linguistic and paralinguistic cues. Existing multitask SLMs typically adopt naive parameter sharing or prompt-based conditioning without explicitly modeling the differences in information composition required by each task. Such designs risk task interference and performance degradation, especially under limited data conditions. To address these limitations, we propose HarmoniFuse, a component-selective and prompt-adaptive framework for multi-task speech language modeling. HarmoniFuse is designed to harmonize heterogeneous task demands by selecting and fusing task-relevant components of speech representations. Specifically, it integrates a gated speech encoder to extract task-specific acoustic features and a prompt-adaptive dynamic fusion module to aggregate transformer layers based on task characteristics. In addition, a batch-interleaved training strategy enables leveraging separate ASR and SER datasets without requiring joint annotation. Experimental results demonstrate that HarmoniFuse improves both ASR and SER performance, offering a scalable and robust solution for multitask speech understanding under realistic data constraints.


【24】SoundCompass: Navigating Target Sound Extraction With Effective Directional Clue Integration In Complex Acoustic Scenes
标题:SoundCompass:在复杂声学场景中通过有效的方向线索集成进行导航目标声音提取
链接:https://arxiv.org/abs/2509.18561

作者:Dayun Choi, Jung-Woo Choi
备注:5 pages, 4 figures, submitted to ICASSP 2026
摘要:目标声音提取(TSE)的最新进展利用来自到达方向(DoA)的方向线索,其表示在任何声学场景中可用的声音的固有空间属性。然而,以前的基于DoA的方法依赖于手工制作的特征或离散编码,这丢失了细粒度的空间信息并限制了适应性。我们提出了SoundCompass,一个有效的方向线索集成框架为中心的频谱成对相互作用(SPIN)模块,捕捉跨通道的空间相关性在复杂的频谱域,以保留完整的空间信息,在多通道信号。表示在空间相关性方面的输入特征与表示为球面谐波(SH)编码的DoA线索融合。融合是在重叠的频率子带上进行的,继承了以前的带分离架构中报告的优点。我们还将迭代细化策略,推理链(CoI),在TSE框架中,递归地融合DoA与从先前的推理阶段估计的声音事件激活。实验表明,SoundCompass结合SPIN,SH嵌入和CoI,可以在不同的信号类别和空间配置中鲁棒地提取目标源。
摘要:Recent advances in target sound extraction (TSE) utilize directional clues derived from direction of arrival (DoA), which represent an inherent spatial property of sound available in any acoustic scene. However, previous DoA-based methods rely on hand-crafted features or discrete encodings, which lose fine-grained spatial information and limit adaptability. We propose SoundCompass, an effective directional clue integration framework centered on a Spectral Pairwise INteraction (SPIN) module that captures cross-channel spatial correlations in the complex spectrogram domain to preserve full spatial information in multichannel signals. The input feature expressed in terms of spatial correlations is fused with a DoA clue represented as spherical harmonics (SH) encoding. The fusion is carried out across overlapping frequency subbands, inheriting the benefits reported in the previous band-split architectures. We also incorporate the iterative refinement strategy, chain-of-inference (CoI), in the TSE framework, which recursively fuses DoA with sound event activation estimated from the previous inference stage. Experiments demonstrate that SoundCompass, combining SPIN, SH embedding, and CoI, robustly extracts target sources across diverse signal classes and spatial configurations.


【25】No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS
标题:韵律无可验证奖励:迈向DTS中偏好引导的韵律学习
链接:https://arxiv.org/abs/2509.18531

作者:Seungyoun Shin, Dongha Ahn, Jiwoo Kim, Sungwook Jeon
备注:submitted to ICASSP 2026
摘要:最近的工作报告在神经文本到语音(TTS)与组相对策略优化(GRPO)的收益。然而,在缺乏可验证的奖励\textit{韵律}的情况下,在转录导向信号(CER/NLL)上训练的GRPO降低了错误率,但将韵律压缩成单调的、不自然的语音;增加说话者相似性进一步破坏了训练的稳定性,降低了CER。我们用一个迭代直接偏好优化(DPO)方案来解决这个问题,该方案每轮只使用几百个人类标记的偏好对来直接优化韵律自然度,同时正则化到当前模型。在\textbf{KoCC-TTS}上,这是一个真实的韩国呼叫中心交互的策划数据集,捕获了面向任务的对话,我们的方法获得了最高的人类偏好(ELO),具有竞争力的CER,优于GRPO和强大的商业基线。这些结果表明,当韵律不能自动奖励,\textit{人类偏好优化}提供了一个实用的和数据有效的路径自然和强大的TTS。演示页面位于\href{https:tts.ch.dev}
摘要:Recent work reports gains in neural text-to-speech (TTS) with Group Relative Policy Optimization (GRPO). However, in the absence of a verifiable reward for \textit{prosody}, GRPO trained on transcription-oriented signals (CER/NLL) lowers error rates yet collapses prosody into monotone, unnatural speech; adding speaker-similarity further destabilizes training and degrades CER. We address this with an \textit{iterative Direct Preference Optimization (DPO)} scheme that uses only a few hundred human-labeled preference pairs per round to directly optimize prosodic naturalness while regularizing to the current model. On \textbf{KoCC-TTS}, a curated dataset of authentic Korean call center interactions capturing task-oriented dialogues, our method attains the highest human preference (ELO) with competitive CER, outperforming GRPO and strong commercial baselines. These results suggest that when prosody cannot be rewarded automatically, \textit{human preference optimization} offers a practical and data-efficient path to natural and robust TTS. The demo page is available at \href{https://tts.ch.dev}


【26】Qubit Instrumentation of Entanglement
标题:量子比特纠缠仪器
链接:https://arxiv.org/abs/2509.18340

作者:Mark Carney
备注:28 pages, 4 figures, book chapter
摘要:本章和其中描述的实验探讨了如何通过物理纠缠来表现甚至模拟“人类纠缠”。为了实现这一点,两个音乐家之间的“音调中心性”的概念通过嵌入式设备(Raspberry Pi Pico)捕获,并作为参数传递到嵌入式设备上发生的量子模拟中。然后,这些模拟的结果被编码回MIDI,并发送到玩家的乐器。音乐家的音调越接近,他们的乐器就越会被缠绕在一起。|\Phi^+ \rangle$态,它们离得越远,它们的仪器就越会纠缠在一个|\Psi^+ \rangle$ state。其目的是创建相关的随机参数--在两种工具上相同-或反相关- \r {即}在演奏者的音调关系的影响下,这些具有这些特殊性质的随机参数为量子音乐的表达增加了一个新的维度。这个概念是通过实验实现的,并提供了完整的代码和示例输出。这项工作旨在为音乐家探索和体验自己音乐体验的量子仿真铺平道路,为纠缠合奏的未来增加新的细微差别和可能性。
摘要:This chapter and the experiments described within explore how `human entanglement' might be represented and even emulated by physical entanglement. To achieve this, a notion of `tonal centrality' between two musicians is captured via MIDI and passed as a parameter into a quantum simulation taking place on an embedded device (a Raspberry Pi Pico). The results of these simulations are then coded back into MIDI and sent to the players' instruments. The closer the musicians' tonality is, the more their instruments will be entangled in a $|\Phi^+ \rangle$ state, and the further away they are the more their instruments will be entangled in a $|\Psi^+ \rangle$ state. The intention is to create random parameters that are correlative - \emph{i.e.} the same on both instruments - or anti-correlative - \emph{i.e.} the bit-wise opposite of each other, influenced by the tonal relationship from the players. These random parameters sharing these particular properties add a new dimension for quantum-musical expression. This concept was realised experimentally, and the full code and sample outputs are provided. This work aims to pave the way for musicians to explore and experience quantum emulations of their own musical experiences, adding a new nuance and possibilities for the future of \emph{entangled ensembles.}


【27】Automated Analysis of Naturalistic Recordings in Early Childhood: Applications, Challenges, and Opportunities
标题:幼儿自然记录的自动分析:应用、挑战和机遇
链接:https://arxiv.org/abs/2509.18235

作者: Jialu Li Member, Marvin Lavechin, Xulin Fan, Nancy L. McElwain, Alejandrina Cristia, Paola Garcia-Perera, Mark Hasegawa-Johnson
备注:Accepted to IEEE Signal Processing Magazine
摘要:自然主义录音捕捉真实世界环境中的音频,参与者在没有研究人员或实验协议干扰的情况下自然地表现出来。自然主义的长形式录音通过捕捉参与者日常生活中长时间的自发和持续的互动来扩展这一概念,通常跨越数小时甚至数天。自然主义录音已被广泛用于研究儿童的行为,包括他们如何与他人在他们的环境中互动,在心理学,教育,认知科学和临床研究领域。这些录音提供了一种不引人注目的方式来观察儿童在现实世界中的设置超出控制和约束的实验环境。语音技术和机器学习的进步为研究人员自动和系统地分析儿童的大规模自然录音提供了第一步。尽管机器学习模型的准确性并不完美,但这些工具仍然为揭示儿童认知和社会发展的重要见解提供了宝贵的机会。几个关键的语音技术包括说话人日记,发声分类,从成人的字数估计,说话人验证,和语码转换的语言日记。这些技术中的大多数主要是为成年人开发的,而专门应用于儿童的语音技术仍处于探索阶段。为了填补这一空白,我们讨论了推进这些技术以分析儿童早期发育(<3岁)期间的自然记录的当前进展、挑战和机遇。我们致力于激励信号处理社区,促进跨学科合作,以进一步发展这一新兴技术,并应对其独特的挑战和机遇。
摘要:Naturalistic recordings capture audio in real-world environments where participants behave naturally without interference from researchers or experimental protocols. Naturalistic long-form recordings extend this concept by capturing spontaneous and continuous interactions over extended periods, often spanning hours or even days, in participants' daily lives. Naturalistic recordings have been extensively used to study children's behaviors, including how they interact with others in their environment, in the fields of psychology, education, cognitive science, and clinical research. These recordings provide an unobtrusive way to observe children in real-world settings beyond controlled and constrained experimental environments. Advancements in speech technology and machine learning have provided an initial step for researchers to automatically and systematically analyze large-scale naturalistic recordings of children. Despite the imperfect accuracy of machine learning models, these tools still offer valuable opportunities to uncover important insights into children's cognitive and social development. Several critical speech technologies involved include speaker diarization, vocalization classification, word count estimate from adults, speaker verification, and language diarization for code-switching. Most of these technologies have been primarily developed for adults, and speech technologies applied to children specifically are still vastly under-explored. To fill this gap, we discuss current progress, challenges, and opportunities in advancing these technologies to analyze naturalistic recordings of children during early development (<3 years of age). We strive to inspire the signal processing community and foster interdisciplinary collaborations to further develop this emerging technology and address its unique challenges and opportunities.


eess.AS音频处理


【1】Audio-Based Pedestrian Detection in the Presence of Vehicular Noise
标题:存在车辆噪音时基于音频的行人检测
链接:https://arxiv.org/abs/2509.19295

作者:Yonghyun Kim, Chaeyeon Han, Akash Sarode, Noah Posner, Subhrajit Guhathakurta, Alexander Lerch
备注:Accepted to the 10th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE), 2025
摘要:基于音频的行人检测是一项具有挑战性的任务,迄今为止,仅在噪声受限的环境中进行了探索。我们提出了一个新的数据集,结果,并详细分析了最先进的基于音频的行人检测中存在的车辆噪声。在我们的研究中,我们进行了三项分析:(i)噪声和噪声限制环境之间的跨数据集评估,(ii)评估噪声数据对模型性能的影响,突出声学背景的影响,以及(iii)评估模型对域外声音的预测鲁棒性。新数据集是一个全面的1321小时路边数据集。它结合了丰富的交通音景。每个记录包括与帧级行人注释和1fps视频缩略图同步的16kHz音频。
摘要:Audio-based pedestrian detection is a challenging task and has, thus far, only been explored in noise-limited environments. We present a new dataset, results, and a detailed analysis of the state-of-the-art in audio-based pedestrian detection in the presence of vehicular noise. In our study, we conduct three analyses: (i) cross-dataset evaluation between noisy and noise-limited environments, (ii) an assessment of the impact of noisy data on model performance, highlighting the influence of acoustic context, and (iii) an evaluation of the model's predictive robustness on out-of-domain sounds. The new dataset is a comprehensive 1321-hour roadside dataset. It incorporates traffic-rich soundscapes. Each recording includes 16kHz audio synchronized with frame-level pedestrian annotations and 1fps video thumbnails.


【2】MUSHRA-1S: A scalable and sensitive test approach for evaluating top-tier speech processing systems
标题:MUHRA-1 S:用于评估顶级语音处理系统的可扩展且敏感的测试方法
链接:https://arxiv.org/abs/2509.19219

作者:Laura Lechler, Ivana Balic
备注:Submitted to ICASSP 2026; under review
摘要:评估国家的最先进的语音系统需要可扩展的和敏感的评估方法来检测微妙的,但不可接受的文物。标准MUSHRA是敏感的,但缺乏可扩展性,而ACR扩展良好,但失去了敏感性和饱和在一个高质量。为了解决这个问题,我们引入了MUSHRA 1S,这是一种单刺激变体,可以针对固定的锚点和参考一次对一个系统进行评级。在我们的实验中,MUSHRA 1S比ACR更接近标准MUSHRA,包括在ACR饱和的高质量状态下。MUSHRA 1S还能有效识别特定偏差,并通过固定上下文来减少范围均衡偏差。总的来说,MUSHRA 1S结合了MUSHRA级别的灵敏度和类似ACR的可扩展性,使其成为顶级语音处理系统基准测试的强大和可扩展的解决方案。
摘要:Evaluating state-of-the-art speech systems necessitates scalable and sensitive evaluation methods to detect subtle but unacceptable artifacts. Standard MUSHRA is sensitive but lacks scalability, while ACR scales well but loses sensitivity and saturates at a high quality. To address this, we introduce MUSHRA 1S, a single-stimulus variant that rates one system at a time against a fixed anchor and reference. Across our experiments, MUSHRA 1S matches standard MUSHRA more closely than ACR, including in the high-quality regime, where ACR saturates. MUSHRA 1S also effectively identifies specific deviations and reduces range-equalizing biases by fixing context. Overall, MUSHRA 1S combines MUSHRA level sensitivity with ACR like scalability, making it a robust and scalable solution for benchmarking top-tier speech processing systems.


【3】Improving Test-Time Performance of RVQ-based Neural Codecs
标题:提高基于RVQ的神经编解码器的测试时间性能
链接:https://arxiv.org/abs/2509.19186

作者:Hyeongju Kim, Junhyeok Lee, Jacob Morton, Juheon Lee, Jinhyeok Yang
备注:5 pages, preprint
摘要:残差矢量量化(RVQ)技术在神经音频编解码器的最新进展中起着核心作用。这些模型有效地合成高保真音频从有限数量的代码,由于量化级别之间的层次结构。在本文中,我们提出了一种编码算法,以进一步提高合成质量的RVQ为基础的神经编解码器在测试时。首先,我们指出的次优性质的量化矢量产生的传统方法。我们证明,量化误差可以通过选择一组不同的代码来减轻。随后,我们提出了我们的编码算法,旨在确定一组离散代码,实现较低的量化误差。然后,我们将所提出的方法应用于预先训练的模型,并使用不同的指标评估其有效性。实验结果表明,该方法不仅减少了量化误差,而且提高了合成质量。
摘要:The residual vector quantization (RVQ) technique plays a central role in recent advances in neural audio codecs. These models effectively synthesize high-fidelity audio from a limited number of codes due to the hierarchical structure among quantization levels. In this paper, we propose an encoding algorithm to further enhance the synthesis quality of RVQ-based neural codecs at test-time. Firstly, we point out the suboptimal nature of quantized vectors generated by conventional methods. We demonstrate that quantization error can be mitigated by selecting a different set of codes. Subsequently, we present our encoding algorithm, designed to identify a set of discrete codes that achieve a lower quantization error. We then apply the proposed method to pre-trained models and evaluate its efficacy using diverse metrics. Our experimental findings validate that our method not only reduces quantization errors, but also improves synthesis quality.


【4】On-device Internet of Sounds Sonification with Wavetable Synthesis Techniques for Soil Moisture Monitoring in Water Scarcity Contexts
标题:设备上的声音互联网超声处理和波表合成技术用于缺水环境中的土壤湿度监测
链接:https://arxiv.org/abs/2509.19097

作者:Stephen Roddy
备注:8 pages, 5 figures, 6 equations, 5th International Symposium on the Internet of Sounds
摘要:声音化,将数据映射到声音以传达有关原始数据源的信息,正在成为声音表示和传达来自声音互联网(IoS)网络之间交换的复杂数据流的信息的可行策略。本文提出了一种在全球日益严重的水资源短缺的更广泛背景下监测土壤水分水平的IoS声化实现。虽然以前的工作主要集中在物联网网络基础设施的应用程序和服务级别上的声化操作,本文探讨了使用波表合成技术将传感器数据映射到声学参数的设备级声化。设备上的波表声化的方法是形式化的,并提出了一个原型实现和探索的方法是上下文化的土壤水分监测任务之前。
摘要:Sonification, the mapping of data to sound to communicate information about the original data source, is becoming a viable strategy for the sonic representation and communication of information derived from the complex flows of data exchanged across Internet of Sounds (IoS) networks. This paper presents an IoS sonification implementation for monitoring soil moisture levels within the broader context of the globally increasing water scarcity. While previous work has focused on sonifications operating on the applications and services level of the IoS network infrastructure, this paper explores device-level sonification using wavetable synthesis techniques to map sensor data to acoustic parameters. An approach to on-device wavetable sonification is formalized, and a prototype implementation is presented and explored before the approach is contextualised with regard to the soil moisture monitoring tasks.


【5】Training Flow Matching Models with Reliable Labels via Self-Purification
标题:通过自我净化训练具有可靠标签的流匹配模型
链接:https://arxiv.org/abs/2509.19091

作者:Hyeongju Kim, Yechan Yu, June Young Yi, Juheon Lee
备注:5 pages, 3 figures, preprint
摘要:训练数据集本质上是不完美的,通常包含由于人为注释错误、标记模型的限制和其他噪声源而导致的错误标记样本。这种标签污染会显著降低训练模型的性能。在这项工作中,我们介绍了自净化流匹配(SPFM),过滤不可靠的数据流匹配框架内的原则性方法。SPFM在训练过程中使用模型本身识别可疑数据,绕过了对预训练模型或其他模块的需求。我们的实验表明,使用SPFM训练的模型生成的样本准确地遵守指定的条件,即使在嘈杂的标签上训练。此外,我们验证了SPFM在TITW数据集上的鲁棒性,该数据集由野外语音数据组成,实现了超越现有基线的性能。
摘要:Training datasets are inherently imperfect, often containing mislabeled samples due to human annotation errors, limitations of tagging models, and other sources of noise. Such label contamination can significantly degrade the performance of a trained model. In this work, we introduce Self-Purifying Flow Matching (SPFM), a principled approach to filtering unreliable data within the flow-matching framework. SPFM identifies suspicious data using the model itself during the training process, bypassing the need for pretrained models or additional modules. Our experiments demonstrate that models trained with SPFM generate samples that accurately adhere to the specified conditioning, even when trained on noisy labels. Furthermore, we validate the robustness of SPFM on the TITW dataset, which consists of in-the-wild speech data, achieving performance that surpasses existing baselines.


【6】Enhancing Noise Robustness for Neural Speech Codecs through Resource-Efficient Progressive Quantization Perturbation Simulation
标题:通过资源高效的渐进量化微扰模拟增强神经语音编解码器的噪音鲁棒性
链接:https://arxiv.org/abs/2509.19025

作者:Rui-Chen Zheng, Yang Ai, Hui-Peng Du, Zhen-Hua Ling
摘要:噪声鲁棒性仍然是在现实世界的声学场景中部署神经语音编解码器的关键挑战,其中背景噪声通常是不可避免的。我们的一个关键观察是,即使是轻微的输入噪声扰动也会导致量化码字的意外移位,从而降低重建语音的质量。出于这一发现,我们提出了一种新的和资源有效的训练策略,以提高语音编解码器的噪声鲁棒性,直接在量化级模拟这种扰动。我们的方法引入了两个核心机制:(1)距离加权概率前K采样策略,取代了传统的确定性最近邻选择残差矢量量化(RVQ);(2)渐进式训练方案,以受控方式从最后一个量化器到第一个量化器引入扰动。至关重要的是,我们的方法只在干净的语音上训练,不需要任何成对的噪声-干净数据。两个先进的神经语音编解码器,Encodec和WavTokenizer的实验表明,该策略大大提高了噪声条件下的鲁棒性-例如,提高UTMOS从3.475到3.586在15 dB SNR的Encodec-同时也提高了编码质量为干净的语音。
摘要:Noise robustness remains a critical challenge for deploying neural speech codecs in real-world acoustic scenarios where background noise is often inevitable. A key observation we make is that even slight input noise perturbations can cause unintended shifts in quantized codewords, thereby degrading the quality of reconstructed speech. Motivated by this finding, we propose a novel and resource-efficient training strategy to enhance the noise robustness of speech codecs by simulating such perturbations directly at the quantization level. Our approach introduces two core mechanisms: (1) a distance-weighted probabilistic top-K sampling strategy that replaces the conventional deterministic nearest-neighbor selection in residual vector quantization (RVQ); and (2) a progressive training scheme that introduces perturbations from the last to the first quantizer in a controlled manner. Crucially, our method is trained exclusively on clean speech, eliminating the need for any paired noisy-clean data. Experiments on two advanced neural speech codecs, Encodec and WavTokenizer, demonstrate that the proposed strategy substantially improves robustness under noisy conditions-for example, boosting UTMOS from 3.475 to 3.586 at 15 dB SNR on Encodec-while also enhancing coding quality for clean speech.


【7】HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS
标题:HD-PPT:基于指令的TTC的内容和预算偏好令牌的分层解码
链接:https://arxiv.org/abs/2509.19001

作者:Sihang Nie, Xiaofen Xing, Jingyuan Xing, Baiji Liu, Xiangmin Xu
备注:5 pages, 2 figures, submitted to ICASSP2026
摘要:基于大语言模型(LLM)的文语转换(TTS)模型已经达到了很高的自然度。然而,TTS推理的精度控制仍然是一个挑战。虽然基于解释的文本到语音转换(Instruct-TTS)模型被提出,但是由于单级文本指令和多级语音令牌之间的模态间隙,这些模型仍然缺乏细粒度的控制。为了解决这一限制,我们提出了HD-PPT,一个框架,将语音合成转换成一个结构化的,层次化的任务。为了实现细粒度控制,我们引入了一种新的语音编解码器,从复杂的语音令牌中提取不同的语音偏好和内容偏好令牌,由自动语音识别(ASR)和跨语言音频文本预训练(CLAP)目标监督。为了弥合这些令牌的模态差距,我们提出了一种分层解码策略,其中LLM以结构化的顺序生成令牌:首先是语义,然后是细粒度风格,最后是完整的声学表示。大量的实验表明,这种分层模式显着提高了指令的遵守,并实现了最先进的自然,验证了我们的方法,精确和可控的语音合成。音频样本可在https://xxh333.github.io/上获得。
摘要:Large Language Model (LLM)-based Text-to-Speech (TTS) models have already reached a high degree of naturalness. However, the precision control of TTS inference is still challenging. Although instruction-based Text-to-Speech (Instruct-TTS) models are proposed, these models still lack fine-grained control due to the modality gap between single-level text instructions and multilevel speech tokens. To address this limitation, we propose HD-PPT, a framework that transforms speech synthesis into a structured, hierarchical task. To enable fine-grained control, we introduce a novel speech codec to extract distinct prompt-preference and content-preference tokens from the complex speech tokens, supervised by automatic speech recognition (ASR) and cross-lingual audio-text pre-training (CLAP) objectives. To bridge the modality gap of these tokens, we propose a hierarchical decoding strategy, where the LLM generates tokens in a structured order: first semantic, then fine-grained style, and finally complete acoustic representation. Extensive experiments demonstrate that this hierarchical paradigm significantly improves instruction adherence and achieves state-of-the-art naturalness, validating our approach for precise and controllable speech synthesis. Audio samples are available at https://xxh333.github.io/.


【8】Direct Preference Optimization for Speech Autoregressive Diffusion Models
标题:语音自回归扩散模型的直接偏好优化
链接:https://arxiv.org/abs/2509.18928

作者:Zhijun Liu, Dongya Jia, Xiaoqiang Wang, Chenpeng Du, Shuai Wang, Zhuo Chen, Haizhou Li
摘要:自回归扩散模型(ARDM)最近已被应用到语音生成,实现国家的最先进的(SOTA)性能在zero-shot文本到语音。通过自回归生成具有下一个令牌扩散的连续语音令牌,这些模型为下一个令牌预测提供了一种有前途的替代方案,避免了与离散语音令牌化相关的技术复杂性。作为一个相对较新的范例,基于强化学习(RL)的语音ARDM微调的研究仍然有限。在本文中,我们提出了自回归扩散直接偏好优化(ARDM-DPO),以推进这一研究。通过微调最近提出的zero-shot文本到语音模型DiTAR与DPO,我们实现了显着的改善,在语音表达能力和鲁棒性方面的长文本。
摘要:Autoregressive diffusion models (ARDMs) have recently been applied to speech generation, achieving state-of-the-art (SOTA) performance in zero-shot text-to-speech. By autoregressively generating continuous speech tokens with next-token diffusion, these models offer a promising alternative to next-token prediction, avoiding the technical complexities associated with discrete speech tokenization. As a relatively new paradigm, research on reinforcement learning (RL)-based fine-tuning of speech ARDMs remains limited. In this paper, we propose Autoregressive Diffusion-Direct Preference Optimization (ARDM-DPO) to advance this research. By fine-tuning the recently proposed zero-shot text-to-speech model DiTAR with DPO, we achieve significant improvements in terms of speech expressiveness and robustness for long texts.


【9】Generalizability of Predictive and Generative Speech Enhancement Models to Pathological Speakers
标题:预测性和生成性语音增强模型对病理说话者的可推广性
链接:https://arxiv.org/abs/2509.18890

作者:Mingchi Hou, Ante Jukic, Ina Kodrasi
摘要:最先进的语音增强(SE)模型在神经典型语音上实现了很强的性能,但它们的有效性在病理语音上大大降低。在本文中,我们研究了解决预测和生成SE模型这一差距的策略,包括i)使用病理数据从头开始训练模型,ii)使用来自病理说话者的额外数据对神经典型语音进行预训练的模型进行微调,iii)仅使用来自单个病理测试说话者的数据进行说话者特定个性化。我们的研究结果表明,尽管病理语音数据集的大小有限,但SE模型可以在这些数据上成功训练或微调。使用来自几个病态说话者的数据对模型进行微调可以产生最大的性能改进,而特定于说话者的个性化效果较差,这可能是由于每个说话者可用的数据量很小。这些研究结果突出的挑战和潜在的策略,以提高SE性能的病理扬声器。
摘要:State of the art speech enhancement (SE) models achieve strong performance on neurotypical speech, but their effectiveness is substantially reduced for pathological speech. In this paper, we investigate strategies to address this gap for both predictive and generative SE models, including i) training models from scratch using pathological data, ii) finetuning models pretrained on neurotypical speech with additional data from pathological speakers, and iii) speaker specific personalization using only data from the individual pathological test speaker. Our results show that, despite the limited size of pathological speech datasets, SE models can be successfully trained or finetuned on such data. Finetuning models with data from several pathological speakers yields the largest performance improvements, while speaker specific personalization is less effective, likely due to the small amount of data available per speaker. These findings highlight the challenges and potential strategies for improving SE performance for pathological speakers.


【10】Influence of Clean Speech Characteristics on Speech Enhancement Performance
标题:干净语音特征对语音增强性能的影响
链接:https://arxiv.org/abs/2509.18885

作者:Mingchi Hou, Ina Kodrasi
摘要:语音增强(SE)的性能取决于噪声特性和信噪比(SNR),但干净的语音信号本身的内在属性仍然是一个未充分探索的因素。在这项工作中,我们系统地分析了干净的语音特征如何影响增强的难度在多个国家的最先进的SE模型,语言和噪声条件。我们提取一组音调,共振峰,响度和频谱通量的功能,从干净的语音和计算相关性与客观SE指标,包括频率加权分段SNR和PESQ。我们的研究结果表明,共振峰振幅是一致的预测SE性能,更高,更稳定的共振峰导致更大的增强增益。我们进一步证明,性能变化很大,甚至在一个单一的扬声器的话语,突出的重要性,内扬声器声学变化。这些发现为SE挑战提供了新的见解,表明在设计数据集,评估协议和增强模型时应考虑固有的语音特征。
摘要:Speech enhancement (SE) performance is known to depend on noise characteristics and signal to noise ratio (SNR), yet intrinsic properties of the clean speech signal itself remain an underexplored factor. In this work, we systematically analyze how clean speech characteristics influence enhancement difficulty across multiple state of the art SE models, languages, and noise conditions. We extract a set of pitch, formant, loudness, and spectral flux features from clean speech and compute correlations with objective SE metrics, including frequency weighted segmental SNR and PESQ. Our results show that formant amplitudes are consistently predictive of SE performance, with higher and more stable formants leading to larger enhancement gains. We further demonstrate that performance varies substantially even within a single speaker's utterances, highlighting the importance of intraspeaker acoustic variability. These findings provide new insights into SE challenges, suggesting that intrinsic speech characteristics should be considered when designing datasets, evaluation protocols, and enhancement models.


【11】Towards Evaluating Generative Audio: Insights from Neural Audio Codec Embedding Distances
标题:评估生成音频:来自神经音频编解码器嵌入距离的见解
链接:https://arxiv.org/abs/2509.18823

作者:Arijit Biswas, Lars Villemoes
备注:Pre-review version submitted to ICASSP 2026
摘要:神经音频编解码器(NAC)通过学习紧凑的音频表示来实现低比特率压缩,这也可以作为感知质量评估的特征。我们介绍了DACe,这是描述音频编解码器(DAC)的增强型高保真版本,采用平衡采样对各种真实和合成音调数据进行训练。我们系统地比较了Frechet音频距离(FAD)和最大平均离散度(MMD)在MUSHRA测试中的语音,音乐和混合内容。FAD始终优于MMD,并且来自高保真度NAC(如DACe)的嵌入显示出与人类判断更强的相关性。CLAP LAION Music(CLAP-M)和OpenL 3 Mel 128(OpenL 3 - 128 M)嵌入实现了更高的相关性,而NAC嵌入提供了一种实用的zero-shot音频质量评估方法,只需要未编码的音频进行训练。这些结果证明了NAC用于压缩和感知上知情的音频评估的双重效用。
摘要:Neural audio codecs (NACs) achieve low-bitrate compression by learning compact audio representations, which can also serve as features for perceptual quality evaluation. We introduce DACe, an enhanced, higher-fidelity version of the Descript Audio Codec (DAC), trained on diverse real and synthetic tonal data with balanced sampling. We systematically compare Fr\'echet Audio Distance (FAD) and Maximum Mean Discrepancy (MMD) on MUSHRA tests across speech, music, and mixed content. FAD consistently outperforms MMD, and embeddings from higher-fidelity NACs (such as DACe) show stronger correlations with human judgments. While CLAP LAION Music (CLAP-M) and OpenL3 Mel128 (OpenL3-128M) embeddings achieve higher correlations, NAC embeddings provide a practical zero-shot approach to audio quality assessment, requiring only unencoded audio for training. These results demonstrate the dual utility of NACs for compression and perceptually informed audio evaluation.


【12】Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocoders
标题:重新思考时频域神经声码器幅度和阶段的联合估计
链接:https://arxiv.org/abs/2509.18806

作者:Lingling Dai, Andong Li, Tong Lei, Meng Yu, Xiaodong Li, Chengshi Zheng
备注:Submitted to ICASSP 2026
摘要:基于时频域的神经元声码器在合成高保真音频方面已经显示出很好的效果。然而,目前还不清楚有效地预测震级和相位目标的机制。在本文中,我们从两个代表性的T-F域声码器,即Vocos和APNet 2,这属于单流和双流模式的幅度和相位估计,分别。当在大规模数据集上评估它们的性能时,我们意外地观察到APNet 2的严重性能崩溃。为了稳定其性能,在本文中,我们介绍了三个简单而有效的策略,每个策略分别针对拓扑空间,源空间和输出空间。具体来说,我们修改的建筑拓扑结构更好地在拓扑空间中的信息交换,引入先验知识,以促进在源空间中的生成过程,并优化反向传播过程中的参数更新与改进的输出格式在输出空间。实验结果表明,该方法有效地实现了APNet 2中幅度和相位的联合估计,从而弥补了单码流声码器和双流声码器的性能差距。
摘要:Time-frequency (T-F) domain-based neural vocoders have shown promising results in synthesizing high-fidelity audio. Nevertheless, it remains unclear on the mechanism of effectively predicting magnitude and phase targets jointly. In this paper, we start from two representative T-F domain vocoders, namely Vocos and APNet2, which belong to the single-stream and dual-stream modes for magnitude and phase estimation, respectively. When evaluating their performance on a large-scale dataset, we accidentally observe severe performance collapse of APNet2. To stabilize its performance, in this paper, we introduce three simple yet effective strategies, each targeting the topological space, the source space, and the output space, respectively. Specifically, we modify the architectural topology for better information exchange in the topological space, introduce prior knowledge to facilitate the generation process in the source space, and optimize the backpropagation process for parameter updates with an improved output format in the output space. Experimental results demonstrate that our proposed method effectively facilitates the joint estimation of magnitude and phase in APNet2, thus bridging the performance disparities between the single-stream and dual-stream vocoders.


【13】Group Relative Policy Optimization for Text-to-Speech with Large Language Models
标题:具有大型语言模型的文本到语音的群体相对策略优化
链接:https://arxiv.org/abs/2509.18798

作者:Chang Liu, Ya-Jun Hu, Ying-Ying Gao, Shi-Lei Zhang, Zhen-Hua Ling
备注:5 pages,submitted to ICASSP2026
摘要:本文提出了一种基于GRPO的方法,以提高基于大语言模型(LLM)的文本到语音(TTS)模型的性能,从现成的自动语音识别(ASR)模型获得奖励。与以前基于LLM的TTS的强化学习方法相比,我们的方法不需要专门的奖励计算或训练模型。此外,我们设计了一个复合奖励函数,结合字符错误率(CER)和负对数似然(NLL)从ASR模型中获得,提供更多的信息和准确的奖励信号。我们将GRPO微调应用于预训练的基于LLM的TTS模型,并评估其zero-shot TTS性能。实验结果表明,该方法在提高合成语音的可懂度和自然度方面都有较大的改善。消融研究和进一步的分析证实了整合两个奖励成分的有效性。
摘要:This paper proposes a GRPO-based approach to enhance the performance of large language model (LLM)-based text-to-speech (TTS) models by deriving rewards from an off-the-shelf automatic speech recognition (ASR) model. Compared to previous reinforcement learning methods for LLM-based TTS, our method requires no dedicated model for reward computation or training. Moreover, we design a composite reward function that combines character error rate (CER) with negative log-likelihood (NLL) obtained from the ASR model, providing more informative and accurate reward signals. We apply GRPO fine-tuning to pre-trained LLM-based TTS models and evaluate their zero-shot TTS performance. Experimental results show that the proposed method substantially improves both the intelligibility and naturalness of synthesized speech. Ablation studies and further analyses confirm the effectiveness of integrating the two reward components.


【14】FlexSED: Towards Open-Vocabulary Sound Event Detection
标题:FlexMED:迈向开放词汇声音事件检测
链接:https://arxiv.org/abs/2509.18606

作者:Jiarui Hai, Helin Wang, Weizhe Guo, Mounya Elhilali
摘要:尽管最近在能够处理数百个声音类别的大规模声音事件检测(SED)系统中取得了进展,但现有的多类别分类框架仍然从根本上受到限制。它们不能处理自由文本声音查询,而自由文本声音查询可以实现更灵活和用户友好的交互,并且它们缺乏zero-shot功能,并且提供较差的Few-Shot适应性。虽然基于文本查询的分离方法已被探索,他们主要集中在源分离和不适合的SED任务,需要精确的时间定位和高效的检测,在大型和多样化的声音词汇。在本文中,我们提出了FlexSED,一个开放词汇的声音事件检测系统。FlexSED建立在预训练的音频SSL模型和CLAP文本编码器的基础上,引入了编码器-解码器组合和自适应融合策略,以实现从预训练权重进行有效的连续训练。为了确保强大的监督,它还采用了大型语言模型(LLM)来帮助在训练过程中选择事件查询,解决与缺失标签相关的挑战。因此,与AudioSet-Strong上的普通SED模型相比,FlexSED实现了卓越的性能,同时展示了强大的zero-shot和Few-Shot功能。我们发布代码和预训练模型,以支持未来基于FlexSED的研究和应用。
摘要:Despite recent progress in large-scale sound event detection (SED) systems capable of handling hundreds of sound classes, existing multi-class classification frameworks remain fundamentally limited. They cannot process free-text sound queries, which enable more flexible and user-friendly interaction, and they lack zero-shot capabilities and offer poor few-shot adaptability. Although text-query-based separation methods have been explored, they primarily focus on source separation and are ill-suited for SED tasks that require precise temporal localization and efficient detection across large and diverse sound vocabularies. In this paper, we propose FlexSED, an open-vocabulary sound event detection system. FlexSED builds on a pretrained audio SSL model and the CLAP text encoder, introducing an encoder-decoder composition and an adaptive fusion strategy to enable effective continuous training from pretrained weights. To ensure robust supervision, it also employs large language models (LLMs) to assist in event query selection during training, addressing challenges related to missing labels. As a result, FlexSED achieves superior performance compared to vanilla SED models on AudioSet-Strong, while demonstrating strong zero-shot and few-shot capabilities. We release the code and pretrained models to support future research and applications based on FlexSED.


【15】SynSonic: Augmenting Sound Event Detection through Text-to-Audio Diffusion ControlNet and Effective Sample Filtering
标题:SynSonic:通过文本到音频扩散控制网和有效的样本过滤增强声音事件检测
链接:https://arxiv.org/abs/2509.18603

作者:Jiarui Hai, Mounya Elhilali
摘要:由于时间标记数据的稀缺性,数据合成和增强是声音事件检测(SED)的关键。虽然SpecAugment和Mix-up等增强方法可以提高模型性能,但它们仍然受到现有样本多样性的限制。最近的生成模型提供了新的机会,但它们的直接应用到SED是具有挑战性的,由于缺乏精确的时间注释和通过不可靠的过滤引入噪声的风险。为了解决这些挑战,并使生成为基础的增强SED,我们提出了SynSonic,为这项任务量身定制的数据增强方法。SynSonic利用由能量包络ControlNet引导的文本到音频扩散模型来生成时间相干的声音事件。具有双分类器的联合评分过滤策略确保了样本质量,并且我们探索了将其实际集成到训练管道中。实验结果表明,SynSonic提高了复调声音检测分数(PSDS 1和PSDS 2),增强了时间定位和声音类别区分。
摘要:Data synthesis and augmentation are essential for Sound Event Detection (SED) due to the scarcity of temporally labeled data. While augmentation methods like SpecAugment and Mix-up can enhance model performance, they remain constrained by the diversity of existing samples. Recent generative models offer new opportunities, yet their direct application to SED is challenging due to the lack of precise temporal annotations and the risk of introducing noise through unreliable filtering. To address these challenges and enable generative-based augmentation for SED, we propose SynSonic, a data augmentation method tailored for this task. SynSonic leverages text-to-audio diffusion models guided by an energy-envelope ControlNet to generate temporally coherent sound events. A joint score filtering strategy with dual classifiers ensures sample quality, and we explore its practical integration into training pipelines. Experimental results show that SynSonic improves Polyphonic Sound Detection Scores (PSDS1 and PSDS2), enhancing both temporal localization and sound class discrimination.


【16】Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation
标题:教授音频模型推理:源蒸馏和分层蒸馏的统一框架
链接:https://arxiv.org/abs/2509.18579

作者:Runyan Yang, Yuke Si, Yingying Gao, Junlan Feng, Chao Deng, Shilei Zhang
备注:5 pages; submitted to ICASSP 2026
摘要:虽然大型音频语言模型擅长ASR和情感识别等任务,但由于音频和文本之间的模态差距以及缺乏结构化的中间监督,它们仍然难以进行复杂的推理。为了解决这个问题,我们提出了一个统一的知识蒸馏框架,将推理能力从高容量的文本教师模型转移到学生音频模型,同时保留其声学能力。我们的方法引入了两个关键维度:源级蒸馏,它利用文本和声学教师提供互补的特定模态监督;分层蒸馏,它将教师信号与适当的学生层对齐,以提高传输效率。这种双维度策略能够对蒸馏过程进行细粒度控制,有效地弥合了符号推理和语音表示之间的差距。实验结果表明,音频推理性能的显着改善,证明了我们的框架作为音频建模的推理转移解决方案的有效性。
摘要:While large audio language models excel at tasks like ASR and emotion recognition, they still struggle with complex reasoning due to the modality gap between audio and text as well as the lack of structured intermediate supervision. To address this, we propose a unified knowledge distillation framework to transfer reasoning capabilities from a high-capacity textual teacher model to a student audio models while preserving its acoustic competence. Our method introduces two key dimensions: source-wise distillation, which leverages both textual and acoustic teachers to provide complementary modality-specific supervision; and layer-wise distillation, which aligns teacher signals with appropriate student layers to improve transfer efficiency. This dual-dimensional strategy enables fine-grained control over the distillation process, effectively bridging the gap between symbolic reasoning and speech representations. Experimental results show significant improvements in audio reasoning performance, demonstrating the effectiveness of our framework as a reasoning transfer solution for audio modeling.


【17】HarmoniFuse: A Component-Selective and Prompt-Adaptive Framework for Multi-Task Speech Language Modeling
标题:Harmonious:一个用于多任务语音语言建模的对象选择和对象自适应框架
链接:https://arxiv.org/abs/2509.18570

作者:Yuke Si, Runyan Yang, Yingying Gao, Junlan Feng, Chao Deng, Shilei Zhang
备注:5 pages; submitted to ICASSP 2026
摘要:大型语言模型的最新进展促进了能够在共享架构内支持多个语音任务的统一语音语言模型(SLM)的开发。然而,诸如自动语音识别(ASR)和语音情感识别(SER)等任务依赖于不同类型的信息:ASR主要依赖于语言内容,而SER需要整合语言和非语言线索。现有的多任务SLM通常采用朴素的参数共享或基于约束的条件,而没有明确地建模每个任务所需的信息组成的差异。这种设计有任务干扰和性能下降的风险,特别是在有限的数据条件下。为了解决这些局限性,我们提出了Harmoniplane,一个组件选择性和自适应框架的多任务语音语言建模。Harmonietry旨在通过选择和融合语音表示的任务相关组件来协调异构任务需求。具体来说,它集成了一个门控语音编码器,以提取特定于任务的声学特征和一个自适应动态融合模块聚合Transformer层的任务特性的基础上。此外,批量交错训练策略可以利用单独的ASR和SER数据集,而无需联合注释。实验结果表明,谐波改善ASR和SER的性能,提供了一个可扩展的和强大的解决方案,多任务语音理解现实的数据约束下。
摘要:Recent advances in large language models have facilitated the development of unified speech language models (SLMs) capable of supporting multiple speech tasks within a shared architecture. However, tasks such as automatic speech recognition (ASR) and speech emotion recognition (SER) rely on distinct types of information: ASR primarily depends on linguistic content, whereas SER requires the integration of both linguistic and paralinguistic cues. Existing multitask SLMs typically adopt naive parameter sharing or prompt-based conditioning without explicitly modeling the differences in information composition required by each task. Such designs risk task interference and performance degradation, especially under limited data conditions. To address these limitations, we propose HarmoniFuse, a component-selective and prompt-adaptive framework for multi-task speech language modeling. HarmoniFuse is designed to harmonize heterogeneous task demands by selecting and fusing task-relevant components of speech representations. Specifically, it integrates a gated speech encoder to extract task-specific acoustic features and a prompt-adaptive dynamic fusion module to aggregate transformer layers based on task characteristics. In addition, a batch-interleaved training strategy enables leveraging separate ASR and SER datasets without requiring joint annotation. Experimental results demonstrate that HarmoniFuse improves both ASR and SER performance, offering a scalable and robust solution for multitask speech understanding under realistic data constraints.


【18】SoundCompass: Navigating Target Sound Extraction With Effective Directional Clue Integration In Complex Acoustic Scenes
标题:SoundCompass:在复杂声学场景中通过有效的方向线索集成进行导航目标声音提取
链接:https://arxiv.org/abs/2509.18561

作者:Dayun Choi, Jung-Woo Choi
备注:5 pages, 4 figures, submitted to ICASSP 2026
摘要:目标声音提取(TSE)的最新进展利用来自到达方向(DoA)的方向线索,其表示在任何声学场景中可用的声音的固有空间属性。然而,以前的基于DoA的方法依赖于手工制作的特征或离散编码,这丢失了细粒度的空间信息并限制了适应性。我们提出了SoundCompass,一个有效的方向线索集成框架为中心的频谱成对相互作用(SPIN)模块,捕捉跨通道的空间相关性在复杂的频谱域,以保留完整的空间信息,在多通道信号。表示在空间相关性方面的输入特征与表示为球面谐波(SH)编码的DoA线索融合。融合是在重叠的频率子带上进行的,继承了以前的带分离架构中报告的优点。我们还将迭代细化策略,推理链(CoI),在TSE框架中,递归地融合DoA与从先前的推理阶段估计的声音事件激活。实验表明,SoundCompass结合SPIN,SH嵌入和CoI,可以在不同的信号类别和空间配置中鲁棒地提取目标源。
摘要:Recent advances in target sound extraction (TSE) utilize directional clues derived from direction of arrival (DoA), which represent an inherent spatial property of sound available in any acoustic scene. However, previous DoA-based methods rely on hand-crafted features or discrete encodings, which lose fine-grained spatial information and limit adaptability. We propose SoundCompass, an effective directional clue integration framework centered on a Spectral Pairwise INteraction (SPIN) module that captures cross-channel spatial correlations in the complex spectrogram domain to preserve full spatial information in multichannel signals. The input feature expressed in terms of spatial correlations is fused with a DoA clue represented as spherical harmonics (SH) encoding. The fusion is carried out across overlapping frequency subbands, inheriting the benefits reported in the previous band-split architectures. We also incorporate the iterative refinement strategy, chain-of-inference (CoI), in the TSE framework, which recursively fuses DoA with sound event activation estimated from the previous inference stage. Experiments demonstrate that SoundCompass, combining SPIN, SH embedding, and CoI, robustly extracts target sources across diverse signal classes and spatial configurations.


【19】No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS
标题:韵律无可验证奖励:迈向DTS中偏好引导的韵律学习
链接:https://arxiv.org/abs/2509.18531

作者:Seungyoun Shin, Dongha Ahn, Jiwoo Kim, Sungwook Jeon
备注:submitted to ICASSP 2026
摘要:最近的工作报告在神经文本到语音(TTS)与组相对策略优化(GRPO)的收益。然而,在缺乏可验证的奖励\textit{韵律}的情况下,在转录导向信号(CER/NLL)上训练的GRPO降低了错误率,但将韵律压缩成单调的、不自然的语音;增加说话者相似性进一步破坏了训练的稳定性,降低了CER。我们用一个迭代直接偏好优化(DPO)方案来解决这个问题,该方案每轮只使用几百个人类标记的偏好对来直接优化韵律自然度,同时正则化到当前模型。在\textbf{KoCC-TTS}上,这是一个真实的韩国呼叫中心交互的策划数据集,捕获了面向任务的对话,我们的方法获得了最高的人类偏好(ELO),具有竞争力的CER,优于GRPO和强大的商业基线。这些结果表明,当韵律不能自动奖励,\textit{人类偏好优化}提供了一个实用的和数据有效的路径自然和强大的TTS。演示页面位于\href{https:tts.ch.dev}
摘要:Recent work reports gains in neural text-to-speech (TTS) with Group Relative Policy Optimization (GRPO). However, in the absence of a verifiable reward for \textit{prosody}, GRPO trained on transcription-oriented signals (CER/NLL) lowers error rates yet collapses prosody into monotone, unnatural speech; adding speaker-similarity further destabilizes training and degrades CER. We address this with an \textit{iterative Direct Preference Optimization (DPO)} scheme that uses only a few hundred human-labeled preference pairs per round to directly optimize prosodic naturalness while regularizing to the current model. On \textbf{KoCC-TTS}, a curated dataset of authentic Korean call center interactions capturing task-oriented dialogues, our method attains the highest human preference (ELO) with competitive CER, outperforming GRPO and strong commercial baselines. These results suggest that when prosody cannot be rewarded automatically, \textit{human preference optimization} offers a practical and data-efficient path to natural and robust TTS. The demo page is available at \href{https://tts.ch.dev}


【20】Automated Analysis of Naturalistic Recordings in Early Childhood: Applications, Challenges, and Opportunities
标题:幼儿自然记录的自动分析:应用、挑战和机遇
链接:https://arxiv.org/abs/2509.18235

作者:Jialu Li Member, Marvin Lavechin, Xulin Fan, Nancy L. McElwain, Alejandrina Cristia, Paola Garcia-Perera, Mark Hasegawa-Johnson
备注:Accepted to IEEE Signal Processing Magazine
摘要:自然主义录音捕捉真实世界环境中的音频,参与者在没有研究人员或实验协议干扰的情况下自然地表现出来。自然主义的长形式录音通过捕捉参与者日常生活中长时间的自发和持续的互动来扩展这一概念,通常跨越数小时甚至数天。自然主义录音已被广泛用于研究儿童的行为,包括他们如何与他人在他们的环境中互动,在心理学,教育,认知科学和临床研究领域。这些录音提供了一种不引人注目的方式来观察儿童在现实世界中的设置超出控制和约束的实验环境。语音技术和机器学习的进步为研究人员自动和系统地分析儿童的大规模自然录音提供了第一步。尽管机器学习模型的准确性并不完美,但这些工具仍然为揭示儿童认知和社会发展的重要见解提供了宝贵的机会。几个关键的语音技术包括说话人日记,发声分类,从成人的字数估计,说话人验证,和语码转换的语言日记。这些技术中的大多数主要是为成年人开发的,而专门应用于儿童的语音技术仍处于探索阶段。为了填补这一空白,我们讨论了推进这些技术以分析儿童早期发育(<3岁)期间的自然记录的当前进展、挑战和机遇。我们致力于激励信号处理社区,促进跨学科合作,以进一步发展这一新兴技术,并应对其独特的挑战和机遇。
摘要:Naturalistic recordings capture audio in real-world environments where participants behave naturally without interference from researchers or experimental protocols. Naturalistic long-form recordings extend this concept by capturing spontaneous and continuous interactions over extended periods, often spanning hours or even days, in participants' daily lives. Naturalistic recordings have been extensively used to study children's behaviors, including how they interact with others in their environment, in the fields of psychology, education, cognitive science, and clinical research. These recordings provide an unobtrusive way to observe children in real-world settings beyond controlled and constrained experimental environments. Advancements in speech technology and machine learning have provided an initial step for researchers to automatically and systematically analyze large-scale naturalistic recordings of children. Despite the imperfect accuracy of machine learning models, these tools still offer valuable opportunities to uncover important insights into children's cognitive and social development. Several critical speech technologies involved include speaker diarization, vocalization classification, word count estimate from adults, speaker verification, and language diarization for code-switching. Most of these technologies have been primarily developed for adults, and speech technologies applied to children specifically are still vastly under-explored. To fill this gap, we discuss current progress, challenges, and opportunities in advancing these technologies to analyze naturalistic recordings of children during early development (<3 years of age). We strive to inspire the signal processing community and foster interdisciplinary collaborations to further develop this emerging technology and address its unique challenges and opportunities.


【21】Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models
标题:更加关注音频:缓解大型音频语言模型中跨模式注意力的不平衡
链接:https://arxiv.org/abs/2509.18816

作者:Junyu Wang, Ziyang Ma, Zhengding Luo, Tianrui Wang, Meng Ge, Xiaobao Wang, Longbiao Wang
备注:Submitted to ICASSP 2026
摘要:大型音频语言模型(LALM)经常遭受音频-文本注意力不平衡,优先考虑文本而不是声学信息,特别是在Transformer架构的多模态融合层中。这种偏见阻碍了他们充分利用声学线索的能力,导致音频推理任务的表现不佳。为了缓解这一问题,我们提出了一种新的无训练方法\textbf{MATA},该方法动态地推动LALM在自注意力机制中支付\textbf{M}或\textbf {A}注意力\textbf{T}或\textbf{A}音频令牌。具体而言,MATA干预后原始注意力评分,仅针对中间层中的最后一个令牌,而不引入额外的参数或计算开销。在MMAU和MMAR基准测试上的实验证实了MATA的有效性,并具有一致的性能增益。值得注意的是,在MMAR上,MATA使开源模型首次超越了专有的Gemini 2.0 Flash。我们的工作提供了一个有效的解决方案,以减轻注意偏差,并打开了一个新的研究方向,提高音频处理能力的多模态模型。
摘要:Large Audio-Language Models (LALMs) often suffer from audio-textual attention imbalance, prioritizing text over acoustic information, particularly in the multi-modal fusion layers of the Transformer architecture. This bias hinders their ability to fully utilize acoustic cues, causing suboptimal performance on audio reasoning tasks. To mitigate this, we propose \textbf{MATA}, a novel training-free method that dynamically pushes LALMs to pay \textbf{M}ore \textbf{A}ttention \textbf{T}o \textbf{A}udio tokens within the self-attention mechanism. Specifically, MATA intervenes post raw attention scoring, targeting only the last token in intermediate layers without introducing additional parameters or computational overhead. Experiments on the MMAU and MMAR benchmarks confirm MATA's effectiveness, with consistent performance gains. Notably, on MMAR, MATA enables an open-source model to surpass the proprietary Gemini 2.0 Flash for the first time. Our work provides an efficient solution to mitigate attention bias and opens a new research direction for enhancing the audio-processing capabilities of multi-modal models.


【22】Enhancing Automatic Chord Recognition through LLM Chain-of-Thought Reasoning
标题:通过LLM思想链推理增强自动和弦识别
链接:https://arxiv.org/abs/2509.18700

作者:Chih-Cheng Chang, Bo-Yu Chen, Lu-Rong Chen, Li Su
摘要:音乐信息检索(MIR)包括用于分析和理解音乐内容的广泛计算技术,最近的深度学习进步推动了实质性的改进。基于这些进展,本文探讨了大型语言模型(LLM)如何作为一个集成的桥梁,连接和集成来自多个MIR工具的信息,重点是提高自动和弦识别性能。我们提出了一种新的方法,将基于文本的LLM定位为智能协调器,该协调器处理和集成来自各种最先进的MIR工具的输出,包括音乐源分离,关键检测,和弦识别和节拍跟踪。我们的方法将音频衍生的音乐信息转换为文本表示,使LLM能够专门针对和弦识别任务进行推理和校正。我们设计了一个5阶段的思想链框架,允许GPT-4 o系统地分析,比较和完善和弦识别结果,通过利用音乐理论知识来整合不同MIR组件的信息。对三个数据集的实验评估表明,多个评估指标的一致改进,MIREX指标的总体准确率提高了1-2.77%。我们的研究结果表明,LLM可以有效地作为MIR管道中的综合桥梁,为音乐信息检索任务中的多工具协调开辟了新的方向。
摘要:Music Information Retrieval (MIR) encompasses a broad range of computational techniques for analyzing and understanding musical content, with recent deep learning advances driving substantial improvements. Building upon these advances, this paper explores how large language models (LLMs) can serve as an integrative bridge to connect and integrate information from multiple MIR tools, with a focus on enhancing automatic chord recognition performance. We present a novel approach that positions text-based LLMs as intelligent coordinators that process and integrate outputs from diverse state-of-the-art MIR tools-including music source separation, key detection, chord recognition, and beat tracking. Our method converts audio-derived musical information into textual representations, enabling LLMs to perform reasoning and correction specifically for chord recognition tasks. We design a 5-stage chain-of-thought framework that allows GPT-4o to systematically analyze, compare, and refine chord recognition results by leveraging music-theoretical knowledge to integrate information across different MIR components. Experimental evaluation on three datasets demonstrates consistent improvements across multiple evaluation metrics, with overall accuracy gains of 1-2.77% on the MIREX metric. Our findings demonstrate that LLMs can effectively function as integrative bridges in MIR pipelines, opening new directions for multi-tool coordination in music information retrieval tasks.


【23】An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
标题:从掩蔽频谱图进行自我监督音频表示学习的神经架构概述
链接:https://arxiv.org/abs/2509.18691

作者:Sarthak Yadav, Sergios Theodoridis, Zheng-Hua Tan
摘要:近年来,自监督学习在没有标记数据的情况下训练深度神经表示方面引起了极大的兴趣。一种这样的自监督学习方法是掩蔽频谱图建模,其中目标是通过预测输入音频频谱图的移除或隐藏部分来学习语义丰富的上下文表示。以Transformer神经架构为核心,掩蔽声谱图建模已成为学习通用音频表示的主要方法,也称为音频基础模型。同时,解决Transformer架构的问题,特别是底层的缩放点积注意力操作,其与输入序列长度成二次缩放,导致了对递归序列建模方法的新兴趣。其中,选择性结构化状态空间模型(如Mamba)和扩展的长短期记忆(xLSTM)是两种最有前途的方法,已被广泛采用。虽然关于这两个专题的工作不断增加,但目前缺乏对这两个专题的交叉点的适当概述。在本文中,我们对上述研究领域进行了全面的概述,涵盖了掩蔽谱图建模和前面提到的神经序列建模架构Mamba和xLSTM。此外,我们比较了Transformers,Mamba和xLSTM基于掩蔽的声谱图模型在一个统一的,可重复的框架上的十个不同的下游音频分类任务,这将有助于感兴趣的读者作出明智的决定,对邻近应用程序的评估方法的适用性。
摘要:In recent years, self-supervised learning has amassed significant interest for training deep neural representations without labeled data. One such self-supervised learning approach is masked spectrogram modeling, where the objective is to learn semantically rich contextual representations by predicting removed or hidden portions of the input audio spectrogram. With the Transformer neural architecture at its core, masked spectrogram modeling has emerged as the prominent approach for learning general purpose audio representations, a.k.a. audio foundation models. Meanwhile, addressing the issues of the Transformer architecture, in particular the underlying Scaled Dot-product Attention operation, which scales quadratically with input sequence length, has led to renewed interest in recurrent sequence modeling approaches. Among them, Selective structured state space models (such as Mamba) and extended Long Short-Term Memory (xLSTM) are the two most promising approaches which have experienced widespread adoption. While the body of work on these two topics continues to grow, there is currently a lack of an adequate overview encompassing the intersection of these topics. In this paper, we present a comprehensive overview of the aforementioned research domains, covering masked spectrogram modeling and the previously mentioned neural sequence modeling architectures, Mamba and xLSTM. Further, we compare Transformers, Mamba and xLSTM based masked spectrogram models in a unified, reproducible framework on ten diverse downstream audio classification tasks, which will help interested readers to make informed decisions regarding suitability of the evaluated approaches to adjacent applications.


【24】Scalable Evaluation for Audio Identification via Synthetic Latent Fingerprint Generation
标题:基于合成潜在指纹生成的可扩展音频识别评估
链接:https://arxiv.org/abs/2509.18620

作者:Aditya Bhattacharjee, Marco Pasini, Emmanouil Benetos
备注:Under review for International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Barcelona, 2026
摘要:由于缺乏大型公共音乐数据库,现实规模的音频指纹评估受到限制。我们提出了一种无音频的方法,合成的潜在指纹近似真实指纹的分布。我们的方法在预训练的神经音频指纹系统提取的嵌入上训练整流流模型。使用我们的系统生成的合成指纹作为现实的干扰,并在不需要额外的音频的情况下,使检索性能的模拟在大规模。我们通过比较真实数据的分布来评估合成指纹的保真度。我们进一步基准的检索性能在多个国家的最先进的音频指纹识别框架,通过增加真实的参考数据库与合成干扰,并显示与合成干扰获得的缩放趋势密切跟踪与真实干扰。最后,我们规模的合成干扰数据库模型检索性能非常大的数据库,提供了一个实用的度量系统的可扩展性,不依赖于访问音频语料库。
摘要:The evaluation of audio fingerprinting at a realistic scale is limited by the scarcity of large public music databases. We present an audio-free approach that synthesises latent fingerprints which approximate the distribution of real fingerprints. Our method trains a Rectified Flow model on embeddings extracted by pre-trained neural audio fingerprinting systems. The synthetic fingerprints generated using our system act as realistic distractors and enable the simulation of retrieval performance at a large scale without requiring additional audio. We assess the fidelity of synthetic fingerprints by comparing the distributions to real data. We further benchmark the retrieval performances across multiple state-of-the-art audio fingerprinting frameworks by augmenting real reference databases with synthetic distractors, and show that the scaling trends obtained with synthetic distractors closely track those obtained with real distractors. Finally, we scale the synthetic distractor database to model retrieval performance for very large databases, providing a practical metric of system scalability that does not depend on access to audio corpora.


【25】Explore the Reinforcement Learning for the LLM based ASR and TTS system
标题:探索基于LLM的SVR和TTC系统的强化学习
链接:https://arxiv.org/abs/2509.18569

作者:Changfeng Gao, Yabin Li, Keyu An, Zhifu Gao, Zhihao Du, Han Zhao, Xiangang Li
摘要:近年来,大语言模型(LLM)在自动语音识别(ASR)和文本到语音(TTS)系统中发挥了重要作用。虽然强化学习(RL)在基于文本的任务中显着提高了LLM的性能,但由于训练基于音频的模型的复杂性,其在ASR和TTS中的应用仍然没有得到充分的探索。在这项研究中,我们提出了一个轻量级的RL框架,专为基于音频的LLM,可以处理音频输入和生成音频输出。基于这个框架,我们评估了强化学习在ASR和TTS任务上的有效性。对于ASR任务,我们在组相对策略优化(GRPO)框架内实验了不同的基于规则的奖励函数,并研究了RL数据构建的影响。对于TTS任务,我们将GRPO与微分奖励优化(DiffRO)进行比较,并进一步将这两种方法结合起来以提高性能。我们的实验表明,RL可以显着提高ASR和TTS系统的性能,即使在有限的训练数据和少量的优化步骤。
摘要:In recent years, large language models (LLMs) have played an important role in automatic speech recognition (ASR) and text-to-speech (TTS) systems. While reinforcement learning (RL) has significantly enhanced LLM performance in text-based tasks, its application to ASR and TTS remains underexplored due to the complexity of training audio-based models. In this study, we propose a lightweight RL framework tailored for audio-based LLMs that can process audio inputs and generate audio outputs. Based on this framework, we evaluate the effectiveness of reinforcement learning on both ASR and TTS tasks. For the ASR task, we experiment with different rule-based reward functions within the Group Relative Policy Optimization (GRPO) framework and investigate the impact of RL data construction. For the TTS task, we compare GRPO with Differentiable Reward Optimization (DiffRO) and further combine the two approaches to achieve improved performance. Our experiments demonstrate that RL can significantly enhance the performance of both ASR and TTS systems, even with limited training data and a small number of optimization steps.


【26】Discrete-time diffusion-like models for speech synthesis
标题:语音合成的离散时间类扩散模型
链接:https://arxiv.org/abs/2509.18470

作者:Xiaozhou Tan, Minghui Zhao, Mattias Cross, Anton Ragni
摘要:近年来,扩散模型引起了人们的广泛关注。这些模型将语音生成视为一个连续时间过程。为了有效训练,该过程通常限于加性高斯噪声,这是有限的。对于推理,时间通常是离散化的,导致连续训练和离散采样条件之间的不匹配。另一方面,最近提出的离散时间过程通常没有这些限制,可能需要更少的推理步骤,并且在训练/推理条件之间完全一致。本文探讨了一些扩散类离散时间过程,并提出了一些新的变种。这些包括应用加性高斯噪声、乘性高斯噪声、模糊噪声以及模糊和高斯噪声的混合的过程。实验结果表明,离散时间过程提供了可比的主观和客观的语音质量,他们广泛流行的连续对应,更有效和一致的训练和推理模式。
摘要:Diffusion models have attracted a lot of attention in recent years. These models view speech generation as a continuous-time process. For efficient training, this process is typically restricted to additive Gaussian noising, which is limiting. For inference, the time is typically discretized, leading to the mismatch between continuous training and discrete sampling conditions. Recently proposed discrete-time processes, on the other hand, usually do not have these limitations, may require substantially fewer inference steps, and are fully consistent between training/inference conditions. This paper explores some diffusion-like discrete-time processes and proposes some new variants. These include processes applying additive Gaussian noise, multiplicative Gaussian noise, blurring noise and a mixture of blurring and Gaussian noises. The experimental results suggest that discrete-time processes offer comparable subjective and objective speech quality to their widely popular continuous counterpart, with more efficient and consistent training and inference schemas.


【27】Scattering Transformer: A Training-Free Transformer Architecture for Heart Murmur Detection
标题:Scattering Transformer:一种用于心脏杂音检测的免训练Transformer架构
链接:https://arxiv.org/abs/2509.18424

作者:Rami Zewail
摘要:为了解决对熟练的临床医生在心音解释方面的需求,最近关于自动心脏听诊的研究工作已经探索了深度学习方法。这些方法中的大多数都是基于监督学习的,在训练数据有限的情况下总是面临挑战。最近,人们对预训练的自监督音频基础模型用于生物医学最终任务的潜力越来越感兴趣。尽管表现出有希望的结果,这些基础模型通常是计算密集型的。在自动心脏听诊的背景下,本研究通过引入散射Transformer(一种用于心脏杂音检测的新型免训练Transformer架构)来探索这些通用音频基础模型的轻量级替代方案。所提出的方法利用标准的小波散射网络,通过在一个类似于变换的架构中引入上下文依赖关系,而无需任何反向传播。我们在公共CirCor DigiScope数据集上评估了我们的方法,直接将其与领先的通用基础模型进行比较。散射Transformer实现了0.786的加权准确率(WAR)和0.697的未加权平均召回率(UAR),证明了与当代最先进方法相比具有高度竞争力的性能。这项研究建立了散射Transformer作为一个可行的和有前途的替代资源有限的设置。
摘要:In an attempt to address the need for skilled clinicians in heart sound interpretation, recent research efforts on automating cardiac auscultation have explored deep learning approaches. The majority of these approaches have been based on supervised learning that is always challenged in occasions where training data is limited. More recently, there has been a growing interest in potentials of pre-trained self-supervised audio foundation models for biomedical end tasks. Despite exhibiting promising results, these foundational models are typically computationally intensive. Within the context of automatic cardiac auscultation, this study explores a lightweight alternative to these general-purpose audio foundation models by introducing the Scattering Transformer, a novel, training-free transformer architecture for heart murmur detection. The proposed method leverages standard wavelet scattering networks by introducing contextual dependencies in a transformer-like architecture without any backpropagation. We evaluate our approach on the public CirCor DigiScope dataset, directly comparing it against leading general-purpose foundational models. The Scattering Transformer achieves a Weighted Accuracy(WAR) of 0.786 and an Unweighted Average Recall(UAR) of 0.697, demonstrating performance highly competitive with contemporary state of the art methods. This study establishes the Scattering Transformer as a viable and promising alternative in resource-constrained setups.


【28】Identifying birdsong syllables without labelled data
标题:在没有标记数据的情况下识别鸟鸣音节
链接:https://arxiv.org/abs/2509.18412

作者:Melisande Teng, Julien Boussard, David Rolnick, Hugo Larochelle
摘要:识别鸟鸣中的音节序列是应对各种挑战的关键,包括鸟类个体识别和更好地理解动物交流和感觉运动学习。最近,机器学习方法在减轻专家手工标记长音频记录的需求方面表现出了巨大的潜力。然而,它们通常仍然依赖于标记数据的可用性来进行模型训练,限制了对少数物种和数据集的适用性。在这项工作中,我们建立了第一个完全无监督的算法分解成音节序列的鸟鸣录音。我们首先检测音节事件,然后将它们聚类以提取模板-音节表示-然后执行匹配追踪以将记录分解为音节序列。我们评估我们的自动注释对人类标签的数据集孟加拉雀的歌曲,发现我们的无监督的方法实现了高性能。我们还表明,我们的方法可以区分单个鸟类在一个物种内,通过其独特的声音签名,孟加拉雀和另一个物种,大山雀。
摘要:Identifying sequences of syllables within birdsongs is key to tackling a wide array of challenges, including bird individual identification and better understanding of animal communication and sensory-motor learning. Recently, machine learning approaches have demonstrated great potential to alleviate the need for experts to label long audio recordings by hand. However, they still typically rely on the availability of labelled data for model training, restricting applicability to a few species and datasets. In this work, we build the first fully unsupervised algorithm to decompose birdsong recordings into sequences of syllables. We first detect syllable events, then cluster them to extract templates --syllable representations-- before performing matching pursuit to decompose the recording as a sequence of syllables. We evaluate our automatic annotations against human labels on a dataset of Bengalese finch songs and find that our unsupervised method achieves high performance. We also demonstrate that our approach can distinguish individual birds within a species through their unique vocal signatures, for both Bengalese finches and another species, the great tit.


【29】A Dimensional Approach to Canine Bark Analysis for Assistance Dog Seizure Signaling
标题:辅助犬癫痫信号的犬吠分析的维度方法
链接:https://arxiv.org/abs/2509.18375

作者:Hailin Song, Shelley Brady, Tomás Ward, Alan F. Smeaton
摘要:犬发声的标准分类对于辅助犬来说是非常有限的,因为样本数据在犬只之间是稀疏和可变的,并且在道德上限制了对所有树皮类型的捕获。我们重新定义这个问题作为一个连续的回归任务在一个二维的唤醒价空间。我们的方法的核心是一个调整后的暹罗网络,它不是基于二进制相似性训练的,而是基于输入样本对之间的顺序和数字距离训练的。在公共数据集上训练,与回归基线相比,我们的模型在具有挑战性的效价维度上将周转百分比降低了50%。在真实世界数据集上的定性验证证实了学习空间在语义上是有意义的,建立了在严重数据限制下分析犬吠的概念验证。
摘要:Standard classification of canine vocalisations is severely limited for assistance dogs, where sample data is sparse and variable across dogs and where capture of the full range of bark types is ethically constrained. We reframe this problem as a continuous regression task within a two-dimensional arousal-valence space. Central to our approach is an adjusted Siamese Network trained not on binary similarity, but on the ordinal and numeric distance between input sample pairs. Trained on a public dataset, our model reduces Turn-around Percentage by up to 50% on the challenging valence dimension compared to a regression baseline. Qualitative validation on a real-world dataset confirms the learned space is semantically meaningful, establishing a proof-of-concept for analysing canine barking under severe data limitations.


【30】Qubit Instrumentation of Entanglement
标题:量子比特纠缠仪器
链接:https://arxiv.org/abs/2509.18340

作者:Mark Carney
备注:28 pages, 4 figures, book chapter
摘要:本章和其中描述的实验探讨了如何通过物理纠缠来表现甚至模拟“人类纠缠”。为了实现这一点,两个音乐家之间的“音调中心性”的概念通过嵌入式设备(Raspberry Pi Pico)捕获,并作为参数传递到嵌入式设备上发生的量子模拟中。然后,这些模拟的结果被编码回MIDI,并发送到玩家的乐器。音乐家的音调越接近,他们的乐器就越会被缠绕在一起。|\Phi^+ \rangle$态,它们离得越远,它们的仪器就越会纠缠在一个|\Psi^+ \rangle$ state。其目的是创建相关的随机参数--在两种工具上相同-或反相关- \r {即}在演奏者的音调关系的影响下,这些具有这些特殊性质的随机参数为量子音乐的表达增加了一个新的维度。这个概念是通过实验实现的,并提供了完整的代码和示例输出。这项工作旨在为音乐家探索和体验自己音乐体验的量子仿真铺平道路,为纠缠合奏的未来增加新的细微差别和可能性。
摘要:This chapter and the experiments described within explore how `human entanglement' might be represented and even emulated by physical entanglement. To achieve this, a notion of `tonal centrality' between two musicians is captured via MIDI and passed as a parameter into a quantum simulation taking place on an embedded device (a Raspberry Pi Pico). The results of these simulations are then coded back into MIDI and sent to the players' instruments. The closer the musicians' tonality is, the more their instruments will be entangled in a $|\Phi^+ \rangle$ state, and the further away they are the more their instruments will be entangled in a $|\Psi^+ \rangle$ state. The intention is to create random parameters that are correlative - \emph{i.e.} the same on both instruments - or anti-correlative - \emph{i.e.} the bit-wise opposite of each other, influenced by the tonal relationship from the players. These random parameters sharing these particular properties add a new dimension for quantum-musical expression. This concept was realised experimentally, and the full code and sample outputs are provided. This work aims to pave the way for musicians to explore and experience quantum emulations of their own musical experiences, adding a new nuance and possibilities for the future of \emph{entangled ensembles.}


【31】StereoFoley: Object-Aware Stereo Audio Generation from Video
标题:StereoFoley:从视频中生成对象感知立体声音频
链接:https://arxiv.org/abs/2509.18272

作者:Tornike Karchkhadze, Kuan-Lin Chen, Mojtaba (Moji)Heydari, Robert Henzel, Alessandro Toso, Mehrez Souden, Joshua Atkins
摘要:我们提出了StereoFoley,一个视频到音频生成框架,产生语义对齐,时间同步,空间准确的立体声在48 kHz。虽然最近的生成视频到音频模型实现了强大的语义和时间保真度,但它们在很大程度上仍然限于单声道或无法提供对象感知的立体声成像,受到缺乏专业混合的空间准确的视频到音频数据集的限制。首先,我们开发并训练一个从视频生成立体声音频的基础模型,在语义准确性和同步方面都达到了最先进的水平。接下来,为了克服数据集的限制,我们引入了一个合成数据生成管道,该管道将视频分析、对象跟踪和音频合成与动态平移和基于距离的响度控制相结合,从而实现空间精确的对象感知声音。最后,我们在这个合成数据集上微调基础模型,产生清晰的对象-音频对应关系。由于没有既定的指标存在,我们引入立体声对象感知措施,并通过人类听力研究验证它,表现出很强的相关性与感知。这项工作为立体声对象感知视频到音频生成建立了第一个端到端框架,解决了一个关键的差距,并在该领域建立了一个新的基准。
摘要:We present StereoFoley, a video-to-audio generation framework that produces semantically aligned, temporally synchronized, and spatially accurate stereo sound at 48 kHz. While recent generative video-to-audio models achieve strong semantic and temporal fidelity, they largely remain limited to mono or fail to deliver object-aware stereo imaging, constrained by the lack of professionally mixed, spatially accurate video-to-audio datasets. First, we develop and train a base model that generates stereo audio from video, achieving state-of-the-art in both semantic accuracy and synchronization. Next, to overcome dataset limitations, we introduce a synthetic data generation pipeline that combines video analysis, object tracking, and audio synthesis with dynamic panning and distance-based loudness controls, enabling spatially accurate object-aware sound. Finally, we fine-tune the base model on this synthetic dataset, yielding clear object-audio correspondence. Since no established metrics exist, we introduce stereo object-awareness measures and validate it through a human listening study, showing strong correlation with perception. This work establishes the first end-to-end framework for stereo object-aware video-to-audio generation, addressing a critical gap and setting a new benchmark in the field.


【32】MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech
标题:MNV-17:用于语音非言语发声识别的高质量表演性普通话数据集
链接:https://arxiv.org/abs/2509.18196

作者:Jialong Mai, Jinxin Ji, Xiaofen Xing, Chen Yang, Weidong Chen, Jingyuan Xing, Xiangmin Xu
备注:Submitted to ICASSP 2026
摘要:主流的自动语音识别(ASR)系统擅长转录词汇内容,但在很大程度上无法识别嵌入在语音中的非言语发声(NV),如叹息,笑声和咳嗽。这种能力对于全面理解人类交流非常重要,因为NVs传达了关键的情感和意图线索。NV-aware ASR的进展一直受到缺乏高质量,注释良好的数据集的阻碍。为了解决这一差距,我们引入了MNV-17,一个7.55小时的表演性普通话语音数据集。与大多数依赖于基于模型的检测的现有语料库不同,MNV-17的表演性质确保了高保真度,清晰的NV实例。据我们所知,MNV-17提供了最广泛的非语言发声类别,包括17个不同的和平衡的常见NV类。我们在四种主流ASR架构上对MNV-17进行了基准测试,评估了它们在语义转录和NV分类方面的联合性能。数据集和预训练模型检查点将公开提供,以促进未来表达性ASR的研究。
摘要:Mainstream Automatic Speech Recognition (ASR) systems excel at transcribing lexical content, but largely fail to recognize nonverbal vocalizations (NVs) embedded in speech, such as sighs, laughs, and coughs. This capability is important for a comprehensive understanding of human communication, as NVs convey crucial emotional and intentional cues. Progress in NV-aware ASR has been hindered by the lack of high-quality, well-annotated datasets. To address this gap, we introduce MNV-17, a 7.55-hour performative Mandarin speech dataset. Unlike most existing corpora that rely on model-based detection, MNV-17's performative nature ensures high-fidelity, clearly articulated NV instances. To the best of our knowledge, MNV-17 provides the most extensive set of nonverbal vocalization categories, comprising 17 distinct and well-balanced classes of common NVs. We benchmarked MNV-17 on four mainstream ASR architectures, evaluating their joint performance on semantic transcription and NV classification. The dataset and the pretrained model checkpoints will be made publicly available to facilitate future research in expressive ASR.


【33】XMUspeech Systems for the ASVspoof 5 Challenge
标题:ASVspoof 5挑战赛的XMUspeech系统
链接:https://arxiv.org/abs/2509.18102

作者:Wangjie Li, Xingjia Xie, Yishuang Li, Wenhao Guan, Kaidi Wang, Pengyu Ren, Lin Li, Qingyang Hong
摘要:在本文中,我们将我们提交的XMU语音系统提交给ASVspoof 5挑战赛的语音deepfake检测轨道。与之前的挑战相比,ASVspoof 5数据库中的音频持续时间显着增加。我们观察到,仅仅调整输入音频长度就可以大大提高系统性能。为了在多个级别上捕获伪影,我们探索了AASIST、HM-Conformer、Hubert和Wav 2 vec 2在各种输入特征和损失函数下的性能。具体来说,为了获得伪影相关信息,我们在包含欺骗话语的数据集上训练了自监督模型作为特征提取器。我们采用了自适应多尺度特征融合(AMFF)方法,将多个Transformer层的特征与手工特征相结合,以增强检测能力。此外,我们对单类损失函数进行了广泛的实验,并提供了优化的配置,以更好地配合反欺骗任务。我们的融合系统在闭合条件下的minDCF为0.4783,EER为20.45%,在开放条件下的minDCF为0.2245,EER为9.36%。
摘要:In this paper, we present our submitted XMUspeech systems to the speech deepfake detection track of the ASVspoof 5 Challenge. Compared to previous challenges, the audio duration in ASVspoof 5 database has significantly increased. And we observed that merely adjusting the input audio length can substantially improve system performance. To capture artifacts at multiple levels, we explored the performance of AASIST, HM-Conformer, Hubert, and Wav2vec2 with various input features and loss functions. Specifically, in order to obtain artifact-related information, we trained self-supervised models on the dataset containing spoofing utterances as the feature extractors. And we applied an adaptive multi-scale feature fusion (AMFF) method to integrate features from multiple Transformer layers with the hand-crafted feature to enhance the detection capability. In addition, we conducted extensive experiments on one-class loss functions and provided optimized configurations to better align with the anti-spoofing task. Our fusion system achieved a minDCF of 0.4783 and an EER of 20.45% in the closed condition, and a minDCF of 0.2245 and an EER of 9.36% in the open condition.


【34】PoolingVQ: A VQVAE Variant for Reducing Audio Redundancy and Boosting Multi-Modal Fusion in Music Emotion Analysis
标题:PoolingVQ:一种VQVAE变体,用于减少音频冗余并促进音乐情感分析中的多模式融合
链接:https://arxiv.org/abs/2509.11976

作者:Dinghao Zou, Yicheng Gong, Xiaokang Li, Xin Cao, Sunbowen Lee
摘要:多模态音乐情感分析利用音频和音频模态来增强性能。虽然主流方法专注于复杂的特征提取网络,但我们建议缩短音频序列特征的长度以减轻冗余,特别是与ESTA的紧凑表示相比,可以有效地提高任务性能。为了实现这一点,我们开发了PoolingVQ相结合的矢量量化变分自编码器(VQVAE)和空间池,它直接压缩音频特征序列通过码本引导的本地聚合,以减少冗余,然后设计了一个两阶段的共同注意力的方法来融合音频和视频信息。在公共数据集EMOPIA和VGQQ上的实验结果表明,我们的多模态框架实现了最先进的性能,PoolingVQ产生了有效的改进。我们提出的方法的代码可以在匿名GitHub上找到
摘要:Multimodal music emotion analysis leverages both audio and MIDI modalities to enhance performance. While mainstream approaches focus on complex feature extraction networks, we propose that shortening the length of audio sequence features to mitigate redundancy, especially in contrast to MIDI's compact representation, may effectively boost task performance. To achieve this, we developed PoolingVQ by combining Vector Quantized Variational Autoencoder (VQVAE) with spatial pooling, which directly compresses audio feature sequences through codebook-guided local aggregation to reduce redundancy, then devised a two-stage co-attention approach to fuse audio and MIDI information. Experimental results on the public datasets EMOPIA and VGMIDI demonstrate that our multimodal framework achieves state-of-the-art performance, with PoolingVQ yielding effective improvement. Our proposed metho's code is available at Anonymous GitHub


机器翻译由腾讯交互翻译提供,仅供参考