微信公众号:arXiv_Daily
cs.SD语音
【1】Time delay embeddings to characterize the timbre of musical instruments using Topological Data Analysis: a study on synthetic and real data
标题:使用Topological Data分析来描述乐器音色的时间延迟嵌入:对合成和真实数据的研究
链接:https://arxiv.org/abs/2510.19435
摘要:音色使我们能够区分声音,即使它们具有相同的音高和响度,在音乐,乐器识别和语音中发挥着重要作用。传统的方法,如频率分析或机器学习,往往忽略了声音的细微特征。拓扑数据分析(TDA)可以捕获复杂的模式,但其应用到音色一直有限,部分原因是不清楚如何有效地表示声音的TDA。在这项研究中,我们研究了不同的时间延迟嵌入如何影响TDA结果。我们使用合成和真实音频信号来识别增强谐波结构检测的时间延迟。我们的研究结果表明,特定的延迟,相关的分数的基本周期,允许TDA揭示关键的谐波特征,并区分整数和非整数谐波。该方法对合成和真实乐器声音有效,并为未来的工作开辟了道路,可以使用更高维的嵌入和额外的持久性统计将其扩展到更复杂的声音。
摘要:Timbre allows us to distinguish between sounds even when they share the same pitch and loudness, playing an important role in music, instrument recognition, and speech. Traditional approaches, such as frequency analysis or machine learning, often overlook subtle characteristics of sound. Topological Data Analysis (TDA) can capture complex patterns, but its application to timbre has been limited, partly because it is unclear how to represent sound effectively for TDA. In this study, we investigate how different time delay embeddings affect TDA results. Using both synthetic and real audio signals, we identify time delays that enhance the detection of harmonic structures. Our findings show that specific delays, related to fractions of the fundamental period, allow TDA to reveal key harmonic features and distinguish between integer and non-integer harmonics. The method is effective for synthetic and real musical instrument sounds and opens the way for future works, which could extend it to more complex sounds using higher-dimensional embeddings and additional persistence statistics.
【2】AMAuT: A Flexible and Efficient Multiview Audio Transformer Framework Trained from Scratch
标题:AMauT:一个灵活有效的多视图音频Transformer框架,从Scratch训练而来
链接:https://arxiv.org/abs/2510.19368
摘要:最近的基础模型,SSAST,EAT,HuBERT,Qwen-Audio和Audio Flamingo,在标准音频基准测试中获得了顶级结果,但受到固定输入速率和持续时间的限制,阻碍了它们的可重用性。本文介绍了增强驱动的多视图音频Transformer(AMAuT),这是一个从头开始训练的框架,它消除了对预训练权重的依赖,同时支持任意采样率和音频长度。AMauT集成了四个关键组件:(1)增强驱动的多视图学习,以实现鲁棒性;(2)conv 1 + conv 7 + conv 1一维CNN瓶颈,以实现稳定的时间编码;(3)双CLS + TAL令牌,用于双向上下文表示;以及(4)测试时自适应/增强(TTA ^2),以提高推理可靠性。在AudioMNIST、SpeechCommands V1和V2、VocalSound和CochlScene五个公共基准测试上的实验表明,AMauT的准确率高达99.8%,而消耗的GPU时间不到可比预训练模型所需的3%。因此,AMauT为大型预训练模型提供了一种高效灵活的替代方案,使最先进的音频分类在计算受限的设置中变得可访问。
摘要:Recent foundational models, SSAST, EAT, HuBERT, Qwen-Audio, and Audio Flamingo, achieve top-tier results across standard audio benchmarks but are limited by fixed input rates and durations, hindering their reusability. This paper introduces the Augmentation-driven Multiview Audio Transformer (AMAuT), a training-from-scratch framework that eliminates the dependency on pre-trained weights while supporting arbitrary sample rates and audio lengths. AMAuT integrates four key components: (1) augmentation-driven multiview learning for robustness, (2) a conv1 + conv7 + conv1 one-dimensional CNN bottleneck for stable temporal encoding, (3) dual CLS + TAL tokens for bidirectional context representation, and (4) test-time adaptation/augmentation (TTA^2) to improve inference reliability. Experiments on five public benchmarks, AudioMNIST, SpeechCommands V1 & V2, VocalSound, and CochlScene, show that AMAuT achieves accuracies up to 99.8% while consuming less than 3% of the GPU hours required by comparable pre-trained models. Thus, AMAuT presents a highly efficient and flexible alternative to large pre-trained models, making state-of-the-art audio classification accessible in computationally constrained settings.
【3】Steering Autoregressive Music Generation with Recursive Feature Machines
标题:使用回归特征机引导自回归音乐生成
链接:https://arxiv.org/abs/2510.19127
摘要:可控的音乐生成仍然是一个重大挑战,现有的方法通常需要模型重新训练或引入听觉伪像。我们引入了MusicRFM,这是一个适应递归特征机(RFM)的框架,通过直接引导其内部激活来实现对冻结的、预训练的音乐模型的细粒度、可解释的控制。RFM分析模型的内部梯度,以产生可解释的“概念方向”,或激活空间中对应于音符或和弦等音乐属性的特定轴。我们首先训练轻量级RFM探测器来发现MusicGen隐藏状态中的这些方向;然后,在推理过程中,我们将它们注入模型中,以实时指导生成过程,而无需每步优化。我们提出了先进的机制,这种控制,包括动态的,随时间变化的时间表和方法,同时执行多个音乐属性。我们的方法成功地导航控制和生成质量之间的权衡:我们可以提高生成目标音符的准确性从0.23到0.82,而文本提示遵守保持在约0.02的未转向基线,展示了有效的控制与最小的影响提示保真度。我们发布代码以鼓励在音乐领域对RFM进行进一步探索。
摘要:Controllable music generation remains a significant challenge, with existing methods often requiring model retraining or introducing audible artifacts. We introduce MusicRFM, a framework that adapts Recursive Feature Machines (RFMs) to enable fine-grained, interpretable control over frozen, pre-trained music models by directly steering their internal activations. RFMs analyze a model's internal gradients to produce interpretable "concept directions", or specific axes in the activation space that correspond to musical attributes like notes or chords. We first train lightweight RFM probes to discover these directions within MusicGen's hidden states; then, during inference, we inject them back into the model to guide the generation process in real-time without per-step optimization. We present advanced mechanisms for this control, including dynamic, time-varying schedules and methods for the simultaneous enforcement of multiple musical properties. Our method successfully navigates the trade-off between control and generation quality: we can increase the accuracy of generating a target musical note from 0.23 to 0.82, while text prompt adherence remains within approximately 0.02 of the unsteered baseline, demonstrating effective control with minimal impact on prompt fidelity. We release code to encourage further exploration on RFMs in the music domain.
【4】The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS
标题:MUSE基准:在音频LLMS中探索音乐感知和听觉关系推理
链接:https://arxiv.org/abs/2510.19055
备注:5 pages, 2 figures, 2 tables
摘要:多模态大型语言模型(MLLM)已经证明了音频理解的能力,但目前的评估可能会掩盖关系推理的根本弱点。我们介绍了音乐理解和结构评估(MUSE)基准,这是一个开源资源,有10个任务,旨在探索基本的音乐感知技能。我们评估了四个SOTA模型(Gemini Pro和Flash,Qwen2.5-Omni和Audio-Flamingo 3)对一个大的人类基线(N=200)。我们的研究结果揭示了SOTA能力的广泛差异以及与人类专家的持续差距。虽然双子座专业成功的基本感知,Qwen和音频火烈鸟3执行或接近机会,暴露严重的感知缺陷。此外,我们发现思想链(CoT)提示提供了不一致的,往往是有害的结果。我们的工作为评估不变的音乐表示和推动更强大的AI系统的开发提供了关键工具。
摘要:Multimodal Large Language Models (MLLMs) have demonstrated capabilities in audio understanding, but current evaluations may obscure fundamental weaknesses in relational reasoning. We introduce the Music Understanding and Structural Evaluation (MUSE) Benchmark, an open-source resource with 10 tasks designed to probe fundamental music perception skills. We evaluate four SOTA models (Gemini Pro and Flash, Qwen2.5-Omni, and Audio-Flamingo 3) against a large human baseline (N=200). Our results reveal a wide variance in SOTA capabilities and a persistent gap with human experts. While Gemini Pro succeeds on basic perception, Qwen and Audio Flamingo 3 perform at or near chance, exposing severe perceptual deficits. Furthermore, we find Chain-of-Thought (CoT) prompting provides inconsistent, often detrimental results. Our work provides a critical tool for evaluating invariant musical representations and driving development of more robust AI systems.
【5】Relative Transfer Matrix Estimator using Covariance Subtraction
标题:使用协方差减法的相对转移矩阵估计
链接:https://arxiv.org/abs/2510.19439
摘要:最近引入的相对传递矩阵(ReTM)作为多个接收器和源的相对传递函数的推广,在噪声环境中应用于语音增强和说话人分离时显示出有前途的性能。利用多通道录音的协方差矩阵来盲估计声源的ReTM,对于实际应用是非常有益的。在本文中,我们使用协方差减法提出了一个灵活的和实际可行的方法来估计一组独立的声源的ReTM。为了展示该方法的通用性,我们通过混响条件下的扬声器分离应用程序对其进行了验证。在低信噪比水平下,与现有的基于ReTM和基于相对传递函数的估计器相比,在模拟和现实生活中的环境中,分离性能进行评估。
摘要:The Relative Transfer Matrix (ReTM), recently introduced as a generalization of the relative transfer function for multiple receivers and sources, shows promising performance when applied to speech enhancement and speaker separation in noisy environments. Blindly estimating the ReTM of sound sources by exploiting the covariance matrices of multichannel recordings is highly beneficial for practical applications. In this paper, we use covariance subtraction to present a flexible and practically viable method for estimating the ReTM for a select set of independent sound sources. To show the versatility of the method, we validated it through a speaker separation application under reverberant conditions. Separation performance is evaluated at low signal-to-noise ratio levels in comparison with existing ReTM-based and relative transfer function-based estimators, in both simulated and real-life environments.
【6】EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection
标题:EchoFake:用于实际语音Deepfake检测的回放感知数据集
链接:https://arxiv.org/abs/2510.19414
摘要:语音深度伪造的日益流行引起了人们的严重关注,特别是在电话欺诈和身份盗窃等现实场景中。虽然许多反欺骗系统在实验室生成的合成语音上表现出了良好的性能,但当遇到物理重放攻击时,它们往往会失败,这是一种在实际环境中使用的常见且低成本的攻击形式。我们的实验表明,在现有数据集上训练的模型表现出严重的性能下降,在重放音频上评估时,平均准确率下降到59.6%。为了弥合这一差距,我们提出了EchoFake,这是一个综合数据集,包括来自13,000多名扬声器的超过120小时的音频,具有尖端的zero-shot文本到语音(TTS)语音和在各种设备和真实环境设置下收集的物理重放录音。此外,我们评估了三个基线检测模型,并表明在EchoFake上训练的模型在数据集上实现了较低的平均EER,这表明泛化能力更好。通过引入与现实世界部署相关的更多实际挑战,EchoFake为推进欺骗检测方法提供了更现实的基础。
摘要:The growing prevalence of speech deepfakes has raised serious concerns, particularly in real-world scenarios such as telephone fraud and identity theft. While many anti-spoofing systems have demonstrated promising performance on lab-generated synthetic speech, they often fail when confronted with physical replay attacks-a common and low-cost form of attack used in practical settings. Our experiments show that models trained on existing datasets exhibit severe performance degradation, with average accuracy dropping to 59.6% when evaluated on replayed audio. To bridge this gap, we present EchoFake, a comprehensive dataset comprising more than 120 hours of audio from over 13,000 speakers, featuring both cutting-edge zero-shot text-to-speech (TTS) speech and physical replay recordings collected under varied devices and real-world environmental settings. Additionally, we evaluate three baseline detection models and show that models trained on EchoFake achieve lower average EERs across datasets, indicating better generalization. By introducing more practical challenges relevant to real-world deployment, EchoFake offers a more realistic foundation for advancing spoofing detection methods.
【1】VBx for End-to-End Neural and Clustering-based Diarization
标题:VB x用于端到端神经和基于迭代的扩展
链接:https://arxiv.org/abs/2510.19572
备注:Submitted to ICASSP 2026
摘要:我们提出了改进的两阶段端到端的神经日记向量聚类(EEND-VC)框架中的发言人日记。第一阶段采用基于一致性的EEND模型与WavLM功能推断短窗口内的帧级扬声器活动。然后,在第二阶段中,通过跨窗口聚类扬声器嵌入来导出全局扬声器的身份和计数。这项工作的重点是改善第二阶段,我们过滤不可靠的嵌入短段和重新分配后,聚类。我们还集成了VBx聚类,以提高鲁棒性时,扬声器的数量是大的,个人发言持续时间是有限的。在没有微调EEND模型或调整每个数据集的聚类参数的情况下,对跨越多个域的复合基准进行评估。尽管如此,该系统推广良好,并匹配或超过最近的最先进的性能。
摘要:We present improvements to speaker diarization in the two-stage end-to-end neural diarization with vector clustering (EEND-VC) framework. The first stage employs a Conformer-based EEND model with WavLM features to infer frame-level speaker activity within short windows. The identities and counts of global speakers are then derived in the second stage by clustering speaker embeddings across windows. The focus of this work is to improve the second stage; we filter unreliable embeddings from short segments and reassign them after clustering. We also integrate the VBx clustering to improve robustness when the number of speakers is large and individual speaking durations are limited. Evaluation on a compound benchmark spanning multiple domains is conducted without fine-tuning the EEND model or tuning clustering parameters per dataset. Despite this, the system generalizes well and matches or exceeds recent state-of-the-art performance.
【2】Relative Transfer Matrix Estimator using Covariance Subtraction
标题:使用协方差减法的相对转移矩阵估计
链接:https://arxiv.org/abs/2510.19439
摘要:最近引入的相对传递矩阵(ReTM)作为多个接收器和源的相对传递函数的推广,在噪声环境中应用于语音增强和说话人分离时显示出有前途的性能。利用多通道录音的协方差矩阵来盲估计声源的ReTM,对于实际应用是非常有益的。在本文中,我们使用协方差减法提出了一个灵活的和实际可行的方法来估计一组独立的声源的ReTM。为了展示该方法的通用性,我们通过混响条件下的扬声器分离应用程序对其进行了验证。在低信噪比水平下,与现有的基于ReTM和基于相对传递函数的估计器相比,在模拟和现实生活中的环境中,分离性能进行评估。
摘要:The Relative Transfer Matrix (ReTM), recently introduced as a generalization of the relative transfer function for multiple receivers and sources, shows promising performance when applied to speech enhancement and speaker separation in noisy environments. Blindly estimating the ReTM of sound sources by exploiting the covariance matrices of multichannel recordings is highly beneficial for practical applications. In this paper, we use covariance subtraction to present a flexible and practically viable method for estimating the ReTM for a select set of independent sound sources. To show the versatility of the method, we validated it through a speaker separation application under reverberant conditions. Separation performance is evaluated at low signal-to-noise ratio levels in comparison with existing ReTM-based and relative transfer function-based estimators, in both simulated and real-life environments.
【3】EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection
标题:EchoFake:用于实际语音Deepfake检测的回放感知数据集
链接:https://arxiv.org/abs/2510.19414
摘要:语音深度伪造的日益流行引起了人们的严重关注,特别是在电话欺诈和身份盗窃等现实场景中。虽然许多反欺骗系统在实验室生成的合成语音上表现出了良好的性能,但当遇到物理重放攻击时,它们往往会失败,这是一种在实际环境中使用的常见且低成本的攻击形式。我们的实验表明,在现有数据集上训练的模型表现出严重的性能下降,在重放音频上评估时,平均准确率下降到59.6%。为了弥合这一差距,我们提出了EchoFake,这是一个综合数据集,包括来自13,000多名扬声器的超过120小时的音频,具有尖端的zero-shot文本到语音(TTS)语音和在各种设备和真实环境设置下收集的物理重放录音。此外,我们评估了三个基线检测模型,并表明在EchoFake上训练的模型在数据集上实现了较低的平均EER,这表明泛化能力更好。通过引入与现实世界部署相关的更多实际挑战,EchoFake为推进欺骗检测方法提供了更现实的基础。
摘要:The growing prevalence of speech deepfakes has raised serious concerns, particularly in real-world scenarios such as telephone fraud and identity theft. While many anti-spoofing systems have demonstrated promising performance on lab-generated synthetic speech, they often fail when confronted with physical replay attacks-a common and low-cost form of attack used in practical settings. Our experiments show that models trained on existing datasets exhibit severe performance degradation, with average accuracy dropping to 59.6% when evaluated on replayed audio. To bridge this gap, we present EchoFake, a comprehensive dataset comprising more than 120 hours of audio from over 13,000 speakers, featuring both cutting-edge zero-shot text-to-speech (TTS) speech and physical replay recordings collected under varied devices and real-world environmental settings. Additionally, we evaluate three baseline detection models and show that models trained on EchoFake achieve lower average EERs across datasets, indicating better generalization. By introducing more practical challenges relevant to real-world deployment, EchoFake offers a more realistic foundation for advancing spoofing detection methods.
【4】An Efficient Neural Network for Modeling Human Auditory Neurograms for Speech
标题:一种有效的神经网络建模人类语音听觉神经图
链接:https://arxiv.org/abs/2510.19354
摘要:以Bruce等人为例的经典的边缘-边缘模型,2018年,提供高保真模拟,但随机和计算要求高,限制了大规模实验和低延迟使用。先前的神经编码器近似于外围的方面;然而,很少有明确训练来再现确定性的率域神经图,这阻碍了同类评估。我们提出了一个紧凑的卷积编码器,近似布鲁斯平均速率路径和映射音频到多频神经图。我们故意忽略随机尖峰效应,专注于确定性映射(相同输入的相同输出)。该编码器采用计算效率高的设计,实现了与参考信号的紧密对应,同时显著减少了计算量,为听觉神经科学和音频信号处理应用提供了高效的建模和前端处理。
摘要:Classical auditory-periphery models, exemplified by Bruce et al., 2018, provide high-fidelity simulations but are stochastic and computationally demanding, limiting large-scale experimentation and low-latency use. Prior neural encoders approximate aspects of the periphery; however, few are explicitly trained to reproduce the deterministic, rate-domain neurogram , hindering like-for-like evaluation. We present a compact convolutional encoder that approximates the Bruce mean-rate pathway and maps audio to a multi-frequency neurogram. We deliberately omit stochastic spiking effects and focus on a deterministic mapping (identical outputs for identical inputs). Using a computationally efficient design, the encoder achieves close correspondence to the reference while significantly reducing computation, enabling efficient modeling and front-end processing for auditory neuroscience and audio signal processing applications.
【5】Auditory Attention Decoding from Ear-EEG Signals: A Dataset with Dynamic Attention Switching and Rigorous Cross-Validation
标题:耳脑电信号的听觉注意力解码:具有动态注意力切换和严格交叉验证的数据集
链接:https://arxiv.org/abs/2510.19174
摘要:近年来,头皮脑电(EEG)在听觉注意解码(AAD)方面取得了可喜的成果,这促使人们开发了一种灵活、便携的耳-脑电系统cEEGrid。虽然先前基于cEEGrid的研究已经证实了AAD的可行性,但他们往往忽略了注意状态在现实世界中的动态性质。为了解决这个差距,一个新的cEEGrid数据集具有三个并发扬声器分布在三个五个不同的空间位置。新的数据集旨在探测现实场景中的注意力跟踪和切换。嵌套留一法验证-一种比传统的单循环留一法验证更严格的方法-被用来减少EEG复杂的时间动态产生的偏差。四个基于规则的模型进行了评估:维纳滤波器(WF),典型成分分析(CCA),共同的空间模式(CSP)和黎曼几何为基础的分类器(RGC)。在30秒的决策窗口下,WF和CCA模型分别实现了41.5%和41.4%的解码准确度,而CSP和RGC模型在10秒的窗口下分别实现了37.8%和37.6%的准确度。值得注意的是,WF和CCA都成功地跟踪了所有实验任务中的注意状态转换。此外,对于位于上部cEEGrid布局和靠近收听者右耳的电极,观察到更高的解码精度。这些发现强调了动态的,生态有效的范例和严格的验证,在推进AAD研究与cEEGrid的效用。
摘要:Recent promising results in auditory attention decoding (AAD) using scalp electroencephalography (EEG) have motivated the exploration of cEEGrid, a flexible and portable ear-EEG system. While prior cEEGrid-based studies have confirmed the feasibility of AAD, they often neglect the dynamic nature of attentional states in real-world contexts. To address this gap, a novel cEEGrid dataset featuring three concurrent speakers distributed across three of five distinct spatial locations is introduced. The novel dataset is designed to probe attentional tracking and switching in realistic scenarios. Nested leave-one-out validation-an approach more rigorous than conventional single-loop leave-one-out validation-is employed to reduce biases stemming from EEG's intricate temporal dynamics. Four rule-based models are evaluated: Wiener filter (WF), canonical component analysis (CCA), common spatial pattern (CSP) and Riemannian Geometry-based classifier (RGC). With a 30-second decision window, WF and CCA models achieve decoding accuracies of 41.5% and 41.4%, respectively, while CSP and RGC models yield 37.8% and 37.6% accuracies using a 10-second window. Notably, both WF and CCA successfully track attentional state switches across all experimental tasks. Additionally, higher decoding accuracies are observed for electrodes positioned at the upper cEEGrid layout and near the listener's right ear. These findings underscore the utility of dynamic, ecologically valid paradigms and rigorous validation in advancing AAD research with cEEGrid.
【6】StutterZero and StutterFormer: End-to-End Speech Conversion for Stuttering Transcription and Correction
标题:StutterZero和StutterFormer:用于口吃转录和纠正的端到端语音转换
链接:https://arxiv.org/abs/2510.18938
备注:13 pages, 5 figures
摘要:全世界有超过7000万人患有口吃,但大多数自动语音系统会误解不流利的话语或无法准确地转录它们。现有的口吃校正方法依赖于手工特征提取或多级自动语音识别(ASR)和文本到语音(TTS)管道,这些管道将转录与音频重建分开,并且通常会放大失真。这项工作介绍了StutterZero和StutterFormer,这是第一个端到端的波形到波形模型,可以直接将口吃的语音转换为流利的语音,同时联合预测其转录。StutterZero采用了一个带有注意力的卷积双向LSTM编码器-解码器,而StutterFormer则集成了一个具有共享声学语言表示的双流Transformer。这两种架构都是在SEP-28 K和LibriStutter语料库合成的成对口吃流利数据上训练的,并在FluencyBank数据集的未见过的说话者上进行评估。在所有基准测试中,与领先的Whisper-Medium模型相比,StutterZero的单词错误率(WER)降低了24%,语义相似度(BERTScore)提高了31%。StutterFormer取得了更好的结果,WER降低了28%,BERTScore提高了34%。结果验证了直接端到端口吃到流利语音转换的可行性,为包容性人机交互,语音治疗和面向可访问性的AI系统提供了新的机会。
摘要:Over 70 million people worldwide experience stuttering, yet most automatic speech systems misinterpret disfluent utterances or fail to transcribe them accurately. Existing methods for stutter correction rely on handcrafted feature extraction or multi-stage automatic speech recognition (ASR) and text-to-speech (TTS) pipelines, which separate transcription from audio reconstruction and often amplify distortions. This work introduces StutterZero and StutterFormer, the first end-to-end waveform-to-waveform models that directly convert stuttered speech into fluent speech while jointly predicting its transcription. StutterZero employs a convolutional-bidirectional LSTM encoder-decoder with attention, whereas StutterFormer integrates a dual-stream Transformer with shared acoustic-linguistic representations. Both architectures are trained on paired stuttered-fluent data synthesized from the SEP-28K and LibriStutter corpora and evaluated on unseen speakers from the FluencyBank dataset. Across all benchmarks, StutterZero had a 24% decrease in Word Error Rate (WER) and a 31% improvement in semantic similarity (BERTScore) compared to the leading Whisper-Medium model. StutterFormer achieved better results, with a 28% decrease in WER and a 34% improvement in BERTScore. The results validate the feasibility of direct end-to-end stutter-to-fluent speech conversion, offering new opportunities for inclusive human-computer interaction, speech therapy, and accessibility-oriented AI systems.
【7】RIR-Mega: a large-scale simulated room impulse response dataset for machine learning and room acoustics modeling
标题:RIR-Mega:用于机器学习和房间声学建模的大规模模拟房间脉冲响应数据集
链接:https://arxiv.org/abs/2510.18917
备注:8 pages, 3 figures
摘要:室内脉冲响应是混响消除、鲁棒语音识别、声源定位和室内声学估计的核心资源。我们提出了RIR-巨型,一个大型的集合,模拟RIR描述了一个紧凑的,机器友好的元数据模式和分布的简单工具进行验证和重用。该数据集附带了一个拥抱脸数据集加载器,用于元数据检查和校验和的脚本,以及一个参考回归基线,可以从波形中预测RT 60等目标。在36,000和4,000个样本的训练和验证分割上,轻量级时间和频谱特征的小型随机森林达到接近0.013 s的平均绝对误差和接近0.022 s的均方根误差。我们在Hugging Face上托管了一个包含1,000个线性阵列RIR和3,000个圆形阵列RIR的子集,用于流式传输和快速测试,并在Zenodo上保存了完整的50,000个RIR存档。数据集和代码是公开的,以支持可重复的研究。
摘要:Room impulse responses are a core resource for dereverberation, robust speech recognition, source localization, and room acoustics estimation. We present RIR-Mega, a large collection of simulated RIRs described by a compact, machine friendly metadata schema and distributed with simple tools for validation and reuse. The dataset ships with a Hugging Face Datasets loader, scripts for metadata checks and checksums, and a reference regression baseline that predicts RT60 like targets from waveforms. On a train and validation split of 36,000 and 4,000 examples, a small Random Forest on lightweight time and spectral features reaches a mean absolute error near 0.013 s and a root mean square error near 0.022 s. We host a subset with 1,000 linear array RIRs and 3,000 circular array RIRs on Hugging Face for streaming and quick tests, and preserve the complete 50,000 RIR archive on Zenodo. The dataset and code are public to support reproducible studies.
【8】Which Evaluation for Which Model? A Taxonomy for Speech Model Assessment
标题:对哪个模型进行哪种评估?语音模型评估的分类
链接:https://arxiv.org/abs/2510.19509
备注:57 pages (26 main, 25 appendix, 6 references)
摘要:语音基础模型最近在广泛的任务中取得了显着的能力。然而,他们的评价仍然脱节的任务和模型类型。不同的模型擅长语音处理的不同方面,因此需要不同的评估协议。本文提出了一个统一的分类法,解决了这个问题:哪种评估是适合于哪种模式?该分类定义了三个正交轴:被测量的\textbf{evaluation aspect},尝试任务所需的模型功能,以及执行任务所需的任务或协议要求。我们沿着这些轴对广泛的现有评估和基准进行分类,跨越表示学习,语音生成和交互式对话等领域。通过将每个评估映射到模型暴露的能力(例如,语音生成,实时处理)及其方法学要求(例如,微调数据,人工判断),分类法提供了一个原则性的框架,使模型与适当的评估方法相一致。它还揭示了系统性的差距,如韵律,互动或推理的覆盖范围有限,突出了未来基准设计的优先事项。总的来说,这项工作提供了一个概念基础和实践指南,选择,解释和扩展语音模型的评估。
摘要:Speech foundation models have recently achieved remarkable capabilities across a wide range of tasks. However, their evaluation remains disjointed across tasks and model types. Different models excel at distinct aspects of speech processing and thus require different evaluation protocols. This paper proposes a unified taxonomy that addresses the question: Which evaluation is appropriate for which model? The taxonomy defines three orthogonal axes: the \textbf{evaluation aspect} being measured, the model capabilities required to attempt the task, and the task or protocol requirements needed to perform it. We classify a broad set of existing evaluations and benchmarks along these axes, spanning areas such as representation learning, speech generation, and interactive dialogue. By mapping each evaluation to the capabilities a model exposes (e.g., speech generation, real-time processing) and to its methodological demands (e.g., fine-tuning data, human judgment), the taxonomy provides a principled framework for aligning models with suitable evaluation methods. It also reveals systematic gaps, such as limited coverage of prosody, interaction, or reasoning, that highlight priorities for future benchmark design. Overall, this work offers a conceptual foundation and practical guide for selecting, interpreting, and extending evaluations of speech models.
【9】Re-evaluating Minimum Bayes Risk Decoding for Automatic Speech Recognition
标题:自动语音识别中的最小Bayes风险解码重新评估
链接:https://arxiv.org/abs/2510.19471
摘要:最近的研究表明,基于样本的最小贝叶斯风险(MBR)解码在文本到文本生成任务中优于波束搜索,例如机器翻译,文本摘要和图像字幕。另一方面,波束搜索是语音到文本任务的当前实践,例如自动语音识别(ASR)和语音翻译(ST)。鉴于MBR解码在文本到文本生成任务中是有效的,因此有理由期望它对语音到文本任务也是有效的。在本文中,我们评估MBR解码的ASR和ST任务的英语和日语使用耳语及其衍生模型。我们观察到,MBR解码的准确性优于波束搜索在大多数实验设置,我们已经评估。结果表明,MBR解码是一种很有前途的方法,离线ASR和ST任务,需要高精度。该代码可在https://github.com/CyberAgentAILab/mbr-for-asr上获得
摘要:Recent work has shown that sample-based Minimum Bayes Risk (MBR) decoding outperforms beam search in text-to-text generation tasks, such as machine translation, text summarization, and image captioning. On the other hand, beam search is the current practice for speech-to-text tasks such as automatic speech recognition (ASR) and Speech Translation (ST). Given that MBR decoding is effective in text-to-text generation tasks, it is reasonable to expect it to also be effective for speech-to-text tasks. In this paper, we evaluate MBR decoding for ASR and ST tasks on English and Japanese using Whisper and its derivative models. We observe that the accuracy of MBR decoding outperforms that of beam search in most of the experimental settings we have evaluated. The results show that MBR decoding is a promising method for offline ASR and ST tasks that require high accuracy. The code is available at https://github.com/CyberAgentAILab/mbr-for-asr
【10】Steering Autoregressive Music Generation with Recursive Feature Machines
标题:使用回归特征机引导自回归音乐生成
链接:https://arxiv.org/abs/2510.19127
摘要:可控的音乐生成仍然是一个重大挑战,现有的方法通常需要模型重新训练或引入听觉伪像。我们引入了MusicRFM,这是一个适应递归特征机(RFM)的框架,通过直接引导其内部激活来实现对冻结的、预训练的音乐模型的细粒度、可解释的控制。RFM分析模型的内部梯度,以产生可解释的“概念方向”,或激活空间中对应于音符或和弦等音乐属性的特定轴。我们首先训练轻量级RFM探测器来发现MusicGen隐藏状态中的这些方向;然后,在推理过程中,我们将它们注入模型中,以实时指导生成过程,而无需每步优化。我们提出了先进的机制,这种控制,包括动态的,随时间变化的时间表和方法,同时执行多个音乐属性。我们的方法成功地导航控制和生成质量之间的权衡:我们可以提高生成目标音符的准确性从0.23到0.82,而文本提示遵守保持在约0.02的未转向基线,展示了有效的控制与最小的影响提示保真度。我们发布代码以鼓励在音乐领域对RFM进行进一步探索。
摘要:Controllable music generation remains a significant challenge, with existing methods often requiring model retraining or introducing audible artifacts. We introduce MusicRFM, a framework that adapts Recursive Feature Machines (RFMs) to enable fine-grained, interpretable control over frozen, pre-trained music models by directly steering their internal activations. RFMs analyze a model's internal gradients to produce interpretable "concept directions", or specific axes in the activation space that correspond to musical attributes like notes or chords. We first train lightweight RFM probes to discover these directions within MusicGen's hidden states; then, during inference, we inject them back into the model to guide the generation process in real-time without per-step optimization. We present advanced mechanisms for this control, including dynamic, time-varying schedules and methods for the simultaneous enforcement of multiple musical properties. Our method successfully navigates the trade-off between control and generation quality: we can increase the accuracy of generating a target musical note from 0.23 to 0.82, while text prompt adherence remains within approximately 0.02 of the unsteered baseline, demonstrating effective control with minimal impact on prompt fidelity. We release code to encourage further exploration on RFMs in the music domain.
【11】The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS
标题:MUSE基准:在音频LLMS中探索音乐感知和听觉关系推理
链接:https://arxiv.org/abs/2510.19055
备注:5 pages, 2 figures, 2 tables
摘要:多模态大型语言模型(MLLM)已经证明了音频理解的能力,但目前的评估可能会掩盖关系推理的根本弱点。我们介绍了音乐理解和结构评估(MUSE)基准,这是一个开源资源,有10个任务,旨在探索基本的音乐感知技能。我们评估了四个SOTA模型(Gemini Pro和Flash,Qwen2.5-Omni和Audio-Flamingo 3)对一个大的人类基线(N=200)。我们的研究结果揭示了SOTA能力的广泛差异以及与人类专家的持续差距。虽然双子座专业成功的基本感知,Qwen和音频火烈鸟3执行或接近机会,暴露严重的感知缺陷。此外,我们发现思想链(CoT)提示提供了不一致的,往往是有害的结果。我们的工作为评估不变的音乐表示和推动更强大的AI系统的开发提供了关键工具。
摘要:Multimodal Large Language Models (MLLMs) have demonstrated capabilities in audio understanding, but current evaluations may obscure fundamental weaknesses in relational reasoning. We introduce the Music Understanding and Structural Evaluation (MUSE) Benchmark, an open-source resource with 10 tasks designed to probe fundamental music perception skills. We evaluate four SOTA models (Gemini Pro and Flash, Qwen2.5-Omni, and Audio-Flamingo 3) against a large human baseline (N=200). Our results reveal a wide variance in SOTA capabilities and a persistent gap with human experts. While Gemini Pro succeeds on basic perception, Qwen and Audio Flamingo 3 perform at or near chance, exposing severe perceptual deficits. Furthermore, we find Chain-of-Thought (CoT) prompting provides inconsistent, often detrimental results. Our work provides a critical tool for evaluating invariant musical representations and driving development of more robust AI systems.
机器翻译由腾讯交互翻译提供,仅供参考
