微信公众号:arXiv_Daily
cs.SD语音
标题:使用Whisper的语音置信度检测半监督框架
链接:https://arxiv.org/abs/2605.12387
备注:12 pages, 9 Figures, Submitted to IEEE Transactions on Audio, Speech and Language Processing
摘要:
摘要:
【2】Poly-SVC: Polyphony-Aware Singing Voice Conversion with Harmonic Modeling
标题:Poly-SRC:利用和声建模的多音感知歌唱声音转换链接:https://arxiv.org/abs/2605.12310
备注:Accepted by ICASSP 2026
摘要:
摘要:
【3】STRUM: A Spectral Transcription and Rhythm Understanding Model for End-to-End Generation of Playable Rhythm-Game Charts
标题:SEARCH:用于端到端生成可玩节奏游戏图表的光谱转录和节奏理解模型链接:https://arxiv.org/abs/2605.12135
备注:9 pages, 4 figures, 3 tables. Code and models: https://github.com/
摘要:我们提出了一个音频到图表的管道,将原始录音转换为可播放的克隆英雄/ YARG图表,用于鼓,吉他,贝司,人声和键,而无需任何Oracle元数据。CNOM是一个多阶段的混合:一个两阶段的CRNN起始检测器和一个六模型集成分类器鼓;神经起始检测器与吉他和低音单声道音高跟踪;单词对齐ASR的人声;和光谱键盘检测键。我们在一个30首歌曲的信封基准上进行评估,该基准是通过一个单一的音频质量标准筛选候选歌曲而构建的-htdemucs_6s源分离后的中值1秒鼓干RMS。在此基准测试中,在每首歌曲全局偏移搜索的+/- 100 ms容差下,CNORM实现了鼓起始F1 = 0.838,低音F1 = 0.694,吉他F1 = 0.651,人声F1 = 0.539。我们报告了一个完整的消融7鼓管道组件配对每首歌Wilcoxon测试,地面实况音频定时分布在社区克隆英雄图表的分析,和每类的鼓分类混淆矩阵。代码、模型权重和完整的基准清单都已发布。
摘要:We present STRUM (Spectral Transcription and Rhythm Understanding Model), an audio-to-chart pipeline that converts raw recordings into playable Clone Hero / YARG charts for drums, guitar, bass, vocals, and keys without any oracle metadata. STRUM is a multi-stage hybrid: a two-stage CRNN onset detector and a six-model ensemble classifier for drums; neural onset detectors with monophonic pitch tracking for guitar and bass; word-aligned ASR for vocals; and spectral keyboard detection for keys. We evaluate on a 30-song in-envelope benchmark constructed by screening candidate songs on a single audio-quality criterion -- the median 1-second drum-stem RMS after htdemucs_6s source separation. On this benchmark STRUM achieves drums onset F1 = 0.838, bass F1 = 0.694, guitar F1 = 0.651, and vocals F1 = 0.539 at a +/- 100 ms tolerance with per-song global offset search. We report a complete ablation of seven drum-pipeline components with paired per-song Wilcoxon tests, an analysis of ground-truth-to-audio timing distributions in community Clone Hero charts, and a per-class confusion matrix for the drum classifier. Code, model weights, and the full benchmark manifest are released.
【4】AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling
标题:AuDirector:沉浸式音频讲故事的自我反思闭环框架链接:https://arxiv.org/abs/2605.11866
摘要:尽管在文本和视觉生成方面取得了进展,但创建连贯的长篇音频叙事仍然具有挑战性。现有的框架通常表现出诸如与语音性能不匹配的字符设置、不充分的自我校正机制以及有限的人类交互等限制。为了应对这些挑战,我们提出了AuDirector,一个自反射的闭环多智能体框架。具体而言,它涉及到一个身份感知的预生产机制,将叙事文本转换为字符配置文件和话语级的情感指令,以检索合适的语音候选人和指导表达性语音合成,从而促进上下文一致的语音适应。为了提高质量,协作合成和校正模块引入了闭环自校正机制,以系统地审计和再生有缺陷的音频组件。此外,人工引导的交互式优化模块通过解释自然语言反馈以交互式地优化底层脚本来促进用户控制。实验表明,AuDirector在结构连贯性、情感表现力和声学保真度方面与最先进的基线相比具有优异的性能。音频样本可以在https://anonymous-itsh.github.io/上找到。
摘要:Despite advances in text and visual generation, creating coherent long-form audio narratives remains challenging. Existing frameworks often exhibit limitations such as mismatched character settings with voice performance, insufficient self-correction mechanisms, and limited human interactivity. To address these challenges, we propose AuDirector, a self-reflective closed-loop multi-agent framework. Specifically, it involves an Identity-Aware Pre-production mechanism that transforms narrative texts into character profiles and utterance-level emotional instructions to retrieve suitable voice candidates and guide expressive speech synthesis, thereby promoting context-aligned voice adaptation. To enhance quality, a Collaborative Synthesis and Correction module introduces a closed-loop self-correction mechanism to systematically audit and regenerate defective audio components. Furthermore, a Human-Guided Interactive Refinement module facilitates user control by interpreting natural language feedback to interactively refine the underlying scripts. Experiments demonstrate that AuDirector achieves superior performance compared to state-of-the-art baselines in structural coherence, emotional expressiveness, and acoustic fidelity. Audio samples can be found at https://anonymous-itsh.github.io/.
【5】Exploring Token-Space Manipulation in Latent Audio Tokenizers
标题:探索潜在音频令牌器中的令牌空间操纵链接:https://arxiv.org/abs/2605.11192
摘要:神经音频编解码器为语音生成和操纵提供紧凑的离散表示。然而,大多数编解码器将令牌组织为帧级序列,使得难以研究或干预全局变化因素。在这项工作中,我们提出了用于令牌空间编辑的潜在音频令牌器(LATTE),它将一组固定的可学习潜在令牌添加到音频特征序列中,并仅保留这些令牌用于量化和解码。这种设计产生了一个紧凑的,非时间对齐的瓶颈,其中每个令牌可以聚合整个话语的全局信息。我们表明,由此产生的令牌化保留竞争力的重建质量在低比特率的语音编码设置,同时使简单的令牌空间干预。特别是,我们发现,在话语之间交换选定的潜在标记位置可以修改全局属性,如说话者身份和背景噪声,我们评估这些干预对语音转换和去噪任务。我们的研究结果表明,紧凑的潜在音频标记器可以支持可控的音频操作,而无需在特定于任务的编辑模型中进行监督。
摘要:Neural audio codecs provide compact discrete representations for speech generation and manipulation. However, most codecs organize tokens as frame-level sequences, making it difficult to study or intervene on global factors of variation. In this work, we propose the Latent Audio Tokenizer for Token-space Editing (LATTE) that appends a fixed set of learnable latent tokens to the audio feature sequence and retains only these tokens for quantization and decoding. This design produces a compact, non-temporally aligned bottleneck in which each token can aggregate global information across the full utterance. We show that the resulting tokenizer preserves competitive reconstruction quality in low-bitrate speech coding settings while enabling simple token-space interventions. In particular, we find that swapping selected latent token positions between utterances can modify global attributes, such as speaker identity and background noise, and we evaluate these interventions on voice conversion and denoising tasks. Our results suggest that compact latent audio tokenizers can support controllable audio manipulation without supervision in task-specific editing models.
【6】AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling
标题:AffectCodec:用于表达性语音建模的描述保持神经语音编解码器链接:https://arxiv.org/abs/2605.11098
备注:Accepted to ACL Findings 2026
摘要:神经语音编解码器为语音语言模型提供离散表示,但是情感线索在量化期间经常被降级。现有的编解码器主要优化声学重建,使情感表现力在表示层面上建模不足。我们提出了一个情感引导的神经语音编解码器,明确保留情感信息,同时保持语义保真度和韵律自然。我们的框架结合了情感语义引导的潜在调制,关系保持情感语义蒸馏,和情感加权语义对齐,以保留在压缩下的情感显着线索。在语音重建、情感识别和下游文本到语音生成方面的广泛评估表明,在不牺牲内容准确性的情况下,情感一致性和感知质量得到了改善。
摘要:Neural speech codecs provide discrete representations for speech language models, but emotional cues are often degraded during quantization. Existing codecs mainly optimize acoustic reconstruction, leaving emotion expressiveness insufficiently modeled at the representation level. We propose an emotion-guided neural speech codec that explicitly preserves emotional information while maintaining semantic fidelity and prosodic naturalness. Our framework combines emotion-semantic guided latent modulation, relation-preserving emotional-semantic distillation, and emotion-weighted semantic alignment to retain emotionally salient cues under compression. Extensive evaluations across speech reconstruction, emotion recognition, and downstream text-to-speech generation demonstrate improved emotion consistency and perceptual quality without sacrificing content accuracy.
【7】The SMC Blind Spot: A Failure Mode Analysis of State-of-the-Art Beat Tracking
标题:SMC盲点:最先进节拍跟踪的故障模式分析链接:https://arxiv.org/abs/2605.12287
备注:6 pages, 3 figures. Technical report on beat tracking failure modes; prepared for ISMIR 2026
摘要:在过去的二十年里,音乐节拍跟踪的任务已经从启发式起始点检测算法过渡到高性能的深度神经网络(DNN)。尽管基于DNN的节拍跟踪模型在主流的、抽象的数据集上实现了近乎完美的性能,但SMC数据集顽固地产生了较低的F-measure分数。通过测试最先进的模型在SMC数据集中检测单个轨道上的节拍的能力,我们确定了三种不同的故障模式:倍频程错误、连续性错误和完全跟踪故障,其中所有指标都低于0.3。我们发现,国家的最先进的模型往往会产生“自信,但错误的”激活。此外,我们表明,标准DBN的默认最低速度为55 BPM,防止它推断正确的速度为21%的SMC轨道,迫使双节奏的预测慢音乐。通过揭示这些基本的疏忽,我们为改进节拍和强拍检测提供了具体的方向,特别强调了训练数据的多样化和多假设节奏估计。
摘要:Over the past two decades, the task of musical beat tracking has transitioned from heuristic onset detection algorithms to highly capable deep neural networks (DNN). Although DNN-based beat tracking models achieve near-perfect performance on mainstream, percussive datasets, the SMC dataset has stubbornly yielded low F-measure scores. By testing how well state-of-the-art models detect beats on individual tracks in the SMC dataset, we identify three distinct failure modes: octave errors, continuity errors, and complete tracking failure where all metrics fall below 0.3. We reveal that state-of-the-art models tend to generate "confident-but-wrong" activations. Furthermore, we show that the standard DBN's default minimum tempo of 55 BPM prevents it from inferring the correct tempo for 21\% of SMC tracks, forcing double-tempo predictions on slow music. By exposing such fundamental oversights, we provide concrete directions for improving beat and downbeat detection, specifically emphasizing training data diversification and multi-hypothesis tempo estimation.
【8】Adaptive Diagonal Loading using Krylov Subspaces for Robust Beamforming
标题:使用Krylov子空间的自适应对角加载用于鲁棒束形成链接:https://arxiv.org/abs/2605.11286
备注:5 pages, 8 figures
摘要:可靠的自适应波束形成对于在高动态声学环境中工作的大型麦克风阵列至关重要。在以快速移动的说话者和说话者为特征的场景中,用于估计空间相关矩阵的可用样本支持通常是快照不足的。这种缺陷降低了白噪声增益(WNG),导致严重的目标信号抵消。为了保证波束形成的稳定性和鲁棒性,我们之前提出了一种自适应对角加载方法,该方法利用Kantorovich不等式来保证WNG严格保持在指定的范围内。然而,准确地确定最小的必要负载水平需要计算空间相关矩阵的极端特征值,这对于大型阵列来说是一个计算昂贵的$\mathcal{O}(M^3)$操作。在本文中,我们介绍了一个高效的$\mathcal{O}(kM^2)$估计技术,使用Lanczos迭代,以建立一个小的Krylov子空间。通过将相关矩阵投影到维度为k \ll M$的三对角矩阵上,我们提取快速收敛到确切的极端特征值的Ritz值。我们的评估表明,这种Lanczos加速方法实现了与精确特征值分解(EVD)相同的性能,以一小部分计算成本确保了最佳干扰抑制和严格的WNG遵守。
摘要:Reliable adaptive beamforming is critical for large microphone arrays operating in highly dynamic acoustic environments. In scenarios characterized by fast-moving talkers and interferers, the available sample support for estimating the spatial correlation matrix is often snapshot-deficient. This deficiency degrades the White Noise Gain (WNG), leading to severe target signal cancellation. To ensure stable and robust beamforming, we previously proposed an adaptive diagonal loading method that leverages the Kantorovich inequality to guarantee the WNG remains strictly within specified bounds. However, accurately determining the smallest necessary loading level requires calculating the extreme eigenvalues of the spatial correlation matrix, a computationally expensive $\mathcal{O}(M^3)$ operation for large arrays. In this paper, we introduce a highly efficient $\mathcal{O}(kM^2)$ estimation technique using Lanczos iterations to build a small Krylov subspace. By projecting the correlation matrix onto a tridiagonal matrix of dimension $k \ll M$, we extract Ritz values that rapidly converge to the exact extreme eigenvalues. Our evaluations demonstrate that this Lanczos-accelerated approach achieves performance identical to exact Eigenvalue Decomposition (EVD), ensuring optimal interference suppression and strict WNG adherence at a fraction of the computational cost.
标题:SMC盲点:最先进节拍跟踪的故障模式分析
链接:https://arxiv.org/abs/2605.12287
备注:6 pages, 3 figures. Technical report on beat tracking failure modes; prepared for ISMIR 2026
摘要:在过去的二十年里,音乐节拍跟踪的任务已经从启发式起始点检测算法过渡到高性能的深度神经网络(DNN)。尽管基于DNN的节拍跟踪模型在主流的、抽象的数据集上实现了近乎完美的性能,但SMC数据集顽固地产生了较低的F-measure分数。通过测试最先进的模型在SMC数据集中检测单个轨道上的节拍的能力,我们确定了三种不同的故障模式:倍频程错误,连续性错误和完全跟踪故障,其中所有指标都低于0.3。我们发现,国家的最先进的模型往往会产生“自信,但错误的”激活。此外,我们表明,标准DBN的默认最低速度为55 BPM,防止它推断正确的速度为21%的SMC轨道,迫使双节奏的预测慢音乐。通过揭示这些基本的疏忽,我们为改进节拍和强拍检测提供了具体的方向,特别强调了训练数据的多样化和多假设节奏估计。
摘要:Over the past two decades, the task of musical beat tracking has transitioned from heuristic onset detection algorithms to highly capable deep neural networks (DNN). Although DNN-based beat tracking models achieve near-perfect performance on mainstream, percussive datasets, the SMC dataset has stubbornly yielded low F-measure scores. By testing how well state-of-the-art models detect beats on individual tracks in the SMC dataset, we identify three distinct failure modes: octave errors, continuity errors, and complete tracking failure where all metrics fall below 0.3. We reveal that state-of-the-art models tend to generate "confident-but-wrong" activations. Furthermore, we show that the standard DBN's default minimum tempo of 55 BPM prevents it from inferring the correct tempo for 21\% of SMC tracks, forcing double-tempo predictions on slow music. By exposing such fundamental oversights, we provide concrete directions for improving beat and downbeat detection, specifically emphasizing training data diversification and multi-hypothesis tempo estimation.
【2】Too Good to Be True: A Study on Modern Automatic Speech Recognition for the Evaluation of Speech Enhancement
标题:好得令人难以置信:用于语音增强评估的现代自动语音识别研究链接:https://arxiv.org/abs/2605.12107
摘要:语音增强(SE)系统通常使用各种仪器度量来评估。使用自动语音识别(ASR)系统来评估SE性能在文献中很常见,通常是在字错误率(WER)方面。然而,WER分数在很大程度上取决于ASR系统和文本规范化管道的选择。在本文中,我们研究了现代ASR模型与人类识别增强语音的相关性。听力实验表明,现代ASR模型与大规模的噪声训练和嵌入式语言模型与人类WER的相关性比简单的模型更高,换能器模型提供了最可靠的transmittance。然而,我们也表明,这些模型的鲁棒性噪声和使用的背景下,可以是没有信息的声学增强性能的评估为重点。
摘要:Speech enhancement (SE) systems are typically evaluated using a variety of instrumental metrics. The use of automatic speech recognition (ASR) systems to evaluate SE performance is common in literature, usually in terms of word error rate (WER). However, WER scores depend heavily on the choice of ASR system and text normalization pipeline. In this paper, we investigate how modern ASR models correlate with human recognition of enhanced speech. A listening experiment reveals that modern ASR models with large-scale noisy training and embedded language models correlate more with human WER than simpler ones, with a transducer model providing the most reliable transcriptions. Nevertheless, we also show that these models' robustness to noise and use of context can be uninformative to an acoustics-focused evaluation of enhancement performance.
【3】Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model
标题:迈向细粒度多维语音理解:数据管道、基准和模型链接:https://arxiv.org/abs/2605.12036
摘要:虽然语音大语言模型(LLM)在基本语音识别等传统任务中表现出色,但它们缺乏细粒度的多维感知。这一缺陷在他们努力解开复杂特征(如微声学线索、声学场景和非语言信号)时表现得很明显。这导致对真实世界语音的不完全理解从根本上阻碍了感知和移情的下一代语音系统的发展。在其核心,这种持续的感知限制主要源于三个相互作用的因素:稀缺的高质量的表达数据,缺乏多维属性的细粒度建模,以及依赖于有限的覆盖范围,粗粒度基准。我们通过三个支柱来应对这些挑战:首先,我们强大的数据管理管道解决了复杂的声学环境和长音频时间戳对齐挑战,从视听源中提取高质量的自发语音语料库。其次,我们构建了FMSU-Bench,这是一个涵盖14个语音属性维度的开创性基准,以严格评估当前模型的细粒度,多维语音理解能力。第三,通过我们策划的语料库,我们介绍了FM-Speech。在解耦属性建模和渐进式课程微调框架的驱动下,它大大提升了细粒度,多维的声学感知。对FMSU-Bench的广泛评估表明,当前的语音LLM仍然需要在多维,细粒度的理解方面进行显着改进。相比之下,FM-Speech大大优于当前的开源模型,为现实世界的语音理解建立了一个强大的范例。
摘要:While speech Large Language Models (LLMs) excel at conventional tasks like basic speech recognition, they lack fine-grained, multi-dimensional perception. This deficiency is evident in their struggle to disentangle complex features like micro-acoustic cues, acoustic scenes, and paralinguistic signals. This resulting incomplete comprehension of real-world speech fundamentally bottlenecks the development of perceptive and empathetic next-generation speech systems. At its core, this persistent perceptual limitation primarily stems from three interacting factors: scarce high-quality expressive data, absent fine-grained modeling for multi-dimensional attributes, and reliance on restricted coverage, coarse-grained benchmarks. We address these challenges through three pillars: First, our robust data curation pipeline resolves complex acoustic environments and long-audio timestamp alignment challenges to extract a high-quality spontaneous speech corpus from audiovisual sources. Second, we construct FMSU-Bench, a pioneering benchmark covering 14 speech attribute dimensions to rigorously assess the fine-grained, multi-dimensional speech understanding capabilities of current models. Third, empowered by our curated corpus, we introduce FM-Speech. Driven by a decoupled attribute modeling and progressive curriculum fine-tuning framework, it substantially elevates fine-grained, multi-dimensional acoustic perception. Extensive evaluations on FMSU-Bench reveal that current speech LLMs still require significant improvement in multi-dimensional, fine-grained understanding. In contrast, FM-Speech substantially outperforms current open-source models, establishing a robust paradigm for real-world speech understanding.
【4】Chunkwise Aligners for Streaming Speech Recognition
标题:流语音识别的Chunkwise对齐器链接:https://arxiv.org/abs/2605.11422
摘要:我们提出了Chunkwise Aligner,一种用于流式自动语音识别(ASR)的新架构。虽然Transducer是流式ASR的标准模型,但由于需要计算所有可能的音频标签对齐,因此其训练成本很高。最近引入的Aligner通过丢弃显式对齐来降低此成本,但此修改使其不适合流式传输。我们的方法通过将音频划分为块并将每个标签与其块的最左侧帧对齐来克服这一限制,而块之间的转换由学习的块结束概率来管理。实验表明,Chunkwise Aligner不仅在离线和流媒体场景中匹配Transducer的准确性,而且还提供了卓越的训练和解码效率。
摘要:We propose the Chunkwise Aligner, a novel architecture for streaming automatic speech recognition (ASR). While the Transducer is the standard model for streaming ASR, its training is costly due to the need to compute all possible audio-label alignments. The recently introduced Aligner reduces this cost by discarding explicit alignments, but this modification makes it unsuitable for streaming. Our approach overcomes this limitation by dividing the audio into chunks and aligning each label to the leftmost frames of its chunk, whereas transitions between chunks are managed by a learned end-of-chunk probability. Experiments show that the Chunkwise Aligner not only matches the Transducer's accuracy in both offline and streaming scenarios, but also offers superior training and decoding efficiencies.
【5】Adaptive Diagonal Loading using Krylov Subspaces for Robust Beamforming
标题:使用Krylov子空间的自适应对角加载用于鲁棒束形成链接:https://arxiv.org/abs/2605.11286
备注:5 pages, 8 figures
摘要:可靠的自适应波束形成对于在高动态声学环境中工作的大型麦克风阵列至关重要。在以快速移动的说话者和说话者为特征的场景中,用于估计空间相关矩阵的可用样本支持通常是快照不足的。这种缺陷降低了白噪声增益(WNG),导致严重的目标信号抵消。为了保证波束形成的稳定性和鲁棒性,我们之前提出了一种自适应对角加载方法,该方法利用Kantorovich不等式来保证WNG严格保持在指定的范围内。然而,准确地确定最小的必要负载水平需要计算空间相关矩阵的极端特征值,这对于大型阵列来说是一个计算昂贵的$\mathcal{O}(M^3)$操作。在本文中,我们介绍了一个高效的$\mathcal{O}(kM^2)$估计技术,使用Lanczos迭代,以建立一个小的Krylov子空间。通过将相关矩阵投影到维度为k \ll M$的三对角矩阵上,我们提取快速收敛到确切的极端特征值的Ritz值。我们的评估表明,这种Lanczos加速方法实现了与精确特征值分解(EVD)相同的性能,以一小部分计算成本确保了最佳干扰抑制和严格的WNG遵守。
摘要:Reliable adaptive beamforming is critical for large microphone arrays operating in highly dynamic acoustic environments. In scenarios characterized by fast-moving talkers and interferers, the available sample support for estimating the spatial correlation matrix is often snapshot-deficient. This deficiency degrades the White Noise Gain (WNG), leading to severe target signal cancellation. To ensure stable and robust beamforming, we previously proposed an adaptive diagonal loading method that leverages the Kantorovich inequality to guarantee the WNG remains strictly within specified bounds. However, accurately determining the smallest necessary loading level requires calculating the extreme eigenvalues of the spatial correlation matrix, a computationally expensive $\mathcal{O}(M^3)$ operation for large arrays. In this paper, we introduce a highly efficient $\mathcal{O}(kM^2)$ estimation technique using Lanczos iterations to build a small Krylov subspace. By projecting the correlation matrix onto a tridiagonal matrix of dimension $k \ll M$, we extract Ritz values that rapidly converge to the exact extreme eigenvalues. Our evaluations demonstrate that this Lanczos-accelerated approach achieves performance identical to exact Eigenvalue Decomposition (EVD), ensuring optimal interference suppression and strict WNG adherence at a fraction of the computational cost.
【6】STRUM: A Spectral Transcription and Rhythm Understanding Model for End-to-End Generation of Playable Rhythm-Game Charts
标题:SEARCH:用于端到端生成可玩节奏游戏图表的光谱转录和节奏理解模型链接:https://arxiv.org/abs/2605.12135
备注:9 pages, 4 figures, 3 tables. Code and models: https://github.com/
摘要:我们提出了一个音频到图表的管道,将原始录音转换为可播放的克隆英雄/ YARG图表,用于鼓,吉他,贝司,人声和键,而无需任何Oracle元数据。CNOM是一个多阶段的混合:一个两阶段的CRNN起始检测器和一个六模型集成分类器鼓;神经起始检测器与吉他和低音单声道音高跟踪;单词对齐ASR的人声;和光谱键盘检测键。我们在一个30首歌曲的信封基准上进行评估,该基准是通过一个单一的音频质量标准筛选候选歌曲而构建的-htdemucs_6s源分离后的中值1秒鼓干RMS。在此基准测试中,在每首歌曲全局偏移搜索的+/- 100 ms容差下,CNORM实现了鼓起始F1 = 0.838,低音F1 = 0.694,吉他F1 = 0.651,人声F1 = 0.539。我们报告了一个完整的消融7鼓管道组件配对每首歌Wilcoxon测试,地面实况音频定时分布在社区克隆英雄图表的分析,和每类的鼓分类混淆矩阵。代码、模型权重和完整的基准清单都已发布。
摘要:We present STRUM (Spectral Transcription and Rhythm Understanding Model), an audio-to-chart pipeline that converts raw recordings into playable Clone Hero / YARG charts for drums, guitar, bass, vocals, and keys without any oracle metadata. STRUM is a multi-stage hybrid: a two-stage CRNN onset detector and a six-model ensemble classifier for drums; neural onset detectors with monophonic pitch tracking for guitar and bass; word-aligned ASR for vocals; and spectral keyboard detection for keys. We evaluate on a 30-song in-envelope benchmark constructed by screening candidate songs on a single audio-quality criterion -- the median 1-second drum-stem RMS after htdemucs_6s source separation. On this benchmark STRUM achieves drums onset F1 = 0.838, bass F1 = 0.694, guitar F1 = 0.651, and vocals F1 = 0.539 at a +/- 100 ms tolerance with per-song global offset search. We report a complete ablation of seven drum-pipeline components with paired per-song Wilcoxon tests, an analysis of ground-truth-to-audio timing distributions in community Clone Hero charts, and a per-class confusion matrix for the drum classifier. Code, model weights, and the full benchmark manifest are released.
【7】Mixture-of-Experts Framework for Field-of-View Enhanced Signal-Dependent Binauralization of Moving Talkers
标题:移动说话者视野增强的信号相关双耳化专家混合框架链接:https://arxiv.org/abs/2509.13548
备注:5 pages, 3 figures
摘要:我们提出了一种新的混合专家框架的视场增强双耳信号匹配。我们的方法可以实现动态空间音频渲染,适应连续的说话者运动,允许用户强调或抑制来自选定方向的声音,同时保留自然的双耳线索。与依赖于显式到达方向估计或在高保真度立体声域中操作的传统方法不同,我们的信号依赖框架使用隐式定位以在线方式组合多个双耳滤波器。这允许实时跟踪和增强移动声源,支持增强和虚拟现实中的语音聚焦、降噪和世界锁定音频等应用。该方法与阵列几何形状无关,为下一代消费音频设备中的空间音频捕获和个性化回放提供了灵活的解决方案。
摘要:We propose a novel mixture of experts framework for field-of-view enhancement in binaural signal matching. Our approach enables dynamic spatial audio rendering that adapts to continuous talker motion, allowing users to emphasize or suppress sounds from selected directions while preserving natural binaural cues. Unlike traditional methods that rely on explicit direction-of-arrival estimation or operate in the Ambisonics domain, our signal-dependent framework combines multiple binaural filters in an online manner using implicit localization. This allows for real-time tracking and enhancement of moving sound sources, supporting applications such as speech focus, noise reduction, and world-locked audio in augmented and virtual reality. The method is agnostic to array geometry offering a flexible solution for spatial audio capture and personalized playback in next-generation consumer audio devices.
机器翻译由腾讯交互翻译提供,仅供参考
