今日论文合集:cs.SD语音6篇,eess.AS音频处理0篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
标题:EVA-Bench:评估语音代理的新端到端框架
链接:https://arxiv.org/abs/2605.13841
作者:Tara Bogavelli,Gabrielle Gauthier Melançon,Katrina Stankiewicz,Oluwanifemi Bamgbose,Fanny Riols,Hoang H. Nguyen,Raghav Mehndiratta,Lindsay Devon Brin,Joseph Marinier,Hari Subramani,Anil Madamala,Sridhar Krishna Nemala,Srinivas Sunkara
备注:Work in progress
摘要:语音代理是一种人工智能系统,可以进行口头对话以完成任务,越来越多地部署在企业应用程序中。然而,目前还没有一个基准能够共同解决两个核心评估挑战:生成逼真的模拟对话,以及在整个语音特定故障模式范围内测量质量。我们提出EVA-Bench,一个端到端的评估框架,解决这两个问题。在模拟方面,EVA-Bench通过动态多轮对话编排机器人对机器人的音频对话,并自动模拟验证,检测用户模拟器错误并在评分前适当地重新生成对话。在测量方面,EVA-Bench引入了两个复合指标:EVA-A(准确性),捕获任务完成,忠诚度和音频级别的语音保真度; EVA-X(体验),捕获会话进展,口语简洁性和话轮转换时间。这两个指标都适用于不同的代理架构,从而实现直接的跨架构比较。EVA-Bench包括跨三个企业域的213个场景,用于口音和噪声鲁棒性的受控扰动套件,以及区分峰值和可靠能力的pass@1,pass@k,pass^k测量。在跨越所有三种架构的12个系统中,我们发现:(1)没有系统在EVA-A pass@1和EVA-X pass@1上同时超过0.5;(2)峰值和可靠性能差异很大(EVA-A上的中值通过@k -通过^k间隙为0.44);以及(3)重音和噪声扰动暴露出相当大的鲁棒性差距,其影响在体系结构,系统,和度量(平均值高达0.314)。我们在开源许可证下发布完整的框架,评估套件和基准测试数据。
摘要:Voice agents, artificial intelligence systems that conduct spoken conversations to complete tasks, are increasingly deployed across enterprise applications. However, no existing benchmark jointly addresses two core evaluation challenges: generating realistic simulated conversations, and measuring quality across the full scope of voice-specific failure modes. We present EVA-Bench, an end-to-end evaluation framework that addresses both. On the simulation side, EVA-Bench orchestrates bot-to-bot audio conversations over dynamic multi-turn dialogues, with automatic simulation validation that detects user simulator error and appropriately regenerates conversations before scoring. On the measurement side, EVA-Bench introduces two composite metrics: EVA-A (Accuracy), capturing task completion, faithfulness, and audio-level speech fidelity; and EVA-X (Experience), capturing conversation progression, spoken conciseness, and turn-taking timing. Both metrics apply to different agent architectures, enabling direct cross-architecture comparison. EVA-Bench includes 213 scenarios across three enterprise domains, a controlled perturbation suite for accent and noise robustness, and pass@1, pass@k, pass^k measurements that distinguish peak from reliable capability. Across 12 systems spanning all three architectures, we find: (1) no system simultaneously exceeds 0.5 on both EVA-A pass@1 and EVA-X pass@1; (2) peak and reliable performance diverge substantially (median pass@k - pass^k gap of 0.44 on EVA-A); and (3) accent and noise perturbations expose substantial robustness gaps, with effects varying across architectures, systems, and metrics (mean up to 0.314). We release the full framework, evaluation suite, and benchmark data under an open-source license.


【2】NAACA: Training-Free NeuroAuditory Attentive Cognitive Architecture with Oscillatory Working Memory for Salience-Driven Attention Gating

标题:NAACA:免训练的神经听觉注意力认知结构,具有振荡工作记忆,用于突出性驱动的注意力门控
链接:https://arxiv.org/abs/2605.13651
作者:Zhongju Yuan,Geraint Wiggins,Dick Botteldooren
备注:Accepted as a regular paper by ICML 2026
摘要:音频提供了关键的情景线索,但目前的音频语言模型(ALM)面临着注意力瓶颈,在长形式的记录,占主导地位的背景模式可以稀释罕见的,突出的事件。我们介绍了NAACA,一个无训练的神经听觉注意认知架构,将注意力分配重新定义为听觉显著性过滤问题。其核心是OWM,一种神经启发的振荡工作记忆,它保持稳定的吸引子状态,只有当自适应能量波动信号感知显著性时才触发更高层次的认知ALM处理,从而触发更高层次的推理。在XD-Violence上,NAACA将AudioQwen的平均精度(AP)从53.50%提高到70.60%,同时减少了不必要的ALM调用。此外,对世界城市音景(USoW)数据集的定性案例研究表明,OWM捕获了新的事件和子类别的变化,同时对短暂的停顿和周围的城市噪音保持鲁棒性。
摘要:Audio provides critical situational cues, yet current Audio Language Models (ALMs) face an attention bottleneck in long-form recordings where dominant background patterns can dilute rare, salient events. We introduce NAACA, a training-free NeuroAuditory Attentive Cognitive Architecture that reframes attention allocation as an auditory salience filtering problem. At its core is OWM, a neuro-inspired Oscillatory Working Memory that maintains stable attractor-like states and triggers higher-cognition ALM processing only when adaptive energy fluctuations signal perceptual salience, triggering higher-level reasoning. On XD-Violence, NAACA improves AudioQwen's average precision (AP) from 53.50% to 70.60% while reducing unnecessary ALM invocations. Furthermore, qualitative case studies on the Urban Soundscapes of the World (USoW) dataset show that OWM captures novel events and subcategory shifts while remaining robust to transient pauses and ambient urban noise.


【3】Text2Score: Generating Sheet Music From Textual Prompts

标题:文本2Score:从文本注释生成乐谱
链接:https://arxiv.org/abs/2605.13431
作者:Keshav Bhandari,Sungkyun Chang,Abhinaba Roy,Francesca Ronchini,Emmanouil Benetos,Dorien Herremans,Simon Colton
备注:8 pages including references, 1 figure
摘要:由于对齐的文本音乐数据集的稀缺性和自动字幕管道的不可靠性,开发文本驱动的符号音乐生成模型仍然具有挑战性。虽然大多数努力都集中在乐谱上,但乐谱表示在文本驱动的生成中基本上没有得到充分的探索。我们提出了Text 2Score,一个两阶段的框架,包括规划阶段和执行阶段,用于从自然语言提示生成乐谱。通过直接从符号XML数据中获得监督信号,我们提出了一种替代的训练范式,绕过嘈杂或稀缺的文本音乐对。在规划阶段,一个LLM编曲翻译成一个结构化的测量明智的计划,定义音乐属性,如乐器,关键,时间签名,和声等,这个计划,然后消费的生成模型在执行阶段,以产生交错的ABC符号的计划的结构约束条件。为了评估输出质量,我们引入了一个评估框架,涵盖可玩性、可读性、乐器利用率、结构复杂性和及时遵守,并经过专家音乐家的验证。Text 2Score在客观和主观维度上始终优于纯基于LLM的代理框架和三个端到端基线。我们开源了本工作中使用的数据集,代码,评估集和LLM提示;在我们的项目页面上可以找到演示(https://keshavbhandari.github.io/portfolio/text2score)。
摘要:Developing text-driven symbolic music generation models remains challenging due to the scarcity of aligned text-music datasets and the unreliability of automated captioning pipelines. While most efforts have focused on MIDI, sheet music representations are largely underexplored in text-driven generation. We present Text2Score, a two-stage framework comprising a planning stage and an execution stage for generating sheet music from natural language prompts. By deriving supervision signals directly from symbolic XML data, we propose an alternative training paradigm that bypasses noisy or scarce text-music pairs. In the planning stage, an LLM orchestrator translates a natural language prompt into a structured measure-wise plan defining musical attributes such as instruments, key, time signatures, harmony, etc. This plan is then consumed by a generative model in the execution stage to produce interleaved ABC notation conditioned on the plan's structural constraints. To assess output quality, we introduce an evaluation framework covering playability, readability, instrument utilization, structural complexity, and prompt adherence, validated by expert musicians. Text2Score consistently outperforms both a pure LLM-based agentic framework and three end-to-end baselines across objective and subjective dimensions. We open-source the dataset, code, evaluation set and LLM prompts used in this work; a demo is available on our project page (https://keshavbhandari.github.io/portfolio/text2score).


【4】Seconds-Aligned PCA-DAC Latent Diffusion for Symbolic-to-Audio Drum Rendering

标题:秒对齐的PCA-DA潜在扩散用于符号到音频鼓渲染
链接:https://arxiv.org/abs/2605.13404
作者:Konstantinos Soiledis,Maximos Kaliakatsos Papakostas,Dimos Makris,Konstantinos Tsamis
摘要:符号控制鼓生成需要保留明确的事件时序和动态,同时合成声学上合理的波形。我们提出了Sec 2Drum-DAC,一个用于符号到音频鼓渲染的条件延迟扩散模型。该模型的条件下,在物理时间在编解码器帧位置采样的事件特征,并预测冻结的DAC求和码本嵌入,而不是波形样本的标准化主成分坐标。在评估的DAC配置中,72个主成分在规定的SVD阈值下捕获观察到的训练帧求和潜在子空间,从而在波形解码之前产生具有到1024维DAC潜在空间的确定性重建路径的紧凑连续去噪目标。   在1,733个四拍窗口中,PCA扩散改善了确定性PCA回归和符号渲染基线的成对频谱和瞬态度量,而直接回归在相位敏感波形L1上仍然更强。辅助RVQ交叉熵改善了mel误差、起始通量余弦和波形L1上的短步长扩散,根据度量,最有利的权衡发生在6-25个去噪步骤处。
摘要:Symbolic-control drum generation requires preserving explicit event timing and dynamics while synthesizing acoustically plausible waveforms. We present Sec2Drum-DAC, a conditional latent-diffusion model for symbolic-to-audio drum rendering. The model conditions on event features sampled in physical time at codec-frame locations and predicts standardized principal-component coordinates of frozen DAC summed-codebook embeddings rather than waveform samples. In the evaluated DAC configuration, 72 principal components capture the observed training-frame summed-latent subspace under the stated SVD threshold, yielding a compact continuous denoising target with a deterministic reconstruction path to the 1024-dimensional DAC latent space before waveform decoding.   Across 1,733 held-out four-beat windows, PCA diffusion improves paired spectral and transient metrics over deterministic PCA regression and a symbolic rendering baseline, while direct regression remains stronger on phase-sensitive waveform L1. Auxiliary RVQ cross-entropy improves short-step diffusion on mel error, onset-flux cosine, and waveform L1, with the most favorable trade-offs occurring at 6-25 denoising steps depending on the metric.


【5】Bypassing Direct Reconstruction: Speech Detection from MEG via Large-Scale Audio Retrieval

标题:绕过直接重建:通过大规模音频检索从MEG进行语音检测
链接:https://arxiv.org/abs/2605.13099
作者:Boda Xiao,Bo Wang,Heping Cheng
备注:ranked first at LibriBrain Competition 2025 https://neural-processing-lab.github.io/2025-libribrain-competition/prizes/
摘要:从非侵入性的大脑信号中解码语音是具有挑战性的。对于LibriBrain 2025语音检测任务,我们提出了一个新的两步框架,绕过了直接重建。首先,对比学习模型从大规模音频库(LibriVox)中检索给定测试MEG的匹配语音片段。其次,语音检测模型直接从该检索到的音频生成二进制静默/语音序列。通过这种方法,我们的团队Sherlock Holmes在扩展曲目中获得了第一名(F1分数:0.962),这表明利用外部音频数据库是一种非常有效的策略。
摘要:Decoding speech from non-invasive brain signals is challenging. For the LibriBrain 2025 Speech Detection task, we propose a novel two-step framework that bypasses direct reconstruction. First, a contrastive learning model retrieves the matching speech segment for the given test MEG from a large-scale audio library (LibriVox). Second, a speech detection model generates the binary silence/speech sequence directly from this retrieved audio. With this approach, our team Sherlock Holmes achieved first place in the extended track (F1-score: 0.962), demonstrating that leveraging external audio databases is a highly effective strategy.


【6】BioSEN: A Bio-acoustic Signal Enhancement Network for Animal Vocalizations

标题:BioSEN:用于动物发声的生物声学信号增强网络
链接:https://arxiv.org/abs/2605.12534
作者:Tianyu Song,Ton Viet Ta,Ngamta Thamwattana,Hisako Nomura,Linh Thi Hoai Nguyen
摘要:大多数音频增强工作针对人类语音,而生物声学由于嘈杂的录音和动物声音的独特特征而研究较少。为了填补这一空白,我们采用了语音增强方法,并建立了BioSEN,一个为生物声学信号制作的模型。BioSEN有三个模块:用于时频特征提取的多尺度双轴注意单元、用于捕获谐波结构的生物谐波多尺度增强单元、以及   能量自适应选通连接单元,其使用频率权重来防止发声作为噪声被去除。对三个生物声学数据集的测试表明,BioSEN匹配或超过了最先进的语音增强模型,同时使用更少的计算。这些结果显示了BioSEN在生物声学音频增强方面的实力及其对生物多样性监测和保护的承诺。
摘要:Most work in audio enhancement targets human speech, while bioacoustics is less studied due to noisy recordings and the distinct traits of animal sounds. To fill this gap, we adapt speech enhancement methods and build BioSEN, a model made for bioacoustic signals. BioSEN has three modules: a multi-scale dual-axis attention unit for time-frequency feature extraction, a bio-harmonic multi-scale enhancement unit for capturing harmonic structures, and an   energy-adaptive gating connection unit that uses frequency weights to keep vocalizations from being removed as noise. Tests on three bioacoustic datasets show that BioSEN matches or exceeds state-of-the-art speech enhancement models while using far less computation. These results show BioSEN's strength for bioacoustic audio enhancement and its promise for biodiversity monitoring and conservation.


机器翻译由腾讯交互翻译提供,仅供参考