今日论文合集:cs.SD语音12篇,eess.AS音频处理6篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Affective Music Recommendation: A Rollout-Based World Model for Offline Preference Optimization
标题:情感音乐推荐:线下偏好优化的基于推出的世界模型
链接:https://arxiv.org/abs/2605.28810
作者:Audrey Chan,Aaron Labbé,Jacob Lavoie,Jordan Bannister,Arsène Fansi Tchango,Guillaume Lajoie,Laurent Charlin
摘要:从消费者关注、睡眠辅助到临床干预等功能性音乐应用程序都面临着一个独特的推荐问题:成功取决于听众的情感状态,但在线情感实验受到道德约束,特别是对于无法可靠地跳过歌曲或报告痛苦的临床人群。我们描述了AMRS,即部署在LUCID健康和保健平台上的情感音乐推荐系统,该系统为临床用户(主要是患有神经认知疾病的老年人)和消费者健康用户提供能量,专注,平静和睡眠模式。AMRS是围绕一个基于推出的世界模型构建的:一个因果Transformer,在记录的听力数据上进行训练,以联合预测参与度、二进制评级以及自我报告的效价和唤醒。世界模型既可以作为离线策略培训的计算机模拟器,也可以作为部署前的压力测试工具。通过行为克隆初始化的推荐策略通过直接偏好优化(DPO)针对可配置的多目标效用函数进行离线微调。在严格的冷启动协议下,世界模型以可用的保真度预测行为和情感信号; DPO提高了克隆基线上的预测效价和唤醒,同时保持类似的多样性配置文件,并避免贪婪优化产生的分布崩溃。我们的位置的工作作为一个早期部署的情感推荐方法的验证时,在线实验是道德上站不住脚的。
摘要:Functional music applications, from consumer focus and sleep aids to clinical interventions, share a distinctive recommendation problem: success is defined by the listener's affective state, but online experimentation on emotion is ethically constrained, particularly for clinical populations who cannot reliably skip a song or report distress. We describe AMRS, the Affective Music Recommendation System deployed on LUCID's health-and-wellness platforms, which serve clinical users (primarily older adults with neurocognitive conditions) and consumer-wellness users across energize, focus, calm, and sleep modes. AMRS is built around a rollout-based world model: a causal transformer trained on logged listening data to jointly predict engagement, binary rating, and self-reported valence and arousal. The world model serves both as an in-silico simulator for offline policy training and as a stress-testing tool before deployment. A recommender policy initialized by behaviour cloning is fine-tuned offline with Direct Preference Optimization (DPO) against a configurable multi-objective utility function. Under a strict cold-start protocol, the world model predicts both behavioural and affective signals with usable fidelity; DPO improves predicted valence and arousal over the cloned baseline while maintaining a similar diversity profile and avoiding the distributional collapse produced by greedy optimization. We position the work as an early deployed validation of a methodology for affective recommendation when online experimentation is ethically untenable.


【2】Cross-modal characterization of infant cry: validation of a chest-surface accelerometer in extracting acoustic vocal function measures

标题:婴儿哭声的跨模式特征:胸面加速度计在提取声学发声功能测量中的验证
链接:https://arxiv.org/abs/2605.28687
作者:Winko W. An, Saketh Sundar, Lisa Yankowitz, Daryush D. Mehta, Carol L. Wilkinson
摘要:
摘要:


【3】DEMON: Diffusion Engine for Musical Orchestrated Noise

标题:DEMON:音乐干扰噪音的扩散引擎
链接:https://arxiv.org/abs/2605.28657
作者:Ryan Fosdick
备注:15 pages, 3 figures, 15 tables. Project page with audio samples and demo video: this https URL
摘要:
摘要:


【4】EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction

标题:EigeNet:基于几何学的多模式学习,用于少量新颖视图RIR预测
链接:https://arxiv.org/abs/2605.28101
作者:Chong Jing,Zitong Lan,Junan Zhang,Zhizheng Wu
备注:Code available on https://github.com/FEAfeatherTHER/EigeNet
摘要:从稀疏观测预测空间变化的房间脉冲响应(RIR)是沉浸式空间音频渲染的关键但极具挑战性的逆问题。在这项工作中,我们提出了EIGENET,几何知情的多模态框架Few-Shot新的视图RIR预测。其核心是一个跨视图交替注意力Transformer,迭代地细化局部视图内声学结构和全局跨视图空间关系。我们的经验表明,这种架构是能够充分利用多视图多模态上下文,同时进行时空推理的RIR预测。受声线跟踪的启发,我们设计了一个几何信息调制模块,以制定几何特征和RIR功率谱之间的联系。同时,引入辅助损失,将单目标波形预测转化为多任务学习框架。通过消融研究,我们证明了这种设计产生一致的性能增益,无论底层骨干,从而确认其基本的实用性和架构不可知的RIR预测任务的可推广性。在模拟和真实世界的基准测试中,EIGENET在Few-Shot新视图RIR预测和模拟到真实的泛化方面都达到了最先进的性能。代码和检查点可在https://github.com/FEAfeatherTHER/EigeNet上找到。
摘要:Predicting spatially varying Room Impulse Response (RIR) from sparse observations is a critical but highly challenging inverse problem for immersive spatial audio rendering. In this work, we present EIGENET, a geometry-informed multi-modal framework for few-shot novel view RIR prediction. At its core is a Cross-view Alternate-attention Transformer that iteratively refines local intra-view acoustic structures and global cross-view spatial relationships. We empirically demonstrate that this architecture is capable of making full use of the multi-view multi-modal context while performing spatial-temporal reasoning for RIR prediction. Inspired by acoustic ray tracing, we design a geometry-informed modulation block to formulate the connection between geometric features and RIR power spectrum. In the mean time, an auxiliary loss is introduced to transform the single-target waveform prediction into a multi-task learning framework. Through ablation studies, we demonstrate that this design yields consistent performance gains regardless of the underlying backbone, thereby confirming its foundational utility and architecture-agnostic generalizability for RIR prediction task. Evaluated on both simulated and real-world benchmarks, EIGENET achieves both state-of-the-art performance in few-shot novel view RIR prediction and sim-to-real generalization. Codes and checkpoints are available on https://github.com/FEAfeatherTHER/EigeNet.


【5】Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts

标题:从自由形式文本预设中统一合成语音和声音
链接:https://arxiv.org/abs/2605.28063
作者:Yuyue Wang,Xihua Wang,Xin Cheng,Yijing Chen,Ruihua Song
摘要:音频生成已经取得了重大进展,但合成语音和声音自然合成的统一音频仍然是一个挑战。目前的方法要么依赖于不相交的管道,无法捕获细粒度的交互,要么需要结构化的输入和外部文本重写,这限制了自由形式的文本提示的灵活性。在本文中,我们介绍了一个新的任务:自由形式的文本提示到统一的音频生成,其目的是直接合成统一的音频包含语音,声音,以及它们的组合从不受约束的自然语言。为了解决这个问题,我们提出了PlanAudio,一个统一的,基于自回归LLM的框架。首先,它通过利用固有的LLM推理能力而不是传统的文本编码器来简化模型架构。其次,它引入了一个语义潜在的思维链机制,一个隐含的规划机制,桥梁高层次的语义理解和低层次的声学合成。此外,我们还创建了PlanAudio-Bench,这是一个专门用于评估复合音频场景的基准。我们在语音、声音及其合成物的场景中进行评估。结果表明,PlanAudio通常优于现有的管道和统一基线,同时与为单一场景设计的模型保持竞争力。我们的分析进一步揭示了语义潜在CoT优于其他CoT机制,并强调了连续多场景培训课程的重要性。
摘要:Audio generation has made significant progress, yet synthesizing unified audio where speech and sounds are naturally composited remains a challenge. Current methods either rely on disjoint pipelines, which fail to capture fine-grained interactions, or require structured inputs and external text rewriting, which limits the flexibility of free-form text prompts. In this paper, we introduce a new task: Free-Form-Text-Prompt-to-Unified-Audio generation, which aims to directly synthesize unified audio containing speech, sound, and their composites from unconstrained natural language. To address this task, we propose PlanAudio, a unified, autoregressive LLM-based framework. First, it simplifies the model architecture by leveraging intrinsic LLM reasoning capability instead of traditional text encoders. Second, it introduces a semantic latent chain-of-thought mechanism, an implicit planning mechanism that bridges high-level semantic understanding and low-level acoustic synthesis. Furthermore, we create PlanAudio-Bench, a specialized benchmark for evaluating composite audio scenarios. We perform evaluations in the scenarios of speech, sound, and their composites. The results demonstrate that PlanAudio generally outperforms the existing pipeline and unified baselines, while staying competitive with models designed for a single scenario. Our analysis further reveals the superiority of semantic latent CoT over other CoT mechanisms and highlights the importance of continuous multi-scenario training curricula.


【6】MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation

标题:MTAVG-Bench 2.0:诊断多人音频视频生成中电影表现力的失败模式
链接:https://arxiv.org/abs/2605.28035
作者:Haitian Li,Yanghao Zhou,Heyan Huang,Liangji Chen,YiMing Cheng,Xu Liu,Dian Jin,Jiajun Xu,Jingyun Liao,Tian Lan,Ziqin Zhou,Yueying Liu,Yu Bai,Changsen Yuan,Jinxing Zhou,Xian-Ling Mao,Xuefeng Chen,Yousheng Feng
摘要:近年来,多说话者音频视频生成(MTAVG)模型在口型同步和视听对齐等基本指标上表现出了良好的性能。然而,这些指标仍然不足以评估场景级生成中的电影表现力。在多角色场景中,生成模型必须超越视听现实主义,以传达连贯的角色表演和其他更高层次的电影质量。为了填补这一空白,我们引入MTAVG-Bench 2.0,一个用于诊断多人音频视频生成中电影表现力故障模式的基准。与之前主要关注基本多回合对话质量的设置不同,MTAVG-Bench 2.0以短剧和场景级生成为目标,并建立了一个涵盖表演、叙事、氛围和视听语言的高级失败分类。基于此分类,我们构建了超过10,000个问答评估实例,以及用于短剧级别评估和故障模式时间定位的子集,以系统地评估Omni大型语言模型诊断高级视听故障的能力。实验结果表明,商业omni模型,如双子座大大优于其他评估,但即使是最强大的模型继续与复杂的故障在我们的基准斗争。这些结果表明,MTAVG-Bench 2.0提供了一个系统的基准故障诊断电影多说话者音频视频生成。
摘要:In recent years, Multi-Talker Audio-Video Generation (MTAVG) models have shown promising performance on fundamental metrics such as lip-sync and audio-visual alignment. However, these metrics remain insufficient for assessing cinematic expressiveness in scene-level generation. In multi-character scenes, generation models must go beyond audio-visual realism to convey coherent character performance and other higher-level cinematic qualities. To fill this gap, we introduce MTAVG-Bench 2.0, a benchmark for diagnosing failure modes of cinematic expressiveness in multi-talker audio-video generation. Unlike prior settings that mainly focus on the quality of basic multi-turn dialogue, MTAVG-Bench 2.0 targets short-drama and scene-level generation, and establishes a high-level failure taxonomy spanning acting, narrative, atmosphere, and audio-visual language. Based on this taxonomy, we construct more than 10,000 question-answering evaluation instances, together with subsets for short-drama-level assessment and temporal localization of failure modes, to systematically evaluate the ability of omni large language models to diagnose high-level audio-visual failures. Experimental results show that commercial omni models such as Gemini substantially outperform other evaluators, yet even the strongest models continue to struggle with complex failures in our benchmark. These results demonstrate that MTAVG-Bench 2.0 provides a systematic benchmark for failure diagnosis in cinematic multi-talker audio-video generation.


【7】VoiceGiraffe: A Benchmark for Extreme Long-Context Audio-Language Understanding

标题:VoiceGiraffe:极长上下文音频语言理解的基准
链接:https://arxiv.org/abs/2605.27976
作者:Jashin Ye,Dongxiao Wang,Yixuan Ye,Sashuai Zhou,Weihuang Lin,Mingyang Han,Kunpeng Wang,Zeyu Yuan,Boyu Li,Haoxiang Shi,Jingchen Shu,Jun Song,Bo Zheng
备注:Benchmark Project: https://github.com/LivingFutureLab/VoiceGiraffe
摘要:虽然大型音频语言模型(LALM)在秒级或分钟级音频处理方面取得了显着进展,但理解小时级音频仍然是一个根本瓶颈。现有的基准主要依赖于短片段或人工连接的片段,未能忠实地评估LALM在现实世界的情况下,如播客和冗长的演讲中的长距离信息理解能力。为了解决这一差距,我们引入了VoiceGiraffe,这是一种新颖的基准测试,旨在严格评估长上下文环境下各种真实场景、模式和语言的LALM。它包括1500个精心策划的三元组,它们被构造成单跳感知和多跳推理的双层分类。我们评估了一套广泛的开源和专有的LALM对人类的表现。结果强调了三个基本发现。首先,VoiceGiraffe仍然极具挑战性,远未饱和。其次,我们表明,没有一个单一的推理范式普遍占主导地位。E2E推理使具有本地长上下文音频理解的模型受益,级联字幕聚合稳定了被小时级音频淹没的小型模型,而与外部LLM的推理增强级联有助于较弱的模型,但可能会阻碍较强的专有系统。第三,我们揭示了远程内存持久性作为一个关键的瓶颈。LALM更擅长回答需要连接显著因果线索的问题,而不是那些需要在长音频中持续跟踪稀疏事件的问题,而人类则表现出相反的模式。这些发现将VoiceGiraffe定位为长形式音频理解的具有挑战性和诊断性的测试平台,强调了对具有持久记忆和强大的长距离聚合的LALM的需求。
摘要:While large audio language models (LALMs) have achieved remarkable progress in audio processing at the second- or minute-level scale, understanding hour-level audio remains a fundamental bottleneck. Existing benchmarks predominantly rely on short clips or artificially concatenated segments, failing to faithfully assess LALM capacity for long-range information comprehension in real-world scenarios such as podcasts and lengthy speeches. To address this gap, we introduce VoiceGiraffe, a novel benchmark designed to rigorously evaluate LALMs across diverse real-world scenarios, modalities, and languages under long-context settings. It comprises 1500 curated triplets structured into a dual-level taxonomy of single-hop perception and multi-hop reasoning. We evaluate a broad suite of open-source and proprietary LALMs against human performance. Results underscore three fundamental findings. First, VoiceGiraffe remains highly challenging and far from saturation. Second, we show that no single inference paradigm universally dominates. The E2E inference benefits models with native long-context audio understanding, cascaded caption aggregation stabilizes small models overwhelmed by hour-scale audio, and reasoning-enhanced cascading with external LLM helps weaker models but can bottleneck stronger proprietary systems. Third, we reveal long-range memory persistence as a key bottleneck. LALMs are better at answering questions that require connecting salient causal cues than those requiring sustained tracking of sparse events across long audio, whereas humans show the opposite pattern. These findings position VoiceGiraffe as a challenging and diagnostic testbed for long-form audio understanding, highlighting the need for LALMs with persistent memory and robust long-range aggregation.


【8】From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection

标题:从说话到唱歌:视听Deepfake检测的新挑战
链接:https://arxiv.org/abs/2605.27944
作者:Ke Liu,Jiwei Wei,Wenyu Zhang,Shuchang Zhou,Ruikun Chai,Yutao Dai,Chaoning Zhang,Yang Yang
备注:Accepted by ICML 2026
摘要:随着视听生成模型的快速发展,可靠的伪造检测变得越来越重要。现有的视听深度伪造检测方法通常依赖于跨模态不一致性。在歌唱,有节奏的发声削弱了这种耦合,并引入了一个非平凡的域转移,大大降低检测性能。我们使用节奏感知生成模型构建Singing Head DeepFake(SHDF)数据集,以填补歌唱基准测试中的空白。为了应对跨场景域的变化,我们提出了一个文本引导的视听伪造检测(T-AVFD)框架,概括了说话和唱歌的情况。T-AVFD包括面部真实性模式学习器和多模态微分权重学习模块。模式学习器将面部特征与多粒度文本描述对齐,以学习可概括的真实性模式。权重学习模块保留了固有的视听一致性,并通过差分加权自适应地将其与真实性模式相结合。对多个说话头深度伪造数据集和SHDF的广泛实验表明,与现有基线相比,在各种扰动下都有一致的改进和强大的鲁棒性。
摘要:With rapid advances in audio-visual generative models, reliable forgery detection becomes increasingly critical. Existing methods for audio-visual deepfake detection typically rely on cross-modal inconsistencies. In singing, rhythmic vocalization weakens this coupling and introduces a nontrivial domain shift, substantially degrading detection performance. We construct the Singing Head DeepFake (SHDF) dataset using rhythm-aware generative models to fill the gap in singing benchmarks. To cope with cross-scenario domain shifts, we propose a Text-guided Audio-Visual Forgery Detection (T-AVFD) framework that generalizes across both talking and singing scenarios. T-AVFD comprises a facial authenticity pattern learner and a multi-modal differential weight learning module. The pattern learner aligns facial features with multi-granularity textual descriptions to learn generalizable authenticity patterns. The weight learning module preserves intrinsic audio-visual consistency and adaptively integrates it with authenticity patterns via differential weighting. Extensive experiments on multiple talking head deepfake datasets and SHDF show consistent improvements over existing baselines and strong robustness under diverse perturbations.


【9】Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text

标题:大胜AudioGen:从文本生成连贯音频场景的统一模型
链接:https://arxiv.org/abs/2605.27838
作者:Jiahao Mei,Heinrich Dinkel,Yadong Niu,Xingwei Sun,Gang Li,Yifan Liao,Jiahao Zhou,Junbo Zhang,Jian Luan,Mengyue Wu
摘要:音频生成长期以来一直是碎片化的,语音,音乐和音效由特定领域的模型产生,无法从单个描述中联合生成连贯的音频场景。关键的障碍是对真实世界混合音频的细粒度监督不足,以及对并发音频组件建模的声学表示有限。我们提出了大声AudioGen,一个统一的框架,用于从文本中生成一般的混合音频场景。Dasheng AudioGen引入了结构化的多视图字幕,将复杂的声学场景显式地解耦为互补的描述视图,从而实现对音频层的细粒度控制。此外,我们采用了一个高维统一的语义声学表示作为共享的潜在空间。它注入语义先验,促进跨模态训练收敛,而其高维特征空间提供了足够的能力,有效地解开和融合并发音频分量。通过这些设计,简单的流匹配DiT实现了高质量的端到端音频场景生成。我们还建立了一个全面的评估管道音频场景生成。实验表明,大圣AudioGen在混合音频类别中实现了接近真实世界录音的性能,同时在单一类型生成任务中与专业模型保持竞争力。演示可在https://nieeim.github.io/Dasheng-AudioGen-Web/上获得。
摘要:Audio generation has long been fragmented, with speech, music, and sound effects produced by domain-specific models that fail to jointly generate coherent audio scenes from a single description. The key obstacles are insufficient fine-grained supervision for real-world mixed audio and limited acoustic representations for modeling concurrent audio components. We present Dasheng AudioGen, a unified framework for generating general mixed-audio scenes from text. Dasheng AudioGen introduces structured multi-view captions, which explicitly decouple complex acoustic scenes into complementary description views, thereby enabling fine-grained control over audio layers. Furthermore, we employ a high-dimensional unified semantic-acoustic representation as the shared latent space. It injects semantic priors that facilitate cross-modal training convergence, while its high-dimensional feature space provides sufficient capacity to disentangle and fuse concurrent audio components effectively. With these designs, a simple flow-matching DiT achieves high-quality end-to-end audio scene generation. We also establish a comprehensive evaluation pipeline for audio scene generation. Experiments demonstrate that Dasheng AudioGen achieves performance approaching real-world recordings in mixed-audio categories, while remaining competitive with specialized models in single-type generation tasks. Demos are available at https://nieeim.github.io/Dasheng-AudioGen-Web/.


【10】Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox

标题:音频LL是听还是读?使用VoxPartridge分析和缓解副语言障碍
链接:https://arxiv.org/abs/2605.27772
作者:Jiacheng Pang,Ashutosh Chaubey,Mohammad Soleymani
备注:Accepted as a conference paper at ICML 2026. Project page: https://voxparadox.github.io/
摘要:音频大语言模型(Audio LLM)在语音理解任务上表现出很强的性能,但它们理解非语言信息的能力仍然有限。为了系统地量化这一问题,我们引入了VoxPartial,这是一个对抗性的基准测试,有2,000个经过验证的例子,跨越10个非语言学任务,使用受控的语音合成来故意不匹配转录声明和说话风格,从而能够直接测量语音非语言学理解。对一组不同的音频LLM的评估显示,声学地面真相的准确性一直很低,并且有很强的倾向于遵循语言暗示(不正确)的答案。为了理解这种差距的原因,我们进行了逐层探测,发现(i)在更深的编码器层和编码器-LLM接口中,语言提示可能会降低,(ii)即使在音频令牌中有这样的提示,语言模型也经常忽略它们。为了解决这些问题,我们提出了提示条件层混合器(PCLM),它根据输入提示自适应地组合来自多个音频层的信息,并将其与直接偏好优化(DPO)配对,以显式地优先选择声学支持的选项,而不是语言隐含的替代方案。这些方法大大提高了Audio LLM的语言理解能力,将Audio Flamingo 3在VoxPartial上从17.40%提高到65.20%,在MMSU语言子集上从37.74%提高到54.78%。我们的项目页面可以在https://voxparadox.github.io/上找到。
摘要:Audio large language models (Audio LLMs) demonstrate strong performance on speech understanding tasks, yet their ability to understand paralinguistic information remains limited. To systematically quantify this issue, we introduce VoxParadox, an adversarial benchmark with 2,000 verified examples, spanning 10 paralinguistic tasks, created with controlled speech synthesis to intentionally mismatch transcript claims and speaking style, enabling direct measurement of speech paralinguistic understanding. Evaluation of a diverse set of Audio LLMs reveals consistently low accuracy on acoustic ground truth and a strong tendency to follow language-implied (incorrect) answers. To understand the cause of this gap, we perform layer-wise probing and find that (i) paralinguistic cues can degrade in deeper encoder layers and at the encoder--LLM interface, and (ii) even when such cues are available in audio tokens, the language model frequently ignores them. To address these problems, we propose Prompt-Conditioned Layer Mixer (PCLM), which adaptively combines information from multiple audio layers based on the input prompt, and pair it with Direct Preference Optimization (DPO) to explicitly prefer acoustically supported options over language-implied alternatives. These methods substantially improve Audio LLM paralinguistic understanding, improving Audio Flamingo 3 from 17.40% to 65.20% on VoxParadox, and from 37.74% to 54.78% on MMSU paralinguistic subset. Our project page is available at https://voxparadox.github.io/.


【11】Audio-Mind: An Auditable Agentic Framework for Audio Understanding

标题:Audio-Mind:音频理解的可审计抽象框架
链接:https://arxiv.org/abs/2605.28480
作者:Yucheng Wang,Jing Peng,Hanqi Li,Chenghao Wang,Wenming Tu,Yu Xi,Zhaokai Sun,Kai Yu,Shuai Wang
摘要:音频代理通过将音频问题分解为工具调用、中间证据和迭代推理步骤来扩展大型音频语言模型(LALM)。然而,随着LALM变得越来越强大,关键的挑战从启用工具的使用转移到确定代理证据获取何时真正有利于音频理解。我们提出了Audio-Mind,一个用于音频理解中的条件证据获取的可审计和可插入框架。Audio-Mind动态地将强大的前端与计划指导的工具使用相结合,在初始证据足够时保留前端判断,同时为未解决的证据缺口问题获取有界外部证据。在MMAR和MSU-Bench上的实验表明,Audio-Mind优于先前的音频代理基线,在MMAR上达到80.4%的准确率,在MSU-Bench上达到82.8%的准确率。一个匹配的主干比较突出了为什么这种设计很重要:在强大的音频前端下,当工作流不保留前端的整体音频基础判断时,代理分解可能成为编排瓶颈。除了准确性之外,Audio-Mind还可以生成更高质量、可审计的推理痕迹,这些痕迹可以揭示不确定性、工具证据和答案依据,为更可靠的音频QA注释和错误分析提供潜在基础。
摘要:Audio agents extend large audio-language models (LALMs) by decomposing audio questions into tool calls, intermediate evidence, and iterative reasoning steps. However, as LALMs become stronger, the key challenge shifts from enabling tool use to determining when agentic evidence acquisition genuinely benefits audio understanding. We propose Audio-Mind, an auditable and pluggable framework for conditional evidence acquisition in audio understanding. Audio-Mind dynamically combines a strong frontend with planner-guided tool use, preserving frontend judgment when initial evidence is sufficient while acquiring bounded external evidence for questions with unresolved evidence gaps. Experiments on MMAR and MSU-Bench show that Audio-Mind outperforms prior audio-agent baselines, reaching 80.4% accuracy on MMAR and 82.8% accuracy on MSU-Bench. A matched-backbone comparison highlights why this design matters: under strong audio frontends, agentic decomposition can become an orchestration bottleneck when the workflow does not preserve the frontend's holistic audio-grounded judgment. Beyond accuracy, Audio-Mind produces higher-quality, auditable reasoning traces that expose uncertainty, tool evidence, and answer rationales, offering a potential basis for more reliable audio-QA annotation and error analysis.


【12】LoSATok: Low-dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation

标题:LoSAtok:用于跨域音频理解和生成的低维语义声学令牌器
链接:https://arxiv.org/abs/2605.27840
作者:Zhisheng Zhang,Xiang Li,Yixuan Zhou,Jing Peng,Guoyang Zeng,Zhiyong Wu
摘要:音频标记器是统一音频理解和生成的基础。理解需要高级语义,而生成需要语义和声学细节。现有的统一标记器在高维连续潜伏期中联合编码两者,这增加了用于生成的扩散Transformers(DiT)的建模负担。我们提出了LoSA托克,一个低维的音频tokenizer跨域音频理解和生成。出于观察,1280维的语义编码器功能是可压缩的,我们引入了一个语义瓶颈,将它们压缩到128个维度,定期提出的时间关系损失的时间特征的一致性。我们进一步设计了一个双层次的语义监督方法,利用高,低维的语义信号,使标记器共同捕捉语义和声学细节在一个紧凑的潜在空间。对语音、音乐和普通音频的实验表明,SemBo保留了较强的低维语义能力,LoSA托克与几种语义表示相比保留了有竞争力的理解性能,同时不断提高语音、音乐和音频生成的DiT建模性能。这些结果表明,LoSA托克的低维表示可以有效地支持音频的理解和生成。我们的代码在https://github.com/wxzyd123/LoSATok上提供。
摘要:Audio tokenizers are fundamental to unifying audio understanding and generation. Understanding requires high-level semantics, while generation demands semantic and acoustic details. Existing unified tokenizers jointly encode both in high-dimensional continuous latents, which increases the modeling burden of Diffusion Transformers (DiTs) for generation. We propose LoSATok, a low-dimensional audio tokenizer for cross-domain audio understanding and generation. Motivated by the observation that 1280-dimensional semantic encoder features are compressible, we introduce a Semantic Bottleneck that compresses them into 128 dimensions, regularized by the proposed time-relation loss for temporal feature consistency. We further design a dual-level semantic supervision method that leverages both high- and low-dimensional semantic signals, enabling the tokenizer to jointly capture semantics and acoustic details within a compact latent space. Experiments on speech, music, and general audio show that SemBo preserves strong low-dimensional semantic capacity and LoSATok retains competitive understanding performance compared with several semantic representations, while consistently improving DiT modeling performance on speech, music, and audio generation. These results demonstrate that LoSATok's low-dimensional representations can effectively support audio understanding and generation. Our code is provided at https://github.com/wxzyd123/LoSATok.


eess.AS音频处理


【1】Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios
标题:不同场景下长格式语音生成的全面基准
链接:https://arxiv.org/abs/2605.28618
作者:Changhao Pan,Rui Yang,Han Wang,Zhuan Zhou,Xuming He,Wenxiang Guo,Ziyue Jiang,Ruiqi Li,Yu Zhang,Chenyuhao Wen,Ke Lei,Xiang Yin,Jingyu Lu,Zhiyuan Zhu,Zhou Zhao
备注:Accepted by ACL 2026(Findings). 36pages, 14figures
摘要:语音生成的最新进展,使高保真度的合成,但系统的评估模型在长期的背景条件下仍然在很大程度上未被探索。基于两个原因,一个针对长文本语音的全面评估基准是必不可少的:1)现有的测试场景通常局限于有限的领域,与下游的各种应用程序产生了巨大的差距; 2)现有的度量忽略了关键的长文本因素,如一致性和连贯性,无法可靠地概括。为此,我们提出了Swanbench-Speech,这是一个全面的基准,它将长格式的语音质量分解为特定的,分离的维度。SwanBench-Speech有三个关键属性。1)丰富的演讲场景:SwanBench-Speech专注于长格式语音生成和对话生成,涵盖声学、语义和表现力挑战,由17种常见语音场景的1,101个样本组成; 2)全面评估维度:沿着声学、语义和表现力轴,SwanBench-Speech定义了一个自动评估协议,具有七个度量标准,以提供全面、准确和标准化的评估; 3)有价值的见解:通过大量的实验,我们发现,目前的模型仍然在高度表达的情况下挣扎,并表现出显着的差距,在一致性和层次结构相比,真正的录音。
摘要:Recent advances in speech generation have enabled high-fidelity synthesis, yet systematic evaluation of models under long-context conditions remains largely underexplored. A comprehensive evaluation benchmark for long-form speech is indispensable for two reasons: 1) existing test scenarios are often confined to limited domains, creating a significant gap with the diverse downstream applications; 2) existing metrics overlook critical long-text factors such as consistency and coherence, failing to generalize reliably. To this end, we propose Swanbench-Speech, a comprehensive benchmark that decomposes long-form speech quality into specific, disentangled dimensions. SwanBench-Speech has three key properties. 1) Rich speech scenarios: Focusing on long-form speech generation and dialog generation, SwanBench-Speech covers acoustics, semantics, and expressiveness challenges, and consists of 1,101 samples spanning 17 common speech scenarios; 2) Comprehensive evaluation dimensions: Along the acoustics, semantics, and expressiveness axes, SwanBench-Speech defines an automated evaluation protocol with seven metrics to provide a comprehensive, accurate, and standardized assessment; 3) Valuable Insights: Through extensive experiments, we reveal that current models still struggle in highly expressive scenarios and exhibit a notable gap in consistency and hierarchy compared to real recordings.


【2】Audio-Mind: An Auditable Agentic Framework for Audio Understanding

标题:Audio-Mind:音频理解的可审计抽象框架
链接:https://arxiv.org/abs/2605.28480
作者:Yucheng Wang,Jing Peng,Hanqi Li,Chenghao Wang,Wenming Tu,Yu Xi,Zhaokai Sun,Kai Yu,Shuai Wang
摘要:音频代理通过将音频问题分解为工具调用、中间证据和迭代推理步骤来扩展大型音频语言模型(LALM)。然而,随着LALM变得越来越强大,关键的挑战从启用工具的使用转移到确定代理证据获取何时真正有利于音频理解。我们提出了Audio-Mind,一个用于音频理解中的条件证据获取的可审计和可插入框架。Audio-Mind动态地将强大的前端与计划指导的工具使用相结合,在初始证据足够时保留前端判断,同时为未解决的证据缺口问题获取有界外部证据。在MMAR和MSU-Bench上的实验表明,Audio-Mind优于先前的音频代理基线,在MMAR上达到80.4%的准确率,在MSU-Bench上达到82.8%的准确率。一个匹配的主干比较突出了为什么这种设计很重要:在强大的音频前端下,当工作流不保留前端的整体音频基础判断时,代理分解可能成为编排瓶颈。除了准确性之外,Audio-Mind还可以生成更高质量、可审计的推理痕迹,这些痕迹可以揭示不确定性、工具证据和答案依据,为更可靠的音频QA注释和错误分析提供潜在基础。
摘要:Audio agents extend large audio-language models (LALMs) by decomposing audio questions into tool calls, intermediate evidence, and iterative reasoning steps. However, as LALMs become stronger, the key challenge shifts from enabling tool use to determining when agentic evidence acquisition genuinely benefits audio understanding. We propose Audio-Mind, an auditable and pluggable framework for conditional evidence acquisition in audio understanding. Audio-Mind dynamically combines a strong frontend with planner-guided tool use, preserving frontend judgment when initial evidence is sufficient while acquiring bounded external evidence for questions with unresolved evidence gaps. Experiments on MMAR and MSU-Bench show that Audio-Mind outperforms prior audio-agent baselines, reaching 80.4% accuracy on MMAR and 82.8% accuracy on MSU-Bench. A matched-backbone comparison highlights why this design matters: under strong audio frontends, agentic decomposition can become an orchestration bottleneck when the workflow does not preserve the frontend's holistic audio-grounded judgment. Beyond accuracy, Audio-Mind produces higher-quality, auditable reasoning traces that expose uncertainty, tool evidence, and answer rationales, offering a potential basis for more reliable audio-QA annotation and error analysis.


【3】I Hear, Therefore I Trust: A Socio-Technical Investigation of Humans as Synthetic Speech Detectors

标题:我听到,故我相信:人类作为合成语音检测器的社会技术调查
链接:https://arxiv.org/abs/2605.28064
作者:Lelia Erscoi,Tomi Kinnunen
备注:To be included in Odyssey 2026: The Speaker and Language Recognition Workshop, Session 4.2, 23-26 June, Lisbon, Portugal
摘要:自动deepfake检测已经受到了相当大的研究关注,但人类实际遇到合成语音的社会技术环境仍然知之甚少。我们研究了语音deepfake检测作为一个感知和上下文的过程,提出了一个本地化任务,其中47名参与者在三个操纵的信任线索下标记了真实的,完全合成的和部分合成的话语中的可疑合成片段:教学框架,情感启动和出处标签。参与者提供了机械性,表现力,可理解性,清晰度,冷静度和评估信心的质量评级。话语类别是检测准确性和感知质量的主要决定因素;信任线索没有产生主要影响,但激发了检测行为。完全合成的语音被检测到低于机会的水平。质量评级跟踪话语类型,表明隐式歧视的公开检测失败。
摘要:Automatic deepfake detection has received considerable research attention, yet the socio-technical environment in which humans actually encounter synthetic speech remains poorly understood. We investigate voice deepfake detection as a perceptual and contextual process, presenting a localization task in which 47 participants marked suspected synthetic segments across authentic, fully synthetic, and partially synthetic utterances under three manipulated trust cues: instructional framing, affective priming, and provenance labeling. Participants provided quality ratings on mechanicalness, expressiveness, intelligibility, clarity, calmness, and confidence of evaluation. Utterance class was the primary determinant of detection accuracy and perceptual quality; trust cues produced no main effects but motivated detection behavior. Fully synthetic speech was detected at below-chance levels. Quality ratings tracked utterance type, indicating implicit discrimination where overt detection failed.


【4】LoSATok: Low-dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation

标题:LoSAtok:用于跨域音频理解和生成的低维语义声学令牌器
链接:https://arxiv.org/abs/2605.27840
作者:Zhisheng Zhang,Xiang Li,Yixuan Zhou,Jing Peng,Guoyang Zeng,Zhiyong Wu
摘要:音频标记器是统一音频理解和生成的基础。理解需要高级语义,而生成需要语义和声学细节。现有的统一标记器在高维连续潜伏期中联合编码两者,这增加了用于生成的扩散Transformers(DiT)的建模负担。我们提出了LoSA托克,一个低维的音频tokenizer跨域音频理解和生成。出于观察,1280维的语义编码器功能是可压缩的,我们引入了一个语义瓶颈,将它们压缩到128个维度,定期提出的时间关系损失的时间特征的一致性。我们进一步设计了一个双层次的语义监督方法,利用高,低维的语义信号,使标记器共同捕捉语义和声学细节在一个紧凑的潜在空间。对语音、音乐和普通音频的实验表明,SemBo保留了较强的低维语义能力,LoSA托克与几种语义表示相比保留了有竞争力的理解性能,同时不断提高语音、音乐和音频生成的DiT建模性能。这些结果表明,LoSA托克的低维表示可以有效地支持音频的理解和生成。我们的代码在https://github.com/wxzyd123/LoSATok上提供。
摘要:Audio tokenizers are fundamental to unifying audio understanding and generation. Understanding requires high-level semantics, while generation demands semantic and acoustic details. Existing unified tokenizers jointly encode both in high-dimensional continuous latents, which increases the modeling burden of Diffusion Transformers (DiTs) for generation. We propose LoSATok, a low-dimensional audio tokenizer for cross-domain audio understanding and generation. Motivated by the observation that 1280-dimensional semantic encoder features are compressible, we introduce a Semantic Bottleneck that compresses them into 128 dimensions, regularized by the proposed time-relation loss for temporal feature consistency. We further design a dual-level semantic supervision method that leverages both high- and low-dimensional semantic signals, enabling the tokenizer to jointly capture semantics and acoustic details within a compact latent space. Experiments on speech, music, and general audio show that SemBo preserves strong low-dimensional semantic capacity and LoSATok retains competitive understanding performance compared with several semantic representations, while consistently improving DiT modeling performance on speech, music, and audio generation. These results demonstrate that LoSATok's low-dimensional representations can effectively support audio understanding and generation. Our code is provided at https://github.com/wxzyd123/LoSATok.


【5】Diffusion Large Language Models for Visual Speech Recognition

标题:用于视觉语音识别的扩散大语言模型
链接:https://arxiv.org/abs/2605.28456
作者:Jeong Hun Yeo,Chae Won Kim,Hyeongseop Rha,Yong Man Ro
备注:Code: https://github.com/JeongHun0716/dllm-vsr
摘要:现有的视觉语音识别(VSR)系统通常依赖于从左到右的自回归解码,这可能会在足够的上下文可用之前强制对视觉上模糊的标记进行过早的决策。我们提出了DLLM-VSR,据我们所知,第一个基于扩散大语言模型(DLLM)的VSR框架,制定转录迭代掩蔽去噪与灵活的顺序解码。通过基于置信度的解蔽,DLLM-VSR提前提交高置信度的位置,并使用提交的令牌作为双向上下文来细化不明确的位置。为了使DLLM适应VSR,我们引入了一个两阶段的掩蔽去噪训练策略,该策略将视觉到文本内容对齐与长度建模分开。我们进一步观察到与oracle长度解码的性能差距,假设访问真实的转录本长度,这表明减少目标长度的不确定性可以提高基于DLLM的VSR。为了减少这一差距,我们开发了长度引导的候选人解码,它使用视频持续时间来构建合理的成绩单长度假设,在多个假设下解码,并使用长度可扩展性和解码置信度对候选人进行重新排序。该方法仅使用其标记的训练数据在LRS 3上实现了19.5\%的最新WER。
摘要:Existing Visual Speech Recognition (VSR) systems commonly rely on left-to-right autoregressive decoding, which can force premature decisions on visually ambiguous tokens before sufficient context is available. We propose DLLM-VSR, to the best of our knowledge, the first Diffusion Large Language Model (DLLM)-based VSR framework, formulating transcription as iterative masked denoising with flexible-order decoding. With confidence-based unmasking, DLLM-VSR commits high-confidence positions early and uses the committed tokens as bidirectional context to refine ambiguous ones. To adapt DLLMs to VSR, we introduce a two-stage masked-denoising training strategy that separates visual-to-text content alignment from length modeling. We further observe a performance gap with oracle-length decoding, which assumes access to the true transcript length, indicating that reducing target-length uncertainty can improve DLLM-based VSR. To reduce this gap, we develop length-guided candidate decoding, which uses video duration to construct plausible transcript-length hypotheses, decodes under multiple hypotheses, and reranks candidates using length plausibility and decoding confidence. The proposed method achieves a state-of-the-art WER of 19.5\% on LRS3 using only its labeled training data.


【6】MoDAl: Self-Supervised Neural Modality Discovery via Decorrelation for Speech Neuroprosthesis

标题:MoDAl:通过去相关的语音神经假体的自监督神经模态发现
链接:https://arxiv.org/abs/2605.00025
作者:Yuanhao Chen,Peter Chin
摘要:语音神经假体系统在没有听觉输出的情况下从神经活动中解码预期的语音,为有语音障碍的个体提供了恢复沟通的途径。目前的方法主要从运动皮层区域解码,丢弃其他区域-如44区,布罗卡区的一部分-可能编码补充语言信息。我们介绍了MoDAl(模态解相关和对齐),一个框架,发现互补的神经模态,通过两个目标的相互作用,在一个共享的投影空间。对比损失将几个并行大脑编码器中的每一个与预训练的大型语言模型(LLM)的文本嵌入对齐,而去相关损失则阻止编码器合并为重复表示。我们证明,这些目标是在生产的紧张局势:对比对齐诱导传递模态聚结,去相关必须抵消的框架,发现不同的神经语言学模式。在Brain-to-Text Benchmark '24上,与之前的最佳端到端方法相比,MoDAl将字错误率(WER)从26.3%降低到21.6%,其中来自合并之前丢弃的区域44信号的增益完全来自去相关机制。对所发现的模式的分析揭示了功能专业化:编码器接收区域44输入捕获结构和句法特性(句子长度,语法语音,wh-词),与布罗卡区的神经语言学理解一致。
摘要:Speech neuroprosthesis systems decode intended speech from neural activity in the absence of audible output, offering a path to restoring communication for individuals with speech-impairing conditions. Current approaches decode predominantly from motor cortical areas, discarding others -- such as area 44, part of Broca's area -- that may encode complementary linguistic information. We introduce MoDAl (Modality Decorrelation and Alignment), a framework that discovers complementary neural modalities through the interplay of two objectives in a shared projection space. A contrastive loss aligns each of several parallel brain encoders with the text embeddings of a pretrained large language model (LLM), while a decorrelation loss prevents the encoders from coalescing to duplicative representations. We prove that these objectives are in productive tension: Contrastive alignment induces transitive modality coalescence, which decorrelation must counteract for the framework to discover diverse neurolinguistic modalities. On the Brain-to-Text Benchmark '24, MoDAl reduces word error rate (WER) from 26.3% to 21.6% compared to the previous best end-to-end method, with the gain from incorporating previously discarded area 44 signals arising entirely from the decorrelation mechanism. Analysis of the discovered modalities reveals functional specialization: Encoders receiving area 44 input capture structural and syntactic properties (sentence length, grammatical voice, wh-words), consistent with the neurolinguistic understanding of Broca's area.


机器翻译由腾讯交互翻译提供,仅供参考