今日论文合集:cs.SD语音14篇,eess.AS音频处理8篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】A Manual Bar-by-Bar Tempo Measurement Protocol for Polyphonic Chamber Music Recordings: Design, Validation, and Application to Beethoven's Piano and Cello Sonatas
标题:复调室内乐录音的手动逐小节节奏测量协议:贝多芬钢琴和大提琴奏鸣曲的设计、验证和应用
链接:https://arxiv.org/abs/2604.15278
作者:Ignasi Sole
摘要:经验性的演奏分析依赖于从录音中准确提取速度数据,然而,为单声道音频或现代录音室条件设计的标准计算工具在应用于历史复调室内乐时系统地失败了。本文记录了贝多芬五首钢琴和大提琴奏鸣曲(Op.~)的二重奏录音中自动节拍检测软件的失败5个~ 1和~2; Op.~ 69; Op.~ 102个~ 1和~2),并提出了一个正式的手动替代方案:一个累积的lap-timer协议,产生的酒吧级每分钟的节拍数据与毫秒的分辨率。该协议是与一位专门从事超大规模集成电路设计的工程师跨学科合作开发的,它依赖于一个累积的时间戳架构,该架构可以防止错误累积,允许内部自验证,并捕获自动化工具系统抑制或误读的表现性时序现象(rubato,fermata,accelerandi,ritardandi)。BPM公式的数学推导,电子表格的数据结构,和误差特性的完整介绍。应用于1930- 2012年的100多个运动水平记录,该协议生成了一个数据集,随后通过tempographs,具有样条平滑概率密度函数的直方图,脊线图和组合图表进行可视化。本文认为,人工注释不是一种方法上的撤退,而是在面对复调历史记录的具体挑战时,对计算工具固有局限性的原则性回应。完整的数据集和分析代码是公开的。
摘要:Empirical performance analysis depends on the accurate extraction of tempo data from recordings, yet standard computational tools, designed for monophonic audio or modern studio conditions, fail systematically when applied to historical polyphonic chamber music. This paper documents the failure of automated beat-detection software on duo recordings of Beethoven's five piano and cello sonatas (Op.~5 Nos.~1 and~2; Op.~69; Op.~102 Nos.~1 and~2), and presents a formalised manual alternative: a cumulative lap-timer protocol that yields bar-level beats-per-minute data with millisecond resolution. The protocol, developed in cross-disciplinary collaboration with an engineer specialising in VLSI design, rests on a cumulative timestamp architecture that prevents error accumulation, permits internal self-validation, and captures expressive timing phenomena (rubato, fermatas, accelerandi, ritardandi) that automated tools systematically suppress or misread. The mathematical derivation of the BPM formula, the spreadsheet data structure, and the error characterisation are presented in full. Applied to over one hundred movement-level recordings spanning 1930--2012, the protocol generated a dataset subsequently visualised through tempographs, histograms with spline-smoothed probability density functions, ridgeline plots, and combination charts. The paper argues that manual annotation is not a methodological retreat but a principled response to the intrinsic limitations of computational tools when faced with the specific challenges of polyphonic historical recordings. The complete dataset and analysis code are publicly available.

【2】ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
标题:Control Foley:统一且可控的视频到音频生成,具有跨模式冲突处理
链接:https://arxiv.org/abs/2604.15086
作者:Jianxuan Yang,Xinyue Guo,Zhi Cheng,Kai Wang,Lipan Zhang,Jinjie Hu,Qiang Ji,Yihua Cao,Yihao Meng,Zhaoyue Cui,Mengmei Liu,Meng Meng,Jian Luan
摘要:视频到音频(V2 A)生成的最新进展使得能够从视觉内容合成高质量的音频,但实现鲁棒和细粒度的可控性仍然具有挑战性。现有的方法遭受弱的文本可控性下的视觉-文本冲突和不精确的风格控制,由于纠缠的时间和音色信息的参考音频。此外,缺乏标准化的基准限制了系统的评价。   我们提出了ControlFoley,一个统一的多模式V2 A框架,可以精确控制视频,文本和参考音频。我们引入了一种联合视觉编码范式,将CLIP与时空视听编码器集成在一起,以提高对齐和文本可控性。我们进一步提出了时间-音色解耦,以抑制冗余的时间线索,同时保留有区别的音色功能。此外,我们设计了一个模态鲁棒的训练计划,统一的多模态表示对齐(REPA)和随机模态丢弃。我们还提出了VGGSound-TVC,在不同程度的视觉-文本冲突下评估文本可控性的基准。   大量的实验表明,在多个V2 A任务,包括文本引导,文本控制和音频控制生成的最先进的性能。ControlFoley在跨模态冲突下实现了卓越的可控性,同时保持了强大的同步和音频质量,与工业V2 A系统相比,显示出具有竞争力或更好的性能。   代码、模型、数据集和演示可在https://yjx-research.github.io/ControlFoley/上获得。
摘要:Recent advances in video-to-audio (V2A) generation enable high-quality audio synthesis from visual content, yet achieving robust and fine-grained controllability remains challenging. Existing methods suffer from weak textual controllability under visual-text conflict and imprecise stylistic control due to entangled temporal and timbre information in reference audio. Moreover, the lack of standardized benchmarks limits systematic evaluation.   We propose ControlFoley, a unified multimodal V2A framework that enables precise control over video, text, and reference audio. We introduce a joint visual encoding paradigm that integrates CLIP with a spatio-temporal audio-visual encoder to improve alignment and textual controllability. We further propose temporal-timbre decoupling to suppress redundant temporal cues while preserving discriminative timbre features. In addition, we design a modality-robust training scheme with unified multimodal representation alignment (REPA) and random modality dropout. We also present VGGSound-TVC, a benchmark for evaluating textual controllability under varying degrees of visual-text conflict.   Extensive experiments demonstrate state-of-the-art performance across multiple V2A tasks, including text-guided, text-controlled, and audio-controlled generation. ControlFoley achieves superior controllability under cross-modal conflict while maintaining strong synchronization and audio quality, and shows competitive or better performance compared to an industrial V2A system.   Code, models, datasets, and demos are available at: https://yjx-research.github.io/ControlFoley/.

【3】From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench
标题:从被动到主动:通过ProVoice-Bench评估语音代理的主动性
链接:https://arxiv.org/abs/2604.15037
作者:Ke Xu,Yuhao Wang,Yu Wang
摘要:LLM代理的最新进展正在逐渐从反应性,基于文本的范式转向主动,多模式交互。然而,现有基准主要侧重于被动反应,忽视了主动干预和监测的复杂性。为了弥合这一差距,我们引入了ProVoice-Bench,这是第一个专门为主动语音代理设计的评估框架,具有四个新的任务。通过利用多阶段数据合成管道,我们为严格的测试精选了1,182个高质量样本。我们对最先进的多模态LLM的评估揭示了显着的性能差距,特别是在过触发和推理能力方面。这些发现突出了当前模型的局限性,并为开发更自然、更有上下文感知的主动代理提供了路线图。
摘要:Recent advancements in LLM agents are gradually shifting from reactive, text-based paradigms toward proactive, multimodal interaction. However, existing benchmarks primarily focus on reactive responses, overlooking the complexities of proactive intervention and monitoring. To bridge this gap, we introduce ProVoice-Bench, the first evaluation framework specifically designed for proactive voice agents, featuring four novel tasks. By leveraging a multi-stage data synthesis pipeline, we curate 1,182 high-quality samples for rigorous testing. Our evaluation of state-of-the-art Multimodal LLMs reveals a significant performance gap, particularly regarding over-triggering and reasoning capabilities. These findings highlight the limitations of current models and offer a roadmap for developing more natural, context-aware proactive agents.

【4】Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding
标题:听、理解和推理:基于感知的音频理解混合推理
链接:https://arxiv.org/abs/2604.14806
作者:Jieyi Wang,Yazhe Niu,Dexuan Xu,Zhongyu Wei
摘要:最近的大型音频语言模型在音频理解方面表现出了令人印象深刻的能力。然而,他们经常遭受感知错误,而可靠的音频推理是不可能实现的,如果没有首先在结构化的听觉场景中建立模型的感知。受听觉场景分析的启发,我们首先介绍了一个感知感知问题分类(PAQA)数据集。PAQA实现了一种分层解耦策略,将语音从环境声音中分离出来,并区分多个说话者,为训练提供明确的感知推理。在此基础上,我们提出了HyPeR,一个两阶段的混合感知推理框架。在第一阶段,我们对PAQA模型进行微调,以感知复杂音频中的声学属性。在第二阶段,我们利用GRPO来完善模型的内部审议。我们还引入了PAUSE令牌来促进声学模糊阶段的潜在计算,并设计感知一致性奖励来将推理原理与原始音频对齐。跨基准测试的实验表明,HyPeR实现了基础模型的绝对改进,性能与大规模模型相当,强调了混合感知接地推理的有效性,以实现强大的多扬声器音频理解。
摘要:Recent Large Audio Language Models have demonstrated impressive capabilities in audio understanding. However, they often suffer from perceptual errors, while reliable audio reasoning is unattainable without first grounding the model's perception in structured auditory scenes. Inspired by Auditory Scene Analysis, we first introduce a Perception-Aware Question Answering (PAQA) dataset. PAQA implements a hierarchical decoupling strategy that separates speech from environmental sound and distinguishes multiple speakers, providing explicit perceptual reasoning for training. Building on this, we propose HyPeR, a two-stage Hybrid Perception-Reasoning framework. In Stage I, we finetune the model on PAQA to perceive acoustic attributes in complex audio. In Stage II, we leverage GRPO to refine the model's internal deliberation. We also introduce PAUSE tokens to facilitate latent computation during acoustically ambiguous phases and design perceptual consistency reward to align reasoning rationales with raw audio. Experiments across benchmarks demonstrate that HyPeR achieves absolute improvements over the base model, with performance comparable to large-scale models, stressing the effectiveness of hybrid perception-grounded reasoning for robust and multi-speaker audio understanding.

【5】Geo2Sound: A Scalable Geo-Aligned Framework for Soundscape Generation from Satellite Imagery
标题:Geo 2Sound:一个可扩展的地理对齐框架,用于从卫星图像生成声景
链接:https://arxiv.org/abs/2604.14707
作者:Kunlin Wu,Yanning Wang,Haofeng Tan,Boyi Chen,Teng Fei,Xianping Ma,Yang Yue,Zan Zhou,Xiaofeng Liu
备注:15 pages, 4 figures, 4 tables. Includes supplementary material and SatSound-Bench dataset details
摘要:最近的图像到音频模型在以对象为中心的视觉场景上表现出令人印象深刻的性能。然而,它们在卫星图像中的应用仍然受到自上而下视图的复杂、广域语义模糊性的限制。虽然卫星图像提供了一个独特的可扩展的来源,全球声景生成,匹配这些意见,真正的声学环境与独特的空间结构是固有的困难。为了应对这一挑战,我们引入Geo 2Sound,一个新的任务和框架,用于从卫星图像生成地理上逼真的音景。具体来说,Geo 2Sound将结构地理空间属性建模、语义假设扩展和地声对齐结合在一个统一的框架中。轻量级分类器将头顶场景总结为紧凑的地理属性,多个面向声音的语义假设用于生成不同的声学上合理的候选者,并且地理声学对准模块将地理属性投影到声学嵌入空间中并识别与候选者集合最一致的候选者。此外,我们还建立了SatSound-Bench,这是第一个基准,包括超过20,000个高质量的成对卫星图像,文本描述和真实世界的音频记录,这些数据是从10多个国家收集的,并由三个公共数据集补充。实验表明,Geo 2Sound实现了1.765的SOTA FAD,比最强基线高出50.0%。人类评估进一步证实了现实主义(26.5%)和语义对齐的实质性收益,验证了我们的高保真度合成。项目页面和源代码:https://github.com/Blanketzzz/Geo2Sound
摘要:Recent image-to-audio models have shown impressive performance on object-centric visual scenes. However, their application to satellite imagery remains limited by the complex, wide-area semantic ambiguity of top-down views. While satellite imagery provides a uniquely scalable source for global soundscape generation, matching these views to real acoustic environments with unique spatial structures is inherently difficult. To address this challenge, we introduce Geo2Sound, a novel task and framework for generating geographically realistic soundscapes from satellite imagery. Specifically, Geo2Sound combines structural geospatial attributes modeling, semantic hypothesis expansion, and geo-acoustic alignment in a unified framework. A lightweight classifier summarizes overhead scenes into compact geographic attributes, multiple sound-oriented semantic hypotheses are used to generate diverse acoustically plausible candidates, and a geo-acoustic alignment module projects geographic attributes into the acoustic embedding space and identifies the candidate most consistent with the candidate sets. Moreover, we establish SatSound-Bench, the first benchmark comprising over 20k high-quality paired satellite images, text descriptions, and real-world audio recordings, collected from the field across more than 10 countries and complemented by three public datasets. Experiments show that Geo2Sound achieves a SOTA FAD of 1.765, outperforming the strongest baseline by 50.0%. Human evaluations further confirm substantial gains in both realism (26.5%) and semantic alignment, validating our high-fidelity synthesis on scale. Project page and source code: https://github.com/Blanketzzz/Geo2Sound

【6】ClariCodec: Optimising Neural Speech Codes for 200bps Communication using Reinforcement Learning
标题:ClariCodec:使用强化学习优化神经语音代码以实现200 Mbps通信
链接:https://arxiv.org/abs/2604.14654
作者:Junyi Wang,Chi Zhang,Jing Qian,Haifeng Luo,Hao Wang,Zengrui Jin,Chao Zhang
摘要:在卫星和水下信道等带宽受限的通信中,语音通常必须以超低比特率传输,其中可懂度是主要目标。在这种极端的压缩水平下,用声学重建损失训练的编解码器倾向于将比特分配给感知细节,导致字错误率(WER)的大幅下降。本文提出了ClariCodec,这是一种以200比特每秒(bps)的速度运行的神经语音编解码器,它将量化重新制定为随机策略,从而实现基于强化学习(RL)的可懂度优化。具体而言,编码器使用WER驱动的奖励进行微调,而声学重建管道保持冻结。即使没有RL,ClariCodec在200 bps的LibriSpeech测试干净集上也实现了3.68%的WER,已经与以更高比特率运行的编解码器竞争。进一步的RL微调将WER降低到3.20%(测试-干净)和8.93%(测试-其他),对应于13%的相对降低,同时保持感知质量。
摘要:In bandwidth-constrained communication such as satellite and underwater channels, speech must often be transmitted at ultra-low bitrates where intelligibility is the primary objective. At such extreme compression levels, codecs trained with acoustic reconstruction losses tend to allocate bits to perceptual detail, leading to substantial degradation in word error rate (WER). This paper proposes ClariCodec, a neural speech codec operating at 200 bit per second (bps) that reformulates quantisation as a stochastic policy, enabling reinforcement learning (RL)-based optimisation of intelligibility. Specifically, the encoder is fine-tuned using WER-driven rewards while the acoustic reconstruction pipeline remains frozen. Even without RL, ClariCodec achieves 3.68% WER on the LibriSpeech test-clean set at 200 bps, already competitive with codecs operating at higher bitrates. Further RL fine-tuning reduces WER to 3.20% on test-clean and 8.93% on test-other, corresponding to a 13% relative reduction while preserving perceptual quality.

【7】The Acoustic Camouflage Phenomenon: Re-evaluating Speech Features for Financial Risk Prediction
标题:声学Camerage现象:重新评估用于金融风险预测的语音特征
链接:https://arxiv.org/abs/2604.14619
作者:Dhruvin Dungrani,Disha Dungrani
摘要:在计算语言学中,从语音信号中检测认知负荷和欺骗是一个重要的研究领域。最近的努力试图将这些声学框架应用于公司盈利电话,以预测灾难性的股票市场波动。在这项研究中,我们经验性地研究了声学特征提取(音高,抖动和犹豫)的限制时,适用于训练有素的扬声器在野外电话会议环境。利用双流后期融合架构,我们将基于声学的流与基线自然语言处理(NLP)流进行对比。孤立的NLP模型对尾部风险下行事件的召回率为66.25%。令人惊讶的是,通过后期融合整合声学特征显着降低了性能,将召回率降低到47.08%。我们将这种退化确定为声学伪装,其中媒体训练的声音调节引入了矛盾的噪音,破坏了多模态元学习者。我们将这些发现作为语音处理在高风险金融预测中应用的边界条件。
摘要:In computational paralinguistics, detecting cognitive load and deception from speech signals is a heavily researched domain. Recent efforts have attempted to apply these acoustic frameworks to corporate earnings calls to predict catastrophic stock market volatility. In this study, we empirically investigate the limits of acoustic feature extraction (pitch, jitter, and hesitation) when applied to highly trained speakers in in-the-wild teleconference environments. Utilizing a two-stream late-fusion architecture, we contrast an acoustic-based stream with a baseline Natural Language Processing (NLP) stream. The isolated NLP model achieved a recall of 66.25% for tail-risk downside events. Surprisingly, integrating acoustic features via late fusion significantly degraded performance, reducing recall to 47.08%. We identify this degradation as Acoustic Camouflage, where media-trained vocal regulation introduces contradictory noise that disrupts multimodal meta-learners. We present these findings as a boundary condition for speech processing applications in high-stakes financial forecasting.

【8】Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection
标题:通过上下文不可知和不可感知的听觉提示注入劫持大型音频语言模型
链接:https://arxiv.org/abs/2604.14604
作者:Meng Chen,Kun Wang,Li Lu,Jiaheng Zhang,Tianwei Zhang
备注:Accepted by IEEE S&P 2026
摘要:现代大型音频语言模型(LALM)通过紧密集成音频和文本来支持智能语音交互。然而,这种集成将攻击面扩展到文本之外,并在连续的高维音频通道中引入漏洞。虽然之前的工作研究了音频越狱,但恶意音频注入和下游行为操纵的安全风险仍然没有得到充分研究。在这项工作中,我们揭示了一个以前被忽视的威胁,听觉提示注入,在现实的限制下,音频数据只访问和强大的感知隐形。为了系统地分析这种威胁,我们提出了\textit{AudioHijack},这是一个通用框架,可以生成与上下文无关的和不可感知的对抗性音频来劫持LALM。\textit{AudioHijack}采用基于采样的梯度估计,在不同模型之间进行端到端优化,绕过不可区分的音频标记化。通过注意力监督和多上下文训练,它将模型注意力转向对抗性音频,并推广到看不见的用户上下文。我们还设计了一种卷积混合方法,将扰动调制为自然混响,使用户高度感知不到。在13个最先进的LALM上进行的广泛实验显示,在6个不当行为类别中进行了一致的劫持,在高声学保真度的不可见用户上下文上实现了79%-96%的平均成功率。真实世界的研究表明,来自Mistral AI和Microsoft Azure的商业语音代理可以被诱导代表用户执行未经授权的操作。这些发现暴露了LALM的关键漏洞,并强调了对专门防御的迫切需要。
摘要:Modern Large audio-language models (LALMs) power intelligent voice interactions by tightly integrating audio and text. This integration, however, expands the attack surface beyond text and introduces vulnerabilities in the continuous, high-dimensional audio channel. While prior work studied audio jailbreaks, the security risks of malicious audio injection and downstream behavior manipulation remain underexamined. In this work, we reveal a previously overlooked threat, auditory prompt injection, under realistic constraints of audio data-only access and strong perceptual stealth. To systematically analyze this threat, we propose \textit{AudioHijack}, a general framework that generates context-agnostic and imperceptible adversarial audio to hijack LALMs. \textit{AudioHijack} employs sampling-based gradient estimation for end-to-end optimization across diverse models, bypassing non-differentiable audio tokenization. Through attention supervision and multi-context training, it steers model attention toward adversarial audio and generalizes to unseen user contexts. We also design a convolutional blending method that modulates perturbations into natural reverberation, making them highly imperceptible to users. Extensive experiments on 13 state-of-the-art LALMs show consistent hijacking across 6 misbehavior categories, achieving average success rates of 79\%-96\% on unseen user contexts with high acoustic fidelity. Real-world studies demonstrate that commercial voice agents from Mistral AI and Microsoft Azure can be induced to execute unauthorized actions on behalf of users. These findings expose critical vulnerabilities in LALMs and highlight the urgent need for dedicated defense.

【9】TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation
标题:TurboTalk:逐步蒸馏,实现音频驱动的会说话的阿凡达一代
链接:https://arxiv.org/abs/2604.14580
作者:Xiangyu Liu,Feng Gao,Xiaomei Zhang,Yong Zhang,Xiaoming Wei,Zhen Lei,Xiangyu Zhu
摘要:现有的音频驱动的视频数字人生成模型依赖于多步去噪,导致大量的计算开销,严重限制了它们在现实世界中的部署。虽然一步蒸馏方法可以显着加速推理,但它们通常会受到训练不稳定性的影响。为了解决这一挑战,我们提出了TurboTalk,一个两阶段的渐进式蒸馏框架,有效地压缩多步音频驱动的视频扩散模型到一个单步生成器。我们首先采用分布匹配蒸馏来获得一个强大而稳定的4步学生,然后通过对抗蒸馏逐步将去噪步骤从4个减少到1个。为了确保在极端步骤减少下的稳定训练,我们引入了渐进时间步采样策略和自比较对抗目标,该目标提供了一个中间对抗参考,可以稳定渐进蒸馏。我们的方法实现了视频说话化身的一步生成,推理速度提高了120倍,同时保持了高生成质量。
摘要:Existing audio-driven video digital human generation models rely on multi-step denoising, resulting in substantial computational overhead that severely limits their deployment in real-world settings. While one-step distillation approaches can significantly accelerate inference, they often suffer from training instability. To address this challenge, we propose TurboTalk, a two-stage progressive distillation framework that effectively compresses a multi-step audio-driven video diffusion model into a single-step generator. We first adopt Distribution Matching Distillation to obtain a strong and stable 4-step student, and then progressively reduce the denoising steps from 4 to 1 through adversarial distillation. To ensure stable training under extreme step reduction, we introduce a progressive timestep sampling strategy and a self-compare adversarial objective that provides an intermediate adversarial reference that stabilizes progressive distillation. Our method achieve single-step generation of video talking avatar, boosting inference speed by 120 times while maintaining high generation quality.

【10】VoxSafeBench: Not Just What Is Said, but Who, How, and Where
标题:VoxSafeBench:不仅仅是说了什么,还有谁、如何以及在哪里
链接:https://arxiv.org/abs/2604.14548
作者:Yuxiang Wang,Hongyu Liu,Yijiang Xu,Qinke Ni,Li Wang,Wan Lin,Kunyu Feng,Dekun Chen,Xu Tan,Lei Wang,Jie Shi,Zhizheng Wu
摘要:随着语音语言模型(SLM)从个人设备过渡到共享的多用户环境,它们的响应必须远远超过单词本身。谁在说话,他们的声音如何,以及对话发生在哪里,每一个都可以把一个原本善意的请求变成一个不安全,不公平或侵犯隐私的请求。然而,现有的基准主要集中在基本的音频理解上,孤立地研究个体风险,或者将固有有害的内容与仅由于其声学背景而变得有问题的内容混为一谈。我们介绍VoxSafeBench,在第一个基准,共同评估社会对准SLM在三个方面:安全性,公平性和隐私。VoxSafeBench采用两层设计:第一层使用匹配的文本和音频输入评估以内容为中心的风险,而第二层则针对音频条件风险,其中转录是良性的,但适当的响应取决于扬声器,语言提示或周围环境。为了验证Tier2,我们包括中间感知探头,并确认前沿SLM可以成功地检测到这些声学线索,但仍然无法对它们采取适当的行动。在22个双语覆盖的任务中,我们发现,在文本上表现出强大的保障措施往往会在语音中降低:安全意识下降,因为扬声器和场景条件的风险,当人口统计差异被口头传达时,公平性受到侵蚀,当上下文线索到达声学时,隐私保护就会动摇。总之,这些结果暴露了一个普遍的语音接地的差距:目前的SLM经常认识到相关的社会规范的文本,但不能应用它时,决定性的线索必须接地的讲话。代码和数据可在https://amphionteam.github.io/VoxSafeBench_demopage/上公开获取
摘要:As speech language models (SLMs) transition from personal devices into shared, multi-user environments, their responses must account for far more than the words alone. Who is speaking, how they sound, and where the conversation takes place can each turn an otherwise benign request into one that is unsafe, unfair, or privacy-violating. Existing benchmarks, however, largely focus on basic audio comprehension, study individual risks in isolation, or conflate content that is inherently harmful with content that only becomes problematic due to its acoustic context. We introduce VoxSafeBench, among the first benchmarks to jointly evaluate social alignment in SLMs across three dimensions: safety, fairness, and privacy. VoxSafeBench adopts a Two-Tier design: Tier1 evaluates content-centric risks using matched text and audio inputs, while Tier2 targets audio-conditioned risks in which the transcript is benign but the appropriate response hinges on the speaker, paralinguistic cues, or the surrounding environment. To validate Tier2, we include intermediate perception probes and confirm that frontier SLMs can successfully detect these acoustic cues yet still fail to act on them appropriately. Across 22 tasks with bilingual coverage, we find that safeguards appearing robust on text often degrade in speech: safety awareness drops for speaker- and scene-conditioned risks, fairness erodes when demographic differences are conveyed vocally, and privacy protections falter when contextual cues arrive acoustically. Together, these results expose a pervasive speech grounding gap: current SLMs frequently recognize the relevant social norm in text but fail to apply it when the decisive cue must be grounded in speech. Code and data are publicly available at: https://amphionteam.github.io/VoxSafeBench_demopage/

【11】Disentangled Dual-Branch Graph Learning for Conversational Emotion Recognition
标题:用于对话情绪识别的解纠缠双分支图学习
链接:https://arxiv.org/abs/2604.14204
作者:Chengling Guo,Yuntao Shou,Tao Meng,Wei Ai,Yun Tan,Keqin Li
备注:16 pages
摘要:会话中的多模态情感识别旨在通过在上下文中联合建模文本、声学和视觉线索来推断话语级情感。尽管最近取得了进展,关键的挑战仍然存在,包括冗余的跨模态信息,不完善的语义对齐,以及高阶扬声器交互的建模不足。为了解决这些问题,我们提出了一个框架,结合双空间特征解纠缠与双分支图学习。共享编码器和模态特定编码器用于分离模态不变表示和模态特定表示。不变特征由傅立叶图神经网络建模,以捕获全局一致性和互补模式,并具有频域对比目标以增强可区分性。并行地,在模态特定特征上构造说话者感知超图来建模高阶交互,以及说话者一致性约束来保持连贯的语义。最后,将两个分支融合用于话语级情感预测。IEMOCAP和MELD上的实验表明,该方法在强基线下具有更好的性能,验证了其有效性。
摘要:Multimodal emotion recognition in conversations aims to infer utterance-level emotions by jointly modeling textual, acoustic, and visual cues within context. Despite recent progress, key challenges remain, including redundant cross-modal information, imperfect semantic alignment, and insufficient modeling of high-order speaker interactions. To address these issues, we propose a framework that combines dual-space feature disentanglement with dual-branch graph learning. A shared encoder and modality-specific encoders are used to separate modality-invariant and modality-specific representations. The invariant features are modeled by a Fourier graph neural network to capture global consistency and complementary patterns, with a frequency-domain contrastive objective to enhance discriminability. In parallel, a speaker-aware hypergraph is constructed over modality-specific features to model high-order interactions, along with a speaker-consistency constraint to maintain coherent semantics. Finally, the two branches are fused for utterance-level emotion prediction. Experiments on IEMOCAP and MELD demonstrate that the proposed method achieves superior performance over strong baselines, validating its effectiveness.

【12】From Black Box to Glass Box: Cross-Model ASR Disagreement to Prioto Review in Ambient AI Scribe Documentation
标题:从黑匣子到玻璃盒:跨模型ASB对Ambient AI Scribe文档中Prioto审查的分歧
链接:https://arxiv.org/abs/2604.14152
作者:Abdolamir Karbalaie,Fernando Seoane,Farhad Abtahi
摘要:环境人工智能“抄写”系统有望减少临床文档负担,但如果不仔细审查,自动语音识别(ASR)错误可能会被忽视,并且高质量的人类参考转录通常无法用于校准不确定性。我们研究了异构ASR系统之间的跨模型分歧是否可以作为一个无参考的不确定性信号,以优先考虑医疗转录工作流程中的人类验证。使用50个公开的医学教育音频片段(8小时14分钟),我们用8个ASR系统转录了每个片段,这些系统涵盖了商业API和开源引擎。我们对齐多模型输出,建立共识伪参考,并使用多数强度度量量化标记级协议;我们进一步通过类型(内容与标点符号/格式)表征分歧,并通过留一模型(刀切)共识评分评估每个模型的协议。模型间可靠性较低(ICC[2,1] = 0.131),表明系统间存在异质失效模式。在76,398个被评估的代币头寸中,72.1%显示出近乎一致的一致性(7-8个模型),而2.5%属于高风险范围(0-3个模型),高风险质量在口音组中的比例从0.7%到11.4%不等。低一致性区域因内容不一致而丰富,在高风险质量的五分位数中,内容分数从53.9%增加到73.9%。这些结果表明,跨模型不一致提供了一个稀疏的,可定位的信号,可以在没有人类验证的参考的情况下显示潜在不可靠的转录本跨度,从而实现有针对性的审查;标记区域的临床准确性仍有待建立。
摘要:Ambient AI "scribe" systems promise to reduce clinical documentation burden, but automatic speech recognition (ASR) errors can remain unnoticed without careful review, and high-quality human reference transcripts are often unavailable for calibrating uncertainty. We investigate whether cross-model disagreement among heterogeneous ASR systems can act as a reference-free uncertainty signal to prioritize human verification in medical transcription workflows. Using 50 publicly available medical education audio clips (8 h 14 min), we transcribed each clip with eight ASR systems spanning commercial APIs and open-source engines. We aligned multi-model outputs, built consensus pseudo-references, and quantified token-level agreement using a majority-strength metric; we further characterized disagreements by type (content vs. punctuation/formatting) and assessed per-model agreement via leave-one-model-out (jackknife) consensus scoring. Inter-model reliability was low (ICC[2,1] = 0.131), indicating heterogeneous failure modes across systems. Across 76,398 evaluated token positions, 72.1% showed near-unanimous agreement (7-8 models), while 2.5% fell into high-risk bands (0-3 models), with high-risk mass varying from 0.7% to 11.4% across accent groups. Low-agreement regions were enriched for content disagreements, with the content fraction increasing from 53.9% to 73.9% across quintiles of high-risk mass. These results suggest that cross-model disagreement provides a sparse, localizable signal that can surface potentially unreliable transcript spans without human-verified references, enabling targeted review; clinical accuracy of flagged regions remains to be established.

【13】Enhancing time-frequency resolution with optimal transport and barycentric fusion of multiple spectrogram
标题:通过多个谱图的最佳传输和重心融合增强时频分辨率
链接:https://arxiv.org/abs/2604.15055
作者:David Valdivia,Elsa Cazelles,Cédric Févotte
备注:main text: 13 pages, 8 figures. supplementary material: 3 pages, 3 figures
摘要:时频表示,如短时傅立叶变换(STFT),是分析非平稳信号的基本工具。然而,它们在时间和频率上实现尖锐定位的能力本质上受到Gabor-Heisenberg不确定性原理的限制。在本文中,我们通过引入一种方法来产生超分辨率光谱图,通过融合两个或多个光谱图与不同的分辨率来解决这个限制。具体来说,我们使用最优传输(OT)发散计算超分辨率谱图作为输入谱图的重心。与现有的融合方法不同,我们的方法不需要输入谱图共享相同的时频网格。相反,可以使用任何STFT参数计算输入频谱图,并且可以在任意用户指定的网格上定义所得到的超分辨率频谱图。我们探讨了基于不同运输成本的OT差异。值得注意的是,我们引入了一种新的运输成本,保留时间-频率几何,同时显着降低计算复杂度相比,标准Wasserstein重心。我们采用不平衡OT框架,并推导出一个新的块优化-最小化算法的有效重心计算。我们验证所提出的方法控制合成信号和录音语音使用定量和定性评价。结果表明,我们的方法结合了最好的本地化特性的输入频谱图,并优于一个无监督的国家的最先进的融合方法。
摘要:Time-frequency representations, such as the short-time Fourier transform (STFT), are fundamental tools for analyzing non-stationary signals. However, their ability to achieve sharp localization in both time and frequency is inherently limited by the Gabor-Heisenberg uncertainty principle. In this paper, we address this limitation by introducing a method to generate super-resolution spectrograms through the fusion of two or more spectrograms with varying resolutions. Specifically, we compute the super-resolution spectrogram as the barycenter of input spectrograms using optimal transport (OT) divergences. Unlike existing fusion approaches, our method does not require the input spectrograms to share the same time-frequency grid. Instead, the input spectrograms can be computed using any STFT parameters, and the resulting super-resolution spectrogram can be defined on an arbitrary user-specified grid. We explore various OT divergences based on different transportation costs. Notably, we introduce a novel transportation cost that preserves time-frequency geometry while significantly reducing computational complexity compared to standard Wasserstein barycenters. We adopt the unbalanced OT framework and derive a new block majorization-minimization algorithm for efficient barycenter computation. We validate the proposed method on controlled synthetic signals and recorded speech using both quantitative and qualitative evaluations. The results show that our approach combines the best localization properties of the input spectrograms and outperforms an unsupervised state-of-the-art fusion method.

【14】Gaussian Process Regression of Steering Vectors With Physics-Aware Deep Composite Kernels for Augmented Listening
标题:具有物理感知深度复合核的引导载体的高斯过程回归以增强听力
链接:https://arxiv.org/abs/2509.02571
作者:Diego Di Carlo,Shoichi Koyama,Nugraha Aditya Arie,Fontaine Mathieu,Bando Yoshiaki,Yoshii Kazuyoshi
摘要:本文研究了用于增强收听的引导向量在频率和麦克风/源位置上的连续表示(例如,空间滤波和双耳渲染),使得能够对再现声场进行用户参数化控制。导向矢量通常用于表示作为查找方向的函数的麦克风阵列的空间响应。这些量的基本代数表示假设一个理想化的环境不能处理声场的散射效应。因此,可以收集在专用设施中测量的实际导向矢量的离散集合,并且超分辨(即,upsample)。最近,物理感知的深度学习方法已有效地用于此目的。然而,这种确定性的超分辨率,遭受过拟合问题,由于测量空间上的非均匀的不确定性。为了解决这个问题,我们将基于神经场(NF)的表达表示集成到基于高斯过程(GP)的原则概率框架中。具体来说,我们提出了一个物理感知的复合内核,模型的定向入射波和随后的散射效果。通过综合对比实验,验证了该方法在数据不足情况下的有效性.在下游任务中,如语音增强和双耳渲染,使用SPECTRA挑战的模拟数据,预言性能达到不到十倍的测量。
摘要:This paper investigates continuous representations of steering vectors over frequency and microphone/source positions for augmented listening (e.g., spatial filtering and binaural rendering), enabling user-parameterized control of the reproduced sound field. Steering vectors have typically been used for representing the spatial response of a microphone array as a function of the look-up direction. The basic algebraic representation of these quantities assuming an idealized environment cannot deal with the scattering effect of the sound field. One may thus collect a discrete set of real steering vectors measured in dedicated facilities and super-resolve (i.e., upsample) them. Recently, physics-aware deep learning methods have been effectively used for this purpose. Such deterministic super-resolution, however, suffers from the overfitting problem due to the non-uniform uncertainty over the measurement space. To solve this problem, we integrate an expressive representation based on the neural field (NF) into the principled probabilistic framework based on the Gaussian process (GP). Specifically, we propose a physics-aware composite kernel that models the directional incoming waves and the subsequent scattering effect. Our comprehensive comparative experiment showed the effectiveness of the proposed method under data insufficiency conditions. In downstream tasks such as speech enhancement and binaural rendering using the simulated data of the SPEAR challenge, the oracle performances were attained with less than ten times fewer measurements.


eess.AS音频处理


【1】UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations
标题:UniPASE:高保真和低幻觉通用语音增强生成模型
链接:https://arxiv.org/abs/2604.14606
作者:Xiaobin Rong,Zheng Wang,Yushi Wang,Jun Gao,Jing Lu
备注:Submitted to IEEE TASLP
摘要:通用语音增强(USE)的目的是恢复语音信号从不同的失真在多个采样率。我们提出了UniPASE,为使用量身定制的低幻觉PASE框架的扩展。其核心是DeWavLM-Omni,这是一个统一的表示级增强模块,通过对大规模监督多失真数据集的知识蒸馏从WavLM进行微调。该模块直接将降级的波形转换为干净且语言上忠实的语音表示,确保以最小的语言幻觉进行鲁棒增强。基于这些增强的语音表示,适配器生成包含丰富声学细节的增强的声学表示,神经声码器使用该增强的声学表示来重建对应的高保真16 kHz波形。然后,PostNet将波形转换为48 kHz,然后将其重新转换为原始速率,从而能够以多种采样速率无缝处理输入和输出。在多个评估数据集上的实验结果表明,与现有的最先进的模型相比,UniPASE实现了更好的或有竞争力的性能。该模型也是我们提交2026年紧急挑战的基础,该挑战在客观评估中获得第一名。源代码和音频演示可以在https://github.com/xiaobin-rong/unipase/上获得。
摘要:Universal speech enhancement (USE) aims to restore speech signals from diverse distortions across multiple sampling rates. We propose UniPASE, an extension of the low-hallucination PASE framework tailored for USE. At its core is DeWavLM-Omni, a unified representation-level enhancement module fine-tuned from WavLM via knowledge distillation on a large-scale supervised multi-distortion dataset. This module directly converts degraded waveforms into clean and linguistically faithful phonetic representations, ensuring robust enhancement with minimal linguistic hallucination. Based on these enhanced phonetic representations, an Adapter generates enhanced acoustic representations containing rich acoustic details, which a neural Vocoder uses to reconstruct corresponding high-fidelity 16-kHz waveforms. A PostNet then converts the waveforms to 48~kHz before resampling them to their original rates, enabling seamless handling of inputs and outputs at multiple sampling rates. Experimental results on several evaluation datasets, covering sub-tasks and full tasks, demonstrate that UniPASE achieves superior or competitive performance compared with existing state-of-the-art models. The proposed model also serves as the backbone of our submission to the URGENT 2026 Challenge, which achieved 1st place in the objective evaluation. The source code and audio demos are available at https://github.com/xiaobin-rong/unipase/.

【2】Who is Speaking or Who is Depressed? A Controlled Study of Speaker Leakage in Speech-Based Depression Detection
标题:谁在说话,谁在沮丧?基于语音的抑郁检测中说话人泄漏的对照研究
链接:https://arxiv.org/abs/2604.14354
作者:Hsiang-Chen Yeh,Luqi Sun,Aurosweta Mahapatra,Shreeram Suresh Chandra,Emily Mower Provost,Berrak Sisman
备注:Submitted to Interspeech 2026
摘要:这项研究调查是否基于语音的抑郁症检测模型学习抑郁症相关的声学生物标志物,而是依赖于扬声器的身份线索。使用DAIC-WOZ数据集,我们提出了一种数据分割策略,该策略控制训练集和测试集之间的说话人重叠,同时保持训练大小恒定,并评估三种不同复杂度的模型。结果表明,扬声器重叠显着提高性能,而准确性急剧下降,看不见的扬声器。即使使用域对抗神经网络,仍然存在很大的性能差距。这些发现表明,抑郁相关的功能提取当前的语音模型是高度纠缠与说话人身份。因此,传统的评估协议可能会高估泛化和临床实用性,强调需要严格的说话人独立的评估。
摘要:This study investigates whether speech-based depression detection models learn depression-related acoustic biomarkers or instead rely on speaker identity cues. Using the DAIC-WOZ dataset, we propose a data-splitting strategy that controls speaker overlap between training and test sets while keeping the training size constant, and evaluate three models of varying complexity. Results show that speaker overlap significantly boosts performance, whereas accuracy drops sharply on unseen speakers. Even with a Domain-Adversarial Neural Network, a substantial performance gap remains. These findings indicate that depression-related features extracted by current speech models are highly entangled with speaker identity. Conventional evaluation protocols may therefore overestimate generalization and clinical utility, highlighting the need for strictly speaker-independent evaluation.

【3】HARNESS: Lightweight Distilled Arabic Speech Foundation Models
标题:HARNESS:轻量级蒸馏阿拉伯语基础模型
链接:https://arxiv.org/abs/2604.14186
作者:Vrunda N. Sukhadia,Shammur Absar Chowdhury
备注:8 pages, 2 figures
摘要:大型自监督语音(SSL)模型实现了强大的下游性能,但它们的大小限制了在资源受限环境中的部署。我们提出了HArnESS,一个以阿拉伯语为中心的自监督语音模型家族,通过迭代自蒸馏从头开始训练,以及轻量级的学生变体,这些变体在自动语音识别(ASR),方言识别(DID)和语音情感识别(SER)上提供了强大的准确性-效率权衡。我们的方法从一个大型的阿拉伯语-英语双语教师开始,逐步将其知识提炼成压缩的学生模型,同时保留阿拉伯语相关的声学和非语言表示。进一步研究了基于PCA的教师监督信号的压缩,以更好地匹配浅瘦学生的容量。与HuBERT和XLS-R相比,HArnESS始终提高了阿拉伯语下游任务的性能,而压缩模型在大量结构简化下仍具有竞争力。这些结果将HArnESS定位为现实世界语音应用的实用且可访问的以阿拉伯语为中心的SSL基础。
摘要:Large self-supervised speech (SSL) models achieve strong downstream performance, but their size limits deployment in resource-constrained settings. We present HArnESS, an Arabic-centric self-supervised speech model family trained from scratch with iterative self-distillation, together with lightweight student variants that offer strong accuracy-efficiency trade-offs on Automatic Speech Recognition (ASR), Dialect Identification (DID), and Speech Emotion Recognition (SER). Our approach begins with a large bilingual Arabic-English teacher and progressively distills its knowledge into compressed student models while preserving Arabic-relevant acoustic and paralinguistic representations. We further study PCA-based compression of the teacher supervision signal to better match the capacity of shallow and thin students. Compared with HuBERT and XLS-R, HArnESS consistently improves performance on Arabic downstream tasks, while the compressed models remain competitive under substantial structural reduction. These results position HArnESS as a practical and accessible Arabic-centric SSL foundation for real-world speech applications.

【4】A Manual Bar-by-Bar Tempo Measurement Protocol for Polyphonic Chamber Music Recordings: Design, Validation, and Application to Beethoven's Piano and Cello Sonatas
标题:复调室内乐录音的手动逐小节节奏测量协议:贝多芬钢琴和大提琴奏鸣曲的设计、验证和应用
链接:https://arxiv.org/abs/2604.15278
作者:Ignasi Sole
摘要:经验性的演奏分析依赖于从录音中准确提取速度数据,然而,为单声道音频或现代录音室条件设计的标准计算工具在应用于历史复调室内乐时系统地失败了。本文记录了贝多芬五首钢琴和大提琴奏鸣曲(Op.~)的二重奏录音中自动节拍检测软件的失败5个~ 1和~2; Op.~ 69; Op.~ 102个~ 1和~2),并提出了一个正式的手动替代方案:一个累积的lap-timer协议,产生的酒吧级每分钟的节拍数据与毫秒的分辨率。该协议是与一位专门从事超大规模集成电路设计的工程师跨学科合作开发的,它依赖于一个累积的时间戳架构,该架构可以防止错误累积,允许内部自验证,并捕获自动化工具系统抑制或误读的表现性时序现象(rubato,fermata,accelerandi,ritardandi)。BPM公式的数学推导,电子表格的数据结构,和误差特性的完整介绍。应用于1930- 2012年的100多个运动水平记录,该协议生成了一个数据集,随后通过tempographs,具有样条平滑概率密度函数的直方图,脊线图和组合图表进行可视化。本文认为,人工注释不是一种方法上的撤退,而是在面对复调历史记录的具体挑战时,对计算工具固有局限性的原则性回应。完整的数据集和分析代码是公开的。
摘要:Empirical performance analysis depends on the accurate extraction of tempo data from recordings, yet standard computational tools, designed for monophonic audio or modern studio conditions, fail systematically when applied to historical polyphonic chamber music. This paper documents the failure of automated beat-detection software on duo recordings of Beethoven's five piano and cello sonatas (Op.~5 Nos.~1 and~2; Op.~69; Op.~102 Nos.~1 and~2), and presents a formalised manual alternative: a cumulative lap-timer protocol that yields bar-level beats-per-minute data with millisecond resolution. The protocol, developed in cross-disciplinary collaboration with an engineer specialising in VLSI design, rests on a cumulative timestamp architecture that prevents error accumulation, permits internal self-validation, and captures expressive timing phenomena (rubato, fermatas, accelerandi, ritardandi) that automated tools systematically suppress or misread. The mathematical derivation of the BPM formula, the spreadsheet data structure, and the error characterisation are presented in full. Applied to over one hundred movement-level recordings spanning 1930--2012, the protocol generated a dataset subsequently visualised through tempographs, histograms with spline-smoothed probability density functions, ridgeline plots, and combination charts. The paper argues that manual annotation is not a methodological retreat but a principled response to the intrinsic limitations of computational tools when faced with the specific challenges of polyphonic historical recordings. The complete dataset and analysis code are publicly available.

【5】The Acoustic Camouflage Phenomenon: Re-evaluating Speech Features for Financial Risk Prediction
标题:声学Camerage现象:重新评估用于金融风险预测的语音特征
链接:https://arxiv.org/abs/2604.14619
作者:Dhruvin Dungrani,Disha Dungrani
摘要:在计算语言学中,从语音信号中检测认知负荷和欺骗是一个重要的研究领域。最近的努力试图将这些声学框架应用于公司盈利电话,以预测灾难性的股票市场波动。在这项研究中,我们经验性地研究了声学特征提取(音高,抖动和犹豫)的限制时,适用于训练有素的扬声器在野外电话会议环境。利用双流后期融合架构,我们将基于声学的流与基线自然语言处理(NLP)流进行对比。孤立的NLP模型对尾部风险下行事件的召回率为66.25%。令人惊讶的是,通过后期融合整合声学特征显着降低了性能,将召回率降低到47.08%。我们将这种退化确定为声学伪装,其中媒体训练的声音调节引入了矛盾的噪音,破坏了多模态元学习者。我们将这些发现作为语音处理在高风险金融预测中应用的边界条件。
摘要:In computational paralinguistics, detecting cognitive load and deception from speech signals is a heavily researched domain. Recent efforts have attempted to apply these acoustic frameworks to corporate earnings calls to predict catastrophic stock market volatility. In this study, we empirically investigate the limits of acoustic feature extraction (pitch, jitter, and hesitation) when applied to highly trained speakers in in-the-wild teleconference environments. Utilizing a two-stream late-fusion architecture, we contrast an acoustic-based stream with a baseline Natural Language Processing (NLP) stream. The isolated NLP model achieved a recall of 66.25% for tail-risk downside events. Surprisingly, integrating acoustic features via late fusion significantly degraded performance, reducing recall to 47.08%. We identify this degradation as Acoustic Camouflage, where media-trained vocal regulation introduces contradictory noise that disrupts multimodal meta-learners. We present these findings as a boundary condition for speech processing applications in high-stakes financial forecasting.

【6】VoxSafeBench: Not Just What Is Said, but Who, How, and Where
标题:VoxSafeBench:不仅仅是说了什么,还有谁、如何以及在哪里
链接:https://arxiv.org/abs/2604.14548
作者:Yuxiang Wang,Hongyu Liu,Yijiang Xu,Qinke Ni,Li Wang,Wan Lin,Kunyu Feng,Dekun Chen,Xu Tan,Lei Wang,Jie Shi,Zhizheng Wu
摘要:随着语音语言模型(SLM)从个人设备过渡到共享的多用户环境,它们的响应必须远远超过单词本身。谁在说话,他们的声音如何,以及对话发生在哪里,每一个都可以把一个原本善意的请求变成一个不安全,不公平或侵犯隐私的请求。然而,现有的基准主要集中在基本的音频理解上,孤立地研究个体风险,或者将固有有害的内容与仅由于其声学背景而变得有问题的内容混为一谈。我们介绍VoxSafeBench,在第一个基准,共同评估社会对准SLM在三个方面:安全性,公平性和隐私。VoxSafeBench采用两层设计:第一层使用匹配的文本和音频输入评估以内容为中心的风险,而第二层则针对音频条件风险,其中转录是良性的,但适当的响应取决于扬声器,语言提示或周围环境。为了验证Tier2,我们包括中间感知探头,并确认前沿SLM可以成功地检测到这些声学线索,但仍然无法对它们采取适当的行动。在22个双语覆盖的任务中,我们发现,在文本上表现出强大的保障措施往往会在语音中降低:安全意识下降,因为扬声器和场景条件的风险,当人口统计差异被口头传达时,公平性受到侵蚀,当上下文线索到达声学时,隐私保护就会动摇。总之,这些结果暴露了一个普遍的语音接地的差距:目前的SLM经常认识到相关的社会规范的文本,但不能应用它时,决定性的线索必须接地的讲话。代码和数据可在https://amphionteam.github.io/VoxSafeBench_demopage/上公开获取
摘要:As speech language models (SLMs) transition from personal devices into shared, multi-user environments, their responses must account for far more than the words alone. Who is speaking, how they sound, and where the conversation takes place can each turn an otherwise benign request into one that is unsafe, unfair, or privacy-violating. Existing benchmarks, however, largely focus on basic audio comprehension, study individual risks in isolation, or conflate content that is inherently harmful with content that only becomes problematic due to its acoustic context. We introduce VoxSafeBench, among the first benchmarks to jointly evaluate social alignment in SLMs across three dimensions: safety, fairness, and privacy. VoxSafeBench adopts a Two-Tier design: Tier1 evaluates content-centric risks using matched text and audio inputs, while Tier2 targets audio-conditioned risks in which the transcript is benign but the appropriate response hinges on the speaker, paralinguistic cues, or the surrounding environment. To validate Tier2, we include intermediate perception probes and confirm that frontier SLMs can successfully detect these acoustic cues yet still fail to act on them appropriately. Across 22 tasks with bilingual coverage, we find that safeguards appearing robust on text often degrade in speech: safety awareness drops for speaker- and scene-conditioned risks, fairness erodes when demographic differences are conveyed vocally, and privacy protections falter when contextual cues arrive acoustically. Together, these results expose a pervasive speech grounding gap: current SLMs frequently recognize the relevant social norm in text but fail to apply it when the decisive cue must be grounded in speech. Code and data are publicly available at: https://amphionteam.github.io/VoxSafeBench_demopage/

【7】Disentangled Dual-Branch Graph Learning for Conversational Emotion Recognition
标题:用于对话情绪识别的解纠缠双分支图学习
链接:https://arxiv.org/abs/2604.14204
作者:Chengling Guo,Yuntao Shou,Tao Meng,Wei Ai,Yun Tan,Keqin Li
备注:16 pages
摘要:会话中的多模态情感识别旨在通过在上下文中联合建模文本、声学和视觉线索来推断话语级情感。尽管最近取得了进展,关键的挑战仍然存在,包括冗余的跨模态信息,不完善的语义对齐,以及高阶扬声器交互的建模不足。为了解决这些问题,我们提出了一个框架,结合双空间特征解纠缠与双分支图学习。共享编码器和模态特定编码器用于分离模态不变表示和模态特定表示。不变特征由傅立叶图神经网络建模,以捕获全局一致性和互补模式,并具有频域对比目标以增强可区分性。并行地,在模态特定特征上构造说话者感知超图来建模高阶交互,以及说话者一致性约束来保持连贯的语义。最后,将两个分支融合用于话语级情感预测。IEMOCAP和MELD上的实验表明,该方法在强基线下具有更好的性能,验证了其有效性。
摘要:Multimodal emotion recognition in conversations aims to infer utterance-level emotions by jointly modeling textual, acoustic, and visual cues within context. Despite recent progress, key challenges remain, including redundant cross-modal information, imperfect semantic alignment, and insufficient modeling of high-order speaker interactions. To address these issues, we propose a framework that combines dual-space feature disentanglement with dual-branch graph learning. A shared encoder and modality-specific encoders are used to separate modality-invariant and modality-specific representations. The invariant features are modeled by a Fourier graph neural network to capture global consistency and complementary patterns, with a frequency-domain contrastive objective to enhance discriminability. In parallel, a speaker-aware hypergraph is constructed over modality-specific features to model high-order interactions, along with a speaker-consistency constraint to maintain coherent semantics. Finally, the two branches are fused for utterance-level emotion prediction. Experiments on IEMOCAP and MELD demonstrate that the proposed method achieves superior performance over strong baselines, validating its effectiveness.

【8】From Black Box to Glass Box: Cross-Model ASR Disagreement to Prioto Review in Ambient AI Scribe Documentation
标题:从黑匣子到玻璃盒:跨模型ASB对Ambient AI Scribe文档中Prioto审查的分歧
链接:https://arxiv.org/abs/2604.14152
作者:Abdolamir Karbalaie,Fernando Seoane,Farhad Abtahi
摘要:环境人工智能“抄写”系统有望减少临床文档负担,但如果不仔细审查,自动语音识别(ASR)错误可能会被忽视,并且高质量的人类参考转录通常无法用于校准不确定性。我们研究了异构ASR系统之间的跨模型分歧是否可以作为一个无参考的不确定性信号,以优先考虑医疗转录工作流程中的人类验证。使用50个公开的医学教育音频片段(8小时14分钟),我们用8个ASR系统转录了每个片段,这些系统涵盖了商业API和开源引擎。我们对齐多模型输出,建立共识伪参考,并使用多数强度度量量化标记级协议;我们进一步通过类型(内容与标点符号/格式)表征分歧,并通过留一模型(刀切)共识评分评估每个模型的协议。模型间可靠性较低(ICC[2,1] = 0.131),表明系统间存在异质失效模式。在76,398个被评估的代币头寸中,72.1%显示出近乎一致的一致性(7-8个模型),而2.5%属于高风险范围(0-3个模型),高风险质量在口音组中的比例从0.7%到11.4%不等。低一致性区域因内容不一致而丰富,在高风险质量的五分位数中,内容分数从53.9%增加到73.9%。这些结果表明,跨模型不一致提供了一个稀疏的,可定位的信号,可以在没有人类验证的参考的情况下显示潜在不可靠的转录本跨度,从而实现有针对性的审查;标记区域的临床准确性仍有待建立。
摘要:Ambient AI "scribe" systems promise to reduce clinical documentation burden, but automatic speech recognition (ASR) errors can remain unnoticed without careful review, and high-quality human reference transcripts are often unavailable for calibrating uncertainty. We investigate whether cross-model disagreement among heterogeneous ASR systems can act as a reference-free uncertainty signal to prioritize human verification in medical transcription workflows. Using 50 publicly available medical education audio clips (8 h 14 min), we transcribed each clip with eight ASR systems spanning commercial APIs and open-source engines. We aligned multi-model outputs, built consensus pseudo-references, and quantified token-level agreement using a majority-strength metric; we further characterized disagreements by type (content vs. punctuation/formatting) and assessed per-model agreement via leave-one-model-out (jackknife) consensus scoring. Inter-model reliability was low (ICC[2,1] = 0.131), indicating heterogeneous failure modes across systems. Across 76,398 evaluated token positions, 72.1% showed near-unanimous agreement (7-8 models), while 2.5% fell into high-risk bands (0-3 models), with high-risk mass varying from 0.7% to 11.4% across accent groups. Low-agreement regions were enriched for content disagreements, with the content fraction increasing from 53.9% to 73.9% across quintiles of high-risk mass. These results suggest that cross-model disagreement provides a sparse, localizable signal that can surface potentially unreliable transcript spans without human-verified references, enabling targeted review; clinical accuracy of flagged regions remains to be established.


机器翻译由腾讯交互翻译提供,仅供参考