今日论文合集:cs.SD语音29篇,eess.AS音频处理10篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora
标题:MimicLM:通过伪并行语音库的自回归建模实现Zero-Shot语音模仿
链接:https://arxiv.org/abs/2604.11552

作者:Tao Feng,Yuxiang Wang,Yuancheng Wang,Xueyao Zhang,Dekun Chen,Chaoren Wang,Xun Guan,Zhizheng Wu
摘要:语音模仿的目的是在保持语言内容的同时,将源语音转换为与参考说话人的音色和说话风格相匹配的语音。一种直接的方法是在(源、参考、目标)的三元组上进行训练,其中源和目标共享相同的内容,但目标匹配参考的语音特征,然而这样的数据极其稀缺。现有的方法要么采用精心设计的解纠缠架构来绕过这种数据稀缺性,要么利用外部系统来合成伪并行训练数据。然而,前者需要复杂的模型设计,而后者在使用合成语音作为训练目标时面临质量天花板。为了解决这些限制,我们提出了MimicLM,它采用了一种新的方法,使用合成语音作为训练源,同时保留真实的录音作为目标。这种设计使模型能够直接从真实的语音分布中学习,打破了合成质量的上限。在这种数据构建方法的基础上,我们结合了交织文本音频建模来指导内容准确的语音的生成,并应用具有偏好对齐的后训练来减轻在合成数据上训练时固有的分布失配。实验表明,MimicLM实现了卓越的语音模仿质量与一个简单而有效的架构,显着优于现有的方法在自然,同时保持竞争力的相似性分数在说话人的身份,口音和情感维度。
摘要:Voice imitation aims to transform source speech to match a reference speaker's timbre and speaking style while preserving linguistic content. A straightforward approach is to train on triplets of (source, reference, target), where source and target share the same content but target matches the reference's voice characteristics, yet such data is extremely scarce. Existing approaches either employ carefully designed disentanglement architectures to bypass this data scarcity or leverage external systems to synthesize pseudo-parallel training data. However, the former requires intricate model design, and the latter faces a quality ceiling when synthetic speech is used as training targets. To address these limitations, we propose MimicLM, which takes a novel approach by using synthetic speech as training sources while retaining real recordings as targets. This design enables the model to learn directly from real speech distributions, breaking the synthetic quality ceiling. Building on this data construction approach, we incorporate interleaved text-audio modeling to guide the generation of content-accurate speech and apply post-training with preference alignment to mitigate the inherent distributional mismatch when training on synthetic data. Experiments demonstrate that MimicLM achieves superior voice imitation quality with a simple yet effective architecture, significantly outperforming existing methods in naturalness while maintaining competitive similarity scores across speaker identity, accent, and emotion dimensions.


【2】Ti-Audio: The First Multi-Dialectal End-to-End Speech LLM for Tibetan
标题:Ti-Audio:第一个针对藏族的多方言端到端语音LLM
链接:https://arxiv.org/abs/2604.11110

作者:Jialing Wang,Yue Zhao,Yuhao Zhang,Jing Yu,Shaosai Li,Zhanchen Dai,Benyou Wang,Haizhou Li
摘要:近年来,语音大语言模型(Speech-LLM)取得了重大进展,极大地增强了多模态交互能力,但在低资源和方言多样性环境下的应用仍面临挑战。藏文数据的严重缺乏,加上其主要方言(藏藏、安多和康)之间的语音差异,是这一挑战的一个主要例子。本文提出了Ti-Audio,第一个多方言端到端语音LLM藏语。为了有效地对齐语音和文本,我们引入了一个动态Q-Former适配器,它可以从可变长度的语音中提取基本的声学特征,即使在数据有限的情况下也能确保稳定的跨模态对齐。在数据层面,我们利用相关方言之间的互助来缓解数据稀缺,并采用基于温度的采样策略来最大限度地发挥这种协同作用。实验结果表明,Ti-Audio在藏语自动语音识别和语音翻译基准测试中达到了最先进的性能。我们的工作验证了跨方言合作的有效性,并提供了一个可扩展的范例,在低资源的情况下,语音LLM的发展。
摘要:Recent advances in Speech Large Language Models (Speech-LLMs) have made significant progress, greatly enhancing multimodal interaction capabilities.However, their application in low-resource and dialect-diverse environments still faces challenges. The severe scarcity of Tibetan data, coupled with the phonetic differences among its major dialects (Ü-Tsang, Amdo, and Kham), is a prime example of this challenge. This paper proposes Ti-Audio, the first multi-dialectal end-to-end Speech-LLM for Tibetan. To efficiently align speech and text, we introduce a Dynamic Q-Former Adapter that extracts essential acoustic features from variable-length speech, ensuring stable cross-modal alignment even with limited data. At the data level, we leverage mutual assistance among related dialects to alleviate data scarcity and employ a temperature-based sampling strategy to maximize this synergy. Experimental results demonstrate that Ti-Audio achieves state-of-the-art performance on Tibetan benchmarks for automatic speech recognition and speech translation. Our work validates the effectiveness of cross-dialectal cooperation and provides a scalable paradigm for the development of Speech-LLM in low-resource scenarios.


【3】ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing
标题:ActorMind:模仿人类演员推理进行演讲角色扮演
链接:https://arxiv.org/abs/2604.11103

作者:Xi Chen,Wei Xue,Yike Guo
摘要:角色扮演越来越受到关注,因为它为人机交互提供了坚实的基础,并促进了社会学研究。然而,目前的研究仅限于语篇形式,忽视了在日常生活中起主导作用的言语,从而限制了真正的角色扮演。为了弥合这一差距,我们概念化和基准语音角色扮演通过ActorMindBench,我们提出了一个相应的推理框架,称为ActorMind。具体而言,(1)语音角色扮演使模型能够根据其角色,场景和口语对话提供具有个性化语言特征的自发响应。(2)ActorMindBench是一个分层基准,包括具有7,653个话语的话语级内容,具有313个场景的场景级内容和具有6个角色的角色级内容。(3)ActorMind是一个现成的、多智能体的、链式推理框架,它模拟了人类演员在剧院中的表演。具体来说,ActorMind首先通过Eye Agent读取其分配的角色描述,然后通过Ear Agent在上下文口语对话中识别情感线索。随后,大脑Agent生成一个描述性的情感状态,最后,嘴部Agent提供注入相应情感状态的脚本。实验结果证明了ActorMind在增强语音角色扮演方面的有效性。
摘要:Role-playing has garnered rising attention as it provides a strong foundation for human-machine interaction and facilitates sociological research. However, current work is confined to textual modalities, neglecting speech, which plays a predominant role in daily life, thus limiting genuine role-playing. To bridge this gap, we conceptualize and benchmark speech role-playing through ActorMindBench, and we present a corresponding reasoning framework, called ActorMind. Specifically, (1) Speech Role-Playing enables models to deliver spontaneous responses with personalized verbal traits based on their role, the scene, and spoken dialogue. (2) ActorMindBench is a hierarchical benchmark comprises Utterance-Level content with 7,653 utterances, Scene-Level content with 313 scenes, and Role-Level content with 6 roles. (3) ActorMind is an off-the-shelf, multi-agent, chain-of-though style reasoning framework that emulates how human actors perform in theaters. Concretely, ActorMind first reads its assigned role description via Eye Agent, then comprehends emotional cues within contextual spoken dialogues through Ear Agent. Subsequently, Brain Agent generates a descriptive emotional state, and finally, Mouth Agent delivers the scripts infused with corresponding emotion state. Experimental results demonstrate the effectiveness of ActorMind in enhancing speech role-playing.


【4】Efficient Training for Cross-lingual Speech Language Models
标题:跨语言语音语言模型的高效训练
链接:https://arxiv.org/abs/2604.11096

作者:Yan Zhou,Qingkai Fang,Yun Hong,Yang Feng
备注:Accepted to Findings of ACL 2026
摘要:目前,大型语言模型(LLM)主要关注文本模态。为了实现更自然的人机交互,语音LLM正在出现,但由于数据有限以及难以扩展到更多语言,构建有效的端到端语音LLM仍然具有挑战性。在本文中,我们介绍了跨语言的语音语言模型(CSLM),一个有效的训练方法的跨语言的语音LLM的基础上离散的语音令牌。我们提出了一种新的对齐策略,通过持续的预训练实现跨模态和跨语言对齐。通过在语音-文本交错模态链生成过程之后进行指令微调,我们以更细的粒度增强模态对齐,从而提高生成质量并减少延迟。CSLM同时对齐不同的模态和语言,而不需要大量的语音数据,从而表现出良好的语言可扩展性。跨通道任务、单语言会话任务和跨语言会话任务的测试结果表明CSLM具有较强的跨通道对齐能力和通用任务能力。(Code网址:https://github.com/ictnlp/CSLM)
摘要:Currently, large language models (LLMs) predominantly focus on the text modality. To enable more natural human-AI interaction, speech LLMs are emerging, but building effective end-to-end speech LLMs remains challenging due to limited data and the difficulty in expanding to more languages. In this paper, we introduce Cross-lingual Speech Language Model (CSLM), an efficient training method for cross-lingual speech LLMs based on discrete speech tokens. We propose a novel alignment strategy that achieves cross-modal and cross-lingual alignment through continual pre-training. By conducting instruction fine-tuning following a speech-text interleaved chain-of-modality generation process, we enhance modal alignment at a finer granularity, thereby improving generation quality and reducing latency. CSLM aligns different modalities and languages simultaneously without the need for massive speech data, thus exhibiting good language scalability. Evaluations on cross-modal tasks, mono-lingual conversational tasks, and cross-lingual conversational tasks demonstrate CSLM's strong cross-modal alignment capabilities and general task abilities. (Code is available at: https://github.com/ictnlp/CSLM)


【5】LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation
标题:LaDA-Band:声乐到伴奏一代的语言扩散模型
链接:https://arxiv.org/abs/2604.11052

作者:Qi Wang,Zhexu Shen,Meng Chen,Guoxin Yu,Chaoxu Pang,Weifeng Zhao,Wenjiang Zhou
备注:Submitted to ACMMM 2026. Under review
摘要:声乐到伴奏(V2 A)生成旨在将原始声乐录音转换为完全编排的伴奏,本质上需要共同解决伴奏三难:保持声学真实性,保持与声乐音轨的全局一致性,并在整首歌曲中产生动态配器。现有的开源方法通常在这些目标之间做出妥协。连续潜在生成模型可以捕获长的音乐跨度,但通常难以保留细粒度的声学细节。相比之下,离散自回归模型保持局部保真度,但遭受单向生成和误差积累在扩展的上下文中。我们提出了LaDA波段,一个端到端的框架,引入离散掩蔽扩散的V2 A任务。我们的方法将V2 A生成公式化为离散掩蔽扩散,即,一种全局的、非自回归的去噪公式,其将离散音频编解码器令牌的代表性优势与全序列双向上下文建模相结合。这种设计提高了长距离结构一致性和时间同步性,同时保留了清晰的声学细节。在此基础上,LaDA波段进一步引入了一个双轨前缀条件反射架构,一个辅助的锚定伴奏区域的标记检测目标,以及一个两阶段的渐进课程,以扩展离散掩蔽扩散到完整的歌曲声乐伴奏生成。在学术和现实基准上进行的大量实验表明,LaDA-Band在现有基准上不断提高声学真实性,全局一致性和动态编排,同时即使没有辅助参考音频也能保持强大的性能。代码和音频样本可在https://github.com/Duoluoluos/TME-LaDA-Band上获得。
摘要:Vocal-to-accompaniment (V2A) generation, which aims to transform a raw vocal recording into a fully arranged accompaniment, inherently requires jointly addressing an accompaniment trilemma: preserving acoustic authenticity, maintaining global coherence with the vocal track, and producing dynamic orchestration across a full song. Existing open-source approaches typically make compromises among these goals. Continuous-latent generation models can capture long musical spans but often struggle to preserve fine-grained acoustic detail. In contrast, discrete autoregressive models retain local fidelity but suffer from unidirectional generation and error accumulation in extended contexts. We present LaDA-Band, an end-to-end framework that introduces Discrete Masked Diffusion to the V2A task. Our approach formulates V2A generation as Discrete Masked Diffusion, i.e., a global, non-autoregressive denoising formulation that combines the representational advantages of discrete audio codec tokens with full-sequence bidirectional context modeling. This design improves long-range structural consistency and temporal synchronization while preserving crisp acoustic details. Built on this formulation, LaDA-Band further introduces a dual-track prefix-conditioning architecture, an auxiliary replaced-token detection objective for weakly anchored accompaniment regions, and a two-stage progressive curriculum to scale Discrete Masked Diffusion to full-song vocal-to-accompaniment generation. Extensive experiments on both academic and real-world benchmarks show that LaDA-Band consistently improves acoustic authenticity, global coherence, and dynamic orchestration over existing baselines, while maintaining strong performance even without auxiliary reference audio. Codes and audio samples are available at https://github.com/Duoluoluos/TME-LaDA-Band .


【6】Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
标题:音频Flamingo Next:下一代语音、声音和音乐开放音频语言模型
链接:https://arxiv.org/abs/2604.10905

作者:Sreyan Ghosh,Arushi Goel,Kaousheik Jayakumar,Lasha Koroshinadze,Nishit Anand,Zhifeng Kong,Siddharth Gururani,Sang-gil Lee,Jaehyeon Kim,Aya Aljafari,Chao-Han Huck Yang,Sungwon Kim,Ramani Duraiswami,Dinesh Manocha,Mohammad Shoeybi,Bryan Catanzaro,Ming-Yu Liu,Wei Ping
备注:Project website: https://afnext-umd-nvidia.github.io/
摘要:Audio Flamingo Next(AF-Next)是Audio Flamingo系列中功能最强大的下一代大型音频语言模型,旨在促进对语音、环境声音和音乐的理解和推理。与Audio Flamingo 3相比,AF-Next引入了:(i)更强大的基础音频语言模型,可显著提高各种音频理解任务的准确性;(ii)可扩展的策略,用于构建超出现有学术基准的大规模音频理解和推理数据;(iii)支持长达30分钟的长而复杂的音频输入;以及(iv)时间音频思想链,一种新的推理范式,其明确地将中间推理步骤基于长音频中的时间戳,从而实现细粒度的时间对齐和改进的可解释性。为了实现这些功能,我们首先对Audio Flamingo 3进行系统分析,以确定音频理解和推理方面的关键差距。然后,我们策划和扩展总计超过100万小时的新的大规模数据集,以解决这些限制,并扩展现有的AudioSkills-XL,LongAudio-XL,AF-Think和AF-Chat数据集。AF-Next使用基于训练的策略进行训练,包括训练前、训练中和训练后阶段。在20个音频理解和推理基准测试中进行的广泛实验,包括具有挑战性的长音频任务,表明AF-Next的性能远远优于类似大小的开放模型,并且与更大的开放重量和封闭模型保持高度竞争力,有时甚至超过它们。除了基准性能之外,AF-Next还表现出强大的现实效用,并可以很好地转移到看不见的任务,突出了其鲁棒性和泛化能力。除了所有的数据、代码和方法,我们还开源了AF-Next的3个变体,包括AF-Next-Instruct、AF-Next-Think和AF-Next-Captioner。
摘要:We present Audio Flamingo Next (AF-Next), the next-generation and most capable large audio-language model in the Audio Flamingo series, designed to advance understanding and reasoning over speech, environmental sounds and music. Compared to Audio Flamingo 3, AF-Next introduces: (i) a stronger foundational audio-language model that significantly improves accuracy across diverse audio understanding tasks; (ii) scalable strategies for constructing large-scale audio understanding and reasoning data beyond existing academic benchmarks; (iii) support for long and complex audio inputs up to 30 minutes; and (iv) Temporal Audio Chain-of-Thought, a new reasoning paradigm that explicitly grounds intermediate reasoning steps to timestamps in long audio, enabling fine-grained temporal alignment and improved interpretability. To enable these capabilities, we first conduct a systematic analysis of Audio Flamingo 3 to identify key gaps in audio understanding and reasoning. We then curate and scale new large-scale datasets totaling over 1 million hours to address these limitations and expand the existing AudioSkills-XL, LongAudio-XL, AF-Think and AF-Chat datasets. AF-Next is trained using a curriculum-based strategy spanning pre-training, mid-training and post-training stages. Extensive experiments across 20 audio understanding and reasoning benchmarks, including challenging long-audio tasks, show that AF-Next outperforms similarly sized open models by large margins and remains highly competitive with and sometimes surpasses, much larger open-weight and closed models. Beyond benchmark performance, AF-Next exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability. In addition to all data, code and methods, we open-source 3 variants of AF-Next, including AF-Next-Instruct, AF-Next-Think and AF-Next-Captioner.


【7】MeloTune: On-Device Arousal Learning and Peer-to-Peer Mood Coupling for Proactive Music Curation
标题:MeloButton:设备上唤醒学习和点对点情绪耦合,实现主动式音乐策划
链接:https://arxiv.org/abs/2604.10815

作者:Hongwei Xu
备注:31 pages, 1 figures, 3 tables
摘要:MeloTune是一款部署在iPhone上的音乐代理,它将网状记忆协议(MMP)和符号向量注意力融合(SVAF)实例化为具有点对点情绪耦合的情感感知音乐策展的制作系统。每个设备运行两个封闭形式的连续时间(CfC)网络:一个私人的网络级CfC,预测罗素环上的短期情感轨迹,并驱动主动策展,以及MMP层6的共享网格运行时CfC,它集成了来自共听对等体的认知记忆块(CMB)。CfC隐藏状态永远不会越过电线;只有结构化CMB才能做到。个人唤醒功能(PAF)取代了从音频强度到心理唤醒的标准线性映射,并根据每个听众的学习调整,从行为信号(跳过,完成,喜爱,音量)以及用户声明的情绪和机器推理之间的漂移进行训练。相同的音轨接收针对不同听众的不同唤醒预测。该模型(94,552个参数)的轨迹MAE为0.414,模式准确率为96.6%,意图准确率为69.4%。来自现场部署会话的PAF证据(11个流派的46个观察结果)表明,学习循环是端到端的,在22个观察结果后,pop达到完全置信度。所有推理都通过CoreML在设备上运行。据我们所知,这是MMP/SVAF在消费者移动硬件上的首次生产部署。附带的SDK(sym-swift v0.3.78、SYMCore v0.3.7)强制执行严格的协议一致性。音乐是案例研究;基底是贡献。
摘要:MeloTune is an iPhone-deployed music agent that instantiates the Mesh Memory Protocol (MMP) and Symbolic-Vector Attention Fusion (SVAF) as a production system for affect-aware music curation with peer-to-peer mood coupling. Each device runs two closed-form continuous-time (CfC) networks: a private listener-level CfC that predicts a short-horizon affective trajectory on Russell's circumplex and drives proactive curation, and a shared mesh-runtime CfC at MMP Layer 6 that integrates Cognitive Memory Blocks (CMBs) from co-listening peers. CfC hidden states never cross the wire; only structured CMBs do. A Personal Arousal Function (PAF) replaces the standard linear mapping from audio intensity to psychological arousal with a per-listener learned adjustment, trained from behavioral signals (skip, completion, favorite, volume) and from drift between user-declared mood and machine inference. The same track receives different arousal predictions for different listeners. The model (94,552 parameters) achieves trajectory MAE 0.414, pattern accuracy 96.6%, and intent accuracy 69.4% on held-out validation. PAF evidence from a live deployment session (46 observations across 11 genres) demonstrates that the learning loop operates end-to-end, with pop reaching full confidence after 22 observations. All inference runs on-device via CoreML. To our knowledge, this is the first production deployment of MMP/SVAF on consumer mobile hardware. The accompanying SDK (sym-swift v0.3.78, SYMCore v0.3.7) enforces strict protocol conformance. Music is the case study; the substrate is the contribution.


【8】BlasBench: An Open Benchmark for Irish Speech Recognition
标题:BlasBench:爱尔兰语音识别的开放基准
链接:https://arxiv.org/abs/2604.10736

作者:Jyoutir Raj,John Conway
备注:8 pages, 4 tables, 3 appendices. Code and data: https://github.com/jyoutir/blasbench
摘要:没有开放的爱尔兰特定的基准比较终端用户ASR系统下共享爱尔兰意识的评估协议。为了解决这个问题,我们发布了BlasBench,这是一个开放的评估工具,具有爱尔兰语感知的文本规范化功能,可以保留流行语、lenition和eclipsis。我们在Common Voice ga-IE和FLEURS ga-IE上对四个架构系列的12个系统进行了基准测试。所有Whisper变体均超过100% WER。最好的开放模型(omniASR LLM 7 B)在Common Voice上实现了30.65%的WER,在FLEURS上实现了39.09%。我们注意到,在Common Voice上微调的模型在FLEURS上丢失了33-43个WER点,揭示了一个对单数据集评估不可见的泛化差距。
摘要:No open Irish-specific benchmark compares end-user ASR systems under a shared Irish-aware evaluation protocol. To solve this, we release BlasBench, an open evaluation harness with Irish-aware text normalisation that preserves fadas, lenition, and eclipsis. We benchmark 12 systems across four architecture families on Common Voice ga-IE and FLEURS ga-IE. All Whisper variants exceed 100% WER. The best open model (omniASR LLM 7B) achieves 30.65% WER on Common Voice and 39.09% on FLEURS. We noticed models fine-tuned on Common Voice lose 33-43 WER points on FLEURS, revealing a generalisation gap that is invisible to single-dataset evaluation.


【9】Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
标题:Audio-Omni:将多模态理解扩展到多功能音频生成和编辑
链接:https://arxiv.org/abs/2604.10708

作者:Zeyue Tian,Binxin Yang,Zhaoyang Liu,Jiexuan Zhang,Ruibin Yuan,Hubery Yin,Qifeng Chen,Chen Li,Jing Lv,Wei Xue,Yike Guo
摘要:多模态模型的最新进展促进了音频理解、生成和编辑的快速发展。然而,这些功能通常由专门的模型来解决,从而使能够无缝集成所有三个任务的真正统一框架的开发未得到充分探索。虽然一些开创性的工作已经探索了统一的音频理解和生成,但它们通常仍然局限于特定的领域。为了解决这个问题,我们引入了Audio-Omni,这是第一个端到端框架,可以统一通用声音,音乐和语音领域的生成和编辑,并具有集成的多模态理解功能。我们的架构协同冻结多模态大语言模型的高层次推理与可训练的扩散Transformer高保真合成。为了克服音频编辑中关键数据的稀缺性,我们构建了AudioEdit,这是一个新的大规模数据集,包含超过一百万个精心策划的编辑对。大量的实验表明,Audio-Omni在一系列基准测试中实现了最先进的性能,优于先前的统一方法,同时实现了与专业专家模型相当或更高的性能。除了其核心功能之外,Audio-Omni还具有显着的继承能力,包括知识增强推理生成,上下文生成和用于音频生成的zero-shot跨语言控制,突出了通用生成音频智能的有希望的方向。代码、模型和数据集将在https://zeyuet.github.io/Audio-Omni上公开发布。
摘要:Recent progress in multimodal models has spurred rapid advances in audio understanding, generation, and editing. However, these capabilities are typically addressed by specialized models, leaving the development of a truly unified framework that can seamlessly integrate all three tasks underexplored. While some pioneering works have explored unifying audio understanding and generation, they often remain confined to specific domains. To address this, we introduce Audio-Omni, the first end-to-end framework to unify generation and editing across general sound, music, and speech domains, with integrated multi-modal understanding capabilities. Our architecture synergizes a frozen Multimodal Large Language Model for high-level reasoning with a trainable Diffusion Transformer for high-fidelity synthesis. To overcome the critical data scarcity in audio editing, we construct AudioEdit, a new large-scale dataset comprising over one million meticulously curated editing pairs. Extensive experiments demonstrate that Audio-Omni achieves state-of-the-art performance across a suite of benchmarks, outperforming prior unified approaches while achieving performance on par with or superior to specialized expert models. Beyond its core capabilities, Audio-Omni exhibits remarkable inherited capabilities, including knowledge-augmented reasoning generation, in-context generation, and zero-shot cross-lingual control for audio generation, highlighting a promising direction toward universal generative audio intelligence. The code, model, and dataset will be publicly released on https://zeyuet.github.io/Audio-Omni.


【10】Multimodal Dataset Normalization and Perceptual Validation for Music-Taste Correspondences
标题:音乐品味对应的多模式数据集规范化和感知验证
链接:https://arxiv.org/abs/2604.10632

作者:Matteo Spanio,Valentina Frezzato,Antonio Rodà
备注:Submitted to SMC2026
摘要:收集大型的、对齐的跨模态数据集用于音乐风味研究是困难的,因为感知实验是昂贵的,而且设计得很小。我们通过两个互补的实验来解决这个瓶颈。第一个测试是否音频风味的相关性,功能的重要性排名,和潜在的因素结构转移从实验的音轨集合(257轨道与人类注释)到一个大型的FM衍生语料库(49,300段合成标签)。第二个验证计算风味目标-来自食品化学通过一个可重复的管道-对人类的感知在一个在线的听众研究(49~参与者,20~轨道)。两个实验的结果趋于一致:定量迁移分析证实跨模态结构在不同的监督机制中得到了保留,感知评估显示计算目标和听众评级之间存在显著的一致性(排列p<0.0001$,Mantel $r=0.45$,Procrustes $m^2=0.51$)。总之,这些研究结果支持的结论是,声音调味效果存在于合成FMA注释。我们发布数据集和配套代码,以支持可重复的跨模态AI研究。
摘要:Collecting large, aligned cross-modal datasets for music-flavor research is difficult because perceptual experiments are costly and small by design. We address this bottleneck through two complementary experiments. The first tests whether audio-flavor correlations, feature-importance rankings, and latent-factor structure transfer from an experimental soundtracks collection (257~tracks with human annotations) to a large FMA-derived corpus ($\sim$49,300 segments with synthetic labels). The second validates computational flavor targets -- derived from food chemistry via a reproducible pipeline -- against human perception in an online listener study (49~participants, 20~tracks). Results from both experiments converge: the quantitative transfer analysis confirms that cross-modal structure is preserved across supervision regimes, and the perceptual evaluation shows significant alignment between computational targets and listener ratings (permutation $p<0.0001$, Mantel $r=0.45$, Procrustes $m^2=0.51$). Together, these findings support the conclusion that sonic seasoning effects are present in synthetic FMA annotations. We release datasets and companion code to support reproducible cross-modal AI research.


【11】BMdataset: A Musicologically Curated LilyPond Dataset
标题:BM数据集:一个音乐学策划的LilyPond数据集
链接:https://arxiv.org/abs/2604.10628

作者:Matteo Spanio,Ilay Guler,Antonio Rodà
备注:Submitted to SMC2026
摘要:符号音乐的研究几乎完全依赖于基于MIDI的数据集;基于文本的雕刻格式,如LilyPond,仍然没有被探索用于音乐理解。我们介绍了BM数据集,这是一个由专家直接从原始巴洛克手稿中转录的393个LilyPond乐谱(2,646个乐章)的音乐学策划数据集,元数据涵盖作曲家,音乐形式,乐器和部分属性。在此资源的基础上,我们介绍了LilyBERT(权重可以在https://huggingface.co/csc-unipd/lilybert上找到),这是一种基于CodeBERT的编码器,通过115个LilyPond特定令牌的词汇扩展和掩码语言模型预训练来适应符号音乐。对域外Mutopia语料库的线性探测表明,尽管其大小适中(约9000万个标记),但单独对BM数据集进行微调在作曲家和风格分类方面都优于对完整的PDMX语料库(约15 B个标记)进行连续预训练,这表明小型,专业策划的数据集可以比大型,嘈杂的语料库更有效地理解音乐。将广泛的预训练与特定领域的微调相结合,总体上产生了最好的结果(84.3%的作曲家准确率),证实了这两种数据体系是互补的。我们发布了数据集、标记器和模型,以建立LilyPond上表示学习的基线。
摘要:Symbolic music research has relied almost exclusively on MIDI-based datasets; text-based engraving formats such as LilyPond remain unexplored for music understanding. We present BMdataset, a musicologically curated dataset of 393 LilyPond scores (2,646 movements) transcribed by experts directly from original Baroque manuscripts, with metadata covering composer, musical form, instrumentation, and sectional attributes. Building on this resource, we introduce LilyBERT (weights can be found at https://huggingface.co/csc-unipd/lilybert), a CodeBERT-based encoder adapted to symbolic music through vocabulary extension with 115 LilyPond-specific tokens and masked language model pre-training. Linear probing on the out-of-domain Mutopia corpus shows that, despite its modest size (~90M tokens), fine-tuning on BMdataset alone outperforms continuous pre-training on the full PDMX corpus (~15B tokens) for both composer and style classification, demonstrating that small, expertly curated datasets can be more effective than large, noisy corpora for music understanding. Combining broad pre-training with domain-specific fine-tuning yields the best results overall (84.3% composer accuracy), confirming that the two data regimes are complementary. We release the dataset, tokenizer, and model to establish a baseline for representation learning on LilyPond.


【12】Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark
标题:知道要强调什么:有话语条件的文本到语音基准
链接:https://arxiv.org/abs/2604.10580

作者:Arnon Turetzky,Avihu Dekel,Hagai Aronowitz,Ron Hoory,Yossi Adi
备注:Preprint
摘要:口语的意义往往不仅取决于说了什么,还取决于哪个词被强调了。同一个句子可以传达纠正、对比或澄清,这取决于强调的位置。虽然现代文语转换(TTS)系统产生表达性的语音,目前还不清楚他们是否推断上下文适当的压力,从话语本身。为了解决这一差距,我们提出了上下文感知的压力TTS(CAST),一个基准评估上下文条件的词级压力在TTS。项目被定义为对比上下文对:相同的句子与不同的上下文配对,需要不同的重读单词。我们评估了最先进的系统,并发现了一个一致的差距:纯文本语言模型可靠地恢复预期的压力从上下文,但TTS系统经常无法实现它的讲话。我们发布了基准,评估框架,建设管道和合成语料库,以支持未来的工作上下文感知语音合成。
摘要:Spoken meaning often depends not only on what is said, but also on which word is emphasized. The same sentence can convey correction, contrast, or clarification depending on where emphasis falls. Although modern text-to-speech (TTS) systems generate expressive speech, it remains unclear whether they infer contextually appropriate stress from discourse alone. To address this gap, we present Context-Aware Stress TTS (CAST), a benchmark for evaluating context-conditioned word-level stress in TTS. Items are defined as contrastive context pairs: identical sentences paired with distinct contexts requiring different stressed words. We evaluate state-of-the-art systems and find a consistent gap: text-only language models reliably recover the intended stress from context, yet TTS systems frequently fail to realize it in speech. We release the benchmark, evaluation framework, construction pipeline and a synthetic corpus to support future work on context-aware speech synthesis.


【13】VidAudio-Bench: Benchmarking V2A and VT2A Generation across Four Audio Categories
标题:VidAudio-Bench:对四种音频类别的V2 A和VT 2A世代进行基准测试
链接:https://arxiv.org/abs/2604.10542

作者:Qian Zhang,Yuqin Cao,Yixuan Gao,Xiongkuo Min
摘要:视频到音频(V2 A)生成对于沉浸式多媒体体验至关重要,但其评估仍有待深入研究。现有的基准通常在统一的协议下评估不同的音频类型,忽略了不同音频类别的细粒度要求。为了解决这一差距,我们提出了VidAudio-Bench,这是一个用于V2 A评估的多任务基准,具有四个关键特征:(1)广泛的覆盖范围:它包括四个代表性的音频类别-音效,音乐,演讲和唱歌-在V2 A和视频文本到音频(VT 2A)设置下。(2)广泛的评估:它包括1,634个视频文本对和基准11个最先进的生成模型。(3)综合评分:它引入了13个特定于任务的无参考指标,以系统地评估音频质量,视频-音频一致性和文本-音频一致性。(4)人类一致性:它通过主观研究验证所有指标,证明与人类偏好的高度一致性。实验结果表明,目前的V2 A模型在语音和歌唱方面的表现不如音效。我们的VT 2A结果进一步强调了指令遵循和视觉接地生成之间的根本紧张关系:更强的视觉条件反射改善了视频-音频对齐,但通常以生成预期的音频类别为代价。这些发现将VidAudio-Bench确立为诊断V2 A系统的全面且可扩展的框架,并为多模态音频生成提供了新的见解。
摘要:Video-to-Audio (V2A) generation is essential for immersive multimedia experiences, yet its evaluation remains underexplored. Existing benchmarks typically assess diverse audio types under a unified protocol, overlooking the fine-grained requirements of distinct audio categories. To address this gap, we propose VidAudio-Bench, a multi-task benchmark for V2A evaluation with four key features: (1) Broad Coverage: It encompasses four representative audio categories - sound effects, music, speech, and singing - under both V2A and Video-Text-to-Audio (VT2A) settings. (2) Extensive Evaluation: It comprises 1,634 video-text pairs and benchmarks 11 state-of-the-art generation models. (3) Comprehensive Metrics: It introduces 13 task-specific, reference-free metrics to systematically assess audio quality, video-audio consistency, and text-audio consistency. (4) Human Alignment: It validates all metrics through subjective studies, demonstrating strong consistency with human preferences. Experimental results reveal that current V2A models perform poorly in speech and singing compared to sound effects. Our VT2A results further highlight a fundamental tension between instruction following and visually grounded generation: stronger visual conditioning improves video-audio alignment, but often at the cost of generating the intended audio category. These findings establish VidAudio-Bench as a comprehensive and scalable framework for diagnosing V2A systems and provide new insights into multimodal audio generation.


【14】Cross-Cultural Bias in Mel-Scale Representations: Evidence and Alternatives from Speech and Music
标题:梅尔音阶表示中的跨文化偏见:来自言语和音乐的证据和替代方案
链接:https://arxiv.org/abs/2604.10503

作者:Shivam Chauhan,Ajay Pundhir
备注:5 pages, 3 figures, 4 tables. Accepted at ICASSP 2026
摘要:现代音频系统普遍采用源自20世纪40年代西方心理声学研究的梅尔音阶表示,可能会编码造成系统性能差异的文化偏见。我们提出了一个全面的评估音频前端的跨文化偏见,比较梅尔规模的功能与可学习的替代品(叶,SincNet)和心理声学变体(ERB,树皮,CQT)在语音识别(11种语言),音乐分析(6个集合),和欧洲声学场景分类(10个欧洲城市)。我们的受控实验隔离了前端贡献,同时保持架构和训练协议最小化和恒定。结果表明,梅尔规模的功能产生31.2%的WER的音调语言相比,18.7%的非音调语言(12.5%的差距),并显示15.7%的F1退化之间的西方和非西方音乐。替代表示显著地减少了这些差异:LEAF通过自适应频率分配将语音间隙减少了34%,CQT实现了音乐性能间隙减少了52%,ERB尺度滤波仅用1%的计算开销将差异减少了31%。我们还发布了FairAudioBench,支持跨文化评估,并证明自适应频率分解为公平的音频处理提供了实用的途径。这些发现揭示了基础信号处理选择如何传播偏见,为开发包容性音频系统提供了重要指导。
摘要:Modern audio systems universally employ mel-scale representations derived from 1940s Western psychoacoustic studies, potentially encoding cultural biases that create systematic performance disparities. We present a comprehensive evaluation of cross-cultural bias in audio front-ends, comparing mel-scale features with learnable alternatives (LEAF, SincNet) and psychoacoustic variants (ERB, Bark, CQT) across speech recognition (11 languages), music analysis (6 collections), and European acoustic scene classification (10 European cities). Our controlled experiments isolate front-end contributions while holding architecture and training protocols minimal and constant. Results demonstrate that mel-scale features yield 31.2% WER for tonal languages compared to 18.7% for non-tonal languages (12.5% gap), and show 15.7% F1 degradation between Western and non-Western music. Alternative representations significantly reduce these disparities: LEAF reduces the speech gap by 34% through adaptive frequency allocation, CQT achieves 52% reduction in music performance gaps, and ERB-scale filtering cuts disparities by 31% with only 1% computational overhead. We also release FairAudioBench, enabling cross-cultural evaluation, and demonstrate that adaptive frequency decomposition offers practical paths toward equitable audio processing. These findings reveal how foundational signal processing choices propagate bias, providing crucial guidance for developing inclusive audio systems.


【15】Whisper-AuT: Domain-Adapted Audio Encoder for Efficient Audio-LLM Training
标题:Whisper-AuT:用于高效音频LLM训练的域自适应音频编码器
链接:https://arxiv.org/abs/2604.10438

作者:Jielin Qiu,Ming Zhu,Wenting Zhao,Zhiwei Liu,Liangwei Yang,Zixiang Chen,Roshan Ram,Akshara Prabhakar,Juntao Tan,Rithesh Murthy,Shelby Heinecke,Caiming Xiong,Silvio Savarese,Huan Wang
摘要:音频原生大型语言模型(audio-LLM)通常使用Whisper作为其音频编码器。然而,Whisper只接受语音数据的训练,对音乐和环境声音的表现很弱。这迫使下游音频LLM通过对大规模非语音数据的广泛训练进行补偿。我们提出了Whisper-AuT,这是一种域适应的音频编码器,通过对语音(80%),环境声音(10%)和音乐(10%)的策划混合物进行微调而获得。完整的编码器-解码器使用seq 2seq字幕目标进行端到端训练;然后丢弃解码器,只保留编码器。线性探头评估显示,与原始Whisperlarge-v3编码器相比,Whisper-AuT在ESC-50(环境声音)上实现了+23.0%,在GTZAN(音乐类型)上实现了+5.0%,在语音命令(关键字定位)上实现了+0.7%。Whisper-AuT被设计为音频LLM架构中Whisper的直接替代品,其目标是通过为非语音域提供更强的初始音频表示来降低下游训练成本。
摘要:Audio-native large language models (audio-LLMs) commonly use Whisper as their audio encoder. However, Whisper was trained exclusively on speech data, producing weak representations for music and environmental sound. This forces downstream audio-LLMs to compensate through extensive training on large-scale non-speech data. We present Whisper-AuT, a domain-adapted audio encoder obtained by fine-tuning Whisper-large-v3 on a curated mixture of speech (80%), environmental sound (10%), and music (10%) totaling approximately 20M samples. The full encoder-decoder is trained end-to-end with a seq2seq captioning objective; the decoder is then discarded and only the encoder is retained. Linear probe evaluations show that Whisper-AuT achieves +23.0% on ESC-50 (environmental sound), +5.0% on GTZAN (music genre), and +0.7% on Speech Commands (keyword spotting) compared to the original Whisperlarge-v3 encoder. Whisper-AuT is designed as a drop-in replacement for Whisper in audio-LLM architectures, with the goal of reducing downstream training cost by providing stronger initial audio representations for non-speech domains.


【16】Sign-to-Speech Prosody Transfer via Sign Reconstruction-based GAN
标题:基于符号重建的GAN的符号到语音韵律传输
链接:https://arxiv.org/abs/2604.10413

作者:Toranosuke Manabe,Yuto Shibata,Shinnosuke Takamichi,Yoshimitsu Aoki
备注:Accepted to ICPR 2026
摘要:深度学习模型改进了手语到文本的翻译,使非签名者更容易理解签名消息。当目标是口头通信时,一种简单的方法是将签名消息转换为文本,然后通过文本到语音(TTS)合成语音。然而,这种两阶段的管道不可避免地将文本视为瓶颈表示,导致原本在签名中传达的丰富的非语言信息丢失。为了解决这个问题,我们提出了一个新的任务,\endash {Sign-to-Speech Prosody Transfer},其目的是捕捉手语中表达的全局韵律细微差别,并直接将它们集成到合成语音中。一个主要的挑战是,对齐符号和语音需要专业知识,使得注释成本极高,并阻止大型并行语料库的构建。为了克服这一点,我们引入了SignRecGAN,这是一个可扩展的训练框架,它通过对抗性学习和重建损失来利用单峰数据集,而不使用跨模态注释。此外,我们提出了一个新的模型架构,保留了现有的TTS模型的表达能力,同时使符号导出的韵律注入到合成语音。大量的实验表明,所提出的方法可以合成语音,忠实地反映了手语的情感内容,从而打开了新的可能性,更自然的手语交流。我们的代码将在接受后提供。
摘要:Deep learning models have improved sign language-to-text translation and made it easier for non-signers to understand signed messages. When the goal is spoken communication, a naive approach is to convert signed messages into text and then synthesize speech via Text-to-Speech (TTS). However, this two-stage pipeline inevitably treat text as a bottleneck representation, causing the loss of rich non-verbal information originally conveyed in the signing. To address this limitation, we propose a novel task, \emph{Sign-to-Speech Prosody Transfer}, which aims to capture the global prosodic nuances expressed in sign language and directly integrate them into synthesized speech. A major challenge is that aligning sign and speech requires expert knowledge, making annotation extremely costly and preventing the construction of large parallel corpora. To overcome this, we introduce \emph{SignRecGAN}, a scalable training framework that leverages unimodal datasets without cross-modal annotations through adversarial learning and reconstruction losses. Furthermore, we propose \emph{S2PFormer}, a new model architecture that preserves the expressive power of existing TTS models while enabling the injection of sign-derived prosody into the synthesized speech. Extensive experiments demonstrate that the proposed method can synthesize speech that faithfully reflects the emotional content of sign language, thereby opening new possibilities for more natural sign language communication. Our code will be available upon acceptance.


【17】Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware Kernels
标题:超越独白:具有对话音频上下文感知核心的交互式谈话-聆听化身生成
链接:https://arxiv.org/abs/2604.10367

作者:Yuzhe Weng,Haotian Wang,Xinyi Yu,Xiaoyan Wu,Haoran Xu,Shan He,Jun Du
摘要:音频驱动的人类视频生成在独白场景中取得了显着的成功,这主要是由强大的视频生成基础模型的进步所驱动的。超越独白,真实的人类交流本质上是一个全双工的交互过程,要求虚拟代理不仅要表达自己的语音,还要对传入的对话音频做出自然的反应。大多数现有的方法只是将传统的音频驱动范例扩展到收听场景。然而,依赖于严格的帧到帧对齐使得模型对长距离会话动态的响应刚性,而直接引入全局注意力则灾难性地降低了嘴唇同步。认识到说话和倾听行为之间独特的时间尺度差异,我们引入了多头高斯内核,将这种物理直觉作为渐进的时间归纳偏差显式地注入到模型中。在此基础上,我们构建了一个全双工的交互式虚拟代理能够同时处理双流音频输入的谈话和倾听。此外,我们引入了一个严格清理的说话-倾听数据集VoxHear,具有完美解耦的语音和背景音轨。大量的实验表明,我们的方法成功地融合了强大的时间对齐与深上下文语义,设置一个新的国家的最先进的生成高度自然和响应全双工交互式数字人。该项目的网页可在https://warmcongee.github.io/beyond-monologue/上查阅。
摘要:Audio-driven human video generation has achieved remarkable success in monologue scenarios, largely driven by advancements in powerful video generation foundation models. Moving beyond monologues, authentic human communication is inherently a full-duplex interactive process, requiring virtual agents not only to articulate their own speech but also to react naturally to incoming conversational audio. Most existing methods simply extend conventional audio-driven paradigms to listening scenarios. However, relying on strict frame-to-frame alignment renders the model's response to long-range conversational dynamics rigid, whereas directly introducing global attention catastrophically degrades lip synchronization. Recognizing the unique temporal Scale Discrepancy between talking and listening behaviors, we introduce a multi-head Gaussian kernel to explicitly inject this physical intuition into the model as a progressive temporal inductive bias. Building upon this, we construct a full-duplex interactive virtual agent capable of simultaneously processing dual-stream audio inputs for both talking and listening. Furthermore, we introduce a rigorously cleaned Talking-Listening dataset VoxHear featuring perfectly decoupled speech and background audio tracks. Extensive experiments demonstrate that our approach successfully fuses strong temporal alignment with deep contextual semantics, setting a new state-of-the-art for generating highly natural and responsive full-duplex interactive digital humans. The project page is available at https://warmcongee.github.io/beyond-monologue/ .


【18】Descriptor-Injected Cross-Modal Learning: A Systematic Exploration of Audio-MIDI Alignment via Spectral and Melodic Features
标题:描述符注入的跨模式学习:通过频谱和旋律特征对音频与音频对齐的系统探索
链接:https://arxiv.org/abs/2604.10283

作者:Mariano Fernández Méndez
备注:26 pages, 11 figures, 20 tables. Companion paper to "Harmonic Information Theory: Foundations" (2026). Code: https://github.com/AlterMundi/Phideus
摘要:音频记录和符号音乐表征(Symbolic Music Representation,缩写为MIDI)之间的跨模态检索仍然具有挑战性,因为连续波形和离散事件序列编码相同性能的不同方面。我们研究描述符注入,增强特定于模态的编码器与手工制作的域功能,作为跨越这一差距的桥梁。在一个包含13个减速器-机构组合、6个体系结构家族和3个训练计划的三阶段活动中,最佳配置在5个独立种子中的平均S达到84.0%,将无减速器基线提高了8.8个百分点。因果关系消除表明,基于八度带能量动态的音频描述符A4驱动了顶级双模型的增益,而MIDI描述符D4尽管改善了训练动态,但仅具有较弱的推理时间效应。我们还引入了反向交叉注意,其中描述符令牌查询编码器功能,减少注意操作相对于标准制定,同时保持竞争力。CKA分析表明,描述符大大增加了音频压缩Transformer层对齐,表明代表性的收敛,而不是简单的功能级联。扰动分析识别高频倍频程带作为主要的鉴别信号。所有实验均使用MAESTRO v3.0.0,并采用控制作曲家和作品相似性的评估方案。
摘要:Cross-modal retrieval between audio recordings and symbolic music representations (MIDI) remains challenging because continuous waveforms and discrete event sequences encode different aspects of the same performance. We study descriptor injection, the augmentation of modality-specific encoders with hand-crafted domain features, as a bridge across this gap. In a three-phase campaign covering 13 descriptor-mechanism combinations, 6 architectural families, and 3 training schedules, the best configuration reaches a mean S of 84.0 percent across five independent seeds, improving the descriptor-free baseline by 8.8 percentage points. Causal ablation shows that the audio descriptor A4, based on octave-band energy dynamics, drives the gain in the top dual models, while the MIDI descriptor D4 has only a weak inference-time effect despite improving training dynamics. We also introduce reverse cross-attention, where descriptor tokens query encoder features, reducing attention operations relative to the standard formulation while remaining competitive. CKA analysis shows that descriptors substantially increase audio-MIDI transformer layer alignment, indicating representational convergence rather than simple feature concatenation. Perturbation analysis identifies high-frequency octave bands as the dominant discriminative signal. All experiments use MAESTRO v3.0.0 with an evaluation protocol controlling for composer and piece similarity.


【19】Learning to Attend to Depression-Related Patterns: An Adaptive Cross-Modal Gating Network for Depression Detection
标题:学习关注抑郁相关模式:用于抑郁检测的自适应跨模式门控网络
链接:https://arxiv.org/abs/2604.10181

作者:Hangbin Yu,Yudong Yang,Rongfeng Su,Nan Yan,Lan Wang
摘要:使用具有声学和文本模态的语音信号自动检测抑郁症是一种很有前途的早期诊断方法。抑郁症相关的模式在语音中表现出稀疏性:诊断相关的特征出现在特定的片段中,而不是均匀分布。然而,大多数现有的方法平等地对待所有帧,假设抑郁相关的信息是均匀分布的,从而忽略了这种稀疏性。为了解决这个问题,我们提出了一种基于自适应跨模态门控(ACMG)的抑郁检测网络,该网络自适应地在两种模态之间重新分配帧级权重,从而能够选择性地关注抑郁相关的片段。实验结果表明,与ACMG的抑郁症检测系统优于基线没有它。可视化分析进一步证实,ACMG自动出席临床有意义的模式,包括低能量的声学段和文本段含有负面情绪。
摘要:Automatic depression detection using speech signals with acoustic and textual modalities is a promising approach for early diagnosis. Depression-related patterns exhibit sparsity in speech: diagnostically relevant features occur in specific segments rather than being uniformly distributed. However, most existing methods treat all frames equally, assuming depression-related information is uniformly distributed and thus overlooking this sparsity. To address this issue, we proposes a depression detection network based on Adaptive Cross-Modal Gating (ACMG) that adaptively reassigns frame-level weights across both modalities, enabling selective attention to depression-related segments. Experimental results show that the depression detection system with ACMG outperforms baselines without it. Visualization analyses further confirm that ACMG automatically attends to clinically meaningful patterns, including low-energy acoustic segments and textual segments containing negative sentiments.


【20】From Speech to Profile: A Protocol-Driven LLM Agent for Psychological Profile Generation
标题:从语音到个人资料:用于心理个人资料生成的协议驱动LLM代理
链接:https://arxiv.org/abs/2604.10161

作者:Xingjian Yang,Yudong Yang,Zhixing Guo,Yongjie Zhou,Nan Yan,Lan Wang
摘要:从结构上记录抑郁症患者的心理特征对心理治疗至关重要。大型语言模型可用于从咨询演讲中总结轮廓,但由于演讲长度过长,多方互动和非结构化聊天,它可能会遭受长上下文遗忘并产生无法验证的幻觉。据此,我们提出了一个StreamProfile,一个流媒体框架,增量处理咨询语音,通过将其存储在分层证据存储器中从ASR transmittance中提取证据,然后根据PM+心理干预进行临床推理的思想链管道。最终的轮廓是严格从这些证据合成的,使每一个索赔都有迹可循。对真实世界的青少年咨询语音的实验表明,所提出的StreamProfile系统可以准确地生成配置文件,并防止幻觉。
摘要:The psychological profile that structurally documents the case of a depression patient is essential for psychotherapy. Large language models can be applied to summarize the profiles from counseling speech, however, it may suffer from long-context forgetting and produce unverifiable hallucinations, due to overlong length of speech, multi-party interactions and unstructured chatting. Hereby, we propose a StreamProfile, a streaming framework that processes counseling speech incrementally, extracts evidences grounded from ASR transcriptions by storing it in a Hierarchical Evidence Memory, and then performs a Chain-of-Thought pipeline according to PM+ psychological intervention for clinical reasoning. The final profile is synthesized strictly from those evidences, making every claim traceable. Experiments on real-world teenager counseling speech have shown that the proposed StreamProfile system can accurately generate the profiles and prevent hallucination.


【21】ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models
标题:ASPIRin:用于全复式语音语言模型中交互优化强化学习的动作空间投影
链接:https://arxiv.org/abs/2604.10065

作者:Chi-Yuan Hsiao,Ke-Han Lu,Yu-Kuan Fu,Guan-Ting Lin,Hsiao-Tsung Hung,Hung-yi Lee
摘要:端到端全双工语音语言模型(SLM)需要精确的话轮转换以实现自然交互。然而,通过标准的原始令牌强化学习(RL)优化时间动态会降低语义质量,导致严重的生成崩溃和重复。我们提出了ASPIRin,这是一个交互优化的RL框架,它明确地解释了什么时候说什么。使用动作空间投影,ASPIRin将文本词汇映射到粗粒度的二进制状态(主动语音与非主动沉默)。通过应用具有基于规则的奖励的组相对策略优化(GRPO),它平衡了用户中断和响应延迟。实证评估表明,ASPIRin优化了话轮转换、反向引导和暂停处理的交互性。至关重要的是,将计时与标记选择隔离开来,可以保持语义一致性,并将重复的n元语法的部分与标准GRPO相比减少了50%以上,有效地消除了退化重复。
摘要:End-to-end full-duplex Speech Language Models (SLMs) require precise turn-taking for natural interaction. However, optimizing temporal dynamics via standard raw-token reinforcement learning (RL) degrades semantic quality, causing severe generative collapse and repetition. We propose ASPIRin, an interactivity-optimized RL framework that explicitly decouples when to speak from what to say. Using Action Space Projection, ASPIRin maps the text vocabulary into a coarse-grained binary state (active speech vs. inactive silence). By applying Group Relative Policy Optimization (GRPO) with rule-based rewards, it balances user interruption and response latency. Empirical evaluations show ASPIRin optimizes interactivity across turn-taking, backchanneling, and pause handling. Crucially, isolating timing from token selection preserves semantic coherence and reduces the portion of duplicate n-grams by over 50% compared to standard GRPO, effectively eliminating degenerative repetition.


【22】Cross-Validated Cross-Channel Self-Attention and Denoising for Automatic Modulation Classification
标题:交叉验证的跨通道自注意和去噪用于自动调制分类
链接:https://arxiv.org/abs/2604.10054

作者:Prakash Suman,Yanzhen Qu
摘要:这项研究解决了深度学习自动调制分类(AMC)模型的一个关键限制,该模型在高信噪比(SNR)下表现良好,但在噪声条件下性能下降,因为传统的特征提取抑制了区分结构和干扰。我们的目标是开发一种保留特征的去噪方法,以减轻调制类分离的损失。提出了一种深度学习AMC模型,该模型包含一个跨通道自注意块来捕获同相和正交分量之间的依赖关系,以及双路径深度残差收缩去噪块来抑制噪声。使用RML2018.01a数据集的实验采用了24种调制类型和26种SNR水平的分层采样。结果表明,去噪深度强烈影响鲁棒性在低和中等信噪比。与基准模型PET-CGDNN、MCLDNN和DAE相比,所提出的模型在-8 dB到+2 dB SNR范围内实现了显著的精度提高,分别提高了3%、2.3%和14%。交叉验证证实了该模型的鲁棒性,平均准确率为62.6%,宏观精度为65.8%,宏观召回率为62.6%,宏观F1得分为62.9%。该架构通过将基带建模形式化为正交子问题并引入跨信道注意力作为广义复杂交互算子来推进干扰感知AMC,消融确认了特征保留去噪在中低SNR下的鲁棒性的关键作用。
摘要:This study addresses a key limitation in deep learning Automatic Modulation Classification (AMC) models, which perform well at high signal-to-noise ratios (SNRs) but degrade under noisy conditions due to conventional feature extraction suppressing both discriminative structure and interference. The goal was to develop a feature-preserving denoising method that mitigates the loss of modulation class separation. A deep learning AMC model was proposed, incorporating a cross-channel self-attention block to capture dependencies between in-phase and quadrature components, along with dual-path deep residual shrinkage denoising blocks to suppress noise. Experiments using the RML2018.01a dataset employed stratified sampling across 24 modulation types and 26 SNR levels. Results showed that denoising depth strongly influences robustness at low and moderate SNRs. Compared to benchmark models PET-CGDNN, MCLDNN, and DAE, the proposed model achieved notable accuracy improvements across -8 dB to +2 dB SNR, with increases of 3%, 2.3%, and 14%, respectively. Cross-validation confirmed the model's robustness, yielding a mean accuracy of 62.6%, macro precision of 65.8%, macro-recall of 62.6%, and macro-F1 score of 62.9%. The architecture advances interference-aware AMC by formalizing baseband modeling as orthogonal subproblems and introducing cross-channel attention as a generalized complex interaction operator, with ablations confirming the critical role of feature-preserving denoising for robustness at low-to-medium SNR.


【23】Masked Contrastive Pre-Training Improves Music Audio Key Detection
标题:掩蔽对比预训练改进了音乐音频密钥检测
链接:https://arxiv.org/abs/2604.10021

作者:Ori Yonay,Tracy Hammond,Tianbao Yang
备注:Code and models available at github.com/echo-cipher/keymyna
摘要:自我监督的音乐基础模型在关键检测上表现不佳,这需要音高敏感的表示。在这项工作中,我们提出了第一个系统的研究,表明自监督预训练的设计直接影响音高敏感性,并证明了掩蔽对比嵌入独特地使最先进的(SOTA)性能在监督设置中的关键检测。首先,我们发现在Mel谱图上进行基于掩蔽的对比预训练后的线性评估会导致开箱即用的音乐键检测的竞争性能。这使我们训练浅,但广泛的多层感知器(MLP)的功能提取从我们的基础模型,导致SOTA性能,而不需要复杂的数据增强策略。我们进一步分析了鲁棒性,并经验表明,学习表示自然编码常见的增强。我们的研究建立了自我监督的预训练作为音高敏感的MIR任务的有效方法,并为设计和探索音乐基础模型提供了见解。
摘要:Self-supervised music foundation models underperform on key detection, which requires pitch-sensitive representations. In this work, we present the first systematic study showing that the design of self-supervised pretraining directly impacts pitch sensitivity, and demonstrate that masked contrastive embeddings uniquely enable state-of-the-art (SOTA) performance in key detection in the supervised setting. First, we discover that linear evaluation after masking-based contrastive pretraining on Mel spectrograms leads to competitive performance on music key detection out of the box. This leads us to train shallow but wide multi-layer perceptrons (MLPs) on features extracted from our base model, leading to SOTA performance without the need for sophisticated data augmentation policies. We further analyze robustness and show empirically that the learned representations naturally encode common augmentations. Our study establishes self-supervised pretraining as an effective approach for pitch-sensitive MIR tasks and provides insights for designing and probing music foundation models.


【24】MAGE: Modality-Agnostic Music Generation and Editing
标题:MAGE:情态不可知的音乐生成和编辑
链接:https://arxiv.org/abs/2604.09803

作者:Muhammad Usama Saleem,Tejasvi Ravi,Tianyu Xu,Rajeev Nongpiur,Ishan Chatterjee,Mayur Jagdishbhai Patel,Pu Wang
摘要:多模态音乐创作需要能够从高级线索生成音频并以有针对性的方式编辑现有混合的模型。然而,大多数多模态音乐系统都是为单一任务和固定的提示界面而构建的,当指导模糊不清、时间错位或部分缺失时,它们的条件反射就变得脆弱。常见的附加融合或特征串联进一步削弱了跨模态的基础,通常会在生成和编辑过程中导致快速漂移和虚假的音乐内容。我们提出了MAGE,模态不可知的框架,统一的多模态音乐生成和混合接地编辑在一个单一的连续的潜在配方。在其核心,MAGE使用一个受控的多模态通量形成器,一个基于流的Transformer,学习可控的潜在轨迹,在任何可用的条件子集下进行合成和编辑。为了改善接地,我们引入视听连接对齐选择时间上一致的视觉证据的音频时间轴,和交叉门控调制机制,应用乘法控制对齐的视觉和文本线索的音频潜伏,抑制不支持的组件,而不是注入它们。最后,我们使用动态模态掩蔽课程进行训练,该课程将模型暴露于纯文本、纯视觉、联合多模态和混合引导设置,从而在缺失模态的情况下实现鲁棒推理,而无需训练单独的模型。MUSIC基准测试的实验表明,MAGE支持有效的多模式引导音乐生成和有针对性的编辑,实现有竞争力的质量,同时提供针对实用音乐工作流程量身定制的轻量级且灵活的界面。
摘要:Multimodal music creation requires models that can both generate audio from high-level cues and edit existing mixtures in a targeted manner. Yet most multimodal music systems are built for a single task and a fixed prompting interface, making their conditioning brittle when guidance is ambiguous, temporally misaligned, or partially missing. Common additive fusion or feature concatenation further weakens cross-modal grounding, often causing prompt drift and spurious musical content during generation and editing. We propose MAGE, a modality-agnostic framework that unifies multimodal music generation and mixture-grounded editing within a single continuous latent formulation. At its core, MAGE uses a Controlled Multimodal FluxFormer, a flow-based Transformer that learns controllable latent trajectories for synthesis and editing under any available subset of conditions. To improve grounding, we introduce Audio-Visual Nexus Alignment to select temporally consistent visual evidence for the audio timeline, and a cross-gated modulation mechanism that applies multiplicative control from aligned visual and textual cues to the audio latents, suppressing unsupported components rather than injecting them. Finally, we train with a dynamic modality-masking curriculum that exposes the model to text-only, visual-only, joint multimodal, and mixture-guided settings, enabling robust inference under missing modalities without training separate models. Experiments on the MUSIC benchmark show that MAGE supports effective multimodal-guided music generation and targeted editing, achieving competitive quality while offering a lightweight and flexible interface tailored to practical music workflows.


【25】Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering
标题:Jamendo-MT-QA:多轨比较音乐问题回答的基准
链接:https://arxiv.org/abs/2604.09721

作者:Junyoung Koh,Jaeyun Lee,Soo Yong Kim,Gyu Hyeong Choi,Jung In Koh,Jordan Phillips,Yeonjin Lee,Min Song
备注:ACL 2026 Findings
摘要:最近关于音乐问答(Music-QA)的工作主要集中在单音轨理解上,其中模型使用其标签,标题或元数据回答关于单个音频片段的问题。然而,听众经常以比较的方式描述音乐,现有的基准并没有系统地评估多个轨道的推理。在Jamendo-QA数据集的基础上,我们介绍了Jamendo-MT-QA,这是一个多轨比较问题回答的数据集和基准。从Jamendo上的Creative Commons许可轨道,我们构建了36,519个比较QA项目,超过12,173个轨道对,每对产生三种问题类型:是/否,简短回答和重复级别问题。我们描述了一个LLM辅助管道,用于生成和过滤比较问题,并使用自动度量和LLM作为法官评估基准代表性的音频语言模型。
摘要:Recent work on music question answering (Music-QA) has primarily focused on single-track understanding, where models answer questions about an individual audio clip using its tags, captions, or metadata. However, listeners often describe music in comparative terms, and existing benchmarks do not systematically evaluate reasoning across multiple tracks. Building on the Jamendo-QA dataset, we introduce Jamendo-MT-QA, a dataset and benchmark for multi-track comparative question answering. From Creative Commons-licensed tracks on Jamendo, we construct 36,519 comparative QA items over 12,173 track pairs, with each pair yielding three question types: yes/no, short-answer, and sentence-level questions. We describe an LLM-assisted pipeline for generating and filtering comparative questions, and benchmark representative audio-language models using both automatic metrics and LLM-as-a-Judge evaluation.


【26】Real-Time Voicemail Detection in Telephony Audio Using Temporal Speech Activity Features
标题:使用时间语音活动特征进行电话音频中的实时语音邮件检测
链接:https://arxiv.org/abs/2604.09675

作者:Kumar Saurav
备注:16 pages, 5 tables. Preprint
摘要:人工智能呼叫系统必须实时区分语音邮件问候和真人应答,以避免浪费座席交互和掉线。我们提出了一种轻量级的方法,从预先训练的神经语音活动检测器(VAD)的语音活动模式中提取15个时间特征,然后使用基于浅树的集成进行分类。在总共764个电话录音的两个评估集上,该系统实现了96.1%的准确率(734/764),其中专家标记的测试集为99.3%(139/140),生产集为95.4%(595/624)。在超过77,000个调用的生产验证中,它保持了0.3%的假阳性率和1.3%的假阴性率。端到端推理在没有GPU的商用双核CPU上在46 ms内完成,支持380+并发WebSocket调用。在我们对3,780多个模型、特征和阈值组合的搜索中,特征重要性集中在三个时间变量上。添加转录关键字或基于蜂鸣音的功能并没有改善最佳实时配置,反而大大增加了延迟。我们的研究结果表明,时间语音模式是一个强有力的信号区分语音邮件问候从真人的答案。
摘要:Outbound AI calling systems must distinguish voicemail greetings from live human answers in real time to avoid wasted agent interactions and dropped calls. We present a lightweight approach that extracts 15 temporal features from the speech activity pattern of a pre-trained neural voice activity detector (VAD), then classifies with a shallow tree-based ensemble. Across two evaluation sets totaling 764 telephony recordings, the system achieves a combined 96.1% accuracy (734/764), with 99.3% (139/140) on an expert-labeled test set and 95.4% (595/624) on a held-out production set. In production validation over 77,000 calls, it maintained a 0.3% false positive rate and 1.3% false negative rate. End-to-end inference completes in 46 ms on a commodity dual-core CPU with no GPU, supporting 380+ concurrent WebSocket calls. In our search over 3,780 model, feature, and threshold combinations, feature importance was concentrated in three temporal variables. Adding transcription keywords or beep-based features did not improve the best real-time configuration and increased latency substantially. Our results suggest that temporal speech patterns are a strong signal for distinguishing voicemail greetings from live human answers.


【27】Sink or SWIM: Tackling Real-Time ASR at Scale
标题:水槽或游泳:大规模解决实时ZR问题
链接:https://arxiv.org/abs/2601.17097

作者:Federico Bruzzone,Walter Cazzola,Matteo Brancaleoni,Dario Pellegrino
备注:14 pages, 7 figures
摘要:实时自动语音识别系统越来越多地集成到交互式应用中,从语音助理到实时转录服务。然而,扩展这些系统以支持多个并发客户端,同时保持低延迟和高准确性仍然是一个重大挑战。在这项工作中,我们提出了SWIM,这是一种建立在OpenAI Whisper模型之上的新型实时ASR系统,可以实现真正的模型级并行化,以实现可扩展的多语言转录。SWIM支持多个并发音频流,而无需修改底层模型。它引入了一种缓冲区合并策略,在确保高效资源使用的同时保持转录保真度。我们评估SWIM在多客户端设置-扩展到20个并发用户-并表明,它提供了准确的实时transmittance在英语,意大利语和西班牙语,同时保持低延迟和高吞吐量。虽然Whisper-Streaming在单客户端、仅英语设置中实现了约8.2%的单词错误率和约3.4秒的平均延迟,但SWIM将此功能扩展到多语言、多客户端环境。它保持了相当的准确性和显著更低的延迟(5个客户端约2.4秒),并继续有效地扩展到20个并发客户端,而不会降低转录质量和提高整体吞吐量。我们的方法通过提高动态多用户环境中的鲁棒性和效率来推进可扩展的ASR。
摘要:Real-time automatic speech recognition systems are increasingly integrated into interactive applications, from voice assistants to live transcription services. However, scaling these systems to support multiple concurrent clients while maintaining low latency and high accuracy remains a major challenge. In this work, we present SWIM, a novel real-time ASR system built on top of OpenAI's Whisper model that enables true model-level parallelization for scalable, multilingual transcription. SWIM supports multiple concurrent audio streams without modifying the underlying model. It introduces a buffer merging strategy that maintains transcription fidelity while ensuring efficient resource usage. We evaluate SWIM in multi-client settings -- scaling up to 20 concurrent users -- and show that it delivers accurate real-time transcriptions in English, Italian, and Spanish, while maintaining low latency and high throughput. While Whisper-Streaming achieves a word error rate of approximately 8.2% with an average delay of approximately 3.4 s in a single-client, English-only setting, SWIM extends this capability to multilingual, multi-client environments. It maintains comparable accuracy with significantly lower delay -- around 2.4 s with 5 clients -- and continues to scale effectively up to 20 concurrent clients without degrading transcription quality and increasing overall throughput. Our approach advances scalable ASR by improving robustness and efficiency in dynamic, multi-user environments.


【28】HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models
标题:HumDial-EIBench:音频语言模型的人类记录的多回合情商基准
链接:https://arxiv.org/abs/2604.11594

作者:Shuiyuan Wang,Zhixian Zhao,Hongfei Yue,Chengyou Wang,Shuai Wang,Hui Bu,Xin Xu,Lei Xie
摘要:评估音频语言模型(ALM)的情绪智力(EI)是至关重要的。然而,现有的基准主要依赖于合成语音,仅限于单圈交互,并严重依赖于开放式评分。本文提出了HumDial-EIBench,一个全面的基准评估ALMs的EI。它使用ICASSP 2026 HumDial挑战赛中的真实记录的人类对话,将情感跟踪和因果推理重新制定为带有对抗性干扰的多项选择题,减轻认知任务的主观评分偏差。它保留了移情反应的生成,并引入了一个声学语义冲突任务,以评估对矛盾的多模态信号的鲁棒性。对八个ALMs的评估表明,大多数模型都难以进行多轮情感跟踪和隐式因果推理。此外,所有模型都表现出解耦的文本和声学移情,以及在跨模态冲突期间严重的文本主导偏见。
摘要:Evaluating the emotional intelligence (EI) of audio language models (ALMs) is critical. However, existing benchmarks mostly rely on synthesized speech, are limited to single-turn interactions, and depend heavily on open-ended scoring. This paper proposes HumDial-EIBench, a comprehensive benchmark for evaluating ALMs' EI. Using real-recorded human dialogues from the ICASSP 2026 HumDial Challenge, it reformulates emotional tracking and causal reasoning into multiple-choice questions with adversarial distractors, mitigating subjective scoring bias for cognitive tasks. It retains the generation of empathetic responses and introduces an acoustic-semantic conflict task to assess robustness against contradictory multimodal signals. Evaluations of eight ALMs reveal that most models struggle with multi-turn emotional tracking and implicit causal reasoning. Furthermore, all models exhibit decoupled textual and acoustic empathy, alongside a severe text-dominance bias during cross-modal conflicts.


【29】Speech-preserving active noise control: a deep learning approach in reverberant environments
标题:语音保留主动噪音控制:回响环境中的深度学习方法
链接:https://arxiv.org/abs/2604.10979

作者:Shuning Dai
备注:89 pages, 17 figures, master's dissertation
摘要:传统的有源噪声控制(ANC)系统大多基于FxLMS算法,但此类算法依赖于线性假设,并且通常在处理宽带非平稳噪声或非线性声学路径方面受到限制。不仅如此,传统的方法是将所有信号一起消除,降噪往往会意外地破坏语音信号,影响正常通信。为了解决这些问题,本研究提出了一种语音保留深度学习ANC系统,旨在实现稳定的降噪,同时在复杂的声学环境中有效地保留语音。   本研究建立了一个端到端的控制架构,其核心采用卷积递归网络(CRN)。该结构使用长短期记忆(LSTM)网络来捕获声信号的时间相关特征。结合复谱映射技术,有效地解决了非线性失真问题。为了在去除噪声的同时保留有用的语音,本研究还设计了一个特殊的语音保留损失函数。该设计指导模型通过识别频谱结构的特性来选择性地保留目标语音,同时抑制环境噪声。此外,为了验证该系统在真实场景中的有效性,我们采用图像源法(ISM)构建了一个高保真的声学仿真环境,也模拟了真实的混响效果。   实验结果表明,所提出的深度ANC系统实现了显着优于传统的FxLMS算法的降噪,特别是对于非平稳噪声,如人群串音。同时,基于PESQ和STOI的评估证实,该系统保持了目标语音的自然度和可懂度。
摘要:Traditional Active Noise Control (ANC) systems are mostly based on FxLMS algorithms, but such algorithms rely on linear assumptions and are often limited in handling broadband non-stationary noise or nonlinear acoustic paths. Not only that, the traditional method is used to eliminating all signals together, and noise reduction often accidentally damages the voice signal and affects normal communication. To tackle these issues, this study proposes a speech preserving deep learning ANC system, which aims to achieve stable noise reduction while effectively retaining speech in a complex acoustic environment.   This study builds an end-to-end control architecture, the core of which adopts a Convolutional Recurrent Network (CRN). The structure uses the long short-term memory (LSTM) network to capture the time-related characteristics of acoustic signals. Combined with complex spectrum mapping (CSM) technology, the nonlinear distortion problem is effectively solved. In order to retain useful voice while removing noise, this study also designs a special voice retention loss function. This design guidance model selectively retains the target voice while suppressing environmental noise by identifying the characteristics of the spectrum structure. In addition, in order to verify whether the system is effective in real scenes, we use the Image Source Method (ISM) to build a high-fidelity acoustic simulation environment, which also simulates the real reverberation effect.   Experimental results demonstrate that the proposed Deep ANC system achieves significantly better noise reduction than the traditional FxLMS algorithm, especially for non-stationary noises like crowd babble. Meanwhile, PESQ and STOI based evaluations confirm that the system preserves both the naturalness and intelligibility of the target speech.


eess.AS音频处理


【1】HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models
标题:HumDial-EIBench:音频语言模型的人类记录的多回合情商基准
链接:https://arxiv.org/abs/2604.11594

作者:Shuiyuan Wang,Zhixian Zhao,Hongfei Yue,Chengyou Wang,Shuai Wang,Hui Bu,Xin Xu,Lei Xie
摘要:评估音频语言模型(ALM)的情绪智力(EI)是至关重要的。然而,现有的基准主要依赖于合成语音,仅限于单圈交互,并严重依赖于开放式评分。本文提出了HumDial-EIBench,一个全面的基准评估ALMs的EI。它使用ICASSP 2026 HumDial挑战赛中的真实记录的人类对话,将情感跟踪和因果推理重新制定为带有对抗性干扰的多项选择题,减轻认知任务的主观评分偏差。它保留了移情反应的生成,并引入了一个声学语义冲突任务,以评估对矛盾的多模态信号的鲁棒性。对八个ALMs的评估表明,大多数模型都难以进行多轮情感跟踪和隐式因果推理。此外,所有模型都表现出解耦的文本和声学移情,以及在跨模态冲突期间严重的文本主导偏见。
摘要:Evaluating the emotional intelligence (EI) of audio language models (ALMs) is critical. However, existing benchmarks mostly rely on synthesized speech, are limited to single-turn interactions, and depend heavily on open-ended scoring. This paper proposes HumDial-EIBench, a comprehensive benchmark for evaluating ALMs' EI. Using real-recorded human dialogues from the ICASSP 2026 HumDial Challenge, it reformulates emotional tracking and causal reasoning into multiple-choice questions with adversarial distractors, mitigating subjective scoring bias for cognitive tasks. It retains the generation of empathetic responses and introduces an acoustic-semantic conflict task to assess robustness against contradictory multimodal signals. Evaluations of eight ALMs reveal that most models struggle with multi-turn emotional tracking and implicit causal reasoning. Furthermore, all models exhibit decoupled textual and acoustic empathy, alongside a severe text-dominance bias during cross-modal conflicts.


【2】Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS
标题:使用语音感知LLMS的说话人归因自动语音识别
链接:https://arxiv.org/abs/2604.11269

作者:Hagai Aronowitz,Zvi Kons,Avihu Dekel,George Saon,Ron Hoory
备注:\c{opyright} 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
摘要:说话者属性自动语音识别(SAA)通过将相对说话者身份标签直接并入转录本(例如,[发言人1]:,[发言人2]:)。在这项工作中,我们扩展了Granite-speech的功能,Granite-speech是一种最先进的语音感知大型语言模型(LLM),最初是为了转录和翻译而训练的。我们证明,它可以有效地适应SAA只有最小的架构变化。我们的核心贡献是引入说话人聚类识别标签(例如,[扬声器1集群42]:),其与SAA联合训练以显著提高准确性。为了解决训练数据的局限性,我们提出了一种数据增强方法,该方法使用人工连接的多说话者对话。我们的方法在多个基准进行评估,并显示出优越的性能相比,传统的管道,依次执行扬声器diarization,然后ASR。
摘要:Speaker-Attributed Automatic Speech Recognition (SAA) enhances traditional ASR systems by incorporating relative speaker identity tags directly into the transcript (e.g., [Speaker 1]:, [Speaker 2]:). In this work, we extend the capabilities of Granite-speech, a state-of-the-art speech-aware Large Language Model (LLM) originally trained for transcription and translation. We demonstrate that it can be effectively adapted for SAA with only minimal architectural changes. Our core contribution is the introduction of speaker cluster identification tags (e.g., [Speaker 1 cluster 42]:) which are jointly trained with SAA to significantly improve accuracy. To address limitations in training data, we propose a data augmentation method that uses artificially concatenated multi-speaker conversations. Our approach is evaluated across multiple benchmarks and shows superior performance compared to conventional pipelines that sequentially perform speaker diarization followed by ASR.


【3】Teaching the Teachers: Boosting unsupervised domain adaptation in speech recognition by ensemble update
标题:教老师:通过集成更新增强语音识别中的无监督域自适应
链接:https://arxiv.org/abs/2604.11256

作者:Rehan Ahmad,Muhammad Umar Farooq,Qihang Feng,Thomas Hain
摘要:语音识别系统经常与未包含在训练中的数据域进行斗争。为了解决这个问题,无监督的领域适应已经探索了集成和多阶段的师生培训方法,减少单词错误率。尽管有所改进,但错误率仍然远远高于有监督的域内训练。这项工作提出了一种更有效的策略,通过同时更新教师模型的集合以及单个学生模型,消除了对顺序模型训练的需要。联合更新改善了学生模型的单词错误率,使逐步增强的教师模型受益。实验进行了三个标记的源数据集,即AMI,WSJ,LS 360,和一个未标记的目标域,即SwitchBoard。实验结果表明,该方法在Switchboard eval00测试集上的WER提高了4.6%,优于多阶段和迭代训练方法。
摘要:Speech recognition systems often struggle with data domains that have not been included in the training. To address this, unsupervised domain adaptation has been explored with ensemble and multi-stage teacher-student training methods reducing the word error rate. Despite improvements, the error rate remains much higher than that achieved with supervised in-domain training. This work proposes a more efficient strategy by simultaneously updating the ensemble of teacher models along with the single student model eliminating the need for sequential models training. The joint update improves the word error rate of the student model, benefiting the progressively enhanced teacher models. Experiments are conducted with three labelled source datasets, namely AMI, WSJ, LS360, and one unlabeled target domain i.e. SwitchBoard. The results show that the proposed method improves the WER by 4.6% on the Switchboard eval00 test set, thus outperforming multi-stage and iterative training methods.


【4】Direction-Preserving MIMO Speech Enhancement Using a Neural Covariance Estimator
标题:使用神经协方差估计器的方向保持的MMO语音增强
链接:https://arxiv.org/abs/2604.11179

作者:Thomas Deppisch
摘要:多通道语音增强作为前端处理被广泛应用于麦克风阵列处理系统中。虽然大多数现有方法产生单个增强信号,但方向保持多输入多输出(MIMO)方法的目标改为提供保持方向属性的增强多信道信号,从而实现下游应用,例如波束成形、双耳渲染和到达方向估计。在这项工作中,我们提出了一个全盲的,方向保持MIMO语音增强方法的基础上神经估计的空间噪声协方差矩阵。轻量级OnlineSpatialNet估计频域噪声协方差的尺度归一化Cholesky因子,该因子与方向保持MIMO Wiener滤波器相结合以增强语音,同时保留目标和残余噪声的空间特性。与以往依赖于预言信息或基于掩码的单输出系统协方差估计的方法相比,所提出的方法直接以低计算复杂度精确多通道协方差估计为目标。实验结果表明,改进的语音增强,协方差估计能力,并在下游任务的性能超过基于掩码的基线,接近Oracle性能与显着更少的参数和计算成本。
摘要:Multichannel speech enhancement is widely used as a front-end in microphone array processing systems. While most existing approaches produce a single enhanced signal, direction-preserving multiple-input multiple-output (MIMO) methods instead aim to provide enhanced multichannel signals that retain directional properties, enabling downstream applications such as beamforming, binaural rendering, and direction-of-arrival estimation. In this work, we propose a fully blind, direction-preserving MIMO speech enhancement method based on neural estimation of the spatial noise covariance matrix. A lightweight OnlineSpatialNet estimates a scale-normalized Cholesky factor of the frequency-domain noise covariance, which is combined with a direction-preserving MIMO Wiener filter to enhance speech while preserving the spatial characteristics of both target and residual noise. In contrast to prior approaches relying on oracle information or mask-based covariance estimation for single-output systems, the proposed method directly targets accurate multichannel covariance estimation with low computational complexity. Experimental results show improved speech enhancement, covariance estimation capability, and performance in downstream tasks over a mask-based baseline, approaching oracle performance with significantly fewer parameters and computational cost.


【5】Speech-preserving active noise control: a deep learning approach in reverberant environments
标题:语音保留主动噪音控制:回响环境中的深度学习方法
链接:https://arxiv.org/abs/2604.10979

作者:Shuning Dai
备注:89 pages, 17 figures, master's dissertation
摘要:传统的有源噪声控制(ANC)系统大多基于FxLMS算法,但此类算法依赖于线性假设,并且通常在处理宽带非平稳噪声或非线性声学路径方面受到限制。不仅如此,传统的方法是将所有信号一起消除,降噪往往会意外地破坏语音信号,影响正常通信。为了解决这些问题,本研究提出了一种语音保留深度学习ANC系统,旨在实现稳定的降噪,同时在复杂的声学环境中有效地保留语音。   本研究建立了一个端到端的控制架构,其核心采用卷积递归网络(CRN)。该结构使用长短期记忆(LSTM)网络来捕获声信号的时间相关特征。结合复谱映射技术,有效地解决了非线性失真问题。为了在去除噪声的同时保留有用的语音,本研究还设计了一个特殊的语音保留损失函数。该设计指导模型通过识别频谱结构的特性来选择性地保留目标语音,同时抑制环境噪声。此外,为了验证该系统在真实场景中的有效性,我们采用图像源法(ISM)构建了一个高保真的声学仿真环境,也模拟了真实的混响效果。   实验结果表明,所提出的深度ANC系统实现了显着优于传统的FxLMS算法的降噪,特别是对于非平稳噪声,如人群串音。同时,基于PESQ和STOI的评估证实,该系统保持了目标语音的自然度和可懂度。
摘要:Traditional Active Noise Control (ANC) systems are mostly based on FxLMS algorithms, but such algorithms rely on linear assumptions and are often limited in handling broadband non-stationary noise or nonlinear acoustic paths. Not only that, the traditional method is used to eliminating all signals together, and noise reduction often accidentally damages the voice signal and affects normal communication. To tackle these issues, this study proposes a speech preserving deep learning ANC system, which aims to achieve stable noise reduction while effectively retaining speech in a complex acoustic environment.   This study builds an end-to-end control architecture, the core of which adopts a Convolutional Recurrent Network (CRN). The structure uses the long short-term memory (LSTM) network to capture the time-related characteristics of acoustic signals. Combined with complex spectrum mapping (CSM) technology, the nonlinear distortion problem is effectively solved. In order to retain useful voice while removing noise, this study also designs a special voice retention loss function. This design guidance model selectively retains the target voice while suppressing environmental noise by identifying the characteristics of the spectrum structure. In addition, in order to verify whether the system is effective in real scenes, we use the Image Source Method (ISM) to build a high-fidelity acoustic simulation environment, which also simulates the real reverberation effect.   Experimental results demonstrate that the proposed Deep ANC system achieves significantly better noise reduction than the traditional FxLMS algorithm, especially for non-stationary noises like crowd babble. Meanwhile, PESQ and STOI based evaluations confirm that the system preserves both the naturalness and intelligibility of the target speech.


【6】Toward using Speech to Sense Student Emotion in Remote Learning Environments
标题:在远程学习环境中使用语音感知学生情绪
链接:https://arxiv.org/abs/2604.09881

作者:Sargam Vyas,Bogdan Vlasenko,André Mayoraz,Egon Werlen,Per Bergamin,Mathew Magimai. -Doss
摘要:随着多模态通信技术的进步,诸如远程大学的远程学习环境正在增加。远程学习通常是异步进行的。因此,与面对面的课堂教学不同,这缺乏足够的情感线索,使学习成为一种愉快的体验。出于对情感预测的语言学语音处理社区的进步,在本文中,我们探讨了使用语音感知学生的情绪,建立在基于语音的自我控制任务,以帮助有效的远程学习。更确切地说,我们调查:(一)是否通过自我控制任务获得的语音沿效价,唤醒和优势维度表现出可感知的变化?以及(b)这些维度情感变化是否可以被自动预测?我们通过开发一个包含自发独白言语的数据集来解决这两个研究问题,这些独白言语是作为对自我控制任务的开放反应而获得的,并通过对该数据集进行主观听众评估和自动维度情绪预测研究。我们的研究表明,基于语音的自我控制任务可以作为远程学习环境中感知学生情绪的一种手段。这为在远程学习循环中无缝集成双语言语音处理技术开辟了潜在的场所,以便通过教学设计和反馈生成来增强学习体验。
摘要:With advancements in multimodal communication technologies, remote learning environments such as, distance universities are increasing. Remote learning typically happens asynchronously. As a consequence, unlike face-to-face in-person classroom teaching, this lacks availability of sufficient emotional cues for making learning a pleasant experience. Motivated by advances made in the paralinguistic speech processing community on emotion prediction, in this paper we explore use of speech for sensing students' emotions by building upon speech-based self-control tasks developed to aid effective remote learning. More precisely, we investigate: (a) whether speech acquired through self-control tasks exhibit perceptible variation along valence, arousal, and dominance dimensions? and (b) whether those dimensional emotion variations can be automatically predicted? We address these two research questions by developing a dataset containing spontaneous monologue speech acquired as open responses to self-control tasks and by carrying out subjective listener evaluations and automatic dimensional emotion prediction studies on that dataset. Our investigations indicate that speech-based self-control tasks can be a means to sense student emotion in remote learning environment. This opens potential venues to seamlessly integrate paralinguistic speech processing technologies in the remote learning loop for enhancing learning experiences through instructional design and feedback generation.


【7】Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
标题:音频Flamingo Next:下一代语音、声音和音乐开放音频语言模型
链接:https://arxiv.org/abs/2604.10905

作者:Sreyan Ghosh,Arushi Goel,Kaousheik Jayakumar,Lasha Koroshinadze,Nishit Anand,Zhifeng Kong,Siddharth Gururani,Sang-gil Lee,Jaehyeon Kim,Aya Aljafari,Chao-Han Huck Yang,Sungwon Kim,Ramani Duraiswami,Dinesh Manocha,Mohammad Shoeybi,Bryan Catanzaro,Ming-Yu Liu,Wei Ping
备注:Project website: https://afnext-umd-nvidia.github.io/
摘要:Audio Flamingo Next(AF-Next)是Audio Flamingo系列中功能最强大的下一代大型音频语言模型,旨在促进对语音、环境声音和音乐的理解和推理。与Audio Flamingo 3相比,AF-Next引入了:(i)更强大的基础音频语言模型,可显著提高各种音频理解任务的准确性;(ii)可扩展的策略,用于构建超出现有学术基准的大规模音频理解和推理数据;(iii)支持长达30分钟的长而复杂的音频输入;以及(iv)时间音频思想链,一种新的推理范式,其明确地将中间推理步骤基于长音频中的时间戳,从而实现细粒度的时间对齐和改进的可解释性。为了实现这些功能,我们首先对Audio Flamingo 3进行系统分析,以确定音频理解和推理方面的关键差距。然后,我们策划和扩展总计超过100万小时的新的大规模数据集,以解决这些限制,并扩展现有的AudioSkills-XL,LongAudio-XL,AF-Think和AF-Chat数据集。AF-Next使用基于训练的策略进行训练,包括训练前、训练中和训练后阶段。在20个音频理解和推理基准测试中进行的广泛实验,包括具有挑战性的长音频任务,表明AF-Next的性能远远优于类似大小的开放模型,并且与更大的开放重量和封闭模型保持高度竞争力,有时甚至超过它们。除了基准性能之外,AF-Next还表现出强大的现实效用,并可以很好地转移到看不见的任务,突出了其鲁棒性和泛化能力。除了所有的数据、代码和方法,我们还开源了AF-Next的3个变体,包括AF-Next-Instruct、AF-Next-Think和AF-Next-Captioner。
摘要:We present Audio Flamingo Next (AF-Next), the next-generation and most capable large audio-language model in the Audio Flamingo series, designed to advance understanding and reasoning over speech, environmental sounds and music. Compared to Audio Flamingo 3, AF-Next introduces: (i) a stronger foundational audio-language model that significantly improves accuracy across diverse audio understanding tasks; (ii) scalable strategies for constructing large-scale audio understanding and reasoning data beyond existing academic benchmarks; (iii) support for long and complex audio inputs up to 30 minutes; and (iv) Temporal Audio Chain-of-Thought, a new reasoning paradigm that explicitly grounds intermediate reasoning steps to timestamps in long audio, enabling fine-grained temporal alignment and improved interpretability. To enable these capabilities, we first conduct a systematic analysis of Audio Flamingo 3 to identify key gaps in audio understanding and reasoning. We then curate and scale new large-scale datasets totaling over 1 million hours to address these limitations and expand the existing AudioSkills-XL, LongAudio-XL, AF-Think and AF-Chat datasets. AF-Next is trained using a curriculum-based strategy spanning pre-training, mid-training and post-training stages. Extensive experiments across 20 audio understanding and reasoning benchmarks, including challenging long-audio tasks, show that AF-Next outperforms similarly sized open models by large margins and remains highly competitive with and sometimes surpasses, much larger open-weight and closed models. Beyond benchmark performance, AF-Next exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability. In addition to all data, code and methods, we open-source 3 variants of AF-Next, including AF-Next-Instruct, AF-Next-Think and AF-Next-Captioner.


【8】Multimodal Dataset Normalization and Perceptual Validation for Music-Taste Correspondences
标题:音乐品味对应的多模式数据集规范化和感知验证
链接:https://arxiv.org/abs/2604.10632

作者:Matteo Spanio,Valentina Frezzato,Antonio Rodà
备注:Submitted to SMC2026
摘要:收集大型的、对齐的跨模态数据集用于音乐风味研究是困难的,因为感知实验是昂贵的,而且设计得很小。我们通过两个互补的实验来解决这个瓶颈。第一个测试是否音频风味的相关性,功能的重要性排名,和潜在的因素结构转移从实验的音轨集合(257轨道与人类注释)到一个大型的FM衍生语料库(49,300段合成标签)。第二个验证计算风味目标-来自食品化学通过一个可重复的管道-对人类的感知在一个在线的听众研究(49~参与者,20~轨道)。两个实验的结果趋于一致:定量迁移分析证实跨模态结构在不同的监督机制中得到了保留,感知评估显示计算目标和听众评级之间存在显著的一致性(排列p<0.0001$,Mantel $r=0.45$,Procrustes $m^2=0.51$)。总之,这些研究结果支持的结论是,声音调味效果存在于合成FMA注释。我们发布数据集和配套代码,以支持可重复的跨模态AI研究。
摘要:Collecting large, aligned cross-modal datasets for music-flavor research is difficult because perceptual experiments are costly and small by design. We address this bottleneck through two complementary experiments. The first tests whether audio-flavor correlations, feature-importance rankings, and latent-factor structure transfer from an experimental soundtracks collection (257~tracks with human annotations) to a large FMA-derived corpus ($\sim$49,300 segments with synthetic labels). The second validates computational flavor targets -- derived from food chemistry via a reproducible pipeline -- against human perception in an online listener study (49~participants, 20~tracks). Results from both experiments converge: the quantitative transfer analysis confirms that cross-modal structure is preserved across supervision regimes, and the perceptual evaluation shows significant alignment between computational targets and listener ratings (permutation $p<0.0001$, Mantel $r=0.45$, Procrustes $m^2=0.51$). Together, these findings support the conclusion that sonic seasoning effects are present in synthetic FMA annotations. We release datasets and companion code to support reproducible cross-modal AI research.


【9】ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models
标题:ASPIRin:用于全复式语音语言模型中交互优化强化学习的动作空间投影
链接:https://arxiv.org/abs/2604.10065

作者:Chi-Yuan Hsiao,Ke-Han Lu,Yu-Kuan Fu,Guan-Ting Lin,Hsiao-Tsung Hung,Hung-yi Lee
摘要:端到端全双工语音语言模型(SLM)需要精确的话轮转换以实现自然交互。然而,通过标准的原始令牌强化学习(RL)优化时间动态会降低语义质量,导致严重的生成崩溃和重复。我们提出了ASPIRin,这是一个交互优化的RL框架,它明确地解释了什么时候说什么。使用动作空间投影,ASPIRin将文本词汇映射到粗粒度的二进制状态(主动语音与非主动沉默)。通过应用具有基于规则的奖励的组相对策略优化(GRPO),它平衡了用户中断和响应延迟。实证评估表明,ASPIRin优化了话轮转换、反向引导和暂停处理的交互性。至关重要的是,将计时与标记选择隔离开来,可以保持语义一致性,并将重复的n元语法的部分与标准GRPO相比减少了50%以上,有效地消除了退化重复。
摘要:End-to-end full-duplex Speech Language Models (SLMs) require precise turn-taking for natural interaction. However, optimizing temporal dynamics via standard raw-token reinforcement learning (RL) degrades semantic quality, causing severe generative collapse and repetition. We propose ASPIRin, an interactivity-optimized RL framework that explicitly decouples when to speak from what to say. Using Action Space Projection, ASPIRin maps the text vocabulary into a coarse-grained binary state (active speech vs. inactive silence). By applying Group Relative Policy Optimization (GRPO) with rule-based rewards, it balances user interruption and response latency. Empirical evaluations show ASPIRin optimizes interactivity across turn-taking, backchanneling, and pause handling. Crucially, isolating timing from token selection preserves semantic coherence and reduces the portion of duplicate n-grams by over 50% compared to standard GRPO, effectively eliminating degenerative repetition.


【10】Regularized Entropy Information Adaptation with Temporal-Awareness Networks for Simultaneous Speech Translation
标题:基于时间感知网络的规则化信息自适应语音同步翻译
链接:https://arxiv.org/abs/2604.09916

作者:Joseph Liu,Nameer Hirschkind,Xiao Yu,Mahesh Kumar Nandwana
备注:Under review at Interspeech 2026
摘要:同时语音翻译(SimulST)需要平衡高翻译质量和低延迟。最近的工作介绍了REINA,这是一种基于估计读取更多音频的信息增益来训练读/写策略的方法。然而,我们发现,基于信息的政策往往缺乏时间背景,导致政策偏向于阅读大部分的音频开始写之前。我们使用两种不同的策略来改进REINA:监督对齐网络(REINA-SAN)和时间步长增强网络(REINA-TAN)。我们的研究结果表明,虽然这两种方法显着优于基线和解决稳定性问题,REINA-TAN提供了一个稍微优越的帕累托前沿流效率,而REINA-SAN提供了更强大的鲁棒性对'读循环'。应用于Whisper,这两种方法都提高了流媒体效率的帕累托边界,通过标准化流媒体效率(NoSE)分数测量,比现有的竞争基准高出7.1%。
摘要:Simultaneous Speech Translation (SimulST) requires balancing high translation quality with low latency. Recent work introduced REINA, a method that trains a Read/Write policy based on estimating the information gain of reading more audio. However, we find that information-based policies often lack temporal context, leading the policy to bias itself toward reading most of the audio before starting to write. We improve REINA using two distinct strategies: a supervised alignment network (REINA-SAN) and a timestep-augmented network (REINA-TAN). Our results demonstrate that while both methods significantly outperform the baseline and resolve stability issues, REINA-TAN provides a slightly superior Pareto frontier for streaming efficiency, whereas REINA-SAN offers more robustness against 'read loops'. Applied to Whisper, both methods improve the pareto frontier of streaming efficiency as measured by Normalized Streaming Efficiency (NoSE) scores up to 7.1% over existing competitive baselines.


机器翻译由腾讯交互翻译提供,仅供参考