微信公众号:arXiv_Daily
cs.SD语音
【1】The Sonar Moment: Benchmarking Audio-Language Models in Audio Geo-Localization
标题:声纳时刻:音频地理本地化中音频语言模型基准测试
链接:https://arxiv.org/abs/2601.03227
摘要:地理定位旨在推断给定信号的地理来源。在计算机视觉中,地理定位已经成为组合推理的一个苛刻的基准,并且与公共安全有关。相比之下,音频地理定位的进展受到缺乏高质量音频位置对的限制。为了解决这一差距,我们推出了AGL 1K,这是首个音频语言模型(ALM)的音频地理定位基准,覆盖72个国家和地区。为了从众包平台中提取可靠的本地化样本,我们提出了音频本地化指标,该指标量化了每个录音的信息量,产生了1,444个策划的音频片段。对16个ALM的评估表明,ALM已经具有音频地理定位能力。我们发现,闭源模型大大优于开源模型,语言线索往往占主导地位,作为一个脚手架的预测。我们进一步分析了ALMs的推理痕迹,区域偏见,错误的原因,以及本地化度量的可解释性。总的来说,AGL1K为音频地理定位建立了基准,并可能使ALM具有更好的地理空间推理能力。
摘要:Geo-localization aims to infer the geographic origin of a given signal. In computer vision, geo-localization has served as a demanding benchmark for compositional reasoning and is relevant to public safety. In contrast, progress on audio geo-localization has been constrained by the lack of high-quality audio-location pairs. To address this gap, we introduce AGL1K, the first audio geo-localization benchmark for audio language models (ALMs), spanning 72 countries and territories. To extract reliably localizable samples from a crowd-sourced platform, we propose the Audio Localizability metric that quantifies the informativeness of each recording, yielding 1,444 curated audio clips. Evaluations on 16 ALMs show that ALMs have emerged with audio geo-localization capability. We find that closed-source models substantially outperform open-source models, and that linguistic clues often dominate as a scaffold for prediction. We further analyze ALMs' reasoning traces, regional bias, error causes, and the interpretability of the localizability metric. Overall, AGL1K establishes a benchmark for audio geo-localization and may advance ALMs with better geospatial reasoning capability.
【2】Segment-Aware Conditioning for Training-Free Intra-Utterance Emotion and Duration Control in Text-to-Speech
标题:文本转语音中免训练言语内情绪和持续时间控制的分段感知条件反射
链接:https://arxiv.org/abs/2601.03170
备注:24 pages, 8 figures, 7 tables, 3 lists
摘要:虽然可控的文本到语音(TTS)已经取得了显着的进展,大多数现有的方法仍然局限于话语间的控制,使细粒度的话语内的表达具有挑战性,由于它们依赖于非公开的数据集或复杂的多阶段训练。在本文中,我们提出了一个无训练的可控框架预训练zero-shot TTS,使话语内的情感和持续时间的表达。具体来说,我们提出了一个段感知的情绪调节策略,结合因果掩蔽与单调流对齐过滤,以隔离情绪调节和时间表掩码转换,使顺利的话语内的情绪变化,同时保持全局语义一致性。在此基础上,我们进一步提出了一个片段感知的持续时间转向策略,结合本地持续时间嵌入转向与全局EOS logit调制,允许本地持续时间调整,同时确保全局一致的终止。为了消除段级手动提示工程的需要,我们构建了一个30,000个样本的多情感和持续时间注释的文本数据集,以实现基于LLM的自动提示构建。大量的实验表明,我们的无训练方法不仅实现了最先进的多情感和持续时间控制的话语内一致性,但也保持了基本的TTS模型的基线水平的语音质量。音频样本可在https://aclanonymous111.github.io/TED-TTS-DemoPage/上获得。
摘要:While controllable Text-to-Speech (TTS) has achieved notable progress, most existing methods remain limited to inter-utterance-level control, making fine-grained intra-utterance expression challenging due to their reliance on non-public datasets or complex multi-stage training. In this paper, we propose a training-free controllable framework for pretrained zero-shot TTS to enable intra-utterance emotion and duration expression. Specifically, we propose a segment-aware emotion conditioning strategy that combines causal masking with monotonic stream alignment filtering to isolate emotion conditioning and schedule mask transitions, enabling smooth intra-utterance emotion shifts while preserving global semantic coherence. Based on this, we further propose a segment-aware duration steering strategy to combine local duration embedding steering with global EOS logit modulation, allowing local duration adjustment while ensuring globally consistent termination. To eliminate the need for segment-level manual prompt engineering, we construct a 30,000-sample multi-emotion and duration-annotated text dataset to enable LLM-based automatic prompt construction. Extensive experiments demonstrate that our training-free method not only achieves state-of-the-art intra-utterance consistency in multi-emotion and duration control, but also maintains baseline-level speech quality of the underlying TTS model. Audio samples are available at https://aclanonymous111.github.io/TED-TTS-DemoPage/.
【3】Interpretable All-Type Audio Deepfake Detection with Audio LLMs via Frequency-Time Reinforcement Learning
标题:通过频率-时间强化学习,使用音频LLM进行可解释的全类型音频Deepfake检测
链接:https://arxiv.org/abs/2601.02983
摘要:音频大语言模型(ALLM)的最新进展使得高质量的合成音频被广泛使用,增加了语音、环境声音、歌声和音乐中恶意音频deepfake的风险。因此,现实世界的音频深度伪造检测(ADD)需要所有类型的检测器,这些检测器可以在异构音频中进行泛化,并提供可解释的决策。鉴于ALLM强大的多任务泛化能力,我们首先研究了它们在监督微调(SFT)和强化微调(RFT)下对所有类型ADD的性能。然而,仅使用二进制真/假标签的SFT倾向于将模型简化为黑盒分类器,从而牺牲了可解释性。与此同时,在稀疏监督下的香草RFT容易受到奖励黑客的攻击,并可能产生幻觉,没有根据的理由。为了解决这个问题,我们提出了一个自动注释和抛光管道,构建频率-时间结构化的思想链(CoT)的理由,产生约340 K冷启动演示。在CoT数据的基础上,我们提出了频率时间组相对策略优化(FT-GRPO),这是一个两阶段的训练范式,它用SFT冷启动ALLM,然后在基于规则的频率-时间约束下应用GRPO。实验表明,FT-GRPO实现国家的最先进的性能对所有类型的ADD,同时产生可解释的,FT接地的理由。数据和代码可在线获取。
摘要:Recent advances in audio large language models (ALLMs) have made high-quality synthetic audio widely accessible, increasing the risk of malicious audio deepfakes across speech, environmental sounds, singing voice, and music. Real-world audio deepfake detection (ADD) therefore requires all-type detectors that generalize across heterogeneous audio and provide interpretable decisions. Given the strong multi-task generalization ability of ALLMs, we first investigate their performance on all-type ADD under both supervised fine-tuning (SFT) and reinforcement fine-tuning (RFT). However, SFT using only binary real/fake labels tends to reduce the model to a black-box classifier, sacrificing interpretability. Meanwhile, vanilla RFT under sparse supervision is prone to reward hacking and can produce hallucinated, ungrounded rationales. To address this, we propose an automatic annotation and polishing pipeline that constructs Frequency-Time structured chain-of-thought (CoT) rationales, producing ~340K cold-start demonstrations. Building on CoT data, we propose Frequency Time-Group Relative Policy Optimization (FT-GRPO), a two-stage training paradigm that cold-starts ALLMs with SFT and then applies GRPO under rule-based frequency-time constraints. Experiments demonstrate that FT-GRPO achieves state-of-the-art performance on all-type ADD while producing interpretable, FT-grounded rationales. The data and code are available online.
【4】MoE Adapter for Large Audio Language Models: Sparsity, Disentanglement, and Gradient-Conflict-Free
标题:适用于大型音频语言模型的MoE适配器:稀疏性、解纠缠和无冲突
链接:https://arxiv.org/abs/2601.02967
备注:13 pages, 5 figures
摘要:将大语言模型的输入模态扩展到音频域是实现多模态感知的关键。然而,众所周知,声学信息本质上是\textit {异构}的,纠缠的属性,如语音,音乐和环境背景。现有的研究仅限于一个密集的,参数共享的适配器来模拟这些不同的模式,这会导致优化过程中的\textit {梯度冲突},因为不同属性所需的参数更新相互矛盾。为了解决这个问题,我们引入了\textit {\textbf {MoE-Adapter}},一个稀疏的混合专家~(MoE)架构,旨在解耦声学信息。具体而言,它采用了一种动态门控机制,将音频令牌路由到捕获互补特征子空间的专业专家,同时保留全局上下文的共享专家,从而减轻梯度冲突并实现细粒度特征学习。综合实验表明,MoE-Adapter在音频语义和语言任务上都取得了优异的性能,在计算成本相当的情况下,始终优于密集的线性基线。此外,我们将发布相关的代码和模型,以方便未来的研究。
摘要:Extending the input modality of Large Language Models~(LLMs) to the audio domain is essential for achieving comprehensive multimodal perception. However, it is well-known that acoustic information is intrinsically \textit{heterogeneous}, entangling attributes such as speech, music, and environmental context. Existing research is limited to a dense, parameter-shared adapter to model these diverse patterns, which induces \textit{gradient conflict} during optimization, as parameter updates required for distinct attributes contradict each other. To address this limitation, we introduce the \textit{\textbf{MoE-Adapter}}, a sparse Mixture-of-Experts~(MoE) architecture designed to decouple acoustic information. Specifically, it employs a dynamic gating mechanism that routes audio tokens to specialized experts capturing complementary feature subspaces while retaining shared experts for global context, thereby mitigating gradient conflicts and enabling fine-grained feature learning. Comprehensive experiments show that the MoE-Adapter achieves superior performance on both audio semantic and paralinguistic tasks, consistently outperforming dense linear baselines with comparable computational costs. Furthermore, we will release the related code and models to facilitate future research.
【5】The World is Not Mono: Enabling Spatial Understanding in Large Audio-Language Models
标题:世界不是单一的:在大型音频语言模型中实现空间理解
链接:https://arxiv.org/abs/2601.02954
摘要:现有的大型音频语言模型将世界视为“单声道”-忽略通用声学场景分析所需的关键空间维度(“何处”)的单一音频流。为了弥合这一差距,我们首先介绍了一个层次框架的听觉场景分析(ASA)。在这个框架的指导下,我们引入了一个系统,使Qwen 2-Audio这样的模型能够理解和推理复杂的声学世界。我们的框架通过三个核心贡献实现了这一点:首先,我们构建了一个大规模的合成双耳音频数据集,以提供丰富的空间线索。其次,我们设计了一个混合特征投影仪,它利用并行的语义和空间编码器来提取解耦的表示。这些不同的流通过密集融合机制集成,确保模型接收声学场景的整体视图。最后,我们采用渐进式训练课程,从监督微调(SFT)推进到通过组相对策略优化(GRPO)的强化学习,以显式地发展模型的推理能力。在我们的综合基准上,该模型表现出较强的空间理解能力。通过实现这种空间感知,我们的工作为利用大型模型的强大推理能力进行整体声学场景分析提供了一条清晰的途径,从“单”语义识别推进到空间智能。
摘要:Existing large audio-language models perceive the world as "mono" -- a single stream of audio that ignores the critical spatial dimension ("where") required for universal acoustic scene analysis. To bridge this gap, we first introduce a hierarchical framework for Auditory Scene Analysis (ASA). Guided by this framework, we introduce a system that enables models like Qwen2-Audio to understand and reason about the complex acoustic world. Our framework achieves this through three core contributions: First, we build a large-scale, synthesized binaural audio dataset to provide the rich spatial cues. Second, we design a hybrid feature projector, which leverages parallel semantic and spatial encoders to extract decoupled representations. These distinct streams are integrated via a dense fusion mechanism, ensuring the model receives a holistic view of the acoustic scene. Finally, we employ a progressive training curriculum, advancing from supervised fine-tuning (SFT) to reinforcement learning via Group Relative Policy Optimization (GRPO), to explicitly evolve the model's capabilities towards reasoning. On our comprehensive benchmark, the model demonstrates comparatively strong capability for spatial understanding. By enabling this spatial perception, our work provides a clear pathway for leveraging the powerful reasoning abilities of large models towards holistic acoustic scene analysis, advancing from "mono" semantic recognition to spatial intelligence.
【6】Vulnerabilities of Audio-Based Biometric Authentication Systems Against Deepfake Speech Synthesis
标题:基于音频的生物识别认证系统针对Deepfake语音合成的漏洞
链接:https://arxiv.org/abs/2601.02914
摘要:随着音频deepfake从研究工件过渡到广泛使用的商业工具,强大的生物特征认证在高风险行业面临着紧迫的安全威胁。本文对基于大规模语音合成数据集的最新说话人认证系统进行了系统的实证评估,揭示了两个主要的安全漏洞:1)在非常小的样本上训练的现代语音克隆模型可以很容易地绕过商业说话人认证系统;以及2)反欺骗检测器努力在不同的音频合成方法上进行推广,导致域内性能和真实世界鲁棒性之间的显著差距。这些发现要求重新考虑安全措施,并强调需要架构创新,自适应防御以及向多因素身份验证过渡。
摘要:As audio deepfakes transition from research artifacts to widely available commercial tools, robust biometric authentication faces pressing security threats in high-stakes industries. This paper presents a systematic empirical evaluation of state-of-the-art speaker authentication systems based on a large-scale speech synthesis dataset, revealing two major security vulnerabilities: 1) modern voice cloning models trained on very small samples can easily bypass commercial speaker verification systems; and 2) anti-spoofing detectors struggle to generalize across different methods of audio synthesis, leading to a significant gap between in-domain performance and real-world robustness. These findings call for a reconsideration of security measures and stress the need for architectural innovations, adaptive defenses, and the transition towards multi-factor authentication.
【7】SPO-CLAPScore: Enhancing CLAP-based alignment prediction system with Standardize Preference Optimization, for the first XACLE Challenge
标题:SPO-CLAPScore:通过标准化偏好优化增强基于CLAP的对齐预测系统,参加首届XACLE挑战赛
链接:https://arxiv.org/abs/2601.02900
备注:https://github.com/ttakano398/SPO-CLAPScore
摘要:第一个XACLE挑战(x到音频对齐挑战)解决了与人类感知的音频文本语义对齐相关的自动评估指标的关键需求。在本文中,我们描述了提交给XACLE挑战赛的“Takano_UTokyo_03”系统。我们的方法利用了基于CLAPScore的架构,并集成了一种名为标准化偏好优化(SPO)的新型训练方法。SPO对每个听众提供的原始对齐分数进行分析,使模型能够学习相对偏好并减轻个人评分偏差的影响。此外,我们采用听众筛选,以排除听众与不一致的评级。实验结果表明,SPO和听者筛选都有效地提高了与人类判断的相关性。我们的系统在挑战中获得了第6名,斯皮尔曼等级相关系数(SRCC)为0.6142,显示出与顶级系统的边缘差距内的竞争力。该代码可在https://github.com/ttakano398/SPO-CLAPScore上获得。
摘要:The first XACLE Challenge (x-to-audio alignment challenge) addresses the critical need for automatic evaluation metrics that correlate with human perception of audio-text semantic alignment. In this paper, we describe the "Takano_UTokyo_03" system submitted to XACLE Challenge. Our approach leverages a CLAPScore-based architecture integrated with a novel training method called Standardized Preference Optimization (SPO). SPO standardizes the raw alignment scores provided by each listener, enabling the model to learn relative preferences and mitigate the impact of individual scoring biases. Additionally, we employ listener screening to exclude listeners with inconsistent ratings. Experimental evaluations demonstrate that both SPO and listener screening effectively improve the correlation with human judgment. Our system achieved 6th place in the challenge with a Spearman's rank correlation coefficient (SRCC) of 0.6142, demonstrating competitive performance within a marginal gap from the top-ranked systems. The code is available at https://github.com/ttakano398/SPO-CLAPScore.
【8】UniSRCodec: Unified and Low-Bitrate Single Codebook Codec with Sub-Band Reconstruction
标题:UniSRCodec:统一的低比特率单码本子带重构编解码器
链接:https://arxiv.org/abs/2601.02776
备注:6 pages, 2 figures, and 3 tables
摘要:神经音频编解码器(NAC)可以通过执行紧凑的压缩和重构来减少传输开销,这也旨在弥合连续和离散信号之间的差距。现有的NAC可以分为两类:多码本和单码本编解码器。多码本编解码器面临诸如结构复杂性和难以适应下游任务的挑战,而单码本编解码器虽然结构上更简单,但遭受低保真度、统一音频的无效建模以及不能支持高频音频的建模。我们提出了UniSRCodec,一个单码本编解码器能够支持高采样率,低带宽,高保真度,和统一的。分析了基于波形的压缩效率低的问题,提出了基于Mel-频谱图的时频压缩方法,并与声码器相结合,恢复出原始音频的相位信息。此外,我们提出了一个子带重建技术,以实现高品质的压缩在低和高频段。主观和客观实验结果表明,UniSRCodec在令牌速率仅为40的情况下,在跨域单码本编解码器中达到了最先进的(SOTA)性能,其重建质量与某些多码本方法相当。我们的演示页面可以在https://wxzyd123.github.io/unisrcodec上找到。
摘要:Neural Audio Codecs (NACs) can reduce transmission overhead by performing compact compression and reconstruction, which also aim to bridge the gap between continuous and discrete signals. Existing NACs can be divided into two categories: multi-codebook and single-codebook codecs. Multi-codebook codecs face challenges such as structural complexity and difficulty in adapting to downstream tasks, while single-codebook codecs, though structurally simpler, suffer from low-fidelity, ineffective modeling of unified audio, and an inability to support modeling of high-frequency audio. We propose the UniSRCodec, a single-codebook codec capable of supporting high sampling rate, low-bandwidth, high fidelity, and unified. We analyze the inefficiency of waveform-based compression and introduce the time and frequency compression method using the Mel-spectrogram, and cooperate with a Vocoder to recover the phase information of the original audio. Moreover, we propose a sub-band reconstruction technique to achieve high-quality compression across both low and high frequency bands. Subjective and objective experimental results demonstrate that UniSRCodec achieves state-of-the-art (SOTA) performance among cross-domain single-codebook codecs with only a token rate of 40, and its reconstruction quality is comparable to that of certain multi-codebook methods. Our demo page is available at https://wxzyd123.github.io/unisrcodec.
【9】Omni2Sound: Towards Unified Video-Text-to-Audio Generation
标题:Omni 2Sound:迈向统一的视频文本到音频生成
链接:https://arxiv.org/abs/2601.02731
摘要:训练集成视频到音频(V2 A)、文本到音频(T2 A)和联合视频-文本到音频(VT 2A)生成的统一模型提供了显著的应用灵活性,但面临两个未探索的基本挑战:(1)缺乏具有紧密A-V-T对齐的高质量音频字幕,导致多模态条件之间的严重语义冲突,以及(2)跨任务和任务内竞争,表现为不利的V2 A-T2 A性能权衡和VT 2A任务中的模态偏差。首先,为了解决数据稀缺问题,我们引入了SoundAtlas,这是一个大规模的数据集(47万对),在质量上明显优于现有的基准甚至人类专家。它由一个新颖的代理管道提供支持,它集成了视觉到语言压缩,以减轻MLLM的视觉偏见,初级-高级代理汉诺,以降低5倍的成本,以及严格的事后过滤,以确保保真度。因此,SoundAtlas提供了语义丰富和时间详细的字幕,具有紧密的V-A-T对齐。其次,我们提出了Omni 2Sound,一个支持灵活输入方式的统一VT 2A扩散模型。为了解决固有的跨任务和任务内竞争,我们设计了一个三阶段的多任务渐进式训练计划,将跨任务竞争转化为联合优化,并减轻VT 2A任务中的模态偏差,同时保持视听对齐和屏幕外音频生成的忠实性。最后,我们构建了VGGSound-Omni,这是一个用于统一评估的综合基准,包括具有挑战性的屏幕外曲目。通过标准的DiT主干,Omni 2Sound在单个模型中实现了所有三个任务的统一SOTA性能,在具有异构输入条件的基准测试中表现出强大的泛化能力。该项目的页面位于https://swapforward.github.io/Omni2Sound。
摘要:Training a unified model integrating video-to-audio (V2A), text-to-audio (T2A), and joint video-text-to-audio (VT2A) generation offers significant application flexibility, yet faces two unexplored foundational challenges: (1) the scarcity of high-quality audio captions with tight A-V-T alignment, leading to severe semantic conflict between multimodal conditions, and (2) cross-task and intra-task competition, manifesting as an adverse V2A-T2A performance trade-off and modality bias in the VT2A task. First, to address data scarcity, we introduce SoundAtlas, a large-scale dataset (470k pairs) that significantly outperforms existing benchmarks and even human experts in quality. Powered by a novel agentic pipeline, it integrates Vision-to-Language Compression to mitigate visual bias of MLLMs, a Junior-Senior Agent Handoff for a 5 times cost reduction, and rigorous Post-hoc Filtering to ensure fidelity. Consequently, SoundAtlas delivers semantically rich and temporally detailed captions with tight V-A-T alignment. Second, we propose Omni2Sound, a unified VT2A diffusion model supporting flexible input modalities. To resolve the inherent cross-task and intra-task competition, we design a three-stage multi-task progressive training schedule that converts cross-task competition into joint optimization and mitigates modality bias in the VT2A task, maintaining both audio-visual alignment and off-screen audio generation faithfulness. Finally, we construct VGGSound-Omni, a comprehensive benchmark for unified evaluation, including challenging off-screen tracks. With a standard DiT backbone, Omni2Sound achieves unified SOTA performance across all three tasks within a single model, demonstrating strong generalization across benchmarks with heterogeneous input conditions. The project page is at https://swapforward.github.io/Omni2Sound.
【10】Multi-channel multi-speaker transformer for speech recognition
标题:用于语音识别的多通道多扬声器Transformer
链接:https://arxiv.org/abs/2601.02688
备注:Proc. INTERSPEECH 2023, 5 pages
摘要:随着远程会议和车载语音助手的发展,远场多说话人语音识别已经成为一个研究热点。最近,已经提出了多通道Transformer(MCT),其展示了Transformer对远场声学环境进行建模的能力。然而,由于扬声器之间的干扰,MCT不能从混合输入音频中编码每个扬声器的高维声学特征。在此基础上,本文提出了一种用于远场多说话人语音合成的多通道多说话人Transformer(M2 Former)。在SMS-WSJ基准测试上的实验表明,M2 Former在相对字错误率降低方面分别优于神经波束形成器、MCT、具有变换平均级联和基于多通道深度聚类的端到端系统的双路径RNN 9.2%、14.3%、24.9%和52.2%。
摘要:With the development of teleconferencing and in-vehicle voice assistants, far-field multi-speaker speech recognition has become a hot research topic. Recently, a multi-channel transformer (MCT) has been proposed, which demonstrates the ability of the transformer to model far-field acoustic environments. However, MCT cannot encode high-dimensional acoustic features for each speaker from mixed input audio because of the interference between speakers. Based on these, we propose the multi-channel multi-speaker transformer (M2Former) for far-field multi-speaker ASR in this paper. Experiments on the SMS-WSJ benchmark show that the M2Former outperforms the neural beamformer, MCT, dual-path RNN with transform-average-concatenate and multi-channel deep clustering based end-to-end systems by 9.2%, 14.3%, 24.9%, and 52.2% respectively, in terms of relative word error rate reduction.
【11】A Music Information Retrieval Approach to Classify Sub-Genres in Role Playing Games
标题:角色扮演游戏子类型分类的音乐信息检索方法
链接:https://arxiv.org/abs/2601.02591
备注:3 pages, 1 figure. D. Hwang, X. Cai, E. Melcer, and E. Carstensdottir, A Music Information Retrieval Approach to Classify Sub-Genres in Role Playing Games, in Extended Abstracts for the Late-Breaking Demo Session of the 25th Int. Society for Music Information Retrieval Conf., San Francisco, United States, 2024
摘要:电子游戏音乐(VGM)通常与电影音乐一样被研究,主要集中在其与媒体类型相关的理论功能上。然而,到目前为止,我们还没有发现任何系统的方法来分析VGM中几种已识别的游戏类型的可量化音乐特征。因此,我们从角色扮演游戏(RPG)的三个子类型的游戏中提取VGM的音乐特征,然后假设不同的音乐特征是如何与每种类型的感知和描绘相关的。这种观察到的相关性可以用于进一步表明这样的特征与预期的讲故事元素或与子流派相关联的游戏机制相关。
摘要:Video game music (VGM) is often studied under the same lens as film music, which largely focuses on its theoretical functionality with relation to the identified genres of the media. However, till date, we are unaware of any systematic approach that analyzes the quantifiable musical features in VGM across several identified game genres. Therefore, we extracted musical features from VGM in games from three sub-genres of Role-Playing Games (RPG), and then hypothesized how different musical features are correlated to the perceptions and portrayals of each genre. This observed correlation may be used to further suggest such features are relevant to the expected storytelling elements or play mechanics associated with the sub-genre.
【12】Understanding Human Perception of Music Plagiarism Through a Computational Approach
标题:通过计算方法理解人类对音乐抄袭的看法
链接:https://arxiv.org/abs/2601.02586
备注:3 pages, D. Hwang and H. Hwang, Understanding Human Perception of Music Plagiarism Through a Computational Approach, in Extended Abstracts for the Late-Breaking Demo Session of the 25th Int. Society for Music Information Retrieval Conf., San Francisco, United States, 2024
摘要:音乐相似性检测算法种类繁多,而现实世界中关于音乐剽窃的讨论往往基于观众的感知。因此,我们的目标是进行一项研究,以检查人类感知音乐剽窃的关键标准,重点是相似性分析中常用的三个音乐特征:旋律,节奏和和弦进行。在确定了人类在感知音乐相似性时使用的关键特征和变化水平之后,我们提出了一个LLM作为判断框架,该框架采用系统的,逐步的方法,利用提取这种高级属性的模块。
摘要:There is a wide variety of music similarity detection algorithms, while discussions about music plagiarism in the real world are often based on audience perceptions. Therefore, we aim to conduct a study to examine the key criteria of human perception of music plagiarism, focusing on the three commonly used musical features in similarity analysis: melody, rhythm, and chord progression. After identifying the key features and levels of variation humans use in perceiving musical similarity, we propose a LLM-as-a-judge framework that applies a systematic, step-by-step approach, drawing on modules that extract such high-level attributes.
【13】Dynamic Quantization Error Propagation in Encoder-Decoder ASR Quantization
标题:编解码器ASB量化中的动态量化误差传播
链接:https://arxiv.org/abs/2601.02455
备注:9 pages, 4 figures, 3 tables
摘要:在内存受限的边缘设备上运行自动语音识别(ASR)模型需要高效压缩。虽然逐层后训练量化是有效的,但其遭受误差累积,特别是在编码器-解码器架构中。现有的解决方案,如量化误差传播(QEP),由于模型的异质性,在编码器中处理声学特征,同时在解码器中生成文本,因此对于ASR来说是次优的。为了解决这个问题,我们提出了动态量化误差传播(FADE)的细粒度Alpha,它自适应地控制跨层纠错和局部量化之间的权衡。实验表明,FADE显着提高了稳定性,减少跨运行的性能差异,同时超过基线的平均WER。
摘要:Running Automatic Speech Recognition (ASR) models on memory-constrained edge devices requires efficient compression. While layer-wise post-training quantization is effective, it suffers from error accumulation, especially in encoder-decoder architectures. Existing solutions like Quantization Error Propagation (QEP) are suboptimal for ASR due to the model's heterogeneity, processing acoustic features in the encoder while generating text in the decoder. To address this, we propose Fine-grained Alpha for Dynamic Quantization Error Propagation (FADE), which adaptively controls the trade-off between cross-layer error correction and local quantization. Experiments show that FADE significantly improves stability by reducing performance variance across runs, while simultaneously surpassing baselines in mean WER.
【14】VocalBridge: Latent Diffusion-Bridge Purification for Defeating Perturbation-Based Voiceprint Defenses
标题:VocalBridge:潜在的扩散桥净化,击败基于微扰的声纹防御
链接:https://arxiv.org/abs/2601.02444
摘要:语音合成技术的快速发展,包括文本到语音(TTS)和语音转换(VC),加剧了与语音克隆相关的安全和隐私问题。最近的防御措施试图通过在语音中嵌入保护性扰动来防止未经授权的克隆,以在保持可理解性的同时模糊说话者的身份。然而,对手可以应用先进的净化技术来消除这些干扰,恢复真实的声学特征,并重新生成可克隆的声音。尽管这种攻击越来越现实,但现有防御在自适应净化下的鲁棒性仍然没有得到充分的研究。 大多数现有的净化方法被设计用于对抗自动语音识别(ASR)系统中的对抗性噪声,而不是说话人验证或语音克隆管道。因此,它们无法抑制定义说话人身份的细粒度声学线索,并且通常对说话人验证攻击(SVA)无效。为了解决这些限制,我们提出了扩散桥(VocalBridge),这是一个净化框架,可以在EnCodec潜在空间中学习从扰动语音到干净语音的潜在映射。该模型使用具有余弦噪声时间表的时间条件化1D U-Net,实现了高效的无转录纯化,同时保留了说话者区分结构。我们进一步介绍了耳语指导的音素变体,采用轻量级的时间指导,而不需要地面实况成绩单。实验结果表明,我们的方法始终优于现有的净化方法在恢复受保护的语音克隆的声音。我们的研究结果证明了当前基于扰动的防御的脆弱性,并强调需要更强大的保护机制来应对不断变化的语音克隆和说话者验证威胁。
摘要:The rapid advancement of speech synthesis technologies, including text-to-speech (TTS) and voice conversion (VC), has intensified security and privacy concerns related to voice cloning. Recent defenses attempt to prevent unauthorized cloning by embedding protective perturbations into speech to obscure speaker identity while maintaining intelligibility. However, adversaries can apply advanced purification techniques to remove these perturbations, recover authentic acoustic characteristics, and regenerate cloneable voices. Despite the growing realism of such attacks, the robustness of existing defenses under adaptive purification remains insufficiently studied. Most existing purification methods are designed to counter adversarial noise in automatic speech recognition (ASR) systems rather than speaker verification or voice cloning pipelines. As a result, they fail to suppress the fine-grained acoustic cues that define speaker identity and are often ineffective against speaker verification attacks (SVA). To address these limitations, we propose Diffusion-Bridge (VocalBridge), a purification framework that learns a latent mapping from perturbed to clean speech in the EnCodec latent space. Using a time-conditioned 1D U-Net with a cosine noise schedule, the model enables efficient, transcript-free purification while preserving speaker-discriminative structure. We further introduce a Whisper-guided phoneme variant that incorporates lightweight temporal guidance without requiring ground-truth transcripts. Experimental results show that our approach consistently outperforms existing purification methods in recovering cloneable voices from protected speech. Our findings demonstrate the fragility of current perturbation-based defenses and highlight the need for more robust protection mechanisms against evolving voice-cloning and speaker verification threats.
【15】Quantifying Quanvolutional Neural Networks Robustness for Speech in Healthcare Applications
标题:量化量子卷积神经网络在医疗保健应用中的语音鲁棒性
链接:https://arxiv.org/abs/2601.02432
摘要:基于语音的机器学习系统对噪声敏感,使情感识别和语音病理检测的可靠部署复杂化。我们评估了混合量子机器学习模型的鲁棒性,quanvolutional神经网络(QNN)对经典卷积神经网络(CNN)在四种声学腐败(高斯噪声,音调偏移,时间偏移和速度变化)下的干净训练/腐败测试制度。使用AVFAD(语音病理学)和TESS(语音情感),我们将三种QNN模型(随机,基本,强)与简单的CNN基线(CNN-Base),ResNet-18和VGG-16进行了比较,使用准确性和腐败指标(CE,mCE,RCE,RmCE),并分析了架构因素(电路复杂性或深度,收敛性)以及每个情感的鲁棒性。QNN通常在音调偏移、时间偏移和速度变化下优于CNN-Base(在严重的时间偏移下,CE/RCE降低高达22%),而CNN-Base对高斯噪声保持更强的弹性。在量子电路中,QNN-Basic在AVFAD上实现了最好的整体鲁棒性,QNN-Random在TESS上表现最强。恐惧是最强大的(在严重腐败下的准确率为80-90%),中性可以在强高斯噪声下崩溃(5.5%的准确率),而快乐最容易受到音调,时间和速度失真的影响。QNN的收敛速度也比CNN-Base快6倍。据我们所知,这是对常见非对抗性声学破坏下的QNN语音鲁棒性的系统研究,表明浅纠缠量子前端可以提高噪声弹性,而对加性噪声的敏感性仍然是一个挑战。
摘要:Speech-based machine learning systems are sensitive to noise, complicating reliable deployment in emotion recognition and voice pathology detection. We evaluate the robustness of a hybrid quantum machine learning model, quanvolutional neural networks (QNNs) against classical convolutional neural networks (CNNs) under four acoustic corruptions (Gaussian noise, pitch shift, temporal shift, and speed variation) in a clean-train/corrupted-test regime. Using AVFAD (voice pathology) and TESS (speech emotion), we compare three QNN models (Random, Basic, Strongly) to a simple CNN baseline (CNN-Base), ResNet-18 and VGG-16 using accuracy and corruption metrics (CE, mCE, RCE, RmCE), and analyze architectural factors (circuit complexity or depth, convergence) alongside per-emotion robustness. QNNs generally outperform the CNN-Base under pitch shift, temporal shift, and speed variation (up to 22% lower CE/RCE at severe temporal shift), while the CNN-Base remains more resilient to Gaussian noise. Among quantum circuits, QNN-Basic achieves the best overall robustness on AVFAD, and QNN-Random performs strongest on TESS. Emotion-wise, fear is most robust (80-90% accuracy under severe corruptions), neutral can collapse under strong Gaussian noise (5.5% accuracy), and happy is most vulnerable to pitch, temporal, and speed distortions. QNNs also converge up to six times faster than the CNN-Base. To our knowledge, this is a systematic study of QNN robustness for speech under common non-adversarial acoustic corruptions, indicating that shallow entangling quantum front-ends can improve noise resilience while sensitivity to additive noise remains a challenge.
【16】WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables
标题:WearVox:一个以自我为中心的可穿戴设备多通道语音助理基准
链接:https://arxiv.org/abs/2601.02391
摘要:人工智能眼镜等可穿戴设备正在将语音助手转变为始终可用的免提协作者,与日常生活无缝集成,但它们也带来了挑战,如受运动和噪音影响的以自我为中心的音频,快速的微交互,以及区分设备导向语音和背景对话的需求。现有的基准在很大程度上忽略了这些复杂性,而是专注于干净或通用的对话音频。为了弥合这一差距,我们提出了WearVox,这是第一个旨在严格评估现实可穿戴场景中语音助手的基准测试。WearVox包括通过人工智能眼镜收集的3,842个多通道,以自我为中心的音频记录,涉及五个不同的任务,包括搜索接地QA,封闭式QA,侧边谈话拒绝,工具调用和语音翻译,涵盖广泛的室内和室外环境和声学条件。每个记录都伴随着丰富的元数据,从而能够在现实世界的约束下对模型性能进行细致入微的分析。我们对领先的专有和开源语音大语言模型(SLLM)进行了基准测试,发现大多数实时SLLM在WearVox上的准确率从29%到59%不等,在嘈杂的户外音频上的性能大幅下降,强调了基准测试的难度和现实性。此外,我们进行了一个案例研究,两个新的SLLM执行推理与单通道和多通道音频,表明多通道音频输入显着增强模型的鲁棒性,环境噪声和改善设备之间的歧视定向和背景语音。我们的研究结果强调了空间音频线索对上下文感知语音助手的至关重要性,并将WearVox建立为推进可穿戴语音AI研究的综合测试平台。
摘要:Wearable devices such as AI glasses are transforming voice assistants into always-available, hands-free collaborators that integrate seamlessly with daily life, but they also introduce challenges like egocentric audio affected by motion and noise, rapid micro-interactions, and the need to distinguish device-directed speech from background conversations. Existing benchmarks largely overlook these complexities, focusing instead on clean or generic conversational audio. To bridge this gap, we present WearVox, the first benchmark designed to rigorously evaluate voice assistants in realistic wearable scenarios. WearVox comprises 3,842 multi-channel, egocentric audio recordings collected via AI glasses across five diverse tasks including Search-Grounded QA, Closed-Book QA, Side-Talk Rejection, Tool Calling, and Speech Translation, spanning a wide range of indoor and outdoor environments and acoustic conditions. Each recording is accompanied by rich metadata, enabling nuanced analysis of model performance under real-world constraints. We benchmark leading proprietary and open-source speech Large Language Models (SLLMs) and find that most real-time SLLMs achieve accuracies on WearVox ranging from 29% to 59%, with substantial performance degradation on noisy outdoor audio, underscoring the difficulty and realism of the benchmark. Additionally, we conduct a case study with two new SLLMs that perform inference with single-channel and multi-channel audio, demonstrating that multi-channel audio inputs significantly enhance model robustness to environmental noise and improve discrimination between device-directed and background speech. Our results highlight the critical importance of spatial audio cues for context-aware voice assistants and establish WearVox as a comprehensive testbed for advancing wearable voice AI research.
【1】Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
标题:迈向细粒度和多粒度对比语音预训练
链接:https://arxiv.org/abs/2601.03065
摘要:建模细粒度的说话风格对于语言-语音表示预训练仍然具有挑战性,因为现有的语音-文本模型通常使用粗略的字幕或特定于任务的监督进行训练,并且不可用可扩展的细粒度风格注释。我们提出了FCaps,这是一个具有细粒度自由文本风格描述的大规模数据集,包含47 k小时的语音和19 M细粒度字幕,这些字幕通过一个新颖的端到端管道进行注释,该管道直接将音频中的详细字幕接地,从而避免了现有级联管道中基于LLM的重写所引起的错误传播。使用LLM-as-a-judge的评估表明,我们的注释在正确性,覆盖率和自然性方面超过了现有的级联注释。基于FCaps,我们提出了CLSP,这是一种对比语言-语音预训练模型,集成了全局和细粒度监督,实现了跨多个粒度的统一表示。大量的实验表明,CLSP学习细粒度和多粒度的语音文本表示,在全球和细粒度的语音文本检索,zero-shot语言分类和语音风格相似性评分,可靠地执行,与人类的判断很强的对齐。所有资源都将公开提供。
摘要:Modeling fine-grained speaking styles remains challenging for language-speech representation pre-training, as existing speech-text models are typically trained with coarse captions or task-specific supervision, and scalable fine-grained style annotations are unavailable. We present FCaps, a large-scale dataset with fine-grained free-text style descriptions, encompassing 47k hours of speech and 19M fine-grained captions annotated via a novel end-to-end pipeline that directly grounds detailed captions in audio, thereby avoiding the error propagation caused by LLM-based rewriting in existing cascaded pipelines. Evaluations using LLM-as-a-judge demonstrate that our annotations surpass existing cascaded annotations in terms of correctness, coverage, and naturalness. Building on FCaps, we propose CLSP, a contrastive language-speech pre-trained model that integrates global and fine-grained supervision, enabling unified representations across multiple granularities. Extensive experiments demonstrate that CLSP learns fine-grained and multi-granular speech-text representations that perform reliably across global and fine-grained speech-text retrieval, zero-shot paralinguistic classification, and speech style similarity scoring, with strong alignment to human judgments. All resources will be made publicly available.
【2】XLSR-MamBo: Scaling the Hybrid Mamba-Attention Backbone for Audio Deepfake Detection
标题:XLSR-Mambo:缩放混合Mamba-Attention骨干用于音频Deepfake检测
链接:https://arxiv.org/abs/2601.02944
备注:11 pages, 3 figures
摘要:先进的语音合成技术已经实现了高度逼真的语音生成,带来了安全风险,促使人们研究音频深度伪造检测(ADD)。虽然状态空间模型(SSM)提供线性复杂度,但纯因果SSM架构通常难以捕获全局频域伪影所需的基于内容的检索。为了解决这个问题,我们探索混合架构的扩展特性,提出XLSR-Mambo,一个模块化的框架集成了XLSR前端协同曼巴注意骨干。我们系统地评估了四种使用高级SSM变体的拓扑设计,Mamba,Mamba 2,Hydra和Gated DeltaNet。实验结果表明,Mambo-3-Hydra-N3配置在ASVspoof 2021 LA,DF和In-the-Wild基准测试中与其他最先进的系统相比具有竞争力的性能。这种性能得益于Hydra的原生双向建模,它比以前的作品中采用的启发式双分支策略更有效地捕获整体时间依赖关系。此外,对DFADD数据集的评估表明,对基于扩散和流匹配的合成方法具有强大的泛化能力。至关重要的是,我们的分析表明,扩展骨干深度有效地减轻了在较浅的模型中观察到的性能差异和不稳定性。这些结果表明,混合框架的能力,以捕捉伪语音信号,提供了一个有效的方法,ADD。
摘要:Advanced speech synthesis technologies have enabled highly realistic speech generation, posing security risks that motivate research into audio deepfake detection (ADD). While state space models (SSMs) offer linear complexity, pure causal SSMs architectures often struggle with the content-based retrieval required to capture global frequency-domain artifacts. To address this, we explore the scaling properties of hybrid architectures by proposing XLSR-MamBo, a modular framework integrating an XLSR front-end with synergistic Mamba-Attention backbones. We systematically evaluate four topological designs using advanced SSM variants, Mamba, Mamba2, Hydra, and Gated DeltaNet. Experimental results demonstrate that the MamBo-3-Hydra-N3 configuration achieves competitive performance compared to other state-of-the-art systems on the ASVspoof 2021 LA, DF, and In-the-Wild benchmarks. This performance benefits from Hydra's native bidirectional modeling, which captures holistic temporal dependencies more efficiently than the heuristic dual-branch strategies employed in prior works. Furthermore, evaluations on the DFADD dataset demonstrate robust generalization to unseen diffusion- and flow-matching-based synthesis methods. Crucially, our analysis reveals that scaling backbone depth effectively mitigates the performance variance and instability observed in shallower models. These results demonstrate the hybrid framework's ability to capture artifacts in spoofed speech signals, providing an effective method for ADD.
【3】Vclip: Face-based Speaker Generation by Face-voice Association Learning
标题:Vclip:通过面部语音关联学习生成基于面部的说话者
链接:https://arxiv.org/abs/2601.02753
备注:work done in 2023
摘要:本文讨论了基于人脸的语音合成的任务,这是一种个性化的语音合成,其中合成的声音被限制为与参考人脸图像在感知上匹配。由于缺乏TTS质量的视听语料库,以前的方法遭受低合成质量或域不匹配引起的知识转移计划。本文提出了一种名为Vclip的新方法,该方法利用CLIP编码器对噪声视听数据的面部语义知识来有效地学习人脸和语音之间的关联,在Voxceleb测试集上实现了89.63%的跨模态验证AUC得分。该方法使用基于检索的策略,结合基于GMM的说话人生成模块,用于下游TTS系统,以产生可能的目标说话人给定的参考图像。实验结果表明,所提出的Vclip系统结合检索步骤,可以弥合人脸和语音特征之间的差距,基于人脸的语音合成。使用从下游TTS提取的反馈信息有助于合成与参考人脸紧密匹配的语音。演示可在www.example.com获得。
摘要:This paper discusses the task of face-based speech synthesis, a kind of personalized speech synthesis where the synthesized voices are con- strained to perceptually match with a reference face image. Due to the lack of TTS-quality audio-visual corpora, previous approaches suffer from either low synthesis quality or domain mismatch induced by a knowledge transfer scheme. This paper proposes a new approach called Vclip that utilizes the facial-semantic knowledge of the CLIP encoder on noisy audio-visual data to learn the association between face and voice efficiently, achieving 89.63% cross-modal verification AUC score on Voxceleb testset. The proposed method then uses a retrieval-based strategy, combined with GMM-based speaker generation module for a downstream TTS system, to produce probable target speakers given reference images. Experimental results demonstrate that the proposed Vclip system in conjunction with the retrieval step can bridge the gap between face and voice features for face-based speech synthesis. And using the feedback information distilled from downstream TTS helps to synthesize voices that match closely with reference faces. Demos available at sos1sos2sixteen.github.io/vclip.
【4】Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models
标题:大型音频语言模型中的描述敏感神经元的发现和因果验证
链接:https://arxiv.org/abs/2601.03115
备注:16 pages, 6 figures
摘要:情感是口语交流的一个核心维度,然而,我们仍然缺乏一个现代大型音频语言模型(LALM)如何内部编码它的机制。我们提出了LALM中情绪敏感神经元(ESN)的第一个神经元水平的可解释性研究,并提供了因果关系的证据,证明这些单位存在于Qwen2.5-Omni,Kimi-Audio和Audio Flamingo 3中。在这三个广泛使用的开源模型中,我们比较了基于频率、熵、幅度和对比度的神经元选择器在多个情感识别基准上的表现。使用推理时间干预,我们揭示了一致的情绪特异性签名:消融为给定情绪选择的神经元不成比例地降低了对该情绪的识别,同时在很大程度上保留了其他类别,而基于增益的放大则将预测转向目标情绪。这些影响出现在适度的识别数据和规模与干预力度系统。我们进一步观察到,ESN表现出不均匀的逐层聚类与部分跨数据集传输。总之,我们的研究结果提供了一个因果关系,神经元水平的帐户的情绪决定LALM和突出有针对性的神经元干预作为一个可操作的处理可控的情感行为。
摘要:Emotion is a central dimension of spoken communication, yet, we still lack a mechanistic account of how modern large audio-language models (LALMs) encode it internally. We present the first neuron-level interpretability study of emotion-sensitive neurons (ESNs) in LALMs and provide causal evidence that such units exist in Qwen2.5-Omni, Kimi-Audio, and Audio Flamingo 3. Across these three widely used open-source models, we compare frequency-, entropy-, magnitude-, and contrast-based neuron selectors on multiple emotion recognition benchmarks. Using inference-time interventions, we reveal a consistent emotion-specific signature: ablating neurons selected for a given emotion disproportionately degrades recognition of that emotion while largely preserving other classes, whereas gain-based amplification steers predictions toward the target emotion. These effects arise with modest identification data and scale systematically with intervention strength. We further observe that ESNs exhibit non-uniform layer-wise clustering with partial cross-dataset transfer. Taken together, our results offer a causal, neuron-level account of emotion decisions in LALMs and highlight targeted neuron interventions as an actionable handle for controllable affective behaviors.
【5】MoE Adapter for Large Audio Language Models: Sparsity, Disentanglement, and Gradient-Conflict-Free
标题:适用于大型音频语言模型的MoE适配器:稀疏性、解纠缠和无冲突
链接:https://arxiv.org/abs/2601.02967
备注:13 pages, 5 figures
摘要:将大语言模型的输入模态扩展到音频域是实现多模态感知的关键。然而,众所周知,声学信息本质上是\textit{异构}的,纠缠的属性,如语音,音乐和环境背景。现有的研究仅限于一个密集的,参数共享的适配器来模拟这些不同的模式,这会导致优化过程中的\textit{梯度冲突},因为不同属性所需的参数更新相互矛盾。为了解决这个问题,我们引入了\textit{\textbf{MoE-Adapter}},一个稀疏的混合专家~(MoE)架构,旨在解耦声学信息。具体而言,它采用了一种动态门控机制,将音频令牌路由到捕获互补特征子空间的专业专家,同时保留全局上下文的共享专家,从而减轻梯度冲突并实现细粒度特征学习。综合实验表明,MoE-Adapter在音频语义和语言任务上都取得了优异的性能,在计算成本相当的情况下,始终优于密集的线性基线。此外,我们将发布相关的代码和模型,以方便未来的研究。
摘要:Extending the input modality of Large Language Models~(LLMs) to the audio domain is essential for achieving comprehensive multimodal perception. However, it is well-known that acoustic information is intrinsically \textit{heterogeneous}, entangling attributes such as speech, music, and environmental context. Existing research is limited to a dense, parameter-shared adapter to model these diverse patterns, which induces \textit{gradient conflict} during optimization, as parameter updates required for distinct attributes contradict each other. To address this limitation, we introduce the \textit{\textbf{MoE-Adapter}}, a sparse Mixture-of-Experts~(MoE) architecture designed to decouple acoustic information. Specifically, it employs a dynamic gating mechanism that routes audio tokens to specialized experts capturing complementary feature subspaces while retaining shared experts for global context, thereby mitigating gradient conflicts and enabling fine-grained feature learning. Comprehensive experiments show that the MoE-Adapter achieves superior performance on both audio semantic and paralinguistic tasks, consistently outperforming dense linear baselines with comparable computational costs. Furthermore, we will release the related code and models to facilitate future research.
【6】SPO-CLAPScore: Enhancing CLAP-based alignment prediction system with Standardize Preference Optimization, for the first XACLE Challenge
标题:SPO-CLAPScore:通过标准化偏好优化增强基于CLAP的对齐预测系统,参加首届XACLE挑战赛
链接:https://arxiv.org/abs/2601.02900
备注:https://github.com/ttakano398/SPO-CLAPScore
摘要:第一个XACLE挑战(x到音频对齐挑战)解决了与人类感知的音频文本语义对齐相关的自动评估指标的关键需求。在本文中,我们描述了提交给XACLE挑战赛的“Takano_UTokyo_03”系统。我们的方法利用了基于CLAPScore的架构,并集成了一种名为标准化偏好优化(SPO)的新型训练方法。SPO对每个听众提供的原始对齐分数进行分析,使模型能够学习相对偏好并减轻个人评分偏差的影响。此外,我们采用听众筛选,以排除听众与不一致的评级。实验结果表明,SPO和听者筛选都有效地提高了与人类判断的相关性。我们的系统在挑战中获得了第6名,斯皮尔曼等级相关系数(SRCC)为0.6142,显示出与顶级系统的边缘差距内的竞争力。该代码可在https://github.com/ttakano398/SPO-CLAPScore上获得。
摘要:The first XACLE Challenge (x-to-audio alignment challenge) addresses the critical need for automatic evaluation metrics that correlate with human perception of audio-text semantic alignment. In this paper, we describe the "Takano_UTokyo_03" system submitted to XACLE Challenge. Our approach leverages a CLAPScore-based architecture integrated with a novel training method called Standardized Preference Optimization (SPO). SPO standardizes the raw alignment scores provided by each listener, enabling the model to learn relative preferences and mitigate the impact of individual scoring biases. Additionally, we employ listener screening to exclude listeners with inconsistent ratings. Experimental evaluations demonstrate that both SPO and listener screening effectively improve the correlation with human judgment. Our system achieved 6th place in the challenge with a Spearman's rank correlation coefficient (SRCC) of 0.6142, demonstrating competitive performance within a marginal gap from the top-ranked systems. The code is available at https://github.com/ttakano398/SPO-CLAPScore.
【7】Dynamic Quantization Error Propagation in Encoder-Decoder ASR Quantization
标题:编解码器ASB量化中的动态量化误差传播
链接:https://arxiv.org/abs/2601.02455
备注:9 pages, 4 figures, 3 tables
摘要:在内存受限的边缘设备上运行自动语音识别(ASR)模型需要高效压缩。虽然逐层后训练量化是有效的,但其遭受误差累积,特别是在编码器-解码器架构中。现有的解决方案,如量化误差传播(QEP),由于模型的异质性,在编码器中处理声学特征,同时在解码器中生成文本,因此对于ASR来说是次优的。为了解决这个问题,我们提出了动态量化误差传播(FADE)的细粒度Alpha,它自适应地控制跨层纠错和局部量化之间的权衡。实验表明,FADE显着提高了稳定性,减少跨运行的性能差异,同时超过基线的平均WER。
摘要:Running Automatic Speech Recognition (ASR) models on memory-constrained edge devices requires efficient compression. While layer-wise post-training quantization is effective, it suffers from error accumulation, especially in encoder-decoder architectures. Existing solutions like Quantization Error Propagation (QEP) are suboptimal for ASR due to the model's heterogeneity, processing acoustic features in the encoder while generating text in the decoder. To address this, we propose Fine-grained Alpha for Dynamic Quantization Error Propagation (FADE), which adaptively controls the trade-off between cross-layer error correction and local quantization. Experiments show that FADE significantly improves stability by reducing performance variance across runs, while simultaneously surpassing baselines in mean WER.
【8】VocalBridge: Latent Diffusion-Bridge Purification for Defeating Perturbation-Based Voiceprint Defenses
标题:VocalBridge:潜在的扩散桥净化,击败基于微扰的声纹防御
链接:https://arxiv.org/abs/2601.02444
摘要:语音合成技术的快速发展,包括文本到语音(TTS)和语音转换(VC),加剧了与语音克隆相关的安全和隐私问题。最近的防御措施试图通过在语音中嵌入保护性扰动来防止未经授权的克隆,以在保持可理解性的同时模糊说话者的身份。然而,对手可以应用先进的净化技术来消除这些干扰,恢复真实的声学特征,并重新生成可克隆的声音。尽管这种攻击越来越现实,但现有防御在自适应净化下的鲁棒性仍然没有得到充分的研究。 大多数现有的净化方法被设计用于对抗自动语音识别(ASR)系统中的对抗性噪声,而不是说话人验证或语音克隆管道。因此,它们无法抑制定义说话人身份的细粒度声学线索,并且通常对说话人验证攻击(SVA)无效。为了解决这些局限性,我们提出了扩散桥(VocalBridge),一个净化框架,它在EnCodec潜在空间中学习从扰动到干净语音的潜在映射。该模型使用具有余弦噪声时间表的时间条件化1D U-Net,实现了高效的无转录纯化,同时保留了说话者区分结构。我们进一步介绍了耳语指导的音素变体,采用轻量级的时间指导,而不需要地面实况成绩单。实验结果表明,我们的方法始终优于现有的净化方法在恢复受保护的语音克隆的声音。我们的研究结果证明了当前基于扰动的防御的脆弱性,并强调了需要更强大的保护机制来应对不断发展的语音克隆和说话人验证威胁。
摘要:The rapid advancement of speech synthesis technologies, including text-to-speech (TTS) and voice conversion (VC), has intensified security and privacy concerns related to voice cloning. Recent defenses attempt to prevent unauthorized cloning by embedding protective perturbations into speech to obscure speaker identity while maintaining intelligibility. However, adversaries can apply advanced purification techniques to remove these perturbations, recover authentic acoustic characteristics, and regenerate cloneable voices. Despite the growing realism of such attacks, the robustness of existing defenses under adaptive purification remains insufficiently studied. Most existing purification methods are designed to counter adversarial noise in automatic speech recognition (ASR) systems rather than speaker verification or voice cloning pipelines. As a result, they fail to suppress the fine-grained acoustic cues that define speaker identity and are often ineffective against speaker verification attacks (SVA). To address these limitations, we propose Diffusion-Bridge (VocalBridge), a purification framework that learns a latent mapping from perturbed to clean speech in the EnCodec latent space. Using a time-conditioned 1D U-Net with a cosine noise schedule, the model enables efficient, transcript-free purification while preserving speaker-discriminative structure. We further introduce a Whisper-guided phoneme variant that incorporates lightweight temporal guidance without requiring ground-truth transcripts. Experimental results show that our approach consistently outperforms existing purification methods in recovering cloneable voices from protected speech. Our findings demonstrate the fragility of current perturbation-based defenses and highlight the need for more robust protection mechanisms against evolving voice-cloning and speaker verification threats.
【9】Quantifying Quanvolutional Neural Networks Robustness for Speech in Healthcare Applications
标题:量化量子卷积神经网络在医疗保健应用中的语音鲁棒性
链接:https://arxiv.org/abs/2601.02432
摘要:基于语音的机器学习系统对噪声敏感,使情感识别和语音病理检测的可靠部署复杂化。我们评估了混合量子机器学习模型的鲁棒性,quanvolutional神经网络(QNN)对经典卷积神经网络(CNN)在四种声学腐败(高斯噪声,音调偏移,时间偏移和速度变化)下的干净训练/腐败测试制度。使用AVFAD(语音病理学)和TESS(语音情感),我们将三种QNN模型(随机,基本,强)与简单的CNN基线(CNN-Base),ResNet-18和VGG-16进行了比较,使用准确性和腐败指标(CE,mCE,RCE,RmCE),并分析了架构因素(电路复杂性或深度,收敛性)以及每个情感的鲁棒性。QNN通常在音调偏移、时间偏移和速度变化下优于CNN-Base(在严重的时间偏移下,CE/RCE降低高达22%),而CNN-Base对高斯噪声保持更强的弹性。在量子电路中,QNN-Basic在AVFAD上实现了最好的整体鲁棒性,QNN-Random在TESS上表现最强。恐惧是最强大的(在严重腐败下的准确率为80-90%),中性可以在强高斯噪声下崩溃(5.5%的准确率),而快乐最容易受到音调,时间和速度失真的影响。QNN的收敛速度也比CNN-Base快6倍。据我们所知,这是对常见非对抗性声学破坏下的QNN语音鲁棒性的系统研究,表明浅纠缠量子前端可以提高噪声弹性,而对加性噪声的敏感性仍然是一个挑战。
摘要:Speech-based machine learning systems are sensitive to noise, complicating reliable deployment in emotion recognition and voice pathology detection. We evaluate the robustness of a hybrid quantum machine learning model, quanvolutional neural networks (QNNs) against classical convolutional neural networks (CNNs) under four acoustic corruptions (Gaussian noise, pitch shift, temporal shift, and speed variation) in a clean-train/corrupted-test regime. Using AVFAD (voice pathology) and TESS (speech emotion), we compare three QNN models (Random, Basic, Strongly) to a simple CNN baseline (CNN-Base), ResNet-18 and VGG-16 using accuracy and corruption metrics (CE, mCE, RCE, RmCE), and analyze architectural factors (circuit complexity or depth, convergence) alongside per-emotion robustness. QNNs generally outperform the CNN-Base under pitch shift, temporal shift, and speed variation (up to 22% lower CE/RCE at severe temporal shift), while the CNN-Base remains more resilient to Gaussian noise. Among quantum circuits, QNN-Basic achieves the best overall robustness on AVFAD, and QNN-Random performs strongest on TESS. Emotion-wise, fear is most robust (80-90% accuracy under severe corruptions), neutral can collapse under strong Gaussian noise (5.5% accuracy), and happy is most vulnerable to pitch, temporal, and speed distortions. QNNs also converge up to six times faster than the CNN-Base. To our knowledge, this is a systematic study of QNN robustness for speech under common non-adversarial acoustic corruptions, indicating that shallow entangling quantum front-ends can improve noise resilience while sensitivity to additive noise remains a challenge.
【10】WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables
标题:WearVox:一个以自我为中心的可穿戴设备多通道语音助理基准
链接:https://arxiv.org/abs/2601.02391
摘要:人工智能眼镜等可穿戴设备正在将语音助手转变为始终可用的免提协作者,与日常生活无缝集成,但它们也带来了挑战,如受运动和噪音影响的以自我为中心的音频,快速的微交互,以及区分设备导向语音和背景对话的需求。现有的基准在很大程度上忽略了这些复杂性,而是专注于干净或通用的对话音频。为了弥合这一差距,我们提出了WearVox,这是第一个旨在严格评估现实可穿戴场景中语音助手的基准测试。WearVox包括通过人工智能眼镜收集的3,842个多通道,以自我为中心的音频记录,涉及五个不同的任务,包括搜索接地QA,封闭式QA,侧边谈话拒绝,工具调用和语音翻译,涵盖广泛的室内和室外环境和声学条件。每个记录都伴随着丰富的元数据,从而能够在现实世界的约束下对模型性能进行细致入微的分析。我们对领先的专有和开源语音大语言模型(SLLM)进行了基准测试,发现大多数实时SLLM在WearVox上的准确率从29%到59%不等,在嘈杂的户外音频上的性能大幅下降,强调了基准测试的难度和现实性。此外,我们进行了一个案例研究,两个新的SLLM执行推理与单通道和多通道音频,表明多通道音频输入显着增强模型的鲁棒性,环境噪声和改善设备之间的歧视定向和背景语音。我们的研究结果强调了空间音频线索对上下文感知语音助手的至关重要性,并将WearVox建立为推进可穿戴语音AI研究的综合测试平台。
摘要:Wearable devices such as AI glasses are transforming voice assistants into always-available, hands-free collaborators that integrate seamlessly with daily life, but they also introduce challenges like egocentric audio affected by motion and noise, rapid micro-interactions, and the need to distinguish device-directed speech from background conversations. Existing benchmarks largely overlook these complexities, focusing instead on clean or generic conversational audio. To bridge this gap, we present WearVox, the first benchmark designed to rigorously evaluate voice assistants in realistic wearable scenarios. WearVox comprises 3,842 multi-channel, egocentric audio recordings collected via AI glasses across five diverse tasks including Search-Grounded QA, Closed-Book QA, Side-Talk Rejection, Tool Calling, and Speech Translation, spanning a wide range of indoor and outdoor environments and acoustic conditions. Each recording is accompanied by rich metadata, enabling nuanced analysis of model performance under real-world constraints. We benchmark leading proprietary and open-source speech Large Language Models (SLLMs) and find that most real-time SLLMs achieve accuracies on WearVox ranging from 29% to 59%, with substantial performance degradation on noisy outdoor audio, underscoring the difficulty and realism of the benchmark. Additionally, we conduct a case study with two new SLLMs that perform inference with single-channel and multi-channel audio, demonstrating that multi-channel audio inputs significantly enhance model robustness to environmental noise and improve discrimination between device-directed and background speech. Our results highlight the critical importance of spatial audio cues for context-aware voice assistants and establish WearVox as a comprehensive testbed for advancing wearable voice AI research.
机器翻译由腾讯交互翻译提供,仅供参考
