微信公众号:arXiv_Daily
cs.SD语音
【1】VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
标题:VoiceAssistant-Eval:在听力、口语和观看方面对人工智能助理进行基准测试
链接:https://arxiv.org/abs/2509.22651
摘要:The growing capabilities of large language models and multimodal systems have spurred interest in voice-first AI assistants, yet existing benchmarks are inadequate for evaluating the full range of these systems' capabilities. We introduce VoiceAssistant-Eval, a comprehensive benchmark designed to assess AI assistants across listening, speaking, and viewing. VoiceAssistant-Eval comprises 10,497 curated examples spanning 13 task categories. These tasks include natural sounds, music, and spoken dialogue for listening; multi-turn dialogue, role-play imitation, and various scenarios for speaking; and highly heterogeneous images for viewing. To demonstrate its utility, we evaluate 21 open-source models and GPT-4o-Audio, measuring the quality of the response content and speech, as well as their consistency. The results reveal three key findings: (1) proprietary models do not universally outperform open-source models; (2) most models excel at speaking tasks but lag in audio understanding; and (3) well-designed smaller models can rival much larger ones. Notably, the mid-sized Step-Audio-2-mini (7B) achieves more than double the listening accuracy of LLaMA-Omni2-32B-Bilingual. However, challenges remain: multimodal (audio plus visual) input and role-play voice imitation tasks are difficult for current models, and significant gaps persist in robustness and safety alignment. VoiceAssistant-Eval identifies these gaps and establishes a rigorous framework for evaluating and guiding the development of next-generation AI assistants. Code and data will be released at https://mathllm.github.io/VoiceAssistantEval/ .
【2】MDAR: A Multi-scene Dynamic Audio Reasoning Benchmark
标题:MDAR:多场景动态音频推理基准
链接:https://arxiv.org/abs/2509.22461
摘要:从音频(包括语音、非语言提示、环境声音和音乐)中推理的能力对于AI代理在现实世界中有效交互至关重要。现有的基准测试主要集中在静态或单场景设置上,并且不能完全捕获多个扬声器、展开事件和异构音频源交互的场景。为了应对这些挑战,我们引入了MDAR,这是一个用于评估复杂,多场景和动态演变的音频推理任务模型的基准。MDAR包含3,000个精心策划的问答对,与不同的音频片段相关联,涵盖五类复杂推理和三种问题类型。我们对MDAR上的26种最先进的音频语言模型进行了基准测试,并观察到它们在复杂的推理任务中表现出局限性。在单项选择题上,Qwen2.5-Omni(开源)达到了76.67%的准确率,而GPT-4 o Audio(闭源)达到了68.47%;然而,GPT-4 o Audio在更具挑战性的多项选择和开放式任务上大大优于Qwen2.5-Omni。在所有三种问题类型中,没有模型达到80%的性能。这些发现强调了MDAR带来的独特挑战及其作为推进音频推理研究的基准的价值。代码和基准可以在https://github.com/luckyerr/MDAR上找到。
摘要:The ability to reason from audio, including speech, paralinguistic cues, environmental sounds, and music, is essential for AI agents to interact effectively in real-world scenarios. Existing benchmarks mainly focus on static or single-scene settings and do not fully capture scenarios where multiple speakers, unfolding events, and heterogeneous audio sources interact. To address these challenges, we introduce MDAR, a benchmark for evaluating models on complex, multi-scene, and dynamically evolving audio reasoning tasks. MDAR comprises 3,000 carefully curated question-answer pairs linked to diverse audio clips, covering five categories of complex reasoning and spanning three question types. We benchmark 26 state-of-the-art audio language models on MDAR and observe that they exhibit limitations in complex reasoning tasks. On single-choice questions, Qwen2.5-Omni (open-source) achieves 76.67% accuracy, whereas GPT-4o Audio (closed-source) reaches 68.47%; however, GPT-4o Audio substantially outperforms Qwen2.5-Omni on the more challenging multiple-choice and open-ended tasks. Across all three question types, no model achieves 80% performance. These findings underscore the unique challenges posed by MDAR and its value as a benchmark for advancing audio reasoning research.Code and benchmark can be found at https://github.com/luckyerr/MDAR.
【3】From Coarse to Fine: Recursive Audio-Visual Semantic Enhancement for Speech Separation
标题:从粗到细:语音分离的递进视听语义增强
链接:https://arxiv.org/abs/2509.22425
摘要:Audio-visual speech separation aims to isolate each speaker's clean voice from mixtures by leveraging visual cues such as lip movements and facial features. While visual information provides complementary semantic guidance, existing methods often underexploit its potential by relying on static visual representations. In this paper, we propose CSFNet, a Coarse-to-Separate-Fine Network that introduces a recursive semantic enhancement paradigm for more effective separation. CSFNet operates in two stages: (1) Coarse Separation, where a first-pass estimation reconstructs a coarse audio waveform from the mixture and visual input; and (2) Fine Separation, where the coarse audio is fed back into an audio-visual speech recognition (AVSR) model together with the visual stream. This recursive process produces more discriminative semantic representations, which are then used to extract refined audio. To further exploit these semantics, we design a speaker-aware perceptual fusion block to encode speaker identity across modalities, and a multi-range spectro-temporal separation network to capture both local and global time-frequency patterns. Extensive experiments on three benchmark datasets and two noisy datasets show that CSFNet achieves state-of-the-art (SOTA) performance, with substantial coarse-to-fine improvements, validating the necessity and effectiveness of our recursive semantic enhancement framework.
【4】Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach
标题:毫不费力的图像到音乐生成:一种可解释的基于RAG的VLM方法
链接:https://arxiv.org/abs/2509.22378
摘要:Recently, Image-to-Music (I2M) generation has garnered significant attention, with potential applications in fields such as gaming, advertising, and multi-modal art creation. However, due to the ambiguous and subjective nature of I2M tasks, most end-to-end methods lack interpretability, leaving users puzzled about the generation results. Even methods based on emotion mapping face controversy, as emotion represents only a singular aspect of art. Additionally, most learning-based methods require substantial computational resources and large datasets for training, hindering accessibility for common users. To address these challenges, we propose the first Vision Language Model (VLM)-based I2M framework that offers high interpretability and low computational cost. Specifically, we utilize ABC notation to bridge the text and music modalities, enabling the VLM to generate music using natural language. We then apply multi-modal Retrieval-Augmented Generation (RAG) and self-refinement techniques to allow the VLM to produce high-quality music without external training. Furthermore, we leverage the generated motivations in text and the attention maps from the VLM to provide explanations for the generated results in both text and image modalities. To validate our method, we conduct both human studies and machine evaluations, where our method outperforms others in terms of music quality and music-image consistency, indicating promising results. Our code is available at https://github.com/RS2002/Image2Music .
【5】Cross-Dialect Bird Species Recognition with Dialect-Calibrated Augmentation
标题:使用方言校准增强的跨方言鸟类物种识别
链接:https://arxiv.org/abs/2509.22317
摘要:Dialect variation hampers automatic recognition of bird calls collected by passive acoustic monitoring. We address the problem on DB3V, a three-region, ten-species corpus of 8-s clips, and propose a deployable framework built on Time-Delay Neural Networks (TDNNs). Frequency-sensitive normalisation (Instance Frequency Normalisation and a gated Relaxed-IFN) is paired with gradient-reversal adversarial training to learn region-invariant embeddings. A multi-level augmentation scheme combines waveform perturbations, Mixup for rare classes, and CycleGAN transfer that synthesises Region 2 (Interior Plains)-style audio, , with Dialect-Calibrated Augmentation (DCA) softly down-weighting synthetic samples to limit artifacts. The complete system lifts cross-dialect accuracy by up to twenty percentage points over baseline TDNNs while preserving in-region performance. Grad-CAM and LIME analyses show that robust models concentrate on stable harmonic bands, providing ecologically meaningful explanations. The study demonstrates that lightweight, transparent, and dialect-resilient bird-sound recognition is attainable.
【6】High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling
标题:通过视觉引导生成建模跨不同类别的高质量声音分离
链接:https://arxiv.org/abs/2509.22063
摘要:我们提出了DAVIS,一个基于扩散的视听分离框架,通过生成学习解决了视听声源分离任务。现有的方法通常将声音分离作为基于掩码的回归问题,取得了显着的进展。然而,它们在捕获高质量分离不同类别声音所需的复杂数据分布方面面临局限性。相比之下,DAVIS通过利用有效的生成建模范式来规避这些问题,特别是去噪扩散概率模型(DDPM)和最近的流量匹配(FM),集成在一个专门的分离U-Net架构中。我们的框架通过直接从噪声分布中合成所需的分离声谱图来操作,同时以混合音频输入和相关的视觉信息为条件。其生成目标的固有性质使DAVIS特别擅长为不同的声音类别制作高质量的声音分离。我们目前的DAVIS的比较评估,包括其DDPM和流量匹配的变种,对领先的方法标准AVE和MUSIC数据集。结果证实,这两种变体在分离质量上都超过了现有的方法,突出了我们的生成框架在处理视听源分离任务方面的有效性。
摘要:We propose DAVIS, a Diffusion-based Audio-VIsual Separation framework that solves the audio-visual sound source separation task through generative learning. Existing methods typically frame sound separation as a mask-based regression problem, achieving significant progress. However, they face limitations in capturing the complex data distribution required for high-quality separation of sounds from diverse categories. In contrast, DAVIS circumvents these issues by leveraging potent generative modeling paradigms, specifically Denoising Diffusion Probabilistic Models (DDPM) and the more recent Flow Matching (FM), integrated within a specialized Separation U-Net architecture. Our framework operates by synthesizing the desired separated sound spectrograms directly from a noise distribution, conditioned concurrently on the mixed audio input and associated visual information. The inherent nature of its generative objective makes DAVIS particularly adept at producing high-quality sound separations for diverse sound categories. We present comparative evaluations of DAVIS, encompassing both its DDPM and Flow Matching variants, against leading methods on the standard AVE and MUSIC datasets. The results affirm that both variants surpass existing approaches in separation quality, highlighting the efficacy of our generative framework for tackling the audio-visual source separation task.
【7】Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
标题:理解和交谈:通过双语言建模进行文本到语音合成
链接:https://arxiv.org/abs/2509.22062
摘要:现有的基于大语言模型(LLM)的自回归(AR)文本到语音(TTS)系统虽然达到了最先进的质量,但仍然面临着严峻的挑战。这种基于LLM的范例的基础是通过神经音频编解码器将连续语音波形离散化为离散令牌序列。然而,单码本建模非常适合于文本LLM,但遭受显著的信息丢失;通常通过残差矢量量化(RVQ)生成的分层声学令牌通常缺乏明确的语义结构,从而给模型带来沉重的学习负担。此外,自回归过程固有地易受误差累积的影响,这会降低发电稳定性。为了解决这些局限性,我们提出了CaT-TTS,一个新的框架,强大的和语义接地zero-shot合成。首先,我们介绍了S3 Codec,一个分裂的RVQ编解码器,通过从最先进的ASR模型中提取语义,将显式语言特征注入其主码本,提供了一个结构化的表示,简化了学习任务。其次,我们提出了一个“理解然后生成”的双变压器架构,从渲染的理解。初始“理解”Transformer对文本和音频的语义标记之间的跨模态关系进行建模,以形成高级话语计划。随后的“生成”Transformer执行该计划,自回归合成分层声学标记。最后,为了提高生成稳定性,我们引入了掩蔽音频并行推理(Masked Audio Parallel Inference,简称MFS),这是一种几乎无参数的推理策略,可以动态地指导解码过程以减轻局部错误。
摘要:Existing Large Language Model (LLM) based autoregressive (AR) text-to-speech (TTS) systems, while achieving state-of-the-art quality, still face critical challenges. The foundation of this LLM-based paradigm is the discretization of the continuous speech waveform into a sequence of discrete tokens by neural audio codec. However, single codebook modeling is well suited to text LLMs, but suffers from significant information loss; hierarchical acoustic tokens, typically generated via Residual Vector Quantization (RVQ), often lack explicit semantic structure, placing a heavy learning burden on the model. Furthermore, the autoregressive process is inherently susceptible to error accumulation, which can degrade generation stability. To address these limitations, we propose CaT-TTS, a novel framework for robust and semantically-grounded zero-shot synthesis. First, we introduce S3Codec, a split RVQ codec that injects explicit linguistic features into its primary codebook via semantic distillation from a state-of-the-art ASR model, providing a structured representation that simplifies the learning task. Second, we propose an ``Understand-then-Generate'' dual-Transformer architecture that decouples comprehension from rendering. An initial ``Understanding'' Transformer models the cross-modal relationship between text and the audio's semantic tokens to form a high-level utterance plan. A subsequent ``Generation'' Transformer then executes this plan, autoregressively synthesizing hierarchical acoustic tokens. Finally, to enhance generation stability, we introduce Masked Audio Parallel Inference (MAPI), a nearly parameter-free inference strategy that dynamically guides the decoding process to mitigate local errors.
【8】Decoding Deception: Understanding Automatic Speech Recognition Vulnerabilities in Evasion and Poisoning Attacks
标题:解码欺骗:了解逃避和中毒攻击中的自动语音识别漏洞
链接:https://arxiv.org/abs/2509.22060
摘要:Recent studies have demonstrated the vulnerability of Automatic Speech Recognition systems to adversarial examples, which can deceive these systems into misinterpreting input speech commands. While previous research has primarily focused on white-box attacks with constrained optimizations, and transferability based black-box attacks against commercial Automatic Speech Recognition devices, this paper explores cost efficient white-box attack and non transferability black-box adversarial attacks on Automatic Speech Recognition systems, drawing insights from approaches such as Fast Gradient Sign Method and Zeroth-Order Optimization. Further, the novelty of the paper includes how poisoning attack can degrade the performances of state-of-the-art models leading to misinterpretation of audio signals. Through experimentation and analysis, we illustrate how hybrid models can generate subtle yet impactful adversarial examples with very little perturbation having Signal Noise Ratio of 35dB that can be generated within a minute. These vulnerabilities of state-of-the-art open source model have practical security implications, and emphasize the need for adversarial security.
【9】WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
标题:WAVE:通过多模式LLM学习统一和多功能视听嵌入
链接:https://arxiv.org/abs/2509.21990
摘要:While embeddings from multimodal large language models (LLMs) excel as general-purpose representations, their application to dynamic modalities like audio and video remains underexplored. We introduce WAVE (\textbf{u}nified \& \textbf{v}ersatile \textbf{a}udio-\textbf{v}isual \textbf{e}mbeddings), the first LLM-based embedding that creates a unified representation space for text, audio, and video modalities. WAVE employs a novel hierarchical feature fusion strategy and a joint multi-modal, multi-task training approach to enable two key capabilities: any-to-any cross-modal retrieval and the generation of prompt-aware embeddings tailored to user instructions. Experimentally, WAVE sets a new state-of-the-art on the MMEB-v2 video benchmark and achieves superior results in audio and video-to-audio retrieval. Its prompt-aware nature also yields remarkable performance in multimodal question answering, significantly outperforming existing embedding models. Ablation studies validate our joint training strategy, demonstrating improved performance across all modalities. With a newly introduced benchmark for versatile audio-visual learning, WAVE opens up broad possibilities for cross-modal, any-to-any applications. Our code, checkpoints, and data will be released.
【10】A Parallel Ultra-Low Power Silent Speech Interface based on a Wearable, Fully-dry EMG Neckband
标题:基于可穿戴、全干EMG项圈的并行超低功耗无声语音接口
链接:https://arxiv.org/abs/2509.21964
摘要:我们提出了一个可穿戴的,全干燥的,超低功耗的EMG系统,用于无声语音识别,集成到一个纺织品颈带,以实现舒适,非侵入性的使用。该系统具有14个全差分EMG通道,基于BioGAP-Ultra平台,可实现超低功耗(22 mW)生物信号采集和无线传输。我们评估其性能的8个语音命令下发声和无声的清晰度,实现平均分类准确率分别为87$\pm$3%和68$\pm$3%,与5倍CV的方法。为了模拟日常生活条件,我们引入会话到会话的变化,通过重新定位会话之间的颈带,实现了64$\pm$18%和54$\pm$7%的发音和无声实验,分别留一个会话的准确性。这些结果突出了所提出的方法的鲁棒性和节能无声语音解码的承诺。
摘要:We present a wearable, fully-dry, and ultra-low power EMG system for silent speech recognition, integrated into a textile neckband to enable comfortable, non-intrusive use. The system features 14 fully-differential EMG channels and is based on the BioGAP-Ultra platform for ultra-low power (22 mW) biosignal acquisition and wireless transmission. We evaluate its performance on eight speech commands under both vocalized and silent articulation, achieving average classification accuracies of 87$\pm$3% and 68$\pm$3% respectively, with a 5-fold CV approach. To mimic everyday-life conditions, we introduce session-to-session variability by repositioning the neckband between sessions, achieving leave-one-session-out accuracies of 64$\pm$18% and 54$\pm$7% for the vocalized and silent experiments, respectively. These results highlight the robustness of the proposed approach and the promise of energy-efficient silent-speech decoding.
【11】Text2Move: Text-to-moving sound generation via trajectory prediction and temporal alignment
标题:文本2Move:通过轨迹预测和时间对齐生成文本到运动的声音
链接:https://arxiv.org/abs/2509.21919
【12】Lightweight Front-end Enhancement for Robust ASR via Frame Resampling and Sub-Band Pruning
标题:通过帧重排序和子带修剪实现鲁棒ASB的轻量级前端增强
链接:https://arxiv.org/abs/2509.21833
摘要:自动语音识别(ASR)的最新进展已经取得了显着的进步,而在嘈杂的环境中的鲁棒性仍然具有挑战性。虽然语音增强(SE)前端被广泛用于减轻噪声作为ASR的预处理步骤,但它们通常会引入不可忽略的计算开销。本文提出了在不影响ASR性能的情况下降低SE计算成本的优化方法。我们的方法集成了逐层帧重采样和渐进子带修剪。帧恢复对层内的输入进行下采样,利用残差连接来减轻信息丢失。同时,子带修剪逐步排除信息量较少的频带,进一步减少计算需求。在合成和真实世界的噪声数据集上进行的大量实验表明,与标准BSRNN相比,我们的系统将SE计算开销降低了66以上,同时保持了强大的ASR性能。
【13】Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models
标题:用声音思考:音频思维链在大型音频语言模型中实现多模式推理
链接:https://arxiv.org/abs/2509.21749
【14】Noise-to-Notes: Diffusion-based Generation and Refinement for Automatic Drum Transcription
标题:噪音笔记:基于扩散的自动鼓转录生成和细化
链接:https://arxiv.org/abs/2509.21739
【15】Frustratingly Easy Zero-Day Audio DeepFake Detection via Retrieval Augmentation and Profile Matching
标题:令人沮丧的简单零日音频DeepFake检测通过检索增强和配置文件匹配
链接:https://arxiv.org/abs/2509.21728
【16】MusicWeaver: Coherent Long-Range and Editable Music Generation from a Beat-Aligned Structural Plan
标题:MusicWeaver:从与节拍一致的结构计划中产生连贯的长期和可编辑的音乐一代
链接:https://arxiv.org/abs/2509.21714
摘要:当前的音乐生成器捕获本地纹理,但通常无法对远程结构进行建模,导致非节拍输出,弱部分过渡和有限的编辑功能。我们提出了MusicWeaver,一个音乐生成模型的节拍对齐的结构计划的条件。该计划作为输入提示和生成的音乐之间的可编辑中间体,保留全局形式并支持专业的本地化编辑。MusicWeaver由一个计划器和一个基于扩散的生成器组成,前者将提示转换为编码音乐形式和作曲线索的结构计划,后者在计划的指导下合成音乐。为了评估生成和编辑质量,我们引入了两个指标:用于评估长期形式和时间的结构一致性得分(SCS)和用于测量实现计划编辑的准确性的编辑保真度得分(EFS)。实验表明,MusicWeaver实现了最先进的保真度和可控性,产生的音乐更接近人类创作的作品。音乐结果可以在我们的项目页面上找到:https://musicweaver.github.io/。
【17】Guiding Audio Editing with Audio Language Model
标题:用音频语言模型指导音频编辑
链接:https://arxiv.org/abs/2509.21625
【18】Preserving Russek's "Summermood" Using Reality Check and a DeltaLab DL-4 Approximation
链接:https://arxiv.org/abs/2509.21560
摘要:作为对正在进行的努力,以保持现场表演的电声成分的贡献,我们提出了一个纯数据补丁的集合,以保存和执行安东尼奥·罗塞克的作品“夏日心情”低音长笛和现场电子。这篇文章最初是为DeltaLab DL-4延迟架单元写的,包含了DL-4特有的分数标记。在这里,我们在Pure Data中近似了DL-4的声音和独特功能,然后通过比较乐谱和两个正式录音的设置来改进我们的实现,以更好地匹配演奏该作品的单元。DL-4仿真被集成到一个补丁中,用于基于Pure Piece的实时性能,并使用Reality Check框架对Pure Data进行回归测试。使用这个补丁库,Summermood可以在不使用现已停产的DL-4的情况下恢复实时旋转。这些补丁将不断进行测试,以确保该作品在计算机环境中可播放,并随着Pure Data编程语言的更新。
【19】Real-time implementation of vibrato transfer as an audio effect
标题:实时实现颤音传输作为音频效果
链接:https://arxiv.org/abs/2509.21544
摘要:最近引入了一种用于基于颤音的真实示例导出延迟函数的算法,并且该算法可以用于执行颤音转移,其中使用延迟线将目标信号的颤音模式赋予到传入声音上。该算法包含计算限制实时实现的方法。在这里,提出了一个实时的近似,它采用了一个有效的基本频率估计算法和时域多相IIR滤波器,近似的分析信号。颤音转移算法进一步补充了一种建议的方法来转移目标声音的幅度调制,移动这种方法超出了典型的基于延迟的颤音效果的能力。这里详细介绍了对原始算法的实时修改,并作为VST插件实现的源代码提供。该算法在声音设计、声音变形和合成声音的实时颤音控制中具有音频效果的应用。
【20】Shortcut Flow Matching for Speech Enhancement: Step-Invariant flows via single stage training
标题:语音增强的MIDI流匹配:通过单阶段训练的步进不变流
链接:https://arxiv.org/abs/2509.21522
摘要:基于扩散的生成模型在语音增强(SE)的感知质量方面取得了最先进的性能。然而,它们的迭代性质需要大量的神经功能评估(NFE),这对实时应用构成了挑战。相反,流匹配提供了一个更有效的替代方案,通过学习一个直接的矢量场,使高质量的合成,在短短几个步骤,使用确定性常微分方程~(ODE)求解器。因此,我们介绍了用于语音增强(SFMSE),一种新的方法,训练一个单一的,步不变的模型的单流匹配。通过在一个阶段的训练过程中对目标时间步长的速度场进行调节,SFMSE可以执行单步、几步或多步去噪,而无需任何架构更改或微调。我们的研究结果表明,单步SFMSE推理在消费级GPU上实现了0.013的实时因子(RTF),同时提供了与需要60个NFE的强扩散基线相当的感知质量。这项工作还提供了随机性在训练和推理中的作用的实证分析,弥合了高质量的生成SE和低延迟约束之间的差距。
【21】Golden Tonnetz
标题:金托尼茨
链接:https://arxiv.org/abs/2509.21428
摘要:音乐概念已经被几何学和音调所代表。例如,在半音圈中,12个音调由圆上的12个点表示,而在Tonnetz中,和声之间的关系由三角形网格表示。最近,我们已经表明,在正二十面体上的几个安排的音调可以与半音音阶,全音音阶,主要的音调,并通过黄金分割率的小调。在这里,我们研究音乐和黄金分割率之间的另一种联系。我们发现,存在着一个安排的7个音调的黄金三角形,可以代表一个给定的大/小调规模和它的主音,属音和下属和弦的黄金三角形。应用这一发现,我们提出了“黄金音调”,它表示所有的大/小调音阶和三和弦的黄金三角形或gnomons,也表示相对的,平行的,并在新黎曼理论的引导音交换转换的黄金三角形和gnomons之间的转换。
【22】Speaker Anonymisation for Speech-based Suicide Risk Detection
标题:基于言语的自杀风险检测的发言人匿名化
链接:https://arxiv.org/abs/2509.22148
【23】Speak Your Mind: The Speech Continuation Task as a Probe of Voice-Based Model Bias
标题:说出你的想法:语音延续任务作为基于语音的模型偏见的探索
链接:https://arxiv.org/abs/2509.22061
摘要:语音接续(SC)的任务是生成一个连贯的延伸的口语提示,同时保持语义上下文和扬声器的身份。因为SC被限制到单个音频流,所以它提供了比对话更直接的设置来探测语音基础模型中的偏差。在这项工作中,我们提出了第一个系统的评估SC的偏见,调查如何性别和发声类型(呼吸,吱吱作响,最终吱吱作响)影响延续行为。我们评估了三个最近的模型:SpiritLM(基础和表达),VAE-GSLM和SpeechGPT跨扬声器相似性,语音质量保持和基于文本的偏见指标。结果表明,虽然说话人的相似性和连贯性仍然是一个挑战,文本评价揭示了显着的模式和性别的相互作用:一旦连贯性足够高(VAE-GSLM),性别效应出现在文本指标,如机构和句子极性。此外,延续转向模态发声更强烈的女性提示比男性的,揭示了系统的语音质量的偏见。这些研究结果突出SC作为一个控制探针的社会相关的代表性偏见的语音基础模型,并建议它将成为一个越来越翔实的诊断连续质量的提高。
【24】AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook
标题:A紫外线:使用单嵌套代码簿教授音频通用载体量化
链接:https://arxiv.org/abs/2509.21968
摘要:我们提出了AUV,一个统一的神经音频编解码器与一个单一的码本,这使得一个有利的重建语音,并进一步扩展到一般的音频,包括声乐,音乐和声音。AUV能够以700 bps左右的比特率处理任何16 kHz混合域音频段。为了实现这一点,我们使用嵌套的特定于域的分区来指导matryoshka码本,分配相应的教师模型来执行蒸馏,所有这些都在单阶段训练中进行。一个符合风格的编码器-解码器架构与STFT功能作为音频表示,产生更好的音频质量。综合评估表明,AUV表现出相当的音频重建能力,以国家的最先进的特定领域的单层量化器编解码器,展示了音频通用矢量量化与一个单一的码本的潜力。预训练模型和演示样本可在https://swivid.github.io/AUV/上获得。
【25】HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
标题:HuLA:具有多任务学习的韵律感知反欺骗,用于表达性和情感合成语音
链接:https://arxiv.org/abs/2509.21676
摘要:目前的反欺骗系统仍然容易受到表达和情感的合成语音,因为他们很少利用韵律作为一个有区别的线索。韵律是人类表达能力和情感的核心,人类本能地使用韵律线索,如F0模式和有声/无声结构来区分自然语音和合成语音。在本文中,我们提出了HuLA,一个两阶段的韵律感知多任务学习框架的欺骗检测。在第1阶段,在真实语音上训练自监督学习(SSL)骨干,并进行F0预测和有声/无声分类的辅助任务,增强其捕获类似于人类感知学习的自然韵律变化的能力。在第2阶段,该模型针对真实和合成数据上的欺骗检测和韵律任务进行了联合优化,利用韵律感知来检测自然和表达性合成语音之间的不匹配。实验表明,HuLA在具有挑战性的域外数据集上始终优于强大的基线,包括表达,情感和跨语言攻击。这些结果表明,显式韵律监督,结合SSL嵌入,大大提高了对先进的合成语音攻击的鲁棒性。
【26】AUDDT: Audio Unified Deepfake Detection Benchmark Toolkit
标题:AUDT:音频统一Deepfake检测基准工具包
链接:https://arxiv.org/abs/2509.21597
【27】Multi-Speaker DOA Estimation in Binaural Hearing Aids using Deep Learning and Speaker Count Fusion
标题:基于深度学习和说话人计数融合的双耳助听器多说话人DOA估计
链接:https://arxiv.org/abs/2509.21382
备注:5 pages, 2 figures, submitted to IEEE ICASSP 2026
摘要:为了提取目标说话人语音,到达方向(DOA)估计对于在嘈杂的多说话人环境中操作的双耳助听器至关重要。在为此任务开发的解决方案中,利用麦克风信号之间的频谱相位差和幅度比的深度学习卷积递归神经网络(CRNN)模型是一个流行的选择。在本文中,我们探讨了添加源计数信息的多源DOA估计。首先考虑使用联合多源DOA估计和源计数的双任务训练。然后,我们考虑使用的源计数作为一个辅助功能,在一个独立的DOA估计系统,其中活跃的源(0,1,或2+)的数量被集成到CRNN架构,通过早期,中期和后期融合策略。使用真正的双耳录音进行实验。结果表明,双任务训练虽然有利于信源计数预测,但并没有提高DOA估计性能。然而,地面实况(oracle)源计数用作辅助功能显着提高独立的DOA估计性能,后期融合产生高达14%的平均F1分数比基线CRNN。这突出了在双耳助听器中使用源计数估计用于鲁棒DOA估计的潜力。
【1】Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
标题:Semantic-VAE:用于更好语音合成的语义对齐潜在表示
链接:https://arxiv.org/abs/2509.22167
【2】Towards Cross-Task Suicide Risk Detection via Speech LLM
标题:通过语音LLM实现跨任务自杀风险检测
链接:https://arxiv.org/abs/2509.22153
【3】Speaker Anonymisation for Speech-based Suicide Risk Detection
标题:基于言语的自杀风险检测的发言人匿名化
链接:https://arxiv.org/abs/2509.22148
【4】Speak Your Mind: The Speech Continuation Task as a Probe of Voice-Based Model Bias
标题:说出你的想法:语音延续任务作为基于语音的模型偏见的探索
链接:https://arxiv.org/abs/2509.22061
摘要:语音接续(SC)的任务是生成一个连贯的延伸的口语提示,同时保持语义上下文和扬声器的身份。因为SC被限制到单个音频流,所以它提供了比对话更直接的设置来探测语音基础模型中的偏差。在这项工作中,我们提出了第一个系统的评估SC的偏见,调查如何性别和发声类型(呼吸,吱吱作响,最终吱吱作响)影响延续行为。我们评估了三个最近的模型:SpiritLM(基础和表达),VAE-GSLM和SpeechGPT跨扬声器相似性,语音质量保持和基于文本的偏见指标。结果表明,虽然说话人的相似性和连贯性仍然是一个挑战,文本评价揭示了显着的模式和性别的相互作用:一旦连贯性足够高(VAE-GSLM),性别效应出现在文本指标,如机构和句子极性。此外,延续转向模态发声更强烈的女性提示比男性的,揭示了系统的语音质量的偏见。这些研究结果突出SC作为一个控制探针的社会相关的代表性偏见的语音基础模型,并建议它将成为一个越来越翔实的诊断连续质量的提高。
【5】AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook
标题:A紫外线:使用单嵌套代码簿教授音频通用载体量化
链接:https://arxiv.org/abs/2509.21968
摘要:我们提出了AUV,一个统一的神经音频编解码器与一个单一的码本,这使得一个有利的重建语音,并进一步扩展到一般的音频,包括声乐,音乐和声音。AUV能够以700 bps左右的比特率处理任何16 kHz混合域音频段。为了实现这一点,我们使用嵌套的特定于域的分区来指导matryoshka码本,分配相应的教师模型来执行蒸馏,所有这些都在单阶段训练中进行。一个符合风格的编码器-解码器架构与STFT功能作为音频表示,产生更好的音频质量。综合评估表明,AUV表现出相当的音频重建能力,以国家的最先进的特定领域的单层量化器编解码器,展示了音频通用矢量量化与一个单一的码本的潜力。预训练模型和演示样本可在https://swivid.github.io/AUV/上获得。
【6】A Parallel Ultra-Low Power Silent Speech Interface based on a Wearable, Fully-dry EMG Neckband
标题:基于可穿戴、全干EMG项圈的并行超低功耗无声语音接口
链接:https://arxiv.org/abs/2509.21964
摘要:我们提出了一个可穿戴的,全干燥的,超低功耗的EMG系统,用于无声语音识别,集成到一个纺织品颈带,以实现舒适,非侵入性的使用。该系统具有14个全差分EMG通道,基于BioGAP-Ultra平台,可实现超低功耗(22 mW)生物信号采集和无线传输。我们评估其性能的8个语音命令下发声和无声的清晰度,实现平均分类准确率分别为87$\pm$3%和68$\pm$3%,与5倍CV的方法。为了模拟日常生活条件,我们引入会话到会话的变化,通过重新定位会话之间的颈带,实现了64$\pm$18%和54$\pm$7%的发音和无声实验,分别留一个会话的准确性。这些结果突出了所提出的方法的鲁棒性和节能无声语音解码的承诺。
【7】IPDnet2: an efficient and improved inter-channel phase difference estimation network for sound source localization
标题:IPDnet 2:一种高效且改进的通道间相差估计网络,用于声音源定位
链接:https://arxiv.org/abs/2509.21900
【8】FastEnhancer: Speed-Optimized Streaming Neural Speech Enhancement
标题:FastEnhancer:速度优化的流媒体神经语音增强
链接:https://arxiv.org/abs/2509.21867
【9】HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
标题:HuLA:具有多任务学习的韵律感知反欺骗,用于表达性和情感合成语音
链接:https://arxiv.org/abs/2509.21676
摘要:目前的反欺骗系统仍然容易受到表达和情感的合成语音,因为他们很少利用韵律作为一个有区别的线索。韵律是人类表达能力和情感的核心,人类本能地使用韵律线索,如F0模式和有声/无声结构来区分自然语音和合成语音。在本文中,我们提出了HuLA,一个两阶段的韵律感知多任务学习框架的欺骗检测。在第1阶段,在真实语音上训练自监督学习(SSL)骨干,并进行F0预测和有声/无声分类的辅助任务,增强其捕获类似于人类感知学习的自然韵律变化的能力。在第2阶段,该模型针对真实和合成数据上的欺骗检测和韵律任务进行了联合优化,利用韵律感知来检测自然和表达性合成语音之间的不匹配。实验表明,HuLA在具有挑战性的域外数据集上始终优于强大的基线,包括表达,情感和跨语言攻击。这些结果表明,显式韵律监督,结合SSL嵌入,大大提高了对先进的合成语音攻击的鲁棒性。
【10】AUDDT: Audio Unified Deepfake Detection Benchmark Toolkit
标题:AUDT:音频统一Deepfake检测基准工具包
链接:https://arxiv.org/abs/2509.21597
【11】Enhanced Generative Machine Listener
标题:增强的生成机器收件箱
链接:https://arxiv.org/abs/2509.21463
【12】ARTI-6: Towards Six-dimensional Articulatory Speech Encoding
标题:ARTI-6:迈向六维发音语音编码
链接:https://arxiv.org/abs/2509.21447
【13】Multi-Speaker DOA Estimation in Binaural Hearing Aids using Deep Learning and Speaker Count Fusion
标题:基于深度学习和说话人计数融合的双耳助听器多说话人DOA估计
链接:https://arxiv.org/abs/2509.21382
摘要:为了提取目标说话人语音,到达方向(DOA)估计对于在嘈杂的多说话人环境中操作的双耳助听器至关重要。在为此任务开发的解决方案中,利用麦克风信号之间的频谱相位差和幅度比的深度学习卷积递归神经网络(CRNN)模型是一个流行的选择。在本文中,我们探讨了添加源计数信息的多源DOA估计。首先考虑使用联合多源DOA估计和源计数的双任务训练。然后,我们考虑使用源计数作为独立波达方向估计系统中的辅助特征,其中活动源的数量(0、1或2+)通过早期、中期和后期融合策略集成到CRNN架构中。使用真正的双耳录音进行实验。结果表明,双任务训练虽然有利于信源计数预测,但并没有提高DOA估计性能。然而,地面实况(oracle)源计数用作辅助功能显着提高独立的DOA估计性能,后期融合产生高达14%的平均F1分数比基线CRNN。这突出了在双耳助听器中使用源计数估计用于鲁棒DOA估计的潜力。
【14】Toward a Realistic Encoding Model of Auditory Affective Understanding in the Brain
标题:建立大脑听觉情感理解的现实编码模型
链接:https://arxiv.org/abs/2509.21381
【15】MDAR: A Multi-scene Dynamic Audio Reasoning Benchmark
标题:MDAR:多场景动态音频推理基准
链接:https://arxiv.org/abs/2509.22461
摘要:从音频(包括语音、非语言提示、环境声音和音乐)中推理的能力对于AI代理在现实世界中有效交互至关重要。现有的基准测试主要集中在静态或单场景设置上,并且不能完全捕获多个扬声器、展开事件和异构音频源交互的场景。为了应对这些挑战,我们引入了MDAR,这是一个用于评估复杂,多场景和动态演变的音频推理任务模型的基准。MDAR包含3,000个精心策划的问答对,与不同的音频片段相关联,涵盖五类复杂推理和三种问题类型。我们对MDAR上的26种最先进的音频语言模型进行了基准测试,并观察到它们在复杂的推理任务中表现出局限性。在单项选择题上,Qwen2.5-Omni(开源)达到了76.67%的准确率,而GPT-4 o Audio(闭源)达到了68.47%;然而,GPT-4 o Audio在更具挑战性的多项选择和开放式任务上大大优于Qwen2.5-Omni。在所有三种问题类型中,没有模型达到80%的性能。这些发现强调了MDAR带来的独特挑战及其作为推进音频推理研究的基准的价值。代码和基准可以在https://github.com/luckyerr/MDAR上找到。
【16】Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach
标题:毫不费力的图像到音乐生成:一种可解释的基于RAG的VLM方法
链接:https://arxiv.org/abs/2509.22378
【17】Investigating Faithfulness in Large Audio Language Models
标题:研究大型音频语言模型中的忠实性
链接:https://arxiv.org/abs/2509.22363
【18】Cross-Dialect Bird Species Recognition with Dialect-Calibrated Augmentation
标题:使用方言校准增强的跨方言鸟类物种识别
链接:https://arxiv.org/abs/2509.22317
【19】Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
标题:理解和交谈:通过双语言建模进行文本到语音合成
链接:https://arxiv.org/abs/2509.22062
摘要:现有的基于大语言模型(LLM)的自回归(AR)文本到语音(TTS)系统虽然达到了最先进的质量,但仍然面临着严峻的挑战。这种基于LLM的范例的基础是通过神经音频编解码器将连续语音波形离散化为离散令牌序列。然而,单码本建模非常适合于文本LLM,但遭受显著的信息丢失;通常通过残差矢量量化(RVQ)生成的分层声学令牌通常缺乏明确的语义结构,从而给模型带来沉重的学习负担。此外,自回归过程固有地易受误差累积的影响,这会降低生成稳定性。为了解决这些局限性,我们提出了CaT-TTS,一个新的框架,强大的和语义接地zero-shot合成。首先,我们介绍了S3 Codec,一个分裂的RVQ编解码器,通过从最先进的ASR模型中提取语义,将显式语言特征注入其主码本,提供了一个结构化的表示,简化了学习任务。其次,我们提出了一个“理解然后生成”的双变压器架构,从渲染的理解。初始“理解”Transformer对文本和音频的语义标记之间的跨模态关系进行建模,以形成高级话语计划。随后的“生成”Transformer执行该计划,自回归合成分层声学标记。最后,为了提高生成稳定性,我们引入了掩蔽音频并行推理(Masked Audio Parallel Inference,简称MFS),这是一种几乎无参数的推理策略,可以动态地指导解码过程以减轻局部错误。
【20】Text2Move: Text-to-moving sound generation via trajectory prediction and temporal alignment
标题:文本2Move:通过轨迹预测和时间对齐生成文本到运动的声音
链接:https://arxiv.org/abs/2509.21919
【21】Noise-to-Notes: Diffusion-based Generation and Refinement for Automatic Drum Transcription
标题:噪音笔记:基于扩散的自动鼓转录生成和细化
链接:https://arxiv.org/abs/2509.21739
【22】Align2Speak: Improving TTS for Low Resource Languages via ASR-Guided Online Preference Optimization
标题:Alignn 2Speak:通过ASB引导的在线偏好优化改进低资源语言的TTC
链接:https://arxiv.org/abs/2509.21718
摘要:由于成对文本和语音数据的稀缺性,为低资源语言开发高质量的文本到语音(TTS)系统具有挑战性。相比之下,由于大规模的多语言预训练工作,这些语言的自动语音识别(ASR)模型通常更容易获得。我们提出了一个框架的基础上组相对策略优化(GRPO),以适应自回归,多语言的TTS模型,以新的语言。我们的方法首先建立了一个语言不可知的TTS合成的基础,通过训练多语言的基线与国际音标(IPA)令牌。接下来,我们在新语言的有限配对数据上对该模型进行微调,以捕获目标语言的韵律特征。最后,我们应用GRPO仅使用未配对的文本和扬声器提示来优化模型,并由来自预训练ASR,扬声器验证和音频质量估计模型的多目标奖励指导。实验表明,该管道在低资源语言中产生可理解和说话者一致的语音,大大优于单独的微调。此外,我们基于GRPO的框架还提高了高资源语言的TTS性能,超越了直接偏好优化(DPO)等离线对齐方法,从而获得了卓越的可懂度,扬声器相似性和音频质量。
【23】Guiding Audio Editing with Audio Language Model
标题:用音频语言模型指导音频编辑
链接:https://arxiv.org/abs/2509.21625
【24】Real-time implementation of vibrato transfer as an audio effect
标题:实时实现颤音传输作为音频效果
链接:https://arxiv.org/abs/2509.21544
摘要:最近引入了一种用于基于颤音的真实示例导出延迟函数的算法,并且该算法可以用于执行颤音转移,其中使用延迟线将目标信号的颤音模式赋予到传入声音上。该算法包含计算限制实时实现的方法。在这里,提出了一个实时的近似,它采用了一个有效的基本频率估计算法和时域多相IIR滤波器,近似的分析信号。颤音转移算法进一步补充了一种建议的方法来转移目标声音的幅度调制,移动这种方法超出了典型的基于延迟的颤音效果的能力。这里详细介绍了对原始算法的实时修改,并作为VST插件实现的源代码提供。该算法在声音设计、声音变形和合成声音的实时颤音控制中具有音频效果的应用。
【25】Golden Tonnetz
标题:金托尼茨
链接:https://arxiv.org/abs/2509.21428
摘要:音乐概念已经被几何学和音调所代表。例如,在半音圈中,12个音调由圆上的12个点表示,而在Tonnetz中,和声之间的关系由三角形网格表示。最近,我们已经表明,在正二十面体上的几个安排的音调可以与半音音阶,全音音阶,主要的音调,并通过黄金分割率的小调。在这里,我们研究音乐和黄金分割率之间的另一种联系。我们发现,存在着一个安排的7个音调的黄金三角形,可以代表一个给定的大/小调规模和它的主音,属音和下属和弦的黄金三角形。应用这一发现,我们提出了“黄金音调”,它表示所有的大/小调音阶和三和弦的黄金三角形或gnomons,也表示相对的,平行的,并在新黎曼理论的引导音交换转换的黄金三角形和gnomons之间的转换。
机器翻译由腾讯交互翻译提供,仅供参考
