今日论文合集:cs.SD语音12篇,eess.AS音频处理13篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Leveraging Whisper Embeddings for Audio-based Lyrics Matching
标题:利用Whisper嵌入进行音频歌词匹配
链接:https://arxiv.org/abs/2510.08176

作者:Eleonora Mancini, Joan Serrà, Paolo Torroni, Yuki Mitsufuji
摘要:基于音频的歌词匹配可以是其他基于内容的检索方法的一个有吸引力的替代方案,但现有的方法往往受到有限的再现性和不一致的基线。在这项工作中,我们介绍了WEALY,一个完全可再现的管道,利用Whisper解码器嵌入歌词匹配任务。WEALY建立了强大而透明的基线,同时还探索了整合文本和声学特征的多模态扩展。通过对标准数据集的广泛实验,我们证明WEALY实现了与缺乏可重复性的最先进方法相当的性能。此外,我们还提供了关于语言鲁棒性、损失函数和嵌入策略的消融研究和分析。这项工作为未来的研究提供了一个可靠的基准,并强调了语音技术在音乐信息检索任务中的潜力。
摘要:Audio-based lyrics matching can be an appealing alternative to other content-based retrieval approaches, but existing methods often suffer from limited reproducibility and inconsistent baselines. In this work, we introduce WEALY, a fully reproducible pipeline that leverages Whisper decoder embeddings for lyrics matching tasks. WEALY establishes robust and transparent baselines, while also exploring multimodal extensions that integrate textual and acoustic features. Through extensive experiments on standard datasets, we demonstrate that WEALY achieves a performance comparable to state-of-the-art methods that lack reproducibility. In addition, we provide ablation studies and analyses on language robustness, loss functions, and embedding strategies. This work contributes a reliable benchmark for future research, and underscores the potential of speech technologies for music information retrieval tasks.


【2】Detecting and Mitigating Insertion Hallucination in Video-to-Audio Generation
标题:检测和缓解视频转音频生成中的插入幻觉
链接:https://arxiv.org/abs/2510.08078

作者:Liyang Chen, Hongkai Chen, Yujun Cai, Sifan Li, Qingwen Ye, Yiwei Wang
摘要:视频到音频生成在自动合成视频声音方面取得了显著的进步。然而,现有的评估指标,侧重于语义和时间对齐,忽略了一个关键的故障模式:模型往往产生的声学事件,特别是语音和音乐,没有相应的视觉源。我们将这种现象称为插入幻觉,并将其确定为由数据集偏差驱动的系统性风险,例如屏幕外声音的普遍存在,而当前的指标仍然完全未检测到。为了解决这一挑战,我们首先开发了一个系统的评估框架,采用多个音频事件检测器的多数投票合奏。我们还引入了两个新的指标来量化这个问题的患病率和严重程度:IH@vid(具有幻觉的视频的比例)和IH@dur(幻觉持续时间的比例)。在此基础上,我们提出了后验特征校正,这是一种新型的免训练推理时间方法,可以减轻IH。PFC在两个过程中运行:它首先生成初始音频输出以检测幻觉片段,然后在掩蔽这些时间戳的相应视频特征后重新生成音频。在几个主流V2 A基准测试上的实验首先揭示了最先进的模型遭受严重的IH。相比之下,我们的PFC方法平均将幻觉的患病率和持续时间降低了50%以上,而不会降低,并且在某些情况下甚至改善了音频质量和时间同步的传统指标。我们的工作是第一个正式定义、系统测量和有效缓解插入幻觉的工作,为更可靠、更忠实的V2 A模型铺平了道路。
摘要:Video-to-Audio generation has made remarkable strides in automatically synthesizing sound for video. However, existing evaluation metrics, which focus on semantic and temporal alignment, overlook a critical failure mode: models often generate acoustic events, particularly speech and music, that have no corresponding visual source. We term this phenomenon Insertion Hallucination and identify it as a systemic risk driven by dataset biases, such as the prevalence of off-screen sounds, that remains completely undetected by current metrics. To address this challenge, we first develop a systematic evaluation framework that employs a majority-voting ensemble of multiple audio event detectors. We also introduce two novel metrics to quantify the prevalence and severity of this issue: IH@vid (the fraction of videos with hallucinations) and IH@dur (the fraction of hallucinated duration). Building on this, we propose Posterior Feature Correction, a novel training-free inference-time method that mitigates IH. PFC operates in a two-pass process: it first generates an initial audio output to detect hallucinated segments, and then regenerates the audio after masking the corresponding video features at those timestamps. Experiments on several mainstream V2A benchmarks first reveal that state-of-the-art models suffer from severe IH. In contrast, our PFC method reduces both the prevalence and duration of hallucinations by over 50\% on average, without degrading, and in some cases even improving, conventional metrics for audio quality and temporal synchronization. Our work is the first to formally define, systematically measure, and effectively mitigate Insertion Hallucination, paving the way for more reliable and faithful V2A models.


【3】Attribution-by-design: Ensuring Inference-Time Provenance in Generative Music Systems
标题:设计归因:确保生成音乐系统中的推理时间来源
链接:https://arxiv.org/abs/2510.08062

作者:Fabio Morreale, Wiebke Hutiri, Joan Serrà, Alice Xiang, Yuki Mitsufuji
摘要:人工智能音乐的兴起正在稀释版税池,并揭示了现有薪酬框架的结构性缺陷,挑战了音乐产业中完善的艺术家薪酬体系。现有的补偿解决方案,如零碎的许可协议,缺乏可扩展性和技术严谨性,而目前的数据归属机制只提供不确定的估计,很少在实践中实施。本文介绍了一个生成音乐基础设施的框架,该框架以直接归属、透明的版税分配以及对艺术家和版权持有人的粒度控制为中心。我们区分训练集和推理集,这使我们能够提出两种互补形式的归因:训练时归因和推理时归因。我们在这里赞成推理时间归因,因为它使直接的,可验证的补偿,每当艺术家的目录是用来条件生成的输出。此外,用户还可以根据特定的歌曲来调节世代,并获得有关归属和允许使用的透明信息。我们的方法提供了一个道德和实用的解决方案,以满足人工智能生成音乐时代对强大补偿机制的迫切需求,确保出处和公平性嵌入生成系统的核心。
摘要:The rise of AI-generated music is diluting royalty pools and revealing structural flaws in existing remuneration frameworks, challenging the well-established artist compensation systems in the music industry. Existing compensation solutions, such as piecemeal licensing agreements, lack scalability and technical rigour, while current data attribution mechanisms provide only uncertain estimates and are rarely implemented in practice. This paper introduces a framework for a generative music infrastructure centred on direct attribution, transparent royalty distribution, and granular control for artists and rights' holders. We distinguish ontologically between the training set and the inference set, which allows us to propose two complementary forms of attribution: training-time attribution and inference-time attribution. We here favour inference-time attribution, as it enables direct, verifiable compensation whenever an artist's catalogue is used to condition a generated output. Besides, users benefit from the ability to condition generations on specific songs and receive transparent information about attribution and permitted usage. Our approach offers an ethical and practical solution to the pressing need for robust compensation mechanisms in the era of AI-generated music, ensuring that provenance and fairness are embedded at the core of generative systems.


【4】Personality-Enhanced Multimodal Depression Detection in the Elderly
标题:老年人的个性增强多模式抑郁检测
链接:https://arxiv.org/abs/2510.08004

作者:Honghong Wang, Jing Deng, Rong Zheng
备注:6 pages,2 figures,accepted by ACM Multimedia Asia 2025
摘要:本文介绍了我们在ACM MM 2025上对多模态个性感知抑郁检测(MPDD)挑战的解决方案。我们提出了一个多模态抑郁症检测模型,在老年人,包括人格特征。我们引入了一种基于共同注意力机制的多特征融合方法,以有效地将LLDs,MFCC和Wav2Vec特征集成在音频模态中。对于视频模态,我们结合了从OpenFace,ResNet和DenseNet中提取的表示来构建一个全面的视觉特征集。认识到人格在抑郁症检测中的关键作用,我们设计了一个互动模块,捕捉人格特质和多模态特征之间的关系。来自MPDD老年抑郁症检测跟踪的实验结果表明,我们的方法显着提高了性能,为老年人群中多模态抑郁症检测的未来研究提供了有价值的见解。
摘要:This paper presents our solution to the Multimodal Personality-aware Depression Detection (MPDD) challenge at ACM MM 2025. We propose a multimodal depression detection model in the Elderly that incorporates personality characteristics. We introduce a multi-feature fusion approach based on a co-attention mechanism to effectively integrate LLDs, MFCCs, and Wav2Vec features in the audio modality. For the video modality, we combine representations extracted from OpenFace, ResNet, and DenseNet to construct a comprehensive visual feature set. Recognizing the critical role of personality in depression detection, we design an interaction module that captures the relationships between personality traits and multimodal features. Experimental results from the MPDD Elderly Depression Detection track demonstrate that our method significantly enhances performance, providing valuable insights for future research in multimodal depression detection among elderly populations.


【5】IntMeanFlow: Few-step Speech Generation with Integral Velocity Distillation
标题:IntMeanFlow:利用积分速度蒸馏的几步语音生成
链接:https://arxiv.org/abs/2510.07979

作者:Wei Wang, Rong Cao, Yi Guo, Zhengyang Chen, Kuan Chen, Yuanyuan Huo
摘要:基于流的生成模型极大地提高了文语转换(TTS)合成质量,但推理速度仍然受到迭代采样过程和多函数求值(NFE)的限制。最近的MeanFlow模型通过模拟平均速度而不是瞬时速度来加速生成。然而,其直接应用于TTS遇到了挑战,包括雅可比向量积(JVP)的GPU内存开销和自引导过程导致的训练不稳定性。为了解决这些问题,我们引入了IntMeanFlow,一个具有积分速度蒸馏的几步语音生成框架。通过在时间间隔内用教师的瞬时速度近似平均速度,IntMeanFlow消除了对JVP和自引导的需要,提高了稳定性并减少了GPU内存使用。我们还提出了最优步长采样搜索(O3 S)算法,该算法确定了特定于模型的最优采样步长,在不增加额外推理开销的情况下提高了语音合成。实验表明,IntMeanFlow实现了1-NFE推理的token-to-spectrogram和3-NFE的文本到spectrogram的任务,同时保持高质量的合成。演示示例可在https://vvwangvv.github.io/intmeanflow上获得。
摘要:Flow-based generative models have greatly improved text-to-speech (TTS) synthesis quality, but inference speed remains limited by the iterative sampling process and multiple function evaluations (NFE). The recent MeanFlow model accelerates generation by modeling average velocity instead of instantaneous velocity. However, its direct application to TTS encounters challenges, including GPU memory overhead from Jacobian-vector products (JVP) and training instability due to self-bootstrap processes. To address these issues, we introduce IntMeanFlow, a framework for few-step speech generation with integral velocity distillation. By approximating average velocity with the teacher's instantaneous velocity over a temporal interval, IntMeanFlow eliminates the need for JVPs and self-bootstrap, improving stability and reducing GPU memory usage. We also propose the Optimal Step Sampling Search (O3S) algorithm, which identifies the model-specific optimal sampling steps, improving speech synthesis without additional inference overhead. Experiments show that IntMeanFlow achieves 1-NFE inference for token-to-spectrogram and 3-NFE for text-to-spectrogram tasks while maintaining high-quality synthesis. Demo samples are available at https://vvwangvv.github.io/intmeanflow.


【6】ACMID: Automatic Curation of Musical Instrument Dataset for 7-Stem Music Source Separation
标题:ACMID:用于7-Stem音乐源分离的乐器数据集的自动处理
链接:https://arxiv.org/abs/2510.07840

作者:Ji Yu, Yang shuo, Xu Yuetonghui, Liu Mengmei, Ji Qiang, Han Zerui
摘要:目前的音乐源分离(MSS)方法大多依赖于监督学习,受到训练数据数量和质量的限制。虽然网络爬虫可以带来丰富的数据,但平台级的音轨标注往往会导致元数据的不匹配,从而影响“音频-标签”对的准确获取。为了解决这个问题,我们提出了ACMID:通过对大量原始数据进行网络抓取生成的MSS数据集,然后通过构建在预先训练的音频编码器上的乐器分类器进行自动清理,该编码器从抓取的曲目中过滤和聚合目标乐器的干净片段,从而产生经过改进的ACMID清理数据集。利用丰富的数据,我们将传统的4-干(声乐/低音/鼓/其他)扩展到7-干(钢琴/鼓/低音/原声吉他/电吉他/弦乐器/管乐器),从而实现高粒度MSS系统。在SOTA MSS模型上的实验表明了两个关键结果:(i)使用ACMID-Cleaned训练的MSS模型与使用ACMID-Uncleaned相比,SDR性能提高了2.39dB,证明了我们的数据清洗过程的有效性;(ii)将ACMID-Cleaned纳入训练,使MSS模型的平均性能提高了1.16dB,证实了我们数据集的价值。我们的数据抓取代码,清洁模型代码和权重可在:https://github.com/scottishfold0621/ACMID.
摘要:Most current music source separation (MSS) methods rely on supervised learning, limited by training data quan- tity and quality. Though web-crawling can bring abundant data, platform-level track labeling often causes metadata mismatches, impeding accurate "audio-label" pair acquisi- tion. To address this, we present ACMID: a dataset for MSS generated through web crawling of extensive raw data, fol- lowed by automatic cleaning via an instrument classifier built on a pre-trained audio encoder that filters and aggregates clean segments of target instruments from the crawled tracks, resulting in the refined ACMID-Cleaned dataset. Leverag- ing abundant data, we expand the conventional classifica- tion from 4-stem (Vocal/Bass/Drums/Others) to 7-stem (Pi- ano/Drums/Bass/Acoustic Guitar/Electric Guitar/Strings/Wind- Brass), enabling high granularity MSS systems. Experiments on SOTA MSS model demonstrates two key results: (i) MSS model trained with ACMID-Cleaned achieved a 2.39dB improvement in SDR performance compared to that with ACMID-Uncleaned, demostrating the effectiveness of our data cleaning procedure; (ii) incorporating ACMID-Cleaned to training enhances MSS model's average performance by 1.16dB, confirming the value of our dataset. Our data crawl- ing code, cleaning model code and weights are available at: https://github.com/scottishfold0621/ACMID.


【7】IsoSignVid2Aud: Sign Language Video to Audio Conversion without Text Intermediaries
标题:IsoSignVid2Aud:无需文本中介的手语视频到音频转换
链接:https://arxiv.org/abs/2510.07837

作者:Harsh Kavediya, Vighnesh Nayak, Bheeshm Sharma, Balamurugan Palaniappan
备注:Accepted in AIML-Systems-2025
摘要:手语到口语的音频翻译对于将听力和语言障碍的人与他人联系起来非常重要。我们认为手语视频与孤立的符号序列,而不是连续的语法签名。这样的视频在教育应用和签名提示界面中是有用的。为此,我们提出了IsoSignVid 2Aud,这是一种新型的端到端框架,可以将手语视频与可能的非语法连续符号序列转换为语音,而不需要中间文本表示,提供即时的通信优势,同时避免多阶段翻译系统中固有的延迟和级联错误。我们的方法结合了一个基于I3 D的特征提取模块与一个专门的特征变换网络和音频生成管道,利用一种新的非最大抑制(NMS)算法的时间检测的非语法连续序列中的符号。实验结果表明,在ASL-Citizen-1500和WLASL-100数据集上的性能具有竞争力,Top-1准确率分别为72.01\%和78.67\%,音频质量指标(PESQ:2.67,STOI:0.73)表明语音输出清晰。代码可从以下网址获得:https://github.com/BheeshmSharma/IsoSignVid2Aud_AIMLsystems-2025。
摘要:Sign language to spoken language audio translation is important to connect the hearing- and speech-challenged humans with others. We consider sign language videos with isolated sign sequences rather than continuous grammatical signing. Such videos are useful in educational applications and sign prompt interfaces. Towards this, we propose IsoSignVid2Aud, a novel end-to-end framework that translates sign language videos with a sequence of possibly non-grammatic continuous signs to speech without requiring intermediate text representation, providing immediate communication benefits while avoiding the latency and cascading errors inherent in multi-stage translation systems. Our approach combines an I3D-based feature extraction module with a specialized feature transformation network and an audio generation pipeline, utilizing a novel Non-Maximal Suppression (NMS) algorithm for the temporal detection of signs in non-grammatic continuous sequences. Experimental results demonstrate competitive performance on ASL-Citizen-1500 and WLASL-100 datasets with Top-1 accuracies of 72.01\% and 78.67\%, respectively, and audio quality metrics (PESQ: 2.67, STOI: 0.73) indicating intelligible speech output. Code is available at: https://github.com/BheeshmSharma/IsoSignVid2Aud_AIMLsystems-2025.


【8】INFER : Learning Implicit Neural Frequency Response Fields for Confined Car Cabin
标题:INBER:学习密闭车厢的隐式神经频率响应场
链接:https://arxiv.org/abs/2510.07442

作者:Harshvardhan C. Takawale, Nirupam Roy, Phil Brown
摘要:空间声学的精确建模对于在密闭的共振环境(如汽车舱)中实现沉浸式和可理解的音频至关重要。目前的调谐方法是手动的、硬件密集型的和静态的,不能考虑频率选择行为和动态变化,如乘客存在或座椅调整。为了解决这个问题,我们提出了INFER:隐式神经频率响应场,这是一个频域神经框架,它联合取决于源和接收器的位置,方向,以直接学习封闭的谐振环境(如汽车舱)内的复值频率响应场。我们介绍了当前神经声学建模方法的三个关键创新:(1)新的端到端频域前向模型,直接学习3D空间中的频率响应场和频率特定衰减;(2)感知和硬件感知的频谱监督,强调关键的听觉频带,并淡化不稳定的交叉区域;以及(3)基于物理的Kramers-Kronig一致性约束,其正则化频率相关衰减和延迟。我们评估我们的方法在多个汽车舱收集的真实数据。我们的方法在模拟和真实世界的汽车数据集上的性能均显著优于时域和混合域基线,将平均幅度和相位重建误差分别降低了39%和51%以上。INFER为汽车空间的神经声学建模提供了新的最先进的技术
摘要:Accurate modeling of spatial acoustics is critical for immersive and intelligible audio in confined, resonant environments such as car cabins. Current tuning methods are manual, hardware-intensive, and static, failing to account for frequency selective behaviors and dynamic changes like passenger presence or seat adjustments. To address this issue, we propose INFER: Implicit Neural Frequency Response fields, a frequency-domain neural framework that is jointly conditioned on source and receiver positions, orientations to directly learn complex-valued frequency response fields inside confined, resonant environments like car cabins. We introduce three key innovations over current neural acoustic modeling methods: (1) novel end-to-end frequency-domain forward model that directly learns the frequency response field and frequency-specific attenuation in 3D space; (2) perceptual and hardware-aware spectral supervision that emphasizes critical auditory frequency bands and deemphasizes unstable crossover regions; and (3) a physics-based Kramers-Kronig consistency constraint that regularizes frequency-dependent attenuation and delay. We evaluate our method over real-world data collected in multiple car cabins. Our approach significantly outperforms time- and hybrid-domain baselines on both simulated and real-world automotive datasets, cutting average magnitude and phase reconstruction errors by over 39% and 51%, respectively. INFER sets a new state-of-the-art for neural acoustic modeling in automotive spaces


【9】AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
标题:AV-EMO-Reasoning:在具有视听线索的全模式LLMS中对情感推理能力进行基准测试
链接:https://arxiv.org/abs/2510.07355

作者:Krish Patel, Dingkun Zhou, Ajay Kankipati, Akshaj Gupta, Zeyi Austin Li, Mohul Shukla, Vibhor Narang, Sara Kofman, Zongli Ye, Grace Wang, Xiaoyu Shi, Tingle Li, Guan-Ting Lin, Kan Jen Cheng, Huang-Cheng Chou, Jiachen Lian, Gopala Anumanchipalli
摘要:通过语音和面部传达的情感塑造了人机交互中的参与和上下文。尽管全模态大型语言模型(LLM)取得了快速进展,但利用视听线索对情感推理的整体评估仍然有限。为了解决这一差距,我们引入了AV-EMO-Reasoning,这是一个旨在系统评估LLM情感一致性的基准。该框架利用了一个策划,单轮和多轮合成视听语料库与现实世界的一套,并根据连续,分类和感知指标进行评估。与领先的LLM的实验表明,视觉线索可靠地提高了情感的一致性超过音频的基线。此外,LLM可以利用视听线索来生成更多的情感感知语音。模型表现出互补的优势,在度量的家庭,表明自动分数捕获方面不同的感知判断。通过发布系统的评估基准,AV-EMO-Reasoning为评估情感感知对话提供了可重复的标准,并朝着更自然、自适应的人机交互方向迈进。
摘要:Emotions conveyed through voice and face shape engagement and context in human-AI interaction. Despite rapid progress in omni-modal large language models (LLMs), the holistic evaluation of emotional reasoning with audiovisual cues remains limited. To address this gap, we introduce AV-EMO-Reasoning, a benchmark designed to systematically assess emotional coherence in LLMs. The framework leverages a curated, single- and multi-turn synthetic audiovisual corpus with a real-world set and is assessed under continuous, categorical, and perceptual metrics. Experiments with leading LLMs show that visual cues reliably improve emotional coherence over audio-only baselines. Moreover, LLMs can leverage audio-visual cues to generate more emotion-aware speech. Models exhibit complementary strengths across metric families, indicating that automatic scores capture facets distinct from perceptual judgments. By releasing a systematic evaluation benchmark, AV-EMO-Reasoning offers a reproducible standard for evaluating emotion-aware dialogue and advances toward more natural, adaptive human-AI interaction.


【10】Audio-Visual Separation with Hierarchical Fusion and Representation Alignment
标题:采用分层融合和表示对齐的视听分离
链接:https://arxiv.org/abs/2510.07326

作者:Han Hu, Dongheng Lin, Qiming Huang, Yuqi Hou, Hyung Jin Chang, Jianbo Jiao
摘要:自监督视听源分离利用音频和视觉模态之间的自然相关性来分离混合音频信号。在这项工作中,我们首先系统地分析了现有的多模态融合方法的视听分离任务的性能,表明不同的融合策略的性能是密切相关的声音的特性:中间融合更适合于处理短,瞬态的声音,而后期融合更有效地捕捉持续和谐波丰富的声音。因此,我们提出了一个分层融合策略,有效地整合了两个融合阶段。此外,通过合并高质量的外部音频表示,而不是仅仅依靠音频分支来独立学习它们,可以使训练变得更容易。为了探索这一点,我们提出了一种表示对齐方法,该方法将音频编码器的潜在特征与从预训练的音频模型中提取的嵌入对齐。在MUSIC、MUSIC-21和VGGSound数据集上进行的大量实验表明,我们的方法取得了最先进的结果,超越了自监督环境下的现有方法。我们进一步分析了表示对齐对音频特征的影响,表明它减少了音频和视觉模态之间的模态差距。
摘要:Self-supervised audio-visual source separation leverages natural correlations between audio and vision modalities to separate mixed audio signals. In this work, we first systematically analyse the performance of existing multimodal fusion methods for audio-visual separation task, demonstrating that the performance of different fusion strategies is closely linked to the characteristics of the sound: middle fusion is better suited for handling short, transient sounds, while late fusion is more effective for capturing sustained and harmonically rich sounds. We thus propose a hierarchical fusion strategy that effectively integrates both fusion stages. In addition, training can be made easier by incorporating high-quality external audio representations, rather than relying solely on the audio branch to learn them independently. To explore this, we propose a representation alignment approach that aligns the latent features of the audio encoder with embeddings extracted from pre-trained audio models. Extensive experiments on MUSIC, MUSIC-21 and VGGSound datasets demonstrate that our approach achieves state-of-the-art results, surpassing existing methods under the self-supervised setting. We further analyse the impact of representation alignment on audio features, showing that it reduces modality gap between the audio and visual modalities.


【11】MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
标题:MeanVC:通过Mean Flow实现轻量级和流媒体Zero-Shot语音转换
链接:https://arxiv.org/abs/2510.08392

作者:Guobin Ma, Jixun Yao, Ziqian Ning, Yuepeng Jiang, Lingxin Xiong, Lei Xie, Pengcheng Zhu
摘要:Zero-shot语音转换(VC)旨在将音色从源说话人转移到任何看不见的目标说话人,同时保留语言内容。不断增长的应用场景需要具有流式推理功能的模型。这就迫切需要同时快速、轻量和高保真的模型。然而,现有的流式传输方法通常依赖于自回归(AR)或非自回归(NAR)框架,其要么需要大的参数大小来实现强大的性能,要么难以推广到看不见的扬声器。在这项研究中,我们提出了MeanVC,一个轻量级的流zero-shot VC方法。MeanVC引入了一个具有块自回归去噪策略的扩散Transformer,结合了AR和NAR范例的优势,以实现高效的流处理。通过引入平均流,MeanVC在训练过程中回归平均速度场,通过直接从流轨迹的起点映射到终点,使zero-shot VC在单个采样步骤中具有优异的语音质量和说话人相似性。此外,我们还结合了扩散对抗后训练,以减轻过度平滑并进一步提高语音质量。实验结果表明,MeanVC显着优于现有的zero-shot流式VC系统,实现更高的效率和更少的参数,更好的转换质量。音频演示和代码可在https://aslp-lab.github.io/MeanVC上公开获得。
摘要:Zero-shot voice conversion (VC) aims to transfer timbre from a source speaker to any unseen target speaker while preserving linguistic content. Growing application scenarios demand models with streaming inference capabilities. This has created a pressing need for models that are simultaneously fast, lightweight, and high-fidelity. However, existing streaming methods typically rely on either autoregressive (AR) or non-autoregressive (NAR) frameworks, which either require large parameter sizes to achieve strong performance or struggle to generalize to unseen speakers. In this study, we propose MeanVC, a lightweight and streaming zero-shot VC approach. MeanVC introduces a diffusion transformer with a chunk-wise autoregressive denoising strategy, combining the strengths of both AR and NAR paradigms for efficient streaming processing. By introducing mean flows, MeanVC regresses the average velocity field during training, enabling zero-shot VC with superior speech quality and speaker similarity in a single sampling step by directly mapping from the start to the endpoint of the flow trajectory. Additionally, we incorporate diffusion adversarial post-training to mitigate over-smoothing and further enhance speech quality. Experimental results demonstrate that MeanVC significantly outperforms existing zero-shot streaming VC systems, achieving superior conversion quality with higher efficiency and significantly fewer parameters. Audio demos and code are publicly available at https://aslp-lab.github.io/MeanVC.


【12】A time-causal and time-recursive analogue of the Gabor transform
标题:Gabor变换的时间因果和时间回归模拟
链接:https://arxiv.org/abs/2308.14512

作者:Tony Lindeberg
备注:31 pages, 7 figures, 7 tables, 1 algorithm
摘要:本文提出了一种时间因果模拟的Gabor滤波器,以及时间因果和时间递归模拟的Gabor变换,其中提出的时间因果表示服从时间尺度协方差和级联属性与简化内核的时间尺度。这些构造背后的动机是使理论上有充分根据的时间-频率分析在多个时间尺度上的实时情况下,或物理或生物建模的情况下,当未来不能访问,和非因果访问未来的伽柏滤波因此是不可行的时间-频率分析的系统。   我们开发了这些表示的理论,通过将Gabor滤波中的高斯核替换为时间因果核(称为时间因果极限核)来获得,该时间因果核保证了时间因果情况下从更精细到更粗糙尺度级别的简化属性,类似于高斯核可以被证明可以保证非因果时域。以这些方式,所提出的时间-频率表示保证在待分析的信号或物理或生物现象中的特征尺度可能大幅变化的情况下,在多个尺度上进行有根据的处理,并且另外,时间-频率分析中的所有步骤必须是完全时间因果的。
摘要:This paper presents a time-causal analogue of the Gabor filter, as well as a both time-causal and time-recursive analogue of the Gabor transform, where the proposed time-causal representations obey both temporal scale covariance and a cascade property with a simplifying kernel over temporal scales. The motivation behind these constructions is to enable theoretically well-founded time-frequency analysis over multiple temporal scales for real-time situations, or for physical or biological modelling situations, when the future cannot be accessed, and the non-causal access to future in Gabor filtering is therefore not viable for a time-frequency analysis of the system.   We develop the theory for these representations, obtained by replacing the Gaussian kernel in Gabor filtering with a time-causal kernel, referred to as the time-causal limit kernel, which guarantees simplification properties from finer to coarser levels of scales in a time-causal situation, similar as the Gaussian kernel can be shown to guarantee over a non-causal temporal domain. In these ways, the proposed time-frequency representations guarantee well-founded treatment over multiple scales, in situations when the characteristic scales in the signals, or physical or biological phenomena, to be analyzed may vary substantially, and additionally all steps in the time-frequency analysis have to be fully time-causal.


eess.AS音频处理


【1】MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
标题:MeanVC:通过Mean Flow实现轻量级和流媒体Zero-Shot语音转换
链接:https://arxiv.org/abs/2510.08392

作者:Guobin Ma, Jixun Yao, Ziqian Ning, Yuepeng Jiang, Lingxin Xiong, Lei Xie, Pengcheng Zhu
摘要:Zero-shot语音转换(VC)旨在将音色从源说话人转移到任何看不见的目标说话人,同时保留语言内容。不断增长的应用场景需要具有流式推理功能的模型。这就迫切需要同时快速、轻量和高保真的模型。然而,现有的流式传输方法通常依赖于自回归(AR)或非自回归(NAR)框架,其要么需要大的参数大小来实现强大的性能,要么难以推广到看不见的扬声器。在这项研究中,我们提出了MeanVC,一个轻量级的流zero-shot VC方法。MeanVC引入了一个具有块自回归去噪策略的扩散Transformer,结合了AR和NAR范例的优势,以实现高效的流处理。通过引入平均流,MeanVC在训练过程中回归平均速度场,通过直接从流轨迹的起点映射到终点,使zero-shot VC在单个采样步骤中具有优异的语音质量和说话人相似性。此外,我们还结合了扩散对抗后训练,以减轻过度平滑并进一步提高语音质量。实验结果表明,MeanVC显着优于现有的zero-shot流式VC系统,实现更高的效率和更少的参数,更好的转换质量。音频演示和代码可在https://aslp-lab.github.io/MeanVC上公开获得。
摘要:Zero-shot voice conversion (VC) aims to transfer timbre from a source speaker to any unseen target speaker while preserving linguistic content. Growing application scenarios demand models with streaming inference capabilities. This has created a pressing need for models that are simultaneously fast, lightweight, and high-fidelity. However, existing streaming methods typically rely on either autoregressive (AR) or non-autoregressive (NAR) frameworks, which either require large parameter sizes to achieve strong performance or struggle to generalize to unseen speakers. In this study, we propose MeanVC, a lightweight and streaming zero-shot VC approach. MeanVC introduces a diffusion transformer with a chunk-wise autoregressive denoising strategy, combining the strengths of both AR and NAR paradigms for efficient streaming processing. By introducing mean flows, MeanVC regresses the average velocity field during training, enabling zero-shot VC with superior speech quality and speaker similarity in a single sampling step by directly mapping from the start to the endpoint of the flow trajectory. Additionally, we incorporate diffusion adversarial post-training to mitigate over-smoothing and further enhance speech quality. Experimental results demonstrate that MeanVC significantly outperforms existing zero-shot streaming VC systems, achieving superior conversion quality with higher efficiency and significantly fewer parameters. Audio demos and code are publicly available at https://aslp-lab.github.io/MeanVC.


【2】DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
标题:DialoSpeech:使用LLM和流匹配的双扬声器对话生成
链接:https://arxiv.org/abs/2510.08373

作者:Hanke Xie, Dake Guo, Chengyou Wang, Yue Li, Wenjie Tian, Xinfa Zhu, Xinsheng Wang, Xiulin Li, Guanqiong Miao, Bo Liu, Lei Xie
摘要:文本到语音(TTS)合成的最新进展,特别是那些利用大型语言模型(LLM),显着提高了表现力和自然度。然而,生成类人的交互式对话语音仍然是一个挑战。目前的系统面临的限制,由于缺乏双轨数据和困难,在实现自然,上下文连贯性,和intermingdynamics,如话轮转换,重叠的语音,和扬声器的一致性,在多轮对话。为了解决这些挑战,我们提出了DialoSpeech,一个双轨架构,结合了一个大的语言模型与分块流匹配的表达,人性化的对话语音合成。DialoSpeech生成自然的多话轮对话,具有连贯的说话人话轮和自然重叠,支持中文和英文以及跨语言语音合成。我们引入了一个数据处理管道来构建双轨对话数据集,促进可扩展的训练和实验验证。实验表明,我们的模型优于基线,为生成类似人类的口语对话提供了解决方案。音频样本可在https://tiamojames.github.io/DialoSpeech上获得
摘要:Recent advances in text-to-speech (TTS) synthesis, particularly those leveraging large language models (LLMs), have significantly improved expressiveness and naturalness. However, generating human-like, interactive dialogue speech remains challenging. Current systems face limitations due to the scarcity of dual-track data and difficulties in achieving naturalness, contextual coherence, and interactional dynamics, such as turn-taking, overlapping speech, and speaker consistency, in multi-turn conversations. To address these challenges, we propose DialoSpeech, a dual-track architecture combining a large language model with Chunked Flow Matching for expressive, human-like dialogue speech synthesis. DialoSpeech generates natural multi-turn conversations with coherent speaker turns and natural overlaps, supporting both Chinese and English and cross-lingual speech synthesis. We introduce a data processing pipeline to construct dual-track dialogue datasets, facilitating scalable training and experimental validation. Experiments show that our model outperforms baselines, offering a solution for generating human-like spoken dialogues. Audio samples are available at https://tiamojames.github.io/DialoSpeech


【3】Pseudo2Real: Task Arithmetic for Pseudo-Label Correction in Automatic Speech Recognition
标题:伪2Real:自动语音识别中伪标签纠正的任务算法
链接:https://arxiv.org/abs/2510.08047

作者:Yi-Cheng Lin, Yu-Hsuan Li Liang, Hsuan Su, Tzu-Quan Lin, Shang-Tse Chen, Yun-Nung Chen, Hung-yi Lee
摘要:域转移下的鲁棒ASR至关重要,因为现实世界的系统会遇到不可见的口音和具有有限标记数据的域。虽然伪标签提供了一个实用的解决方案,但它经常会引入系统性的、特定于口音的错误,而过滤无法修复这些错误。我们要问:我们如何在没有目标基础事实的情况下纠正这些反复出现的偏见?我们提出了一个简单的参数空间校正:在包含真实和伪标签数据的源域中,两个ASR模型从相同的初始化进行微调,一个在真实标签上,另一个在伪标签上,它们的权重差形成了一个捕获伪标签偏差的校正向量。当应用于伪标记目标模型时,该向量增强了识别,使用Whisper微型模型在AfriSpeech-200上跨10种非洲口音实现了高达35%的相对单词错误率(WER)降低。
摘要:Robust ASR under domain shift is crucial because real-world systems encounter unseen accents and domains with limited labeled data. Although pseudo-labeling offers a practical workaround, it often introduces systematic, accent-specific errors that filtering fails to fix. We ask: How can we correct these recurring biases without target ground truth? We propose a simple parameter-space correction: in a source domain containing both real and pseudo-labeled data, two ASR models are fine-tuned from the same initialization, one on ground-truth labels and the other on pseudo-labels, and their weight difference forms a correction vector that captures pseudo-label biases. When applied to a pseudo-labeled target model, this vector enhances recognition, achieving up to a 35% relative Word Error Rate (WER) reduction on AfriSpeech-200 across ten African accents with the Whisper tiny model.


【4】Bloodroot: When Watermarking Turns Poisonous For Stealthy Backdoor
标题:Bloodroot:当水印对隐形后门有毒时
链接:https://arxiv.org/abs/2510.07909

作者:Kuan-Yu Chen, Yi-Cheng Lin, Jeng-Lin Li, Jian-Jiun Ding
备注:5 pages, 3 figures
摘要:后门数据中毒是保护所有权和抵御恶意攻击的关键技术。在训练数据中嵌入隐藏的触发器可以操纵模型输出,实现出处验证,并阻止未经授权的使用。然而,当前的音频后门方法是次优的,因为中毒的音频通常表现出降低的感知质量,这对人类听众是明显的。这项工作探讨了音频水印在实现成功中毒的内在隐蔽性和有效性。我们提出了一种新的水印触发器概念,通过对抗性LoRA微调集成到Bloodroot后门框架中,从而提高了感知质量,同时实现了更高的触发成功率和干净样本精度。语音识别(SR)和说话人识别(SID)数据集上的实验表明,基于水印的中毒仍然有效的声学滤波和模型修剪。提出的Bloodroot后门框架不仅保护了数据到模型的所有权,而且很好地揭示了对抗性滥用的风险。
摘要:Backdoor data poisoning is a crucial technique for ownership protection and defending against malicious attacks. Embedding hidden triggers in training data can manipulate model outputs, enabling provenance verification, and deterring unauthorized use. However, current audio backdoor methods are suboptimal, as poisoned audio often exhibits degraded perceptual quality, which is noticeable to human listeners. This work explores the intrinsic stealthiness and effectiveness of audio watermarking in achieving successful poisoning. We propose a novel Watermark-as-Trigger concept, integrated into the Bloodroot backdoor framework via adversarial LoRA fine-tuning, which enhances perceptual quality while achieving a much higher trigger success rate and clean-sample accuracy. Experiments on speech recognition (SR) and speaker identification (SID) datasets show that watermark-based poisoning remains effective under acoustic filtering and model pruning. The proposed Bloodroot backdoor framework not only secures data-to-model ownership, but also well reveals the risk of adversarial misuse.


【5】Guitar Tone Morphing by Diffusion-based Model
标题:基于扩散模型的吉他音色变形
链接:https://arxiv.org/abs/2510.07908

作者:Kuan-Yu Chen, Kuan-Lin Chen, Yu-Chieh Yu, Jian-Jiun Ding
备注:5 pages
摘要:在音乐信息检索中,乐器尤其是电吉他的音色建模和转换由于其音色的丰富性和表达的灵活性而受到越来越多的关注。音调变形可以在不同的吉他声音之间平滑过渡,让音乐家更自由地探索新的纹理和个性化他们的表演。本研究探讨了基于学习的吉他音调变形方法,从LoRA微调开始,以提高有限数据上的模型性能。此外,我们介绍了一种更简单的方法,命名为球面插值使用Music2Latent。它比更复杂的微调方法产生更好的结果。实验表明,该架构产生更平滑,更自然的音调过渡,使其成为一个实用和有效的工具,音乐制作和实时音频效果。
摘要:In Music Information Retrieval (MIR), modeling and transforming the tone of musical instruments, particularly electric guitars, has gained increasing attention due to the richness of the instrument tone and the flexibility of expression. Tone morphing enables smooth transitions between different guitar sounds, giving musicians greater freedom to explore new textures and personalize their performances. This study explores learning-based approaches for guitar tone morphing, beginning with LoRA fine-tuning to improve the model performance on limited data. Moreover, we introduce a simpler method, named spherical interpolation using Music2Latent. It yields significantly better results than the more complex fine-tuning approach. Experiments show that the proposed architecture generates smoother and more natural tone transitions, making it a practical and efficient tool for music production and real-time audio effects.


【6】Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner
标题:Full-Duplex-Bench-v2:具有自动审查员的Duplex对话系统的多回合评估框架
链接:https://arxiv.org/abs/2510.07838

作者:Guan-Ting Lin, Shih-Yun Shan Kuan, Jiatong Shi, Kai-Wei Chang, Siddhant Arora, Shinji Watanabe, Hung-yi Lee
备注:Work in progress
摘要:虽然全双工语音代理通过同时说话和倾听实现自然,低延迟的交互,但它们在多回合设置中的一致性和任务性能仍然未得到充分研究。我们介绍了Full-Duplex-Bench-v2(FDB-v2),这是一个流媒体框架,它与自动检查器集成,可以在两种起搏设置(快速与慢速)下执行阶段性目标。FDB-v2涵盖四个任务系列:日常、纠正、实体跟踪和安全。我们报告的话轮转换流畅性,多轮指令以下,和特定任务的能力。该框架是可扩展的,支持商业API和开源模型。当我们使用FDB-v2测试全双工系统时,当人们同时说话时,它们经常会感到困惑,难以顺利处理更正,有时会忘记正在谈论的人或内容。通过开源的标准化流协议和任务集,FDB-v2可以轻松扩展到新的任务系列,使社区能够定制和加速多圈全双工系统的评估。
摘要:While full-duplex speech agents enable natural, low-latency interaction by speaking and listening simultaneously, their consistency and task performance in multi-turn settings remain underexplored. We introduce Full-Duplex-Bench-v2 (FDB-v2), a streaming framework that integrates with an automated examiner that enforces staged goals under two pacing setups (Fast vs. Slow). FDB-v2 covers four task families: daily, correction, entity tracking, and safety. We report turn-taking fluency, multi-turn instruction following, and task-specific competence. The framework is extensible, supporting both commercial APIs and open source models. When we test full-duplex systems with FDB-v2, they often get confused when people talk at the same time, struggle to handle corrections smoothly, and sometimes lose track of who or what is being talked about. Through an open-sourced, standardized streaming protocol and a task set, FDB-v2 makes it easy to extend to new task families, allowing the community to tailor and accelerate evaluation of multi-turn full-duplex systems.


【7】SALAD-VAE: Semantic Audio Compression with Language-Audio Distillation
标题:SALAD-VAE:采用语音音频蒸馏的语义音频压缩
链接:https://arxiv.org/abs/2510.07592

作者:Sebastian Braun, Hannes Gamper, Dimitra Emmanouilidou
备注:submitted to ICASSP 2026
摘要:现代生成和多模态模型越来越依赖于紧凑的潜在表示,这些表示可以用高保真重建来交换和平衡语义丰富性。我们介绍SALAD-VAE,一个连续的和高度紧凑的语义音频变分自动编码器,它在频域中工作,并实现了最先进的压缩,具有非常低的潜在帧速率(7.8 Hz),同时表面的语义结构,并产生高音频质量。我们增强了标准的VAE语义损失和增强,特别是对比学习和基于CLAP的嵌入蒸馏,使其能够在不同的音频领域推广。SALAD-VAE的计算复杂度明显低于最先进的VAE,与它们的重建质量相匹配,同时在广泛的分类基准上始终优于它们。此外,所提出的附加损失函数提供了训练的CLAP投影层,其可以用于zero-shot音频字幕和分类匹配预训练的CLAP音频文本嵌入。
摘要:Modern generative and multimodal models increasingly rely on compact latent representations that trade and balance semantic richness with high-fidelity reconstruction. We introduce SALAD-VAE, a continuous and highly compact semantic Audio Variational Autoencoder, which operates in the frequency domain and achieves state-of-the-art compression with very low latent frame rate (7.8 Hz) while surfacing semantic structure and producing high audio quality. We enhance the standard VAE semantic losses and augmentation, specifically contrastive learning and CLAP-based embedding distillation, enabling it to generalize across diverse audio domains. With a significantly less computational complex architecture than comparable state-of-the-art VAEs, SALAD-VAE matches their reconstruction quality while it consistently outperforms them on a wide range of classification benchmarks. Furthermore, the proposed additional loss function provides a trained CLAP projection layer, which can be used zero-shot audio captioning and classification matching pretrained CLAP audio-text embeddings.


【8】A time-causal and time-recursive analogue of the Gabor transform
标题:Gabor变换的时间因果和时间回归模拟
链接:https://arxiv.org/abs/2308.14512

作者:Tony Lindeberg
备注:31 pages, 7 figures, 7 tables, 1 algorithm
摘要:本文提出了一种时间因果模拟的Gabor滤波器,以及时间因果和时间递归模拟的Gabor变换,其中提出的时间因果表示服从时间尺度协方差和级联属性与简化内核的时间尺度。这些构造背后的动机是使理论上有充分根据的时间-频率分析在多个时间尺度上的实时情况下,或物理或生物建模的情况下,当未来不能访问,和非因果访问未来的伽柏滤波因此是不可行的时间-频率分析的系统。   我们开发了这些表示的理论,通过将Gabor滤波中的高斯核替换为时间因果核(称为时间因果极限核)来获得,该时间因果核保证了时间因果情况下从更精细到更粗糙尺度级别的简化属性,类似于高斯核可以被证明可以保证非因果时域。以这些方式,所提出的时间-频率表示保证在待分析的信号或物理或生物现象中的特征尺度可能大幅变化的情况下,在多个尺度上进行有根据的处理,并且另外,时间-频率分析中的所有步骤必须是完全时间因果的。
摘要:This paper presents a time-causal analogue of the Gabor filter, as well as a both time-causal and time-recursive analogue of the Gabor transform, where the proposed time-causal representations obey both temporal scale covariance and a cascade property with a simplifying kernel over temporal scales. The motivation behind these constructions is to enable theoretically well-founded time-frequency analysis over multiple temporal scales for real-time situations, or for physical or biological modelling situations, when the future cannot be accessed, and the non-causal access to future in Gabor filtering is therefore not viable for a time-frequency analysis of the system.   We develop the theory for these representations, obtained by replacing the Gaussian kernel in Gabor filtering with a time-causal kernel, referred to as the time-causal limit kernel, which guarantees simplification properties from finer to coarser levels of scales in a time-causal situation, similar as the Gaussian kernel can be shown to guarantee over a non-causal temporal domain. In these ways, the proposed time-frequency representations guarantee well-founded treatment over multiple scales, in situations when the characteristic scales in the signals, or physical or biological phenomena, to be analyzed may vary substantially, and additionally all steps in the time-frequency analysis have to be fully time-causal.


【9】Leveraging Whisper Embeddings for Audio-based Lyrics Matching
标题:利用Whisper嵌入进行音频歌词匹配
链接:https://arxiv.org/abs/2510.08176

作者:Eleonora Mancini, Joan Serrà, Paolo Torroni, Yuki Mitsufuji
摘要:基于音频的歌词匹配可以是其他基于内容的检索方法的一个有吸引力的替代方案,但现有的方法往往受到有限的再现性和不一致的基线。在这项工作中,我们介绍了WEALY,一个完全可再现的管道,利用Whisper解码器嵌入歌词匹配任务。WEALY建立了强大而透明的基线,同时还探索了整合文本和声学特征的多模态扩展。通过对标准数据集的广泛实验,我们证明WEALY实现了与缺乏可重复性的最先进方法相当的性能。此外,我们还提供了关于语言鲁棒性、损失函数和嵌入策略的消融研究和分析。这项工作为未来的研究提供了一个可靠的基准,并强调了语音技术在音乐信息检索任务中的潜力。
摘要:Audio-based lyrics matching can be an appealing alternative to other content-based retrieval approaches, but existing methods often suffer from limited reproducibility and inconsistent baselines. In this work, we introduce WEALY, a fully reproducible pipeline that leverages Whisper decoder embeddings for lyrics matching tasks. WEALY establishes robust and transparent baselines, while also exploring multimodal extensions that integrate textual and acoustic features. Through extensive experiments on standard datasets, we demonstrate that WEALY achieves a performance comparable to state-of-the-art methods that lack reproducibility. In addition, we provide ablation studies and analyses on language robustness, loss functions, and embedding strategies. This work contributes a reliable benchmark for future research, and underscores the potential of speech technologies for music information retrieval tasks.


【10】Personality-Enhanced Multimodal Depression Detection in the Elderly
标题:老年人的个性增强多模式抑郁检测
链接:https://arxiv.org/abs/2510.08004

作者:Honghong Wang, Jing Deng, Rong Zheng
备注:6 pages,2 figures,accepted by ACM Multimedia Asia 2025
摘要:本文介绍了我们在ACM MM 2025上对多模态个性感知抑郁检测(MPDD)挑战的解决方案。我们提出了一个多模态抑郁症检测模型,在老年人,包括人格特征。我们引入了一种基于共同注意力机制的多特征融合方法,以有效地将LLDs,MFCC和Wav2Vec特征集成在音频模态中。对于视频模态,我们结合了从OpenFace,ResNet和DenseNet中提取的表示来构建一个全面的视觉特征集。认识到人格在抑郁症检测中的关键作用,我们设计了一个互动模块,捕捉人格特质和多模态特征之间的关系。来自MPDD老年抑郁症检测跟踪的实验结果表明,我们的方法显着提高了性能,为老年人群中多模态抑郁症检测的未来研究提供了有价值的见解。
摘要:This paper presents our solution to the Multimodal Personality-aware Depression Detection (MPDD) challenge at ACM MM 2025. We propose a multimodal depression detection model in the Elderly that incorporates personality characteristics. We introduce a multi-feature fusion approach based on a co-attention mechanism to effectively integrate LLDs, MFCCs, and Wav2Vec features in the audio modality. For the video modality, we combine representations extracted from OpenFace, ResNet, and DenseNet to construct a comprehensive visual feature set. Recognizing the critical role of personality in depression detection, we design an interaction module that captures the relationships between personality traits and multimodal features. Experimental results from the MPDD Elderly Depression Detection track demonstrate that our method significantly enhances performance, providing valuable insights for future research in multimodal depression detection among elderly populations.


【11】ACMID: Automatic Curation of Musical Instrument Dataset for 7-Stem Music Source Separation
标题:ACMID:用于7-Stem音乐源分离的乐器数据集的自动处理
链接:https://arxiv.org/abs/2510.07840

作者:Ji Yu, Yang shuo, Xu Yuetonghui, Liu Mengmei, Ji Qiang, Han Zerui
摘要:目前的音乐源分离(MSS)方法大多依赖于监督学习,受到训练数据数量和质量的限制。虽然网络爬虫可以带来丰富的数据,但平台级的音轨标注往往会导致元数据的不匹配,从而影响“音频-标签”对的准确获取。为了解决这个问题,我们提出了ACMID:通过对大量原始数据进行网络抓取生成的MSS数据集,然后通过构建在预先训练的音频编码器上的乐器分类器进行自动清理,该编码器从抓取的曲目中过滤和聚合目标乐器的干净片段,从而产生经过改进的ACMID清理数据集。利用丰富的数据,我们将传统的4-干(声乐/低音/鼓/其他)扩展到7-干(钢琴/鼓/低音/原声吉他/电吉他/弦乐器/管乐器),从而实现高粒度MSS系统。在SOTA MSS模型上的实验表明了两个关键结果:(i)使用ACMID-Cleaned训练的MSS模型与使用ACMID-Uncleaned相比,SDR性能提高了2.39dB,证明了我们的数据清洗过程的有效性;(ii)将ACMID-Cleaned纳入训练,使MSS模型的平均性能提高了1.16dB,证实了我们数据集的价值。我们的数据抓取代码,清洁模型代码和权重可在:https://github.com/scottishfold0621/ACMID.
摘要:Most current music source separation (MSS) methods rely on supervised learning, limited by training data quan- tity and quality. Though web-crawling can bring abundant data, platform-level track labeling often causes metadata mismatches, impeding accurate "audio-label" pair acquisi- tion. To address this, we present ACMID: a dataset for MSS generated through web crawling of extensive raw data, fol- lowed by automatic cleaning via an instrument classifier built on a pre-trained audio encoder that filters and aggregates clean segments of target instruments from the crawled tracks, resulting in the refined ACMID-Cleaned dataset. Leverag- ing abundant data, we expand the conventional classifica- tion from 4-stem (Vocal/Bass/Drums/Others) to 7-stem (Pi- ano/Drums/Bass/Acoustic Guitar/Electric Guitar/Strings/Wind- Brass), enabling high granularity MSS systems. Experiments on SOTA MSS model demonstrates two key results: (i) MSS model trained with ACMID-Cleaned achieved a 2.39dB improvement in SDR performance compared to that with ACMID-Uncleaned, demostrating the effectiveness of our data cleaning procedure; (ii) incorporating ACMID-Cleaned to training enhances MSS model's average performance by 1.16dB, confirming the value of our dataset. Our data crawl- ing code, cleaning model code and weights are available at: https://github.com/scottishfold0621/ACMID.


【12】Can Speech LLMs Think while Listening?
标题:言语LL可以边听边思考吗?
链接:https://arxiv.org/abs/2510.07497

作者:Yi-Jen Shih, Desh Raj, Chunyang Wu, Wei Zhou, SK Bong, Yashesh Gaur, Jay Mahadeokar, Ozlem Kalinli, Mike Seltzer
摘要:语音大语言模型(speech LLM)的最新进展已经实现了无缝的口语交互,但这些系统仍然难以完成复杂的推理任务。以前,思想链(CoT)提示或微调已被证明可以显着提高基于文本的LLM的推理能力。在这项工作中,我们研究了CoT微调对多流语音LLM的影响,证明了在一系列口语推理任务中,文本空间中的推理将语音LLM的准确性平均提高了2.4倍。除了准确性之外,语音响应的延迟是与基于语音的代理交互的关键因素。受人类“边听边想”行为的启发,我们提出了一些方法,通过允许模型在用户查询结束之前开始推理来减少推理的额外等待时间。为了实现这一点,我们引入了一个基于熵的度量,“问题完整性”,它作为一个指标,以指导模型的最佳时间开始推理。这种方法提供了更大的控制精度-延迟权衡相比,基于神经网络的方法,在同等的延迟条件下,ARC容易产生4%的精度增益。最后,我们对使用拒绝采样创建的偏好数据使用直接偏好优化(DPO),以进一步推动准确性-延迟帕累托边界,从而在不损失准确性的情况下减少70%的延迟。
摘要:Recent advances in speech large language models (speech LLMs) have enabled seamless spoken interactions, but these systems still struggle with complex reasoning tasks. Previously, chain-of-thought (CoT) prompting or fine-tuning has been to shown to significantly improve the reasoning abilities of text-based LLMs. In this work, we investigate the effect of CoT fine-tuning for multi-stream speech LLMs, demonstrating that reasoning in text space improves the accuracy of speech LLMs by 2.4x, on average, over a suite of spoken reasoning tasks. Beyond accuracy, the latency of the spoken response is a crucial factor for interacting with voice-based agents. Inspired by the human behavior of "thinking while listening," we propose methods to reduce the additional latency from reasoning by allowing the model to start reasoning before the user query has ended. To achieve this, we introduce an entropy-based metric, "question completeness," which acts as an indicator to guide the model on the optimal time to start reasoning. This method provides greater control over the accuracy-latency trade-off compared with heuristic-based approaches and, under equivalent latency conditions, yields a 4% accuracy gain on ARC-Easy. Finally, we use Direct Preference Optimization (DPO) on preference data created using rejection sampling to push the accuracy-latency pareto frontier further, resulting in a 70% reduction in latency without loss in accuracy.


【13】LASER: An LLM-based ASR Scoring and Evaluation Rubric
标题:LAPER:基于LLM的ASC评分和评估版块
链接:https://arxiv.org/abs/2510.07437

作者:Amruta Parulekar, Preethi Jyothi
备注:Accepted to EMNLP 2025
摘要:标准的ASR评估指标,如单词错误率(WER),往往不公平地惩罚形态和句法的细微差别,不显着改变句子语义。我们引入了一个基于LLM的评分规则LASER,它利用了最先进的LLM的上下文学习能力,从具有详细示例的提示中学习。使用Gemini 2.5 Pro的印地语激光评分与人类注释的相关性非常高,达到94%。提示中的印地语例子在分析其他印度语言如马拉地语、卡纳达语和马拉雅拉姆语中的错误时也很有效。我们还演示了如何对来自参考和ASR预测的词对示例进行微调,以接近89%的准确率预测应该应用什么样的惩罚。
摘要:Standard ASR evaluation metrics like Word Error Rate (WER) tend to unfairly penalize morphological and syntactic nuances that do not significantly alter sentence semantics. We introduce an LLM-based scoring rubric LASER that leverages state-of-the-art LLMs' in-context learning abilities to learn from prompts with detailed examples. Hindi LASER scores using Gemini 2.5 Pro achieved a very high correlation score of 94% with human annotations. Hindi examples in the prompt were also effective in analyzing errors in other Indian languages such as Marathi, Kannada and Malayalam. We also demonstrate how a smaller LLM like Llama 3 can be finetuned on word-pair examples derived from reference and ASR predictions to predict what kind of penalty should be applied with close to 89% accuracy.


机器翻译由腾讯交互翻译提供,仅供参考