微信公众号:arXiv_Daily
cs.SD语音
【1】Efficient and Fast Generative-Based Singing Voice Separation using a Latent Diffusion Model
标题:使用潜在扩散模型高效快速的基于生成的歌唱声音分离
链接:https://arxiv.org/abs/2511.20470
备注:Accepted for oral presentation at IJCNN 2025
摘要:从音乐混合物中提取单个元素是音乐创作和实践的宝贵工具。虽然神经网络被优化以将混合频谱图掩蔽或转换为单个源一直是领先的方法,但音乐信号中的源重叠和相关性构成了固有的挑战。此外,访问混合物中的所有源对于训练这些系统是至关重要的,尽管复杂。存在以生成方式解决这些挑战的尝试,然而,分离性能和推理效率仍然有限。在这项工作中,我们研究了扩散模型的潜力,以推进弥合这一差距,专注于生成的歌声分离只依赖于相应的对孤立的人声和混合训练。为了与创意工作流程保持一致,我们利用潜在扩散:系统生成在紧凑的潜在空间中编码的样本,然后将其解码为音频。这可以实现有效的优化和更快的推理。我们的系统仅使用开放数据进行训练。我们优于现有的生成分离系统,并在信号质量测量和干扰去除列表上与非生成系统进行比较。我们提供了一个潜在的编码器的噪声鲁棒性研究,提供其潜在的任务的见解。我们发布了一个模块化的工具包,以进一步研究这个主题。
摘要:Extracting individual elements from music mixtures is a valuable tool for music production and practice. While neural networks optimized to mask or transform mixture spectrograms into the individual source(s) have been the leading approach, the source overlap and correlation in music signals poses an inherent challenge. Also, accessing all sources in the mixture is crucial to train these systems, while complicated. Attempts to address these challenges in a generative fashion exist, however, the separation performance and inference efficiency remain limited. In this work, we study the potential of diffusion models to advance toward bridging this gap, focusing on generative singing voice separation relying only on corresponding pairs of isolated vocals and mixtures for training. To align with creative workflows, we leverage latent diffusion: the system generates samples encoded in a compact latent space, and subsequently decodes these into audio. This enables efficient optimization and faster inference. Our system is trained using only open data. We outperform existing generative separation systems, and level the compared non-generative systems on a list of signal quality measures and on interference removal. We provide a noise robustness study on the latent encoder, providing insights on its potential for the task. We release a modular toolkit for further research on the topic.
【2】Differentiable Attenuation Filters for Feedback Delay Networks
标题:反馈延迟网络的可微分衰减滤波器
链接:https://arxiv.org/abs/2511.20380
摘要:提出了一种基于反馈延迟网络(FDNs)的数字音频混响系统中衰减滤波器的设计方法。我们的方法使用无限脉冲响应(IIR)滤波器的二阶部分(SOS)作为参数均衡器(PEQ),使频率相关的混响衰减的精细控制。与传统的图形均衡器设计,这需要许多滤波器每个延迟线,我们提出了一个可扩展的解决方案,滤波器的数量可以调整。频率、增益和品质因数(Q)参数是延迟线上的共享参数,只有增益根据延迟长度进行调整。这种设计不仅减少了优化参数的数量,而且保持了完全可微性,并与基于梯度的学习框架兼容。利用模拟滤波器设计的原理,我们的方法允许使用监督学习进行高效和准确的滤波器拟合。我们的方法提供了一个灵活和差异化的设计,实现国家的最先进的性能,同时显着降低计算成本。
摘要:We introduce a novel method for designing attenuation filters in digital audio reverberation systems based on Feedback Delay Net- works (FDNs). Our approach uses Second Order Sections (SOS) of Infinite Impulse Response (IIR) filters arranged as parametric equalizers (PEQ), enabling fine control over frequency-dependent reverberation decay. Unlike traditional graphic equalizer designs, which require numerous filters per delay line, we propose a scal- able solution where the number of filters can be adjusted. The fre- quency, gain, and quality factor (Q) parameters are shared parame- ters across delay lines and only the gain is adjusted based on delay length. This design not only reduces the number of optimization parameters, but also remains fully differentiable and compatible with gradient-based learning frameworks. Leveraging principles of analog filter design, our method allows for efficient and accu- rate filter fitting using supervised learning. Our method delivers a flexible and differentiable design, achieving state-of-the-art per- formance while significantly reducing computational cost.
【3】DUO-TOK: Dual-Track Semantic Music Tokenizer for Vocal-Accompaniment Generation
标题:DUO-TOK:用于声乐伴奏生成的双轨语义音乐代币器
链接:https://arxiv.org/abs/2511.20224
备注:17 pages, 5 figures, 8 tables. Project page: https://eps-acoustic-revolution-lab.github.io/DUO_TOK/
摘要:Duo-Tok是一种用于声乐伴奏音乐的源感知双码本标记器,其目标是现代歌词到歌曲系统中重建质量和语言模型(LM)可学习性之间日益增长的紧张关系。现有的编解码器要么优先考虑具有难以建模的声学令牌的高保真重建,要么积极地压缩成LM友好但有损的语义令牌,并且它们很少使令牌化器本身意识到双轨结构。Duo-Tok遵循以SSL为中心的四阶段流水线:我们首先在大规模音频上预训练BEST-RQ风格的编码器,然后使用高斯替换噪声和多任务监督稳定和分解表示,然后冻结编码器以学习基于SimVQ的双码本,并为人声和伴奏进行硬路由,最后在离散令牌上训练潜在扩散解码器。0.75 kbps的Duo-Tok移动了经验重建生成的帕累托边界,在比较的编解码器中实现了最佳的音乐标记AP和最低的词汇归一化LM困惑,同时保持了与最先进的音乐标记器相当的重建质量。
摘要:Duo-Tok is a source-aware dual-codebook tokenizer for vocal-accompaniment music that targets the growing tension between reconstruction quality and language-model (LM) learnability in modern lyrics-to-song systems. Existing codecs either prioritize high-fidelity reconstruction with difficult-to-model acoustic tokens or compress aggressively into semantic tokens that are LM-friendly but lossy, and they rarely make the tokenizer itself aware of dual-track structure. Duo-Tok follows a four-stage, SSL-centered pipeline: we first pretrain a BEST-RQ-style encoder on large-scale audio, then stabilize and factorize the representation with Gaussian replacement noise and multi-task supervision, before freezing the encoder to learn SimVQ-based dual codebooks with hard routing for vocals and accompaniment, and finally training latent diffusion decoders on top of the discrete tokens. Duo-Tok at 0.75 kbps shifts the empirical reconstruction-generation Pareto frontier, achieving the best music-tagging AP and the lowest vocabulary-normalized LM perplexity among compared codecs while maintaining reconstruction quality comparable to state-of-the-art music tokenizers.
【4】Mispronunciation Detection and Diagnosis Without Model Training: A Retrieval-Based Approach
标题:无需模型训练的发音错误检测和诊断:基于检索的方法
链接:https://arxiv.org/abs/2511.20107
摘要:发音错误检测与诊断(MDD)是语言学习和言语治疗的关键。与传统的方法,需要评分模型或训练音素级模型,我们提出了一种新的训练免费的框架,利用检索技术与预训练的自动语音识别模型。我们的方法避免了特定于音素的建模或额外的特定于任务的训练,同时仍然实现了准确的检测和诊断发音错误。在L2-ARCTIC数据集上的实验表明,该方法在避免模型训练复杂性的同时,获得了69.60%的F1分数。
摘要:Mispronunciation Detection and Diagnosis (MDD) is crucial for language learning and speech therapy. Unlike conventional methods that require scoring models or training phoneme-level models, we propose a novel training-free framework that leverages retrieval techniques with a pretrained Automatic Speech Recognition model. Our method avoids phoneme-specific modeling or additional task-specific training, while still achieving accurate detection and diagnosis of pronunciation errors. Experiments on the L2-ARCTIC dataset show that our method achieves a superior F1 score of 69.60% while avoiding the complexity of model training.
【5】Continual Audio Deepfake Detection via Universal Adversarial Perturbation
标题:通过普遍对抗扰动进行连续音频深度伪造检测
链接:https://arxiv.org/abs/2511.19974
摘要:语音合成和语音转换技术的快速发展已经引起了多媒体取证中的重大安全问题。尽管当前的检测模型表现出令人印象深刻的性能,但它们很难保持对不断发展的deepfake攻击的有效性。此外,使用历史训练数据不断微调这些模型会产生大量的计算和存储成本。为了解决这些限制,我们提出了一种新的框架,将通用对抗扰动(UAP)纳入音频深度伪造检测,使模型能够保留历史欺骗分布的知识,而无需直接访问过去的数据。我们的方法在微调过程中将UAP与预训练的自监督音频模型无缝集成。大量的实验验证了我们方法的有效性,展示了它作为音频deepfake检测中持续学习的有效解决方案的潜力。
摘要:The rapid advancement of speech synthesis and voice conversion technologies has raised significant security concerns in multimedia forensics. Although current detection models demonstrate impressive performance, they struggle to maintain effectiveness against constantly evolving deepfake attacks. Additionally, continually fine-tuning these models using historical training data incurs substantial computational and storage costs. To address these limitations, we propose a novel framework that incorporates Universal Adversarial Perturbation (UAP) into audio deepfake detection, enabling models to retain knowledge of historical spoofing distribution without direct access to past data. Our method integrates UAP seamlessly with pre-trained self-supervised audio models during fine-tuning. Extensive experiments validate the effectiveness of our approach, showcasing its potential as an efficient solution for continual learning in audio deepfake detection.
【6】Evaluating Objective Speech Quality Metrics for Neural Audio Codecs
标题:评估神经音频编解码器的客观语音质量指标
链接:https://arxiv.org/abs/2511.19734
摘要:神经音频编解码器最近因其在生成建模中的使用而受到欢迎,因为它们以低比特率提供高保真音频重建。虽然人类听力研究仍然是评估感知质量的黄金标准,但它们既耗时又不切实际。在这项工作中,我们研究了现有的客观质量指标在评估最近的神经音频编解码器的性能的可靠性。为此,我们对高保真语音信号进行了MUSHRA听力测试,并分析了主观分数与广泛使用的客观指标之间的相关性。我们的研究结果表明,虽然一些指标与人类感知一致,但其他指标很难捕捉相关的扭曲。我们的研究结果为使用神经音频编解码器进行语音时选择适当的评估指标提供了实际指导。
摘要:Neural audio codecs have gained recent popularity for their use in generative modeling as they offer high-fidelity audio reconstruction at low bitrates. While human listening studies remain the gold standard for assessing perceptual quality, they are time-consuming and impractical. In this work, we examine the reliability of existing objective quality metrics in assessing the performance of recent neural audio codecs. To this end, we conduct a MUSHRA listening test on high-fidelity speech signals and analyze the correlation between subjective scores and widely used objective metrics. Our results show that, while some metrics align well with human perception, others struggle to capture relevant distortions. Our findings provide practical guidance for selecting appropriate evaluation metrics when using neural audio codecs for speech.
【7】BERT-APC: A Reference-free Framework for Automatic Pitch Correction via Musical Context Inference
标题:BERT-IPC:通过音乐上下文推理自动音调纠正的无参考框架
链接:https://arxiv.org/abs/2511.20006
备注:12 pages, 6 figures, 5 tables
摘要:自动音高校正(APC)通过将音高偏差与预期的音符对齐来增强声乐录音。然而,现有的APC系统要么依赖于参考音高,这限制了它们的实际适用性,要么采用简单的音高估计算法,往往不能保持表现力和自然性。我们提出了BERT-APC,一种新的无参考APC框架,纠正音高错误,同时保持声乐表演的自然表现力。在BERT-APC中,一种新型的固定音高预测器首先从失调的歌声中估计每个音符的感知音高。上下文感知音符音高预测器通过利用被重新利用以并入音乐上下文的音乐语言模型来估计预期音高序列。最后,音符级校正算法修复音高错误,同时保留情感表达的故意音高偏差。此外,我们引入了一个可学习的数据增强策略,通过模拟现实的失谐模式,提高了音乐语言模型的鲁棒性。与最近的两个歌唱声转录模型相比,BERT-APC在音符音高预测方面表现出更好的性能,在原始音高准确度方面,在高度失谐的样本上比第二好的模型ROSVOT高出10.49%p。在MOS测试中,BERT-APC获得了4.32美元的最高分,这显著高于广泛使用的商业APC工具AutoTune(3.22美元)和Melodyne(3.08美元),同时保持了相当的保留表达细微差别的能力。据我们所知,这是第一个利用音乐语言模型来实现符号音乐上下文的无参考音高校正的APC模型。BERT-APC的校正音频样本可在线获得。
摘要:Automatic Pitch Correction (APC) enhances vocal recordings by aligning pitch deviations with the intended musical notes. However, existing APC systems either rely on reference pitches, which limits their practical applicability, or employ simple pitch estimation algorithms that often fail to preserve expressiveness and naturalness. We propose BERT-APC, a novel reference-free APC framework that corrects pitch errors while maintaining the natural expressiveness of vocal performances. In BERT-APC, a novel stationary pitch predictor first estimates the perceived pitch of each note from the detuned singing voice. A context-aware note pitch predictor estimates the intended pitch sequence by leveraging a music language model repurposed to incorporate musical context. Finally, a note-level correction algorithm fixes pitch errors while preserving intentional pitch deviations for emotional expression. In addition, we introduce a learnable data augmentation strategy that improves the robustness of the music language model by simulating realistic detuning patterns. Compared to two recent singing voice transcription models, BERT-APC demonstrated superior performance in note pitch prediction, outperforming the second-best model, ROSVOT, by 10.49%p on highly detuned samples in terms of the raw pitch accuracy. In the MOS test, BERT-APC achieved the highest score of $4.32 \pm 0.15$, which is significantly higher than those of the widely-used commercial APC tools, AutoTune ($3.22 \pm 0.18$) and Melodyne ($3.08 \pm 0.18$), while maintaining a comparable ability to preserve expressive nuances. To the best of our knowledge, this is the first APC model that leverages a music language model to achieve reference-free pitch correction with symbolic musical context. The corrected audio samples of BERT-APC are available online.
【1】BERT-APC: A Reference-free Framework for Automatic Pitch Correction via Musical Context Inference
标题:BERT-IPC:通过音乐上下文推理自动音调纠正的无参考框架
链接:https://arxiv.org/abs/2511.20006
备注:12 pages, 6 figures, 5 tables
摘要:自动音高校正(APC)通过将音高偏差与预期的音符对齐来增强声乐录音。然而,现有的APC系统要么依赖于参考音高,这限制了它们的实际适用性,要么采用简单的音高估计算法,往往不能保持表现力和自然性。我们提出了BERT-APC,一种新的无参考APC框架,纠正音高错误,同时保持声乐表演的自然表现力。在BERT-APC中,一种新颖的固定音高预测器首先从失谐的歌声中估计每个音符的感知音高。上下文感知音符音高预测器通过利用被重新利用以并入音乐上下文的音乐语言模型来估计预期音高序列。最后,音符级校正算法修复音高错误,同时保留情感表达的故意音高偏差。此外,我们引入了一个可学习的数据增强策略,通过模拟现实的失谐模式,提高了音乐语言模型的鲁棒性。与最近的两个歌唱声转录模型相比,BERT-APC在音符音高预测方面表现出更好的性能,在原始音高准确度方面,在高度失谐的样本上比第二好的模型ROSVOT高出10.49%p。在MOS测试中,BERT-APC获得了4.32美元的最高分,这显著高于广泛使用的商业APC工具AutoTune(3.22美元)和Melodyne(3.08美元),同时保持了相当的保留表达细微差别的能力。据我们所知,这是第一个利用音乐语言模型来实现符号音乐上下文的无参考音高校正的APC模型。BERT-APC的校正音频样本可在线获得。
摘要:Automatic Pitch Correction (APC) enhances vocal recordings by aligning pitch deviations with the intended musical notes. However, existing APC systems either rely on reference pitches, which limits their practical applicability, or employ simple pitch estimation algorithms that often fail to preserve expressiveness and naturalness. We propose BERT-APC, a novel reference-free APC framework that corrects pitch errors while maintaining the natural expressiveness of vocal performances. In BERT-APC, a novel stationary pitch predictor first estimates the perceived pitch of each note from the detuned singing voice. A context-aware note pitch predictor estimates the intended pitch sequence by leveraging a music language model repurposed to incorporate musical context. Finally, a note-level correction algorithm fixes pitch errors while preserving intentional pitch deviations for emotional expression. In addition, we introduce a learnable data augmentation strategy that improves the robustness of the music language model by simulating realistic detuning patterns. Compared to two recent singing voice transcription models, BERT-APC demonstrated superior performance in note pitch prediction, outperforming the second-best model, ROSVOT, by 10.49%p on highly detuned samples in terms of the raw pitch accuracy. In the MOS test, BERT-APC achieved the highest score of $4.32 \pm 0.15$, which is significantly higher than those of the widely-used commercial APC tools, AutoTune ($3.22 \pm 0.18$) and Melodyne ($3.08 \pm 0.18$), while maintaining a comparable ability to preserve expressive nuances. To the best of our knowledge, this is the first APC model that leverages a music language model to achieve reference-free pitch correction with symbolic musical context. The corrected audio samples of BERT-APC are available online.
【2】Mispronunciation Detection and Diagnosis Without Model Training: A Retrieval-Based Approach
标题:无需模型训练的发音错误检测和诊断:基于检索的方法
链接:https://arxiv.org/abs/2511.20107
摘要:发音错误检测和诊断(MDD)对于语言学习和言语治疗至关重要。与传统的方法,需要评分模型或训练音素级模型,我们提出了一种新的无训练的框架,利用检索技术与预训练的自动语音识别模型。我们的方法避免了特定于音素的建模或额外的特定于任务的训练,同时仍然实现了准确的检测和诊断发音错误。在L2-ARCTIC数据集上的实验表明,该方法在避免模型训练复杂性的同时,获得了69.60%的F1分数。
摘要:Mispronunciation Detection and Diagnosis (MDD) is crucial for language learning and speech therapy. Unlike conventional methods that require scoring models or training phoneme-level models, we propose a novel training-free framework that leverages retrieval techniques with a pretrained Automatic Speech Recognition model. Our method avoids phoneme-specific modeling or additional task-specific training, while still achieving accurate detection and diagnosis of pronunciation errors. Experiments on the L2-ARCTIC dataset show that our method achieves a superior F1 score of 69.60% while avoiding the complexity of model training.
【3】It Hears, It Sees too: Multi-Modal LLM for Depression Detection By Integrating Visual Understanding into Audio Language Models
标题:它听到了,它也看到了:通过将视觉理解集成到音频语言模型中来检测抑郁症的多模式LLM
链接:https://arxiv.org/abs/2511.19877
摘要:抑郁症是全球最普遍的心理健康疾病之一。近年来,语音、视频和文字记录等多模式数据越来越多地用于开发人工智能辅助抑郁症评估系统。大型语言模型由于其强大的语言理解和泛化能力,进一步推动了这一领域的发展。然而,传统的LLM仍然以文本为中心,无法处理音频和视觉形式中丰富的非语言线索,这些线索是心理健康评估的关键组成部分。虽然多模态LLM提供了一个有前途的方向,但很少有针对心理学应用的。在这项研究中,我们提出了一种新的多模态LLM框架抑郁症检测。我们的方法增强了音频语言模型与视觉理解,并在时间戳级别调整视听功能。这种细粒度对齐改进了跨模态的时间动态建模,同时减少了对大量训练数据和计算资源的需求。DAIC-WoZ数据集上的实验表明,我们的模型优于单模态方法和以前的多模态方法。此外,所提出的框架可以扩展到包含更多的生理信号,为精神健康以外的更广泛的临床应用铺平道路。
摘要:Depression is one of the most prevalent mental health disorders globally. In recent years, multi-modal data, such as speech, video, and transcripts, has been increasingly used to develop AI-assisted depression assessment systems. Large language models have further advanced this field due to their strong language understanding and generalization capabilities. However, conventional LLMs remain text-centric and cannot process the rich non-verbal cues found in audio and visual modalities, which are critical components in mental health evaluation. While multi-modal LLMs offer a promising direction, few are tailored for psychological applications. In this study, we propose a novel multi-modal LLM framework for depression detection. Our approach augments an audio language model with visual understanding and aligns audio-visual features at the timestamp level. This fine-grained alignment improves modeling of temporal dynamics across modalities while reducing the need for extensive training data and computational resources. Experiments on the DAIC-WoZ dataset demonstrate that our model outperforms both single-modality approaches and previous multi-modal methods. Moreover, the proposed framework can be extended to incorporate additional physiological signals, paving the way for broader clinical applications beyond mental health.
机器翻译由腾讯交互翻译提供,仅供参考
