微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
快速导航
1. 语音识别与关键词检测 1 篇
2. 语音合成与声音生成 1 篇
3. 语音增强、降噪与音频修复 1 篇
4. 音乐信息检索与音乐生成 4 篇
5. 数据集、基准与评测 5 篇
6. 安全、隐私与深度伪造音频 3 篇
7. 其他/综合语音音频 5 篇
1. 语音识别与关键词检测 | 1 篇
1. Transcript-Free Lightweight Detection of Alzheimer's Disease from Spontaneous Speech Using Handcrafted MFCC-Dominant Acoustic Biomarkers
使用手工制作的以MFCC为主的声学生物标志物从自发语音中进行无转录本的阿尔茨海默病轻量级检测
AI 总结:研究旨在从自发语音中无转录本检测阿尔茨海默病,利用痴呆症银行匹兹堡语料库,提取99个手工声学-时间特征,用WebRTC VAD分离语音,通过支持向量机评估,结果显示相关线索可促进独立于说话者的AD筛查,为部署研究奠定基础。
链接:https://arxiv.org/abs/2607.10168
作者:Rashin Gholijani Farahani, Azam Bastanfard
英文摘要:It is still hard to find Alzheimer's disease (AD) early, especially when neuroimaging is expensive or tools that depend on language are not available. Spontaneous speech provides a non-invasive signal; however, numerous current methodologies depend on transcripts/ASR or computationally intensive deep models. We offer a simple, audio-only baseline for detecting AD using 176 Cookie Theft recordings from the DementiaBank Pitt corpus (88 AD, 88 controls). WebRTC voice activity detection (VAD) is used to separate speech from non-speech. We take out 99 hand-crafted acoustic-temporal features, including pause and fluency statistics, spectral/prosodic descriptors, and MFCC summaries with Δ and ΔΔ. Evaluation is performed using a stringent speaker-independent GroupShuffleSplit,documenting performance across 30 iterations. A lightweight SVM with an RBF kernel gets an average AUC of 0.674 across runs. For example, a single split has an AUC of 0.742 and an accuracy of 0.657. We also present an exploratory compact-feature analysis utilizing a Top-20 subset ranked by Random Forest importance; since selection is not nested within training splits, these results may be overly optimistic and are not employed for primary conclusions (AUC 0.719). The results indicate that transcript-free spectro-temporal and fluency-related cues can facilitate speaker-independent Alzheimer's disease screening from raw audio, establishing a practical foundation for deployment-oriented research.
2. 语音合成与声音生成 | 1 篇
2. VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
VoxENES 2026:针对大语言模型时代文本转语音和语音转换的语音欺骗检测泛化基准测试
AI 总结:研究针对大语言模型时代TTS和VC系统合成语音与传统基准测试不匹配问题,引入VoxENES 2026基准测试,对八个预训练检测器测试发现性能大幅下降,凸显当前检测器问题,确立该基准为开发音频欺骗对策的实用平台。
链接:https://arxiv.org/abs/2607.11706
机构:University of South Florida(佛罗里达州立大学)
作者:Aastha Sharma, Guangjing Wang
英文摘要:Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world post-processing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish) benchmark of 53,628 audio samples generated using 10 contemporary speech synthesis methods and evaluated under 10 standardized post-processing conditions. Using VoxENES 2026, we benchmark eight pretrained detectors without fine-tuning and observe substantial performance degradation: the best model achieves 28.98\% EER overall, while most perform near or below random chance across modern generators and perturbations. Our results highlight the reliance on brittle artifacts in current detectors and establish VoxENES 2026 as a practical testbed for developing robust audio spoofing countermeasures.
3. 语音增强、降噪与音频修复 | 1 篇
3. Teaching Speech Enhancement Models to Sing: Domain Adaptation from Speech Enhancement to Singing Voice Separation
教语音增强模型唱歌:从语音增强到歌声分离的域适应
AI 总结:针对歌声分离训练数据有限问题,将其作为从语音增强到歌声分离的域适应,研究全量微调和LoRA参数高效微调策略,实验表明两种策略均有效提升歌声分离性能,LoRA微调还能保留语音增强能力,证明适配预训练模型是数据稀缺时训练歌声分离模型的有效策略。
链接:https://arxiv.org/abs/2607.11630
作者:Paul A. Bereuter, Mark D. Plumbley, Alois Sontacchi
英文摘要:State-of-the-art speech enhancement models benefit from large-scale labeled datasets, whereas singing voice separation models suffer from limited available training data. To address this limitation, we formulate singing voice separation as domain adaptation from speech enhancement to singing voice separation. We investigate two fine-tuning strategies: full fine-tuning and parameter-efficient fine-tuning using Low-Rank Adaptation (LoRA) on a discriminative and a generative model. Models with either adaptation strategy outperform the same architectures trained from scratch by 0.29-1.8 dB in Signal-to-Distortion-Ratio. Full fine-tuning yields the highest singing voice separation performance, but catastrophic forgetting degrades speech enhancement performance. LoRA fine-tuning achieves competitive singing voice separation performance while preserving the original speech enhancement capability with only 6-12% additional parameters compared to the base speech enhancement model. Furthermore, the generative model shows improved generalization to an unseen test set. The results demonstrate that adapting pretrained speech enhancement models is an effective strategy for training singing voice separation models in data-scarce scenarios.
4. 音乐信息检索与音乐生成 | 4 篇
4. Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment
利用未配对数据和对比对齐进行预训练的舞蹈音乐生成
AI 总结:针对舞蹈音乐生成中高质量配对数据稀缺问题,提出利用未配对和配对数据的框架,结合预训练单峰编码器、对比预训练及控制网络模块,实验证明该方法提升了舞蹈音乐对齐和音频质量,优于现有技术。
链接:https://arxiv.org/abs/2607.10537
机构:Sony Computer Science Laboratories(索尼计算机科学实验室); Keio University(庆应义塾大学); Georgia Institute of Technology(佐治亚理工学院)
作者:Ryota Kimura, Sangheon Park, Natalia Polouliakh, Taketo Akama
英文摘要:Dance-to-music generation is a promising task for applications such as choreography support and automatic accompaniment, where temporal coordination between body movement and sound is essential. In particular, using human joint positions as the motion representation is attractive because they explicitly capture body dynamics while being lightweight, privacy-preserving, and easy to integrate with motion capture and pose-estimation pipelines. A central challenge in this setting, however, is the scarcity of high-quality paired dance-music data, since collecting accurately synchronized pairs is costly and often constrained by copyright and performance rights. This makes it difficult to train end-to-end models solely from paired data. To address this issue, we propose a dance-conditioned music generation framework that efficiently exploits both unpaired and paired data. Our method combines pretrained unimodal encoders for motion and music, beat-guided contrastive pretraining to align their feature spaces, and a ControlNet-style conditioning module on top of a pretrained text-to-audio diffusion model. Experiments on AIST++ demonstrate that the proposed techniques improve both dance-music alignment and audio quality, as confirmed by quantitative and qualitative evaluations. Compared to a state-of-the-art method, our approach achieves superior dance alignment performance and competitive audio quality. Code is available at https://github.com/kmraven/AudioLDM-ControlNet .
5. MusicMark: A Robust Generative Watermarking Framework for Music Generation
MusicMark:一种用于音乐生成的鲁棒生成式水印框架
AI 总结:针对人工智能音乐生成中水印溯源问题,现有方法多为事后处理且易受攻击。本文提出MusicMark,在生成时将水印嵌入语义潜在空间,通过水印适配器和联合目标训练,实验证明其在多种攻击下优于基线,翻唱攻击中也更具鲁棒性。
链接:https://arxiv.org/abs/2607.11117
机构:Korea University(韩国大学); Inha University(仁荷大学)
作者:Seohwan Yun, Jeeyoung Yun, Yongjin Kim, Juyeon Lee, Sungwoong Kim
英文摘要:AI music generation has rapidly advanced alongside commercial platforms, raising the need for reliable watermarking for provenance and attribution. However, existing audio watermarking research has largely focused on speech, and applying speech-oriented methods to music is challenging due to music's complex structure and rich acoustic texture. Most existing methods are post-hoc, adding imperceptible perturbations after generation rather than embedding watermarks as part of the content. This makes them fragile under transformations and especially vulnerable to neural codec re-synthesis, which can discard imperceptible residual signals. Moreover, since generation and watermarking are decoupled, the watermarking step can be bypassed or omitted, weakening provenance guarantees. To address these issues, we propose MusicMark, which, to the best of our knowledge, is the first generative watermarking framework for music. Specifically, MusicMark embeds watermark messages into the semantic latent space during generation, incorporating the watermark as part of the musical content and ensuring robustness against diverse attacks, particularly neural codec re-synthesis. To this end, we introduce a watermark adapter into a diffusion-based generation model to embed watermark messages across denoising steps. The adapter and detector are trained with a joint objective that preserves fidelity by constraining watermarked latents close to their unwatermarked reference latents, while improving robustness through attack augmentations. Experiments demonstrate that MusicMark substantially outperforms post-hoc baselines across diverse attacks including neural codec re-synthesis, while maintaining comparable generation quality. We further introduce a cover-song attack, converting the singing voice while preserving musical content, and show that MusicMark remains more robust than post-hoc methods.
6. BeatEdit: Symbolic Music Generation as Explicit Editing
BeatEdit:作为显式编辑的符号音乐生成
AI 总结:研究针对符号音乐生成中缺乏选择性修改支持的问题,提出BeatEdit框架,基于BEAT编码,含三种互补机制,共享单一编码和预训练主干,在多项任务中精度和质量更高且高效,揭示编码设计对编辑有效性有重要影响。
链接:https://arxiv.org/abs/2607.11124
机构:School of Future Technology South China University of Technology Guangzhou China(未来技术学院 华南理工大学 广州 中国); South China University of Technology(华南理工大学); Nanjing University(南京大学)
作者:Haoyu Gu, Lekai Qian, Haowu Zhou, Qi Liu, Shuai Wang
英文摘要:Music creation is fundamentally a process of revision. Yet symbolic music generation remains dominated by paradigms that produce complete sequences from scratch, with limited support for selective modification. Edit-based methods have proven effective for text transformation tasks, but remain largely unexplored for symbolic music. We trace this absence to the representational level: conventional event-based music encodings lack the structural properties required by explicit music editing. In contrast, the BEAT encoding, a beat-grid-anchored representation originally designed for autoregressive generation, possesses structural properties amenable to editing. We propose BeatEdit, the first framework for symbolic music generation based on explicit edit operations, recasting generation as producing new content by editing a draft rather than synthesizing from scratch. BeatEdit comprises three complementary mechanisms along an axis of increasing edit density: per-token sequence tagging for error correction, iterative refinement for accompaniment editing, and tag-then-fill for segment completion. All these mechanisms share a single encoding and pre-trained backbone, achieving higher precision and perceptual quality than autoregressive and diffusion methods across all three tasks, while remaining efficient, with single-pass inference completing in under 100 ms. Cross-encoding evaluation further reveals that encoding design substantially influences editing effectiveness, with notable encoding-method interaction effects. Code is available at https://github.com/Haoyu-Gu/BeatEdit-code
7. Qwen-Music Technical Report
Qwen-音乐技术报告
AI 总结:介绍Qwen-音乐,它支持文本到音乐生成和翻唱歌曲生成两个核心任务。集成三个核心组件,通过特定机制建模并渲染。经多语言数据训练及多种训练方式,在多个音乐性和音频质量指标上达领先成果,获专业评估人员青睐。
链接:https://arxiv.org/abs/2607.11699
机构:Qwen Team(通义团队)
作者:Jin Xu, Kangdi Wang, Ruibin Yuan, Shun Lei, Xiong Wang, Xize Cheng, Xueyao Zhang, Yang Zhang, Yiheng Chen, Yongqi Wang, Yue Wang, Zhifang Guo, Zihan Liu, Zijian Lin, Dake Guo, Hangrui Hu, Lei Xie, Linhan Ma, Wei Xue, Wenxiang Guo, Xinfa Zhu, Xipin Wei, Yangze Li, Yuanjun Lv, Yuxuan Wang, Yunfei Chu, Zhiyong Wu
英文摘要:In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical and high-fidelity songs with complete vocal singing. Qwen-Music supports two core tasks: Text to Music Generation, which create entirely new songs from text descriptions, lyrics, and musical attributes, and Cover Song Generation, which reinterprets existing songs with different styles and vocal characteristics. Architecturally, Qwen-Music integrates three core components: Qwen-Music-Tokenizer, Qwen-Music-LLM, and Qwen-Music-Render. Qwen-Music-Tokenizer compresses audio into a 25 Hz single-codebook stream of Music Semantic Tokens that preserve semantic and melodic information for LLM prediction. Based on these tokens, Qwen-Music-LLM performs autoregressive music semantic modeling, with a key novelty being a melody-token-based chain-of-thought (Melody-CoT) mechanism that plans melodies before full-song generation, improving creativity, musicality, structural coherence, and reference-audio-based melody cloning. To overcome the fidelity limitations of discrete semantic tokens, Qwen-Music-Render performs generative stereo rendering, enriching acoustic details and producing high-fidelity stereo waveforms. Finally, we train Qwen-Music-LLM on more than 5 million hours of multilingual music data covering hundreds of languages. We first apply quality-aware pre-training curriculum, then use progressive post-training, comprising supervised initialization, offline DPO, and online GSPO, to further improve musicality and instruction-following ability. Across 600 Chinese and English prompts, Qwen-Music achieves state-of-the-art results in 13 of 16 objective musicality and audio-quality metrics. Professional evaluators also prefer Qwen-Music over leading proprietary systems. For cover song generation, Qwen-Music preserves reference melodies more accurately than leading proprietary systems.
5. 数据集、基准与评测 | 5 篇
8. A Production-Oriented Framework for Evaluation of SFX Generation
一种面向生产的声效生成评估框架
AI 总结:针对工业声音设计中声效生成评估问题,提出面向生产的评估框架,通过确定生产要求、两阶段协议及结合客观指标与人类研究,揭示不同基线的优势权衡,为声效变化评估建立协议并为设计音频生成管道提供基础。
链接:https://arxiv.org/abs/2607.09973
作者:Mélodie Desbos, Yara Bahram, Eric Granger, Mohammadhadi Shateri
英文摘要:Industrial sound design requires audio generation systems that not only produce realistic audio, but also preserve the perceptual identity of a reference, support controllable variation, and remain efficient for practical workflows. Existing evaluations are usually tied to text-to-audio (TTA), unconditional, or task-specific settings, limiting assessment for reference-guided sound effects (SFX) variation. To address this gap, we present a production-oriented evaluation framework for structured comparison of heterogeneous audio generation and editing methods. Our framework identifies nine production requirements and explicitly accounts for differences in model capabilities, enabling comparison under a common production objective. A two-stage protocol is introduced: (1) a reference-guided audio-to-audio (ATA) variation task, in which all methods are evaluated under the same ESC-50 SFX adaptation setup, and (2) capability-specific analyses of native operations such as SFX morphing, temporal and energy alignment, inpainting, and targeted editing. This framework combines objective metrics (including FAD, ImageBind-based reference alignment, and diversity across generated variants), together with a human study of perceptual identity preservation and transient diagnosis. Our study reveals complementary strengths and trade-offs across baselines for different production needs. Among the full-generation baselines evaluated under a shared ATA setting, AudioX provides the strongest overall trade-off between reference alignment and diversity while still supporting SFX morphing. Other baselines remain most suitable for specific editing operations. Our framework establishes a structured evaluation and decision protocol for reference-guided SFX variation and provides a practical basis for designing future unified industrial audio generation pipelines. Audio demos are on the accompanying web page.
9. Breaking the Quality--Intelligibility Trade-off in Streaming Target Speaker Extraction via Deep-Feature-Anchored Preference Optimization
通过深度特征锚定偏好优化打破流式目标说话人提取中的质量-可懂度权衡
AI 总结:研究流式目标说话人提取中的质量-可懂度权衡问题,提出扩大Conformer卷积核及基于WavLM的DPO微调策略,实现了相对可懂度10.9%的提升,同时音频质量和说话人相似度也有改善。
链接:https://arxiv.org/abs/2607.10191
机构:Tsinghua University(清华大学); The Chinese University of Hong Kong(香港中文大学); SenseTime(商汤科技)
作者:Shuhai Peng, Jinjiang Liu, Hui Lu, Liyang Chen, Guiping Zhong, Jiakui Li, Shiyin Kang, Zhiyong Wu
英文摘要:Generative streaming models for Target Speaker Extraction (TSE) commonly exhibit a quality--intelligibility trade-off, wherein naive optimization for perceptual audio quality tends to degrade speech intelligibility, and conversely. We reveal that this trade-off arises not from the constraints of streaming architectures, but from an inappropriate choice of optimization anchor. Directly optimizing against audio quality metrics induces catastrophic reward hacking, where content critical to pronunciation and intelligibility is systematically erased to maximize a proxy score. To break this bottleneck, we propose two complementary improvements: an enlarged Conformer convolution kernel for richer local spectro-temporal modeling, and WavLM-anchored Direct Preference Optimization (DPO) fine-tuning strategy. DPO preference pairs are ranked by WavLM cosine similarity, a deep acoustic feature encoding both phonetic structure and speaker identity, providing an optimization anchor that resists hacking. Under a 560 ms streaming chunk size, the proposed method achieves a 10.9% relative intelligibility improvement (word error rate: 0.138 to 0.123), with marginal simultaneous gains in audio quality and speaker similarity.
10. Graph Representation of RaagBase: A Unique Dataset for Hindustani Music
拉格基础的图形表示:一个独特的北印度音乐数据集
AI 总结:针对北印度音乐拉格聚类未充分探索的问题,提出基于Pt. Bhatkhande作品音符序列的拉格基础数据集,用图形表示拉格结构,通过音符频率分布体现作品相似性,应用聚类技术,实验验证了数据集及表示的有效性。
链接:https://arxiv.org/abs/2607.10229
机构:XIM University(西姆大学)
作者:Chandan Misra, Swarup Chattopadhyay
英文摘要:Raag classification is a fundamental MIR task for Hindustani Music, with applications in recommendation, education, archiving, and intelligent search. However, raag clustering remains underexplored, as most existing approaches rely on annotated audio or labeled datasets. While annotated melodic phrases capture characteristic patterns, complete note sequences preserve temporal structure and contextual dependencies, making them more suitable for data-driven modeling. In this work, we introduce RaagBase, a notation-based text dataset consisting of note sequences from compositions by Pt. Bhatkhande. Furthermore we propose a novel graph-based representation of raag structures by modeling the dominance and absence of notes in compositions. Each composition is represented as a node, and the edges between two compositions corresponds the similarities between them based on the note frequency distribution. Further, we apply established graph clustering techniques to identify groups of similar raag compositions. Experimental results demonstrate highly coherent clusters with strong agreement to ground-truth raag labels, thereby validating both the dataset and the proposed representation. The dataset is publicly available at https://anonymous.4open.science/r/RaagBase-5427.
11. MeloBottleneck: Self-Supervised Melody Skeleton Extraction with a Latent Subsequence Bottleneck
MeloBottleneck:基于潜在子序列瓶颈的自监督旋律骨架提取
AI 总结:研究旋律骨架提取问题,提出MeloBottleneck自监督框架,通过特定算子和训练方式生成旋律骨架,在多种模式下评估,结果显示该方法比伪标签模仿迁移性更强,还能提升基于BM25的片段检索效果。
链接:https://arxiv.org/abs/2607.10233
12. The SonicAGI System for the REAL-TSE Challenge
用于 REAL-TSE 挑战赛的 SonicAGI 系统
AI 总结:本文介绍 SonicAGI 参加 REAL-TSE 挑战赛的情况,采用以数据为中心的方法,在线引入 SwiftNet-Lookahead 控制延迟,离线用 USEF-TFGridNet 权衡质量与保真度,在官方评估中取得较好成绩,证明相关训练和建模对对话式 TSE 有效。
链接:https://arxiv.org/abs/2607.11083
6. 安全、隐私与深度伪造音频 | 3 篇
13. What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection
你所训练的就是你所得到的:音频深度伪造检测中的性别偏见、训练构成及事后缓解
AI 总结:研究音频深度伪造检测中的性别偏见,通过在不同性别构成训练集上训练特定攻击模型,使用ResNet18分类器及评估六种校准方法,发现训练数据构成影响偏差方向,事后方法只能部分缓解差异,强调性别公平需在训练时解决。
链接:https://arxiv.org/abs/2607.09891
14. PC-Mix: Partial-Component Audio Spoofing Detection under Mixed Speech and Environmental Sound Conditions
PC-Mix:混合语音和环境声音条件下的部分组件音频欺骗检测
AI 总结:研究针对部分音频欺骗检测,提出PC-Mix数据集,解决现有基准差距。构建含真实与部分欺骗环境声音组件并与语音信号混合的音频,建立评估协议与联合学习框架,实验表明匹配条件下训练检测欺骗更有效。
链接:https://arxiv.org/abs/2607.10345
15. Evidence Subspace Projection: Measuring How Much Evidence Explains Deepfake Detection in Self-Supervised Speech Models
证据子空间投影:衡量自我监督语音模型中多少证据可解释深度伪造检测
AI 总结:研究如何将自我监督学习模型用于音频深度伪造检测,提出证据子空间投影方法,通过在共享空间表示证据因素和真实性标签,投影决策向量量化证据解释力,经多数据集多设置评估,验证方法并获新见解。
链接:https://arxiv.org/abs/2607.11538
7. 其他/综合语音音频 | 5 篇
16. ARIMA: Reconstruction-Grounded Predictive Representation Learning for Symbolic Music
ARIMA:用于符号音乐的基于重构的预测表示学习
AI 总结:本文提出ARIMA框架,用于符号音乐的基于重构的预测表示学习。它直接从数据学习紧凑窗口表示,编码窗口为潜在表示,训练因果预测器并通过结构化重构确定编码器基础,在多下游任务表现良好,证明相关方法的有效性和重要性。
链接:https://arxiv.org/abs/2607.10003
17. Local Multimodal Music Alignment from Global Supervision
基于全局监督的局部多模态音乐对齐
AI 总结:研究如何利用全局监督实现局部多模态音乐对齐,提出FuSiLi方法,通过基于Sinkhorn的软对齐操作局部特征,结合传统全局相似性微调预训练编码器,在跨模态检索和帧级对齐任务中表现出色,优于多种基线。
链接:https://arxiv.org/abs/2607.10023
18. CHARM: Charge Calibration and Acoustic Rescue for LLM-based Multimodal Sarcasm Detection
CHARM:基于大语言模型的多模态讽刺检测中的电荷校准与声学救援
AI 总结:研究针对大语言模型在讽刺检测中过度预测积极类别及韵律线索利用不足的问题,提出CHARM框架,含双向电荷校准和声学后期融合救援两个模块,无需微调主干,提升检测性能,揭示跨文化韵律解耦,产生可解释的跨语言多模态检测器。
链接:https://arxiv.org/abs/2607.11102
19. Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance
Anysynth:通过上下文学习和非对称分层引导实现零样本乐器克隆
AI 总结:研究旨在实现零样本乐器克隆,提出基于上下文流匹配的无嵌入神经合成器Anysynth,通过特定条件设定让模型动态检索声学细节,实验显示其性能优越,还提出非对称分层CFG优化可控性,推动了零样本乐器克隆发展。
链接:https://arxiv.org/abs/2607.11143
20. Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models
大型音频语言模型中用于声学感知的编码器端神经元识别与增强
AI 总结:研究针对大型音频语言模型在语音非语义属性表现不佳的问题,提出IAAN方法,通过对比音频编码器神经元在真实波形与噪声参考上的激活来评分并增强高分神经元,有效提升了模型在多数据集上的声学感知准确率,开辟了新的推理时改进方向。
链接:https://arxiv.org/abs/2607.11801
