今日论文合集:cs.SD语音6篇,eess.AS音频处理5篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Evaluating Identity Leakage in Speaker De-Identification Systems
标题:说话人去识别系统中的身份泄露评估
链接:https://arxiv.org/abs/2508.14012

作者:Seungmin Seo, Oleg Aulov, Afzal Godil, Kevin Mangold
备注:Submitted to ICASSP 2026
摘要:说话人去识别的目的是隐藏说话人的身份,同时保持底层语音的可理解性。我们引入了一个基准,量化残留的身份泄漏与三个互补的错误率:相等的错误率,累积匹配特征命中率,嵌入空间相似性通过典型相关分析和Procrustes分析测量。评估结果表明,所有国家的最先进的说话人去识别系统泄漏身份信息。在我们的评估中,性能最高的系统的性能仅略好于随机猜测,而性能最低的系统在基于CMC的前50名候选人中达到了45%的命中率。这些发现突出了当前说话人去识别技术中持续存在的隐私风险。
摘要:Speaker de-identification aims to conceal a speaker's identity while preserving intelligibility of the underlying speech. We introduce a benchmark that quantifies residual identity leakage with three complementary error rates: equal error rate, cumulative match characteristic hit rate, and embedding-space similarity measured via canonical correlation analysis and Procrustes analysis. Evaluation results reveal that all state-of-the-art speaker de-identification systems leak identity information. The highest performing system in our evaluation performs only slightly better than random guessing, while the lowest performing system achieves a 45% hit rate within the top 50 candidates based on CMC. These findings highlight persistent privacy risks in current speaker de-identification technologies.


【2】DegDiT: Controllable Audio Generation with Dynamic Event Graph Guided Diffusion Transformer
标题:DegDiT:使用动态事件图引导扩散Transformer的可控音频生成
链接:https://arxiv.org/abs/2508.13786

作者: Yisu Liu, Chenxing Li, Wanqian Zhang, Wenfu Wang, Meng Yu, Ruibo Fu, Zheng Lin, Weiping Wang, Dong Yu
摘要:可控文本到音频生成的目的是从文本描述合成音频,同时满足用户指定的约束,包括事件类型,时间序列,以及开始和偏移时间戳。这使得能够精确控制所生成的音频的内容和时间结构。尽管最近取得了进展,但现有方法仍然面临准确时间定位、开放词汇表可扩展性和实际效率之间的固有权衡。为了解决这些挑战,我们提出了DegDiT,这是一种新颖的动态事件图引导扩散Transformer框架,用于开放词汇表可控音频生成。DegDiT将描述中的事件编码为结构化动态图。每个图中的节点被设计为表示三个方面:语义特征、时间属性和事件间的连接。一个图形Transformer被用来整合这些节点,并产生情境化的事件嵌入,作为扩散模型的指导。为了确保高质量和多样化的训练数据,我们引入了一个质量平衡的数据选择管道,将分层事件注释与多标准质量评分相结合,从而产生具有语义多样性的精选数据集。此外,我们提出了共识偏好优化,通过多个奖励信号之间的共识促进音频生成。在AudioCondition、DESED和AudioTime数据集上进行的大量实验表明,DegDiT在各种客观和主观评估指标上都达到了最先进的性能。
摘要:Controllable text-to-audio generation aims to synthesize audio from textual descriptions while satisfying user-specified constraints, including event types, temporal sequences, and onset and offset timestamps. This enables precise control over both the content and temporal structure of the generated audio. Despite recent progress, existing methods still face inherent trade-offs among accurate temporal localization, open-vocabulary scalability, and practical efficiency. To address these challenges, we propose DegDiT, a novel dynamic event graph-guided diffusion transformer framework for open-vocabulary controllable audio generation. DegDiT encodes the events in the description as structured dynamic graphs. The nodes in each graph are designed to represent three aspects: semantic features, temporal attributes, and inter-event connections. A graph transformer is employed to integrate these nodes and produce contextualized event embeddings that serve as guidance for the diffusion model. To ensure high-quality and diverse training data, we introduce a quality-balanced data selection pipeline that combines hierarchical event annotation with multi-criteria quality scoring, resulting in a curated dataset with semantic diversity. Furthermore, we present consensus preference optimization, facilitating audio generation through consensus among multiple reward signals. Extensive experiments on AudioCondition, DESED, and AudioTime datasets demonstrate that DegDiT achieves state-of-the-art performances across a variety of objective and subjective evaluation metrics.


【3】Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
标题:利用Mamba和全脸视觉进行视听语音增强
链接:https://arxiv.org/abs/2508.13624

作者:Rong Chao, Wenze Ren, You-Jin Li, Kuo-Hsuan Hung, Sung-Feng Huang, Szu-Wei Fu, Wen-Huang Cheng, Yu Tsao
备注:Accepted to Interspeech 2025 Workshop
摘要:最近的基于Mamba的模型已经显示出在语音增强的承诺,有效地建模长距离的时间依赖性。然而,像语音增强Mamba(SEamba)这样的模型仍然局限于单扬声器场景,并且在复杂的多扬声器环境中挣扎,例如鸡尾酒会问题。为了克服这一点,我们引入AVSEMamba,视听语音增强模型,集成了全脸视觉线索与基于Mamba的时间骨干。通过利用时空视觉信息,AVSEMamba能够在具有挑战性的条件下更准确地提取目标语音。在AVSEC-4挑战开发和盲测集上进行评估后,AVSEMamba在语音清晰度(STOI)、感知质量(PESQ)和非侵入性质量(UTMOS)方面优于其他单耳基线,并在单耳排行榜上获得\textbf{第一名}。
摘要:Recent Mamba-based models have shown promise in speech enhancement by efficiently modeling long-range temporal dependencies. However, models like Speech Enhancement Mamba (SEMamba) remain limited to single-speaker scenarios and struggle in complex multi-speaker environments such as the cocktail party problem. To overcome this, we introduce AVSEMamba, an audio-visual speech enhancement model that integrates full-face visual cues with a Mamba-based temporal backbone. By leveraging spatiotemporal visual information, AVSEMamba enables more accurate extraction of target speech in challenging conditions. Evaluated on the AVSEC-4 Challenge development and blind test sets, AVSEMamba outperforms other monaural baselines in speech intelligibility (STOI), perceptual quality (PESQ), and non-intrusive quality (UTMOS), and achieves \textbf{1st place} on the monaural leaderboard.


【4】Is Transfer Learning Necessary for Violin Transcription?
标题:小提琴转录需要迁移学习吗?
链接:https://arxiv.org/abs/2508.13516

作者:Yueh-Po Peng, Ting-Kang Wang, Li Su, Vincent K.M. Cheung
摘要:自动音乐转录(AMT)在钢琴等乐器上取得了显着进展,这主要归功于大规模,高质量数据集的可用性。相比之下,小提琴AMT由于有限的注释数据而仍然未被充分探索。一种常见的方法是为其他下游任务微调预训练模型,但在存在音色和发音差异的情况下,这种转移的有效性仍不清楚。在这项工作中,我们研究了在中等规模的小提琴数据集上从头开始训练是否可以匹配微调的钢琴预训练模型的性能。我们采用了一个钢琴转录架构,没有修改和训练它的MOSA数据集,其中包含约30小时的对齐小提琴录音。我们在URMP和Bach 10上的实验表明,与微调的模型相比,从头开始训练的模型具有竞争力甚至更好的性能。这些研究结果表明,强大的小提琴AMT是可能的,而不依赖于预先训练的钢琴表示,突出了仪器特定的数据收集和增强策略的重要性。
摘要:Automatic music transcription (AMT) has achieved remarkable progress for instruments such as the piano, largely due to the availability of large-scale, high-quality datasets. In contrast, violin AMT remains underexplored due to limited annotated data. A common approach is to fine-tune pretrained models for other downstream tasks, but the effectiveness of such transfer remains unclear in the presence of timbral and articulatory differences. In this work, we investigate whether training from scratch on a medium-scale violin dataset can match the performance of fine-tuned piano-pretrained models. We adopt a piano transcription architecture without modification and train it on the MOSA dataset, which contains about 30 hours of aligned violin recordings. Our experiments on URMP and Bach10 show that models trained from scratch achieved competitive or even superior performance compared to fine-tuned counterparts. These findings suggest that strong violin AMT is possible without relying on pretrained piano representations, highlighting the importance of instrument-specific data collection and augmentation strategies.


【5】MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
标题:MMAU-Pro:音频通用智能整体评估的挑战性和全面基准
链接:https://arxiv.org/abs/2508.13992

作者:Sonal Kumar, Šimon Sedláček, Vaibhavi Lokegaonkar, Fernando López, Wenyi Yu, Nishit Anand, Hyeonggon Ryu, Lichang Chen, Maxim Plička, Miroslav Hlaváček, William Fineas Ellingwood, Sathvik Udupa, Siyuan Hou, Allison Ferner, Sara Barahona, Cecilia Bolaños, Satish Rahi, Laura Herrera-Alarcón, Satvik Dixit, Siddhi Patil, Soham Deshmukh, Lasha Koroshinadze, Yao Liu, Leibny Paola Garcia Perera, Eleni Zanou, Themos Stafylakis, Joon Son Chung, David Harwath, Chao Zhang, Dinesh Manocha, Alicia Lozano-Diez, Santosh Kesiraju, Sreyan Ghosh, Ramani Duraiswami
摘要:音频理解包括语音、非语音声音和音乐对于实现人类水平的智能至关重要。因此,人工智能代理必须表现出整体的音频理解能力,才能被视为普遍智能。然而,全面评估听觉智力仍然具有挑战性。为了解决这一差距,我们推出了MMAU-Pro,这是用于评估AI系统中音频智能的最全面、最严格的基准。MMAU-Pro包含5,305个实例,每个实例都有一个或多个音频与人类专家生成的问答对配对,包括语音,声音,音乐及其组合。与现有的基准不同,MMAU-Pro评估了49种独特技能和多个复杂维度的听觉智能,包括长形式音频理解,空间音频推理,多音频理解等。所有的问题都经过精心设计,需要深思熟虑的多跳推理,包括多项选择和开放式回答格式。重要的是,音频数据直接来源于“野生”,而不是来自已知分布的现有数据集。我们评估了22个领先的开源和专有多模态AI模型,揭示了显著的局限性:即使是最先进的模型,如Gemini 2.5 Flash和Audio Flamingo 3,也只能分别达到59.2%和51.7%的准确率,在多个类别中接近随机性能。我们广泛的分析突出了具体的缺点,并提供了新的见解,为社区提供了可操作的观点,以增强未来AI系统向音频通用智能的发展。基准测试和代码可在https://sonalkum.github.io/mmau-pro上获得。
摘要:Audio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challenging. To address this gap, we introduce MMAU-Pro, the most comprehensive and rigorously curated benchmark for assessing audio intelligence in AI systems. MMAU-Pro contains 5,305 instances, where each instance has one or more audios paired with human expert-generated question-answer pairs, spanning speech, sound, music, and their combinations. Unlike existing benchmarks, MMAU-Pro evaluates auditory intelligence across 49 unique skills and multiple complex dimensions, including long-form audio comprehension, spatial audio reasoning, multi-audio understanding, among others. All questions are meticulously designed to require deliberate multi-hop reasoning, including both multiple-choice and open-ended response formats. Importantly, audio data is sourced directly ``from the wild" rather than from existing datasets with known distributions. We evaluate 22 leading open-source and proprietary multimodal AI models, revealing significant limitations: even state-of-the-art models such as Gemini 2.5 Flash and Audio Flamingo 3 achieve only 59.2% and 51.7% accuracy, respectively, approaching random performance in multiple categories. Our extensive analysis highlights specific shortcomings and provides novel insights, offering actionable perspectives for the community to enhance future AI systems' progression toward audio general intelligence. The benchmark and code is available at https://sonalkum.github.io/mmau-pro.


【6】End-to-End Audio-Visual Learning for Cochlear Implant Sound Coding in Noisy Environments
标题:噪音环境中用于人工晶状体植入物声音编码的端到端视听学习
链接:https://arxiv.org/abs/2508.13576

作者: Meng-Ping Lin, Enoch Hsin-Ho Huang, Shao-Yi Chien, Yu Tsao
备注:6 pages, 4 figures
摘要:人工耳蜗(CI)是一种卓越的生物医学设备,其成功地使患有严重至极重度听力损失的个体通过将语音转换为电刺激信号来感知声音。尽管最近的CI系统的性能有所进步,但在嘈杂或混响条件下的语音理解仍然是一个挑战。深度学习领域最近和正在进行的发展揭示了增强CI声音编码能力的良好机会,不仅通过使用神经网络复制传统信号处理方法,而且通过将视觉线索集成为多模式语音处理的辅助数据。因此,本文介绍了一种新型的噪声抑制CI系统AVSE-ECS,该系统利用视听语音增强(AVSE)模型作为基于深度学习的ElectrodeNet-CS(ECS)声音编码策略的预处理模块。具体而言,联合训练方法被应用到模型AVSE-ECS,一个端到端的CI系统。实验结果表明,该方法优于以前的ECS策略在噪声条件下,提高了客观语音可懂度分数。本研究中的方法和发现证明了使用深度学习将AVSE模块集成到端到端CI系统中的可行性和潜力
摘要:The cochlear implant (CI) is a remarkable biomedical device that successfully enables individuals with severe-to-profound hearing loss to perceive sound by converting speech into electrical stimulation signals. Despite advancements in the performance of recent CI systems, speech comprehension in noisy or reverberant conditions remains a challenge. Recent and ongoing developments in deep learning reveal promising opportunities for enhancing CI sound coding capabilities, not only through replicating traditional signal processing methods with neural networks, but also through integrating visual cues as auxiliary data for multimodal speech processing. Therefore, this paper introduces a novel noise-suppressing CI system, AVSE-ECS, which utilizes an audio-visual speech enhancement (AVSE) model as a pre-processing module for the deep-learning-based ElectrodeNet-CS (ECS) sound coding strategy. Specifically, a joint training approach is applied to model AVSE-ECS, an end-to-end CI system. Experimental results indicate that the proposed method outperforms the previous ECS strategy in noisy conditions, with improved objective speech intelligibility scores. The methods and findings in this study demonstrate the feasibility and potential of using deep learning to integrate the AVSE module into an end-to-end CI system


eess.AS音频处理


【1】MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
标题:MMAU-Pro:音频通用智能整体评估的挑战性和全面基准
链接:https://arxiv.org/abs/2508.13992

作者:Sonal Kumar, Šimon Sedláček, Vaibhavi Lokegaonkar, Fernando López, Wenyi Yu, Nishit Anand, Hyeonggon Ryu, Lichang Chen, Maxim Plička, Miroslav Hlaváček, William Fineas Ellingwood, Sathvik Udupa, Siyuan Hou, Allison Ferner, Sara Barahona, Cecilia Bolaños, Satish Rahi, Laura Herrera-Alarcón, Satvik Dixit, Siddhi Patil, Soham Deshmukh, Lasha Koroshinadze, Yao Liu, Leibny Paola Garcia Perera, Eleni Zanou, Themos Stafylakis, Joon Son Chung, David Harwath, Chao Zhang, Dinesh Manocha, Alicia Lozano-Diez, Santosh Kesiraju, Sreyan Ghosh, Ramani Duraiswami
摘要:音频理解包括语音、非语音声音和音乐对于实现人类水平的智能至关重要。因此,人工智能代理必须表现出整体的音频理解能力,才能被视为普遍智能。然而,全面评估听觉智力仍然具有挑战性。为了解决这一差距,我们推出了MMAU-Pro,这是用于评估AI系统中音频智能的最全面、最严格的基准。MMAU-Pro包含5,305个实例,每个实例都有一个或多个音频与人类专家生成的问答对配对,包括语音,声音,音乐及其组合。与现有的基准不同,MMAU-Pro评估了49种独特技能和多个复杂维度的听觉智能,包括长形式音频理解,空间音频推理,多音频理解等。所有的问题都经过精心设计,需要深思熟虑的多跳推理,包括多项选择和开放式回答格式。重要的是,音频数据直接来源于“野生”,而不是来自已知分布的现有数据集。我们评估了22个领先的开源和专有多模态AI模型,揭示了显著的局限性:即使是最先进的模型,如Gemini 2.5 Flash和Audio Flamingo 3,也只能分别达到59.2%和51.7%的准确率,在多个类别中接近随机性能。我们广泛的分析突出了具体的缺点,并提供了新的见解,为社区提供了可操作的观点,以增强未来AI系统向音频通用智能的发展。基准测试和代码可在https://sonalkum.github.io/mmau-pro上获得。
摘要:Audio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challenging. To address this gap, we introduce MMAU-Pro, the most comprehensive and rigorously curated benchmark for assessing audio intelligence in AI systems. MMAU-Pro contains 5,305 instances, where each instance has one or more audios paired with human expert-generated question-answer pairs, spanning speech, sound, music, and their combinations. Unlike existing benchmarks, MMAU-Pro evaluates auditory intelligence across 49 unique skills and multiple complex dimensions, including long-form audio comprehension, spatial audio reasoning, multi-audio understanding, among others. All questions are meticulously designed to require deliberate multi-hop reasoning, including both multiple-choice and open-ended response formats. Importantly, audio data is sourced directly ``from the wild" rather than from existing datasets with known distributions. We evaluate 22 leading open-source and proprietary multimodal AI models, revealing significant limitations: even state-of-the-art models such as Gemini 2.5 Flash and Audio Flamingo 3 achieve only 59.2% and 51.7% accuracy, respectively, approaching random performance in multiple categories. Our extensive analysis highlights specific shortcomings and provides novel insights, offering actionable perspectives for the community to enhance future AI systems' progression toward audio general intelligence. The benchmark and code is available at https://sonalkum.github.io/mmau-pro.


【2】End-to-End Audio-Visual Learning for Cochlear Implant Sound Coding in Noisy Environments
标题:噪音环境中用于人工晶状体植入物声音编码的端到端视听学习
链接:https://arxiv.org/abs/2508.13576

作者:Meng-Ping Lin, Enoch Hsin-Ho Huang, Shao-Yi Chien, Yu Tsao
备注:6 pages, 4 figures
摘要:人工耳蜗(CI)是一种卓越的生物医学设备,其成功地使患有严重至极重度听力损失的个体通过将语音转换为电刺激信号来感知声音。尽管最近的CI系统的性能有所进步,但在嘈杂或混响条件下的语音理解仍然是一个挑战。深度学习领域最近和正在进行的发展揭示了增强CI声音编码能力的良好机会,不仅通过使用神经网络复制传统信号处理方法,而且通过将视觉线索集成为多模式语音处理的辅助数据。因此,本文介绍了一种新型的噪声抑制CI系统AVSE-ECS,该系统利用视听语音增强(AVSE)模型作为基于深度学习的ElectrodeNet-CS(ECS)声音编码策略的预处理模块。具体而言,联合训练方法被应用到模型AVSE-ECS,一个端到端的CI系统。实验结果表明,该方法优于以前的ECS策略在噪声条件下,提高了客观语音可懂度分数。本研究中的方法和发现证明了使用深度学习将AVSE模块集成到端到端CI系统中的可行性和潜力
摘要:The cochlear implant (CI) is a remarkable biomedical device that successfully enables individuals with severe-to-profound hearing loss to perceive sound by converting speech into electrical stimulation signals. Despite advancements in the performance of recent CI systems, speech comprehension in noisy or reverberant conditions remains a challenge. Recent and ongoing developments in deep learning reveal promising opportunities for enhancing CI sound coding capabilities, not only through replicating traditional signal processing methods with neural networks, but also through integrating visual cues as auxiliary data for multimodal speech processing. Therefore, this paper introduces a novel noise-suppressing CI system, AVSE-ECS, which utilizes an audio-visual speech enhancement (AVSE) model as a pre-processing module for the deep-learning-based ElectrodeNet-CS (ECS) sound coding strategy. Specifically, a joint training approach is applied to model AVSE-ECS, an end-to-end CI system. Experimental results indicate that the proposed method outperforms the previous ECS strategy in noisy conditions, with improved objective speech intelligibility scores. The methods and findings in this study demonstrate the feasibility and potential of using deep learning to integrate the AVSE module into an end-to-end CI system


【3】Rapidly Adapting to New Voice Spoofing: Few-Shot Detection of Synthesized Speech Under Distribution Shifts
标题:快速适应新的语音欺骗:分布漂移下合成语音的Few-Shot检测
链接:https://arxiv.org/abs/2508.13320

作者:Ashi Garg, Zexin Cai, Henry Li Xinyuan, Leibny Paola García-Perera, Kevin Duh, Sanjeev Khudanpur, Matthew Wiesner, Nicholas Andrews
摘要:我们解决了在分布变化下检测合成语音的挑战-来自于看不见的合成方法,扬声器,语言或音频条件-相对于训练数据。Few-Shot学习方法是一种很有前途的方法,可以通过在少量分布样本的基础上快速适应来解决分布变化。我们提出了一个自我关注的原型网络,使更强大的Few-Shot适应。为了评估我们的方法,我们系统地比较了传统的zero-shot检测器和建议的Few-Shot检测器的性能,仔细控制训练条件,在评估时引入分布偏移。在分布偏移妨碍zero-shot性能的情况下,我们提出的Few-Shot自适应技术可以使用10个分布样本快速适应-在日语中实现高达32%的相对EER降低,在ASVspoof 2021 Deepfake数据集上实现20%的相对EER降低。
摘要:We address the challenge of detecting synthesized speech under distribution shifts -- arising from unseen synthesis methods, speakers, languages, or audio conditions -- relative to the training data. Few-shot learning methods are a promising way to tackle distribution shifts by rapidly adapting on the basis of a few in-distribution samples. We propose a self-attentive prototypical network to enable more robust few-shot adaptation. To evaluate our approach, we systematically compare the performance of traditional zero-shot detectors and the proposed few-shot detectors, carefully controlling training conditions to introduce distribution shifts at evaluation time. In conditions where distribution shifts hamper the zero-shot performance, our proposed few-shot adaptation technique can quickly adapt using as few as 10 in-distribution samples -- achieving upto 32% relative EER reduction on deepfakes in Japanese language and 20% relative reduction on ASVspoof 2021 Deepfake dataset.


【4】Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
标题:利用Mamba和全脸视觉进行视听语音增强
链接:https://arxiv.org/abs/2508.13624

作者:Rong Chao, Wenze Ren, You-Jin Li, Kuo-Hsuan Hung, Sung-Feng Huang, Szu-Wei Fu, Wen-Huang Cheng, Yu Tsao
备注:Accepted to Interspeech 2025 Workshop
摘要:最近的基于Mamba的模型已经显示出在语音增强的承诺,有效地建模长距离的时间依赖性。然而,像语音增强Mamba(SEamba)这样的模型仍然局限于单扬声器场景,并且在复杂的多扬声器环境中挣扎,例如鸡尾酒会问题。为了克服这一点,我们引入AVSEMamba,视听语音增强模型,集成了全脸视觉线索与基于Mamba的时间骨干。通过利用时空视觉信息,AVSEMamba能够在具有挑战性的条件下更准确地提取目标语音。在AVSEC-4挑战开发和盲测集上进行评估后,AVSEMamba在语音清晰度(STOI)、感知质量(PESQ)和非侵入性质量(UTMOS)方面优于其他单耳基线,并在单耳排行榜上获得\textbf{第一名}。
摘要:Recent Mamba-based models have shown promise in speech enhancement by efficiently modeling long-range temporal dependencies. However, models like Speech Enhancement Mamba (SEMamba) remain limited to single-speaker scenarios and struggle in complex multi-speaker environments such as the cocktail party problem. To overcome this, we introduce AVSEMamba, an audio-visual speech enhancement model that integrates full-face visual cues with a Mamba-based temporal backbone. By leveraging spatiotemporal visual information, AVSEMamba enables more accurate extraction of target speech in challenging conditions. Evaluated on the AVSEC-4 Challenge development and blind test sets, AVSEMamba outperforms other monaural baselines in speech intelligibility (STOI), perceptual quality (PESQ), and non-intrusive quality (UTMOS), and achieves \textbf{1st place} on the monaural leaderboard.


【5】Is Transfer Learning Necessary for Violin Transcription?
标题:小提琴转录需要迁移学习吗?
链接:https://arxiv.org/abs/2508.13516

作者:Yueh-Po Peng, Ting-Kang Wang, Li Su, Vincent K.M. Cheung
摘要:自动音乐转录(AMT)在钢琴等乐器上取得了显着进展,这主要归功于大规模,高质量数据集的可用性。相比之下,小提琴AMT由于有限的注释数据而仍然未被充分探索。一种常见的方法是为其他下游任务微调预训练模型,但在存在音色和发音差异的情况下,这种转移的有效性仍不清楚。在这项工作中,我们研究了在中等规模的小提琴数据集上从头开始训练是否可以匹配微调的钢琴预训练模型的性能。我们采用了一个钢琴转录架构,没有修改和训练它的MOSA数据集,其中包含约30小时的对齐小提琴录音。我们在URMP和Bach 10上的实验表明,与微调的模型相比,从头开始训练的模型具有竞争力甚至更好的性能。这些研究结果表明,强大的小提琴AMT是可能的,而不依赖于预先训练的钢琴表示,突出了仪器特定的数据收集和增强策略的重要性。
摘要:Automatic music transcription (AMT) has achieved remarkable progress for instruments such as the piano, largely due to the availability of large-scale, high-quality datasets. In contrast, violin AMT remains underexplored due to limited annotated data. A common approach is to fine-tune pretrained models for other downstream tasks, but the effectiveness of such transfer remains unclear in the presence of timbral and articulatory differences. In this work, we investigate whether training from scratch on a medium-scale violin dataset can match the performance of fine-tuned piano-pretrained models. We adopt a piano transcription architecture without modification and train it on the MOSA dataset, which contains about 30 hours of aligned violin recordings. Our experiments on URMP and Bach10 show that models trained from scratch achieved competitive or even superior performance compared to fine-tuned counterparts. These findings suggest that strong violin AMT is possible without relying on pretrained piano representations, highlighting the importance of instrument-specific data collection and augmentation strategies.


机器翻译由腾讯交互翻译提供,仅供参考