今日论文合集:cs.SD语音3篇,eess.AS音频处理2篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Towards Practical Automatic Piano Reduction using BERT with Semi-supervised Learning
标题:基于BERT和半监督学习的实用钢琴自动降阶
链接:https://arxiv.org/abs/2512.21324

作者:Wan Ki Wong,Ka Ho To,Chuck-jee Chau,Lucas Wong,Kevin Y. Yip,Irwin King
摘要:在这项研究中,我们提出了一种新的自动钢琴还原方法与半监督机器学习。钢琴还原是一个重要的音乐转化过程,它有助于音乐家和作曲家将音乐素描作为一种形式进行演奏和分析。这种自动化是一个极具挑战性的研究问题,但可以带来巨大的便利,因为手动进行钢琴还原需要大量的时间和精力。虽然监督机器学习通常是学习输入-输出映射的有用工具,但很难获得大量的标记数据。我们的目标是通过利用半监督学习来解决这个问题,这样就可以利用古典音乐中丰富的可用数据来执行任务,而很少或根本没有标记工作。在这方面,我们制定了一个两步的方法,音乐的简化,其次是协调。我们进一步提出并实现了两种可能的解决方案,利用现有的机器学习框架- MidiBERT。我们表明,我们的解决方案可以输出实际和现实的样本,准确的减少,只需要在后处理小的调整。我们的研究为半监督学习在自动钢琴还原中的应用奠定了基础,未来的研究人员可以参考产生更多最先进的结果。
摘要:In this study, we present a novel automatic piano reduction method with semi-supervised machine learning. Piano reduction is an important music transformation process, which helps musicians and composers as a musical sketch for performances and analysis. The automation of such is a highly challenging research problem but could bring huge conveniences as manually doing a piano reduction takes a lot of time and effort. While supervised machine learning is often a useful tool for learning input-output mappings, it is difficult to obtain a large quantity of labelled data. We aim to solve this problem by utilizing semi-supervised learning, so that the abundant available data in classical music can be leveraged to perform the task with little or no labelling effort. In this regard, we formulate a two-step approach of music simplification followed by harmonization. We further propose and implement two possible solutions making use of an existing machine learning framework -- MidiBERT. We show that our solutions can output practical and realistic samples with an accurate reduction that needs only small adjustments in post-processing. Our study forms the groundwork for the use of semi-supervised learning in automatic piano reduction, where future researchers can take reference to produce more state-of-the-art results.


【2】Foundation Model-based Evaluation of Neuropsychiatric Disorders: A Lifespan-Inclusive, Multi-Modal, and Multi-Lingual Study
标题:基于基础模型的神经精神疾病评估:一项涵盖一生、多模式和多语言的研究
链接:https://arxiv.org/abs/2512.20948

作者:Zhongren Dong,Haotian Guo,Weixiang Xu,Huan Zhao,Zixing Zhang
摘要:神经精神障碍,如阿尔茨海默病(AD),抑郁症和自闭症谱系障碍(ASD),其特征在于语言和声学异常,为早期检测提供了潜在的生物标志物。尽管多模式方法有希望,但多语言泛化和缺乏统一评估框架等挑战仍然存在。为了解决这些差距,我们提出了FEND(基于基础模型的神经精神疾病评估),这是一个综合的多模态框架,集成了语音和文本模态,用于在整个生命周期中检测AD,抑郁症和ASD。利用13种多语言数据集,包括英语,中文,希腊语,法语和荷兰语,我们系统地评估了多模态融合性能。我们的研究结果表明,多模态融合在AD和抑郁症检测中表现出色,但由于数据集的异质性,在ASD中表现不佳。我们还确定模态不平衡作为一个普遍的问题,多模态融合未能超过最好的单模态模型。跨语料库的实验显示强大的性能在任务和语言一致的情况下,但在多语言和任务异构设置显着退化。通过提供广泛的基准和性能影响因素的详细分析,FEND推进了自动化,寿命包容性和多语言神经精神障碍评估领域。我们鼓励研究人员采用FEND框架进行公平比较和可重复研究。
摘要:Neuropsychiatric disorders, such as Alzheimer's disease (AD), depression, and autism spectrum disorder (ASD), are characterized by linguistic and acoustic abnormalities, offering potential biomarkers for early detection. Despite the promise of multi-modal approaches, challenges like multi-lingual generalization and the absence of a unified evaluation framework persist. To address these gaps, we propose FEND (Foundation model-based Evaluation of Neuropsychiatric Disorders), a comprehensive multi-modal framework integrating speech and text modalities for detecting AD, depression, and ASD across the lifespan. Leveraging 13 multi-lingual datasets spanning English, Chinese, Greek, French, and Dutch, we systematically evaluate multi-modal fusion performance. Our results show that multi-modal fusion excels in AD and depression detection but underperforms in ASD due to dataset heterogeneity. We also identify modality imbalance as a prevalent issue, where multi-modal fusion fails to surpass the best mono-modal models. Cross-corpus experiments reveal robust performance in task- and language-consistent scenarios but noticeable degradation in multi-lingual and task-heterogeneous settings. By providing extensive benchmarks and a detailed analysis of performance-influencing factors, FEND advances the field of automated, lifespan-inclusive, and multi-lingual neuropsychiatric disorder assessment. We encourage researchers to adopt the FEND framework for fair comparisons and reproducible research.


【3】SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs
标题:SACodec:低比特率高保真神经语音编解码器的非对称量化
链接:https://arxiv.org/abs/2512.20944

作者:Zhongren Dong,Bin Wang,Jing Han,Haotian Guo,Xiaojun Mo,Yimin Cao,Zixing Zhang
摘要:神经语音编解码器在低比特率时面临一个基本的权衡:保持声学保真度通常会损害语义丰富性。为了解决这个问题,我们介绍SACodec,一种新的编解码器建立在一个不对称的双量化器,采用我们提出的语义验证机制。这种设计策略性地实现了语义和声学细节的量化。语义锚定是通过一个轻量级的投影仪实现的,该投影仪将声学特征与冻结的大规模mHuBERT码本对齐,在保证充分利用码本的同时注入语言先验。随后,对于声学细节,具有SimVQ的残差激活模块使得单层量化器(声学路径)能够忠实地恢复细粒度信息。SACodec仅以1.5 kbps的速度在保真度和语义方面都表现出色,从而建立了一种新的技术水平:主观听力测试证实,其重建质量在感知上与地面实况音频具有高度可比性,而其令牌在下游任务中表现出显著改善的语义丰富性。
摘要:Neural Speech Codecs face a fundamental trade-off at low bitrates: preserving acoustic fidelity often compromises semantic richness. To address this, we introduce SACodec, a novel codec built upon an asymmetric dual-quantizer that employs our proposed Semantic Anchoring mechanism. This design strategically decouples the quantization of Semantic and Acoustic details. The semantic anchoring is achieved via a lightweight projector that aligns acoustic features with a frozen, large-scale mHuBERT codebook, injecting linguistic priors while guaranteeing full codebook utilization. Sequentially, for acoustic details, a residual activation module with SimVQ enables a single-layer quantizer (acoustic path) to faithfully recover fine-grained information. At just 1.5 kbps, SACodec establishes a new state of the art by excelling in both fidelity and semantics: subjective listening tests confirm that its reconstruction quality is perceptually highly comparable to ground-truth audio, while its tokens demonstrate substantially improved semantic richness in downstream tasks.


eess.AS音频处理


【1】USE: A Unified Model for Universal Sound Separation and Extraction
标题:用途:通用声音分离和提取的统一模型
链接:https://arxiv.org/abs/2512.21215

作者:Hongyu Wang,Chenda Li,Xin Zhou,Shuai Wang,Yanmin Qian
备注:Accepted as an oral presentation by AAAI 2026
摘要:声分离和目标声提取是处理复杂声学场景的基本技术。虽然现有的SS方法难以确定未知数量的声源,TSE方法需要精确指定的线索,以实现最佳性能。本文提出了一个统一的框架,协同结合SS和TSE,以克服各自的局限性。我们的架构采用了两个互补的组件:1)编码器-解码器吸引子(EDA)网络,可以自动推断SS的源计数和相应的声学线索,以及2)多模态融合网络,可以精确地解释TSE的各种用户提供的线索(声学,语义或视觉)。通过跨任务一致性约束的联合训练,我们建立了一个统一的潜在空间,这两个范式的桥梁。在推理过程中,系统自适应地工作在完全自主的SS模式或线索驱动的TSE模式。实验表明,在这两个任务中的显着性能,与基线和86%的TSE准确度相比,在SS中的1.4dB SDR改善的显着改善。
摘要:Sound separation (SS) and target sound extraction (TSE) are fundamental techniques for addressing complex acoustic scenarios. While existing SS methods struggle with determining the unknown number of sound sources, TSE approaches require precisely specified clues to achieve optimal performance. This paper proposes a unified framework that synergistically combines SS and TSE to overcome their individual limitations. Our architecture employs two complementary components: 1) An Encoder-Decoder Attractor (EDA) network that automatically infers both the source count and corresponding acoustic clues for SS, and 2) A multi-modal fusion network that precisely interprets diverse user-provided clues (acoustic, semantic, or visual) for TSE. Through joint training with cross-task consistency constraints, we establish a unified latent space that bridges both paradigms. During inference, the system adaptively operates in either fully autonomous SS mode or clue-driven TSE mode. Experiments demonstrate remarkable performance in both tasks, with notable improvements of 1.4 dB SDR improvement in SS compared to baseline and 86\% TSE accuracy.


【2】GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model
标题:GenTES:通过粗到细的生成语言模型增强目标说话人提取
链接:https://arxiv.org/abs/2512.20978

作者:Haoyang Li,Xuyi Zhuang,Azmat Adnan,Ye Ni,Wei Rao,Shreyas Gopal,Eng Siong Chng
摘要:基于语言模型(LM)的生成建模已经成为TSE的一个有前途的方向,提供了改进的泛化和高保真语音的潜力。我们提出了GenTSE,一个两阶段的解码器只生成LM方法TSE:阶段1预测粗糙的语义令牌,阶段2生成精细的声学令牌。分离语义和声学稳定解码,并产生更忠实,内容对齐的目标语音。这两个阶段都使用连续的SSL或编解码器嵌入,提供比离散提示方法更丰富的上下文。为了减少暴露偏差,我们采用了冻结LM条件训练策略,该策略根据早期检查点的预测令牌对LM进行条件训练,以减少教师强迫训练和自回归推理之间的差距。我们进一步采用DPO来更好地将输出与人类的感知偏好相匹配。Libri2Mix上的实验表明,GenTSE在语音质量,可懂度和说话人一致性方面优于以前的基于LM的系统。
摘要:Language Model (LM)-based generative modeling has emerged as a promising direction for TSE, offering potential for improved generalization and high-fidelity speech. We present GenTSE, a two-stage decoder-only generative LM approach for TSE: Stage-1 predicts coarse semantic tokens, and Stage-2 generates fine acoustic tokens. Separating semantics and acoustics stabilizes decoding and yields more faithful, content-aligned target speech. Both stages use continuous SSL or codec embeddings, offering richer context than discretized-prompt methods. To reduce exposure bias, we employ a Frozen-LM Conditioning training strategy that conditions the LMs on predicted tokens from earlier checkpoints to reduce the gap between teacher-forcing training and autoregressive inference. We further employ DPO to better align outputs with human perceptual preferences. Experiments on Libri2Mix show that GenTSE surpasses previous LM-based systems in speech quality, intelligibility, and speaker consistency.


机器翻译由腾讯交互翻译提供,仅供参考