今日论文合集:cs.SD语音6篇,eess.AS音频处理9篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Transformer-Based Rhythm Quantization of Performance MIDI Using Beat Annotations
标题:使用节拍注释的基于转换器的表演节拍量化
链接:https://arxiv.org/abs/2604.22290
作者:Maximilian Wachter,Sebastian Murgul,Michael Heizmann
备注:Accepted to the 5th International Conference on SMART MULTIMEDIA (ICSM), 2025
摘要:节奏转录是记谱级自动音乐转录(AMT)的一个关键子任务。虽然深度学习模型已被广泛用于检测音频和视频表演中的韵律网格,但基于节拍的节奏量化在很大程度上仍未被探索。在这项工作中,我们介绍了一种新的深度学习方法,用于使用先验节拍信息量化性能。我们的方法利用Transformer架构来有效地处理用于训练量化模型的同步得分和性能数据。我们的方法的关键组成部分包括数据集准备,基于节拍的预量化方法,以在统一的框架内调整性能和得分时间,以及为这项任务量身定制的标记器。我们采用基于T5架构的Transformer模型来满足节奏量化的特定要求。该模型使用一组分数级的量化性能的客观评估指标进行评估。通过系统的评估,我们优化了数据表示和模型架构。此外,我们应用性能和分数增强,如换位,音符删除和性能侧时间抖动,以提高模型的鲁棒性。最后,定性分析比较了我们的模型与最先进的概率和深度学习模型在各种示例上的量化性能。我们的模型在ASAP数据集上实现了97.3%的发作F1分数和83.3%的音符值准确度。它可以很好地概括时间签名,包括那些在训练过程中看不到的签名,并生成可读的分数输出。对乐器特定数据集的微调通过捕捉特征节奏和旋律模式进一步提高了性能。这项工作有助于一个强大的和灵活的框架,基于拍频的量化使用Transformer模型。
摘要:Rhythm transcription is a key subtask of notation-level Automatic Music Transcription (AMT). While deep learning models have been extensively used for detecting the metrical grid in audio and MIDI performances, beat-based rhythm quantization remains largely unexplored. In this work, we introduce a novel deep learning approach for quantizing MIDI performances using a priori beat information. Our method leverages the transformer architecture to effectively process synchronized score and performance data for training a quantization model. Key components of our approach include dataset preparation, a beat-based pre-quantization method to align performance and score times within a unified framework, and a MIDI tokenizer tailored for this task. We adapt a transformer model based on the T5 architecture to meet the specific requirements of rhythm quantization. The model is evaluated using a set of score-level metrics designed for objective assessment of quantization performance. Through systematic evaluation, we optimize both data representation and model architecture. Additionally, we apply performance and score augmentations, such as transposition, note deletion, and performance-side time jitter, to enhance the model's robustness. Finally, a qualitative analysis compares our model's quantization performance against state-of-the-art probabilistic and deep-learning models on various example pieces. Our model achieves an onset F1-score of 97.3% and a note value accuracy of 83.3% on the ASAP dataset. It generalizes well across time signatures, including those not seen during training, and produces readable score output. Fine-tuning on instrument-specific datasets further improves performance by capturing characteristic rhythmic and melodic patterns. This work contributes a robust and flexible framework for beat-based MIDI quantization using transformer models.


【2】Spectrographic Portamento Gradient Analysis: A Quantitative Method for Historical Cello Recordings with Application to Beethoven's Piano and Cello Sonatas, 1930--2012

标题:光谱Portamento梯度分析:历史大提琴录音的定量方法,并应用于贝多芬钢琴和大提琴奏鸣曲,1930- 2012年
链接:https://arxiv.org/abs/2604.22037
作者:Ignasi Sole
摘要:弦乐演奏中的Portamento主要是作为一种二进制的存在或不存在现象来研究的,现有的研究测量了出现的频率,以及不太常见的以毫秒为单位的持续时间。本文介绍了第三个定量描述符;滑音幻灯片的光谱梯度,以Hz/秒为单位测量,并使用Sonic Visualizer的旋律谱图层,GIMP像素分析和对谱图已知频率轴的度量校准相结合的协议来演示其测量。梯度捕捉到了单靠持续时间无法捕捉到的东西:音高轨迹的陡峭程度,它编码了幻灯片的表达特征,而与幻灯片的长度无关。适用于的开放措施。特别是因为它们的单声道纹理允许可靠的光谱音高跟踪。该方法产生的梯度值范围从晚期记录的约600 Hz/s到20世纪早期表演的4,000 Hz/s以上。本文进一步记录了一个增益恢复协议,该协议将可分析的语料库扩展到20世纪30年代的模拟录音,其中滑音痕迹在数字传输中很微弱。将该方法应用于1930- 2012年的22个录音语料库,该论文测试了梯度陡度与节奏负相关的假设:较慢的表演产生更陡,更长的幻灯片,而较快的表演产生更浅的幻灯片或根本没有。研究结果支持了这一假设,表明滑音在20世纪的衰落并不是一个从有到无的二元过渡,而是一个持续的过程。
摘要:Portamento in string performance has been studied primarily as a binary presence-or-absence phenomenon, with existing research measuring frequency of occurrence and, less commonly, duration in milliseconds. This paper introduces a third quantitative descriptor; the spectrographic gradient of the portamento slide, measured in Hz/second, and demonstrates its measurement using a protocol combining Sonic Visualizer's melodic spectrogram layer, GIMP pixel analysis, and metric calibration against the spectrogram's known frequency axis. The gradient captures what duration alone cannot: the steepness of the pitch trajectory, which encodes the expressive character of the slide independently of its length. Applied to the opening measures of. Specifically because their monophonic texture permits reliable spectrographic pitch tracking. The method yields gradient values ranging from approximately 600~Hz/s in late-period recordings to over 4,000~Hz/s in early twentieth-century performances. The paper further documents a gain-recovery protocol that extends the analysable corpus to analogue recordings from the 1930s where portamento traces are faint in digital transfer. Applying the method to a corpus of 22 recordings spanning 1930--2012, the paper tests the hypothesis that gradient steepness correlates negatively with tempo: that slower performances produce steeper, longer slides while faster performances produce shallower slides or none at all. The results support this hypothesis, suggesting that the widely documented decline of portamento across the twentieth century is not a binary transition from presence to absence but a continuou


【3】Audio Effect Estimation with DNN-Based Prediction and Search Algorithm

标题:基于DNN的预测和搜索算法的音频效果估计
链接:https://arxiv.org/abs/2604.22276
作者:Youichi Okita,Haruhiro Katayose
备注:Accepted for ICASSP2026
摘要:音频效果在声音设计中起着至关重要的作用。本研究针对音频效果估计的任务,其目的是从湿信号中估计应用效果的配置。针对这个问题的现有方法可以分为预测方法和基于搜索的方法,预测方法使用以数据驱动的方式预先训练的模型,基于搜索的方法基于湿信号重构。在这项研究中,我们提出了一种集成这些方法的新方法:首先,DNN预测干信号和效果配置,然后使用这些预测基于湿信号重构进行搜索。通过在预测阶段中估计干信号,可以使用重构相似性作为目标函数来补充或改进预测。实验结果表明,基于该方法的方法优于单纯基于预测方法的方法。此外,研究结果表明,预测效果类型组合的任务划分,其次是基于搜索的估计顺序和参数是最有效的各种指标。
摘要:Audio effects play an essential role in sound design. This research addresses the task of audio effect estimation, which aims to estimate the configuration of applied effects from a wet signal. Existing approaches to this problem can be categorized into predictive approaches, which use models pre-trained in a data-driven manner, and search-based approaches, which are based on wet signal reconstruction. In this study, we propose a novel approach that integrates these approaches: first, DNNs predict the dry signal and effect configuration, and then a search is performed based on wet signal reconstruction using these predictions. By estimating the dry signal in the prediction stage, it becomes possible to complement or improve the predictions using reconstruction similarity as an objective function. The experimental evaluation showed that methods based on the proposed approach outperformed the method solely based on the predictive approach. Furthermore, the findings suggest that the task division of predicting the effect type combination followed by the search-based estimation of order and parameters was the most effective across various metrics.


【4】UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions

标题:UniSonate:通过文本指令生成语音、音乐和音效的统一模型
链接:https://arxiv.org/abs/2604.22209
作者:Chunyu Qiang,Xiaopeng Wang,Kang Yin,Yuzhe Liang,Yuxin Guo,Teng Ma,Ziyu Zhang,Tianrui Wang,Cheng Gong,Yushen Chen,Ruibo Fu,Chen Zhang,Longbiao Wang,Jianwu Dang
备注:Accepted to ACL 2026 main conference (oral)
摘要:生成音频建模在很大程度上被分割成专门的任务,文本到语音(TTS),文本到音乐(TTM)和文本到音频(TTA),每个都在异构的控制范式下运行。由于结构化语义表示(语音/音乐)和非结构化声学纹理(音效)之间的内在不协调,统一这些模态仍然是一个根本性的挑战。在本文中,我们介绍了UniSonate,一个统一的流匹配框架,能够通过标准化的,无参考的自然语言指令接口合成语音,音乐和音效。为了调和结构差异,我们提出了一种新的动态令牌注入机制,将非结构化的环境声音投射到结构化的时间潜在空间中,从而在音素驱动的多模态扩散Transformer(MM-DiT)中实现精确的持续时间控制。再加上多阶段的课程学习策略,这种方法有效地减轻跨模态优化冲突。大量实验表明,UniSonate在基于语音合成的TTS(WER 1.47%)和TTM(SongEval Coherence 3.18)中实现了最先进的性能,同时在TTA中保持了竞争力的保真度。至关重要的是,我们观察到正迁移,与单任务基线相比,对不同音频数据的联合训练显着增强了结构连贯性和韵律表现力。音频样本可在https://qiangchunyu.github.io/UniSonate/上获得。
摘要:Generative audio modeling has largely been fragmented into specialized tasks, text-to-speech (TTS), text-to-music (TTM), and text-to-audio (TTA), each operating under heterogeneous control paradigms. Unifying these modalities remains a fundamental challenge due to the intrinsic dissonance between structured semantic representations (speech/music) and unstructured acoustic textures (sound effects). In this paper, we introduce UniSonate, a unified flow-matching framework capable of synthesizing speech, music, and sound effects through a standardized, reference-free natural language instruction interface. To reconcile structural disparities, we propose a novel dynamic token injection mechanism that projects unstructured environmental sounds into a structured temporal latent space, enabling precise duration control within a phoneme-driven Multimodal Diffusion Transformer (MM-DiT). Coupled with a multi-stage curriculum learning strategy, this approach effectively mitigates cross-modal optimization conflicts. Extensive experiments demonstrate that UniSonate achieves state-of-the-art performance in instruction-based TTS (WER 1.47%) and TTM (SongEval Coherence 3.18), while maintaining competitive fidelity in TTA. Crucially, we observe positive transfer, where joint training on diverse audio data significantly enhances structural coherence and prosodic expressiveness compared to single-task baselines. Audio samples are available at https://qiangchunyu.github.io/UniSonate/.


【5】Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus

标题:使用自监督学习特征融合推进自动语音识别:Fearless Steps Apollo语料库的案例研究
链接:https://arxiv.org/abs/2604.22203
作者:Szu-Jui Chen,John H. L. Hansen
备注:Accepted to Speech Communication 2026
摘要:使用自监督学习(SSL)模型显著提高了下游语音任务的性能,超越了传统手工制作功能的能力。本研究探讨SSL模型的融合,目的是利用他们各自的优势和完善提取的功能,以实现改进的语音识别模型的自然场景。我们的研究调查了大量的自然主义无畏步骤(FS)APOLLO资源,特别关注FS挑战(FSC)第四阶段语料库,提供了对该数据集的初步分析。此外,我们还结合了CHiME-6数据集来评估各种自然语音场景的性能。在探索先前提出的特征细化损失和融合方法时,我们发现这些方法在FSC Phase-4语料库上效果不佳。为了解决这个问题,我们引入了一种新的深度交叉注意(DCA)融合方法,旨在提高性能,特别是对于FSC Phase-4语料库。我们的目标是促进创建卓越的FS APOLLO社区资源,满足各学科研究人员的多样化需求。所提出的解决方案实现了WER绝对+1.1%的改进,为大量FS APOLLO社区资源提供了有效的元数据创建。
摘要:Using self-supervised learning (SSL) models has significantly improved performance for downstream speech tasks, surpassing the capabilities of traditional hand-crafted features. This study investigates the amalgamation of SSL models, with the aim to leverage both their individual strengths and refine extracted features to achieve improved speech recognition models for naturalistic scenarios. Our research investigates the massive naturalistic Fearless Steps (FS) APOLLO resource, with particular focus on the FS Challenge (FSC) Phase-4 corpus, providing the inaugural analysis of this dataset. Additionally, we incorporate the CHiME-6 dataset to evaluate performance across diverse naturalistic speech scenarios. While exploring previously proposed Feature Refinement Loss and fusion methods, we found these methods to be less effective on the FSC Phase-4 corpus. To address this, we introduce a novel deep cross-attention (DCA) fusion method, designed to elevate performance, especially for the FSC Phase-4 corpus. Our objective is to foster creation of superior FS APOLLO community resources, catering to the diverse needs of researchers across various disciplines. The proposed solution achieves an absolute +1.1% improvement in WER, providing effective meta-data creation for the massive FS APOLLO community resource.


【6】Beyond Acoustic Sparsity and Linguistic Bias: A Prompt-Free Paradigm for Mispronunciation Detection and Diagnosis

标题:超越声学稀疏性和语言偏见:发音错误检测和诊断的无预算范式
链接:https://arxiv.org/abs/2604.22133
作者:Haopeng Geng,Longfei Yang,Xi Chen,Haitong Sun,Daisuke Saito,Nobuaki Minematsu
摘要:发音错误检测和诊断(MDD)需要对细粒度的声学偏差进行建模。然而,目前的ASR派生MDD系统往往面临固有的局限性。特别是,CTC为基础的模型有利于序列水平的比对,忽略了短暂的发音错误的线索,而明确的规范先验偏向于预期目标的预测。为了解决这些瓶颈,我们提出了一个无干扰的框架解耦声学保真度从规范的指导。首先,我们介绍CROTTC,声学模型执行单调,帧级对齐,以准确地捕捉发音偏差。其次,我们通过知识转移原则下的IF策略隐式地注入发音错误信息。实验表明,CROTTC-IF在L2-ARCTIC上获得了71.77%的F1分数,在Iqra'Eval 2排行榜上获得了71.70%的F1分数。通过实证分析,我们证明了从显式先验解耦声学提供了高度鲁棒的MDD。
摘要:Mispronunciation Detection and Diagnosis (MDD) requires modeling fine-grained acoustic deviations. However, current ASR-derived MDD systems often face inherent limitations. In particular, CTC-based models favor sequence-level alignments that neglect transient mispronunciation cues, while explicit canonical priors bias predictions toward intended targets. To address these bottlenecks, we propose a prompt-free framework decoupling acoustic fidelity from canonical guidance. First, we introduce CROTTC, an acoustic model enforcing monotonic, frame-level alignment to accurately capture pronunciation deviations. Second, we implicitly inject mispronunciation information via the IF strategy under the knowledge transfer principle. Experiments show CROTTC-IF achieves a 71.77% F1-score on L2-ARCTIC and 71.70% F1-score on the Iqra'Eval2 leaderboard. With empirical analysis, we demonstrate that decoupling acoustics from explicit priors provides highly robust MDD.


eess.AS音频处理


【1】DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models
标题:DM-ASB:具有大型语言模型的日记感知多扬声器ASB
链接:https://arxiv.org/abs/2604.22467
作者:Li Li,Ming Cheng,Weixin Zhu,Yannan Wang,Juan Liu,Ming Li
摘要:多说话者自动语音识别(ASR)旨在转录涉及多个说话者的会话语音,要求模型不仅要捕获所说的内容,还要捕获是谁说的,有时还要捕获何时说的。最近的Speech-LLM方法已经显示出统一建模的潜力,但联合学习说话人属性,时间结构和词汇识别仍然是困难的和数据密集型的。在目前阶段,利用可靠的发言人日记作为一个明确的结构先验提供了一个实用和有效的方法来简化这项任务。为了有效地利用这些先验知识,我们提出了DM-ASR,一个diarization-aware多扬声器ASR框架,将任务重新定义为多回合对话生成过程。给定音频块和日志化结果,DM-ASR将转录分解为说话者和时间条件查询的序列,每个查询对应于一个时间段中的一个说话者。该公式将多说话人识别转换为一系列结构化的子任务,明确地将说话人时间结构与语言内容解耦,并使日志化线索与大型语言模型的推理能力有效整合。我们还引入了一个可选的单词级时间戳预测机制,该机制将单词和时间戳令牌交织在一起,从而产生更丰富的结构化输出和更好的转录质量。我们的分析表明,日记化系统提供了更可靠的扬声器身份和段级边界,而LLM擅长建模语言内容和远程依赖关系,展示了它们的互补优势。在普通话和英语基准测试上的实验表明,该方法在相对较小的模型和训练数据下实现了较强的性能,同时与现有的统一方法保持竞争力或优于现有的统一方法。
摘要:Multi-speaker automatic speech recognition (ASR) aims to transcribe conversational speech involving multiple speakers, requiring the model to capture not only what was said, but also who said it and sometimes when it was spoken. Recent Speech-LLM approaches have shown the potential of unified modeling for this task, but jointly learning speaker attribution, temporal structure, and lexical recognition remains difficult and data-intensive. At the current stage, leveraging reliable speaker diarization as an explicit structural prior provides a practical and efficient way to simplify this task. To effectively exploit such priors, we propose DM-ASR, a diarization-aware multi-speaker ASR framework that reformulates the task as a multi-turn dialogue generation process. Given an audio chunk and diarization results, DM-ASR decomposes transcription into a sequence of speaker- and time-conditioned queries, each corresponding to one speaker in one time segment. This formulation converts multi-speaker recognition into a series of structured sub-tasks, explicitly decoupling speaker-temporal structure from linguistic content and enabling effective integration of diarization cues with the reasoning capability of large language models. We further introduce an optional word-level timestamp prediction mechanism that interleaves word and timestamp tokens, yielding richer structured outputs and better transcription quality. Our analysis shows that diarization systems provide more reliable speaker identities and segment-level boundaries, while LLMs excel at modeling linguistic content and long-range dependencies, demonstrating their complementary strengths. Experiments on Mandarin and English benchmarks show that the proposed approach achieves strong performance with relatively small models and training data, while remaining competitive with or outperforming existing unified approaches.


【2】Audio Effect Estimation with DNN-Based Prediction and Search Algorithm

标题:基于DNN的预测和搜索算法的音频效果估计
链接:https://arxiv.org/abs/2604.22276
作者:Youichi Okita,Haruhiro Katayose
备注:Accepted for ICASSP2026
摘要:音频效果在声音设计中起着至关重要的作用。本研究针对音频效果估计的任务,其目的是从湿信号中估计应用效果的配置。针对这个问题的现有方法可以分为预测方法和基于搜索的方法,预测方法使用以数据驱动的方式预先训练的模型,基于搜索的方法基于湿信号重构。在这项研究中,我们提出了一种集成这些方法的新方法:首先,DNN预测干信号和效果配置,然后使用这些预测基于湿信号重构进行搜索。通过在预测阶段中估计干信号,可以使用重构相似性作为目标函数来补充或改进预测。实验结果表明,基于该方法的方法优于单纯基于预测方法的方法。此外,研究结果表明,在各种指标中,预测效应类型组合的任务划分以及基于搜索的顺序和参数估计是最有效的。
摘要:Audio effects play an essential role in sound design. This research addresses the task of audio effect estimation, which aims to estimate the configuration of applied effects from a wet signal. Existing approaches to this problem can be categorized into predictive approaches, which use models pre-trained in a data-driven manner, and search-based approaches, which are based on wet signal reconstruction. In this study, we propose a novel approach that integrates these approaches: first, DNNs predict the dry signal and effect configuration, and then a search is performed based on wet signal reconstruction using these predictions. By estimating the dry signal in the prediction stage, it becomes possible to complement or improve the predictions using reconstruction similarity as an objective function. The experimental evaluation showed that methods based on the proposed approach outperformed the method solely based on the predictive approach. Furthermore, the findings suggest that the task division of predicting the effect type combination followed by the search-based estimation of order and parameters was the most effective across various metrics.


【3】Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding

标题:与时间一起倾听:精确的时间意识以实现长篇音频理解
链接:https://arxiv.org/abs/2604.22245
作者:Mingchen Shao,Hang Su,Wenjie Tian,Bingshen Mu,Zhennan Lin,Lichun Fan,Zhenbo Luo,Jian Luan,Lei Xie
摘要:虽然大型音频语言模型(LALM)在短音频上实现了强大的性能,但它们在长形式输入上会下降。这种退化在时间意识任务中更为严重,其中时间对齐随着音频持续时间的增长而变得越来越不准确。我们将这些限制归因于缺乏针对长时间感知的数据、基准和建模方法。为了弥合这一差距,我们首先构建了LAT-Chronicle,这是一个1.2k小时长的音频数据集,具有跨真实世界场景的时间注释。我们进一步开发LAT-Bench,这是第一个经过人类验证的基准测试,支持长达30分钟的音频,同时涵盖三个核心任务:密集音频字幕,时间音频接地和目标音频字幕。利用这些资源,我们提出了LAT-Audio,制定时间意识作为一个渐进的全球到本地的推理范式。首先构建一个全局时间轴作为对齐的时间语义上下文,然后引入带音频思维链(TWA-CoT),通过使用工具结合本地音频信息来执行迭代推理。实验表明,LAT-Audio在长形式音频时间感知任务上优于现有模型,并提高了对输入持续时间的鲁棒性。我们在https://github.com/alanshaoTT/LAT-Audio-Repo上发布数据集,基准和模型,以促进未来的研究。
摘要:While Large Audio Language Models (LALMs) achieve strong performance on short audio, they degrade on long-form inputs. This degradation is more severe in temporal awareness tasks, where temporal alignment becomes increasingly inaccurate as audio duration grows. We attribute these limitations to the lack of data, benchmarks, and modeling approaches tailored for long-form temporal awareness. To bridge this gap, we first construct LAT-Chronicle, a 1.2k hour long-form audio dataset with temporal annotations across real-world scenarios. We further develop LAT-Bench, the first human-verified benchmark supporting audio up to 30 minutes while covering three core tasks: Dense Audio Caption, Temporal Audio Grounding, and Targeted Audio Caption. Leveraging these resources, we propose LAT-Audio, formulating temporal awareness as a progressive global-to-local reasoning paradigm. A global timeline is first constructed as an aligned temporal-semantic context,and the Think-With-Audio Chain-of-Thought (TWA-CoT) is then introduced to perform iterative reasoning by incorporating local audio information via tool use. Experiments show that LAT-Audio surpasses existing models on long-form audio temporal awareness tasks and improves robustness to input duration. We release the dataset, benchmark, and model to facilitate future research at https://github.com/alanshaoTT/LAT-Audio-Repo.


【4】UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions

标题:UniSonate:通过文本指令生成语音、音乐和音效的统一模型
链接:https://arxiv.org/abs/2604.22209
作者:Chunyu Qiang,Xiaopeng Wang,Kang Yin,Yuzhe Liang,Yuxin Guo,Teng Ma,Ziyu Zhang,Tianrui Wang,Cheng Gong,Yushen Chen,Ruibo Fu,Chen Zhang,Longbiao Wang,Jianwu Dang
备注:Accepted to ACL 2026 main conference (oral)
摘要:生成音频建模在很大程度上被分割成专门的任务,文本到语音(TTS),文本到音乐(TTM)和文本到音频(TTA),每个都在异构的控制范式下运行。由于结构化语义表示(语音/音乐)和非结构化声学纹理(音效)之间的内在不和谐,统一这些模态仍然是一个根本性的挑战。在本文中,我们介绍了UniSonate,一个统一的流匹配框架,能够通过标准化的,无参考的自然语言指令接口合成语音,音乐和音效。为了调和结构差异,我们提出了一种新的动态令牌注入机制,将非结构化的环境声音投射到结构化的时间潜在空间中,从而在音素驱动的多模态扩散Transformer(MM-DiT)中实现精确的持续时间控制。再加上多阶段的课程学习策略,这种方法有效地减轻跨模态优化冲突。大量实验表明,UniSonate在基于语音合成的TTS(WER 1.47%)和TTM(SongEval Coherence 3.18)中实现了最先进的性能,同时在TTA中保持了竞争力的保真度。至关重要的是,我们观察到正迁移,与单任务基线相比,对不同音频数据的联合训练显着增强了结构连贯性和韵律表现力。音频样本可在https://qiangchunyu.github.io/UniSonate/上获得。
摘要:Generative audio modeling has largely been fragmented into specialized tasks, text-to-speech (TTS), text-to-music (TTM), and text-to-audio (TTA), each operating under heterogeneous control paradigms. Unifying these modalities remains a fundamental challenge due to the intrinsic dissonance between structured semantic representations (speech/music) and unstructured acoustic textures (sound effects). In this paper, we introduce UniSonate, a unified flow-matching framework capable of synthesizing speech, music, and sound effects through a standardized, reference-free natural language instruction interface. To reconcile structural disparities, we propose a novel dynamic token injection mechanism that projects unstructured environmental sounds into a structured temporal latent space, enabling precise duration control within a phoneme-driven Multimodal Diffusion Transformer (MM-DiT). Coupled with a multi-stage curriculum learning strategy, this approach effectively mitigates cross-modal optimization conflicts. Extensive experiments demonstrate that UniSonate achieves state-of-the-art performance in instruction-based TTS (WER 1.47%) and TTM (SongEval Coherence 3.18), while maintaining competitive fidelity in TTA. Crucially, we observe positive transfer, where joint training on diverse audio data significantly enhances structural coherence and prosodic expressiveness compared to single-task baselines. Audio samples are available at https://qiangchunyu.github.io/UniSonate/.


【5】Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus

标题:使用自监督学习特征融合推进自动语音识别:Fearless Steps Apollo语料库的案例研究
链接:https://arxiv.org/abs/2604.22203
作者:Szu-Jui Chen,John H. L. Hansen
备注:Accepted to Speech Communication 2026
摘要:使用自监督学习(SSL)模型显著提高了下游语音任务的性能,超越了传统手工制作功能的能力。本研究探讨SSL模型的融合,目的是利用他们各自的优势和完善提取的功能,以实现改进的语音识别模型的自然场景。我们的研究调查了大量的自然主义无畏步骤(FS)APOLLO资源,特别关注FS挑战(FSC)第四阶段语料库,提供了对该数据集的初步分析。此外,我们还结合了CHiME-6数据集来评估各种自然语音场景的性能。在探索先前提出的特征细化损失和融合方法时,我们发现这些方法在FSC Phase-4语料库上效果不佳。为了解决这个问题,我们引入了一种新的深度交叉注意(DCA)融合方法,旨在提高性能,特别是对于FSC Phase-4语料库。我们的目标是促进创建卓越的FS APOLLO社区资源,满足各学科研究人员的多样化需求。所提出的解决方案实现了WER绝对+1.1%的改进,为大量FS APOLLO社区资源提供了有效的元数据创建。
摘要:Using self-supervised learning (SSL) models has significantly improved performance for downstream speech tasks, surpassing the capabilities of traditional hand-crafted features. This study investigates the amalgamation of SSL models, with the aim to leverage both their individual strengths and refine extracted features to achieve improved speech recognition models for naturalistic scenarios. Our research investigates the massive naturalistic Fearless Steps (FS) APOLLO resource, with particular focus on the FS Challenge (FSC) Phase-4 corpus, providing the inaugural analysis of this dataset. Additionally, we incorporate the CHiME-6 dataset to evaluate performance across diverse naturalistic speech scenarios. While exploring previously proposed Feature Refinement Loss and fusion methods, we found these methods to be less effective on the FSC Phase-4 corpus. To address this, we introduce a novel deep cross-attention (DCA) fusion method, designed to elevate performance, especially for the FSC Phase-4 corpus. Our objective is to foster creation of superior FS APOLLO community resources, catering to the diverse needs of researchers across various disciplines. The proposed solution achieves an absolute +1.1% improvement in WER, providing effective meta-data creation for the massive FS APOLLO community resource.


【6】Beyond Acoustic Sparsity and Linguistic Bias: A Prompt-Free Paradigm for Mispronunciation Detection and Diagnosis

标题:超越声学稀疏性和语言偏见:发音错误检测和诊断的无预算范式
链接:https://arxiv.org/abs/2604.22133
作者:Haopeng Geng,Longfei Yang,Xi Chen,Haitong Sun,Daisuke Saito,Nobuaki Minematsu
摘要:发音错误检测和诊断(MDD)需要对细粒度的声学偏差进行建模。然而,目前的ASR派生MDD系统往往面临固有的局限性。特别是,CTC为基础的模型有利于序列水平的比对,忽略了短暂的发音错误的线索,而明确的规范先验偏向于预期目标的预测。为了解决这些瓶颈,我们提出了一个无干扰的框架解耦声学保真度从规范的指导。首先,我们介绍CROTTC,声学模型执行单调,帧级对齐,以准确地捕捉发音偏差。其次,我们通过知识转移原则下的IF策略隐式地注入发音错误信息。实验表明,CROTTC-IF在L2-ARCTIC上获得了71.77%的F1分数,在Iqra'Eval 2排行榜上获得了71.70%的F1分数。通过实证分析,我们证明了从显式先验解耦声学提供了高度鲁棒的MDD。
摘要:Mispronunciation Detection and Diagnosis (MDD) requires modeling fine-grained acoustic deviations. However, current ASR-derived MDD systems often face inherent limitations. In particular, CTC-based models favor sequence-level alignments that neglect transient mispronunciation cues, while explicit canonical priors bias predictions toward intended targets. To address these bottlenecks, we propose a prompt-free framework decoupling acoustic fidelity from canonical guidance. First, we introduce CROTTC, an acoustic model enforcing monotonic, frame-level alignment to accurately capture pronunciation deviations. Second, we implicitly inject mispronunciation information via the IF strategy under the knowledge transfer principle. Experiments show CROTTC-IF achieves a 71.77% F1-score on L2-ARCTIC and 71.70% F1-score on the Iqra'Eval2 leaderboard. With empirical analysis, we demonstrate that decoupling acoustics from explicit priors provides highly robust MDD.


【7】Transformer-Based Rhythm Quantization of Performance MIDI Using Beat Annotations

标题:使用节拍注释的基于转换器的表演节拍量化
链接:https://arxiv.org/abs/2604.22290
作者:Maximilian Wachter,Sebastian Murgul,Michael Heizmann
备注:Accepted to the 5th International Conference on SMART MULTIMEDIA (ICSM), 2025
摘要:节奏转录是记谱级自动音乐转录(AMT)的一个关键子任务。虽然深度学习模型已被广泛用于检测音频和视频表演中的韵律网格,但基于节拍的节奏量化在很大程度上仍未被探索。在这项工作中,我们介绍了一种新的深度学习方法,用于使用先验节拍信息量化性能。我们的方法利用Transformer架构来有效地处理用于训练量化模型的同步得分和性能数据。我们的方法的关键组成部分包括数据集准备,基于节拍的预量化方法,以在统一的框架内调整性能和得分时间,以及为这项任务量身定制的标记器。我们采用基于T5架构的Transformer模型来满足节奏量化的特定要求。该模型使用一组分数级的量化性能的客观评估指标进行评估。通过系统的评估,我们优化了数据表示和模型架构。此外,我们应用性能和分数增强,如换位,音符删除和性能侧时间抖动,以提高模型的鲁棒性。最后,定性分析将我们模型的量化性能与各种示例的最新概率和深度学习模型进行比较。我们的模型在ASAP数据集上实现了97.3%的发作F1分数和83.3%的音符值准确度。它可以很好地概括时间签名,包括那些在训练过程中看不到的签名,并生成可读的分数输出。对乐器特定数据集的微调通过捕捉特征节奏和旋律模式进一步提高了性能。这项工作有助于一个强大的和灵活的框架,基于拍频的量化使用Transformer模型。
摘要:Rhythm transcription is a key subtask of notation-level Automatic Music Transcription (AMT). While deep learning models have been extensively used for detecting the metrical grid in audio and MIDI performances, beat-based rhythm quantization remains largely unexplored. In this work, we introduce a novel deep learning approach for quantizing MIDI performances using a priori beat information. Our method leverages the transformer architecture to effectively process synchronized score and performance data for training a quantization model. Key components of our approach include dataset preparation, a beat-based pre-quantization method to align performance and score times within a unified framework, and a MIDI tokenizer tailored for this task. We adapt a transformer model based on the T5 architecture to meet the specific requirements of rhythm quantization. The model is evaluated using a set of score-level metrics designed for objective assessment of quantization performance. Through systematic evaluation, we optimize both data representation and model architecture. Additionally, we apply performance and score augmentations, such as transposition, note deletion, and performance-side time jitter, to enhance the model's robustness. Finally, a qualitative analysis compares our model's quantization performance against state-of-the-art probabilistic and deep-learning models on various example pieces. Our model achieves an onset F1-score of 97.3% and a note value accuracy of 83.3% on the ASAP dataset. It generalizes well across time signatures, including those not seen during training, and produces readable score output. Fine-tuning on instrument-specific datasets further improves performance by capturing characteristic rhythmic and melodic patterns. This work contributes a robust and flexible framework for beat-based MIDI quantization using transformer models.


【8】TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis

标题:TTS-PRism:一种用于细粒度诊断的感知推理和可解释语音模型
链接:https://arxiv.org/abs/2604.22225
作者:Xi Wang,Jie Wang,Xingchen Song,Baijun Song,Jingran Xie,Jiahe Shao,Zijian Lin,Di Wu,Meng Meng,Jian Luan,Zhiyong Wu
备注:Submitted to Interspeech 2026
摘要:虽然生成式文本到语音(TTS)模型接近人类水平的质量,但单一指标无法诊断细粒度的声学伪影或解释感知崩溃。为了解决这个问题,我们提出了TTS-PRISM,一个多维的诊断框架,为普通话。首先,我们建立了一个12维的架构,跨越稳定性,先进的表现力。其次,我们设计了一个具有对抗性扰动和专家锚的有针对性的合成管道,以构建高质量的诊断数据集。第三,模式驱动的指令调优将明确的评分标准和推理嵌入到高效的端到端模型中。在1,600个样本的Gold测试集上的实验表明,TTS-PRISM在人类对齐方面优于通才模型。分析六个TTS范例建立了直观的诊断标志,揭示了细粒度的能力差异。TTS-PRISM是开源的,代码和检查点在https://github.com/xiaomi-research/tts-prism上。
摘要:While generative text-to-speech (TTS) models approach human-level quality, monolithic metrics fail to diagnose fine-grained acoustic artifacts or explain perceptual collapse. To address this, we propose TTS-PRISM, a multi-dimensional diagnostic framework for Mandarin. First, we establish a 12-dimensional schema spanning stability to advanced expressiveness. Second, we design a targeted synthesis pipeline with adversarial perturbations and expert anchors to build a high-quality diagnostic dataset. Third, schema-driven instruction tuning embeds explicit scoring criteria and reasoning into an efficient end-to-end model. Experiments on a 1,600-sample Gold Test Set show TTS-PRISM outperforms generalist models in human alignment. Profiling six TTS paradigms establishes intuitive diagnostic flags that reveal fine-grained capability differences. TTS-PRISM is open-source, with code and checkpoints at https://github.com/xiaomi-research/tts-prism.


【9】Spectrographic Portamento Gradient Analysis: A Quantitative Method for Historical Cello Recordings with Application to Beethoven's Piano and Cello Sonatas, 1930--2012

标题:光谱Portamento梯度分析:历史大提琴录音的定量方法,并应用于贝多芬钢琴和大提琴奏鸣曲,1930- 2012年
链接:https://arxiv.org/abs/2604.22037
作者:Ignasi Sole
摘要:弦乐演奏中的Portamento主要是作为一种二进制的存在或不存在现象来研究的,现有的研究测量了出现的频率,以及不太常见的以毫秒为单位的持续时间。本文介绍了第三个定量描述符;滑音幻灯片的光谱梯度,以Hz/秒为单位测量,并使用Sonic Visualizer的旋律谱图层,GIMP像素分析和对谱图已知频率轴的度量校准相结合的协议来演示其测量。梯度捕捉到了单靠持续时间无法捕捉到的东西:音高轨迹的陡峭程度,它编码了幻灯片的表达特征,而与幻灯片的长度无关。适用于的开放措施。特别是因为它们的单声道纹理允许可靠的光谱音高跟踪。该方法产生的梯度值范围从晚期记录的约600 Hz/s到20世纪早期表演的4,000 Hz/s以上。本文进一步记录了一个增益恢复协议,该协议将可分析的语料库扩展到20世纪30年代的模拟录音,其中滑音痕迹在数字传输中很微弱。将该方法应用于1930- 2012年的22个录音语料库,该论文测试了梯度陡度与节奏负相关的假设:较慢的表演产生更陡,更长的幻灯片,而较快的表演产生更浅的幻灯片或根本没有。研究结果支持了这一假设,表明滑音在20世纪的衰落并不是一个从有到无的二元过渡,而是一个持续的过程。
摘要:Portamento in string performance has been studied primarily as a binary presence-or-absence phenomenon, with existing research measuring frequency of occurrence and, less commonly, duration in milliseconds. This paper introduces a third quantitative descriptor; the spectrographic gradient of the portamento slide, measured in Hz/second, and demonstrates its measurement using a protocol combining Sonic Visualizer's melodic spectrogram layer, GIMP pixel analysis, and metric calibration against the spectrogram's known frequency axis. The gradient captures what duration alone cannot: the steepness of the pitch trajectory, which encodes the expressive character of the slide independently of its length. Applied to the opening measures of. Specifically because their monophonic texture permits reliable spectrographic pitch tracking. The method yields gradient values ranging from approximately 600~Hz/s in late-period recordings to over 4,000~Hz/s in early twentieth-century performances. The paper further documents a gain-recovery protocol that extends the analysable corpus to analogue recordings from the 1930s where portamento traces are faint in digital transfer. Applying the method to a corpus of 22 recordings spanning 1930--2012, the paper tests the hypothesis that gradient steepness correlates negatively with tempo: that slower performances produce steeper, longer slides while faster performances produce shallower slides or none at all. The results support this hypothesis, suggesting that the widely documented decline of portamento across the twentieth century is not a binary transition from presence to absence but a continuou


机器翻译由腾讯交互翻译提供,仅供参考