今日论文合集:cs.SD语音4篇,eess.AS音频处理8篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Step-Audio-EditX Technical Report
标题:Step-Audio-EditX技术报告
链接:https://arxiv.org/abs/2511.03601

作者:Chao Yan, Boyong Wu, Peng Yang, Pengfei Tan, Guoqiang Hu, Yuxin Zhang, Xiangyu (Tony)Zhang, Fei Tian, Xuerui Yang, Xiangyu Zhang, Daxin Jiang, Gang Yu
摘要:Step-Audio-EditX是第一个开源的基于LLM的音频模型,擅长表达和迭代音频编辑,包括情感,说话风格和非语言学以及强大的zero-shot文本到语音(TTS)功能。我们的核心创新在于只利用大利润率的合成数据,这避免了对基于嵌入的先验或辅助模块的需求。这种大幅度的学习方法可以实现迭代控制和跨声音的高表现力,并且代表了传统的代表级解纠缠的基本支点。测试结果表明,Step-Audio-EditX在情感编辑和其他细粒度控制任务方面优于MiniMax-2.6-hd和Doubao-Seed-TTS-2.0。
摘要:We present Step-Audio-EditX, the first open-source LLM-based audio model excelling at expressive and iterative audio editing encompassing emotion, speaking style, and paralinguistics alongside robust zero-shot text-to-speech (TTS) capabilities.Our core innovation lies in leveraging only large-margin synthetic data, which circumvents the need for embedding-based priors or auxiliary modules. This large-margin learning approach enables both iterative control and high expressivity across voices, and represents a fundamental pivot from the conventional focus on representation-level disentanglement. Evaluation results demonstrate that Step-Audio-EditX surpasses both MiniMax-2.6-hd and Doubao-Seed-TTS-2.0 in emotion editing and other fine-grained control tasks.


【2】SyMuPe: Affective and Controllable Symbolic Music Performance
标题:SyMuPe:感人且可控的象征音乐表演
链接:https://arxiv.org/abs/2511.03425

作者:Ilya Borovik, Dmitrii Gavrilev, Vladimir Viro
备注:ACM Multimedia 2025. Extended version with supplementary material
摘要:情感是音乐表演创作和感知的基础。然而,通过机器学习模型实现类似人类的表达和情感,以实现性能渲染仍然是一项具有挑战性的任务。在这项工作中,我们提出了SyMuPe,一个新的框架,用于开发和训练情感和可控的象征性钢琴演奏模型。我们的旗舰模型PianoFlow使用经过训练的条件流匹配来解决各种多掩模性能修复任务。通过设计,它支持无条件生成和填充音乐表演功能。在训练中,我们使用了一个精心策划的、经过清理的数据集,其中包含了2,968小时的对齐乐谱和富有表现力的演奏。对于文本和情感控制,我们集成了钢琴演奏情感分类器,并将PianoFlow与情感加权的Flan-T5文本嵌入作为条件输入进行调整。针对基于transformer的基线和现有模型的客观和主观评估表明,PianoFlow不仅优于其他方法,而且还实现了与人类记录和转录的样本相当的性能质量。对于情绪控制,我们提出并分析了不同的文本条件下产生的样本。开发的模型可以集成到交互式应用程序中,有助于创建更易于访问和参与的音乐表演系统。
摘要:Emotions are fundamental to the creation and perception of music performances. However, achieving human-like expression and emotion through machine learning models for performance rendering remains a challenging task. In this work, we present SyMuPe, a novel framework for developing and training affective and controllable symbolic piano performance models. Our flagship model, PianoFlow, uses conditional flow matching trained to solve diverse multi-mask performance inpainting tasks. By design, it supports both unconditional generation and infilling of music performance features. For training, we use a curated, cleaned dataset of 2,968 hours of aligned musical scores and expressive MIDI performances. For text and emotion control, we integrate a piano performance emotion classifier and tune PianoFlow with the emotion-weighted Flan-T5 text embeddings provided as conditional inputs. Objective and subjective evaluations against transformer-based baselines and existing models show that PianoFlow not only outperforms other approaches, but also achieves performance quality comparable to that of human-recorded and transcribed MIDI samples. For emotion control, we present and analyze samples generated under different text conditioning scenarios. The developed model can be integrated into interactive applications, contributing to the creation of more accessible and engaging music performance systems.


【3】Why Not Put a Microphone Near the Loudspeaker? A New Paradigm for Acoustic Echo Cancellation
标题:为什么不把麦克风放在扬声器附近?一种新的声学回声消除方法
链接:https://arxiv.org/abs/2511.03244

作者:Fei Zhao, Zhong-Qiu Wang
摘要:由于低成本扬声器和复杂的室内声学引起的非线性失真,声学回声消除(AEC)在现实环境中仍然具有挑战性。为了缓解这些问题,我们引入了双麦克风配置,其中辅助参考麦克风放置在扬声器附近以捕获非线性失真的远端信号。虽然这个参考信号被污染的近端语音,我们提出了一个预处理模块的基础上维纳滤波估计压缩的时间-频率掩模来抑制近端组件。该净化的参考信号实现了更有效的线性AEC级,其残余误差信号然后被馈送到深度神经网络以进行联合残余回声和噪声抑制。评估结果表明,我们的方法优于匹配的测试集上的基线方法。为了评估它在强非线性下的鲁棒性,我们进一步在不匹配的数据集上测试它,并观察到它实现了实质性的性能增益。这些结果表明,在实际情况下,非线性失真通常是未知的有效性。
摘要:Acoustic echo cancellation (AEC) remains challenging in real-world environments due to nonlinear distortions caused by low-cost loudspeakers and complex room acoustics. To mitigate these issues, we introduce a dual-microphone configuration, where an auxiliary reference microphone is placed near the loudspeaker to capture the nonlinearly distorted far-end signal. Although this reference signal is contaminated by near-end speech, we propose a preprocessing module based on Wiener filtering to estimate a compressed time-frequency mask to suppress near-end components. This purified reference signal enables a more effective linear AEC stage, whose residual error signal is then fed to a deep neural network for joint residual echo and noise suppression. Evaluation results show that our method outperforms baseline approaches on matched test sets. To evaluate its robustness under strong nonlinearities, we further test it on a mismatched dataset and observe that it achieves substantial performance gains. These results demonstrate its effectiveness in practical scenarios where the nonlinear distortions are typically unknown.


【4】audio2chart: End to End Audio Transcription into playable Guitar Hero charts
标题:audio 2 chart:端到端音频转录到可玩的吉他英雄排行榜
链接:https://arxiv.org/abs/2511.03337

作者:Riccardo Tripodi
摘要:这篇文章介绍了audio2chart,一个直接从原始音频自动生成Guitar Hero风格图表的框架。该任务被形式化为序列预测问题,其中模型被训练以生成与离散时间步长上的音频对齐的离散图表令牌。无条件的基线展示了强大的预测性能,而音频调节的添加在基于准确性的度量上产生了一致的改进。这项工作表明,将音频调节是可行的和有效的,以提高自动图表生成的音符预测。用于训练和推理的完整代码库在GitHub上公开提供,支持对神经图表生成的可重复研究。一系列预训练模型在Hugging Face上发布。
摘要:This work introduces audio2chart, a framework for the automatic generation of Guitar Hero style charts directly from raw audio. The task is formalized as a sequence prediction problem, where models are trained to generate discrete chart tokens aligned with the audio on discrete time steps. An unconditional baseline demonstrates strong predictive performance, while the addition of audio conditioning yields consistent improvements across accuracy based metrics. This work demonstrates that incorporating audio conditioning is both feasible and effective for improving note prediction in automatic chart generation. The complete codebase for training and inference is publicly available on GitHub supporting reproducible research on neural chart generation. A family of pretrained models is released on Hugging Face.


eess.AS音频处理


【1】Seeing What You Say: Expressive Image Generation from Speech
标题:看到你所说的:从言语中生成富有表情的形象
链接:https://arxiv.org/abs/2511.03423

作者:Jiyoung Lee, Song Park, Sanghyuk Chun, Soo-Whan Chung
备注:In progress
摘要:本文提出了VoxStudio,第一个统一的端到端的语音到图像模型,通过联合对齐语言和非语言信息,直接从口头描述中生成富有表现力的图像。其核心是语音信息瓶颈(SIB)模块,该模块将原始语音压缩为紧凑的语义标记,保留韵律和情感细微差别。通过直接对这些令牌进行操作,VoxStudio消除了对额外的语音到文本系统的需要,该系统通常会忽略文本之外的隐藏细节,例如,语气或情绪。我们还发布了Voxtosset,这是一个通过先进的TTS引擎构建的大规模配对情感语音图像数据集,可以生成丰富的表达性话语。在SpokenCOCO、Flickr8kAudio和Voxslave集基准测试上的综合实验证明了我们方法的可行性,并突出了关键挑战,包括情感一致性和语言模糊性,为未来的研究铺平了道路。
摘要:This paper proposes VoxStudio, the first unified and end-to-end speech-to-image model that generates expressive images directly from spoken descriptions by jointly aligning linguistic and paralinguistic information. At its core is a speech information bottleneck (SIB) module, which compresses raw speech into compact semantic tokens, preserving prosody and emotional nuance. By operating directly on these tokens, VoxStudio eliminates the need for an additional speech-to-text system, which often ignores the hidden details beyond text, e.g., tone or emotion. We also release VoxEmoset, a large-scale paired emotional speech-image dataset built via an advanced TTS engine to affordably generate richly expressive utterances. Comprehensive experiments on the SpokenCOCO, Flickr8kAudio, and VoxEmoset benchmarks demonstrate the feasibility of our method and highlight key challenges, including emotional consistency and linguistic ambiguity, paving the way for future research.


【2】Open Source State-Of-the-Art Solution for Romanian Speech Recognition
标题:罗马尼亚语音识别开源最先进的解决方案
链接:https://arxiv.org/abs/2511.03361

作者:Gabriel Pirlogeanu, Alexandru-Lucian Georgescu, Horia Cucu
备注:13th Conference on Speech Technology and Human-Computer Dialogue (SpeD 2025), Cluj-Napoca, Romania
摘要:在这项工作中,我们提出了一个新的国家的最先进的罗马尼亚语自动语音识别(ASR)系统的基础上,NVIDIA的FastConformer架构-探索这里首次在罗马尼亚语的背景下。我们训练我们的模型在一个大的语料库上,大部分是弱监督的transmitting,总共超过2,600小时的语音。利用具有连接时间分类(CTC)和令牌持续时间转换器(TDT)分支的混合解码器,我们评估了一系列解码策略,包括贪婪,ALSD和CTC波束搜索与6-gram令牌级语言模型。我们的系统在所有罗马尼亚评估基准中实现了最先进的性能,包括阅读,自发和特定领域的语音,与以前性能最好的系统相比,相对WER减少了27%。除了提高转录准确性外,我们的方法还展示了实用的解码效率,使其适用于低延迟ASR应用的研究和部署。
摘要:In this work, we present a new state-of-the-art Romanian Automatic Speech Recognition (ASR) system based on NVIDIA's FastConformer architecture--explored here for the first time in the context of Romanian. We train our model on a large corpus of, mostly, weakly supervised transcriptions, totaling over 2,600 hours of speech. Leveraging a hybrid decoder with both Connectionist Temporal Classification (CTC) and Token-Duration Transducer (TDT) branches, we evaluate a range of decoding strategies including greedy, ALSD, and CTC beam search with a 6-gram token-level language model. Our system achieves state-of-the-art performance across all Romanian evaluation benchmarks, including read, spontaneous, and domain-specific speech, with up to 27% relative WER reduction compared to previous best-performing systems. In addition to improved transcription accuracy, our approach demonstrates practical decoding efficiency, making it suitable for both research and deployment in low-latency ASR applications.


【3】audio2chart: End to End Audio Transcription into playable Guitar Hero charts
标题:audio 2 chart:端到端音频转录到可玩的吉他英雄排行榜
链接:https://arxiv.org/abs/2511.03337

作者:Riccardo Tripodi
摘要:这篇文章介绍了audio2chart,一个直接从原始音频自动生成Guitar Hero风格图表的框架。该任务被形式化为序列预测问题,其中模型被训练以生成与离散时间步长上的音频对齐的离散图表令牌。无条件的基线展示了强大的预测性能,而音频调节的添加在基于准确性的度量上产生了一致的改进。这项工作表明,将音频调节是可行的和有效的,以提高自动图表生成的音符预测。用于训练和推理的完整代码库在GitHub上公开提供,支持对神经图表生成的可重复研究。一系列预训练模型在Hugging Face上发布。
摘要:This work introduces audio2chart, a framework for the automatic generation of Guitar Hero style charts directly from raw audio. The task is formalized as a sequence prediction problem, where models are trained to generate discrete chart tokens aligned with the audio on discrete time steps. An unconditional baseline demonstrates strong predictive performance, while the addition of audio conditioning yields consistent improvements across accuracy based metrics. This work demonstrates that incorporating audio conditioning is both feasible and effective for improving note prediction in automatic chart generation. The complete codebase for training and inference is publicly available on GitHub supporting reproducible research on neural chart generation. A family of pretrained models is released on Hugging Face.


【4】TASU: Text-Only Alignment for Speech Understanding
标题:TASU:语音理解的纯文本对齐
链接:https://arxiv.org/abs/2511.03310

作者:Jing Peng, Yi Yang, Xu Li, Yu Xi, Quanwei Tang, Yangui Fang, Junjie Li, Kai Yu
备注:This paper is submitted to ICASSP 2026
摘要:语音大语言模型(Speech LLM)的最新进展为跨各种语音理解任务的统一架构铺平了道路。然而,主流的对齐范式严重依赖于大规模的音频文本配对数据和计算密集型训练,但往往表现出有限的泛化到看不见的领域或任务。为了解决这些限制,我们提出了TASU(纯文本对齐语音理解),一种新的对齐范式,可以利用不成对的文本数据来指导跨模态对齐。实验表明,TASU实现了竞争性zero-shot语音识别。利用这一特性,它可以进一步作为课程学习中的预训练阶段,增强语音识别中的领域泛化。最终,TASU可以将其zero-shot泛化扩展到广泛的语音理解任务,并且在MMSU基准上显著优于包括GLM-4-Voice和Step-Audio在内的突出的语音LLM,从而将TASU确立为语音LLM的高效且可扩展的对齐范例。
摘要:Recent advances in Speech Large Language Models (Speech LLMs) have paved the way for unified architectures across diverse speech understanding tasks. However, prevailing alignment paradigms rely heavily on large-scale audio-text paired data and computationally intensive training, yet often exhibit limited generalization to unseen domains or tasks. To address these limitations, we propose TASU (Text-only Alignment for Speech Understanding), a novel alignment paradigm that can leverage only unpaired text data to guide cross-modal alignment. Experiments show that TASU achieves competitive zero-shot speech recognition. Leveraging this property, it can further function as a pre-training stage in curriculum learning, enhancing domain generalization in speech recognition. Ultimately, TASU can extend its zero-shot generalization to a wide range of speech understanding tasks and notably outperforms prominent Speech LLMs including GLM-4-Voice and Step-Audio on the MMSU benchmark, establishing TASU as an efficient and scalable alignment paradigm for Speech LLMs.


【5】Speech-Based Prioritization for Schizophrenia Intervention
标题:基于言语的精神分裂症干预优先顺序
链接:https://arxiv.org/abs/2511.03086

作者:Gowtham Premananth, Philip Resnik, Sonia Bansal, Deanna L.Kelly, Carol Espy-Wilson
备注:Submitted for ICASSP 2026
摘要:数以百万计的人患有心理健康状况,但由于临床资源有限和劳动密集型评估方法,许多人仍未得到诊断或接受延迟治疗。虽然大多数机器辅助方法侧重于诊断分类,但估计症状严重程度对于优先考虑护理至关重要,特别是在资源有限的环境中。基于语音的人工智能提供了一种可扩展的替代方案,可以实现自动化、连续和远程监控,减少对主观自我报告和耗时评估的依赖。在本文中,我们介绍了一个基于语音的模型,成对比较精神分裂症症状的严重程度,利用发音和声学特征。这些比较用于通过Bradley-Terry模型生成严重程度排名。我们的方法在基于排名的指标上优于以前基于回归的模型,为临床分诊和优先级排序提供了更有效的解决方案。
摘要:Millions of people suffer from mental health conditions, yet many remain undiagnosed or receive delayed care due to limited clinical resources and labor-intensive assessment methods. While most machine-assisted approaches focus on diagnostic classification, estimating symptom severity is essential for prioritizing care, particularly in resource-constrained settings. Speech-based AI provides a scalable alternative by enabling automated, continuous, and remote monitoring, reducing reliance on subjective self-reports and time-consuming evaluations. In this paper, we introduce a speech-based model for pairwise comparison of schizophrenia symptom severity, leveraging articulatory and acoustic features. These comparisons are used to generate severity rankings via the Bradley-Terry model. Our approach outperforms previous regression-based models on ranking-based metrics, offering a more effective solution for clinical triage and prioritization.


【6】Quantifying Articulatory Coordination as a Biomarker for Schizophrenia
标题:量化关节协调作为精神分裂症的生物标志物
链接:https://arxiv.org/abs/2511.03084

作者:Gowtham Premananth, Carol Espy-Wilson
备注:Submitted to ICASSP 2026
摘要:人工智能(AI)和深度学习的进步提高了医疗保健的诊断能力,但有限的可解释性继续阻碍临床应用。精神分裂症是一种复杂的疾病,具有多种症状,包括言语紊乱和社交退缩,需要能够捕捉症状严重程度并提供超越二元诊断的临床有意义的见解的工具。在这里,我们提出了一个可解释的框架,利用发音语音特征,通过特征谱差异图和加权和指数衰减(WSED)来量化声道协调。本征谱图有效地区分了复杂的简单的协调模式,和WSED分数可靠地分开这些群体,与模糊性限制在一个狭窄的范围内接近零。重要的是,WSED评分不仅与总体BPRS严重程度相关,而且与阳性和阴性症状之间的平衡相关,反映了具有明显阳性症状的受试者的更复杂的协调性和更强的阴性症状的相反趋势。这种方法为精神分裂症提供了一个透明的,严重程度敏感的生物标志物,推进了临床可解释的基于语音的评估工具的潜力。
摘要:Advances in artificial intelligence (AI) and deep learning have improved diagnostic capabilities in healthcare, yet limited interpretability continues to hinder clinical adoption. Schizophrenia, a complex disorder with diverse symptoms including disorganized speech and social withdrawal, demands tools that capture symptom severity and provide clinically meaningful insights beyond binary diagnosis. Here, we present an interpretable framework that leverages articulatory speech features through eigenspectra difference plots and a weighted sum with exponential decay (WSED) to quantify vocal tract coordination. Eigenspectra plots effectively distinguished complex from simpler coordination patterns, and WSED scores reliably separated these groups, with ambiguity confined to a narrow range near zero. Importantly, WSED scores correlated not only with overall BPRS severity but also with the balance between positive and negative symptoms, reflecting more complex coordination in subjects with pronounced positive symptoms and the opposite trend for stronger negative symptoms. This approach offers a transparent, severity-sensitive biomarker for schizophrenia, advancing the potential for clinically interpretable speech-based assessment tools.


【7】Step-Audio-EditX Technical Report
标题:Step-Audio-EditX技术报告
链接:https://arxiv.org/abs/2511.03601

作者:Chao Yan, Boyong Wu, Peng Yang, Pengfei Tan, Guoqiang Hu, Yuxin Zhang, Xiangyu (Tony)Zhang, Fei Tian, Xuerui Yang, Xiangyu Zhang, Daxin Jiang, Gang Yu
摘要:Step-Audio-EditX是第一个开源的基于LLM的音频模型,擅长表达和迭代音频编辑,包括情感,说话风格和非语言学以及强大的zero-shot文本到语音(TTS)功能。我们的核心创新在于只利用大利润率的合成数据,这避免了对基于嵌入的先验或辅助模块的需求。这种大幅度的学习方法可以实现迭代控制和跨声音的高表现力,并且代表了传统的代表级解纠缠的基本支点。测试结果表明,Step-Audio-EditX在情感编辑和其他细粒度控制任务方面优于MiniMax-2.6-hd和Doubao-Seed-TTS-2.0。
摘要:We present Step-Audio-EditX, the first open-source LLM-based audio model excelling at expressive and iterative audio editing encompassing emotion, speaking style, and paralinguistics alongside robust zero-shot text-to-speech (TTS) capabilities.Our core innovation lies in leveraging only large-margin synthetic data, which circumvents the need for embedding-based priors or auxiliary modules. This large-margin learning approach enables both iterative control and high expressivity across voices, and represents a fundamental pivot from the conventional focus on representation-level disentanglement. Evaluation results demonstrate that Step-Audio-EditX surpasses both MiniMax-2.6-hd and Doubao-Seed-TTS-2.0 in emotion editing and other fine-grained control tasks.


【8】A Computational Approach to Analyzing Disrupted Language in Schizophrenia: Integrating Surprisal and Coherence Measures
标题:分析精神分裂症语言中断的计算方法:整合惊讶和连贯措施
链接:https://arxiv.org/abs/2511.03089

作者:Gowtham Premananth, Carol Espy-Wilson
备注:Submitted to ICASSP 2026
摘要:语言中断是精神分裂症症状的众所周知的影响之一。它们通常表现为言语混乱和语篇连贯性受损。这些自发性语言产生的异常反映了潜在的认知障碍,并有可能作为精神分裂症症状严重程度和诊断的客观标志。本研究的重点是如何这些语言中断可以在两个计算语言学的措施:句法和语义连贯的特点。通过使用计算模型计算语言的语义连贯性,本研究探讨了精神分裂症患者和健康对照者之间的差异。此外,这项研究提供了进一步了解如何在这些语言措施的语言中断改变不同程度的精神分裂症症状的严重程度。
摘要:Language disruptions are one of the well-known effects of schizophrenia symptoms. They are often manifested as disorganized speech and impaired discourse coherence. These abnormalities in spontaneous language production reflect underlying cognitive disturbances and have the potential to serve as objective markers for symptom severity and diagnosis of schizophrenia. This study focuses on how these language disruptions can be characterized in terms of two computational linguistic measures: surprisal and semantic coherence. By computing surprisal and semantic coherence of language using computational models, this study investigates how they differ between subjects with schizophrenia and healthy controls. Furthermore, this study provides further insight into how language disruptions in terms of these linguistic measures change with varying degrees of schizophrenia symptom severity.


机器翻译由腾讯交互翻译提供,仅供参考