今日论文合集:cs.SD语音8篇,eess.AS音频处理9篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】AudioStory: Generating Long-Form Narrative Audio with Large Language Models
标题:AudioStory:使用大型语言模型生成长篇叙事音频
链接:https://arxiv.org/abs/2508.20088

作者:Yuxin Guo, Teng Wang, Yuying Ge, Shijie Ma, Yixiao Ge, Wei Zou, Ying Shan
摘要:文本到音频(TTA)生成的最新进展擅长合成短音频片段,但与长形式的叙事音频,这需要时间的连贯性和组成推理斗争。为了解决这一差距,我们提出了AudioStory,这是一个统一的框架,它将大型语言模型(LLM)与TTA系统集成在一起,以生成结构化的长篇音频叙事。AudioStory具有强大的推理能力。它采用LLM将复杂的叙事查询分解为具有上下文线索的时间顺序子任务,从而实现连贯的场景转换和情感基调一致性。AudioStory有两个吸引人的特点:(1)解耦桥接机制:AudioStory将LLM扩散器协作分解为两个专门的组件,即,用于事件内语义对齐的桥接查询和用于跨事件一致性保持的残留查询。(2)端到端培训:通过将指令理解和音频生成统一在一个端到端框架内,AudioStory消除了对模块化培训管道的需求,同时增强了组件之间的协同作用。此外,我们建立了一个基准AudioStory-10 K,涵盖了不同的领域,如动画音景和自然声音叙事。大量的实验表明,AudioStory在单音频生成和叙事音频生成方面都具有优越性,在跟随能力和音频保真度方面都超过了之前的TTA基线。我们的代码可在https://github.com/TencentARC/AudioStory上获得
摘要:Recent advances in text-to-audio (TTA) generation excel at synthesizing short audio clips but struggle with long-form narrative audio, which requires temporal coherence and compositional reasoning. To address this gap, we propose AudioStory, a unified framework that integrates large language models (LLMs) with TTA systems to generate structured, long-form audio narratives. AudioStory possesses strong instruction-following reasoning generation capabilities. It employs LLMs to decompose complex narrative queries into temporally ordered sub-tasks with contextual cues, enabling coherent scene transitions and emotional tone consistency. AudioStory has two appealing features: (1) Decoupled bridging mechanism: AudioStory disentangles LLM-diffuser collaboration into two specialized components, i.e., a bridging query for intra-event semantic alignment and a residual query for cross-event coherence preservation. (2) End-to-end training: By unifying instruction comprehension and audio generation within a single end-to-end framework, AudioStory eliminates the need for modular training pipelines while enhancing synergy between components. Furthermore, we establish a benchmark AudioStory-10K, encompassing diverse domains such as animated soundscapes and natural sound narratives. Extensive experiments show the superiority of AudioStory on both single-audio generation and narrative audio generation, surpassing prior TTA baselines in both instruction-following ability and audio fidelity. Our code is available at https://github.com/TencentARC/AudioStory


【2】The IRMA Dataset: A Structured Audio-MIDI Corpus for Iranian Classical Music
标题:IRMA数据集:伊朗古典音乐的结构化音频数据库
链接:https://arxiv.org/abs/2508.19876

作者:Sepideh Shafiei, Shapour Hakam
摘要:我们提出了IRMA数据集(伊朗Radif音频),一个多层次的,开放式的语料库,专为伊朗古典音乐的计算研究,特别强调radif,一个结构化的曲目的模态旋律单位中央教学和性能。该数据集结合了符号化的符号表示,短语级的音频对齐,PDF格式的音乐学transmittance,以及从一系列表演者和学者那里收集的理论信息的比较表。我们概述了多阶段的建设过程中,包括段注释,对齐方法,和一个结构化的标识符代码系统,以参考个人的音乐单位。当前版本包括Karimi的完整radif; Mirza Abdollah的radif中的音频文件和元数据;由Payvar和Fereyduni转录的Davami声乐radif中的精选片段;以及一个专门的部分,其中包括由20世纪着名声乐家表演的tahrir演奏的音频示例。虽然符号和分析组件是在开放获取许可证(CC BY-NC 4.0)下发布的,但一些引用的音频记录和第三方转录引用使用唱片信息,使用户能够独立找到原始材料,等待版权许可。作为学术档案和计算分析的资源,该数据集支持民族音乐学,教育学,符号音频研究,文化遗产保护和人工智能驱动的任务(如自动转录和音乐生成)的应用。我们欢迎合作和反馈,以支持其持续的改进和更广泛的整合到音乐学和机器学习工作流程。
摘要:We present the IRMA Dataset (Iranian Radif MIDI Audio), a multi-level, open-access corpus designed for the computational study of Iranian classical music, with a particular emphasis on the radif, a structured repertoire of modal-melodic units central to pedagogy and performance. The dataset combines symbolic MIDI representations, phrase-level audio-MIDI alignment, musicological transcriptions in PDF format, and comparative tables of theoretical information curated from a range of performers and scholars. We outline the multi-phase construction process, including segment annotation, alignment methods, and a structured system of identifier codes to reference individual musical units. The current release includes the complete radif of Karimi; MIDI files and metadata from Mirza Abdollah's radif; selected segments from the vocal radif of Davami, as transcribed by Payvar and Fereyduni; and a dedicated section featuring audio-MIDI examples of tahrir ornamentation performed by prominent 20th-century vocalists. While the symbolic and analytical components are released under an open-access license (CC BY-NC 4.0), some referenced audio recordings and third-party transcriptions are cited using discographic information to enable users to locate the original materials independently, pending copyright permission. Serving both as a scholarly archive and a resource for computational analysis, this dataset supports applications in ethnomusicology, pedagogy, symbolic audio research, cultural heritage preservation, and AI-driven tasks such as automatic transcription and music generation. We welcome collaboration and feedback to support its ongoing refinement and broader integration into musicological and machine learning workflows.


【3】CompLex: Music Theory Lexicon Constructed by Autonomous Agents for Automatic Music Generation
标题:CompLex:自主智能体构建的音乐理论词典
链接:https://arxiv.org/abs/2508.19603

作者:Zhejing Hu, Yan Liu, Gong Chen, Bruce X.B. Yu
摘要:音乐领域的生成式人工智能已经取得了长足的进步,但它仍然没有达到自然语言处理领域的实质性成就,这主要是由于音乐数据的可用性有限。知识知情的方法已被证明可以提高音乐生成模型的性能,即使只有几条音乐知识被集成。本文试图在人工智能驱动的音乐生成任务中利用综合音乐理论,例如算法作曲和风格转换,传统上需要使用现有技术进行大量手动工作。我们介绍了一种新的自动音乐词典构建模型,生成一个词典,名为CompLex,包括37,432项来自9个手动输入的类别关键字和5个句子提示模板。提出了一种新的多智能体算法来自动检测和消除幻觉。CompLex在三种最先进的文本到音乐生成模型中展示了令人印象深刻的性能改进,包括符号和基于音频的方法。此外,我们评估CompLex的完整性,准确性,非冗余性和可执行性,确认它具有有效的词典的关键特征。
摘要:Generative artificial intelligence in music has made significant strides, yet it still falls short of the substantial achievements seen in natural language processing, primarily due to the limited availability of music data. Knowledge-informed approaches have been shown to enhance the performance of music generation models, even when only a few pieces of musical knowledge are integrated. This paper seeks to leverage comprehensive music theory in AI-driven music generation tasks, such as algorithmic composition and style transfer, which traditionally require significant manual effort with existing techniques. We introduce a novel automatic music lexicon construction model that generates a lexicon, named CompLex, comprising 37,432 items derived from just 9 manually input category keywords and 5 sentence prompt templates. A new multi-agent algorithm is proposed to automatically detect and mitigate hallucinations. CompLex demonstrates impressive performance improvements across three state-of-the-art text-to-music generation models, encompassing both symbolic and audio-based methods. Furthermore, we evaluate CompLex in terms of completeness, accuracy, non-redundancy, and executability, confirming that it possesses the key characteristics of an effective lexicon.


【4】MQAD: A Large-Scale Question Answering Dataset for Training Music Large Language Models
标题:MQAD:用于训练音乐大型语言模型的大规模问题回答数据集
链接:https://arxiv.org/abs/2508.19514

作者:Zhihao Ouyang, Ju-Chiang Wang, Daiyu Zhang, Bin Chen, Shangjie Li, Quan Lin
摘要:问答(QA)是人类理解一段音乐音频的自然方法。然而,对于机器来说,访问涵盖音乐各个方面的大规模数据集是至关重要的,但由于这种类型的公开可用音乐数据的稀缺性,因此具有挑战性。本文介绍了MQAD,一个建立在百万歌曲数据集(MSD)上的音乐QA数据集,包含丰富的音乐特征,包括节拍,和弦,基调,结构,乐器和流派-跨越270,000首曲目,具有近300万个不同的问题和标题。MQAD通过提供详细的随时间变化的音乐信息(如和弦和小节)来区分自己,从而可以探索歌曲中音乐的内在结构。为了编译MQAD,我们的方法利用专门的音乐信息检索(MIR)模型来提取更高级别的音乐特征和大型语言模型(LLM)来生成自然语言QA对。然后,我们利用一个多模态LLM,它集成了LLaMA 2和Whisper架构,以及新的主观指标来评估MQAD的性能。在实验中,我们在MQAD上训练的模型展示了传统音乐音频字幕方法的进步。数据集和代码可在https://github.com/oyzh888/MQAD上获得。
摘要:Question-answering (QA) is a natural approach for humans to understand a piece of music audio. However, for machines, accessing a large-scale dataset covering diverse aspects of music is crucial, yet challenging, due to the scarcity of publicly available music data of this type. This paper introduces MQAD, a music QA dataset built on the Million Song Dataset (MSD), encompassing a rich array of musical features, including beat, chord, key, structure, instrument, and genre -- across 270,000 tracks, featuring nearly 3 million diverse questions and captions. MQAD distinguishes itself by offering detailed time-varying musical information such as chords and sections, enabling exploration into the inherent structure of music within a song. To compile MQAD, our methodology leverages specialized Music Information Retrieval (MIR) models to extract higher-level musical features and Large Language Models (LLMs) to generate natural language QA pairs. Then, we leverage a multimodal LLM that integrates the LLaMA2 and Whisper architectures, along with novel subjective metrics to assess the performance of MQAD. In experiments, our model trained on MQAD demonstrates advancements over conventional music audio captioning approaches. The dataset and code are available at https://github.com/oyzh888/MQAD.


【5】Infant Cry Detection In Noisy Environment Using Blueprint Separable Convolutions and Time-Frequency Recurrent Neural Network
标题:使用蓝图可分离卷积和时频回归神经网络在噪音环境中检测婴儿哭声
链接:https://arxiv.org/abs/2508.19308

作者:Haolin Yu, Yanxiong Li
摘要:婴儿哭声检测是婴儿护理系统的重要组成部分。在本文中,我们提出了一个轻量级的和强大的婴儿哭声检测方法。该方法利用蓝图可分离卷积来降低计算复杂度,并利用时频递归神经网络进行自适应去噪。该方法的总体框架是一个多尺度卷积递归神经网络,通过有效的空间注意机制和对比度感知通道注意模块增强网络,从对数梅尔谱图的输入特征中获取局部和全局信息。采用多个公共数据集来创建多样化和代表性的数据集,并使用环境腐败技术来生成现实世界中遇到的噪声样本。结果表明,我们的方法超过了许多国家的最先进的方法在准确性,F1分数,和复杂性在各种信噪比条件下。代码在https://github.com/fhfjsd1/ICD_MMSP。
摘要:Infant cry detection is a crucial component of baby care system. In this paper, we propose a lightweight and robust method for infant cry detection. The method leverages blueprint separable convolutions to reduce computational complexity, and a time-frequency recurrent neural network for adaptive denoising. The overall framework of the method is structured as a multi-scale convolutional recurrent neural network, which is enhanced by efficient spatial attention mechanism and contrast-aware channel attention module, and acquire local and global information from the input feature of log Mel-spectrogram. Multiple public datasets are adopted to create a diverse and representative dataset, and environmental corruption techniques are used to generate the noisy samples encountered in real-world scenarios. Results show that our method exceeds many state-of-the-art methods in accuracy, F1-score, and complexity under various signal-to-noise ratio conditions. The code is at https://github.com/fhfjsd1/ICD_MMSP.


【6】Beat-Based Rhythm Quantization of MIDI Performances
标题:MIDI表演的基于节拍的节奏量化
链接:https://arxiv.org/abs/2508.19262

作者:Maximilian Wachter, Sebastian Murgul, Michael Heizmann
备注:Accepted to the Late Breaking Demo Papers of the 1st AES International Conference on Artificial Intelligence and Machine Learning for Audio (AIMLA LBDP), 2025
摘要:我们提出了一种基于变换器的节奏量化模型,该模型将节拍和强拍信息整合到度量对齐的人类可读分数中。我们提出了一种基于节拍的预处理方法,将分数和性能数据转换为统一的令牌表示。我们优化我们的模型架构和数据表示,并对钢琴和吉他演奏进行训练。我们的模型超过了基于MUSTER指标的最先进的性能。
摘要:We propose a transformer-based rhythm quantization model that incorporates beat and downbeat information to quantize MIDI performances into metrically-aligned, human-readable scores. We propose a beat-based preprocessing method that transfers score and performance data into a unified token representation. We optimize our model architecture and data representation and train on piano and guitar performances. Our model exceeds state-of-the-art performance based on the MUSTER metric.


【7】MuSpike: A Benchmark and Evaluation Framework for Symbolic Music Generation with Spiking Neural Networks
标题:MuSpike:利用Spiking神经网络生成符号音乐的基准和评估框架
链接:https://arxiv.org/abs/2508.19251

作者:Qian Liang, Menghaoran Tang, Yi Zeng
摘要:人工神经网络在符号音乐生成方面取得了迅速的进展,但在尖峰神经网络(SNN)的生物学上似乎合理的领域中仍然没有得到充分的探索,其中缺乏标准化的基准和全面的评估方法。为了解决这一差距,我们引入了MuSpike,这是一个统一的基准和评估框架,它系统地评估了五个典型数据集上的五个代表性SNN架构(SNN-CNN,SNN-RNN,SNN-LSTM,SNN-GAN和SNN-Transformer),涵盖音调,结构,情感和风格变化。MuSpike强调综合评估,将既定的客观指标与大规模的听力研究相结合。我们提出了新的主观指标,针对音乐的印象,自传体的协会,和个人喜好,捕捉知觉的尺寸往往被忽视在以前的工作。结果表明:(1)不同的SNN模型在评价维度上表现出不同的优势;(2)不同音乐背景的参与者表现出不同的感知模式,专家对人工智能创作的音乐表现出更大的容忍度;以及(3)客观评价和主观评价之间存在明显的不一致,强调了纯统计指标的局限性,并强调了人类感知判断在评估音乐质量方面的价值。MuSpike为符号音乐生成中的SNN模型提供了第一个系统的基准和系统的评估框架,为未来研究生物学上合理的和认知上接地的音乐生成奠定了坚实的基础。
摘要:Symbolic music generation has seen rapid progress with artificial neural networks, yet remains underexplored in the biologically plausible domain of spiking neural networks (SNNs), where both standardized benchmarks and comprehensive evaluation methods are lacking. To address this gap, we introduce MuSpike, a unified benchmark and evaluation framework that systematically assesses five representative SNN architectures (SNN-CNN, SNN-RNN, SNN-LSTM, SNN-GAN and SNN-Transformer) across five typical datasets, covering tonal, structural, emotional, and stylistic variations. MuSpike emphasizes comprehensive evaluation, combining established objective metrics with a large-scale listening study. We propose new subjective metrics, targeting musical impression, autobiographical association, and personal preference, that capture perceptual dimensions often overlooked in prior work. Results reveal that (1) different SNN models exhibit distinct strengths across evaluation dimensions; (2) participants with different musical backgrounds exhibit diverse perceptual patterns, with experts showing greater tolerance toward AI-composed music; and (3) a noticeable misalignment exists between objective and subjective evaluations, highlighting the limitations of purely statistical metrics and underscoring the value of human perceptual judgment in assessing musical quality. MuSpike provides the first systematic benchmark and systemic evaluation framework for SNN models in symbolic music generation, establishing a solid foundation for future research into biologically plausible and cognitively grounded music generation.


【8】FLASepformer: Efficient Speech Separation with Gated Focused Linear Attention Transformer
标题:FLASepformer:具有门控聚焦线性注意力Transformer的高效语音分离
链接:https://arxiv.org/abs/2508.19528

作者:Haoxu Wang, Yiheng Jiang, Gang Qiao, Pengteng Shi, Biao Tian
备注:Accepted by Interspeech 2025
摘要:语音分离一直面临着处理长时间序列的挑战。过去的方法试图减少序列长度并使用Transformer来捕获全局信息。然而,由于注意力模块的二次时间复杂度,内存使用和推理时间仍然随着更长的片段而显着增加。为了解决这个问题,我们引入了Focused Linear Attention,并构建了具有线性复杂度的FLASepformer,以实现高效的语音分离。受SepReformer和TF-Locoformer的启发,我们有两个变体:FLA-SepReformer和FLA-TFLocoformer。我们还添加了一个新的门控模块,以进一步提高性能。在各种数据集上的实验结果表明,FLASepformer具有最先进的性能,内存消耗更少,推理速度更快。FLA-SepReformer-T/B/L将速度提高了2.29倍,1.91倍和1.49倍,GPU内存使用率分别为15.8%,20.9%和31.9%,证明了我们模型的有效性。
摘要:Speech separation always faces the challenge of handling prolonged time sequences. Past methods try to reduce sequence lengths and use the Transformer to capture global information. However, due to the quadratic time complexity of the attention module, memory usage and inference time still increase significantly with longer segments. To tackle this, we introduce Focused Linear Attention and build FLASepformer with linear complexity for efficient speech separation. Inspired by SepReformer and TF-Locoformer, we have two variants: FLA-SepReformer and FLA-TFLocoformer. We also add a new Gated module to improve performance further. Experimental results on various datasets show that FLASepformer matches state-of-the-art performance with less memory consumption and faster inference. FLA-SepReformer-T/B/L increases speed by 2.29x, 1.91x, and 1.49x, with 15.8%, 20.9%, and 31.9% GPU memory usage, proving our model's effectiveness.


eess.AS音频处理


【1】CAVEMOVE: An Acoustic Database for the Study of Voice-enabled Technologies inside Moving Vehicles
标题:CAVEMOVE:用于研究移动车辆内部语音技术的声学数据库
链接:https://arxiv.org/abs/2508.19691

作者:Nikolaos Stefanakis, Marinos Kalaitzakis, Andreas Symiakakis, Stefanos Papadakis, Despoina Pavlidi
摘要:在本文中,我们提出了一个声学数据库,旨在推动和支持移动车辆内的语音技术的研究。记录过程涉及(i)在静态条件下采集的声学脉冲响应的记录,以提供用于对语音和汽车音频分量进行建模的手段;(ii)在广泛的静态和运动条件下的声学噪声的记录。使用两种不同的麦克风配置记录数据,特别是(i)紧凑型麦克风阵列和(ii)分布式麦克风设置。我们简要描述了获取录音的条件,并提供了对Python API的深入了解,该API旨在支持移动车辆内语音技术的研究和开发。这个Python API的第一个版本和部分描述的数据集可以免费下载。
摘要:In this paper, we present an acoustic database, designed to drive and support research on voiced enabled technologies inside moving vehicles. The recording process involves (i) recordings of acoustic impulse responses, acquired under static conditions to provide the means for modeling the speech and car-audio components (ii) recordings of acoustic noise at a wide range of static and in-motion conditions. Data are recorded with two different microphone configurations, particularly (i) a compact microphone array and (ii) a distributed microphone setup. We briefly describe the conditions under which the recordings were acquired, and we provide insight into a Python API that we designed to support the research and development of voice-enabled technologies inside moving vehicles. The first version of this Python API and part of the described dataset are available for free download.


【2】Hybrid Decoding: Rapid Pass and Selective Detailed Correction for Sequence Models
标题:混合解码:序列模型的快速通过和选择性详细纠正
链接:https://arxiv.org/abs/2508.19671

作者:Yunkyu Lim, Jihwan Park, Hyung Yong Kim, Hanbin Lee, Byeong-Yeol Kim
备注:Accepted to ASRU 2025
摘要:最近,基于transformer的编码器-解码器模型在多语言语音识别中表现出了强大的性能。然而,解码器的自回归性质和大尺寸在推理期间引入显著的瓶颈。此外,虽然罕见,但重复可能发生并对识别准确性产生负面影响。为了解决这些挑战,我们提出了一种新的混合解码方法,既加速推理,又避免了重复的问题。我们的方法扩展了Transformer编码器-解码器架构,通过附加一个轻量级的,快速的解码器的预训练的编码器。在推断期间,快速解码器快速地生成输出,该输出然后被验证,并且如果必要的话,由Transformer解码器选择性地校正。这导致更快的解码和针对重复错误的改进的鲁棒性。LibriSpeech和GigaSpeech测试集上的实验表明,由于微调仅限于添加的解码器,我们的方法实现了与基线相当或更好的单词错误率,同时推理速度增加了一倍以上。
摘要:Recently, Transformer-based encoder-decoder models have demonstrated strong performance in multilingual speech recognition. However, the decoder's autoregressive nature and large size introduce significant bottlenecks during inference. Additionally, although rare, repetition can occur and negatively affect recognition accuracy. To tackle these challenges, we propose a novel Hybrid Decoding approach that both accelerates inference and alleviates the issue of repetition. Our method extends the transformer encoder-decoder architecture by attaching a lightweight, fast decoder to the pretrained encoder. During inference, the fast decoder rapidly generates an output, which is then verified and, if necessary, selectively corrected by the Transformer decoder. This results in faster decoding and improved robustness against repetitive errors. Experiments on the LibriSpeech and GigaSpeech test sets indicate that, with fine-tuning limited to the added decoder, our method achieves word error rates comparable to or better than the baseline, while more than doubling the inference speed.


【3】Lightweight speech enhancement guided target speech extraction in noisy multi-speaker scenarios
标题:有噪多说话人场景中轻量级语音增强引导目标语音提取
链接:https://arxiv.org/abs/2508.19583

作者:Ziling Huang, Junnan Wu, Lichun Fan, Zhenbo Luo, Jian Luan, Haixin Guan, Yanhua Long
备注:This paper has been submitted to ICASSP 2026. Copyright 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, including reprinting/republishing, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work. DOI will be added upon IEEE Xplore publication
摘要:目标语音提取(TSE)在单说话人加噪声和两说话人混合等相对简单的情况下已经取得了很好的效果,但在多说话人噪声环境下的效果仍然不理想。为了解决这个问题,我们引入了一个轻量级的语音增强模型,GTCRN,以更好地指导TSE在嘈杂的环境中。基于我们先前的具有竞争力的扬声器嵌入/无编码器框架SEF-PNet,我们提出了两个扩展:LGTSE和D-LGTSE。LGTSE通过在与注册语音进行上下文交互之前对输入噪声语音进行去噪来合并噪声不可知的注册指导,从而减少噪声干扰。D-LGTSE通过在训练过程中利用降噪语音作为额外的噪声输入,扩大噪声条件的动态范围,并使模型能够直接从失真信号中学习,进一步提高了系统对语音失真的鲁棒性。在Libri2Mix数据集上的实验表明,SISDR、PESQ和STOI分别有0.89 dB、0.16 dB和1.97%的显著改善,验证了该方法的有效性。我们的代码可在www.example.com上公开获取。
摘要:Target speech extraction (TSE) has achieved strong performance in relatively simple conditions such as one-speaker-plus-noise and two-speaker mixtures, but its performance remains unsatisfactory in noisy multi-speaker scenarios. To address this issue, we introduce a lightweight speech enhancement model, GTCRN, to better guide TSE in noisy environments. Building on our competitive previous speaker embedding/encoder-free framework SEF-PNet, we propose two extensions: LGTSE and D-LGTSE. LGTSE incorporates noise-agnostic enrollment guidance by denoising the input noisy speech before context interaction with enrollment speech, thereby reducing noise interference. D-LGTSE further improves system robustness against speech distortion by leveraging denoised speech as an additional noisy input during training, expanding the dynamic range of noisy conditions and enabling the model to directly learn from distorted signals. Furthermore, we propose a two-stage training strategy, first with GTCRN enhancement-guided pre-training and then joint fine-tuning, to fully exploit model potential.Experiments on the Libri2Mix dataset demonstrate significant improvements of 0.89 dB in SISDR, 0.16 in PESQ, and 1.97% in STOI, validating the effectiveness of our approach. Our code is publicly available at https://github.com/isHuangZiling/D-LGTSE.


【4】FLASepformer: Efficient Speech Separation with Gated Focused Linear Attention Transformer
标题:FLASepformer:具有门控聚焦线性注意力Transformer的高效语音分离
链接:https://arxiv.org/abs/2508.19528

作者:Haoxu Wang, Yiheng Jiang, Gang Qiao, Pengteng Shi, Biao Tian
备注:Accepted by Interspeech 2025
摘要:语音分离一直面临着处理长时间序列的挑战。过去的方法试图减少序列长度并使用Transformer来捕获全局信息。然而,由于注意力模块的二次时间复杂度,内存使用和推理时间仍然随着更长的片段而显着增加。为了解决这个问题,我们引入了Focused Linear Attention,并构建了具有线性复杂度的FLASepformer,以实现高效的语音分离。受SepReformer和TF-Locoformer的启发,我们有两个变体:FLA-SepReformer和FLA-TFLocoformer。我们还添加了一个新的门控模块,以进一步提高性能。在各种数据集上的实验结果表明,FLASepformer具有最先进的性能,内存消耗更少,推理速度更快。FLA-SepReformer-T/B/L将速度提高了2.29倍,1.91倍和1.49倍,GPU内存使用率分别为15.8%,20.9%和31.9%,证明了我们模型的有效性。
摘要:Speech separation always faces the challenge of handling prolonged time sequences. Past methods try to reduce sequence lengths and use the Transformer to capture global information. However, due to the quadratic time complexity of the attention module, memory usage and inference time still increase significantly with longer segments. To tackle this, we introduce Focused Linear Attention and build FLASepformer with linear complexity for efficient speech separation. Inspired by SepReformer and TF-Locoformer, we have two variants: FLA-SepReformer and FLA-TFLocoformer. We also add a new Gated module to improve performance further. Experimental results on various datasets show that FLASepformer matches state-of-the-art performance with less memory consumption and faster inference. FLA-SepReformer-T/B/L increases speed by 2.29x, 1.91x, and 1.49x, with 15.8%, 20.9%, and 31.9% GPU memory usage, proving our model's effectiveness.


【5】Audio-Visual Feature Synchronization for Robust Speech Enhancement in Hearing Aids
标题:视听特征同步用于助听器中的鲁棒语音增强
链接:https://arxiv.org/abs/2508.19483

作者:Nasir Saleem, Mandar Gogate, Kia Dashtipour, Adeel Hussain, Usman Anwar, Adewale Adetomi, Tughrul Arslan, Amir Hussain
备注:Preprint of the paper presented at Euronoise 2025 Malaga, Spain
摘要:用于助听器中实时语音增强的视听特征同步代表了一种改进语音可懂度和用户体验的渐进方法,特别是在强噪声背景中。这种方法将听觉信号与视觉线索相结合,利用这些模态的互补描述来提高语音清晰度。助听器中实时SE的视听特征同步可以使用高效的特征对准模块进一步优化。在这项研究中,一个轻量级的交叉注意模型通过利用大规模数据和简单的架构来学习鲁棒的视听表示。通过将轻量级交叉注意模型纳入AVSE框架,神经系统动态地强调音频和视觉模态的关键特征,从而实现定义的同步和改善的语音清晰度。所提出的AVSE模型不仅确保了噪声抑制和特征对齐的高性能,而且以最小的延迟(36 ms)和能耗实现了实时处理。AVSEC 3数据集上的评估显示了该模型的效率,在感知质量(PESQ:0.52),可懂度(STOI:19\%)和保真度(SI-SDR:10.10 dB)方面取得了显著的基线增益。
摘要:Audio-visual feature synchronization for real-time speech enhancement in hearing aids represents a progressive approach to improving speech intelligibility and user experience, particularly in strong noisy backgrounds. This approach integrates auditory signals with visual cues, utilizing the complementary description of these modalities to improve speech intelligibility. Audio-visual feature synchronization for real-time SE in hearing aids can be further optimized using an efficient feature alignment module. In this study, a lightweight cross-attentional model learns robust audio-visual representations by exploiting large-scale data and simple architecture. By incorporating the lightweight cross-attentional model in an AVSE framework, the neural system dynamically emphasizes critical features across audio and visual modalities, enabling defined synchronization and improved speech intelligibility. The proposed AVSE model not only ensures high performance in noise suppression and feature alignment but also achieves real-time processing with minimal latency (36ms) and energy consumption. Evaluations on the AVSEC3 dataset show the efficiency of the model, achieving significant gains over baselines in perceptual quality (PESQ:0.52), intelligibility (STOI:19\%), and fidelity (SI-SDR:10.10dB).


【6】TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
标题:TokenVerse++:通过动态任务激活实现灵活的多任务学习
链接:https://arxiv.org/abs/2508.19856

作者:Shashi Kumar, Srikanth Madikeri, Esaú Villatoro-Tello, Sergio Burdisso, Pradeep Rangappa, Andrés Carofilis, Petr Motlicek, Karthik Pandia, Shankar Venkatesan, Kadri Hacioğlu, Andreas Stolcke
备注:Accepted to IEEE ASRU 2025. Copyright©2025 IEEE
摘要:像TokenVerse这样的基于令牌的多任务框架要求所有训练语句都有所有任务的标签,这阻碍了它们利用部分注释数据集和有效扩展的能力。我们提出了TokenVerse++,它在XLSR-Transducer ASR模型的声学嵌入空间中引入了可学习的向量,用于动态任务激活。这种核心机制支持仅针对任务子集标记的话语进行训练,这是TokenVerse的一个关键优势。我们通过成功地将数据集与部分标签集成来证明这一点,特别是针对ASR和额外的任务,语言识别,提高整体性能。TokenVerse++在多个任务上实现了与TokenVerse相当或超过TokenVerse的结果,使其成为更实用的多任务替代方案,而不会牺牲ASR性能。
摘要:Token-based multitasking frameworks like TokenVerse require all training utterances to have labels for all tasks, hindering their ability to leverage partially annotated datasets and scale effectively. We propose TokenVerse++, which introduces learnable vectors in the acoustic embedding space of the XLSR-Transducer ASR model for dynamic task activation. This core mechanism enables training with utterances labeled for only a subset of tasks, a key advantage over TokenVerse. We demonstrate this by successfully integrating a dataset with partial labels, specifically for ASR and an additional task, language identification, improving overall performance. TokenVerse++ achieves results on par with or exceeding TokenVerse across multiple tasks, establishing it as a more practical multitask alternative without sacrificing ASR performance.


【7】CAMÕES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese
标题:CAMCLARIES:欧洲葡萄牙语的全面自动语音识别基准
链接:https://arxiv.org/abs/2508.19721

作者:Carlos Carvalho, Francisco Teixeira, Catarina Botelho, Anna Pompili, Rubén Solera-Ureña, Sérgio Paulo, Mariana Julião, Thomas Rolland, John Mendonça, Diogo Pereira, Isabel Trancoso, Alberto Abad
备注:Accepted to ASRU 2025
摘要:葡萄牙语自动语音识别的现有资源主要集中在巴西葡萄牙语上,而欧洲葡萄牙语(EP)和其他品种则未得到充分开发。为了弥合这一差距,我们介绍CAM\~OES,第一个开放的框架EP和其他葡萄牙品种。它包括(1)一个全面的评估基准,包括跨越多个领域的46小时EP测试数据;和(2)一组最先进的模型。对于后者,我们考虑多个基础模型,评估它们的zero-shot和微调性能,以及从头开始训练的E-Branchformer模型。一组精心策划的425小时EP用于微调和训练。我们的研究结果表明,微调的基础模型和E-branchformer之间的EP性能相当。此外,与最强的zero-shot基础模型相比,性能最佳的模型实现了35%以上的WER相对改善,为EP和其他品种建立了新的最先进的技术水平。
摘要:Existing resources for Automatic Speech Recognition in Portuguese are mostly focused on Brazilian Portuguese, leaving European Portuguese (EP) and other varieties under-explored. To bridge this gap, we introduce CAM\~OES, the first open framework for EP and other Portuguese varieties. It consists of (1) a comprehensive evaluation benchmark, including 46h of EP test data spanning multiple domains; and (2) a collection of state-of-the-art models. For the latter, we consider multiple foundation models, evaluating their zero-shot and fine-tuned performances, as well as E-Branchformer models trained from scratch. A curated set of 425h of EP was used for both fine-tuning and training. Our results show comparable performance for EP between fine-tuned foundation models and the E-Branchformer. Furthermore, the best-performing models achieve relative improvements above 35% WER, compared to the strongest zero-shot foundation model, establishing a new state-of-the-art for EP and other varieties.


【8】Beat-Based Rhythm Quantization of MIDI Performances
标题:MIDI表演的基于节拍的节奏量化
链接:https://arxiv.org/abs/2508.19262

作者:Maximilian Wachter, Sebastian Murgul, Michael Heizmann
备注:Accepted to the Late Breaking Demo Papers of the 1st AES International Conference on Artificial Intelligence and Machine Learning for Audio (AIMLA LBDP), 2025
摘要:我们提出了一种基于变换器的节奏量化模型,该模型将节拍和强拍信息整合到度量对齐的人类可读分数中。我们提出了一种基于节拍的预处理方法,将分数和性能数据转换为统一的令牌表示。我们优化我们的模型架构和数据表示,并对钢琴和吉他演奏进行训练。我们的模型超过了基于MUSTER指标的最先进的性能。
摘要:We propose a transformer-based rhythm quantization model that incorporates beat and downbeat information to quantize MIDI performances into metrically-aligned, human-readable scores. We propose a beat-based preprocessing method that transfers score and performance data into a unified token representation. We optimize our model architecture and data representation and train on piano and guitar performances. Our model exceeds state-of-the-art performance based on the MUSTER metric.


【9】MuSpike: A Benchmark and Evaluation Framework for Symbolic Music Generation with Spiking Neural Networks
标题:MuSpike:利用Spiking神经网络生成符号音乐的基准和评估框架
链接:https://arxiv.org/abs/2508.19251

作者:Qian Liang, Menghaoran Tang, Yi Zeng
摘要:人工神经网络在符号音乐生成方面取得了迅速的进展,但在尖峰神经网络(SNN)的生物学上似乎合理的领域中仍然没有得到充分的探索,其中缺乏标准化的基准和全面的评估方法。为了解决这一差距,我们引入了MuSpike,这是一个统一的基准和评估框架,它系统地评估了五个典型数据集上的五个代表性SNN架构(SNN-CNN,SNN-RNN,SNN-LSTM,SNN-GAN和SNN-Transformer),涵盖音调,结构,情感和风格变化。MuSpike强调综合评估,将既定的客观指标与大规模的听力研究相结合。我们提出了新的主观指标,针对音乐的印象,自传体的协会,和个人喜好,捕捉知觉的尺寸往往被忽视在以前的工作。结果表明:(1)不同的SNN模型在评价维度上表现出不同的优势;(2)不同音乐背景的参与者表现出不同的感知模式,专家对人工智能创作的音乐表现出更大的容忍度;以及(3)客观评价和主观评价之间存在明显的不一致,强调了纯统计指标的局限性,并强调了人类感知判断在评估音乐质量方面的价值。MuSpike为符号音乐生成中的SNN模型提供了第一个系统的基准和系统的评估框架,为未来研究生物学上合理的和认知上接地的音乐生成奠定了坚实的基础。
摘要:Symbolic music generation has seen rapid progress with artificial neural networks, yet remains underexplored in the biologically plausible domain of spiking neural networks (SNNs), where both standardized benchmarks and comprehensive evaluation methods are lacking. To address this gap, we introduce MuSpike, a unified benchmark and evaluation framework that systematically assesses five representative SNN architectures (SNN-CNN, SNN-RNN, SNN-LSTM, SNN-GAN and SNN-Transformer) across five typical datasets, covering tonal, structural, emotional, and stylistic variations. MuSpike emphasizes comprehensive evaluation, combining established objective metrics with a large-scale listening study. We propose new subjective metrics, targeting musical impression, autobiographical association, and personal preference, that capture perceptual dimensions often overlooked in prior work. Results reveal that (1) different SNN models exhibit distinct strengths across evaluation dimensions; (2) participants with different musical backgrounds exhibit diverse perceptual patterns, with experts showing greater tolerance toward AI-composed music; and (3) a noticeable misalignment exists between objective and subjective evaluations, highlighting the limitations of purely statistical metrics and underscoring the value of human perceptual judgment in assessing musical quality. MuSpike provides the first systematic benchmark and systemic evaluation framework for SNN models in symbolic music generation, establishing a solid foundation for future research into biologically plausible and cognitively grounded music generation.


机器翻译由腾讯交互翻译提供,仅供参考