今日论文合集:cs.SD语音15篇,eess.AS音频处理11篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】FoleySpace: Vision-Aligned Binaural Spatial Audio Generation
标题:FoleySpace:视觉对齐的双耳空间音频生成
链接:https://arxiv.org/abs/2508.12918

作者: Lei Zhao, Rujin Chen, Chi Zhang, Xiao-Lei Zhang, Xuelong Li
摘要:最近,随着AIGC的发展,基于深度学习的视频到音频(V2A)技术引起了人们的极大关注。然而,现有的研究大多集中在单声道音频生成,缺乏空间感知,而双耳空间音频生成技术,可以提供更强的沉浸感的探索仍然不足。为了解决这个问题,我们提出了FoleySpace,一个框架的视频到双耳音频生成,产生沉浸式和空间一致的立体声视觉信息的指导下。具体来说,我们开发了一种声源估计方法来确定每个视频帧中的声源2D坐标和深度,然后采用坐标映射机制将2D源位置转换为3D轨迹。该3D轨迹与由预训练的V2A模型生成的单声道音频一起用作扩散模型的调节输入以生成空间一致的双耳音频。为了支持动态声场的生成,我们基于记录的头部相关脉冲响应构建了一个训练数据集,其中包括各种声源移动场景。实验结果表明,该方法在空间感知一致性方面优于现有方法,有效地提高了视听体验的沉浸质量。
摘要:Recently, with the advancement of AIGC, deep learning-based video-to-audio (V2A) technology has garnered significant attention. However, existing research mostly focuses on mono audio generation that lacks spatial perception, while the exploration of binaural spatial audio generation technologies, which can provide a stronger sense of immersion, remains insufficient. To solve this problem, we propose FoleySpace, a framework for video-to-binaural audio generation that produces immersive and spatially consistent stereo sound guided by visual information. Specifically, we develop a sound source estimation method to determine the sound source 2D coordinates and depth in each video frame, and then employ a coordinate mapping mechanism to convert the 2D source positions into a 3D trajectory. This 3D trajectory, together with the monaural audio generated by a pre-trained V2A model, serves as a conditioning input for a diffusion model to generate spatially consistent binaural audio. To support the generation of dynamic sound fields, we constructed a training dataset based on recorded Head-Related Impulse Responses that includes various sound source movement scenarios. Experimental results demonstrate that the proposed method outperforms existing approaches in spatial perception consistency, effectively enhancing the immersive quality of the audio-visual experience.


【2】MATPAC++: Enhanced Masked Latent Prediction for Self-Supervised Audio Representation Learning
标题:MARTPAC ++:用于自我监督音频表示学习的增强掩蔽潜伏预测
链接:https://arxiv.org/abs/2508.12709

作者:Aurian Quelennec, Pierre Chouteau, Geoffroy Peeters, Slim Essid
备注:Under review
摘要:掩蔽潜在预测已经成为自监督学习(SSL)的主要范式,特别是对于一般的音频和音乐表示学习。虽然最近的方法已经证明了强大的性能,预测模块的作用在这样的SSL系统的输出仍然主要被忽视,尽管是至关重要的解决手头的借口任务。特别是,这个模块应该能够处理音频内容中固有的模糊性,特别是当它由多个声源组成时。这项工作提出了一种新的增强:集成多项选择学习(MCL),以显式地建模预测模糊性和提高表示质量。我们建立在最近提出的MATPAC系统之上,用MCL改进其预测和无监督分类任务。我们广泛评估我们的方法,MATPAC++,通过多个下游任务的线性探测和AudioSet上的微调,采用统一的协议,使严格和公平的比较与最先进的SSL方法。结果表明,我们的建议达到国家的最先进的微调AudioSet和整体国家的最先进的下游任务的分数。此外,我们通过专门对音乐数据进行训练来检查领域专业化,我们的模型实现了最先进的性能,效率显著提高。
摘要:Masked latent prediction has emerged as a leading paradigm in self-supervised learning (SSL), especially for general audio and music representation learning. While recent methods have demonstrated strong performance, the role of the predictor module used at the output of such SSL systems remains mainly overlooked, despite being crucial for solving the pretext task at hand. In particular, this module should be able to deal with the ambiguity inherent in audio content, especially when it is composed of multiple sound sources. This work proposes a novel enhancement: integrating Multiple Choice Learning (MCL) to explicitly model prediction ambiguity and improve representation quality. We build on top of the recently proposed MATPAC system, improving its prediction and unsupervised classification pretext tasks with MCL. We extensively evaluate our method, MATPAC++, through both linear probing across multiple downstream tasks and fine-tuning on AudioSet, employing a unified protocol that enables rigorous and fair comparisons with state-of-the-art SSL approaches. Results show that our proposal achieves state-of-the-art when fine-tuned on AudioSet and overall state-of-the-art scores on downstream tasks. Additionally, we examine domain specialisation by training exclusively on music data, where our model achieves state-of-the-art performance with significantly improved efficiency.


【3】Exploring the Feasibility of LLMs for Automated Music Emotion Annotation
标题:LLM用于自动音乐情感标注的可行性研究
链接:https://arxiv.org/abs/2508.12626

作者:Meng Yang, Jon McCormack, Maria Teresa Llano, Wanchao Su
备注:Accepted to be published at ISMIR 2025
摘要:目前的音乐情感标注方法仍然严重依赖于人工标注,这一过程带来了巨大的资源和劳动力负担,严重限制了可用注释数据的规模。本研究探讨了使用大型语言模型(GPT-4 o)进行音乐情感标注的可行性和可靠性。在这项研究中,我们使用GPT-4 o在四象限效价唤醒框架中注释了GiantMIDI-Piano,这是一个经典的钢琴音乐数据集,并与三位人类专家提供的注释进行了比较。我们进行了广泛的评估,以评估GPT生成的音乐情感注释的性能和可靠性,包括标准准确性,加权准确性,占专家间的协议,注释者间的协议指标,和生成的标签的分布相似性。   虽然GPT的注释性能在整体准确性方面低于人类专家,并且在对特定情绪状态进行分类时表现出较少的细微差别,但评分员间可靠性指标表明GPT的可变性仍然在专家之间的自然分歧范围内。这些发现强调了基于GPT的注释的局限性和潜力:尽管其目前的缺点相对于人类的表现,其成本效益和效率使其成为一个有前途的可扩展的音乐情感注释的替代方案。
摘要:Current approaches to music emotion annotation remain heavily reliant on manual labelling, a process that imposes significant resource and labour burdens, severely limiting the scale of available annotated data. This study examines the feasibility and reliability of employing a large language model (GPT-4o) for music emotion annotation. In this study, we annotated GiantMIDI-Piano, a classical MIDI piano music dataset, in a four-quadrant valence-arousal framework using GPT-4o, and compared against annotations provided by three human experts. We conducted extensive evaluations to assess the performance and reliability of GPT-generated music emotion annotations, including standard accuracy, weighted accuracy that accounts for inter-expert agreement, inter-annotator agreement metrics, and distributional similarity of the generated labels.   While GPT's annotation performance fell short of human experts in overall accuracy and exhibited less nuance in categorizing specific emotional states, inter-rater reliability metrics indicate that GPT's variability remains within the range of natural disagreement among experts. These findings underscore both the limitations and potential of GPT-based annotation: despite its current shortcomings relative to human performance, its cost-effectiveness and efficiency render it a promising scalable alternative for music emotion annotation.


【4】Beyond Modality Limitations: A Unified MLLM Approach to Automated Speaking Assessment with Effective Curriculum Learning
标题:超越模式限制:统一的MLLM方法实现自动化口语评估和有效的课程学习
链接:https://arxiv.org/abs/2508.12591

作者:Yu-Hsuan Fang, Tien-Hong Lo, Yao-Ting Sung, Berlin Chen
备注:Accepted at IEEE ASRU 2025
摘要:传统的自动口语评估(ASA)系统表现出固有的模态限制:基于文本的方法缺乏声学信息,而基于音频的方法错过语义上下文。多模态大型语言模型(MLLM)通过在统一框架内同时处理音频和文本,为全面的ASA提供了前所未有的机会。本文首次对综合ASA的MLLM进行了系统的研究,展示了MLLM在内容和语言使用方面的优越性能。然而,对交付方面的评估揭示了独特的挑战,认为这需要专门的培训战略。因此,我们提出了语音优先多模态训练(SFMT),利用课程学习原则,在跨模态协同融合之前建立更强大的语音建模基础。在基准数据集上的一系列实验表明,基于MLLM的系统可以将整体评估性能从PCC值0.783提高到0.846。特别是,SFMT在评估交付方面表现出色,与传统的训练方法相比,绝对准确率提高了4%,这也为ASA铺平了新的道路。
摘要:Traditional Automated Speaking Assessment (ASA) systems exhibit inherent modality limitations: text-based approaches lack acoustic information while audio-based methods miss semantic context. Multimodal Large Language Models (MLLM) offer unprecedented opportunities for comprehensive ASA by simultaneously processing audio and text within unified frameworks. This paper presents a very first systematic study of MLLM for comprehensive ASA, demonstrating the superior performance of MLLM across the aspects of content and language use . However, assessment on the delivery aspect reveals unique challenges, which is deemed to require specialized training strategies. We thus propose Speech-First Multimodal Training (SFMT), leveraging a curriculum learning principle to establish more robust modeling foundations of speech before cross-modal synergetic fusion. A series of experiments on a benchmark dataset show MLLM-based systems can elevate the holistic assessment performance from a PCC value of 0.783 to 0.846. In particular, SFMT excels in the evaluation of the delivery aspect, achieving an absolute accuracy improvement of 4% over conventional training approaches, which also paves a new avenue for ASA.


【5】CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation
标题:CEM-Net:用于情感说话面部生成的交叉情感记忆网络
链接:https://arxiv.org/abs/2508.12368

作者:Pengna Li, Jingwen Fu, Yang Wu, Yuhan Liu, Sanping Zhou, Jinjun Wang
摘要:情感说话人脸生成的目的是在给定的参考图像动画的人脸,并生成一个说话的视频,匹配的内容和情感的驾驶音频。然而,现有的方法忽略了参考图像可能具有与音频情感冲突的强烈情感,导致严重的情感不准确和失真的生成结果。为了解决这个问题,我们引入了一个跨情感记忆网络(CEM-Net),旨在生成与驾驶音频对齐的情感说话的脸时,参考图像表现出强烈的情感。具体而言,音频情感增强模块(AEE)的设计与交叉重建训练策略,以增强音频情感,克服了干扰,从参考图像的情感。其次,由于参考图像无法提供音频情感下说话者面部运动的足够信息,因此利用情感桥接记忆模块(EBM)来补偿缺失的信息。该算法将参考图像情感到音频情感的表情位移引入到存储器中,并以跨情感特征作为查询,在推理时检索匹配位移。大量的实验表明,我们的CEM-Net可以合成具有表达力的,自然的和嘴唇同步的说话人脸视频,具有更好的情感准确性。
摘要:Emotional talking face generation aims to animate a human face in given reference images and generate a talking video that matches the content and emotion of driving audio. However, existing methods neglect that reference images may have a strong emotion that conflicts with the audio emotion, leading to severe emotion inaccuracy and distorted generated results. To tackle the issue, we introduce a cross-emotion memory network(CEM-Net), designed to generate emotional talking faces aligned with the driving audio when reference images exhibit strong emotion. Specifically, an Audio Emotion Enhancement module(AEE) is first devised with the cross-reconstruction training strategy to enhance audio emotion, overcoming the disruption from reference image emotion. Secondly, since reference images cannot provide sufficient facial motion information of the speaker under audio emotion, an Emotion Bridging Memory module(EBM) is utilized to compensate for the lacked information. It brings in expression displacement from the reference image emotion to the audio emotion and stores it in the memory.Given a cross-emotion feature as a query, the matching displacement can be retrieved at inference time. Extensive experiments have demonstrated that our CEM-Net can synthesize expressive, natural and lip-synced talking face videos with better emotion accuracy.


【6】Cross-Modal Knowledge Distillation with Multi-Level Data Augmentation for Low-Resource Audio-Visual Sound Event Localization and Detection
标题:具有多层数据增强的跨模式知识提炼用于低资源视听声音事件定位和检测
链接:https://arxiv.org/abs/2508.12334

备注:34 pages, 7 figures
摘要:本文提出了一种结合多层次数据增强的跨模态知识提取(CMKD)框架,用于低资源音视频(AV)声音事件定位和检测(SELD)。仅音频SELD模型充当教师,通过输出响应和中间特征表示将知识传递到AV学生模型。为了增强学习,通过混合从多个网络层随机选择的特征和为SELD任务定制的相关损失函数来应用数据增强。在DCASE 2023和2024 SELD数据集上的大量实验表明,该方法显著提高了AV SELD性能,在总体指标上相对于基线获得了22%~36%的相对增益。值得注意的是,我们的方法实现了与在更大数据集上训练的教师模型相当或更好的结果,在DCASE 2023和2024 SELD任务上都超过了最先进的方法。
摘要:This work presents a cross-modal knowledge distillation (CMKD) framework combined with multi-level data augmentation for low-resource audio-visual (AV) sound event localization and detection (SELD). An audio-only SELD model acts as the teacher, transferring knowledge to an AV student model through both output responses and intermediate feature representations. To enhance learning, data augmentation is applied by mixing features randomly selected from multiple network layers and associated loss functions tailored to the SELD task. Extensive experiments on the DCASE 2023 and 2024 SELD datasets show that the proposed method significantly improves AV SELD performance, yielding relative gains of 22%~36% in the overall metric over the baseline. Notably, our approach achieves results comparable to or better than teacher models trained on much larger datasets, surpassing state-of-the-art methods on both DCASE 2023 and 2024 SELD tasks.


【7】CarelessWhisper: Turning Whisper into a Causal Streaming Model
标题:CarelessWhisper:将Whisper转变为因果流媒体模型
链接:https://arxiv.org/abs/2508.12301

备注:17 pages, 7 Figures, This work has been submitted to the IEEE for possible publication
摘要:自动语音识别(ASR)已经取得了显著的进展,OpenAI Whisper和NVIDIA Canary等模型在离线转录方面实现了最先进的(SOTA)性能。然而,由于其架构和训练方法的限制,这些模型并不是为流式(在线或实时)转录而设计的。我们提出了一种方法,把Transformer编码器-解码器模型变成一个低延迟的流模型,是不关心未来的上下文。我们提出了一个分析,解释为什么它不是简单的转换编码器-解码器Transformer到低延迟流模型。我们提出的方法通过使用低秩自适应(LoRA)和弱对齐数据集对编码器和解码器进行微调,将现有的(非因果)编码器修改为因果编码器。然后,我们提出了一个更新的推理机制,利用微调因果编码器和解码器产生贪婪和波束搜索解码,并被证明是局部最优的。低延迟块大小(小于300毫秒)的实验表明,我们的微调模型优于现有的非微调流媒体方法在大多数情况下,同时使用较低的复杂性。此外,我们观察到我们的训练过程产生了更好的对齐,从而实现了提取单词级时间戳的简单方法。我们发布了我们的训练和推理代码,以及经过微调的模型,以支持流式ASR的进一步研究和开发。
摘要:Automatic Speech Recognition (ASR) has seen remarkable progress, with models like OpenAI Whisper and NVIDIA Canary achieving state-of-the-art (SOTA) performance in offline transcription. However, these models are not designed for streaming (online or real-time) transcription, due to limitations in their architecture and training methodology. We propose a method to turn the transformer encoder-decoder model into a low-latency streaming model that is careless about future context. We present an analysis explaining why it is not straightforward to convert an encoder-decoder transformer to a low-latency streaming model. Our proposed method modifies the existing (non-causal) encoder to a causal encoder by fine-tuning both the encoder and decoder using Low-Rank Adaptation (LoRA) and a weakly aligned dataset. We then propose an updated inference mechanism that utilizes the fine-tune causal encoder and decoder to yield greedy and beam-search decoding, and is shown to be locally optimal. Experiments on low-latency chunk sizes (less than 300 msec) show that our fine-tuned model outperforms existing non-fine-tuned streaming approaches in most cases, while using a lower complexity. Additionally, we observe that our training process yields better alignment, enabling a simple method for extracting word-level timestamps. We release our training and inference code, along with the fine-tuned models, to support further research and development in streaming ASR.


【8】HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
标题:HuBERT-VIC:基于方差-不变-协方差正则化的语音基础模型抗噪自动语音识别
链接:https://arxiv.org/abs/2508.12292

备注:Accepted at Interspeech 2025
摘要:语音基础模型(SFM)中的噪声鲁棒性一直是一个关键的挑战,因为大多数模型主要是在干净的数据上训练的,当模型暴露于嘈杂的语音时,性能会下降。为了解决这个问题,我们提出了HuBERT-VIC,一个具有方差,不变性和协方差正则化(VICReg)目标的噪声鲁棒SFM。这些目标调整噪声语音表示的统计,使模型能够捕获不同的声学特征,并提高跨不同类型噪声的泛化能力。当应用于HuBERT时,我们的模型在LibriSpeech test-clean上显示出23.3%的相对性能改进,在test-other上显示出13.2%的相对性能改进,与在嘈杂语音上预训练的基线模型相比。
摘要:Noise robustness in speech foundation models (SFMs) has been a critical challenge, as most models are primarily trained on clean data and experience performance degradation when the models are exposed to noisy speech. To address this issue, we propose HuBERT-VIC, a noise-robust SFM with variance, in-variance, and covariance regularization (VICReg) objectives. These objectives adjust the statistics of noisy speech representations, enabling the model to capture diverse acoustic characteristics and improving the generalization ability across different types of noise. When applied to HuBERT, our model shows relative performance improvements of 23.3% on LibriSpeech test-clean and 13.2% on test-other, compared to the baseline model pre-trained on noisy speech.


【9】Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection
标题:探索用于广义异常声音检测的自我监督音频模型
链接:https://arxiv.org/abs/2508.12230

备注:Accepted by TASLP. 15 pages, 7 figures;
摘要:机器异常声音检测(ASD)是跨各种应用的有价值的技术。然而,由于数据采集的挑战和声学环境的复杂性,其泛化性能往往受到限制。受众多领域中大型预训练模型的成功启发,本文介绍了一种强大的ASD模型,该模型利用在大规模语音和音频数据集上训练的自监督预训练模型。尽管预训练数据集和ASD任务之间存在不一致,但我们的研究结果表明,预训练仍然为ASD提供了实质性的好处。为了在有限数据的微调时减轻过拟合并保留学到的知识,我们探索了全连接低秩自适应(LoRA)作为完全微调的替代方案。此外,我们提出了一个机器感知的组适配器模块,它使模型能够在一个统一的框架内捕捉各种机器之间的差异,从而提高ASD系统的泛化性能。为了解决缺失属性标签的挑战,我们设计了一种新的目标函数,该目标函数使用矢量量化动态聚类未属性数据,并通过双层对比学习损失进行优化。在所有基准数据集上对所提出的方法进行了评估,包括DCASE 2020-2024五个ASD挑战,实验结果表明我们的新方法有了显着的改进,并证明了我们提出的策略的有效性。
摘要:Machine anomalous sound detection (ASD) is a valuable technique across various applications. However, its generalization performance is often limited due to challenges in data collection and the complexity of acoustic environments. Inspired by the success of large pre-trained models in numerous fields, this paper introduces a robust ASD model that leverages self-supervised pre-trained models trained on large-scale speech and audio datasets. Although there are inconsistencies between the pre-training datasets and the ASD task, our findings indicate that pre-training still provides substantial benefits for ASD. To mitigate overfitting and retain learned knowledge when fine-tuning with limited data, we explore Fully-Connected Low-Rank Adaptation (LoRA) as an alternative to full fine-tuning. Additionally, we propose a Machine-aware Group Adapter module, which enables the model to capture differences between various machines within a unified framework, thereby enhancing the generalization performance of ASD systems. To address the challenge of missing attribute labels, we design a novel objective function that dynamically clusters unattributed data using vector quantization and optimizes through a dual-level contrastive learning loss. The proposed methods are evaluated on all benchmark datasets, including the DCASE 2020-2024 five ASD challenges, and the experimental results show significant improvements of our new approach and demonstrate the effectiveness of our proposed strategies.


【10】Optimizing Neural Architectures for Hindi Speech Separation and Enhancement in Noisy Environments
标题:优化神经架构以实现高噪环境中印地语语音分离和增强
链接:https://arxiv.org/abs/2508.12009

备注:ICAD 2025
摘要:本文讨论了使用高级神经网络架构进行印地语语音分离和增强的挑战,重点是边缘设备。我们提出了一种改进的方法,利用DEMUCS模型,以克服传统方法的局限性,实现语音清晰度和可懂度的大幅改善。该模型使用U-Net和LSTM层进行微调,在40万个印地语语音片段的数据集上进行训练,并使用ESC-50和MS-SNSD进行增强,以适应不同的声学环境。使用PESQ和STOI指标进行的评估显示出卓越的性能,特别是在极端噪声条件下。为了确保在TWS耳机等资源受限设备上的部署,我们探索了量化技术来降低计算需求。这项研究强调了定制人工智能算法在印度背景下语音处理的有效性,并提出了优化基于边缘的架构的未来方向。
摘要:This paper addresses the challenges of Hindi speech separation and enhancement using advanced neural network architectures, with a focus on edge devices. We propose a refined approach leveraging the DEMUCS model to overcome limitations of traditional methods, achieving substantial improvements in speech clarity and intelligibility. The model is fine-tuned with U-Net and LSTM layers, trained on a dataset of 400,000 Hindi speech clips augmented with ESC-50 and MS-SNSD for diverse acoustic environments. Evaluation using PESQ and STOI metrics shows superior performance, particularly under extreme noise conditions. To ensure deployment on resource-constrained devices like TWS earbuds, we explore quantization techniques to reduce computational requirements. This research highlights the effectiveness of customized AI algorithms for speech processing in Indian contexts and suggests future directions for optimizing edge-based architectures.


【11】Towards Automatic Evaluation and High-Quality Pseudo-Parallel Dataset Construction for Audio Editing: A Human-in-the-Loop Method
标题:迈向音频编辑的自动评估和高质量伪并行数据集构建:人在循环方法
链接:https://arxiv.org/abs/2508.11966

摘要:音频编辑旨在根据文本描述操纵音频内容,支持添加、删除或替换音频事件等任务。尽管最近取得了进展,但缺乏高质量的基准数据集和全面的评估指标仍然是评估音频编辑质量和改进任务本身的主要挑战。在这项工作中,我们提出了一种新的音频编辑任务的方法,将专家知识融入评估和数据集构建过程:1)首先,我们建立了AuditScore,这是第一个用于音频编辑主观评估的综合数据集,由7个代表性的音频编辑框架和23个系统配置生成的6,300多个编辑样本组成。每个样本都由专业评分员就音频编辑质量的三个关键方面进行注释:总体质量、与编辑意图的相关性和对原始功能的忠实性。2)基于这个数据集,我们训练AuditEval,这是第一个为音频编辑任务量身定制的自动MOS风格评分设计的模型。AuditEval解决了客观评估指标的严重缺乏和主观评估在这一领域的高昂成本。3)我们进一步利用AuditEval来评估和过滤大量的合成混合编辑对,通过选择最合理的样本来构建高质量的伪并行数据集。客观实验验证了我们的专家知情过滤策略在产生更高质量数据方面的有效性,同时也揭示了仅依赖客观指标的局限性。数据集、代码和工具可以在https://github.com/NKU-HLT/AuditEval上找到。
摘要:Audio editing aims to manipulate audio content based on textual descriptions, supporting tasks such as adding, removing, or replacing audio events. Despite recent progress, the lack of high-quality benchmark datasets and comprehensive evaluation metrics remains a major challenge for both assessing audio editing quality and improving the task itself. In this work, we propose a novel approach for audio editing task by incorporating expert knowledge into both the evaluation and dataset construction processes: 1) First, we establish AuditScore, the first comprehensive dataset for subjective evaluation of audio editing, consisting of over 6,300 edited samples generated from 7 representative audio editing frameworks and 23 system configurations. Each sample is annotated by professional raters on three key aspects of audio editing quality: overall Quality, Relevance to editing intent, and Faithfulness to original features. 2) Based on this dataset, we train AuditEval, the first model designed for automatic MOS-style scoring tailored to audio editing tasks. AuditEval addresses the critical lack of objective evaluation metrics and the prohibitive cost of subjective assessment in this field. 3) We further leverage AuditEval to evaluate and filter a large amount of synthetically mixed editing pairs, constructing a high-quality pseudo-parallel dataset by selecting the most plausible samples. Objective experiments validate the effectiveness of our expert-informed filtering strategy in yielding higher-quality data, while also revealing the limitations of relying solely on objective metrics. The dataset, codes and tools can be found at: https://github.com/NKU-HLT/AuditEval.


【12】What Matters for Bioacoustic Encoding
标题:生物声学编码的重要性
链接:https://arxiv.org/abs/2508.11845

摘要:生物声学是对生物体产生的声音的研究,在保护、生物多样性监测和行为研究中起着至关重要的作用。该领域的许多任务,如物种、个体和行为分类和检测,都非常适合机器学习。然而,他们往往受到有限的注释数据,突出需要一个通用的生物声学编码器能够提取有用的表示为不同的下游任务。这样的编码器之前已经被提出,但是由于关注于窄范围的物种(通常是鸟类)并且依赖于单个模型架构或训练范例,因此通常在范围上受到限制。此外,它们通常在一小部分任务和数据集上进行评估。在这项工作中,我们提出了一个大规模的实证研究,涵盖了生物声学方面的研究,但以前很少考虑:训练数据的多样性和规模,模型架构和训练配方,以及评估任务和数据集的广度。我们获得的编码器是现有的和建议的基准上的最先进的。我们还确定了训练这些编码器的重要性,以便当有更多数据可用或提出更好的架构时,可以扩展这项工作。具体来说,在26个任务包括物种分类,检测,个体ID和声乐曲目发现的数据集中,我们发现在混合生物声学+通用音频语料库上进行自我监督预训练,然后进行监督后训练,可以产生最强的分布内和分布外性能。我们在这两个阶段的数据多样性的重要性。为了支持正在进行的研究和应用,我们将发布模型检查点。
摘要:Bioacoustics, the study of sounds produced by living organisms, plays a vital role in conservation, biodiversity monitoring, and behavioral studies. Many tasks in this field, such as species, individual, and behavior classification and detection, are well-suited to machine learning. However, they often suffer from limited annotated data, highlighting the need for a general-purpose bioacoustic encoder capable of extracting useful representations for diverse downstream tasks. Such encoders have been proposed before, but are often limited in scope due to a focus on a narrow range of species (typically birds), and a reliance on a single model architecture or training paradigm. Moreover, they are usually evaluated on a small set of tasks and datasets. In this work, we present a large-scale empirical study that covers aspects of bioacoustics that are relevant to research but have previously been scarcely considered: training data diversity and scale, model architectures and training recipes, and the breadth of evaluation tasks and datasets. We obtain encoders that are state-of-the-art on the existing and proposed benchmarks. We also identify what matters for training these encoders, such that this work can be extended when more data are available or better architectures are proposed. Specifically, across 26 datasets with tasks including species classification, detection, individual ID, and vocal repertoire discovery, we find self-supervised pre-training followed by supervised post-training on a mixed bioacoustics + general-audio corpus yields the strongest in- and out-of-distribution performance. We show the importance of data diversity in both stages. To support ongoing research and application, we will release the model checkpoints.


【13】Audio Flamingo Sound-CoT Technical Report: Improving Chain-of-Thought Reasoning in Sound Understanding
标题:音频火烈鸟Sound-CoT技术报告:改进声音理解中的思维链推理
链接:https://arxiv.org/abs/2508.11818

摘要:思维链推理在大型语言模型和视觉语言模型中表现出了显着的改进,但其在音频语言模型中的潜力仍然很大程度上未被开发。在这份技术报告中,我们采取了缩小这一差距的初步措施。为了更好地评估合理的推理,我们提出了AF-推理-评估,一个针对常识推理和区分密切相关的选择的能力的基准。为了准备训练语料库的声音推理能力,我们提出了自动管道转换现有的音频问答和分类数据到显式的推理链,产生AF-CoT-训练与1.24 M样本。我们研究了微调音频火烈鸟系列对AF-CoT-训练的影响,并观察了几个推理基准的相当大的改进,验证了思想链微调对高级声音理解的有效性。
摘要:Chain-of-thought reasoning has demonstrated significant improvements in large language models and vision language models, yet its potential for audio language models remains largely unexplored. In this technical report, we take a preliminary step towards closing this gap. For better assessment of sound reasoning, we propose AF-Reasoning-Eval, a benchmark targeting common-sense reasoning and the ability to discriminate among closely related choices. To prepare training corpus for sound reasoning abilities, we propose automatic pipelines that transform existing audio question answering and classification data into explicit reasoning chains, yielding AF-CoT-Train with 1.24M samples. We study the effect of finetuning Audio Flamingo series on AF-CoT-Train and observe considerable improvements on several reasoning benchmarks, validating the effectiveness of chain-of-thought finetuning on advanced sound understanding.


【14】Music and Artificial Intelligence: Artistic Trends
标题:音乐与人工智能:艺术趋势
链接:https://arxiv.org/abs/2508.11694

摘要:我们研究音乐家如何在单曲、专辑、表演、装置、声音、芭蕾舞、歌剧或配乐等格式中使用人工智能(AI)。我们收集了337个音乐艺术作品,并根据AI使用情况进行分类:AI作曲,联合作曲,声音设计,歌词生成和翻译。我们发现,人工智能被用作一种共同创造的工具,作为一种艺术媒介,并在现场表演和装置。人工智能的创新用途包括探索神秘的美学,多语言和多流派的歌曲发行,以及在线装置等新格式。这项研究提供了当前人工智能音乐实践的全面概述,为新兴的艺术趋势和人工智能音乐家面临的挑战提供了见解。
摘要:We study how musicians use artificial intelligence (AI) across formats like singles, albums, performances, installations, voices, ballets, operas, or soundtracks. We collect 337 music artworks and categorize them based on AI usage: AI composition, co-composition, sound design, lyrics generation, and translation. We find that AI is employed as a co-creative tool, as an artistic medium, and in live performances and installations. Innovative uses of AI include exploring uncanny aesthetics, multilingual and multigenre song releases, and new formats such as online installations. This research provides a comprehensive overview of current AI music practices, offering insights into emerging artistic trends and the challenges faced by AI musicians.


【15】Prediction of Spotify Chart Success Using Audio and Streaming Features
标题:使用音频和流媒体功能预测Spotify排行榜的成功
链接:https://arxiv.org/abs/2508.11632

摘要:Spotify的流媒体图表提供了一个实时的镜头,可以看到音乐的流行程度,驾驶发现,播放列表,甚至收入潜力。了解是什么影响了一首歌在排行榜上的上升,尤其是在早期,可以指导营销工作,投资决策,甚至艺术方向。在这个项目中,我们开发了一个分类管道,根据歌曲的音乐特征和早期参与数据来预测歌曲的排行榜成功。使用所有2024个美国前200名Spotify每日排行榜和Spotify Web API,我们构建了一个包含14,639首独特歌曲的元数据和音频特征的数据集。   该项目分为两个阶段。首先,我们对四个模型进行了基准测试:Logistic回归,K最近邻,随机森林和XGBoost-使用标准的训练测试分割。在第二阶段,我们结合了交叉验证、超参数调优和详细的类级评估,以确保鲁棒性。基于树的模型始终优于其他模型,随机森林和XGBoost的宏观F1得分接近0.95,准确率约为97%。   即使排除了流计数和排名历史,仅在音频属性上训练的模型也保留了预测能力。这些发现验证了基于音频的建模在A&R侦察,播放列表优化和命中预测中的潜力-远在曲目达到临界质量之前。
摘要:Spotify's streaming charts offer a real-time lens into music popularity, driving discovery, playlists, and even revenue potential. Understanding what influences a song's rise in ranks on these charts-especially early on-can guide marketing efforts, investment decisions, and even artistic direction. In this project, we developed a classification pipeline to predict a song's chart success based on its musical characteristics and early engagement data. Using all 2024 U.S. Top 200 Spotify Daily Charts and the Spotify Web API, we built a dataset containing both metadata and audio features for 14,639 unique songs.   The project was structured in two phases. First, we benchmarked four models: Logistic Regression, K Nearest Neighbors, Random Forest, and XGBoost-using a standard train-test split. In the second phase, we incorporated cross-validation, hyperparameter tuning, and detailed class-level evaluation to ensure robustness. Tree-based models consistently outperformed the rest, with Random Forest and XGBoost achieving macro F1-scores near 0.95 and accuracy around 97%.   Even when stream count and rank history were excluded, models trained solely on audio attributes retained predictive power. These findings validate the potential of audio-based modeling in A&R scouting, playlist optimization, and hit forecasting-long before a track reaches critical mass.


eess.AS音频处理


【1】Arabic ASR on the SADA Large-Scale Arabic Speech Corpus with Transformer-Based Models
标题:具有基于转换器的模型的SADA大规模阿拉伯语语音库上的阿拉伯语ASB
链接:https://arxiv.org/abs/2508.12968

摘要:我们探索了几种最先进的自动语音识别(ASR)模型在大规模阿拉伯语语音数据集SADA(阿拉伯语沙特音频数据集)上的性能,该数据集包含来自沙特电视节目的668小时高质量音频。该数据集包括多种方言和环境,特别是一个嘈杂的子集,这使得它对ASR特别具有挑战性。我们在SADA测试集上评估了模型的性能,并探讨了微调、语言模型以及噪声和去噪对其性能的影响。我们发现,性能最好的模型是MMS 1B模型在SADA上进行了微调,使用4-gram语言模型,在SADA测试干净集上实现了40.9%的WER和17.6%的CER。
摘要:We explore the performance of several state-of-the-art automatic speech recognition (ASR) models on a large-scale Arabic speech dataset, the SADA (Saudi Audio Dataset for Arabic), which contains 668 hours of high-quality audio from Saudi television shows. The dataset includes multiple dialects and environments, specifically a noisy subset that makes it particularly challenging for ASR. We evaluate the performance of the models on the SADA test set, and we explore the impact of fine-tuning, language models, as well as noise and denoising on their performance. We find that the best performing model is the MMS 1B model finetuned on SADA with a 4-gram language model that achieves a WER of 40.9\% and a CER of 17.6\% on the SADA test clean set.


【2】Cryfish: On deep audio analysis with Large Language Models
标题:Cryfish:使用大型语言模型进行深度音频分析
链接:https://arxiv.org/abs/2508.12666

备注:None
摘要:最近在基于文本的大型语言模型(LLM)的革命性进展,有助于增长的兴趣,扩展这种模型的能力,多模态感知和理解任务。听力是一种非常需要融入LLM的基本能力。然而,有效地将听力能力整合到LLM中是一个重大的挑战,在于将复杂的听觉任务概括为语音和声音。为了解决这些问题,我们引入了Cryfish,这是我们版本的具有生物学能力的LLM。该模型使用基于变压器的连接器将WavLM音频编码器功能集成到Qwen 2模型中。小龙虾通过专门的训练策略适应各种听觉任务。我们在新的动态SUPERB Phase-2综合多任务基准上评估了该模型,该基准专门为具有自动驾驶能力的模型设计。本文对Cryfish与公开可用的模型进行了深入的分析和详细的比较。
摘要:The recent revolutionary progress in text-based large language models (LLMs) has contributed to the growth of interest in extending capabilities of such models to multimodal perception and understanding tasks. Hearing is an essential capability that is highly desired to be integrated into LLMs. However, effective integrating listening capabilities into LLMs is a significant challenge lying in generalizing complex auditory tasks across speech and sounds. To address these issues, we introduce Cryfish, our version of auditory-capable LLM. The model integrates WavLM audio-encoder features into Qwen2 model using a transformer-based connector. Cryfish is adapted to various auditory tasks through a specialized training strategy. We evaluate the model on the new Dynamic SUPERB Phase-2 comprehensive multitask benchmark specifically designed for auditory-capable models. The paper presents an in-depth analysis and detailed comparison of Cryfish with the publicly available models.


【3】On the Extension of Differential Beamforming Theory to Arbitrary Planar Arrays of First-Order Elements
标题:差分波束形成理论在任意平面一阶阵元上的推广
链接:https://arxiv.org/abs/2508.12403

摘要:小型声学阵列利用空间多样性来实现超越单元件设备的能力,其应用范围从电话会议到沉浸式多媒体。宽带阵列处理的一个关键要求是频率不变的空间响应,这确保了在宽带宽范围内一致的方向性,并防止光谱着色。差分波束形成通过利用小尺寸阵列的紧密间隔的元件之间的压力差提供固有的频率不变的解决方案。然而,传统的方法,假设阵列元件是全向的,而真正的换能器表现出频率相关的方向性,如果没有正确建模,可能会降低性能。为了解决这个问题,我们提出了一个广义的模式匹配框架的频率不变的差分波束形成,适用于无约束的平面阵列的一阶方向性元素。通过将所需的波束图表示为截断的圆谐波展开并将其拟合到实际的元件响应,我们的方法可适应任意的平面几何形状和元件取向。这种方法可以合成任何顺序和转向方向的波束图案,而无需强加严格的布局要求。仿真证实,占传感器的方向性在设计阶段产生准确和强大的性能在不同的频率,几何形状和噪声条件。
摘要:Small-size acoustic arrays exploit spatial diversity to achieve capabilities beyond those of single-element devices, with applications ranging from teleconferencing to immersive multimedia. A key requirement for broadband array processing is a frequency-invariant spatial response, which ensures consistent directivity across wide bandwidths and prevents spectral coloration. Differential beamforming offers an inherently frequency-invariant solution by leveraging pressure differences between closely spaced elements of small-size arrays. Traditional approaches, however, assume the array elements to be omnidirectional, whereas real transducers exhibit frequency-dependent directivity that can degrade performance if not properly modeled. To address this limitation, we propose a generalized modal matching framework for frequency-invariant differential beamforming, applicable to unconstrained planar arrays of first-order directional elements. By representing the desired beampattern as a truncated circular harmonic expansion and fitting it to the actual element responses, our method accommodates arbitrary planar geometries and element orientations. This approach enables the synthesis of beampatterns of any order and steering direction without imposing rigid layout requirements. Simulations confirm that accounting for sensor directivity at the design stage yields accurate and robust performance across varying frequencies, geometries, and noise conditions.


【4】MASSLOC: A Massive Sound Source Localization System based on Direction-of-Arrival Estimation
标题:MASSEARCH:一种基于到达方向估计的大规模声音源定位系统
链接:https://arxiv.org/abs/2508.12024

备注:IEEE Transactions on Instrumentation and Measurement
摘要:声学室内定位提供了高度准确的位置估计的潜力,同时与基于射频(RF)的解决方案相比通常表现出较低的硬件要求。此外,基于角度的定位通过最小化所需的固定锚节点的数量来显著减少安装工作。在这方面的贡献,我们提出了所谓的MASSESSLESS系统,它利用稀疏的二维阵列几何定位和识别大量的并发活动源。此外,互补Zadoff-Chu序列的使用被引入以实现高效的基于波束成形的源识别。这些序列通过呈现频谱平衡的波形,在有利的相关特性和准确的、非同步的到达方向估计之间提供折衷。该系统在受控消声室和混响时间为1.6 s的高混响大厅环境中进行了评估。在实验室环境中,成功的到达方向估计和识别多达14个同时发射源的演示。该系统采用透视n点(PSPs)校准方法,在具有挑战性的混响环境中,动态源移动高达1.9 mps,实现了55.7 mm的中值三维定位误差和0.84度的中值角度误差。在该环境中,还使用总共三个标记演示和评估了多源功能。这些结果表明,即使在具有挑战性的声学条件下,MASSTYLE系统的可扩展性和鲁棒性。
摘要:Acoustic indoor localization offers the potential for highly accurate position estimation while generally exhibiting low hardware requirements compared to Radio Frequency (RF)-based solutions. Furthermore, angular-based localization significantly reduces installation effort by minimizing the number of required fixed anchor nodes. In this contribution, we propose the so-called MASSLOC system, which leverages sparse two-dimensional array geometries to localize and identify a large number of concurrently active sources. Additionally, the use of complementary Zadoff-Chu sequences is introduced to enable efficient, beamforming-based source identification. These sequences provide a trade-off between favorable correlation properties and accurate, unsynchronized direction-of-arrival estimation by exhibiting a spectrally balanced waveform. The system is evaluated in both a controlled anechoic chamber and a highly reverberant lobby environment with a reverberation time of 1.6 s. In a laboratory setting, successful direction-of-arrival estimation and identification of up to 14 simultaneously emitting sources are demonstrated. Adopting a Perspective-n-Point (PnP) calibration approach, the system achieves a median three-dimensional localization error of 55.7 mm and a median angular error of 0.84 deg with dynamic source movement of up to 1.9 mps in the challenging reverberant environment. The multi-source capability is also demonstrated and evaluated in that environment with a total of three tags. These results indicate the scalability and robustness of the MASSLOC system, even under challenging acoustic conditions.


【5】FNH-TTS: A Fast, Natural, and Human-Like Speech Synthesis System with advanced prosodic modeling based on Mixture of Experts
标题:FNH-TTC:一个快速、自然、类人语音合成系统,具有基于混合专家的高级韵律建模
链接:https://arxiv.org/abs/2508.12001

摘要:实现自然的和人类一样的语音合成与低推理成本仍然是语音合成研究的一个主要挑战。本研究的重点是人类韵律模式和合成频谱和谐,解决韵律建模和非自回归模型中的伪影问题的挑战。为了提高韵律建模和合成质量,我们引入了一个新的持续时间预测器的基础上的混合专家旁边的一个新的声码器与两个先进的多尺度鉴别器。我们将这些新的模块集成到VITS系统中,形成了我们的FNH-TTS系统。我们在LJSpeech、VCTK和LibriTTS上的实验证明了该系统在合成质量、音素持续时间预测、声码器结果和合成速度方面的优越性。我们的韵律可视化结果表明,FNH-TTS产生的持续时间预测,更接近自然人比其他系统。
摘要:Achieving natural and human-like speech synthesis with low inference costs remains a major challenge in speech synthesis research. This study focuses on human prosodic patterns and synthesized spectrum harmony, addressing the challenges of prosody modeling and artifact issues in non-autoregressive models. To enhance prosody modeling and synthesis quality, we introduce a new Duration Predictor based on the Mixture of Experts alongside a new Vocoder with two advanced multi-scale discriminators. We integrated the these new modules into the VITS system, forming our FNH-TTS system. Our experiments on LJSpeech, VCTK, and LibriTTS demonstrate the system's superiority in synthesis quality, phoneme duration prediction, Vocoder results, and synthesis speed. Our prosody visualization results show that FNH-TTS produces duration predictions that more closely align with natural human beings than other systems.


【6】CarelessWhisper: Turning Whisper into a Causal Streaming Model
标题:CarelessWhisper:将Whisper转变为因果流媒体模型
链接:https://arxiv.org/abs/2508.12301

备注:17 pages, 7 Figures, This work has been submitted to the IEEE for possible publication
摘要:自动语音识别(ASR)已经取得了显著的进步,OpenAI Whisper和NVIDIA Canary等模型在离线转录方面实现了最先进的(SOTA)性能。然而,由于其架构和训练方法的限制,这些模型并不是为流式(在线或实时)转录而设计的。我们提出了一种方法,把Transformer编码器-解码器模型变成一个低延迟的流模型,是不关心未来的上下文。我们提出了一个分析,解释为什么它不是简单的转换编码器-解码器Transformer到低延迟流模型。我们提出的方法通过使用低秩自适应(LoRA)和弱对齐数据集对编码器和解码器进行微调,将现有的(非因果)编码器修改为因果编码器。然后,我们提出了一个更新的推理机制,利用微调因果编码器和解码器产生贪婪和波束搜索解码,并被证明是局部最优的。低延迟块大小(小于300毫秒)的实验表明,我们的微调模型优于现有的非微调流媒体方法在大多数情况下,同时使用较低的复杂性。此外,我们观察到我们的训练过程产生了更好的对齐,从而实现了提取单词级时间戳的简单方法。我们发布了我们的训练和推理代码,以及经过微调的模型,以支持流式ASR的进一步研究和开发。
摘要:Automatic Speech Recognition (ASR) has seen remarkable progress, with models like OpenAI Whisper and NVIDIA Canary achieving state-of-the-art (SOTA) performance in offline transcription. However, these models are not designed for streaming (online or real-time) transcription, due to limitations in their architecture and training methodology. We propose a method to turn the transformer encoder-decoder model into a low-latency streaming model that is careless about future context. We present an analysis explaining why it is not straightforward to convert an encoder-decoder transformer to a low-latency streaming model. Our proposed method modifies the existing (non-causal) encoder to a causal encoder by fine-tuning both the encoder and decoder using Low-Rank Adaptation (LoRA) and a weakly aligned dataset. We then propose an updated inference mechanism that utilizes the fine-tune causal encoder and decoder to yield greedy and beam-search decoding, and is shown to be locally optimal. Experiments on low-latency chunk sizes (less than 300 msec) show that our fine-tuned model outperforms existing non-fine-tuned streaming approaches in most cases, while using a lower complexity. Additionally, we observe that our training process yields better alignment, enabling a simple method for extracting word-level timestamps. We release our training and inference code, along with the fine-tuned models, to support further research and development in streaming ASR.


【7】HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
标题:HuBERT-VIC:基于方差-不变-协方差正则化的语音基础模型抗噪自动语音识别
链接:https://arxiv.org/abs/2508.12292

备注:Accepted at Interspeech 2025
摘要:语音基础模型(SFM)中的噪声鲁棒性一直是一个关键的挑战,因为大多数模型主要是在干净的数据上训练的,当模型暴露于嘈杂的语音时,性能会下降。为了解决这个问题,我们提出了HuBERT-VIC,一个具有方差,不变性和协方差正则化(VICReg)目标的噪声鲁棒SFM。这些目标调整噪声语音表示的统计,使模型能够捕获不同的声学特征,并提高跨不同类型噪声的泛化能力。当应用于HuBERT时,我们的模型在LibriSpeech test-clean上显示出23.3%的相对性能改进,在test-other上显示出13.2%的相对性能改进,与在嘈杂语音上预训练的基线模型相比。
摘要:Noise robustness in speech foundation models (SFMs) has been a critical challenge, as most models are primarily trained on clean data and experience performance degradation when the models are exposed to noisy speech. To address this issue, we propose HuBERT-VIC, a noise-robust SFM with variance, in-variance, and covariance regularization (VICReg) objectives. These objectives adjust the statistics of noisy speech representations, enabling the model to capture diverse acoustic characteristics and improving the generalization ability across different types of noise. When applied to HuBERT, our model shows relative performance improvements of 23.3% on LibriSpeech test-clean and 13.2% on test-other, compared to the baseline model pre-trained on noisy speech.


【8】What do Speech Foundation Models Learn? Analysis and Applications
标题:言语基金会模型学到什么?分析和应用
链接:https://arxiv.org/abs/2508.12255

备注:Ph.D. Thesis
摘要:语音基础模型(SFM)被设计成用作各种语音处理任务的通用表示。在过去的五年里,越来越多的成功的自我监督和监督预训练模型涌入,在各种下游任务上表现出色。   尽管可持续森林管理动物园不断发展,但我们对它们所获得的知识的理解却落后了。本文提出了一个轻量级的分析框架,使用统计工具和训练免费的任务,调查声学和语言知识编码在SFM层。我们在多个SFM和统计工具中进行了比较研究。我们的研究还表明,分析的见解有具体的影响下游任务的性能。   SFM的有效性最终取决于其在语音应用上的性能。然而,目前尚不清楚这些好处是否延伸到口语理解(SLU)任务,这些任务需要比广泛研究的任务更深入的理解,例如语音识别。对SLU的有限探索主要是由于缺乏相关数据集。为了缓解这一问题,本文贡献的任务,特别是口语命名实体识别(NER)和命名实体本地化(NEL),口语理解评估基准。我们开发了基于SFM的NER和NEL方法,并发现利用SFM的端到端(E2E)模型可以超越传统的级联(语音识别,然后是文本模型)方法。此外,我们评估E2E SLU模型跨SFM和适应策略,以评估对任务性能的影响。   总的来说,这篇论文解决了以前没有回答的问题,可持续森林管理系统,提供工具和数据集,以进一步我们的理解,并使社区作出明智的设计选择,为未来的模型开发和采用。
摘要:Speech foundation models (SFMs) are designed to serve as general-purpose representations for a wide range of speech-processing tasks. The last five years have seen an influx of increasingly successful self-supervised and supervised pre-trained models with impressive performance on various downstream tasks.   Although the zoo of SFMs continues to grow, our understanding of the knowledge they acquire lags behind. This thesis presents a lightweight analysis framework using statistical tools and training-free tasks to investigate the acoustic and linguistic knowledge encoded in SFM layers. We conduct a comparative study across multiple SFMs and statistical tools. Our study also shows that the analytical insights have concrete implications for downstream task performance.   The effectiveness of an SFM is ultimately determined by its performance on speech applications. Yet it remains unclear whether the benefits extend to spoken language understanding (SLU) tasks that require a deeper understanding than widely studied ones, such as speech recognition. The limited exploration of SLU is primarily due to a lack of relevant datasets. To alleviate that, this thesis contributes tasks, specifically spoken named entity recognition (NER) and named entity localization (NEL), to the Spoken Language Understanding Evaluation benchmark. We develop SFM-based approaches for NER and NEL, and find that end-to-end (E2E) models leveraging SFMs can surpass traditional cascaded (speech recognition followed by a text model) approaches. Further, we evaluate E2E SLU models across SFMs and adaptation strategies to assess the impact on task performance.   Collectively, this thesis tackles previously unanswered questions about SFMs, providing tools and datasets to further our understanding and to enable the community to make informed design choices for future model development and adoption.


【9】Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection
标题:探索用于广义异常声音检测的自我监督音频模型
链接:https://arxiv.org/abs/2508.12230

备注:Accepted by TASLP. 15 pages, 7 figures;
摘要:机器异常声音检测(ASD)是跨各种应用的有价值的技术。然而,由于数据采集的挑战和声学环境的复杂性,其泛化性能往往受到限制。受众多领域中大型预训练模型的成功启发,本文介绍了一种强大的ASD模型,该模型利用在大规模语音和音频数据集上训练的自监督预训练模型。尽管预训练数据集和ASD任务之间存在不一致,但我们的研究结果表明,预训练仍然为ASD提供了实质性的好处。为了在有限数据的微调时减轻过拟合并保留学到的知识,我们探索了全连接低秩自适应(LoRA)作为完全微调的替代方案。此外,我们提出了一个机器感知的组适配器模块,它使模型能够在一个统一的框架内捕捉各种机器之间的差异,从而提高ASD系统的泛化性能。为了解决缺失属性标签的挑战,我们设计了一种新的目标函数,该目标函数使用矢量量化动态聚类未属性数据,并通过双层对比学习损失进行优化。在所有基准数据集上对所提出的方法进行了评估,包括DCASE 2020-2024五个ASD挑战,实验结果表明我们的新方法有了显着的改进,并证明了我们提出的策略的有效性。
摘要:Machine anomalous sound detection (ASD) is a valuable technique across various applications. However, its generalization performance is often limited due to challenges in data collection and the complexity of acoustic environments. Inspired by the success of large pre-trained models in numerous fields, this paper introduces a robust ASD model that leverages self-supervised pre-trained models trained on large-scale speech and audio datasets. Although there are inconsistencies between the pre-training datasets and the ASD task, our findings indicate that pre-training still provides substantial benefits for ASD. To mitigate overfitting and retain learned knowledge when fine-tuning with limited data, we explore Fully-Connected Low-Rank Adaptation (LoRA) as an alternative to full fine-tuning. Additionally, we propose a Machine-aware Group Adapter module, which enables the model to capture differences between various machines within a unified framework, thereby enhancing the generalization performance of ASD systems. To address the challenge of missing attribute labels, we design a novel objective function that dynamically clusters unattributed data using vector quantization and optimizes through a dual-level contrastive learning loss. The proposed methods are evaluated on all benchmark datasets, including the DCASE 2020-2024 five ASD challenges, and the experimental results show significant improvements of our new approach and demonstrate the effectiveness of our proposed strategies.


【10】Music and Artificial Intelligence: Artistic Trends
标题:音乐与人工智能:艺术趋势
链接:https://arxiv.org/abs/2508.11694

摘要:我们研究音乐家如何在单曲、专辑、表演、装置、声音、芭蕾舞、歌剧或配乐等格式中使用人工智能(AI)。我们收集了337个音乐艺术作品,并根据AI使用情况进行分类:AI作曲,联合作曲,声音设计,歌词生成和翻译。我们发现,人工智能被用作一种共同创造的工具,作为一种艺术媒介,并在现场表演和装置。人工智能的创新用途包括探索神秘的美学,多语言和多流派的歌曲发行,以及在线装置等新格式。这项研究提供了当前人工智能音乐实践的全面概述,为新兴的艺术趋势和人工智能音乐家面临的挑战提供了见解。
摘要:We study how musicians use artificial intelligence (AI) across formats like singles, albums, performances, installations, voices, ballets, operas, or soundtracks. We collect 337 music artworks and categorize them based on AI usage: AI composition, co-composition, sound design, lyrics generation, and translation. We find that AI is employed as a co-creative tool, as an artistic medium, and in live performances and installations. Innovative uses of AI include exploring uncanny aesthetics, multilingual and multigenre song releases, and new formats such as online installations. This research provides a comprehensive overview of current AI music practices, offering insights into emerging artistic trends and the challenges faced by AI musicians.


【11】Prediction of Spotify Chart Success Using Audio and Streaming Features
标题:使用音频和流媒体功能预测Spotify排行榜的成功
链接:https://arxiv.org/abs/2508.11632

摘要:Spotify的流媒体图表提供了一个实时的镜头,可以看到音乐的流行程度,驾驶发现,播放列表,甚至收入潜力。了解是什么影响了一首歌在排行榜上的上升,尤其是在早期,可以指导营销工作,投资决策,甚至艺术方向。在这个项目中,我们开发了一个分类管道,根据歌曲的音乐特征和早期参与数据来预测歌曲的排行榜成功。使用所有2024个美国前200名Spotify每日排行榜和Spotify Web API,我们构建了一个包含14,639首独特歌曲的元数据和音频特征的数据集。   该项目分为两个阶段。首先,我们对四个模型进行了基准测试:Logistic回归,K最近邻,随机森林和XGBoost-使用标准的训练测试分割。在第二阶段,我们结合了交叉验证、超参数调优和详细的类级评估,以确保鲁棒性。基于树的模型始终优于其他模型,随机森林和XGBoost的宏观F1得分接近0.95,准确率约为97%。   即使排除了流计数和排名历史,仅在音频属性上训练的模型也保留了预测能力。这些发现验证了基于音频的建模在A&R侦察,播放列表优化和命中预测中的潜力-远在曲目达到临界质量之前。
摘要:Spotify's streaming charts offer a real-time lens into music popularity, driving discovery, playlists, and even revenue potential. Understanding what influences a song's rise in ranks on these charts-especially early on-can guide marketing efforts, investment decisions, and even artistic direction. In this project, we developed a classification pipeline to predict a song's chart success based on its musical characteristics and early engagement data. Using all 2024 U.S. Top 200 Spotify Daily Charts and the Spotify Web API, we built a dataset containing both metadata and audio features for 14,639 unique songs.   The project was structured in two phases. First, we benchmarked four models: Logistic Regression, K Nearest Neighbors, Random Forest, and XGBoost-using a standard train-test split. In the second phase, we incorporated cross-validation, hyperparameter tuning, and detailed class-level evaluation to ensure robustness. Tree-based models consistently outperformed the rest, with Random Forest and XGBoost achieving macro F1-scores near 0.95 and accuracy around 97%.   Even when stream count and rank history were excluded, models trained solely on audio attributes retained predictive power. These findings validate the potential of audio-based modeling in A&R scouting, playlist optimization, and hit forecasting-long before a track reaches critical mass.


机器翻译由腾讯交互翻译提供,仅供参考