微信公众号:arXiv_Daily
cs.SD语音
【1】FoleySpace: Vision-Aligned Binaural Spatial Audio Generation
标题:FoleySpace:视觉对齐的双耳空间音频生成
链接:https://arxiv.org/abs/2508.12918
摘要:最近,随着AIGC的发展,基于深度学习的视频到音频(V2A)技术引起了人们的极大关注。然而,现有的研究大多集中在单声道音频生成,缺乏空间感知,而双耳空间音频生成技术,可以提供更强的沉浸感的探索仍然不足。为了解决这个问题,我们提出了FoleySpace,一个框架的视频到双耳音频生成,产生沉浸式和空间一致的立体声视觉信息的指导下。具体来说,我们开发了一种声源估计方法来确定每个视频帧中的声源2D坐标和深度,然后采用坐标映射机制将2D源位置转换为3D轨迹。该3D轨迹与由预训练的V2A模型生成的单声道音频一起用作扩散模型的调节输入以生成空间一致的双耳音频。为了支持动态声场的生成,我们基于记录的头部相关脉冲响应构建了一个训练数据集,其中包括各种声源移动场景。实验结果表明,该方法在空间感知一致性方面优于现有方法,有效地提高了视听体验的沉浸质量。
摘要:Recently, with the advancement of AIGC, deep learning-based video-to-audio (V2A) technology has garnered significant attention. However, existing research mostly focuses on mono audio generation that lacks spatial perception, while the exploration of binaural spatial audio generation technologies, which can provide a stronger sense of immersion, remains insufficient. To solve this problem, we propose FoleySpace, a framework for video-to-binaural audio generation that produces immersive and spatially consistent stereo sound guided by visual information. Specifically, we develop a sound source estimation method to determine the sound source 2D coordinates and depth in each video frame, and then employ a coordinate mapping mechanism to convert the 2D source positions into a 3D trajectory. This 3D trajectory, together with the monaural audio generated by a pre-trained V2A model, serves as a conditioning input for a diffusion model to generate spatially consistent binaural audio. To support the generation of dynamic sound fields, we constructed a training dataset based on recorded Head-Related Impulse Responses that includes various sound source movement scenarios. Experimental results demonstrate that the proposed method outperforms existing approaches in spatial perception consistency, effectively enhancing the immersive quality of the audio-visual experience.
【2】MATPAC++: Enhanced Masked Latent Prediction for Self-Supervised Audio Representation Learning
标题:MARTPAC ++:用于自我监督音频表示学习的增强掩蔽潜伏预测
链接:https://arxiv.org/abs/2508.12709
备注:Under review
摘要:掩蔽潜在预测已经成为自监督学习(SSL)的主要范式,特别是对于一般的音频和音乐表示学习。虽然最近的方法已经证明了强大的性能,预测模块的作用在这样的SSL系统的输出仍然主要被忽视,尽管是至关重要的解决手头的借口任务。特别是,这个模块应该能够处理音频内容中固有的模糊性,特别是当它由多个声源组成时。这项工作提出了一种新的增强:集成多项选择学习(MCL),以显式地建模预测模糊性和提高表示质量。我们建立在最近提出的MATPAC系统之上,用MCL改进其预测和无监督分类任务。我们广泛评估我们的方法,MATPAC++,通过多个下游任务的线性探测和AudioSet上的微调,采用统一的协议,使严格和公平的比较与最先进的SSL方法。结果表明,我们的建议达到国家的最先进的微调AudioSet和整体国家的最先进的下游任务的分数。此外,我们通过专门对音乐数据进行训练来检查领域专业化,我们的模型实现了最先进的性能,效率显著提高。
摘要:Masked latent prediction has emerged as a leading paradigm in self-supervised learning (SSL), especially for general audio and music representation learning. While recent methods have demonstrated strong performance, the role of the predictor module used at the output of such SSL systems remains mainly overlooked, despite being crucial for solving the pretext task at hand. In particular, this module should be able to deal with the ambiguity inherent in audio content, especially when it is composed of multiple sound sources. This work proposes a novel enhancement: integrating Multiple Choice Learning (MCL) to explicitly model prediction ambiguity and improve representation quality. We build on top of the recently proposed MATPAC system, improving its prediction and unsupervised classification pretext tasks with MCL. We extensively evaluate our method, MATPAC++, through both linear probing across multiple downstream tasks and fine-tuning on AudioSet, employing a unified protocol that enables rigorous and fair comparisons with state-of-the-art SSL approaches. Results show that our proposal achieves state-of-the-art when fine-tuned on AudioSet and overall state-of-the-art scores on downstream tasks. Additionally, we examine domain specialisation by training exclusively on music data, where our model achieves state-of-the-art performance with significantly improved efficiency.
【3】Exploring the Feasibility of LLMs for Automated Music Emotion Annotation
标题:LLM用于自动音乐情感标注的可行性研究
链接:https://arxiv.org/abs/2508.12626
备注:Accepted to be published at ISMIR 2025
摘要:目前的音乐情感标注方法仍然严重依赖于人工标注,这一过程带来了巨大的资源和劳动力负担,严重限制了可用注释数据的规模。本研究探讨了使用大型语言模型(GPT-4 o)进行音乐情感标注的可行性和可靠性。在这项研究中,我们使用GPT-4 o在四象限效价唤醒框架中注释了GiantMIDI-Piano,这是一个经典的钢琴音乐数据集,并与三位人类专家提供的注释进行了比较。我们进行了广泛的评估,以评估GPT生成的音乐情感注释的性能和可靠性,包括标准准确性,加权准确性,占专家间的协议,注释者间的协议指标,和生成的标签的分布相似性。 虽然GPT的注释性能在整体准确性方面低于人类专家,并且在对特定情绪状态进行分类时表现出较少的细微差别,但评分员间可靠性指标表明GPT的可变性仍然在专家之间的自然分歧范围内。这些发现强调了基于GPT的注释的局限性和潜力:尽管其目前的缺点相对于人类的表现,其成本效益和效率使其成为一个有前途的可扩展的音乐情感注释的替代方案。
摘要:Current approaches to music emotion annotation remain heavily reliant on manual labelling, a process that imposes significant resource and labour burdens, severely limiting the scale of available annotated data. This study examines the feasibility and reliability of employing a large language model (GPT-4o) for music emotion annotation. In this study, we annotated GiantMIDI-Piano, a classical MIDI piano music dataset, in a four-quadrant valence-arousal framework using GPT-4o, and compared against annotations provided by three human experts. We conducted extensive evaluations to assess the performance and reliability of GPT-generated music emotion annotations, including standard accuracy, weighted accuracy that accounts for inter-expert agreement, inter-annotator agreement metrics, and distributional similarity of the generated labels. While GPT's annotation performance fell short of human experts in overall accuracy and exhibited less nuance in categorizing specific emotional states, inter-rater reliability metrics indicate that GPT's variability remains within the range of natural disagreement among experts. These findings underscore both the limitations and potential of GPT-based annotation: despite its current shortcomings relative to human performance, its cost-effectiveness and efficiency render it a promising scalable alternative for music emotion annotation.
【4】Beyond Modality Limitations: A Unified MLLM Approach to Automated Speaking Assessment with Effective Curriculum Learning
标题:超越模式限制:统一的MLLM方法实现自动化口语评估和有效的课程学习
链接:https://arxiv.org/abs/2508.12591
备注:Accepted at IEEE ASRU 2025
摘要:传统的自动口语评估(ASA)系统表现出固有的模态限制:基于文本的方法缺乏声学信息,而基于音频的方法错过语义上下文。多模态大型语言模型(MLLM)通过在统一框架内同时处理音频和文本,为全面的ASA提供了前所未有的机会。本文首次对综合ASA的MLLM进行了系统的研究,展示了MLLM在内容和语言使用方面的优越性能。然而,对交付方面的评估揭示了独特的挑战,认为这需要专门的培训战略。因此,我们提出了语音优先多模态训练(SFMT),利用课程学习原则,在跨模态协同融合之前建立更强大的语音建模基础。在基准数据集上的一系列实验表明,基于MLLM的系统可以将整体评估性能从PCC值0.783提高到0.846。特别是,SFMT在评估交付方面表现出色,与传统的训练方法相比,绝对准确率提高了4%,这也为ASA铺平了新的道路。
摘要:Traditional Automated Speaking Assessment (ASA) systems exhibit inherent modality limitations: text-based approaches lack acoustic information while audio-based methods miss semantic context. Multimodal Large Language Models (MLLM) offer unprecedented opportunities for comprehensive ASA by simultaneously processing audio and text within unified frameworks. This paper presents a very first systematic study of MLLM for comprehensive ASA, demonstrating the superior performance of MLLM across the aspects of content and language use . However, assessment on the delivery aspect reveals unique challenges, which is deemed to require specialized training strategies. We thus propose Speech-First Multimodal Training (SFMT), leveraging a curriculum learning principle to establish more robust modeling foundations of speech before cross-modal synergetic fusion. A series of experiments on a benchmark dataset show MLLM-based systems can elevate the holistic assessment performance from a PCC value of 0.783 to 0.846. In particular, SFMT excels in the evaluation of the delivery aspect, achieving an absolute accuracy improvement of 4% over conventional training approaches, which also paves a new avenue for ASA.
【5】CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation
标题:CEM-Net:用于情感说话面部生成的交叉情感记忆网络
链接:https://arxiv.org/abs/2508.12368
摘要:情感说话人脸生成的目的是在给定的参考图像动画的人脸,并生成一个说话的视频,匹配的内容和情感的驾驶音频。然而,现有的方法忽略了参考图像可能具有与音频情感冲突的强烈情感,导致严重的情感不准确和失真的生成结果。为了解决这个问题,我们引入了一个跨情感记忆网络(CEM-Net),旨在生成与驾驶音频对齐的情感说话的脸时,参考图像表现出强烈的情感。具体而言,音频情感增强模块(AEE)的设计与交叉重建训练策略,以增强音频情感,克服了干扰,从参考图像的情感。其次,由于参考图像无法提供音频情感下说话者面部运动的足够信息,因此利用情感桥接记忆模块(EBM)来补偿缺失的信息。该算法将参考图像情感到音频情感的表情位移引入到存储器中,并以跨情感特征作为查询,在推理时检索匹配位移。大量的实验表明,我们的CEM-Net可以合成具有表达力的,自然的和嘴唇同步的说话人脸视频,具有更好的情感准确性。
摘要:Emotional talking face generation aims to animate a human face in given reference images and generate a talking video that matches the content and emotion of driving audio. However, existing methods neglect that reference images may have a strong emotion that conflicts with the audio emotion, leading to severe emotion inaccuracy and distorted generated results. To tackle the issue, we introduce a cross-emotion memory network(CEM-Net), designed to generate emotional talking faces aligned with the driving audio when reference images exhibit strong emotion. Specifically, an Audio Emotion Enhancement module(AEE) is first devised with the cross-reconstruction training strategy to enhance audio emotion, overcoming the disruption from reference image emotion. Secondly, since reference images cannot provide sufficient facial motion information of the speaker under audio emotion, an Emotion Bridging Memory module(EBM) is utilized to compensate for the lacked information. It brings in expression displacement from the reference image emotion to the audio emotion and stores it in the memory.Given a cross-emotion feature as a query, the matching displacement can be retrieved at inference time. Extensive experiments have demonstrated that our CEM-Net can synthesize expressive, natural and lip-synced talking face videos with better emotion accuracy.
【6】Cross-Modal Knowledge Distillation with Multi-Level Data Augmentation for Low-Resource Audio-Visual Sound Event Localization and Detection
标题:具有多层数据增强的跨模式知识提炼用于低资源视听声音事件定位和检测
链接:https://arxiv.org/abs/2508.12334
摘要:本文提出了一种结合多层次数据增强的跨模态知识提取(CMKD)框架,用于低资源音视频(AV)声音事件定位和检测(SELD)。仅音频SELD模型充当教师,通过输出响应和中间特征表示将知识传递到AV学生模型。为了增强学习,通过混合从多个网络层随机选择的特征和为SELD任务定制的相关损失函数来应用数据增强。在DCASE 2023和2024 SELD数据集上的大量实验表明,该方法显著提高了AV SELD性能,在总体指标上相对于基线获得了22%~36%的相对增益。值得注意的是,我们的方法实现了与在更大数据集上训练的教师模型相当或更好的结果,在DCASE 2023和2024 SELD任务上都超过了最先进的方法。
摘要:This work presents a cross-modal knowledge distillation (CMKD) framework combined with multi-level data augmentation for low-resource audio-visual (AV) sound event localization and detection (SELD). An audio-only SELD model acts as the teacher, transferring knowledge to an AV student model through both output responses and intermediate feature representations. To enhance learning, data augmentation is applied by mixing features randomly selected from multiple network layers and associated loss functions tailored to the SELD task. Extensive experiments on the DCASE 2023 and 2024 SELD datasets show that the proposed method significantly improves AV SELD performance, yielding relative gains of 22%~36% in the overall metric over the baseline. Notably, our approach achieves results comparable to or better than teacher models trained on much larger datasets, surpassing state-of-the-art methods on both DCASE 2023 and 2024 SELD tasks.
【7】CarelessWhisper: Turning Whisper into a Causal Streaming Model
标题:CarelessWhisper:将Whisper转变为因果流媒体模型
链接:https://arxiv.org/abs/2508.12301
摘要:自动语音识别(ASR)已经取得了显著的进展,OpenAI Whisper和NVIDIA Canary等模型在离线转录方面实现了最先进的(SOTA)性能。然而,由于其架构和训练方法的限制,这些模型并不是为流式(在线或实时)转录而设计的。我们提出了一种方法,把Transformer编码器-解码器模型变成一个低延迟的流模型,是不关心未来的上下文。我们提出了一个分析,解释为什么它不是简单的转换编码器-解码器Transformer到低延迟流模型。我们提出的方法通过使用低秩自适应(LoRA)和弱对齐数据集对编码器和解码器进行微调,将现有的(非因果)编码器修改为因果编码器。然后,我们提出了一个更新的推理机制,利用微调因果编码器和解码器产生贪婪和波束搜索解码,并被证明是局部最优的。低延迟块大小(小于300毫秒)的实验表明,我们的微调模型优于现有的非微调流媒体方法在大多数情况下,同时使用较低的复杂性。此外,我们观察到我们的训练过程产生了更好的对齐,从而实现了提取单词级时间戳的简单方法。我们发布了我们的训练和推理代码,以及经过微调的模型,以支持流式ASR的进一步研究和开发。
摘要:Automatic Speech Recognition (ASR) has seen remarkable progress, with models like OpenAI Whisper and NVIDIA Canary achieving state-of-the-art (SOTA) performance in offline transcription. However, these models are not designed for streaming (online or real-time) transcription, due to limitations in their architecture and training methodology. We propose a method to turn the transformer encoder-decoder model into a low-latency streaming model that is careless about future context. We present an analysis explaining why it is not straightforward to convert an encoder-decoder transformer to a low-latency streaming model. Our proposed method modifies the existing (non-causal) encoder to a causal encoder by fine-tuning both the encoder and decoder using Low-Rank Adaptation (LoRA) and a weakly aligned dataset. We then propose an updated inference mechanism that utilizes the fine-tune causal encoder and decoder to yield greedy and beam-search decoding, and is shown to be locally optimal. Experiments on low-latency chunk sizes (less than 300 msec) show that our fine-tuned model outperforms existing non-fine-tuned streaming approaches in most cases, while using a lower complexity. Additionally, we observe that our training process yields better alignment, enabling a simple method for extracting word-level timestamps. We release our training and inference code, along with the fine-tuned models, to support further research and development in streaming ASR.
【8】HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
标题:HuBERT-VIC:基于方差-不变-协方差正则化的语音基础模型抗噪自动语音识别
链接:https://arxiv.org/abs/2508.12292
摘要:语音基础模型(SFM)中的噪声鲁棒性一直是一个关键的挑战,因为大多数模型主要是在干净的数据上训练的,当模型暴露于嘈杂的语音时,性能会下降。为了解决这个问题,我们提出了HuBERT-VIC,一个具有方差,不变性和协方差正则化(VICReg)目标的噪声鲁棒SFM。这些目标调整噪声语音表示的统计,使模型能够捕获不同的声学特征,并提高跨不同类型噪声的泛化能力。当应用于HuBERT时,我们的模型在LibriSpeech test-clean上显示出23.3%的相对性能改进,在test-other上显示出13.2%的相对性能改进,与在嘈杂语音上预训练的基线模型相比。
摘要:Noise robustness in speech foundation models (SFMs) has been a critical challenge, as most models are primarily trained on clean data and experience performance degradation when the models are exposed to noisy speech. To address this issue, we propose HuBERT-VIC, a noise-robust SFM with variance, in-variance, and covariance regularization (VICReg) objectives. These objectives adjust the statistics of noisy speech representations, enabling the model to capture diverse acoustic characteristics and improving the generalization ability across different types of noise. When applied to HuBERT, our model shows relative performance improvements of 23.3% on LibriSpeech test-clean and 13.2% on test-other, compared to the baseline model pre-trained on noisy speech.
【9】Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection
标题:探索用于广义异常声音检测的自我监督音频模型
链接:https://arxiv.org/abs/2508.12230
摘要:机器异常声音检测(ASD)是跨各种应用的有价值的技术。然而,由于数据采集的挑战和声学环境的复杂性,其泛化性能往往受到限制。受众多领域中大型预训练模型的成功启发,本文介绍了一种强大的ASD模型,该模型利用在大规模语音和音频数据集上训练的自监督预训练模型。尽管预训练数据集和ASD任务之间存在不一致,但我们的研究结果表明,预训练仍然为ASD提供了实质性的好处。为了在有限数据的微调时减轻过拟合并保留学到的知识,我们探索了全连接低秩自适应(LoRA)作为完全微调的替代方案。此外,我们提出了一个机器感知的组适配器模块,它使模型能够在一个统一的框架内捕捉各种机器之间的差异,从而提高ASD系统的泛化性能。为了解决缺失属性标签的挑战,我们设计了一种新的目标函数,该目标函数使用矢量量化动态聚类未属性数据,并通过双层对比学习损失进行优化。在所有基准数据集上对所提出的方法进行了评估,包括DCASE 2020-2024五个ASD挑战,实验结果表明我们的新方法有了显着的改进,并证明了我们提出的策略的有效性。
摘要:Machine anomalous sound detection (ASD) is a valuable technique across various applications. However, its generalization performance is often limited due to challenges in data collection and the complexity of acoustic environments. Inspired by the success of large pre-trained models in numerous fields, this paper introduces a robust ASD model that leverages self-supervised pre-trained models trained on large-scale speech and audio datasets. Although there are inconsistencies between the pre-training datasets and the ASD task, our findings indicate that pre-training still provides substantial benefits for ASD. To mitigate overfitting and retain learned knowledge when fine-tuning with limited data, we explore Fully-Connected Low-Rank Adaptation (LoRA) as an alternative to full fine-tuning. Additionally, we propose a Machine-aware Group Adapter module, which enables the model to capture differences between various machines within a unified framework, thereby enhancing the generalization performance of ASD systems. To address the challenge of missing attribute labels, we design a novel objective function that dynamically clusters unattributed data using vector quantization and optimizes through a dual-level contrastive learning loss. The proposed methods are evaluated on all benchmark datasets, including the DCASE 2020-2024 five ASD challenges, and the experimental results show significant improvements of our new approach and demonstrate the effectiveness of our proposed strategies.
【10】Optimizing Neural Architectures for Hindi Speech Separation and Enhancement in Noisy Environments
标题:优化神经架构以实现高噪环境中印地语语音分离和增强
链接:https://arxiv.org/abs/2508.12009
摘要:本文讨论了使用高级神经网络架构进行印地语语音分离和增强的挑战,重点是边缘设备。我们提出了一种改进的方法,利用DEMUCS模型,以克服传统方法的局限性,实现语音清晰度和可懂度的大幅改善。该模型使用U-Net和LSTM层进行微调,在40万个印地语语音片段的数据集上进行训练,并使用ESC-50和MS-SNSD进行增强,以适应不同的声学环境。使用PESQ和STOI指标进行的评估显示出卓越的性能,特别是在极端噪声条件下。为了确保在TWS耳机等资源受限设备上的部署,我们探索了量化技术来降低计算需求。这项研究强调了定制人工智能算法在印度背景下语音处理的有效性,并提出了优化基于边缘的架构的未来方向。
摘要:This paper addresses the challenges of Hindi speech separation and enhancement using advanced neural network architectures, with a focus on edge devices. We propose a refined approach leveraging the DEMUCS model to overcome limitations of traditional methods, achieving substantial improvements in speech clarity and intelligibility. The model is fine-tuned with U-Net and LSTM layers, trained on a dataset of 400,000 Hindi speech clips augmented with ESC-50 and MS-SNSD for diverse acoustic environments. Evaluation using PESQ and STOI metrics shows superior performance, particularly under extreme noise conditions. To ensure deployment on resource-constrained devices like TWS earbuds, we explore quantization techniques to reduce computational requirements. This research highlights the effectiveness of customized AI algorithms for speech processing in Indian contexts and suggests future directions for optimizing edge-based architectures.
【11】Towards Automatic Evaluation and High-Quality Pseudo-Parallel Dataset Construction for Audio Editing: A Human-in-the-Loop Method
标题:迈向音频编辑的自动评估和高质量伪并行数据集构建:人在循环方法
链接:https://arxiv.org/abs/2508.11966
摘要:Audio editing aims to manipulate audio content based on textual descriptions, supporting tasks such as adding, removing, or replacing audio events. Despite recent progress, the lack of high-quality benchmark datasets and comprehensive evaluation metrics remains a major challenge for both assessing audio editing quality and improving the task itself. In this work, we propose a novel approach for audio editing task by incorporating expert knowledge into both the evaluation and dataset construction processes: 1) First, we establish AuditScore, the first comprehensive dataset for subjective evaluation of audio editing, consisting of over 6,300 edited samples generated from 7 representative audio editing frameworks and 23 system configurations. Each sample is annotated by professional raters on three key aspects of audio editing quality: overall Quality, Relevance to editing intent, and Faithfulness to original features. 2) Based on this dataset, we train AuditEval, the first model designed for automatic MOS-style scoring tailored to audio editing tasks. AuditEval addresses the critical lack of objective evaluation metrics and the prohibitive cost of subjective assessment in this field. 3) We further leverage AuditEval to evaluate and filter a large amount of synthetically mixed editing pairs, constructing a high-quality pseudo-parallel dataset by selecting the most plausible samples. Objective experiments validate the effectiveness of our expert-informed filtering strategy in yielding higher-quality data, while also revealing the limitations of relying solely on objective metrics. The dataset, codes and tools can be found at: https://github.com/NKU-HLT/AuditEval.
【12】What Matters for Bioacoustic Encoding
标题:生物声学编码的重要性
链接:https://arxiv.org/abs/2508.11845
摘要:Bioacoustics, the study of sounds produced by living organisms, plays a vital role in conservation, biodiversity monitoring, and behavioral studies. Many tasks in this field, such as species, individual, and behavior classification and detection, are well-suited to machine learning. However, they often suffer from limited annotated data, highlighting the need for a general-purpose bioacoustic encoder capable of extracting useful representations for diverse downstream tasks. Such encoders have been proposed before, but are often limited in scope due to a focus on a narrow range of species (typically birds), and a reliance on a single model architecture or training paradigm. Moreover, they are usually evaluated on a small set of tasks and datasets. In this work, we present a large-scale empirical study that covers aspects of bioacoustics that are relevant to research but have previously been scarcely considered: training data diversity and scale, model architectures and training recipes, and the breadth of evaluation tasks and datasets. We obtain encoders that are state-of-the-art on the existing and proposed benchmarks. We also identify what matters for training these encoders, such that this work can be extended when more data are available or better architectures are proposed. Specifically, across 26 datasets with tasks including species classification, detection, individual ID, and vocal repertoire discovery, we find self-supervised pre-training followed by supervised post-training on a mixed bioacoustics + general-audio corpus yields the strongest in- and out-of-distribution performance. We show the importance of data diversity in both stages. To support ongoing research and application, we will release the model checkpoints.
【13】Audio Flamingo Sound-CoT Technical Report: Improving Chain-of-Thought Reasoning in Sound Understanding
标题:音频火烈鸟Sound-CoT技术报告:改进声音理解中的思维链推理
链接:https://arxiv.org/abs/2508.11818
摘要:Chain-of-thought reasoning has demonstrated significant improvements in large language models and vision language models, yet its potential for audio language models remains largely unexplored. In this technical report, we take a preliminary step towards closing this gap. For better assessment of sound reasoning, we propose AF-Reasoning-Eval, a benchmark targeting common-sense reasoning and the ability to discriminate among closely related choices. To prepare training corpus for sound reasoning abilities, we propose automatic pipelines that transform existing audio question answering and classification data into explicit reasoning chains, yielding AF-CoT-Train with 1.24M samples. We study the effect of finetuning Audio Flamingo series on AF-CoT-Train and observe considerable improvements on several reasoning benchmarks, validating the effectiveness of chain-of-thought finetuning on advanced sound understanding.
【14】Music and Artificial Intelligence: Artistic Trends
标题:音乐与人工智能:艺术趋势
链接:https://arxiv.org/abs/2508.11694
摘要:We study how musicians use artificial intelligence (AI) across formats like singles, albums, performances, installations, voices, ballets, operas, or soundtracks. We collect 337 music artworks and categorize them based on AI usage: AI composition, co-composition, sound design, lyrics generation, and translation. We find that AI is employed as a co-creative tool, as an artistic medium, and in live performances and installations. Innovative uses of AI include exploring uncanny aesthetics, multilingual and multigenre song releases, and new formats such as online installations. This research provides a comprehensive overview of current AI music practices, offering insights into emerging artistic trends and the challenges faced by AI musicians.
【15】Prediction of Spotify Chart Success Using Audio and Streaming Features
标题:使用音频和流媒体功能预测Spotify排行榜的成功
链接:https://arxiv.org/abs/2508.11632
摘要:Spotify's streaming charts offer a real-time lens into music popularity, driving discovery, playlists, and even revenue potential. Understanding what influences a song's rise in ranks on these charts-especially early on-can guide marketing efforts, investment decisions, and even artistic direction. In this project, we developed a classification pipeline to predict a song's chart success based on its musical characteristics and early engagement data. Using all 2024 U.S. Top 200 Spotify Daily Charts and the Spotify Web API, we built a dataset containing both metadata and audio features for 14,639 unique songs. The project was structured in two phases. First, we benchmarked four models: Logistic Regression, K Nearest Neighbors, Random Forest, and XGBoost-using a standard train-test split. In the second phase, we incorporated cross-validation, hyperparameter tuning, and detailed class-level evaluation to ensure robustness. Tree-based models consistently outperformed the rest, with Random Forest and XGBoost achieving macro F1-scores near 0.95 and accuracy around 97%. Even when stream count and rank history were excluded, models trained solely on audio attributes retained predictive power. These findings validate the potential of audio-based modeling in A&R scouting, playlist optimization, and hit forecasting-long before a track reaches critical mass.
【1】Arabic ASR on the SADA Large-Scale Arabic Speech Corpus with Transformer-Based Models
标题:具有基于转换器的模型的SADA大规模阿拉伯语语音库上的阿拉伯语ASB
链接:https://arxiv.org/abs/2508.12968
摘要:We explore the performance of several state-of-the-art automatic speech recognition (ASR) models on a large-scale Arabic speech dataset, the SADA (Saudi Audio Dataset for Arabic), which contains 668 hours of high-quality audio from Saudi television shows. The dataset includes multiple dialects and environments, specifically a noisy subset that makes it particularly challenging for ASR. We evaluate the performance of the models on the SADA test set, and we explore the impact of fine-tuning, language models, as well as noise and denoising on their performance. We find that the best performing model is the MMS 1B model finetuned on SADA with a 4-gram language model that achieves a WER of 40.9\% and a CER of 17.6\% on the SADA test clean set.
【2】Cryfish: On deep audio analysis with Large Language Models
标题:Cryfish:使用大型语言模型进行深度音频分析
链接:https://arxiv.org/abs/2508.12666
摘要:最近在基于文本的大型语言模型(LLM)的革命性进展,有助于增长的兴趣,扩展这种模型的能力,多模态感知和理解任务。听力是一种非常需要融入LLM的基本能力。然而,有效地将听力能力整合到LLM中是一个重大的挑战,在于将复杂的听觉任务概括为语音和声音。为了解决这些问题,我们引入了Cryfish,这是我们版本的具有生物学能力的LLM。该模型使用基于变压器的连接器将WavLM音频编码器功能集成到Qwen 2模型中。小龙虾通过专门的训练策略适应各种听觉任务。我们在新的动态SUPERB Phase-2综合多任务基准上评估了该模型,该基准专门为具有自动驾驶能力的模型设计。本文对Cryfish与公开可用的模型进行了深入的分析和详细的比较。
摘要:The recent revolutionary progress in text-based large language models (LLMs) has contributed to the growth of interest in extending capabilities of such models to multimodal perception and understanding tasks. Hearing is an essential capability that is highly desired to be integrated into LLMs. However, effective integrating listening capabilities into LLMs is a significant challenge lying in generalizing complex auditory tasks across speech and sounds. To address these issues, we introduce Cryfish, our version of auditory-capable LLM. The model integrates WavLM audio-encoder features into Qwen2 model using a transformer-based connector. Cryfish is adapted to various auditory tasks through a specialized training strategy. We evaluate the model on the new Dynamic SUPERB Phase-2 comprehensive multitask benchmark specifically designed for auditory-capable models. The paper presents an in-depth analysis and detailed comparison of Cryfish with the publicly available models.
【3】On the Extension of Differential Beamforming Theory to Arbitrary Planar Arrays of First-Order Elements
标题:差分波束形成理论在任意平面一阶阵元上的推广
链接:https://arxiv.org/abs/2508.12403
摘要:Small-size acoustic arrays exploit spatial diversity to achieve capabilities beyond those of single-element devices, with applications ranging from teleconferencing to immersive multimedia. A key requirement for broadband array processing is a frequency-invariant spatial response, which ensures consistent directivity across wide bandwidths and prevents spectral coloration. Differential beamforming offers an inherently frequency-invariant solution by leveraging pressure differences between closely spaced elements of small-size arrays. Traditional approaches, however, assume the array elements to be omnidirectional, whereas real transducers exhibit frequency-dependent directivity that can degrade performance if not properly modeled. To address this limitation, we propose a generalized modal matching framework for frequency-invariant differential beamforming, applicable to unconstrained planar arrays of first-order directional elements. By representing the desired beampattern as a truncated circular harmonic expansion and fitting it to the actual element responses, our method accommodates arbitrary planar geometries and element orientations. This approach enables the synthesis of beampatterns of any order and steering direction without imposing rigid layout requirements. Simulations confirm that accounting for sensor directivity at the design stage yields accurate and robust performance across varying frequencies, geometries, and noise conditions.
【4】MASSLOC: A Massive Sound Source Localization System based on Direction-of-Arrival Estimation
标题:MASSEARCH:一种基于到达方向估计的大规模声音源定位系统
链接:https://arxiv.org/abs/2508.12024
摘要:声学室内定位提供了高度准确的位置估计的潜力,同时与基于射频(RF)的解决方案相比通常表现出较低的硬件要求。此外,基于角度的定位通过最小化所需的固定锚节点的数量来显著减少安装工作。在这方面的贡献,我们提出了所谓的MASSESSLESS系统,它利用稀疏的二维阵列几何定位和识别大量的并发活动源。此外,互补Zadoff-Chu序列的使用被引入以实现高效的基于波束成形的源识别。这些序列通过呈现频谱平衡的波形,在有利的相关特性和准确的、非同步的到达方向估计之间提供折衷。该系统在受控消声室和混响时间为1.6 s的高混响大厅环境中进行了评估。在实验室环境中,成功的到达方向估计和识别多达14个同时发射源的演示。该系统采用透视n点(PSPs)校准方法,在具有挑战性的混响环境中,动态源移动高达1.9 mps,实现了55.7 mm的中值三维定位误差和0.84度的中值角度误差。在该环境中,还使用总共三个标记演示和评估了多源功能。这些结果表明,即使在具有挑战性的声学条件下,MASSTYLE系统的可扩展性和鲁棒性。
摘要:Acoustic indoor localization offers the potential for highly accurate position estimation while generally exhibiting low hardware requirements compared to Radio Frequency (RF)-based solutions. Furthermore, angular-based localization significantly reduces installation effort by minimizing the number of required fixed anchor nodes. In this contribution, we propose the so-called MASSLOC system, which leverages sparse two-dimensional array geometries to localize and identify a large number of concurrently active sources. Additionally, the use of complementary Zadoff-Chu sequences is introduced to enable efficient, beamforming-based source identification. These sequences provide a trade-off between favorable correlation properties and accurate, unsynchronized direction-of-arrival estimation by exhibiting a spectrally balanced waveform. The system is evaluated in both a controlled anechoic chamber and a highly reverberant lobby environment with a reverberation time of 1.6 s. In a laboratory setting, successful direction-of-arrival estimation and identification of up to 14 simultaneously emitting sources are demonstrated. Adopting a Perspective-n-Point (PnP) calibration approach, the system achieves a median three-dimensional localization error of 55.7 mm and a median angular error of 0.84 deg with dynamic source movement of up to 1.9 mps in the challenging reverberant environment. The multi-source capability is also demonstrated and evaluated in that environment with a total of three tags. These results indicate the scalability and robustness of the MASSLOC system, even under challenging acoustic conditions.
【5】FNH-TTS: A Fast, Natural, and Human-Like Speech Synthesis System with advanced prosodic modeling based on Mixture of Experts
标题:FNH-TTC:一个快速、自然、类人语音合成系统,具有基于混合专家的高级韵律建模
链接:https://arxiv.org/abs/2508.12001
摘要:Achieving natural and human-like speech synthesis with low inference costs remains a major challenge in speech synthesis research. This study focuses on human prosodic patterns and synthesized spectrum harmony, addressing the challenges of prosody modeling and artifact issues in non-autoregressive models. To enhance prosody modeling and synthesis quality, we introduce a new Duration Predictor based on the Mixture of Experts alongside a new Vocoder with two advanced multi-scale discriminators. We integrated the these new modules into the VITS system, forming our FNH-TTS system. Our experiments on LJSpeech, VCTK, and LibriTTS demonstrate the system's superiority in synthesis quality, phoneme duration prediction, Vocoder results, and synthesis speed. Our prosody visualization results show that FNH-TTS produces duration predictions that more closely align with natural human beings than other systems.
【6】CarelessWhisper: Turning Whisper into a Causal Streaming Model
标题:CarelessWhisper:将Whisper转变为因果流媒体模型
链接:https://arxiv.org/abs/2508.12301
摘要:自动语音识别(ASR)已经取得了显著的进步,OpenAI Whisper和NVIDIA Canary等模型在离线转录方面实现了最先进的(SOTA)性能。然而,由于其架构和训练方法的限制,这些模型并不是为流式(在线或实时)转录而设计的。我们提出了一种方法,把Transformer编码器-解码器模型变成一个低延迟的流模型,是不关心未来的上下文。我们提出了一个分析,解释为什么它不是简单的转换编码器-解码器Transformer到低延迟流模型。我们提出的方法通过使用低秩自适应(LoRA)和弱对齐数据集对编码器和解码器进行微调,将现有的(非因果)编码器修改为因果编码器。然后,我们提出了一个更新的推理机制,利用微调因果编码器和解码器产生贪婪和波束搜索解码,并被证明是局部最优的。低延迟块大小(小于300毫秒)的实验表明,我们的微调模型优于现有的非微调流媒体方法在大多数情况下,同时使用较低的复杂性。此外,我们观察到我们的训练过程产生了更好的对齐,从而实现了提取单词级时间戳的简单方法。我们发布了我们的训练和推理代码,以及经过微调的模型,以支持流式ASR的进一步研究和开发。
摘要:Automatic Speech Recognition (ASR) has seen remarkable progress, with models like OpenAI Whisper and NVIDIA Canary achieving state-of-the-art (SOTA) performance in offline transcription. However, these models are not designed for streaming (online or real-time) transcription, due to limitations in their architecture and training methodology. We propose a method to turn the transformer encoder-decoder model into a low-latency streaming model that is careless about future context. We present an analysis explaining why it is not straightforward to convert an encoder-decoder transformer to a low-latency streaming model. Our proposed method modifies the existing (non-causal) encoder to a causal encoder by fine-tuning both the encoder and decoder using Low-Rank Adaptation (LoRA) and a weakly aligned dataset. We then propose an updated inference mechanism that utilizes the fine-tune causal encoder and decoder to yield greedy and beam-search decoding, and is shown to be locally optimal. Experiments on low-latency chunk sizes (less than 300 msec) show that our fine-tuned model outperforms existing non-fine-tuned streaming approaches in most cases, while using a lower complexity. Additionally, we observe that our training process yields better alignment, enabling a simple method for extracting word-level timestamps. We release our training and inference code, along with the fine-tuned models, to support further research and development in streaming ASR.
【7】HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
标题:HuBERT-VIC:基于方差-不变-协方差正则化的语音基础模型抗噪自动语音识别
链接:https://arxiv.org/abs/2508.12292
摘要:语音基础模型(SFM)中的噪声鲁棒性一直是一个关键的挑战,因为大多数模型主要是在干净的数据上训练的,当模型暴露于嘈杂的语音时,性能会下降。为了解决这个问题,我们提出了HuBERT-VIC,一个具有方差,不变性和协方差正则化(VICReg)目标的噪声鲁棒SFM。这些目标调整噪声语音表示的统计,使模型能够捕获不同的声学特征,并提高跨不同类型噪声的泛化能力。当应用于HuBERT时,我们的模型在LibriSpeech test-clean上显示出23.3%的相对性能改进,在test-other上显示出13.2%的相对性能改进,与在嘈杂语音上预训练的基线模型相比。
摘要:Noise robustness in speech foundation models (SFMs) has been a critical challenge, as most models are primarily trained on clean data and experience performance degradation when the models are exposed to noisy speech. To address this issue, we propose HuBERT-VIC, a noise-robust SFM with variance, in-variance, and covariance regularization (VICReg) objectives. These objectives adjust the statistics of noisy speech representations, enabling the model to capture diverse acoustic characteristics and improving the generalization ability across different types of noise. When applied to HuBERT, our model shows relative performance improvements of 23.3% on LibriSpeech test-clean and 13.2% on test-other, compared to the baseline model pre-trained on noisy speech.
【8】What do Speech Foundation Models Learn? Analysis and Applications
标题:言语基金会模型学到什么?分析和应用
链接:https://arxiv.org/abs/2508.12255
摘要:语音基础模型(SFM)被设计成用作各种语音处理任务的通用表示。在过去的五年里,越来越多的成功的自我监督和监督预训练模型涌入,在各种下游任务上表现出色。 尽管可持续森林管理动物园不断发展,但我们对它们所获得的知识的理解却落后了。本文提出了一个轻量级的分析框架,使用统计工具和训练免费的任务,调查声学和语言知识编码在SFM层。我们在多个SFM和统计工具中进行了比较研究。我们的研究还表明,分析的见解有具体的影响下游任务的性能。 SFM的有效性最终取决于其在语音应用上的性能。然而,目前尚不清楚这些好处是否延伸到口语理解(SLU)任务,这些任务需要比广泛研究的任务更深入的理解,例如语音识别。对SLU的有限探索主要是由于缺乏相关数据集。为了缓解这一问题,本文贡献的任务,特别是口语命名实体识别(NER)和命名实体本地化(NEL),口语理解评估基准。我们开发了基于SFM的NER和NEL方法,并发现利用SFM的端到端(E2E)模型可以超越传统的级联(语音识别,然后是文本模型)方法。此外,我们评估E2E SLU模型跨SFM和适应策略,以评估对任务性能的影响。 总的来说,这篇论文解决了以前没有回答的问题,可持续森林管理系统,提供工具和数据集,以进一步我们的理解,并使社区作出明智的设计选择,为未来的模型开发和采用。
摘要:Speech foundation models (SFMs) are designed to serve as general-purpose representations for a wide range of speech-processing tasks. The last five years have seen an influx of increasingly successful self-supervised and supervised pre-trained models with impressive performance on various downstream tasks. Although the zoo of SFMs continues to grow, our understanding of the knowledge they acquire lags behind. This thesis presents a lightweight analysis framework using statistical tools and training-free tasks to investigate the acoustic and linguistic knowledge encoded in SFM layers. We conduct a comparative study across multiple SFMs and statistical tools. Our study also shows that the analytical insights have concrete implications for downstream task performance. The effectiveness of an SFM is ultimately determined by its performance on speech applications. Yet it remains unclear whether the benefits extend to spoken language understanding (SLU) tasks that require a deeper understanding than widely studied ones, such as speech recognition. The limited exploration of SLU is primarily due to a lack of relevant datasets. To alleviate that, this thesis contributes tasks, specifically spoken named entity recognition (NER) and named entity localization (NEL), to the Spoken Language Understanding Evaluation benchmark. We develop SFM-based approaches for NER and NEL, and find that end-to-end (E2E) models leveraging SFMs can surpass traditional cascaded (speech recognition followed by a text model) approaches. Further, we evaluate E2E SLU models across SFMs and adaptation strategies to assess the impact on task performance. Collectively, this thesis tackles previously unanswered questions about SFMs, providing tools and datasets to further our understanding and to enable the community to make informed design choices for future model development and adoption.
【9】Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection
标题:探索用于广义异常声音检测的自我监督音频模型
链接:https://arxiv.org/abs/2508.12230
摘要:机器异常声音检测(ASD)是跨各种应用的有价值的技术。然而,由于数据采集的挑战和声学环境的复杂性,其泛化性能往往受到限制。受众多领域中大型预训练模型的成功启发,本文介绍了一种强大的ASD模型,该模型利用在大规模语音和音频数据集上训练的自监督预训练模型。尽管预训练数据集和ASD任务之间存在不一致,但我们的研究结果表明,预训练仍然为ASD提供了实质性的好处。为了在有限数据的微调时减轻过拟合并保留学到的知识,我们探索了全连接低秩自适应(LoRA)作为完全微调的替代方案。此外,我们提出了一个机器感知的组适配器模块,它使模型能够在一个统一的框架内捕捉各种机器之间的差异,从而提高ASD系统的泛化性能。为了解决缺失属性标签的挑战,我们设计了一种新的目标函数,该目标函数使用矢量量化动态聚类未属性数据,并通过双层对比学习损失进行优化。在所有基准数据集上对所提出的方法进行了评估,包括DCASE 2020-2024五个ASD挑战,实验结果表明我们的新方法有了显着的改进,并证明了我们提出的策略的有效性。
摘要:Machine anomalous sound detection (ASD) is a valuable technique across various applications. However, its generalization performance is often limited due to challenges in data collection and the complexity of acoustic environments. Inspired by the success of large pre-trained models in numerous fields, this paper introduces a robust ASD model that leverages self-supervised pre-trained models trained on large-scale speech and audio datasets. Although there are inconsistencies between the pre-training datasets and the ASD task, our findings indicate that pre-training still provides substantial benefits for ASD. To mitigate overfitting and retain learned knowledge when fine-tuning with limited data, we explore Fully-Connected Low-Rank Adaptation (LoRA) as an alternative to full fine-tuning. Additionally, we propose a Machine-aware Group Adapter module, which enables the model to capture differences between various machines within a unified framework, thereby enhancing the generalization performance of ASD systems. To address the challenge of missing attribute labels, we design a novel objective function that dynamically clusters unattributed data using vector quantization and optimizes through a dual-level contrastive learning loss. The proposed methods are evaluated on all benchmark datasets, including the DCASE 2020-2024 five ASD challenges, and the experimental results show significant improvements of our new approach and demonstrate the effectiveness of our proposed strategies.
【10】Music and Artificial Intelligence: Artistic Trends
标题:音乐与人工智能:艺术趋势
链接:https://arxiv.org/abs/2508.11694
摘要:We study how musicians use artificial intelligence (AI) across formats like singles, albums, performances, installations, voices, ballets, operas, or soundtracks. We collect 337 music artworks and categorize them based on AI usage: AI composition, co-composition, sound design, lyrics generation, and translation. We find that AI is employed as a co-creative tool, as an artistic medium, and in live performances and installations. Innovative uses of AI include exploring uncanny aesthetics, multilingual and multigenre song releases, and new formats such as online installations. This research provides a comprehensive overview of current AI music practices, offering insights into emerging artistic trends and the challenges faced by AI musicians.
【11】Prediction of Spotify Chart Success Using Audio and Streaming Features
标题:使用音频和流媒体功能预测Spotify排行榜的成功
链接:https://arxiv.org/abs/2508.11632
摘要:Spotify's streaming charts offer a real-time lens into music popularity, driving discovery, playlists, and even revenue potential. Understanding what influences a song's rise in ranks on these charts-especially early on-can guide marketing efforts, investment decisions, and even artistic direction. In this project, we developed a classification pipeline to predict a song's chart success based on its musical characteristics and early engagement data. Using all 2024 U.S. Top 200 Spotify Daily Charts and the Spotify Web API, we built a dataset containing both metadata and audio features for 14,639 unique songs. The project was structured in two phases. First, we benchmarked four models: Logistic Regression, K Nearest Neighbors, Random Forest, and XGBoost-using a standard train-test split. In the second phase, we incorporated cross-validation, hyperparameter tuning, and detailed class-level evaluation to ensure robustness. Tree-based models consistently outperformed the rest, with Random Forest and XGBoost achieving macro F1-scores near 0.95 and accuracy around 97%. Even when stream count and rank history were excluded, models trained solely on audio attributes retained predictive power. These findings validate the potential of audio-based modeling in A&R scouting, playlist optimization, and hit forecasting-long before a track reaches critical mass.
机器翻译由腾讯交互翻译提供,仅供参考
